codex-orchestrator 2.0.11 → 2.0.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (262) hide show
  1. package/CHANGELOG.md +15 -0
  2. package/README.md +25 -51
  3. package/dist/src/index.d.ts +2 -8
  4. package/dist/src/index.d.ts.map +1 -1
  5. package/dist/src/index.js +1 -4
  6. package/dist/src/index.js.map +1 -1
  7. package/dist/src/v2/acceptance-proof.d.ts +46 -31
  8. package/dist/src/v2/acceptance-proof.d.ts.map +1 -1
  9. package/dist/src/v2/acceptance-proof.js +157 -195
  10. package/dist/src/v2/acceptance-proof.js.map +1 -1
  11. package/dist/src/v2/active-attempt.d.ts +94 -0
  12. package/dist/src/v2/active-attempt.d.ts.map +1 -0
  13. package/dist/src/v2/active-attempt.js +200 -0
  14. package/dist/src/v2/active-attempt.js.map +1 -0
  15. package/dist/src/v2/adapters/command.d.ts +6 -0
  16. package/dist/src/v2/adapters/command.d.ts.map +1 -1
  17. package/dist/src/v2/adapters/command.js +43 -2
  18. package/dist/src/v2/adapters/command.js.map +1 -1
  19. package/dist/src/v2/candidate.d.ts +15 -31
  20. package/dist/src/v2/candidate.d.ts.map +1 -1
  21. package/dist/src/v2/candidate.js +7 -29
  22. package/dist/src/v2/candidate.js.map +1 -1
  23. package/dist/src/v2/checked-change.d.ts +3 -2
  24. package/dist/src/v2/checked-change.d.ts.map +1 -1
  25. package/dist/src/v2/checked-change.js +4 -3
  26. package/dist/src/v2/checked-change.js.map +1 -1
  27. package/dist/src/v2/cli-contract.d.ts +1 -1
  28. package/dist/src/v2/cli-contract.d.ts.map +1 -1
  29. package/dist/src/v2/cli-contract.js +4 -6
  30. package/dist/src/v2/cli-contract.js.map +1 -1
  31. package/dist/src/v2/cli.d.ts +8 -0
  32. package/dist/src/v2/cli.d.ts.map +1 -1
  33. package/dist/src/v2/cli.js +13 -0
  34. package/dist/src/v2/cli.js.map +1 -1
  35. package/dist/src/v2/code-review-report.d.ts +10 -18
  36. package/dist/src/v2/code-review-report.d.ts.map +1 -1
  37. package/dist/src/v2/code-review-report.js +63 -60
  38. package/dist/src/v2/code-review-report.js.map +1 -1
  39. package/dist/src/v2/codex-process.d.ts +6 -2
  40. package/dist/src/v2/codex-process.d.ts.map +1 -1
  41. package/dist/src/v2/codex-process.js +25 -9
  42. package/dist/src/v2/codex-process.js.map +1 -1
  43. package/dist/src/v2/config.d.ts +0 -2
  44. package/dist/src/v2/config.d.ts.map +1 -1
  45. package/dist/src/v2/config.js +3 -6
  46. package/dist/src/v2/config.js.map +1 -1
  47. package/dist/src/v2/contained-report-operation.d.ts +41 -196
  48. package/dist/src/v2/contained-report-operation.d.ts.map +1 -1
  49. package/dist/src/v2/contained-report-operation.js +139 -466
  50. package/dist/src/v2/contained-report-operation.js.map +1 -1
  51. package/dist/src/v2/containment.d.ts +1 -0
  52. package/dist/src/v2/containment.d.ts.map +1 -1
  53. package/dist/src/v2/containment.js +12 -2
  54. package/dist/src/v2/containment.js.map +1 -1
  55. package/dist/src/v2/delivery-authority.d.ts +26 -0
  56. package/dist/src/v2/delivery-authority.d.ts.map +1 -0
  57. package/dist/src/v2/delivery-authority.js +44 -0
  58. package/dist/src/v2/delivery-authority.js.map +1 -0
  59. package/dist/src/v2/direct-delivery.d.ts +16 -36
  60. package/dist/src/v2/direct-delivery.d.ts.map +1 -1
  61. package/dist/src/v2/direct-delivery.js +135 -122
  62. package/dist/src/v2/direct-delivery.js.map +1 -1
  63. package/dist/src/v2/immutable-workflow-publisher.d.ts.map +1 -1
  64. package/dist/src/v2/immutable-workflow-publisher.js +3 -1
  65. package/dist/src/v2/immutable-workflow-publisher.js.map +1 -1
  66. package/dist/src/v2/implementation-report.d.ts +3 -1
  67. package/dist/src/v2/implementation-report.d.ts.map +1 -1
  68. package/dist/src/v2/implementation-report.js +17 -4
  69. package/dist/src/v2/implementation-report.js.map +1 -1
  70. package/dist/src/v2/implementation-reviewer.d.ts +41 -12
  71. package/dist/src/v2/implementation-reviewer.d.ts.map +1 -1
  72. package/dist/src/v2/implementation-reviewer.js +114 -42
  73. package/dist/src/v2/implementation-reviewer.js.map +1 -1
  74. package/dist/src/v2/pending-effect-settlement.d.ts +44 -0
  75. package/dist/src/v2/pending-effect-settlement.d.ts.map +1 -0
  76. package/dist/src/v2/pending-effect-settlement.js +69 -0
  77. package/dist/src/v2/pending-effect-settlement.js.map +1 -0
  78. package/dist/src/v2/process-identity.d.ts +45 -0
  79. package/dist/src/v2/process-identity.d.ts.map +1 -0
  80. package/dist/src/v2/process-identity.js +118 -0
  81. package/dist/src/v2/process-identity.js.map +1 -0
  82. package/dist/src/v2/proof-report.d.ts +2 -1
  83. package/dist/src/v2/proof-report.d.ts.map +1 -1
  84. package/dist/src/v2/proof-report.js +10 -4
  85. package/dist/src/v2/proof-report.js.map +1 -1
  86. package/dist/src/v2/review-feedback-coordinator.d.ts +1 -1
  87. package/dist/src/v2/review-feedback-coordinator.d.ts.map +1 -1
  88. package/dist/src/v2/review-feedback-coordinator.js +1 -1
  89. package/dist/src/v2/review-feedback-coordinator.js.map +1 -1
  90. package/dist/src/v2/review-feedback.d.ts +14 -20
  91. package/dist/src/v2/review-feedback.d.ts.map +1 -1
  92. package/dist/src/v2/review-feedback.js +45 -87
  93. package/dist/src/v2/review-feedback.js.map +1 -1
  94. package/dist/src/v2/run-issue.d.ts +129 -88
  95. package/dist/src/v2/run-issue.d.ts.map +1 -1
  96. package/dist/src/v2/run-issue.js +1965 -2381
  97. package/dist/src/v2/run-issue.js.map +1 -1
  98. package/dist/src/v2/run-state-projections.d.ts +84 -0
  99. package/dist/src/v2/run-state-projections.d.ts.map +1 -0
  100. package/dist/src/v2/run-state-projections.js +142 -0
  101. package/dist/src/v2/run-state-projections.js.map +1 -0
  102. package/dist/src/v2/run-store.d.ts +99 -81
  103. package/dist/src/v2/run-store.d.ts.map +1 -1
  104. package/dist/src/v2/run-store.js +245 -542
  105. package/dist/src/v2/run-store.js.map +1 -1
  106. package/dist/src/v2/runtime-assets.d.ts +3 -0
  107. package/dist/src/v2/runtime-assets.d.ts.map +1 -1
  108. package/dist/src/v2/runtime-assets.js +104 -0
  109. package/dist/src/v2/runtime-assets.js.map +1 -1
  110. package/dist/src/v2/runtime.d.ts +56 -44
  111. package/dist/src/v2/runtime.d.ts.map +1 -1
  112. package/dist/src/v2/runtime.js +383 -503
  113. package/dist/src/v2/runtime.js.map +1 -1
  114. package/dist/src/v2/setup.js +0 -2
  115. package/dist/src/v2/setup.js.map +1 -1
  116. package/dist/src/v2/validation-progression.d.ts +70 -0
  117. package/dist/src/v2/validation-progression.d.ts.map +1 -0
  118. package/dist/src/v2/validation-progression.js +247 -0
  119. package/dist/src/v2/validation-progression.js.map +1 -0
  120. package/dist/src/v2/workflow-assets.d.ts +9 -3
  121. package/dist/src/v2/workflow-assets.d.ts.map +1 -1
  122. package/dist/src/v2/workflow-assets.js +256 -43
  123. package/dist/src/v2/workflow-assets.js.map +1 -1
  124. package/internal-workflow/docs/agents/bug-workflow-routing.md +9 -7
  125. package/internal-workflow/docs/agents/coding-skill-routing.md +170 -120
  126. package/internal-workflow/docs/agents/tool-usage.md +23 -12
  127. package/internal-workflow/manifest.json +1 -1
  128. package/internal-workflow/operations/code-review/SKILL.md +34 -15
  129. package/internal-workflow/operations/implementation/SKILL.md +21 -16
  130. package/internal-workflow/profiles/implementer.toml +9 -0
  131. package/internal-workflow/profiles/review_coordinator.toml +9 -0
  132. package/internal-workflow/profiles/spec_reviewer.toml +9 -0
  133. package/internal-workflow/profiles/standards_reviewer.toml +9 -0
  134. package/internal-workflow/schemas/code-review-v1.json +1 -1
  135. package/internal-workflow/schemas/implementation-report-v1.json +1 -1
  136. package/internal-workflow/schemas/proof-report-v1.json +1 -1
  137. package/internal-workflow/skills/bug-root-cause-explainer/SKILL.md +114 -0
  138. package/internal-workflow/skills/bug-root-cause-explainer/agents/openai.yaml +7 -0
  139. package/internal-workflow/skills/bug-root-cause-explainer/evals/evals.json +18 -0
  140. package/internal-workflow/skills/code-review/SKILL.md +84 -306
  141. package/internal-workflow/skills/code-review/agents/openai.yaml +5 -3
  142. package/internal-workflow/skills/code-review/evals/evals.json +83 -0
  143. package/internal-workflow/skills/code-review/references/standards-smells.md +41 -0
  144. package/internal-workflow/skills/diagnosing-bugs/SKILL.md +69 -32
  145. package/internal-workflow/skills/diagnosing-bugs/agents/openai.yaml +2 -2
  146. package/internal-workflow/skills/diagnosing-bugs/evals/evals.json +63 -0
  147. package/internal-workflow/skills/grilling/SKILL.md +51 -0
  148. package/internal-workflow/skills/grilling/agents/openai.yaml +6 -0
  149. package/internal-workflow/skills/grilling/evals/evals.json +47 -0
  150. package/internal-workflow/skills/implement/SKILL.md +135 -0
  151. package/internal-workflow/skills/implement/agents/openai.yaml +6 -0
  152. package/internal-workflow/skills/implement/evals/evals.json +150 -0
  153. package/internal-workflow/skills/plan/SKILL.md +59 -0
  154. package/internal-workflow/skills/plan/agents/openai.yaml +6 -0
  155. package/internal-workflow/skills/plan/evals/evals.json +36 -0
  156. package/internal-workflow/skills/prototype/LOGIC.md +130 -0
  157. package/internal-workflow/skills/prototype/SKILL.md +69 -0
  158. package/internal-workflow/skills/prototype/UI.md +157 -0
  159. package/internal-workflow/skills/prototype/agents/openai.yaml +6 -0
  160. package/internal-workflow/skills/prototype/evals/evals.json +67 -0
  161. package/internal-workflow/skills/research/SKILL.md +110 -0
  162. package/internal-workflow/skills/research/agents/openai.yaml +6 -0
  163. package/internal-workflow/skills/research/evals/evals.json +49 -0
  164. package/internal-workflow/skills/tdd/SKILL.md +72 -67
  165. package/internal-workflow/skills/tdd/agents/openai.yaml +2 -2
  166. package/internal-workflow/skills/tdd/evals/evals.json +12 -0
  167. package/internal-workflow/skills/tdd/mocking.md +48 -1
  168. package/internal-workflow/skills/tdd/refactoring.md +3 -3
  169. package/internal-workflow/skills/tickets-orchestrator/SKILL.md +199 -0
  170. package/internal-workflow/skills/tickets-orchestrator/agents/openai.yaml +6 -0
  171. package/internal-workflow/skills/tickets-orchestrator/evals/evals.json +126 -0
  172. package/internal-workflow/skills/tickets-orchestrator/references/delegate-integrate.md +83 -0
  173. package/internal-workflow/skills/tickets-orchestrator/references/finish-delivery.md +69 -0
  174. package/internal-workflow/skills/tickets-orchestrator/references/stop-completion.md +63 -0
  175. package/internal-workflow/skills/to-spec/SKILL.md +133 -0
  176. package/internal-workflow/skills/to-spec/agents/openai.yaml +6 -0
  177. package/internal-workflow/skills/to-spec/evals/evals.json +24 -0
  178. package/internal-workflow/skills/to-tickets/SKILL.md +189 -0
  179. package/internal-workflow/skills/to-tickets/agents/openai.yaml +6 -0
  180. package/internal-workflow/skills/to-tickets/evals/evals.json +79 -0
  181. package/internal-workflow/skills/to-tickets/references/publishing-details.md +117 -0
  182. package/package.json +1 -1
  183. package/dist/src/v2/proof-store.d.ts +0 -54
  184. package/dist/src/v2/proof-store.d.ts.map +0 -1
  185. package/dist/src/v2/proof-store.js +0 -301
  186. package/dist/src/v2/proof-store.js.map +0 -1
  187. package/dist/src/v2/route-continuations.d.ts +0 -32
  188. package/dist/src/v2/route-continuations.d.ts.map +0 -1
  189. package/dist/src/v2/route-continuations.js +0 -2
  190. package/dist/src/v2/route-continuations.js.map +0 -1
  191. package/dist/src/v2/route-coordinator.d.ts +0 -72
  192. package/dist/src/v2/route-coordinator.d.ts.map +0 -1
  193. package/dist/src/v2/route-coordinator.js +0 -275
  194. package/dist/src/v2/route-coordinator.js.map +0 -1
  195. package/dist/src/v2/route-decision.d.ts +0 -120
  196. package/dist/src/v2/route-decision.d.ts.map +0 -1
  197. package/dist/src/v2/route-decision.js +0 -380
  198. package/dist/src/v2/route-decision.js.map +0 -1
  199. package/dist/src/v2/spec-coordinator.d.ts +0 -73
  200. package/dist/src/v2/spec-coordinator.d.ts.map +0 -1
  201. package/dist/src/v2/spec-coordinator.js +0 -126
  202. package/dist/src/v2/spec-coordinator.js.map +0 -1
  203. package/dist/src/v2/spec-delivery.d.ts +0 -112
  204. package/dist/src/v2/spec-delivery.d.ts.map +0 -1
  205. package/dist/src/v2/spec-delivery.js +0 -336
  206. package/dist/src/v2/spec-delivery.js.map +0 -1
  207. package/dist/src/v2/triage-route.d.ts +0 -68
  208. package/dist/src/v2/triage-route.d.ts.map +0 -1
  209. package/dist/src/v2/triage-route.js +0 -223
  210. package/dist/src/v2/triage-route.js.map +0 -1
  211. package/dist/src/v2/waiting-human-coordinator.d.ts +0 -49
  212. package/dist/src/v2/waiting-human-coordinator.d.ts.map +0 -1
  213. package/dist/src/v2/waiting-human-coordinator.js +0 -509
  214. package/dist/src/v2/waiting-human-coordinator.js.map +0 -1
  215. package/dist/src/v2/waiting-human.d.ts +0 -143
  216. package/dist/src/v2/waiting-human.d.ts.map +0 -1
  217. package/dist/src/v2/waiting-human.js +0 -408
  218. package/dist/src/v2/waiting-human.js.map +0 -1
  219. package/internal-workflow/docs/agents/contract-test-ledger.md +0 -71
  220. package/internal-workflow/docs/agents/review-gates.md +0 -42
  221. package/internal-workflow/docs/agents/review-protocol.md +0 -98
  222. package/internal-workflow/evals/coding-skill-evals.json +0 -373
  223. package/internal-workflow/operations/ambiguity-review/SKILL.md +0 -5
  224. package/internal-workflow/operations/qualification-repair/SKILL.md +0 -17
  225. package/internal-workflow/operations/spec-author/SKILL.md +0 -12
  226. package/internal-workflow/operations/spec-review/SKILL.md +0 -12
  227. package/internal-workflow/operations/triage/SKILL.md +0 -12
  228. package/internal-workflow/profiles/analyst_deep.toml +0 -9
  229. package/internal-workflow/profiles/implementer_standard.toml +0 -9
  230. package/internal-workflow/profiles/proof_agent.toml +0 -8
  231. package/internal-workflow/profiles/reviewer_deep.toml +0 -9
  232. package/internal-workflow/profiles/reviewer_standard.toml +0 -9
  233. package/internal-workflow/schemas/ambiguity-review-v1.json +0 -1
  234. package/internal-workflow/schemas/spec-author-v1.json +0 -1
  235. package/internal-workflow/schemas/spec-review-v1.json +0 -30
  236. package/internal-workflow/schemas/triage-route-v1.json +0 -1
  237. package/internal-workflow/skills/agent-auto/SKILL.md +0 -19
  238. package/internal-workflow/skills/agent-auto/agents/openai.yaml +0 -6
  239. package/internal-workflow/skills/code-debugger/SKILL.md +0 -122
  240. package/internal-workflow/skills/code-debugger/agents/openai.yaml +0 -7
  241. package/internal-workflow/skills/code-review/references/bug-classes.md +0 -56
  242. package/internal-workflow/skills/code-review/references/cleanup-lens.md +0 -52
  243. package/internal-workflow/skills/code-review/references/framework-lenses.md +0 -34
  244. package/internal-workflow/skills/code-review/references/targeted-recipes.md +0 -49
  245. package/internal-workflow/skills/implementation-spec-maker/SKILL.md +0 -107
  246. package/internal-workflow/skills/implementation-spec-maker/agents/openai.yaml +0 -6
  247. package/internal-workflow/skills/implementation-spec-maker/references/source-modes.md +0 -32
  248. package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +0 -146
  249. package/internal-workflow/skills/implementation-spec-review/SKILL.md +0 -131
  250. package/internal-workflow/skills/implementation-spec-review/agents/openai.yaml +0 -6
  251. package/internal-workflow/skills/implementation-spec-review/evals/evals.json +0 -78
  252. package/internal-workflow/skills/implementation-spec-review/references/review-loop.md +0 -121
  253. package/internal-workflow/skills/small-task-implementer/SKILL.md +0 -112
  254. package/internal-workflow/skills/small-task-implementer/agents/openai.yaml +0 -6
  255. package/internal-workflow/skills/spec-implementer/SKILL.md +0 -133
  256. package/internal-workflow/skills/spec-implementer/agents/openai.yaml +0 -6
  257. package/internal-workflow/skills/spec-implementer/evals/evals.json +0 -30
  258. package/internal-workflow/skills/spec-implementer/references/review-loop.md +0 -100
  259. package/internal-workflow/skills/triage/AGENT-BRIEF.md +0 -192
  260. package/internal-workflow/skills/triage/OUT-OF-SCOPE.md +0 -101
  261. package/internal-workflow/skills/triage/SKILL.md +0 -134
  262. package/internal-workflow/skills/triage/agents/openai.yaml +0 -6
@@ -1,133 +0,0 @@
1
- ---
2
- name: "spec-implementer"
3
- description: "Executes approved specs continuously with honest checklist updates, proportional validation, opt-in Git checkpoints, and required review/signoff."
4
- ---
5
-
6
- # Spec Implementer
7
-
8
- Execute one approved implementation spec continuously. Keep its checklist
9
- truthful and stop only at a real authority, evidence, safety, or explicit pause
10
- boundary. Do not redesign approved scope.
11
-
12
- Read before editing:
13
-
14
- - the complete approved spec and applicable repository instructions;
15
- - `references/review-loop.md` for approved-spec review ownership;
16
- - `../../docs/agents/contract-test-ledger.md` only when the spec contains
17
- material contract invariants.
18
-
19
- ## Modes
20
-
21
- The primary route remains `spec-implementer` whenever an approved spec is the
22
- execution authority. Do not report `implementation_size`, TDD, a Contract Test
23
- Ledger, or review as a replacement route; those are execution details inside
24
- this route.
25
-
26
- Compact and full specs use the same direct phase flow. `compact` describes
27
- document density, not implementation size or risk. Full mode adds only the
28
- concrete `Risk Controls`, stop conditions, validation, or coordination contract
29
- already present in the approved spec; it does not add review or reporting
30
- ceremony by itself.
31
-
32
- Use multiple agents only when the spec defines perfectly disjoint write scopes
33
- and one integrator. Otherwise execute single-agent.
34
-
35
- ## Preflight
36
-
37
- 1. Confirm `status`, `spec_mode`, `implementation_size`, `review_profile`,
38
- repository count, authority, scope, and exclusions.
39
- 2. Confirm first-phase targets, required services/env/data/fixtures, commands,
40
- observable proof, protected paths, and rejected approaches.
41
- 3. Stop if execution needs invented paths, symbols, contracts, commands, or
42
- product decisions; if reality differs only in a bounded technical detail,
43
- resolve it from repository evidence and record the adjustment.
44
- 4. Keep an intermediate review checkpoint only when the spec explicitly names
45
- a stable high-risk slice that later work will not invalidate. Move every
46
- ordinary or unstable checkpoint to final review.
47
- 5. Do not create `## Implementation Review State` yet.
48
-
49
- ## Implementation
50
-
51
- For each phase:
52
-
53
- 1. Re-read its scope and preconditions.
54
- 2. For an applicable behavior slice, activate and read `$tdd` before the first
55
- RED, then implement through that skill. A spec-provided failing command or
56
- RED recorder does not replace loading the TDD contract.
57
- 3. Update reached checklist and Contract Test Ledger items at natural
58
- checkpoints; never save all updates for the end.
59
- 4. Run the phase's targeted exit proof.
60
- 5. Record a short `Blocked:` note for any item that cannot complete.
61
- 6. Continue immediately when the exit proof passes and no stop condition
62
- applies.
63
-
64
- Do not add cleanup, comments, helpers, abstractions, retries, flags, fallbacks,
65
- or compatibility paths outside the spec. A small implementation adjustment is
66
- allowed only when it preserves approved behavior and is supported by current
67
- repository evidence.
68
-
69
- ## Git Checkpoints
70
-
71
- Default to no commits. `$spec-implementer` does not authorize Git writes.
72
-
73
- Use per-slice commits only when explicitly authorized and they materially help
74
- a disjoint multi-agent handoff, planned cross-session pause, or approved
75
- rollback boundary. Require a passed slice proof, stage only owned paths, follow
76
- `$commit`, and never push without separate authority. File/slice count and risk
77
- profile alone do not justify checkpoints.
78
-
79
- ## Review And Validation
80
-
81
- Run targeted behavior tests and the smallest affected integration checks first.
82
- Use a full repository suite only when the spec or repository policy requires
83
- it, broad contract fan-out cannot be isolated, or the task is genuinely high.
84
-
85
- Run final `$code-review` only when the spec, repository policy, or
86
- `../../docs/agents/review-gates.md` applies. Ordinary medium work gets one
87
- `reviewer_standard` Full review on the settled diff. High gets two disjoint
88
- `reviewer_deep` lenses. Cleanup stays inside spec/standards review.
89
-
90
- Immediately before the first reviewer launch, create only the minimal
91
- `## Implementation Review State` required by `references/review-loop.md`.
92
- Reconcile a recorded live session before replacement after interruption.
93
-
94
- Repair compatible findings once and rerun affected validation. Coordinator
95
- verification closes ordinary medium/low behavior-preserving repairs. Use
96
- Closure only for critical/high, protected trust/data/concurrency/shared-contract
97
- impact, or invalidated mandatory coverage. Do not restart broad Full review
98
- unless the repair actually invalidated its coverage.
99
-
100
- ## Stop Conditions
101
-
102
- Stop and report the exact blocker when:
103
-
104
- - a precondition, source contract, required proof, or protected path differs
105
- materially from the approved spec;
106
- - exact execution requires a new product, scope, ownership, or risky trade-off
107
- decision;
108
- - required validation or reviewer is unavailable and no approved substitute
109
- exists;
110
- - multi-agent scopes overlap or integration ownership is missing;
111
- - an explicit user pause or halt condition is reached.
112
-
113
- Do not stop merely because the work is broad, review took time, or a medium/low
114
- finding required one repair.
115
-
116
- ## Completion
117
-
118
- Complete only when reached checklist items are reconciled, affected proof and
119
- required review pass, protected paths and rejected approaches remain intact,
120
- and every unfinished item has a concrete status.
121
-
122
- For ordinary medium work report only:
123
-
124
- - behavior/contract implemented;
125
- - review result and repaired/open findings;
126
- - affected validation;
127
- - skipped checks and residual risk;
128
- - changed files and any authorized commits.
129
-
130
- For high, actual Closure, accepted risk, interrupted recovery, or multi-agent
131
- delivery, add the relevant invariants, reviewer coverage, defect IDs, session
132
- recovery, and handoff ownership. Do not create a separate report file unless
133
- the spec or the complexity of that exceptional handoff requires it.
@@ -1,6 +0,0 @@
1
- interface:
2
- display_name: "Spec Implementer"
3
- short_description: "Execute approved specs with lean delivery"
4
- default_prompt: "Use $spec-implementer to execute the approved spec continuously, default to no Git checkpoints, validate proportionately, and launch only stable required review gates."
5
- policy:
6
- allow_implicit_invocation: true
@@ -1,30 +0,0 @@
1
- {
2
- "schema_version": 1,
3
- "skill": "spec-implementer",
4
- "cases": [
5
- {
6
- "id": "state-created-at-launch",
7
- "prompt": "Execute an approved medium spec whose final review has not started yet.",
8
- "expected": ["do not create review state during implementation", "persist minimal state immediately before reviewer launch"],
9
- "forbidden": ["create per-slice review bookkeeping"]
10
- },
11
- {
12
- "id": "medium-one-final-review",
13
- "prompt": "Execute a normal medium spec with several vertical slices and no stable high-risk checkpoint.",
14
- "expected": ["implement continuously", "run one final reviewer_standard on the settled diff"],
15
- "forbidden": ["review after every slice", "run a full repository suite from file count"]
16
- },
17
- {
18
- "id": "closure-stays-affected",
19
- "prompt": "Final review finds one high defect in a shared schema and several ordinary low findings.",
20
- "expected": ["repair once", "Closure verifies only the affected high-risk contract"],
21
- "forbidden": ["restart all Full reviewers"]
22
- },
23
- {
24
- "id": "resume-live-reviewer",
25
- "prompt": "Resume after a reviewer poll timed out while the recorded session is still live.",
26
- "expected": ["reconcile the existing session"],
27
- "forbidden": ["mark it failed from timeout alone", "launch a duplicate reviewer"]
28
- }
29
- ]
30
- }
@@ -1,100 +0,0 @@
1
- # Approved Spec Implementation Review Loop
2
-
3
- This reference owns review orchestration for `$spec-implementer`. Read it only
4
- when executing an approved implementation spec. Shared Full/Closure and defect
5
- mechanics live in `../../../docs/agents/review-protocol.md`.
6
-
7
- Direct work and deterministic issue delivery use normal TDD and review gates;
8
- they must not create Implementation Review State.
9
-
10
- ## Authority
11
-
12
- Only an approved implementation spec may own `## Implementation Review State`.
13
- PRDs, tickets, architecture notes, and chat summaries are not review-state
14
- owners. A substantive spec change returns through artifact review before
15
- implementation continues.
16
-
17
- Use the spec's `review_profile`; if absent, infer it from current evidence:
18
-
19
- - `simple`: narrow change with direct proof;
20
- - `medium`: default for ordinary implementation;
21
- - `high`: material failure consequence (financial side effect,
22
- unauthorized/cross-owner behavior, durable corruption, or materially false
23
- production result) plus an uncertainty amplifier (concurrency or event
24
- ordering, delayed/background callbacks, retry/idempotency/recovery,
25
- ownership transitions, or shared state across consumers).
26
-
27
- Implementation evidence may raise but never lower the approved profile.
28
- Recheck the settled diff immediately before the first reviewer launch and
29
- persist any required raise before launching reviewers.
30
-
31
- ## Default Review Shape
32
-
33
- Implement continuously through vertical slices. Validate each affected behavior
34
- and run one final review on the settled diff when the gate applies.
35
-
36
- - `simple`: one `reviewer_fast` when review is required.
37
- - `medium`: one `reviewer_standard`, one bounded final Full, no intermediate
38
- checkpoint by default.
39
- - `high`: two parallel `reviewer_deep` Full reviews with disjoint correctness
40
- and spec/standards lenses.
41
-
42
- Add an intermediate checkpoint only when the approved spec explicitly names a
43
- stable high-risk slice whose review will remain valid after later work. Do not
44
- review unstable intermediate diffs or create per-slice review cycles.
45
-
46
- Cleanup stays inside the spec/standards lens. A concrete simplification risk may
47
- amplify that lens; size and profile labels alone do not create another gate.
48
-
49
- ## Minimal Durable State
50
-
51
- Do not create review state during preflight or implementation. Immediately
52
- before the first actual reviewer launch, persist:
53
-
54
- - profile, authority path, settled target revision, and assigned lenses;
55
- - launch ID, reviewer/session handle, lineage, and `pending | completed | failed`;
56
- - returned findings, repair revision, affected validation, and Closure need.
57
-
58
- Write `pending` before launch and reconcile that session before replacing it
59
- after interruption or resume. A usable result becomes `completed`; an explicit
60
- failure becomes `failed`. A poll timeout while the session remains live is not
61
- a failure and does not authorize duplicate review.
62
-
63
- Record extended lineage/session history only for `high`, a real intermediate
64
- checkpoint, actual Closure, accepted risk, or interrupted recovery. Normal
65
- medium execution does not keep epochs, pass thresholds, activation counters, or
66
- per-slice handoff bookkeeping.
67
-
68
- ## Findings And Closure
69
-
70
- Root aggregates findings, repairs compatible defects once, and reruns only
71
- affected validation. Coordinator verification closes ordinary medium/low
72
- behavior-preserving findings after confirming the repair matches the failure.
73
-
74
- Use shared-protocol Closure only for critical/high defects, protected
75
- trust/data/concurrency/shared API impact, or invalidated mandatory coverage.
76
- Closure stays with the affected reviewer lineage and repaired targets. Start a
77
- new Full only when the repair invalidated mandatory-lens coverage.
78
-
79
- Do not repeat review without a material change in target, evidence, repair, or
80
- source decision. Stop and surface the actual decision or evidence blocker when
81
- no progress is possible.
82
-
83
- ## Completion
84
-
85
- Run gates in this order:
86
-
87
- 1. affected behavior and integration validation;
88
- 2. applicable final code review;
89
- 3. Closure only when triggered;
90
- 4. repository architecture/build/smoke gates required by policy or the spec;
91
- 5. delivery actions explicitly authorized by the user or workflow.
92
-
93
- Return `Approved` only for the final settled revision when mandatory lenses and
94
- validation are complete and shared protocol state is clear. `Waived` records
95
- skipped coverage but is not approval. `Blocked` requires a concrete authority,
96
- evidence, reviewer, or convergence blocker—not elapsed time or review count.
97
-
98
- For normal medium work report only profile, review result, repaired/open
99
- findings, affected validation, skipped checks, and residual risk. Add extended
100
- session/defect accounting only when the exceptional state above exists.
@@ -1,192 +0,0 @@
1
- # Writing Agent Briefs
2
-
3
- An agent brief is a structured comment posted when a raw incoming issue moves
4
- to `ready-for-agent` through `$triage`. It is the authoritative specification
5
- for that AFK issue; the original body and discussion remain context.
6
-
7
- Do not post this comment for tickets generated by `$to-tickets`. Their approved
8
- issue body is already the authoritative execution brief and they bypass triage.
9
-
10
- ## Principles
11
-
12
- ### Durability over precision
13
-
14
- The issue may sit in `ready-for-agent` for days or weeks. The codebase will change in the meantime. Write the brief so it stays useful even as files are renamed, moved, or refactored.
15
-
16
- - **Do** describe interfaces, types, and behavioral contracts
17
- - **Do** name specific types, function signatures, or config shapes that the agent should look for or modify
18
- - **Don't** reference file paths — they go stale
19
- - **Don't** reference line numbers
20
- - **Don't** assume the current implementation structure will remain the same
21
-
22
- ### Behavioral, not procedural
23
-
24
- Describe **what** the system should do, not **how** to implement it. The agent will explore the codebase fresh and make its own implementation decisions.
25
-
26
- - **Good:** "The `SkillConfig` type should accept an optional `schedule` field of type `CronExpression`"
27
- - **Bad:** "Open src/types/skill.ts and add a schedule field on line 42"
28
- - **Good:** "When a user runs `$triage` with no arguments, they should see a summary of issues needing attention"
29
- - **Bad:** "Add a switch statement in the main handler function"
30
-
31
- ### Complete acceptance criteria
32
-
33
- The agent needs to know when it's done. Every agent brief must have concrete, testable acceptance criteria. Each criterion should be independently verifiable.
34
-
35
- - **Good:** "Running `gh issue list --label needs-triage` returns issues that have been through initial classification"
36
- - **Bad:** "Triage should work correctly"
37
-
38
- ### Explicit scope boundaries
39
-
40
- State what is out of scope. This prevents the agent from gold-plating or making assumptions about adjacent features.
41
-
42
- ### Proportional execution metadata
43
-
44
- Preserve an existing spec gate, verification contract, and material risk notes
45
- without turning the brief into an implementation spec. Omit those conditional
46
- fields when they are not material. High review risk changes proof depth, not
47
- brief length or solution breadth.
48
-
49
- ## Template
50
-
51
- ```markdown
52
- ## Agent Brief
53
-
54
- **Category:** bug / enhancement
55
- **Summary:** one-line description of what needs to happen
56
-
57
- **Current behavior:**
58
- Describe what happens now. For bugs, this is the broken behavior.
59
- For enhancements, this is the status quo the feature builds on.
60
-
61
- **Desired behavior:**
62
- Describe what should happen after the agent's work is complete.
63
- Be specific about edge cases and error conditions.
64
-
65
- **Key interfaces:**
66
- - `TypeName` — what needs to change and why
67
- - `functionName()` return type — what it currently returns vs what it should return
68
- - Config shape — any new configuration options needed
69
-
70
- **Acceptance criteria:**
71
- - [ ] Specific, testable criterion 1
72
- - [ ] Specific, testable criterion 2
73
- - [ ] Specific, testable criterion 3
74
-
75
- **Implementation preparation:**
76
- - direct / compact spec / standard spec
77
- - Reason or accepted implementation spec reference
78
-
79
- **Verification:**
80
- - Behavior proof expected from the implementing agent
81
- - Required automated or live check, or `Not specified`
82
-
83
- **Risk / Review:**
84
- - Omit when no material risk exists
85
- - Otherwise: primary risk, main invariant, review focus, and final handoff expectation
86
-
87
- **Out of scope:**
88
- - Thing that should NOT be changed or addressed in this issue
89
- - Adjacent feature that might seem related but is separate
90
- ```
91
-
92
- ## Examples
93
-
94
- ### Good agent brief (bug)
95
-
96
- ```markdown
97
- ## Agent Brief
98
-
99
- **Category:** bug
100
- **Summary:** Skill description truncation drops mid-word, producing broken output
101
-
102
- **Current behavior:**
103
- When a skill description exceeds 1024 characters, it is truncated at exactly
104
- 1024 characters regardless of word boundaries. This produces descriptions
105
- that end mid-word (e.g. "Use when the user wants to confi").
106
-
107
- **Desired behavior:**
108
- Truncation should break at the last word boundary before 1024 characters
109
- and append "..." to indicate truncation.
110
-
111
- **Key interfaces:**
112
- - The `SkillMetadata` type's `description` field — no type change needed,
113
- but the validation/processing logic that populates it needs to respect
114
- word boundaries
115
- - Any function that reads SKILL.md frontmatter and extracts the description
116
-
117
- **Acceptance criteria:**
118
- - [ ] Descriptions under 1024 chars are unchanged
119
- - [ ] Descriptions over 1024 chars are truncated at the last word boundary
120
- before 1024 chars
121
- - [ ] Truncated descriptions end with "..."
122
- - [ ] The total length including "..." does not exceed 1024 chars
123
-
124
- **Out of scope:**
125
- - Changing the 1024 char limit itself
126
- - Multi-line description support
127
- ```
128
-
129
- ### Good agent brief (enhancement)
130
-
131
- ```markdown
132
- ## Agent Brief
133
-
134
- **Category:** enhancement
135
- **Summary:** Add `.out-of-scope/` directory support for tracking rejected feature requests
136
-
137
- **Current behavior:**
138
- When a feature request is rejected, the issue is closed with a `wontfix` label
139
- and a comment. There is no persistent record of the decision or reasoning.
140
- Future similar requests require the maintainer to recall or search for the
141
- prior discussion.
142
-
143
- **Desired behavior:**
144
- Rejected feature requests should be documented in `.out-of-scope/<concept>.md`
145
- files that capture the decision, reasoning, and links to all issues that
146
- requested the feature. When triaging new issues, these files should be
147
- checked for matches.
148
-
149
- **Key interfaces:**
150
- - Markdown file format in `.out-of-scope/` — each file should have a
151
- `# Concept Name` heading, a `**Decision:**` line, a `**Reason:**` line,
152
- and a `**Prior requests:**` list with issue links
153
- - The triage workflow should read all `.out-of-scope/*.md` files early
154
- and match incoming issues against them by concept similarity
155
-
156
- **Acceptance criteria:**
157
- - [ ] Closing a feature as wontfix creates/updates a file in `.out-of-scope/`
158
- - [ ] The file includes the decision, reasoning, and link to the closed issue
159
- - [ ] If a matching `.out-of-scope/` file already exists, the new issue is
160
- appended to its "Prior requests" list rather than creating a duplicate
161
- - [ ] During triage, existing `.out-of-scope/` files are checked and surfaced
162
- when a new issue matches a prior rejection
163
-
164
- **Out of scope:**
165
- - Automated matching (human confirms the match)
166
- - Reopening previously rejected features
167
- - Bug reports (only enhancement rejections go to `.out-of-scope/`)
168
- ```
169
-
170
- ### Bad agent brief
171
-
172
- ```markdown
173
- ## Agent Brief
174
-
175
- **Summary:** Fix the triage bug
176
-
177
- **What to do:**
178
- The triage thing is broken. Look at the main file and fix it.
179
- The function around line 150 has the issue.
180
-
181
- **Files to change:**
182
- - src/triage/handler.ts (line 150)
183
- - src/types.ts (line 42)
184
- ```
185
-
186
- This is bad because:
187
- - No category
188
- - Vague description ("the triage thing is broken")
189
- - References file paths and line numbers that will go stale
190
- - No acceptance criteria
191
- - No scope boundaries
192
- - No description of current vs desired behavior
@@ -1,101 +0,0 @@
1
- # Out-of-Scope Knowledge Base
2
-
3
- The `.out-of-scope/` directory in a repo stores persistent records of rejected feature requests. It serves two purposes:
4
-
5
- 1. **Institutional memory** — why a feature was rejected, so the reasoning isn't lost when the issue is closed
6
- 2. **Deduplication** — when a new issue comes in that matches a prior rejection, the skill can surface the previous decision instead of re-litigating it
7
-
8
- ## Directory structure
9
-
10
- ```
11
- .out-of-scope/
12
- ├── dark-mode.md
13
- ├── plugin-system.md
14
- └── graphql-api.md
15
- ```
16
-
17
- One file per **concept**, not per issue. Multiple issues requesting the same thing are grouped under one file.
18
-
19
- ## File format
20
-
21
- The file should be written in a relaxed, readable style — more like a short design document than a database entry. Use paragraphs, code samples, and examples to make the reasoning clear and useful to someone encountering it for the first time.
22
-
23
- ```markdown
24
- # Dark Mode
25
-
26
- This project does not support dark mode or user-facing theming.
27
-
28
- ## Why this is out of scope
29
-
30
- The rendering pipeline assumes a single color palette defined in
31
- `ThemeConfig`. Supporting multiple themes would require:
32
-
33
- - A theme context provider wrapping the entire component tree
34
- - Per-component theme-aware style resolution
35
- - A persistence layer for user theme preferences
36
-
37
- This is a significant architectural change that doesn't align with the
38
- project's focus on content authoring. Theming is a concern for downstream
39
- consumers who embed or redistribute the output.
40
-
41
- ```ts
42
- // The current ThemeConfig interface is not designed for runtime switching:
43
- interface ThemeConfig {
44
- colors: ColorPalette; // single palette, resolved at build time
45
- fonts: FontStack;
46
- }
47
- ```
48
-
49
- ## Prior requests
50
-
51
- - #42 — "Add dark mode support"
52
- - #87 — "Night theme for accessibility"
53
- - #134 — "Dark theme option"
54
- ```
55
-
56
- ### Naming the file
57
-
58
- Use a short, descriptive kebab-case name for the concept: `dark-mode.md`, `plugin-system.md`, `graphql-api.md`. The name should be recognizable enough that someone browsing the directory understands what was rejected without opening the file.
59
-
60
- ### Writing the reason
61
-
62
- The reason should be substantive — not "we don't want this" but why. Good reasons reference:
63
-
64
- - Project scope or philosophy ("This project focuses on X; theming is a downstream concern")
65
- - Technical constraints ("Supporting this would require Y, which conflicts with our Z architecture")
66
- - Strategic decisions ("We chose to use A instead of B because...")
67
-
68
- The reason should be durable. Avoid referencing temporary circumstances ("we're too busy right now") — those aren't real rejections, they're deferrals.
69
-
70
- ## When to check `.out-of-scope/`
71
-
72
- During triage (Step 1: Gather context), read all files in `.out-of-scope/`. When evaluating a new issue:
73
-
74
- - Check if the request matches an existing out-of-scope concept
75
- - Matching is by concept similarity, not keyword — "night theme" matches `dark-mode.md`
76
- - If there's a match, surface it to the maintainer: "This is similar to `.out-of-scope/dark-mode.md` — we rejected this before because [reason]. Do you still feel the same way?"
77
-
78
- The maintainer may:
79
-
80
- - **Confirm** — the new issue gets added to the existing file's "Prior requests" list, then closed
81
- - **Reconsider** — the out-of-scope file gets deleted or updated, and the issue proceeds through normal triage
82
- - **Disagree** — the issues are related but distinct, proceed with normal triage
83
-
84
- ## When to write to `.out-of-scope/`
85
-
86
- Only when an **enhancement** (not a bug) is rejected as `wontfix`. The flow:
87
-
88
- 1. Maintainer decides a feature request is out of scope
89
- 2. Check if a matching `.out-of-scope/` file already exists
90
- 3. If yes: append the new issue to the "Prior requests" list
91
- 4. If no: create a new file with the concept name, decision, reason, and first prior request
92
- 5. Post a comment on the issue explaining the decision and mentioning the `.out-of-scope/` file
93
- 6. Close the issue with the `wontfix` label
94
-
95
- ## Updating or removing out-of-scope files
96
-
97
- If the maintainer changes their mind about a previously rejected concept:
98
-
99
- - Delete the `.out-of-scope/` file
100
- - The skill does not need to reopen old issues — they're historical records
101
- - The new issue that triggered the reconsideration proceeds through normal triage