codex-orchestrator 2.0.1 → 2.0.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (211) hide show
  1. package/CHANGELOG.md +22 -0
  2. package/README.md +12 -9
  3. package/dist/src/index.d.ts +10 -0
  4. package/dist/src/index.d.ts.map +1 -1
  5. package/dist/src/index.js +5 -0
  6. package/dist/src/index.js.map +1 -1
  7. package/dist/src/v2/acceptance-proof.d.ts +3 -0
  8. package/dist/src/v2/acceptance-proof.d.ts.map +1 -1
  9. package/dist/src/v2/acceptance-proof.js +2 -8
  10. package/dist/src/v2/acceptance-proof.js.map +1 -1
  11. package/dist/src/v2/adapters/gh-issue-adapter.d.ts +5 -3
  12. package/dist/src/v2/adapters/gh-issue-adapter.d.ts.map +1 -1
  13. package/dist/src/v2/adapters/gh-issue-adapter.js +63 -7
  14. package/dist/src/v2/adapters/gh-issue-adapter.js.map +1 -1
  15. package/dist/src/v2/adapters/issues.d.ts +16 -2
  16. package/dist/src/v2/adapters/issues.d.ts.map +1 -1
  17. package/dist/src/v2/adapters/issues.js +15 -5
  18. package/dist/src/v2/adapters/issues.js.map +1 -1
  19. package/dist/src/v2/adapters/mission-coordinator-lock.d.ts +1 -0
  20. package/dist/src/v2/adapters/mission-coordinator-lock.d.ts.map +1 -1
  21. package/dist/src/v2/adapters/mission-coordinator-lock.js +5 -1
  22. package/dist/src/v2/adapters/mission-coordinator-lock.js.map +1 -1
  23. package/dist/src/v2/candidate-cli.d.ts +4 -0
  24. package/dist/src/v2/candidate-cli.d.ts.map +1 -1
  25. package/dist/src/v2/candidate-cli.js +26 -11
  26. package/dist/src/v2/candidate-cli.js.map +1 -1
  27. package/dist/src/v2/cli-contract.d.ts +1 -1
  28. package/dist/src/v2/cli-contract.d.ts.map +1 -1
  29. package/dist/src/v2/cli-contract.js +10 -0
  30. package/dist/src/v2/cli-contract.js.map +1 -1
  31. package/dist/src/v2/code-review-report.d.ts +66 -0
  32. package/dist/src/v2/code-review-report.d.ts.map +1 -0
  33. package/dist/src/v2/code-review-report.js +259 -0
  34. package/dist/src/v2/code-review-report.js.map +1 -0
  35. package/dist/src/v2/codex-process.d.ts +8 -1
  36. package/dist/src/v2/codex-process.d.ts.map +1 -1
  37. package/dist/src/v2/codex-process.js +11 -0
  38. package/dist/src/v2/codex-process.js.map +1 -1
  39. package/dist/src/v2/config.d.ts +2 -1
  40. package/dist/src/v2/config.d.ts.map +1 -1
  41. package/dist/src/v2/config.js +8 -3
  42. package/dist/src/v2/config.js.map +1 -1
  43. package/dist/src/v2/contained-report-operation.d.ts +100 -0
  44. package/dist/src/v2/contained-report-operation.d.ts.map +1 -0
  45. package/dist/src/v2/contained-report-operation.js +200 -0
  46. package/dist/src/v2/contained-report-operation.js.map +1 -0
  47. package/dist/src/v2/containment.d.ts +6 -0
  48. package/dist/src/v2/containment.d.ts.map +1 -1
  49. package/dist/src/v2/containment.js +40 -1
  50. package/dist/src/v2/containment.js.map +1 -1
  51. package/dist/src/v2/direct-delivery.d.ts +101 -0
  52. package/dist/src/v2/direct-delivery.d.ts.map +1 -0
  53. package/dist/src/v2/direct-delivery.js +547 -0
  54. package/dist/src/v2/direct-delivery.js.map +1 -0
  55. package/dist/src/v2/immutable-workflow-publisher.d.ts +40 -0
  56. package/dist/src/v2/immutable-workflow-publisher.d.ts.map +1 -0
  57. package/dist/src/v2/immutable-workflow-publisher.js +218 -0
  58. package/dist/src/v2/immutable-workflow-publisher.js.map +1 -0
  59. package/dist/src/v2/implementation-reviewer.d.ts +81 -0
  60. package/dist/src/v2/implementation-reviewer.d.ts.map +1 -0
  61. package/dist/src/v2/implementation-reviewer.js +157 -0
  62. package/dist/src/v2/implementation-reviewer.js.map +1 -0
  63. package/dist/src/v2/owner-control-lock.d.ts +41 -0
  64. package/dist/src/v2/owner-control-lock.d.ts.map +1 -0
  65. package/dist/src/v2/owner-control-lock.js +174 -0
  66. package/dist/src/v2/owner-control-lock.js.map +1 -0
  67. package/dist/src/v2/route-continuations.d.ts +32 -0
  68. package/dist/src/v2/route-continuations.d.ts.map +1 -0
  69. package/dist/src/v2/route-continuations.js +2 -0
  70. package/dist/src/v2/route-continuations.js.map +1 -0
  71. package/dist/src/v2/route-coordinator.d.ts +77 -0
  72. package/dist/src/v2/route-coordinator.d.ts.map +1 -0
  73. package/dist/src/v2/route-coordinator.js +370 -0
  74. package/dist/src/v2/route-coordinator.js.map +1 -0
  75. package/dist/src/v2/route-decision.d.ts +129 -0
  76. package/dist/src/v2/route-decision.d.ts.map +1 -0
  77. package/dist/src/v2/route-decision.js +400 -0
  78. package/dist/src/v2/route-decision.js.map +1 -0
  79. package/dist/src/v2/run-issue.d.ts +63 -2
  80. package/dist/src/v2/run-issue.d.ts.map +1 -1
  81. package/dist/src/v2/run-issue.js +906 -91
  82. package/dist/src/v2/run-issue.js.map +1 -1
  83. package/dist/src/v2/run-store.d.ts +25 -1
  84. package/dist/src/v2/run-store.d.ts.map +1 -1
  85. package/dist/src/v2/run-store.js +143 -3
  86. package/dist/src/v2/run-store.js.map +1 -1
  87. package/dist/src/v2/runtime-assets.d.ts +15 -13
  88. package/dist/src/v2/runtime-assets.d.ts.map +1 -1
  89. package/dist/src/v2/runtime-assets.js +263 -416
  90. package/dist/src/v2/runtime-assets.js.map +1 -1
  91. package/dist/src/v2/runtime.d.ts +14 -6
  92. package/dist/src/v2/runtime.d.ts.map +1 -1
  93. package/dist/src/v2/runtime.js +478 -56
  94. package/dist/src/v2/runtime.js.map +1 -1
  95. package/dist/src/v2/setup-cli.d.ts.map +1 -1
  96. package/dist/src/v2/setup-cli.js +1 -0
  97. package/dist/src/v2/setup-cli.js.map +1 -1
  98. package/dist/src/v2/setup-runtime.d.ts.map +1 -1
  99. package/dist/src/v2/setup-runtime.js +20 -72
  100. package/dist/src/v2/setup-runtime.js.map +1 -1
  101. package/dist/src/v2/setup.d.ts +4 -1
  102. package/dist/src/v2/setup.d.ts.map +1 -1
  103. package/dist/src/v2/setup.js +104 -1
  104. package/dist/src/v2/setup.js.map +1 -1
  105. package/dist/src/v2/spec-coordinator.d.ts +85 -0
  106. package/dist/src/v2/spec-coordinator.d.ts.map +1 -0
  107. package/dist/src/v2/spec-coordinator.js +88 -0
  108. package/dist/src/v2/spec-coordinator.js.map +1 -0
  109. package/dist/src/v2/spec-delivery.d.ts +143 -0
  110. package/dist/src/v2/spec-delivery.d.ts.map +1 -0
  111. package/dist/src/v2/spec-delivery.js +401 -0
  112. package/dist/src/v2/spec-delivery.js.map +1 -0
  113. package/dist/src/v2/triage-route.d.ts +68 -0
  114. package/dist/src/v2/triage-route.d.ts.map +1 -0
  115. package/dist/src/v2/triage-route.js +223 -0
  116. package/dist/src/v2/triage-route.js.map +1 -0
  117. package/dist/src/v2/waiting-human-coordinator.d.ts +49 -0
  118. package/dist/src/v2/waiting-human-coordinator.d.ts.map +1 -0
  119. package/dist/src/v2/waiting-human-coordinator.js +509 -0
  120. package/dist/src/v2/waiting-human-coordinator.js.map +1 -0
  121. package/dist/src/v2/waiting-human.d.ts +143 -0
  122. package/dist/src/v2/waiting-human.d.ts.map +1 -0
  123. package/dist/src/v2/waiting-human.js +408 -0
  124. package/dist/src/v2/waiting-human.js.map +1 -0
  125. package/dist/src/v2/workflow-assets.d.ts +90 -0
  126. package/dist/src/v2/workflow-assets.d.ts.map +1 -0
  127. package/dist/src/v2/workflow-assets.js +554 -0
  128. package/dist/src/v2/workflow-assets.js.map +1 -0
  129. package/docs/deep-dive.md +15 -8
  130. package/internal-workflow/docs/agents/artifact-review-loop.md +267 -0
  131. package/internal-workflow/docs/agents/bug-workflow-routing.md +24 -0
  132. package/internal-workflow/docs/agents/coding-skill-routing.md +203 -0
  133. package/internal-workflow/docs/agents/confidence-rubric.md +65 -0
  134. package/internal-workflow/docs/agents/contract-test-ledger.md +60 -0
  135. package/internal-workflow/docs/agents/implementation-review-loop.md +302 -0
  136. package/internal-workflow/docs/agents/review-gates.md +49 -0
  137. package/internal-workflow/docs/agents/review-protocol.md +170 -0
  138. package/internal-workflow/docs/agents/tool-usage.md +88 -0
  139. package/internal-workflow/manifest.json +1 -0
  140. package/internal-workflow/operations/acceptance-proof/SKILL.md +3 -0
  141. package/internal-workflow/operations/ambiguity-review/SKILL.md +3 -0
  142. package/internal-workflow/operations/cleanup-review/SKILL.md +3 -0
  143. package/internal-workflow/operations/code-review/SKILL.md +3 -0
  144. package/internal-workflow/operations/implementation/SKILL.md +3 -0
  145. package/internal-workflow/operations/spec-author/SKILL.md +3 -0
  146. package/internal-workflow/operations/spec-implementation/SKILL.md +3 -0
  147. package/internal-workflow/operations/spec-review/SKILL.md +3 -0
  148. package/internal-workflow/operations/triage/SKILL.md +3 -0
  149. package/internal-workflow/profiles/analyst_deep.toml +9 -0
  150. package/internal-workflow/profiles/implementer_deep.toml +9 -0
  151. package/internal-workflow/profiles/implementer_standard.toml +9 -0
  152. package/internal-workflow/profiles/proof_agent.toml +8 -0
  153. package/internal-workflow/profiles/researcher_standard.toml +9 -0
  154. package/internal-workflow/profiles/reviewer_deep.toml +9 -0
  155. package/internal-workflow/profiles/reviewer_fast.toml +9 -0
  156. package/internal-workflow/profiles/reviewer_standard.toml +9 -0
  157. package/internal-workflow/schemas/ambiguity-review-v1.json +1 -0
  158. package/internal-workflow/schemas/code-review-v1.json +1 -0
  159. package/internal-workflow/schemas/implementation-report-v1.json +1 -0
  160. package/internal-workflow/schemas/proof-report-v1.json +1 -0
  161. package/internal-workflow/schemas/spec-author-v1.json +1 -0
  162. package/internal-workflow/schemas/spec-review-v1.json +30 -0
  163. package/internal-workflow/schemas/triage-route-v1.json +1 -0
  164. package/internal-workflow/skills/acceptance-proof/agents/openai.yaml +6 -0
  165. package/internal-workflow/skills/agent-auto/agents/openai.yaml +6 -0
  166. package/internal-workflow/skills/cleanup-review/SKILL.md +84 -0
  167. package/internal-workflow/skills/cleanup-review/agents/openai.yaml +6 -0
  168. package/internal-workflow/skills/code-review/SKILL.md +257 -0
  169. package/internal-workflow/skills/code-review/agents/openai.yaml +4 -0
  170. package/internal-workflow/skills/code-review/references/bug-classes.md +56 -0
  171. package/internal-workflow/skills/code-review/references/framework-lenses.md +34 -0
  172. package/internal-workflow/skills/code-review/references/targeted-recipes.md +49 -0
  173. package/internal-workflow/skills/codebase-design/DEEPENING.md +35 -0
  174. package/internal-workflow/skills/codebase-design/DESIGN-IT-TWICE.md +50 -0
  175. package/internal-workflow/skills/codebase-design/SKILL.md +82 -0
  176. package/internal-workflow/skills/codebase-design/agents/openai.yaml +6 -0
  177. package/internal-workflow/skills/diagnosing-bugs/SKILL.md +138 -0
  178. package/internal-workflow/skills/diagnosing-bugs/agents/openai.yaml +6 -0
  179. package/internal-workflow/skills/diagnosing-bugs/scripts/hitl-loop.template.sh +41 -0
  180. package/internal-workflow/skills/implementation-spec-maker/SKILL.md +93 -0
  181. package/internal-workflow/skills/implementation-spec-maker/agents/openai.yaml +6 -0
  182. package/internal-workflow/skills/implementation-spec-maker/references/source-modes.md +31 -0
  183. package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +146 -0
  184. package/internal-workflow/skills/implementation-spec-review/SKILL.md +211 -0
  185. package/internal-workflow/skills/implementation-spec-review/agents/openai.yaml +6 -0
  186. package/internal-workflow/skills/research/SKILL.md +107 -0
  187. package/internal-workflow/skills/research/agents/openai.yaml +6 -0
  188. package/internal-workflow/skills/small-task-implementer/SKILL.md +97 -0
  189. package/internal-workflow/skills/small-task-implementer/agents/openai.yaml +6 -0
  190. package/internal-workflow/skills/spec-implementer/SKILL.md +197 -0
  191. package/internal-workflow/skills/spec-implementer/agents/openai.yaml +6 -0
  192. package/internal-workflow/skills/tdd/SKILL.md +59 -0
  193. package/internal-workflow/skills/tdd/agents/openai.yaml +6 -0
  194. package/internal-workflow/skills/tdd/interface-design.md +31 -0
  195. package/internal-workflow/skills/tdd/mocking.md +59 -0
  196. package/internal-workflow/skills/tdd/refactoring.md +10 -0
  197. package/internal-workflow/skills/tdd/tests.md +77 -0
  198. package/internal-workflow/skills/triage/AGENT-BRIEF.md +192 -0
  199. package/internal-workflow/skills/triage/OUT-OF-SCOPE.md +101 -0
  200. package/internal-workflow/skills/triage/SKILL.md +134 -0
  201. package/internal-workflow/skills/triage/agents/openai.yaml +6 -0
  202. package/internal-workflow/skills/ui-evidence-proof/SKILL.md +123 -0
  203. package/internal-workflow/skills/ui-evidence-proof/agents/openai.yaml +6 -0
  204. package/package.json +6 -3
  205. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/SKILL.md +0 -0
  206. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/android.md +0 -0
  207. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/browser.md +0 -0
  208. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/references/ios.md +0 -0
  209. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/tools/android-lease.mjs +0 -0
  210. /package/{internal-skills → internal-workflow/skills}/acceptance-proof/tools/ios-lease.mjs +0 -0
  211. /package/{internal-skills → internal-workflow/skills}/agent-auto/SKILL.md +0 -0
@@ -0,0 +1,59 @@
1
+ ---
2
+ name: tdd
3
+ description: Test-driven development policy gate for implementation, bugfix, and new feature work unless the user explicitly opts out. Use before planning or editing code to shape the first behavior proof, and when the user mentions red-green-refactor, integration tests, or test-first development.
4
+ ---
5
+
6
+ # Test-Driven Development
7
+
8
+ Use short vertical RED -> GREEN cycles. Make each test prove observable behavior through the same public seam real callers use.
9
+
10
+ ## Core Contract
11
+
12
+ - Lock expected behavior from the request, specification, design, bug report, or existing product behavior before changing implementation.
13
+ - Derive expected values from an independent source, never from the production algorithm.
14
+ - Prove RED on the old behavior for the same observable reason the user reported or requested.
15
+ - Add only enough implementation to make the current test pass; do not anticipate later tests.
16
+ - Keep tests stable across behavior-preserving refactors and refactor only while GREEN.
17
+
18
+ Read [tests.md](tests.md) when choosing or reviewing test shape. Read [mocking.md](mocking.md) before introducing test doubles.
19
+
20
+ ## Before the First RED
21
+
22
+ 1. Read local instructions, domain language, existing tests, and relevant ADRs.
23
+ 2. List the prioritized observable behaviors, not implementation steps.
24
+ 3. Select the public seam where callers observe each behavior.
25
+ 4. Ask the user only when the seam changes the public contract, product intent is unclear, or behavior priorities materially conflict.
26
+ 5. For contract-risk changes, create or update the shared [Contract Test Ledger](../../docs/agents/contract-test-ledger.md) and map each invariant to its first failing test or observable proof.
27
+ 6. If no natural public seam exists, consult [interface-design.md](interface-design.md) instead of testing internals.
28
+
29
+ For UI behavior, define proof at the rendered seam: visible content and order, interaction result, semantics, or screenshot when layout direction or scrolling matters.
30
+
31
+ ## RED -> GREEN Cycle
32
+
33
+ For each behavior:
34
+
35
+ 1. **RED:** Write one test through the selected seam.
36
+ 2. Confirm it fails on current behavior for the expected reason. A passing test or an internal-only failure is not valid RED.
37
+ 3. **GREEN:** Add the minimal implementation required for that test.
38
+ 4. Run the proof and update the ledger status to `red`, `green`, or `blocked` with the missing seam or evidence.
39
+
40
+ Keep each cycle to one seam, one behavior, one test, and one minimal implementation. Do not rewrite the test to fit the code. For state, async, lifecycle, retry, cache, or auth defects, include the competing condition when feasible.
41
+
42
+ Handle reviewer repairs inside the same activation only under [bug workflow routing](../../docs/agents/bug-workflow-routing.md); group related cases by protected invariant.
43
+
44
+ ## After GREEN
45
+
46
+ Refactor as a separate review-stage activity, never while RED. Use [refactoring.md](refactoring.md) for candidates and rerun affected tests after each step.
47
+
48
+ ## Cycle Checklist
49
+
50
+ ```text
51
+ [ ] Behavior is proved through the caller's public seam
52
+ [ ] Expected behavior is locked and the expected value is independent
53
+ [ ] RED fails on old behavior for the correct observable reason
54
+ [ ] Test was not fitted to implementation details
55
+ [ ] GREEN uses only the code needed for the current behavior
56
+ [ ] Final outcome and relevant competing condition are proved
57
+ [ ] Contract Test Ledger is current when applicable
58
+ [ ] Refactoring starts only after GREEN
59
+ ```
@@ -0,0 +1,6 @@
1
+ interface:
2
+ display_name: "Test-Driven Development"
3
+ short_description: "Apply behavior-first red-green development"
4
+ default_prompt: "Use $tdd to define the public test seam and implement this behavior through red-green cycles."
5
+ policy:
6
+ allow_implicit_invocation: true
@@ -0,0 +1,31 @@
1
+ # Interface Design for Testability
2
+
3
+ Good interfaces make testing natural:
4
+
5
+ 1. **Accept dependencies, don't create them**
6
+
7
+ ```typescript
8
+ // Testable
9
+ function processOrder(order, paymentGateway) {}
10
+
11
+ // Hard to test
12
+ function processOrder(order) {
13
+ const gateway = new StripeGateway();
14
+ }
15
+ ```
16
+
17
+ 2. **Return results, don't produce side effects**
18
+
19
+ ```typescript
20
+ // Testable
21
+ function calculateDiscount(cart): Discount {}
22
+
23
+ // Hard to test
24
+ function applyDiscount(cart): void {
25
+ cart.total -= discount;
26
+ }
27
+ ```
28
+
29
+ 3. **Small surface area**
30
+ - Fewer methods = fewer tests needed
31
+ - Fewer params = simpler test setup
@@ -0,0 +1,59 @@
1
+ # When to Mock
2
+
3
+ Mock at **system boundaries** only:
4
+
5
+ - External APIs (payment, email, etc.)
6
+ - Databases (sometimes - prefer test DB)
7
+ - Time/randomness
8
+ - File system (sometimes)
9
+
10
+ Don't mock:
11
+
12
+ - Your own classes/modules
13
+ - Internal collaborators
14
+ - Anything you control
15
+
16
+ ## Designing for Mockability
17
+
18
+ At system boundaries, design interfaces that are easy to mock:
19
+
20
+ **1. Use dependency injection**
21
+
22
+ Pass external dependencies in rather than creating them internally:
23
+
24
+ ```typescript
25
+ // Easy to mock
26
+ function processPayment(order, paymentClient) {
27
+ return paymentClient.charge(order.total);
28
+ }
29
+
30
+ // Hard to mock
31
+ function processPayment(order) {
32
+ const client = new StripeClient(process.env.STRIPE_KEY);
33
+ return client.charge(order.total);
34
+ }
35
+ ```
36
+
37
+ **2. Prefer SDK-style interfaces over generic fetchers**
38
+
39
+ Create specific functions for each external operation instead of one generic function with conditional logic:
40
+
41
+ ```typescript
42
+ // GOOD: Each function is independently mockable
43
+ const api = {
44
+ getUser: (id) => fetch(`/users/${id}`),
45
+ getOrders: (userId) => fetch(`/users/${userId}/orders`),
46
+ createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
47
+ };
48
+
49
+ // BAD: Mocking requires conditional logic inside the mock
50
+ const api = {
51
+ fetch: (endpoint, options) => fetch(endpoint, options),
52
+ };
53
+ ```
54
+
55
+ The SDK approach means:
56
+ - Each mock returns one specific shape
57
+ - No conditional logic in test setup
58
+ - Easier to see which endpoints a test exercises
59
+ - Type safety per endpoint
@@ -0,0 +1,10 @@
1
+ # Refactor Candidates
2
+
3
+ After TDD cycle, look for:
4
+
5
+ - **Duplication** → Extract function/class
6
+ - **Long methods** → Break into private helpers (keep tests on public interface)
7
+ - **Shallow modules** → Combine or deepen
8
+ - **Feature envy** → Move logic to where data lives
9
+ - **Primitive obsession** → Introduce value objects
10
+ - **Existing code** the new code reveals as problematic
@@ -0,0 +1,77 @@
1
+ # Good and Bad Tests
2
+
3
+ ## Good Tests
4
+
5
+ **Integration-style**: Test through real interfaces, not mocks of internal parts.
6
+
7
+ ```typescript
8
+ // GOOD: Tests observable behavior
9
+ test("user can checkout with valid cart", async () => {
10
+ const cart = createCart();
11
+ cart.add(product);
12
+ const result = await checkout(cart, paymentMethod);
13
+ expect(result.status).toBe("confirmed");
14
+ });
15
+ ```
16
+
17
+ Characteristics:
18
+
19
+ - Tests behavior users/callers care about
20
+ - Uses public API only
21
+ - Survives internal refactors
22
+ - Describes WHAT, not HOW
23
+ - One logical assertion per test
24
+
25
+ ## Bad Tests
26
+
27
+ **Implementation-detail tests**: Coupled to internal structure.
28
+
29
+ ```typescript
30
+ // BAD: Tests implementation details
31
+ test("checkout calls paymentService.process", async () => {
32
+ const mockPayment = jest.mock(paymentService);
33
+ await checkout(cart, payment);
34
+ expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
35
+ });
36
+ ```
37
+
38
+ Red flags:
39
+
40
+ - Mocking internal collaborators
41
+ - Testing private methods
42
+ - Asserting on call counts/order
43
+ - Test breaks when refactoring without behavior change
44
+ - Test name describes HOW not WHAT
45
+ - Verifying through external means instead of interface
46
+
47
+ ```typescript
48
+ // BAD: Bypasses interface to verify
49
+ test("createUser saves to database", async () => {
50
+ await createUser({ name: "Alice" });
51
+ const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
52
+ expect(row).toBeDefined();
53
+ });
54
+
55
+ // GOOD: Verifies through interface
56
+ test("createUser makes user retrievable", async () => {
57
+ const user = await createUser({ name: "Alice" });
58
+ const retrieved = await getUser(user.id);
59
+ expect(retrieved.name).toBe("Alice");
60
+ });
61
+ ```
62
+
63
+ **Tautological tests**: Expected value restates the implementation, so the test passes by construction.
64
+
65
+ ```typescript
66
+ // BAD: Expected value is recomputed the way the code computes it
67
+ test("calculateTotal sums line items", () => {
68
+ const items = [{ price: 10 }, { price: 5 }];
69
+ const expected = items.reduce((sum, item) => sum + item.price, 0);
70
+ expect(calculateTotal(items)).toBe(expected);
71
+ });
72
+
73
+ // GOOD: Expected value comes from an independent, known literal
74
+ test("calculateTotal sums line items", () => {
75
+ expect(calculateTotal([{ price: 10 }, { price: 5 }])).toBe(15);
76
+ });
77
+ ```
@@ -0,0 +1,192 @@
1
+ # Writing Agent Briefs
2
+
3
+ An agent brief is a structured comment posted when a raw incoming issue moves
4
+ to `ready-for-agent` through `$triage`. It is the authoritative specification
5
+ for that AFK issue; the original body and discussion remain context.
6
+
7
+ Do not post this comment for tickets generated by `$to-tickets`. Their approved
8
+ issue body is already the authoritative execution brief and they bypass triage.
9
+
10
+ ## Principles
11
+
12
+ ### Durability over precision
13
+
14
+ The issue may sit in `ready-for-agent` for days or weeks. The codebase will change in the meantime. Write the brief so it stays useful even as files are renamed, moved, or refactored.
15
+
16
+ - **Do** describe interfaces, types, and behavioral contracts
17
+ - **Do** name specific types, function signatures, or config shapes that the agent should look for or modify
18
+ - **Don't** reference file paths — they go stale
19
+ - **Don't** reference line numbers
20
+ - **Don't** assume the current implementation structure will remain the same
21
+
22
+ ### Behavioral, not procedural
23
+
24
+ Describe **what** the system should do, not **how** to implement it. The agent will explore the codebase fresh and make its own implementation decisions.
25
+
26
+ - **Good:** "The `SkillConfig` type should accept an optional `schedule` field of type `CronExpression`"
27
+ - **Bad:** "Open src/types/skill.ts and add a schedule field on line 42"
28
+ - **Good:** "When a user runs `$triage` with no arguments, they should see a summary of issues needing attention"
29
+ - **Bad:** "Add a switch statement in the main handler function"
30
+
31
+ ### Complete acceptance criteria
32
+
33
+ The agent needs to know when it's done. Every agent brief must have concrete, testable acceptance criteria. Each criterion should be independently verifiable.
34
+
35
+ - **Good:** "Running `gh issue list --label needs-triage` returns issues that have been through initial classification"
36
+ - **Bad:** "Triage should work correctly"
37
+
38
+ ### Explicit scope boundaries
39
+
40
+ State what is out of scope. This prevents the agent from gold-plating or making assumptions about adjacent features.
41
+
42
+ ### Proportional execution metadata
43
+
44
+ Preserve an existing spec gate, verification contract, and material risk notes
45
+ without turning the brief into an implementation spec. Omit those conditional
46
+ fields when they are not material. High review risk changes proof depth, not
47
+ brief length or solution breadth.
48
+
49
+ ## Template
50
+
51
+ ```markdown
52
+ ## Agent Brief
53
+
54
+ **Category:** bug / enhancement
55
+ **Summary:** one-line description of what needs to happen
56
+
57
+ **Current behavior:**
58
+ Describe what happens now. For bugs, this is the broken behavior.
59
+ For enhancements, this is the status quo the feature builds on.
60
+
61
+ **Desired behavior:**
62
+ Describe what should happen after the agent's work is complete.
63
+ Be specific about edge cases and error conditions.
64
+
65
+ **Key interfaces:**
66
+ - `TypeName` — what needs to change and why
67
+ - `functionName()` return type — what it currently returns vs what it should return
68
+ - Config shape — any new configuration options needed
69
+
70
+ **Acceptance criteria:**
71
+ - [ ] Specific, testable criterion 1
72
+ - [ ] Specific, testable criterion 2
73
+ - [ ] Specific, testable criterion 3
74
+
75
+ **Implementation preparation:**
76
+ - direct / compact spec / standard spec
77
+ - Reason or accepted implementation spec reference
78
+
79
+ **Verification:**
80
+ - Behavior proof expected from the implementing agent
81
+ - Required automated or live check, or `Not specified`
82
+
83
+ **Risk / Review:**
84
+ - Omit when no material risk exists
85
+ - Otherwise: primary risk, main invariant, review focus, and final handoff expectation
86
+
87
+ **Out of scope:**
88
+ - Thing that should NOT be changed or addressed in this issue
89
+ - Adjacent feature that might seem related but is separate
90
+ ```
91
+
92
+ ## Examples
93
+
94
+ ### Good agent brief (bug)
95
+
96
+ ```markdown
97
+ ## Agent Brief
98
+
99
+ **Category:** bug
100
+ **Summary:** Skill description truncation drops mid-word, producing broken output
101
+
102
+ **Current behavior:**
103
+ When a skill description exceeds 1024 characters, it is truncated at exactly
104
+ 1024 characters regardless of word boundaries. This produces descriptions
105
+ that end mid-word (e.g. "Use when the user wants to confi").
106
+
107
+ **Desired behavior:**
108
+ Truncation should break at the last word boundary before 1024 characters
109
+ and append "..." to indicate truncation.
110
+
111
+ **Key interfaces:**
112
+ - The `SkillMetadata` type's `description` field — no type change needed,
113
+ but the validation/processing logic that populates it needs to respect
114
+ word boundaries
115
+ - Any function that reads SKILL.md frontmatter and extracts the description
116
+
117
+ **Acceptance criteria:**
118
+ - [ ] Descriptions under 1024 chars are unchanged
119
+ - [ ] Descriptions over 1024 chars are truncated at the last word boundary
120
+ before 1024 chars
121
+ - [ ] Truncated descriptions end with "..."
122
+ - [ ] The total length including "..." does not exceed 1024 chars
123
+
124
+ **Out of scope:**
125
+ - Changing the 1024 char limit itself
126
+ - Multi-line description support
127
+ ```
128
+
129
+ ### Good agent brief (enhancement)
130
+
131
+ ```markdown
132
+ ## Agent Brief
133
+
134
+ **Category:** enhancement
135
+ **Summary:** Add `.out-of-scope/` directory support for tracking rejected feature requests
136
+
137
+ **Current behavior:**
138
+ When a feature request is rejected, the issue is closed with a `wontfix` label
139
+ and a comment. There is no persistent record of the decision or reasoning.
140
+ Future similar requests require the maintainer to recall or search for the
141
+ prior discussion.
142
+
143
+ **Desired behavior:**
144
+ Rejected feature requests should be documented in `.out-of-scope/<concept>.md`
145
+ files that capture the decision, reasoning, and links to all issues that
146
+ requested the feature. When triaging new issues, these files should be
147
+ checked for matches.
148
+
149
+ **Key interfaces:**
150
+ - Markdown file format in `.out-of-scope/` — each file should have a
151
+ `# Concept Name` heading, a `**Decision:**` line, a `**Reason:**` line,
152
+ and a `**Prior requests:**` list with issue links
153
+ - The triage workflow should read all `.out-of-scope/*.md` files early
154
+ and match incoming issues against them by concept similarity
155
+
156
+ **Acceptance criteria:**
157
+ - [ ] Closing a feature as wontfix creates/updates a file in `.out-of-scope/`
158
+ - [ ] The file includes the decision, reasoning, and link to the closed issue
159
+ - [ ] If a matching `.out-of-scope/` file already exists, the new issue is
160
+ appended to its "Prior requests" list rather than creating a duplicate
161
+ - [ ] During triage, existing `.out-of-scope/` files are checked and surfaced
162
+ when a new issue matches a prior rejection
163
+
164
+ **Out of scope:**
165
+ - Automated matching (human confirms the match)
166
+ - Reopening previously rejected features
167
+ - Bug reports (only enhancement rejections go to `.out-of-scope/`)
168
+ ```
169
+
170
+ ### Bad agent brief
171
+
172
+ ```markdown
173
+ ## Agent Brief
174
+
175
+ **Summary:** Fix the triage bug
176
+
177
+ **What to do:**
178
+ The triage thing is broken. Look at the main file and fix it.
179
+ The function around line 150 has the issue.
180
+
181
+ **Files to change:**
182
+ - src/triage/handler.ts (line 150)
183
+ - src/types.ts (line 42)
184
+ ```
185
+
186
+ This is bad because:
187
+ - No category
188
+ - Vague description ("the triage thing is broken")
189
+ - References file paths and line numbers that will go stale
190
+ - No acceptance criteria
191
+ - No scope boundaries
192
+ - No description of current vs desired behavior
@@ -0,0 +1,101 @@
1
+ # Out-of-Scope Knowledge Base
2
+
3
+ The `.out-of-scope/` directory in a repo stores persistent records of rejected feature requests. It serves two purposes:
4
+
5
+ 1. **Institutional memory** — why a feature was rejected, so the reasoning isn't lost when the issue is closed
6
+ 2. **Deduplication** — when a new issue comes in that matches a prior rejection, the skill can surface the previous decision instead of re-litigating it
7
+
8
+ ## Directory structure
9
+
10
+ ```
11
+ .out-of-scope/
12
+ ├── dark-mode.md
13
+ ├── plugin-system.md
14
+ └── graphql-api.md
15
+ ```
16
+
17
+ One file per **concept**, not per issue. Multiple issues requesting the same thing are grouped under one file.
18
+
19
+ ## File format
20
+
21
+ The file should be written in a relaxed, readable style — more like a short design document than a database entry. Use paragraphs, code samples, and examples to make the reasoning clear and useful to someone encountering it for the first time.
22
+
23
+ ```markdown
24
+ # Dark Mode
25
+
26
+ This project does not support dark mode or user-facing theming.
27
+
28
+ ## Why this is out of scope
29
+
30
+ The rendering pipeline assumes a single color palette defined in
31
+ `ThemeConfig`. Supporting multiple themes would require:
32
+
33
+ - A theme context provider wrapping the entire component tree
34
+ - Per-component theme-aware style resolution
35
+ - A persistence layer for user theme preferences
36
+
37
+ This is a significant architectural change that doesn't align with the
38
+ project's focus on content authoring. Theming is a concern for downstream
39
+ consumers who embed or redistribute the output.
40
+
41
+ ```ts
42
+ // The current ThemeConfig interface is not designed for runtime switching:
43
+ interface ThemeConfig {
44
+ colors: ColorPalette; // single palette, resolved at build time
45
+ fonts: FontStack;
46
+ }
47
+ ```
48
+
49
+ ## Prior requests
50
+
51
+ - #42 — "Add dark mode support"
52
+ - #87 — "Night theme for accessibility"
53
+ - #134 — "Dark theme option"
54
+ ```
55
+
56
+ ### Naming the file
57
+
58
+ Use a short, descriptive kebab-case name for the concept: `dark-mode.md`, `plugin-system.md`, `graphql-api.md`. The name should be recognizable enough that someone browsing the directory understands what was rejected without opening the file.
59
+
60
+ ### Writing the reason
61
+
62
+ The reason should be substantive — not "we don't want this" but why. Good reasons reference:
63
+
64
+ - Project scope or philosophy ("This project focuses on X; theming is a downstream concern")
65
+ - Technical constraints ("Supporting this would require Y, which conflicts with our Z architecture")
66
+ - Strategic decisions ("We chose to use A instead of B because...")
67
+
68
+ The reason should be durable. Avoid referencing temporary circumstances ("we're too busy right now") — those aren't real rejections, they're deferrals.
69
+
70
+ ## When to check `.out-of-scope/`
71
+
72
+ During triage (Step 1: Gather context), read all files in `.out-of-scope/`. When evaluating a new issue:
73
+
74
+ - Check if the request matches an existing out-of-scope concept
75
+ - Matching is by concept similarity, not keyword — "night theme" matches `dark-mode.md`
76
+ - If there's a match, surface it to the maintainer: "This is similar to `.out-of-scope/dark-mode.md` — we rejected this before because [reason]. Do you still feel the same way?"
77
+
78
+ The maintainer may:
79
+
80
+ - **Confirm** — the new issue gets added to the existing file's "Prior requests" list, then closed
81
+ - **Reconsider** — the out-of-scope file gets deleted or updated, and the issue proceeds through normal triage
82
+ - **Disagree** — the issues are related but distinct, proceed with normal triage
83
+
84
+ ## When to write to `.out-of-scope/`
85
+
86
+ Only when an **enhancement** (not a bug) is rejected as `wontfix`. The flow:
87
+
88
+ 1. Maintainer decides a feature request is out of scope
89
+ 2. Check if a matching `.out-of-scope/` file already exists
90
+ 3. If yes: append the new issue to the "Prior requests" list
91
+ 4. If no: create a new file with the concept name, decision, reason, and first prior request
92
+ 5. Post a comment on the issue explaining the decision and mentioning the `.out-of-scope/` file
93
+ 6. Close the issue with the `wontfix` label
94
+
95
+ ## Updating or removing out-of-scope files
96
+
97
+ If the maintainer changes their mind about a previously rejected concept:
98
+
99
+ - Delete the `.out-of-scope/` file
100
+ - The skill does not need to reopen old issues — they're historical records
101
+ - The new issue that triggered the reconsideration proceeds through normal triage
@@ -0,0 +1,134 @@
1
+ ---
2
+ name: triage
3
+ description: Move raw incoming issues and configured external PRs through a triage state machine — categorise, verify, grill if needed, and write durable agent briefs. Do not use for planning-context issues or tickets generated by `$to-tickets`; those publish in their final state.
4
+ ---
5
+
6
+ # Triage
7
+
8
+ Move raw incoming issues on the project tracker through a small state machine.
9
+ Do not process issues marked `Artifact: planning-context` or generated
10
+ `$to-tickets` children: their approved issue bodies are already the durable
11
+ contract and their final states are applied at publication.
12
+
13
+ If this repo treats external pull requests as a request surface (see the issue-tracker config), triage covers them too: a PR is an issue with attached code. Resolve a bare `#42` to an issue or PR per the tracker config.
14
+
15
+ Every comment or issue posted to the issue tracker during triage **must** start with this disclaimer:
16
+
17
+ ```
18
+ > *This was generated by AI during triage.*
19
+ ```
20
+
21
+ ## Reference docs
22
+
23
+ - [AGENT-BRIEF.md](AGENT-BRIEF.md) — how to write durable agent briefs
24
+ - [OUT-OF-SCOPE.md](OUT-OF-SCOPE.md) — how the `.out-of-scope/` knowledge base works
25
+
26
+ ## Roles
27
+
28
+ Two **category** roles:
29
+
30
+ - `bug` — something is broken
31
+ - `enhancement` — new feature or improvement
32
+
33
+ Five **state** roles:
34
+
35
+ - `needs-triage` — maintainer needs to evaluate
36
+ - `needs-info` — waiting on reporter for more information
37
+ - `ready-for-agent` — fully specified, ready for an AFK agent
38
+ - `ready-for-human` — needs human implementation
39
+ - `wontfix` — will not be actioned
40
+
41
+ For a PR, the same states read against the attached code: `ready-for-agent` means a brief is attached and an agent should take the next step on the diff; `ready-for-human` means it's ready for a human to merge.
42
+
43
+ Every triaged issue should carry exactly one category role and one state role. If state roles conflict, flag it and ask the maintainer before doing anything else.
44
+
45
+ These are canonical role names — the actual label strings used in the issue tracker may differ. The mapping should have been provided to you - run `$setup-matt-pocock-skills` if not.
46
+
47
+ State transitions: an unlabeled issue normally goes to `needs-triage` first; from there it moves to `needs-info`, `ready-for-agent`, `ready-for-human`, or `wontfix`. `needs-info` returns to `needs-triage` once the reporter replies. The maintainer can override at any time — flag transitions that look unusual and ask before proceeding.
48
+
49
+ ## Invocation
50
+
51
+ The maintainer invokes `$triage` and describes what they want in natural language. Interpret the request and act. Examples:
52
+
53
+ - "Show me anything that needs my attention"
54
+ - "Let's look at #42" (issue or PR)
55
+ - "Move #42 to ready-for-agent"
56
+ - "What's ready for agents to pick up?"
57
+
58
+ ## Show what needs attention
59
+
60
+ Query the issue tracker and present three buckets, oldest first:
61
+
62
+ 1. **Unlabeled** — never triaged.
63
+ 2. **`needs-triage`** — evaluation in progress.
64
+ 3. **`needs-info` with reporter activity since the last triage notes** — needs re-evaluation.
65
+
66
+ When PRs are in scope, include external PRs in these buckets and tag each line `[PR]` or `[issue]`. Discovery surfaces only external PRs; a collaborator's in-flight PR is not triage work unless it was named explicitly.
67
+
68
+ Show counts and a one-line summary per item. Let the maintainer pick.
69
+
70
+ ## Triage a specific issue or PR
71
+
72
+ 1. **Gather context.** Read the full issue or PR (body, comments, labels, author, dates; for a PR, the diff too). Parse any prior triage notes so you don't re-ask resolved questions. Explore the codebase using the project's domain glossary, respecting ADRs in the area. Run two checks against the codebase: redundancy — search for an existing implementation of the requested behavior by domain concept, not just wording, and report where you looked. Prior rejection — read `.out-of-scope/*.md` and surface any prior rejection that resembles this request.
73
+
74
+ 2. **Recommend.** Tell the maintainer your category and state recommendation with reasoning, plus a brief codebase summary relevant to the request — including whether it appears already implemented. Wait for direction.
75
+
76
+ 3. **Verify the claim.** Before any grilling, check that the claim holds up. For a bug, reproduce it from the reporter's steps. For a PR, confirm the diff does what it claims — check it out, run the relevant tests or commands. Report what happened: confirmed, failed, or insufficient detail. A confirmed verification makes a much stronger brief.
77
+
78
+ 4. **Grill (if needed).** If the request needs fleshing out, run `$grilling`. It stress-tests the decisions against code and domain docs and records resolved terminology or ADRs inline. Use `$domain-modeling` separately when the task is primarily to establish or repair the domain model.
79
+
80
+ 5. **Apply the outcome:**
81
+ - `ready-for-agent` — post an agent brief comment ([AGENT-BRIEF.md](AGENT-BRIEF.md)); include the repo architecture check for code changes.
82
+ - `ready-for-human` — same structure as an agent brief, but note why it can't be delegated (judgment calls, external access, design decisions, manual testing).
83
+ - `needs-info` — post triage notes (template below).
84
+ - `wontfix` — close, with the comment depending on why:
85
+ - Already implemented — point to where it lives; do not write to `.out-of-scope/`.
86
+ - Rejected bug — polite explanation, then close.
87
+ - Rejected enhancement — write to `.out-of-scope/`, link to it from a comment, then close ([OUT-OF-SCOPE.md](OUT-OF-SCOPE.md)).
88
+ - `needs-triage` — apply the role. Optional comment if there's partial progress.
89
+
90
+ ## Quick state override
91
+
92
+ If the maintainer says "move #42 to ready-for-agent", trust them and apply the role directly. Confirm what you're about to do (role changes, comment, close), then act. Skip grilling. If moving to `ready-for-agent` without a grilling session, ask whether they want to write an agent brief.
93
+
94
+ ## Proportional ready-for-agent policy
95
+
96
+ When preparing issues for agents, classify the size first, then choose the handoff shape.
97
+
98
+ Issue size and review risk are independent. A narrow high-risk issue may keep a
99
+ compact brief with stronger proof fields; a broad issue needs separate slices
100
+ only when ownership, release, blockers, or observable behavior differ.
101
+
102
+ - If several issues describe one deterministic bugfix in the same ownership area, recommend merging them or assigning them as one batch before marking all of them `ready-for-agent`.
103
+ - Agent briefs for narrow fixes should encourage one coherent PR when the code, tests, and verification belong together.
104
+ - For medium/large work, do not compress unrelated workflows into one agent brief. Keep slices separate when they have separate ownership, release sequencing, acceptance criteria, or validation paths.
105
+ - Do not add or imply an implementation-spec requirement unless the issue already has an unresolved `Spec gate` or the acceptance criteria are not deterministic enough for an AFK agent.
106
+ - Live mutation, production repair, or human approval can stay as a separate `ready-for-human` or blocked issue, but only when it truly cannot be verified safely inside the implementation issue.
107
+ - Do not add feature flags, telemetry systems, dashboards, rollout machinery, compatibility paths, or generic fallbacks to make an issue look safer; preserve approved scope.
108
+
109
+ ## Spec-gate policy
110
+
111
+ Triage does not create implementation specs. When an issue has a `Spec gate`,
112
+ verification contract, or material `Risk / Review` section, preserve its execution metadata in the agent brief and do not expand it into a file-by-file implementation plan. Omit conditional execution fields when they are not material.
113
+
114
+ ## Needs-info template
115
+
116
+ ```markdown
117
+ ## Triage Notes
118
+
119
+ **What we've established so far:**
120
+
121
+ - point 1
122
+ - point 2
123
+
124
+ **What we still need from you (@reporter):**
125
+
126
+ - question 1
127
+ - question 2
128
+ ```
129
+
130
+ Capture everything resolved during grilling under "established so far" so the work isn't lost. Questions must be specific and actionable, not "please provide more info".
131
+
132
+ ## Resuming a previous session
133
+
134
+ If prior triage notes exist on the issue or PR, read them, check whether the reporter has answered any outstanding questions, and present an updated picture before continuing. Don't re-ask resolved questions.
@@ -0,0 +1,6 @@
1
+ interface:
2
+ display_name: "Triage"
3
+ short_description: "Triage raw incoming issues and PRs"
4
+ default_prompt: "Use $triage to verify, classify, and prepare raw incoming issues or configured external PRs; skip generated planning tickets."
5
+ policy:
6
+ allow_implicit_invocation: false