@frankzhang2026/opencode-android-orchestrator 0.7.0 → 0.8.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (61) hide show
  1. package/CHANGELOG.md +32 -1
  2. package/README.md +58 -22
  3. package/THIRD_PARTY_NOTICES.md +15 -11
  4. package/dist/index.d.ts +3 -3
  5. package/dist/index.d.ts.map +1 -1
  6. package/dist/index.js +3 -3
  7. package/dist/index.js.map +1 -1
  8. package/dist/installer/adaptive-templates.d.ts +2 -5
  9. package/dist/installer/adaptive-templates.d.ts.map +1 -1
  10. package/dist/installer/adaptive-templates.js.map +1 -1
  11. package/dist/installer/agents-config.d.ts +5 -0
  12. package/dist/installer/agents-config.d.ts.map +1 -1
  13. package/dist/installer/agents-config.js +44 -8
  14. package/dist/installer/agents-config.js.map +1 -1
  15. package/dist/installer/index.d.ts +2 -2
  16. package/dist/installer/index.d.ts.map +1 -1
  17. package/dist/installer/index.js +2 -2
  18. package/dist/installer/index.js.map +1 -1
  19. package/dist/installer/opencode-config.d.ts +3 -4
  20. package/dist/installer/opencode-config.d.ts.map +1 -1
  21. package/dist/installer/opencode-config.js +1 -3
  22. package/dist/installer/opencode-config.js.map +1 -1
  23. package/dist/installer/upgrade.d.ts.map +1 -1
  24. package/dist/installer/upgrade.js +117 -46
  25. package/dist/installer/upgrade.js.map +1 -1
  26. package/dist/plugin/index.d.ts +2 -0
  27. package/dist/plugin/index.d.ts.map +1 -1
  28. package/dist/plugin/index.js +32 -0
  29. package/dist/plugin/index.js.map +1 -1
  30. package/docs/MIGRATION.md +45 -26
  31. package/docs/SECURITY.md +32 -14
  32. package/docs/TROUBLESHOOTING.md +31 -9
  33. package/package.json +3 -1
  34. package/resources/third-party/superpowers-v6.2.0/LICENSE +21 -0
  35. package/resources/third-party/superpowers-v6.2.0/PROVENANCE.md +21 -0
  36. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-brainstorming/SKILL.md +43 -0
  37. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-systematic-debugging/SKILL.md +288 -0
  38. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-systematic-debugging/references/condition-based-waiting-example.ts +158 -0
  39. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-systematic-debugging/references/condition-based-waiting.md +115 -0
  40. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-systematic-debugging/references/defense-in-depth.md +122 -0
  41. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-systematic-debugging/references/root-cause-tracing.md +169 -0
  42. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-systematic-debugging/scripts/find-polluter.sh +72 -0
  43. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-test-driven-development/SKILL.md +327 -0
  44. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-test-driven-development/references/writing-good-tests.md +197 -0
  45. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-verification-before-completion/SKILL.md +125 -0
  46. package/resources/third-party/superpowers-v6.2.0/skills/android-orchestrator-writing-plans/SKILL.md +50 -0
  47. package/templates/.opencode/agents/scheduled-coder.md +3 -4
  48. package/templates/.opencode/agents/scheduled-planner.md +5 -5
  49. package/templates/.opencode/agents/scheduled-reviewer.md +1 -2
  50. package/templates/.opencode/skills/scheduled-quality-coder/SKILL.md +7 -6
  51. package/templates/.opencode/skills/scheduled-quality-orchestrator/SKILL.md +2 -1
  52. package/templates/.opencode/skills/scheduled-quality-reviewer/SKILL.md +1 -1
  53. package/templates/README.md +6 -2
  54. package/templates/automation/config.json +6 -9
  55. package/templates/automation/config.schema.json +2 -11
  56. package/templates/automation/task-contract.schema.json +7 -7
  57. package/templates/automation/tasks/TASK-TEMPLATE.json.example +5 -5
  58. package/templates/scripts/automation/lib.sh +1 -3
  59. package/templates/scripts/automation/preflight.sh +0 -3
  60. package/templates/scripts/automation/tests/run-tests.sh +11 -14
  61. package/templates/scripts/automation/validate-contract.sh +7 -2
@@ -0,0 +1,327 @@
1
+ ---
2
+ name: android-orchestrator-test-driven-development
3
+ description: Apply strict RED-GREEN-REFACTOR while implementing an approved Android Orchestrator task
4
+ license: MIT
5
+ compatibility: opencode
6
+ metadata:
7
+ upstream: obra/superpowers@v6.2.0
8
+ workflow: scheduled-coding
9
+ ---
10
+
11
+ # Test-Driven Development (TDD)
12
+
13
+ ## Overview
14
+
15
+ Write the test first. Watch it fail. Write minimal code to pass.
16
+
17
+ **Core principle:** If you didn't watch the test fail, you don't know if it tests the right thing.
18
+
19
+ **Violating the letter of the rules is violating the spirit of the rules.**
20
+
21
+ ## When to Use
22
+
23
+ **Always:**
24
+ - New features
25
+ - Bug fixes
26
+ - Refactoring
27
+ - Behavior changes
28
+
29
+ **Exceptions (ask your human partner):**
30
+ - Throwaway prototypes
31
+ - Generated code
32
+ - Configuration files
33
+
34
+ Thinking "skip TDD just this once"? Stop. That's rationalization.
35
+
36
+ ## The Iron Law
37
+
38
+ ```
39
+ NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST
40
+ ```
41
+
42
+ Write code before the test? Delete it. Start over.
43
+
44
+ **No exceptions:**
45
+ - Don't keep it as "reference"
46
+ - Don't "adapt" it while writing tests
47
+ - Don't look at it
48
+ - Delete means delete
49
+
50
+ Implement fresh from tests. Period.
51
+
52
+ ## Red-Green-Refactor
53
+
54
+ ```dot
55
+ digraph tdd_cycle {
56
+ rankdir=LR;
57
+ red [label="RED\nWrite failing test", shape=box, style=filled, fillcolor="#ffcccc"];
58
+ verify_red [label="Verify fails\ncorrectly", shape=diamond];
59
+ green [label="GREEN\nMinimal code", shape=box, style=filled, fillcolor="#ccffcc"];
60
+ verify_green [label="Verify passes\nAll green", shape=diamond];
61
+ refactor [label="REFACTOR\nClean up", shape=box, style=filled, fillcolor="#ccccff"];
62
+ next [label="Next", shape=ellipse];
63
+
64
+ red -> verify_red;
65
+ verify_red -> green [label="yes"];
66
+ verify_red -> red [label="wrong\nfailure"];
67
+ green -> verify_green;
68
+ verify_green -> refactor [label="yes"];
69
+ verify_green -> green [label="no"];
70
+ refactor -> verify_green [label="stay\ngreen"];
71
+ verify_green -> next;
72
+ next -> red;
73
+ }
74
+ ```
75
+
76
+ ### RED - Write Failing Test
77
+
78
+ Write one minimal test showing what should happen.
79
+
80
+ <Good>
81
+ ```typescript
82
+ test('retries failed operations 3 times', async () => {
83
+ let attempts = 0;
84
+ const operation = () => {
85
+ attempts++;
86
+ if (attempts < 3) throw new Error('fail');
87
+ return 'success';
88
+ };
89
+
90
+ const result = await retryOperation(operation);
91
+
92
+ expect(result).toBe('success');
93
+ expect(attempts).toBe(3);
94
+ });
95
+ ```
96
+ Clear name, tests real behavior, one thing
97
+ </Good>
98
+
99
+ <Bad>
100
+ ```typescript
101
+ test('retry works', async () => {
102
+ const mock = jest.fn()
103
+ .mockRejectedValueOnce(new Error())
104
+ .mockRejectedValueOnce(new Error())
105
+ .mockResolvedValueOnce('success');
106
+ await retryOperation(mock);
107
+ expect(mock).toHaveBeenCalledTimes(3);
108
+ });
109
+ ```
110
+ Vague name, tests mock not code
111
+ </Bad>
112
+
113
+ **Requirements:**
114
+ - One behavior
115
+ - Clear name
116
+ - Real code (no mocks unless unavoidable)
117
+
118
+ ### Verify RED - Watch It Fail
119
+
120
+ **MANDATORY. Never skip.**
121
+
122
+ ```bash
123
+ npm test path/to/test.test.ts
124
+ ```
125
+
126
+ Confirm:
127
+ - Test fails (not errors)
128
+ - Failure message is expected
129
+ - Fails because feature missing (not typos)
130
+
131
+ **Test passes?** You're testing existing behavior. Fix test.
132
+
133
+ **Test errors?** Fix error, re-run until it fails correctly.
134
+
135
+ ### GREEN - Minimal Code
136
+
137
+ Write simplest code to pass the test.
138
+
139
+ <Good>
140
+ ```typescript
141
+ async function retryOperation<T>(fn: () => Promise<T>): Promise<T> {
142
+ for (let i = 0; i < 3; i++) {
143
+ try {
144
+ return await fn();
145
+ } catch (e) {
146
+ if (i === 2) throw e;
147
+ }
148
+ }
149
+ throw new Error('unreachable');
150
+ }
151
+ ```
152
+ Just enough to pass
153
+ </Good>
154
+
155
+ <Bad>
156
+ ```typescript
157
+ async function retryOperation<T>(
158
+ fn: () => Promise<T>,
159
+ options?: {
160
+ maxRetries?: number;
161
+ backoff?: 'linear' | 'exponential';
162
+ onRetry?: (attempt: number) => void;
163
+ }
164
+ ): Promise<T> {
165
+ // YAGNI
166
+ }
167
+ ```
168
+ Over-engineered
169
+ </Bad>
170
+
171
+ Don't add features, refactor other code, or "improve" beyond the test.
172
+
173
+ ### Verify GREEN - Watch It Pass
174
+
175
+ **MANDATORY.**
176
+
177
+ ```bash
178
+ npm test path/to/test.test.ts
179
+ ```
180
+
181
+ Confirm:
182
+ - Test passes
183
+ - Other tests still pass
184
+ - Output pristine (no errors, warnings)
185
+
186
+ **Test fails?** Fix code, not test.
187
+
188
+ **Other tests fail?** Fix now.
189
+
190
+ ### REFACTOR - Clean Up
191
+
192
+ After green only:
193
+ - Remove duplication
194
+ - Improve names
195
+ - Extract helpers
196
+
197
+ Keep tests green. Don't add behavior.
198
+
199
+ ### Repeat
200
+
201
+ Next failing test for next feature.
202
+
203
+ ## Good Tests
204
+
205
+ | Quality | Good | Bad |
206
+ |---------|------|-----|
207
+ | **Minimal** | One thing. "and" in name? Split it. | `test('validates email and domain and whitespace')` |
208
+ | **Clear** | Name describes behavior | `test('test1')` |
209
+ | **Shows intent** | Demonstrates desired API | Obscures what code should do |
210
+
211
+ When writing or changing any test, read
212
+ [references/writing-good-tests.md](references/writing-good-tests.md) for the
213
+ rules that keep tests honest:
214
+ - Name the production change that would make the test fail — before writing it
215
+ - Assert on real behavior, never on mock behavior
216
+ - Keep test-only code in test utilities, out of production classes
217
+ - Understand a dependency's side effects before mocking it
218
+
219
+ ## Common Rationalizations
220
+
221
+ | Excuse | Reality |
222
+ |--------|---------|
223
+ | "Too simple to test" | Simple code breaks. Test takes 30 seconds. |
224
+ | "I'll test after" | Tests written after pass immediately — which proves nothing. They may test the wrong thing, test the implementation instead of the behavior, or miss the edge case you forgot. You never watched it fail, so you never proved it can catch the bug. Test-first forces that failure. |
225
+ | "Tests after achieve same goals (spirit not ritual)" | Tests-after answer "what does this do?"; tests-first answer "what should this do?" Tests written after are biased by the code you already wrote — you verify the cases you remembered, not the ones you'd have discovered. Coverage without proof the tests work. |
226
+ | "Already manually tested" | Manual testing is ad-hoc: no record of what you covered, no way to re-run it when the code changes, easy to forget cases under pressure. "Worked when I tried it" ≠ comprehensive. Automated tests run the same way every time. |
227
+ | "Deleting X hours is wasteful" | Sunk cost fallacy — that time is already spent either way. The real choice: rewrite with TDD (high confidence) vs. keep it and bolt tests on after (low confidence, likely bugs). Keeping code you can't trust is the waste. |
228
+ | "Keep as reference, write tests first" | You'll adapt it. That's testing after. Delete means delete. |
229
+ | "Need to explore first" | Fine. Throw away exploration, start with TDD. |
230
+ | "Test hard = design unclear" | Listen to test. Hard to test = hard to use. |
231
+ | "TDD will slow me down" | TDD IS the pragmatic path: catches bugs before commit, prevents regressions, lets you refactor without fear. "Pragmatic" shortcuts mean debugging in production — slower, not faster. |
232
+ | "Manual test faster" | Manual doesn't prove edge cases. You'll re-test every change. |
233
+ | "Existing code has no tests" | You're improving it. Add tests for existing code. |
234
+
235
+ ## Red Flags - STOP and Start Over
236
+
237
+ - Code before test
238
+ - Test after implementation
239
+ - Test passes immediately
240
+ - Can't explain why test failed
241
+ - Tests added "later"
242
+ - Rationalizing "just this once"
243
+ - "I already manually tested it"
244
+ - "Tests after achieve the same purpose"
245
+ - "It's about spirit not ritual"
246
+ - "Keep as reference" or "adapt existing code"
247
+ - "Already spent X hours, deleting is wasteful"
248
+ - "TDD is dogmatic, I'm being pragmatic"
249
+ - "This is different because..."
250
+
251
+ **All of these mean: Delete code. Start over with TDD.**
252
+
253
+ ## Example: Bug Fix
254
+
255
+ **Bug:** Empty email accepted
256
+
257
+ **RED**
258
+ ```typescript
259
+ test('rejects empty email', async () => {
260
+ const result = await submitForm({ email: '' });
261
+ expect(result.error).toBe('Email required');
262
+ });
263
+ ```
264
+
265
+ **Verify RED**
266
+ ```bash
267
+ $ npm test
268
+ FAIL: expected 'Email required', got undefined
269
+ ```
270
+
271
+ **GREEN**
272
+ ```typescript
273
+ function submitForm(data: FormData) {
274
+ if (!data.email?.trim()) {
275
+ return { error: 'Email required' };
276
+ }
277
+ // ...
278
+ }
279
+ ```
280
+
281
+ **Verify GREEN**
282
+ ```bash
283
+ $ npm test
284
+ PASS
285
+ ```
286
+
287
+ **REFACTOR**
288
+ Extract validation for multiple fields if needed.
289
+
290
+ ## Verification Checklist
291
+
292
+ Before marking work complete:
293
+
294
+ - [ ] Every new function/method has a test
295
+ - [ ] Watched each test fail before implementing
296
+ - [ ] Each test failed for expected reason (feature missing, not typo)
297
+ - [ ] Wrote minimal code to pass each test
298
+ - [ ] All tests pass
299
+ - [ ] Output pristine (no errors, warnings)
300
+ - [ ] Tests use real code (mocks only if unavoidable)
301
+ - [ ] Edge cases and errors covered
302
+
303
+ Can't check all boxes? You skipped TDD. Start over.
304
+
305
+ ## When Stuck
306
+
307
+ | Problem | Solution |
308
+ |---------|----------|
309
+ | Don't know how to test | Write wished-for API. Write assertion first. Ask your human partner. |
310
+ | Test too complicated | Design too complicated. Simplify interface. |
311
+ | Must mock everything | Code too coupled. Use dependency injection. |
312
+ | Test setup huge | Extract helpers. Still complex? Simplify design. |
313
+
314
+ ## Debugging Integration
315
+
316
+ Bug found? Write failing test reproducing it. Follow TDD cycle. Test proves fix and prevents regression.
317
+
318
+ Never fix bugs without a test.
319
+
320
+ ## Final Rule
321
+
322
+ ```
323
+ Production code → test exists and failed first
324
+ Otherwise → not TDD
325
+ ```
326
+
327
+ No exceptions without your human partner's permission.
@@ -0,0 +1,197 @@
1
+ # Writing Good Tests
2
+
3
+ **Load this reference when:** writing or changing tests, adding mocks, or
4
+ adding cleanup/helper methods for tests.
5
+
6
+ ## Overview
7
+
8
+ A test exists to catch a specific break. Two principles govern everything
9
+ here:
10
+
11
+ ```
12
+ 1. Every test names the break it catches
13
+ 2. Every test exercises the real thing
14
+ ```
15
+
16
+ Strict TDD produces both naturally: a test written first and watched
17
+ failing against real code has already proven it can fail, and only earns
18
+ a mock when the real dependency proves slow or external.
19
+
20
+ ## Principle 1: Name the Break
21
+
22
+ Before writing the test body, answer: **what production change should
23
+ make this test fail — and is that change a bug or a decision?** A test
24
+ earns its place by catching a wrong branch, missing side effect, wrong
25
+ argument, boundary case, or broken contract.
26
+
27
+ **Derive expectations independently.** Use literals and hand-checked
28
+ fixtures; table-driven tests with literal `want` values are the preferred
29
+ shape. An expectation computed by the code under test — or its helpers —
30
+ passes no matter what that code does:
31
+
32
+ ```typescript
33
+ // ❌ Mirror assertion: the same builder computes both sides — always true
34
+ const expected = buildSearchQuery({ tag: 'urgent' });
35
+ expect(buildSearchQuery({ tag: 'urgent' })).toBe(expected);
36
+
37
+ // ✅ Hand-derived literal
38
+ expect(buildSearchQuery({ tag: 'urgent' })).toBe('tag:"urgent"');
39
+ ```
40
+
41
+ **No change detectors.** If only intentional decisions can fail a test —
42
+ a constant's value, exact message wording, private structure — it fires
43
+ on redesign and sleeps through bugs. Test the behavior that depends on
44
+ the decision: not `expect(MAX_RETRIES).toBe(5)` but "a failing call is
45
+ retried 5 times and the 6th attempt never happens."
46
+
47
+ **Behavior, not text.** Asserting that a script, skill, or config
48
+ contains an exact line proves only that the source is the source. Run
49
+ scripts against controlled inputs and assert outputs, side effects, or
50
+ exit codes. Documents that instruct agents are tested by the consuming
51
+ agent's behavior; prose for humans earns no test at all.
52
+
53
+ **Your code, not the framework.** Test the contract your code makes at
54
+ its boundaries — the route you register, the query you emit, the payload
55
+ you produce. Upstream mechanics are their maintainers' tests to write
56
+ (the classic: asserting your router invokes a registered handler — that
57
+ is the framework's test, not yours). When upstream behavior genuinely
58
+ surprised you, write one narrow characterization test naming the
59
+ assumption. The same boundary applies inside your code: constructors,
60
+ getters, constants, and trivial forwarding earn tests only when they
61
+ validate, normalize, default, derive, enforce, or cause side effects —
62
+ otherwise assert the first consumer-visible result that depends on them.
63
+
64
+ ### Gate Function
65
+
66
+ ```
67
+ BEFORE writing the test body:
68
+ Name the production change that would make this test fail.
69
+
70
+ Cannot name one → redesign around an observable behavior
71
+ "The source text changed" → run the artifact and assert its effects
72
+ Only intentional decisions → change detector; test the behavior
73
+ that depends on the decision
74
+
75
+ Confirm the expected value is derived without the code under test.
76
+ IF it reuses the code's logic or helpers:
77
+ Replace it with a literal or hand-checked fixture
78
+ ```
79
+
80
+ ## Principle 2: Exercise the Real Thing
81
+
82
+ **The mock earns no assertions.** A mock assertion passes when the mock
83
+ is present and fails when it is absent — it says nothing about the
84
+ component. Assert the real component's behavior; if the mock is what you
85
+ are checking, unmock it or delete the assertion.
86
+
87
+ ```typescript
88
+ // ✅ Real behavior
89
+ expect(screen.getByRole('navigation')).toBeInTheDocument();
90
+
91
+ // ❌ Mock existence
92
+ expect(screen.getByTestId('sidebar-mock')).toBeInTheDocument();
93
+ ```
94
+
95
+ **your human partner's correction:** "Are we testing the behavior of a
96
+ mock?"
97
+
98
+ **Mock at the right level.** Learn every side effect of the real method
99
+ before replacing it; mock the slow or external operation and keep what
100
+ the test depends on real. When unsure, run the test against the real
101
+ implementation first and observe what actually needs to happen.
102
+
103
+ ```typescript
104
+ // ❌ The mock swallows the config write that duplicate detection reads
105
+ vi.mock('ToolCatalog', () => ({
106
+ discoverAndCacheTools: vi.fn().mockResolvedValue(undefined)
107
+ }));
108
+
109
+ // ✅ Mock only the slow server startup; the config write stays real
110
+ vi.mock('MCPServerManager');
111
+ ```
112
+
113
+ **Make doubles specific.** When arguments, call counts, or ordering are
114
+ part of the contract, assert them — a fake that accepts anything verifies
115
+ nothing. Give each branch (success, error, malformed) its own fixture or
116
+ spy, so the wrong branch cannot satisfy the expectation.
117
+
118
+ **Mirror real data completely.** Mock the complete structure as it exists
119
+ in reality — all documented fields — not just the ones your test reads.
120
+ Partial mocks fail silently when downstream code reads an omitted field:
121
+ the test passes while integration breaks.
122
+
123
+ **Production classes carry production methods only.** Cleanup that only
124
+ tests need lives in test utilities, never as a `destroy()` on the
125
+ production class. Ask: is this method called only from tests? Does this
126
+ class own this resource's lifecycle? Wrong answers → test utility.
127
+
128
+ **Prefer real components over complex mocks.** When mock setup outgrows
129
+ the test logic, mocks miss methods the real components have, or tests
130
+ break when the mock changes, switch to an integration test with real
131
+ components. **your human partner's question:** "Do we need to be using a
132
+ mock here?"
133
+
134
+ ### Gate Function
135
+
136
+ ```
137
+ BEFORE adding a mock or test helper:
138
+ List the real method's side effects; keep the ones the test
139
+ depends on real — mock the slow/external level below them.
140
+
141
+ Mock responses mirror the complete real structure.
142
+
143
+ A method only tests call lives in test utilities, not production.
144
+
145
+ About to assert on the mock itself?
146
+ Unmock it or delete the assertion.
147
+ ```
148
+
149
+ ## Tests Ship With the Implementation
150
+
151
+ The TDD cycle — failing test, minimal implementation, refactor — is what
152
+ "complete" means. Ship the tests the behavior needs and only those:
153
+ trivial code and human prose earn none, and a test written to satisfy
154
+ process costs maintenance forever.
155
+
156
+ ## The Mutation Check
157
+
158
+ Before finishing, mentally mutate the production code; at least one test
159
+ should fail for each realistic mutation:
160
+
161
+ - Wrong constant or argument
162
+ - Wrong branch handler
163
+ - Missing state change or side effect
164
+ - Empty or default return
165
+ - Missing validation for zero, empty, nil, unauthorized, or malformed input
166
+
167
+ A mutation nothing catches marks the behavior as unprotected — or the
168
+ test as tautological.
169
+
170
+ ## Quick Reference
171
+
172
+ | When you... | Do |
173
+ |-------------|-----|
174
+ | Write any test | Name the break it catches — a bug, not a decision |
175
+ | Build an expected value | Derive it by hand; never with the code under test |
176
+ | Test a script or document | Run it / pressure-test its consumer; never grep its text |
177
+ | Reach for a dependency test | Test your boundary contract, not their documented mechanics |
178
+ | Want to assert on a mocked element | Test the real component, or unmock it |
179
+ | Are about to mock a method | Learn its side effects; mock the slow/external level |
180
+ | Build a mock response | Mirror the real structure completely |
181
+ | Need cleanup only tests use | Put it in test utilities |
182
+ | Watch mock setup balloon | Switch to an integration test with real components |
183
+ | Finish a test file | Run the mutation check |
184
+
185
+ ## Warning Signs
186
+
187
+ - Setup and assertion share the same object, guaranteeing equality
188
+ - The test can fail only through a panic, crash, or missing selector
189
+ - The test fails on every intentional change, never on accidental breakage
190
+ - Expected values are hidden behind loops, builders, or helpers
191
+ - The test greps source text, or asserts a removed symbol stays removed
192
+ - The test would still matter if only the framework remained
193
+ - The test exists for coverage, checking no side effect or outcome
194
+ - An assertion checks a `*-mock` test ID, or fails if you remove the mock
195
+ - A method is called only from test files
196
+ - Mock setup is more than half the test, or you can't explain why the mock is needed
197
+ - Mocking "just to be safe"
@@ -0,0 +1,125 @@
1
+ ---
2
+ name: android-orchestrator-verification-before-completion
3
+ description: Require fresh command evidence before an Android Orchestrator Coder or Reviewer reports completion or approval
4
+ license: MIT
5
+ compatibility: opencode
6
+ metadata:
7
+ upstream: obra/superpowers@v6.2.0
8
+ workflow: scheduled-coding
9
+ ---
10
+
11
+ # Verification Before Completion
12
+
13
+ ## Overview
14
+
15
+ **Core principle:** Evidence before claims, always.
16
+
17
+ **Violating the letter of this rule is violating the spirit of this rule.**
18
+
19
+ ## The Iron Law
20
+
21
+ ```
22
+ NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE
23
+ ```
24
+
25
+ If you haven't run the verification command in this message, you cannot claim it passes.
26
+
27
+ ## The Gate Function
28
+
29
+ ```
30
+ BEFORE claiming any status or expressing satisfaction:
31
+
32
+ 1. IDENTIFY: What command proves this claim?
33
+ 2. RUN: Execute the FULL command (fresh, complete)
34
+ 3. READ: Full output, check exit code, count failures
35
+ 4. VERIFY: Does output confirm the claim?
36
+ - If NO: State actual status with evidence
37
+ - If YES: State claim WITH evidence
38
+ 5. ONLY THEN: Make the claim
39
+
40
+ Skip any step = lying, not verifying
41
+ ```
42
+
43
+ ## Common Failures
44
+
45
+ | Claim | Requires | Not Sufficient |
46
+ |-------|----------|----------------|
47
+ | Tests pass | Test command output: 0 failures | Previous run, "should pass" |
48
+ | Linter clean | Linter output: 0 errors | Partial check, extrapolation |
49
+ | Build succeeds | Build command: exit 0 | Linter passing, logs look good |
50
+ | Bug fixed | Test original symptom: passes | Code changed, assumed fixed |
51
+ | Regression test works | Red-green cycle verified | Test passes once |
52
+ | Agent completed | VCS diff shows changes | Agent reports "success" |
53
+ | Requirements met | Line-by-line checklist | Tests passing |
54
+
55
+ ## Red Flags - STOP
56
+
57
+ - Using "should", "probably", "seems to"
58
+ - Expressing satisfaction before verification ("Great!", "Perfect!", "Done!", etc.)
59
+ - About to commit/push/PR without verification
60
+ - Trusting agent success reports
61
+ - Relying on partial verification
62
+ - Thinking "just this once"
63
+ - Tired and wanting work over
64
+ - **ANY wording implying success without having run verification**
65
+
66
+ ## Rationalization Prevention
67
+
68
+ | Excuse | Reality |
69
+ |--------|---------|
70
+ | "Should work now" | RUN the verification |
71
+ | "I'm confident" | Confidence ≠ evidence |
72
+ | "Just this once" | No exceptions |
73
+ | "Linter passed" | Linter ≠ compiler |
74
+ | "Agent said success" | Verify independently |
75
+ | "I'm tired" | Exhaustion ≠ excuse |
76
+ | "Partial check is enough" | Partial proves nothing |
77
+ | "Different words so rule doesn't apply" | Spirit over letter |
78
+
79
+ ## Key Patterns
80
+
81
+ **Tests:**
82
+ ```
83
+ ✅ [Run test command] [See: 34/34 pass] "All tests pass"
84
+ ❌ "Should pass now" / "Looks correct"
85
+ ```
86
+
87
+ **Regression tests (TDD Red-Green):**
88
+ ```
89
+ ✅ Write → Run (pass) → Revert fix → Run (MUST FAIL) → Restore → Run (pass)
90
+ ❌ "I've written a regression test" (without red-green verification)
91
+ ```
92
+
93
+ **Build:**
94
+ ```
95
+ ✅ [Run build] [See: exit 0] "Build passes"
96
+ ❌ "Linter passed" (linter doesn't check compilation)
97
+ ```
98
+
99
+ **Requirements:**
100
+ ```
101
+ ✅ Re-read plan → Create checklist → Verify each → Report gaps or completion
102
+ ❌ "Tests pass, phase complete"
103
+ ```
104
+
105
+ **Agent delegation:**
106
+ ```
107
+ ✅ Agent reports success → Check VCS diff → Verify changes → Report actual state
108
+ ❌ Trust agent report
109
+ ```
110
+
111
+ ## When To Apply
112
+
113
+ **ALWAYS before:**
114
+ - ANY variation of success/completion claims
115
+ - ANY expression of satisfaction
116
+ - ANY positive statement about work state
117
+ - Committing, PR creation, task completion
118
+ - Moving to next task
119
+ - Delegating to agents
120
+
121
+ **Rule applies to:**
122
+ - Exact phrases
123
+ - Paraphrases and synonyms
124
+ - Implications of success
125
+ - ANY communication suggesting completion/correctness
@@ -0,0 +1,50 @@
1
+ ---
2
+ name: android-orchestrator-writing-plans
3
+ description: Convert an approved Android Orchestrator proposal into an implementation-ready plan that matches its sealed task contract
4
+ license: MIT
5
+ compatibility: opencode
6
+ metadata:
7
+ upstream: obra/superpowers@v6.2.0
8
+ workflow: scheduled-coding
9
+ ---
10
+
11
+ # Android Orchestrator writing plans
12
+
13
+ Write a plan detailed enough for the restricted Coder to implement without
14
+ guessing. The approved proposal is the scope ceiling.
15
+
16
+ ## Required output
17
+
18
+ Create only the plan path selected by `scheduled-quality-orchestrator`:
19
+ `docs/plans/<TASK-ID>.md`. The matching JSON contract is created separately at
20
+ `automation/tasks/<TASK-ID>.json`; both artifacts must describe the same task.
21
+
22
+ The plan must include:
23
+
24
+ - task ID and title;
25
+ - current and desired observable behavior;
26
+ - acceptance criteria and edge cases;
27
+ - exact files to create or modify;
28
+ - implementation sequence with concrete symbols and interfaces;
29
+ - the first failing behavior test and expected RED reason;
30
+ - focused and full verification commands;
31
+ - allowed paths, forbidden paths, maximum changed-file count, and non-goals;
32
+ - device/emulator policy and any residual risk.
33
+
34
+ ## Plan quality
35
+
36
+ - Use small ordered steps: test, verify RED, minimal implementation, verify
37
+ GREEN, then the configured quality gate.
38
+ - Derive expected values independently from production code.
39
+ - Do not use placeholders such as TODO, TBD, “add suitable handling”, or
40
+ “similar to the previous step”.
41
+ - Keep names, signatures, resources, and paths consistent across every step
42
+ and with the task contract.
43
+ - Do not introduce unapproved refactors or dependencies.
44
+
45
+ ## Handoff
46
+
47
+ Do not offer alternate execution modes, dispatch subagents, commit, or start
48
+ implementation. Return control to `scheduled-quality-orchestrator`, which
49
+ validates and seals the plan and contract before the separate Coder and
50
+ Reviewer sessions can run.