@polderlabs/bizar 10.23.21 → 10.23.23

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (159) hide show
  1. package/AGENTS.md +17 -14
  2. package/README.md +209 -67
  3. package/cli/banner.mjs +1 -1
  4. package/cli/commands/models.mjs +44 -3
  5. package/cli/commands/validate.mjs +6 -7
  6. package/cli/install/banner.mjs +1 -1
  7. package/cli/install/interactive-setup.mjs +13 -1
  8. package/cli/install/paths.mjs +2 -10
  9. package/cli/provision.mjs +21 -14
  10. package/cli/utils.mjs +2 -11
  11. package/config/claude/CLAUDE.md +17 -14
  12. package/config/claude/agents/bizar-accessibility-architect.md +153 -0
  13. package/config/claude/agents/bizar-agent-evaluator.md +210 -0
  14. package/config/claude/agents/bizar-architect.md +224 -0
  15. package/config/claude/agents/bizar-build-error-resolver.md +127 -0
  16. package/config/claude/agents/bizar-chief-of-staff.md +164 -0
  17. package/config/claude/agents/bizar-code-architect.md +84 -0
  18. package/config/claude/agents/bizar-code-explorer.md +82 -0
  19. package/config/claude/agents/bizar-code-reviewer.md +327 -0
  20. package/config/claude/agents/bizar-code-simplifier.md +60 -0
  21. package/config/claude/agents/bizar-comment-analyzer.md +58 -0
  22. package/config/claude/agents/bizar-conversation-analyzer.md +65 -0
  23. package/config/claude/agents/bizar-cpp-build-resolver.md +103 -0
  24. package/config/claude/agents/bizar-cpp-reviewer.md +85 -0
  25. package/config/claude/agents/bizar-csharp-reviewer.md +114 -0
  26. package/config/claude/agents/bizar-dart-build-resolver.md +214 -0
  27. package/config/claude/agents/bizar-database-reviewer.md +104 -0
  28. package/config/claude/agents/bizar-django-build-resolver.md +256 -0
  29. package/config/claude/agents/bizar-django-reviewer.md +173 -0
  30. package/config/claude/agents/bizar-doc-updater.md +120 -0
  31. package/config/claude/agents/bizar-docs-lookup.md +81 -0
  32. package/config/claude/agents/bizar-end-to-end-runner.md +120 -0
  33. package/config/claude/agents/bizar-fastapi-reviewer.md +83 -0
  34. package/config/claude/agents/bizar-flutter-reviewer.md +256 -0
  35. package/config/claude/agents/bizar-fsharp-reviewer.md +113 -0
  36. package/config/claude/agents/bizar-gan-evaluator.md +236 -0
  37. package/config/claude/agents/bizar-gan-generator.md +144 -0
  38. package/config/claude/agents/bizar-gan-planner.md +112 -0
  39. package/config/claude/agents/bizar-go-build-resolver.md +107 -0
  40. package/config/claude/agents/bizar-go-reviewer.md +89 -0
  41. package/config/claude/agents/bizar-harmonyos-app-resolver.md +186 -0
  42. package/config/claude/agents/bizar-harness-optimizer.md +59 -0
  43. package/config/claude/agents/bizar-healthcare-reviewer.md +96 -0
  44. package/config/claude/agents/bizar-homelab-architect.md +111 -0
  45. package/config/claude/agents/bizar-java-build-resolver.md +279 -0
  46. package/config/claude/agents/bizar-java-reviewer.md +194 -0
  47. package/config/claude/agents/bizar-kotlin-build-resolver.md +131 -0
  48. package/config/claude/agents/bizar-kotlin-reviewer.md +172 -0
  49. package/config/claude/agents/bizar-loop-operator.md +49 -0
  50. package/config/claude/agents/bizar-marketing-agent.md +163 -0
  51. package/config/claude/agents/bizar-mle-reviewer.md +166 -0
  52. package/config/claude/agents/bizar-network-architect.md +110 -0
  53. package/config/claude/agents/bizar-network-config-reviewer.md +110 -0
  54. package/config/claude/agents/bizar-network-troubleshooter.md +132 -0
  55. package/config/claude/agents/bizar-opensource-forker.md +211 -0
  56. package/config/claude/agents/bizar-opensource-packager.md +262 -0
  57. package/config/claude/agents/bizar-opensource-sanitizer.md +201 -0
  58. package/config/claude/agents/bizar-performance-optimizer.md +459 -0
  59. package/config/claude/agents/bizar-php-reviewer.md +113 -0
  60. package/config/claude/agents/bizar-planner.md +225 -0
  61. package/config/claude/agents/bizar-pr-test-analyzer.md +58 -0
  62. package/config/claude/agents/bizar-python-reviewer.md +111 -0
  63. package/config/claude/agents/bizar-pytorch-build-resolver.md +133 -0
  64. package/config/claude/agents/bizar-rag-pipeline-reviewer.md +71 -0
  65. package/config/claude/agents/bizar-react-build-resolver.md +219 -0
  66. package/config/claude/agents/bizar-react-reviewer.md +171 -0
  67. package/config/claude/agents/bizar-refactor-cleaner.md +98 -0
  68. package/config/claude/agents/bizar-rust-build-resolver.md +161 -0
  69. package/config/claude/agents/bizar-rust-reviewer.md +107 -0
  70. package/config/claude/agents/bizar-security-reviewer.md +121 -0
  71. package/config/claude/agents/bizar-seo-specialist.md +75 -0
  72. package/config/claude/agents/bizar-silent-failure-hunter.md +63 -0
  73. package/config/claude/agents/bizar-spec-miner.md +221 -0
  74. package/config/claude/agents/bizar-swift-build-resolver.md +174 -0
  75. package/config/claude/agents/bizar-swift-reviewer.md +120 -0
  76. package/config/claude/agents/bizar-tdd-guide.md +104 -0
  77. package/config/claude/agents/bizar-type-design-analyzer.md +54 -0
  78. package/config/claude/agents/bizar-typescript-reviewer.md +128 -0
  79. package/config/claude/agents/bizar-vue-reviewer.md +210 -0
  80. package/config/claude/agents/office-manager.md +13 -15
  81. package/config/claude/commands/bizar.md +6 -5
  82. package/config/claude/commands/plow-through.md +3 -2
  83. package/config/claude/commands/quick.md +14 -14
  84. package/config/claude/commands/team.md +5 -2
  85. package/config/claude/hooks/agent-model-guard.mjs +2 -2
  86. package/config/claude/hooks/sessionend-recall.mjs +1 -9
  87. package/config/claude/hooks/sessionstart-prime.mjs +2 -2
  88. package/config/claude/hooks/thinking-route.mjs +0 -1
  89. package/config/claude/hooks/worker-suggest.mjs +12 -24
  90. package/config/claude/hooks/workflow-route-guard.mjs +4 -3
  91. package/config/claude/hooks/workflow-route-state.mjs +1 -1
  92. package/config/skills/brainstorming/SKILL.md +253 -0
  93. package/config/skills/brainstorming/scripts/frame-template.html +213 -0
  94. package/config/skills/brainstorming/scripts/helper.js +167 -0
  95. package/config/skills/brainstorming/scripts/server.cjs +723 -0
  96. package/config/skills/brainstorming/scripts/start-server.sh +209 -0
  97. package/config/skills/brainstorming/scripts/stop-server.sh +120 -0
  98. package/config/skills/brainstorming/spec-document-reviewer-prompt.md +49 -0
  99. package/config/skills/brainstorming/visual-companion.md +299 -0
  100. package/config/skills/dispatching-parallel-agents/SKILL.md +170 -0
  101. package/config/skills/executing-plans/SKILL.md +67 -0
  102. package/config/skills/finishing-a-development-branch/SKILL.md +228 -0
  103. package/config/skills/receiving-code-review/SKILL.md +208 -0
  104. package/config/skills/requesting-code-review/SKILL.md +98 -0
  105. package/config/skills/requesting-code-review/code-reviewer.md +181 -0
  106. package/config/skills/subagent-driven-development/SKILL.md +571 -0
  107. package/config/skills/subagent-driven-development/implementer-prompt.md +154 -0
  108. package/config/skills/subagent-driven-development/re-review-prompt.md +115 -0
  109. package/config/skills/subagent-driven-development/scripts/review-package +46 -0
  110. package/config/skills/subagent-driven-development/scripts/sdd-workspace +40 -0
  111. package/config/skills/subagent-driven-development/scripts/task-brief +41 -0
  112. package/config/skills/subagent-driven-development/task-reviewer-prompt.md +207 -0
  113. package/config/skills/systematic-debugging/CREATION-LOG.md +119 -0
  114. package/config/skills/systematic-debugging/SKILL.md +286 -0
  115. package/config/skills/systematic-debugging/condition-based-waiting-example.ts +158 -0
  116. package/config/skills/systematic-debugging/condition-based-waiting.md +115 -0
  117. package/config/skills/systematic-debugging/defense-in-depth.md +122 -0
  118. package/config/skills/systematic-debugging/find-polluter.sh +72 -0
  119. package/config/skills/systematic-debugging/root-cause-tracing.md +169 -0
  120. package/config/skills/systematic-debugging/test-academic.md +14 -0
  121. package/config/skills/systematic-debugging/test-pressure-1.md +58 -0
  122. package/config/skills/systematic-debugging/test-pressure-2.md +68 -0
  123. package/config/skills/systematic-debugging/test-pressure-3.md +69 -0
  124. package/config/skills/test-driven-development/SKILL.md +323 -0
  125. package/config/skills/test-driven-development/writing-good-tests.md +198 -0
  126. package/config/skills/using-git-worktrees/SKILL.md +170 -0
  127. package/config/skills/using-superpowers/SKILL.md +66 -0
  128. package/config/skills/using-superpowers/references/antigravity-tools.md +23 -0
  129. package/config/skills/using-superpowers/references/codex-tools.md +108 -0
  130. package/config/skills/using-superpowers/references/gemini-tools.md +63 -0
  131. package/config/skills/using-superpowers/references/hermes-tools.md +56 -0
  132. package/config/skills/using-superpowers/references/pi-tools.md +16 -0
  133. package/config/skills/verification-before-completion/SKILL.md +123 -0
  134. package/config/skills/writing-plans/SKILL.md +174 -0
  135. package/config/skills/writing-plans/plan-document-reviewer-prompt.md +49 -0
  136. package/config/skills/writing-skills/SKILL.md +682 -0
  137. package/config/skills/writing-skills/anthropic-best-practices.md +1150 -0
  138. package/config/skills/writing-skills/examples/CLAUDE_MD_TESTING.md +189 -0
  139. package/config/skills/writing-skills/graphviz-conventions.dot +172 -0
  140. package/config/skills/writing-skills/persuasion-principles.md +187 -0
  141. package/config/skills/writing-skills/render-graphs.js +169 -0
  142. package/config/skills/writing-skills/testing-skills-with-subagents.md +384 -0
  143. package/config/trigger-patterns.json +1 -1
  144. package/config/workflows/bizar-debug.js +1 -0
  145. package/config/workflows/bizar-implement.js +1 -0
  146. package/config/workflows/bizar-research.js +1 -0
  147. package/config/workflows/ultracode-research.js +1 -0
  148. package/config/workflows/ultracode-review.js +1 -0
  149. package/config/workflows/ultracode.js +1 -0
  150. package/package.json +1 -1
  151. package/packages/sdk/dist/version.d.ts +1 -1
  152. package/packages/sdk/dist/version.js +1 -1
  153. package/packages/sdk/package.json +1 -1
  154. package/config/claude/commands/migrate.md +0 -18
  155. package/config/claude/commands/tailscale-serve.md +0 -14
  156. package/config/claude/commands/tier.md +0 -31
  157. package/config/claude/commands/upgrade-defaults.md +0 -34
  158. package/config/claude/commands/use-default.md +0 -12
  159. package/config/claude/commands/use-premium.md +0 -12
@@ -0,0 +1,323 @@
1
+ ---
2
+ name: test-driven-development
3
+ description: Use when implementing any feature or bugfix, before writing implementation code
4
+ ---
5
+
6
+ > **Bizar compatibility:** This procedure is fully integrated into Bizar. Bizar system, repository, model-routing, autonomy, and approval policy control whenever they differ. Do not add a separate mandatory approval gate, recurse into dispatch, or override the active Bizar workflow.
7
+
8
+
9
+ # Test-Driven Development (TDD)
10
+
11
+ ## Overview
12
+
13
+ Write the test first. Watch it fail. Write minimal code to pass.
14
+
15
+ **Core principle:** If you didn't watch the test fail, you don't know if it tests the right thing.
16
+
17
+ **Violating the letter of the rules is violating the spirit of the rules.**
18
+
19
+ ## When to Use
20
+
21
+ **Always:**
22
+ - New features
23
+ - Bug fixes
24
+ - Refactoring
25
+ - Behavior changes
26
+
27
+ **Exceptions (ask your human partner):**
28
+ - Throwaway prototypes
29
+ - Generated code
30
+ - Configuration files
31
+
32
+ Thinking "skip TDD just this once"? Stop. That's rationalization.
33
+
34
+ ## The Iron Law
35
+
36
+ ```
37
+ NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST
38
+ ```
39
+
40
+ Write code before the test? Delete it. Start over.
41
+
42
+ **No exceptions:**
43
+ - Don't keep it as "reference"
44
+ - Don't "adapt" it while writing tests
45
+ - Don't look at it
46
+ - Delete means delete
47
+
48
+ Implement fresh from tests. Period.
49
+
50
+ ## Red-Green-Refactor
51
+
52
+ ```dot
53
+ digraph tdd_cycle {
54
+ rankdir=LR;
55
+ red [label="RED\nWrite failing test", shape=box, style=filled, fillcolor="#ffcccc"];
56
+ verify_red [label="Verify fails\ncorrectly", shape=diamond];
57
+ green [label="GREEN\nMinimal code", shape=box, style=filled, fillcolor="#ccffcc"];
58
+ verify_green [label="Verify passes\nAll green", shape=diamond];
59
+ refactor [label="REFACTOR\nClean up", shape=box, style=filled, fillcolor="#ccccff"];
60
+ next [label="Next", shape=ellipse];
61
+
62
+ red -> verify_red;
63
+ verify_red -> green [label="yes"];
64
+ verify_red -> red [label="wrong\nfailure"];
65
+ green -> verify_green;
66
+ verify_green -> refactor [label="yes"];
67
+ verify_green -> green [label="no"];
68
+ refactor -> verify_green [label="stay\ngreen"];
69
+ verify_green -> next;
70
+ next -> red;
71
+ }
72
+ ```
73
+
74
+ ### RED - Write Failing Test
75
+
76
+ Write one minimal test showing what should happen.
77
+
78
+ <Good>
79
+ ```typescript
80
+ test('retries failed operations 3 times', async () => {
81
+ let attempts = 0;
82
+ const operation = () => {
83
+ attempts++;
84
+ if (attempts < 3) throw new Error('fail');
85
+ return 'success';
86
+ };
87
+
88
+ const result = await retryOperation(operation);
89
+
90
+ expect(result).toBe('success');
91
+ expect(attempts).toBe(3);
92
+ });
93
+ ```
94
+ Clear name, tests real behavior, one thing
95
+ </Good>
96
+
97
+ <Bad>
98
+ ```typescript
99
+ test('retry works', async () => {
100
+ const mock = jest.fn()
101
+ .mockRejectedValueOnce(new Error())
102
+ .mockRejectedValueOnce(new Error())
103
+ .mockResolvedValueOnce('success');
104
+ await retryOperation(mock);
105
+ expect(mock).toHaveBeenCalledTimes(3);
106
+ });
107
+ ```
108
+ Vague name, tests mock not code
109
+ </Bad>
110
+
111
+ **Requirements:**
112
+ - One behavior
113
+ - Clear name
114
+ - Real code (no mocks unless unavoidable)
115
+
116
+ ### Verify RED - Watch It Fail
117
+
118
+ **MANDATORY. Never skip.**
119
+
120
+ ```bash
121
+ npm test path/to/test.test.ts
122
+ ```
123
+
124
+ Confirm:
125
+ - Test fails (not errors)
126
+ - Failure message is expected
127
+ - Fails because feature missing (not typos)
128
+
129
+ **Test passes?** You're testing existing behavior. Fix test.
130
+
131
+ **Test errors?** Fix error, re-run until it fails correctly.
132
+
133
+ ### GREEN - Minimal Code
134
+
135
+ Write simplest code to pass the test.
136
+
137
+ <Good>
138
+ ```typescript
139
+ async function retryOperation<T>(fn: () => Promise<T>): Promise<T> {
140
+ for (let i = 0; i < 3; i++) {
141
+ try {
142
+ return await fn();
143
+ } catch (e) {
144
+ if (i === 2) throw e;
145
+ }
146
+ }
147
+ throw new Error('unreachable');
148
+ }
149
+ ```
150
+ Just enough to pass
151
+ </Good>
152
+
153
+ <Bad>
154
+ ```typescript
155
+ async function retryOperation<T>(
156
+ fn: () => Promise<T>,
157
+ options?: {
158
+ maxRetries?: number;
159
+ backoff?: 'linear' | 'exponential';
160
+ onRetry?: (attempt: number) => void;
161
+ }
162
+ ): Promise<T> {
163
+ // YAGNI
164
+ }
165
+ ```
166
+ Over-engineered
167
+ </Bad>
168
+
169
+ Don't add features, refactor other code, or "improve" beyond the test.
170
+
171
+ ### Verify GREEN - Watch It Pass
172
+
173
+ **MANDATORY.**
174
+
175
+ ```bash
176
+ npm test path/to/test.test.ts
177
+ ```
178
+
179
+ Confirm:
180
+ - Test passes
181
+ - Other tests still pass
182
+ - Output pristine (no errors, warnings)
183
+
184
+ **Test fails?** Fix code, not test.
185
+
186
+ **Other tests fail?** Fix now.
187
+
188
+ ### REFACTOR - Clean Up
189
+
190
+ After green only:
191
+ - Remove duplication
192
+ - Improve names
193
+ - Extract helpers
194
+
195
+ Keep tests green. Don't add behavior.
196
+
197
+ ### Repeat
198
+
199
+ Next failing test for next feature.
200
+
201
+ ## Good Tests
202
+
203
+ | Quality | Good | Bad |
204
+ |---------|------|-----|
205
+ | **Minimal** | One thing. "and" in name? Split it. | `test('validates email and domain and whitespace')` |
206
+ | **Clear** | Name describes behavior | `test('test1')` |
207
+ | **Shows intent** | Demonstrates desired API | Obscures what code should do |
208
+
209
+ When writing or changing any test, read [writing-good-tests.md](writing-good-tests.md) for the rules that keep tests honest:
210
+ - Name the production change that would make the test fail — before writing it
211
+ - Assert on real behavior, never on mock behavior
212
+ - Keep test-only code in test utilities, out of production classes
213
+ - Understand a dependency's side effects before mocking it
214
+
215
+ ## Common Rationalizations
216
+
217
+ | Excuse | Reality |
218
+ |--------|---------|
219
+ | "Too simple to test" | Simple code breaks. Test takes 30 seconds. |
220
+ | "I'll test after" | Tests written after pass immediately — which proves nothing. They may test the wrong thing, test the implementation instead of the behavior, or miss the edge case you forgot. You never watched it fail, so you never proved it can catch the bug. Test-first forces that failure. |
221
+ | "Tests after achieve same goals (spirit not ritual)" | Tests-after answer "what does this do?"; tests-first answer "what should this do?" Tests written after are biased by the code you already wrote — you verify the cases you remembered, not the ones you'd have discovered. Coverage without proof the tests work. |
222
+ | "Already manually tested" | Manual testing is ad-hoc: no record of what you covered, no way to re-run it when the code changes, easy to forget cases under pressure. "Worked when I tried it" ≠ comprehensive. Automated tests run the same way every time. |
223
+ | "Deleting X hours is wasteful" | Sunk cost fallacy — that time is already spent either way. The real choice: rewrite with TDD (high confidence) vs. keep it and bolt tests on after (low confidence, likely bugs). Keeping code you can't trust is the waste. |
224
+ | "Keep as reference, write tests first" | You'll adapt it. That's testing after. Delete means delete. |
225
+ | "Need to explore first" | Fine. Throw away exploration, start with TDD. |
226
+ | "Test hard = design unclear" | Listen to test. Hard to test = hard to use. |
227
+ | "TDD will slow me down" | TDD IS the pragmatic path: catches bugs before commit, prevents regressions, lets you refactor without fear. "Pragmatic" shortcuts mean debugging in production — slower, not faster. |
228
+ | "Manual test faster" | Manual doesn't prove edge cases. You'll re-test every change. |
229
+ | "Existing code has no tests" | You're improving it. Add tests for existing code. |
230
+
231
+ ## Red Flags - STOP and Start Over
232
+
233
+ - Code before test
234
+ - Test after implementation
235
+ - Test passes immediately
236
+ - Can't explain why test failed
237
+ - Tests added "later"
238
+ - Rationalizing "just this once"
239
+ - "I already manually tested it"
240
+ - "Tests after achieve the same purpose"
241
+ - "It's about spirit not ritual"
242
+ - "Keep as reference" or "adapt existing code"
243
+ - "Already spent X hours, deleting is wasteful"
244
+ - "TDD is dogmatic, I'm being pragmatic"
245
+ - "This is different because..."
246
+
247
+ **All of these mean: Delete code. Start over with TDD.**
248
+
249
+ ## Example: Bug Fix
250
+
251
+ **Bug:** Empty email accepted
252
+
253
+ **RED**
254
+ ```typescript
255
+ test('rejects empty email', async () => {
256
+ const result = await submitForm({ email: '' });
257
+ expect(result.error).toBe('Email required');
258
+ });
259
+ ```
260
+
261
+ **Verify RED**
262
+ ```bash
263
+ $ npm test
264
+ FAIL: expected 'Email required', got undefined
265
+ ```
266
+
267
+ **GREEN**
268
+ ```typescript
269
+ function submitForm(data: FormData) {
270
+ if (!data.email?.trim()) {
271
+ return { error: 'Email required' };
272
+ }
273
+ // ...
274
+ }
275
+ ```
276
+
277
+ **Verify GREEN**
278
+ ```bash
279
+ $ npm test
280
+ PASS
281
+ ```
282
+
283
+ **REFACTOR**
284
+ Extract validation for multiple fields if needed.
285
+
286
+ ## Verification Checklist
287
+
288
+ Before marking work complete:
289
+
290
+ - [ ] Every new function/method has a test
291
+ - [ ] Watched each test fail before implementing
292
+ - [ ] Each test failed for expected reason (feature missing, not typo)
293
+ - [ ] Wrote minimal code to pass each test
294
+ - [ ] All tests pass
295
+ - [ ] Output pristine (no errors, warnings)
296
+ - [ ] Tests use real code (mocks only if unavoidable)
297
+ - [ ] Edge cases and errors covered
298
+
299
+ Can't check all boxes? You skipped TDD. Start over.
300
+
301
+ ## When Stuck
302
+
303
+ | Problem | Solution |
304
+ |---------|----------|
305
+ | Don't know how to test | Write wished-for API. Write assertion first. Ask your human partner. |
306
+ | Test too complicated | Design too complicated. Simplify interface. |
307
+ | Must mock everything | Code too coupled. Use dependency injection. |
308
+ | Test setup huge | Extract helpers. Still complex? Simplify design. |
309
+
310
+ ## Debugging Integration
311
+
312
+ Bug found? Write failing test reproducing it. Follow TDD cycle. Test proves fix and prevents regression.
313
+
314
+ Never fix bugs without a test.
315
+
316
+ ## Final Rule
317
+
318
+ ```
319
+ Production code → test exists and failed first
320
+ Otherwise → not TDD
321
+ ```
322
+
323
+ No exceptions without your human partner's permission.
@@ -0,0 +1,198 @@
1
+ # Writing Good Tests
2
+
3
+ **Load this reference when:** writing or changing tests, adding mocks, or
4
+ adding cleanup/helper methods for tests.
5
+
6
+ ## Overview
7
+
8
+ A test exists to catch a specific break. Two principles govern everything
9
+ here:
10
+
11
+ ```
12
+ 1. Every test names the break it catches
13
+ 2. Every test exercises the real thing
14
+ ```
15
+
16
+ Strict TDD produces both naturally: a test written first and watched
17
+ failing against real code has already proven it can fail, and only earns
18
+ a mock when the real dependency proves slow or external.
19
+
20
+ ## Principle 1: Name the Break
21
+
22
+ Before writing the test body, answer: **what production change should
23
+ make this test fail — and is that change a bug or a decision?** A test
24
+ earns its place by catching a wrong branch, missing side effect, wrong
25
+ argument, boundary case, or broken contract.
26
+
27
+ **Derive expectations independently.** Use literals and hand-checked
28
+ fixtures; table-driven tests with literal `want` values are the preferred
29
+ shape. An expectation computed by the code under test — or its helpers —
30
+ passes no matter what that code does:
31
+
32
+ ```typescript
33
+ // ❌ Mirror assertion: the same builder computes both sides — always true
34
+ const expected = buildSearchQuery({ tag: 'urgent' });
35
+ expect(buildSearchQuery({ tag: 'urgent' })).toBe(expected);
36
+
37
+ // ✅ Hand-derived literal
38
+ expect(buildSearchQuery({ tag: 'urgent' })).toBe('tag:"urgent"');
39
+ ```
40
+
41
+ **No change detectors.** If only intentional decisions can fail a test —
42
+ a constant's value, exact message wording, private structure — it fires
43
+ on redesign and sleeps through bugs. Test the behavior that depends on
44
+ the decision: not `expect(MAX_RETRIES).toBe(5)` but "a failing call is
45
+ retried 5 times and the 6th attempt never happens."
46
+
47
+ **Behavior, not text.** Asserting that a script, skill, or config
48
+ contains an exact line proves only that the source is the source. Run
49
+ scripts against controlled inputs and assert outputs, side effects, or
50
+ exit codes. Documents that instruct agents are tested by the consuming
51
+ agent's behavior (superpowers:writing-skills); prose for humans earns no
52
+ test at all.
53
+
54
+ **Your code, not the framework.** Test the contract your code makes at
55
+ its boundaries — the route you register, the query you emit, the payload
56
+ you produce. Upstream mechanics are their maintainers' tests to write
57
+ (the classic: asserting your router invokes a registered handler — that
58
+ is the framework's test, not yours). When upstream behavior genuinely
59
+ surprised you, write one narrow characterization test naming the
60
+ assumption. The same boundary applies inside your code: constructors,
61
+ getters, constants, and trivial forwarding earn tests only when they
62
+ validate, normalize, default, derive, enforce, or cause side effects —
63
+ otherwise assert the first consumer-visible result that depends on them.
64
+
65
+ ### Gate Function
66
+
67
+ ```
68
+ BEFORE writing the test body:
69
+ Name the production change that would make this test fail.
70
+
71
+ Cannot name one → redesign around an observable behavior
72
+ "The source text changed" → run the artifact and assert its effects
73
+ Only intentional decisions → change detector; test the behavior
74
+ that depends on the decision
75
+
76
+ Confirm the expected value is derived without the code under test.
77
+ IF it reuses the code's logic or helpers:
78
+ Replace it with a literal or hand-checked fixture
79
+ ```
80
+
81
+ ## Principle 2: Exercise the Real Thing
82
+
83
+ **The mock earns no assertions.** A mock assertion passes when the mock
84
+ is present and fails when it is absent — it says nothing about the
85
+ component. Assert the real component's behavior; if the mock is what you
86
+ are checking, unmock it or delete the assertion.
87
+
88
+ ```typescript
89
+ // ✅ Real behavior
90
+ expect(screen.getByRole('navigation')).toBeInTheDocument();
91
+
92
+ // ❌ Mock existence
93
+ expect(screen.getByTestId('sidebar-mock')).toBeInTheDocument();
94
+ ```
95
+
96
+ **your human partner's correction:** "Are we testing the behavior of a
97
+ mock?"
98
+
99
+ **Mock at the right level.** Learn every side effect of the real method
100
+ before replacing it; mock the slow or external operation and keep what
101
+ the test depends on real. When unsure, run the test against the real
102
+ implementation first and observe what actually needs to happen.
103
+
104
+ ```typescript
105
+ // ❌ The mock swallows the config write that duplicate detection reads
106
+ vi.mock('ToolCatalog', () => ({
107
+ discoverAndCacheTools: vi.fn().mockResolvedValue(undefined)
108
+ }));
109
+
110
+ // ✅ Mock only the slow server startup; the config write stays real
111
+ vi.mock('MCPServerManager');
112
+ ```
113
+
114
+ **Make doubles specific.** When arguments, call counts, or ordering are
115
+ part of the contract, assert them — a fake that accepts anything verifies
116
+ nothing. Give each branch (success, error, malformed) its own fixture or
117
+ spy, so the wrong branch cannot satisfy the expectation.
118
+
119
+ **Mirror real data completely.** Mock the complete structure as it exists
120
+ in reality — all documented fields — not just the ones your test reads.
121
+ Partial mocks fail silently when downstream code reads an omitted field:
122
+ the test passes while integration breaks.
123
+
124
+ **Production classes carry production methods only.** Cleanup that only
125
+ tests need lives in test utilities, never as a `destroy()` on the
126
+ production class. Ask: is this method called only from tests? Does this
127
+ class own this resource's lifecycle? Wrong answers → test utility.
128
+
129
+ **Prefer real components over complex mocks.** When mock setup outgrows
130
+ the test logic, mocks miss methods the real components have, or tests
131
+ break when the mock changes, switch to an integration test with real
132
+ components. **your human partner's question:** "Do we need to be using a
133
+ mock here?"
134
+
135
+ ### Gate Function
136
+
137
+ ```
138
+ BEFORE adding a mock or test helper:
139
+ List the real method's side effects; keep the ones the test
140
+ depends on real — mock the slow/external level below them.
141
+
142
+ Mock responses mirror the complete real structure.
143
+
144
+ A method only tests call lives in test utilities, not production.
145
+
146
+ About to assert on the mock itself?
147
+ Unmock it or delete the assertion.
148
+ ```
149
+
150
+ ## Tests Ship With the Implementation
151
+
152
+ The TDD cycle — failing test, minimal implementation, refactor — is what
153
+ "complete" means. Ship the tests the behavior needs and only those:
154
+ trivial code and human prose earn none, and a test written to satisfy
155
+ process costs maintenance forever.
156
+
157
+ ## The Mutation Check
158
+
159
+ Before finishing, mentally mutate the production code; at least one test
160
+ should fail for each realistic mutation:
161
+
162
+ - Wrong constant or argument
163
+ - Wrong branch handler
164
+ - Missing state change or side effect
165
+ - Empty or default return
166
+ - Missing validation for zero, empty, nil, unauthorized, or malformed input
167
+
168
+ A mutation nothing catches marks the behavior as unprotected — or the
169
+ test as tautological.
170
+
171
+ ## Quick Reference
172
+
173
+ | When you... | Do |
174
+ |-------------|-----|
175
+ | Write any test | Name the break it catches — a bug, not a decision |
176
+ | Build an expected value | Derive it by hand; never with the code under test |
177
+ | Test a script or document | Run it / pressure-test its consumer; never grep its text |
178
+ | Reach for a dependency test | Test your boundary contract, not their documented mechanics |
179
+ | Want to assert on a mocked element | Test the real component, or unmock it |
180
+ | Are about to mock a method | Learn its side effects; mock the slow/external level |
181
+ | Build a mock response | Mirror the real structure completely |
182
+ | Need cleanup only tests use | Put it in test utilities |
183
+ | Watch mock setup balloon | Switch to an integration test with real components |
184
+ | Finish a test file | Run the mutation check |
185
+
186
+ ## Warning Signs
187
+
188
+ - Setup and assertion share the same object, guaranteeing equality
189
+ - The test can fail only through a panic, crash, or missing selector
190
+ - The test fails on every intentional change, never on accidental breakage
191
+ - Expected values are hidden behind loops, builders, or helpers
192
+ - The test greps source text, or asserts a removed symbol stays removed
193
+ - The test would still matter if only the framework remained
194
+ - The test exists for coverage, checking no side effect or outcome
195
+ - An assertion checks a `*-mock` test ID, or fails if you remove the mock
196
+ - A method is called only from test files
197
+ - Mock setup is more than half the test, or you can't explain why the mock is needed
198
+ - Mocking "just to be safe"
@@ -0,0 +1,170 @@
1
+ ---
2
+ name: using-git-worktrees
3
+ description: Use when starting feature work that needs isolation from current workspace or before executing implementation plans - ensures an isolated workspace exists via native tools or git worktree fallback
4
+ ---
5
+
6
+ > **Bizar compatibility:** This procedure is fully integrated into Bizar. Bizar system, repository, model-routing, autonomy, and approval policy control whenever they differ. Do not add a separate mandatory approval gate, recurse into dispatch, or override the active Bizar workflow.
7
+
8
+
9
+ # Using Git Worktrees
10
+
11
+ ## Overview
12
+
13
+ Ensure work happens in an isolated workspace. Prefer your platform's native worktree tools. Fall back to manual git worktrees only when no native tool is available.
14
+
15
+ **Core principle:** Detect existing isolation first. Then use native tools. Then fall back to git. Never fight the harness.
16
+
17
+ **Announce at start:** "I'm using the using-git-worktrees skill to set up an isolated workspace."
18
+
19
+ ## Step 0: Detect Existing Isolation
20
+
21
+ **Before creating anything, check if you are already in an isolated workspace.**
22
+
23
+ ```bash
24
+ GIT_DIR=$(cd "$(git rev-parse --git-dir)" 2>/dev/null && pwd -P)
25
+ GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" 2>/dev/null && pwd -P)
26
+ BRANCH=$(git branch --show-current)
27
+ ```
28
+
29
+ **Submodule guard:** `GIT_DIR != GIT_COMMON` is also true inside git submodules. Before concluding "already in a worktree," verify you are not in a submodule:
30
+
31
+ ```bash
32
+ # If this returns a path, you're in a submodule, not a worktree — treat as normal repo
33
+ git rev-parse --show-superproject-working-tree 2>/dev/null
34
+ ```
35
+
36
+ **If `GIT_DIR != GIT_COMMON` (and not a submodule):** You are already in a linked worktree. Skip to Step 2 (Project Setup). Do NOT create another worktree.
37
+
38
+ Report with branch state:
39
+ - On a branch: "Already in isolated workspace at `<path>` on branch `<name>`."
40
+ - Detached HEAD: "Already in isolated workspace at `<path>` (detached HEAD, externally managed). Branch creation needed at finish time."
41
+
42
+ **If `GIT_DIR == GIT_COMMON` (or in a submodule):** You are in a normal repo checkout.
43
+
44
+ Has the user already indicated their worktree preference in your instructions? If not, ask for consent before creating a worktree:
45
+
46
+ > "Would you like me to set up an isolated worktree? It protects your current branch from changes."
47
+
48
+ Honor any existing declared preference without asking. If the user declines consent, work in place and skip to Step 2.
49
+
50
+ ## Step 1: Create Isolated Workspace
51
+
52
+ **You have two mechanisms. Try them in this order.**
53
+
54
+ ### 1a. Native Worktree Tools (preferred)
55
+
56
+ The user has asked for an isolated workspace (Step 0 consent). Do you already have a way to create a worktree? It might be a tool with a name like `EnterWorktree`, `WorktreeCreate`, a `/worktree` command, or a `--worktree` flag. If you do, use it and skip to Step 2.
57
+
58
+ Native tools handle directory placement, branch creation, and cleanup automatically. Using `git worktree add` when you have a native tool creates phantom state your harness can't see or manage.
59
+
60
+ Only proceed to Step 1b if you have no native worktree tool available.
61
+
62
+ ### 1b. Git Worktree Fallback
63
+
64
+ **Only use this if Step 1a does not apply** — you have no native worktree tool available. Create a worktree manually using git.
65
+
66
+ #### Directory Selection
67
+
68
+ Follow this priority order. Explicit user preference always beats observed filesystem state.
69
+
70
+ 1. **Check your instructions for a declared worktree directory preference.** If the user has already specified one, use it without asking.
71
+
72
+ 2. **Check for an existing project-local worktree directory:**
73
+ ```bash
74
+ ls -d .worktrees 2>/dev/null # Preferred (hidden)
75
+ ls -d worktrees 2>/dev/null # Alternative
76
+ ```
77
+ If found, use it. If both exist, `.worktrees` wins.
78
+
79
+ 3. **If there is no other guidance available**, default to `.worktrees/` at the project root.
80
+
81
+ #### Safety Verification (project-local directories only)
82
+
83
+ **MUST verify directory is ignored before creating worktree:**
84
+
85
+ ```bash
86
+ git check-ignore -q .worktrees 2>/dev/null || git check-ignore -q worktrees 2>/dev/null
87
+ ```
88
+
89
+ **If NOT ignored:** Add to .gitignore, commit the change, then proceed.
90
+
91
+ **Why critical:** Prevents accidentally committing worktree contents to repository.
92
+
93
+ #### Create the Worktree
94
+
95
+ ```bash
96
+ # Determine path based on chosen location
97
+ path="$LOCATION/$BRANCH_NAME"
98
+
99
+ git worktree add "$path" -b "$BRANCH_NAME"
100
+ cd "$path"
101
+ ```
102
+
103
+ **Sandbox fallback:** If `git worktree add` fails with a permission error (sandbox denial), tell the user the sandbox blocked worktree creation and you're working in the current directory instead. Then run setup and baseline tests in place.
104
+
105
+ ## Step 2: Project Setup
106
+
107
+ Auto-detect and run appropriate setup:
108
+
109
+ ```bash
110
+ # Node.js
111
+ if [ -f package.json ]; then npm install; fi
112
+
113
+ # Rust
114
+ if [ -f Cargo.toml ]; then cargo build; fi
115
+
116
+ # Python
117
+ if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
118
+ if [ -f pyproject.toml ]; then poetry install; fi
119
+
120
+ # Go
121
+ if [ -f go.mod ]; then go mod download; fi
122
+ ```
123
+
124
+ ## Step 3: Verify Clean Baseline
125
+
126
+ Run tests to ensure workspace starts clean:
127
+
128
+ ```bash
129
+ # Use project-appropriate command
130
+ npm test / cargo test / pytest / go test ./...
131
+ ```
132
+
133
+ **If tests fail:** Report failures, ask whether to proceed or investigate.
134
+
135
+ **If tests pass:** Report ready.
136
+
137
+ ### Report
138
+
139
+ ```
140
+ Worktree ready at <full-path>
141
+ Tests passing (<N> tests, 0 failures)
142
+ Ready to implement <feature-name>
143
+ ```
144
+
145
+ ## Quick Reference
146
+
147
+ | Situation | Action |
148
+ |-----------|--------|
149
+ | Already in linked worktree | Skip creation (Step 0) |
150
+ | In a submodule | Treat as normal repo (Step 0 guard) |
151
+ | Native worktree tool available | Use it (Step 1a) |
152
+ | No native tool | Git worktree fallback (Step 1b) |
153
+ | `.worktrees/` exists | Use it (verify ignored) |
154
+ | `worktrees/` exists | Use it (verify ignored) |
155
+ | Both exist | Use `.worktrees/` |
156
+ | Neither exists | Check instruction file, then default `.worktrees/` |
157
+ | Directory not ignored | Add to .gitignore + commit |
158
+ | Permission error on create | Sandbox fallback, work in place |
159
+ | Tests fail during baseline | Report failures + ask |
160
+ | No package.json/Cargo.toml | Skip dependency install |
161
+
162
+ ## Common Rationalizations
163
+
164
+ | Excuse | Reality |
165
+ |--------|---------|
166
+ | "I'm obviously not in a worktree — no need to check" | Run Step 0. Harness-created isolation and submodules both fool eyeballing; the detection commands settle it. |
167
+ | "`git worktree add` is quicker than hunting for a native tool" | A native tool (e.g. `EnterWorktree`) owns placement, branching, and cleanup. Bypassing it is the #1 mistake — it creates phantom state your harness can't see or manage. |
168
+ | "The worktree directory is surely ignored already" | Run `git check-ignore`. An unignored worktree directory commits the whole tree into the repo. |
169
+ | "Any directory name works" | Explicit instructions beat an existing project-local directory, which beats the `.worktrees/` default. |
170
+ | "The workspace is fresh — baseline tests can wait" | A dirty baseline makes every later failure ambiguous. Run the tests now; proceeding past failures is your human partner's call. |