@brainervirus/workit-claude-code 3.0.0 → 5.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (39) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/README.md +1 -1
  3. package/agents/implementer.md +25 -15
  4. package/agents/reviewer.md +21 -15
  5. package/agents/verifier.md +27 -18
  6. package/assets/templates/plan-template.md +17 -18
  7. package/assets/templates/spec-template.md +4 -3
  8. package/dist/workit-hook.js +107 -122
  9. package/dist/workit.js +4187 -2775
  10. package/package.json +3 -3
  11. package/skills/bdd/SKILL.md +35 -38
  12. package/skills/continue/SKILL.md +53 -0
  13. package/skills/debug/SKILL.md +41 -45
  14. package/skills/deslop/SKILL.md +35 -34
  15. package/skills/fanout/SKILL.md +62 -0
  16. package/skills/fanout/references/brief.md +56 -0
  17. package/skills/implement/SKILL.md +48 -53
  18. package/skills/review/SKILL.md +42 -60
  19. package/skills/review/references/impact.md +24 -0
  20. package/skills/shape/SKILL.md +71 -0
  21. package/skills/shape/references/diagrams.md +17 -0
  22. package/skills/shape/references/knowledge.md +58 -0
  23. package/skills/shape/references/mockups.md +15 -0
  24. package/skills/shape/references/slicing.md +42 -0
  25. package/skills/ship/SKILL.md +52 -0
  26. package/skills/test-audit/SKILL.md +10 -11
  27. package/skills/verify-app/SKILL.md +63 -0
  28. package/skills/verify-app/references/template.md +49 -0
  29. package/assets/templates/execution-contract.md +0 -40
  30. package/skills/babysit/SKILL.md +0 -46
  31. package/skills/behavioral-tdd/SKILL.md +0 -65
  32. package/skills/blast-radius/SKILL.md +0 -35
  33. package/skills/challenge/SKILL.md +0 -56
  34. package/skills/diagram/SKILL.md +0 -36
  35. package/skills/green-run/SKILL.md +0 -33
  36. package/skills/handoff/SKILL.md +0 -46
  37. package/skills/mockup/SKILL.md +0 -32
  38. package/skills/plan/SKILL.md +0 -54
  39. package/skills/steer/SKILL.md +0 -48
@@ -1,65 +0,0 @@
1
- ---
2
- name: behavioral-tdd
3
- description: Use when policy identifies behavior, side effects, permissions, or data handling that may change and a regression boundary is needed
4
- ---
5
-
6
- # Behavioral TDD
7
-
8
- Test the observable behavior at a stable boundary, not the implementation shape.
9
- Use this method when assessment selects the `testing` dimension.
10
-
11
-
12
- ## Method
13
-
14
- 1. Inspect the task requirement, current candidate, intended behavior, and real
15
- verification entry point with shared `task`, `policy`, and `evidence` operations.
16
- 2. State one behavior and its observable result. Choose the narrowest stable
17
- boundary a caller or user depends on; avoid private helpers and incidental
18
- representations.
19
- 3. Write one vertical RED slice that fails for the missing behavior and run it
20
- through the CLI so the failure is observed: `workit check test` (the repo's
21
- configured `test` check; `npx -y @brainervirus/workit-cli check test` when
22
- `workit` is not on PATH). Implement the smallest change, then run the same
23
- check GREEN. Close accepts only a fresh passing run of the configured check
24
- that `workit check` observed; a recorded "tests pass" is a note, and an
25
- ad-hoc `workit check -- <cmd>` never satisfies the gate. If no `test` check
26
- is configured or detected, add `workit.checks.json` (committed) with it.
27
- 4. Add only another slice for a distinct behavior or risk. Any edit makes the
28
- observed check stale: re-run `workit check test` before closing.
29
-
30
- ## Reject noisy tests
31
-
32
- - A dependency/version-pin assertion is not behavioral evidence.
33
- - A test that mirrors branches, private calls, or exact implementation structure
34
- is coupled to internals; replace it with the public effect.
35
- - Duplicate assertions and tests that add no distinct failure signal are noise;
36
- delete them.
37
- - Do not claim a passing test satisfies a different requirement.
38
- - Banned: tautologies (asserts what the code says, not what it must do),
39
- ghost loops (assert inside a possibly-empty loop), smoke-only renders,
40
- type-only or CSS-class coupling. If the test still passes when every
41
- imported function returns undefined, rewrite the assertion or delete it.
42
- `workit test-audit --diff` flags these; triage them with workit-test-audit.
43
- Turning acceptance criteria into scenarios and seams is workit-bdd.
44
-
45
- Run RED/GREEN through `workit check`, which records the observed result on the
46
- current task; shared `evidence` operations are for notes and non-test evidence.
47
- Do not add a second lifecycle, approval chain, or test workflow outside the
48
- current task state.
49
-
50
- ## Common mistakes
51
-
52
- | Mistake | Correction |
53
- | --- | --- |
54
- | "The pin changed, so assert the new string" | Exercise the affected consumer behavior. |
55
- | "The code is obvious" | A small vertical slice still proves the contract. |
56
- | Keeping a passing test after the boundary moved | Re-run `workit check test` on the current tree. |
57
-
58
-
59
- ## In Claude Code
60
-
61
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
62
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
63
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
64
- and `implementer` agents take independent verification, fresh-context
65
- review and isolated implementation.
@@ -1,35 +0,0 @@
1
- ---
2
- name: blast-radius
3
- description: Use when a small-looking change could break something else, before close or merge
4
- ---
5
-
6
- # Blast radius beyond the diff
7
-
8
- A small diff is not a small risk. Prove the one fact it is safe because of,
9
- with runnable proof — not assertion.
10
-
11
-
12
- ## Method
13
-
14
- 1. List what the change touches: callers, shared state, contracts, config,
15
- migrations. Grep every caller of each touched function.
16
- 2. For each: state the one fact it is safe because of (type boundary,
17
- existing test, unreachable path) plus how to run the proof.
18
- 3. Run the proofs. Unproven claims stay labeled UNPROVEN in findings —
19
- never silently treated as safe.
20
- 4. Fix at the shared root (one guard where all callers route through),
21
- not per caller.
22
-
23
- ## Completion
24
-
25
- Blast-radius note in evidence or review: each risk with fact + proof
26
- command, or an UNPROVEN finding for what could not be proven.
27
-
28
-
29
- ## In Claude Code
30
-
31
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
32
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
33
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
34
- and `implementer` agents take independent verification, fresh-context
35
- review and isolated implementation.
@@ -1,56 +0,0 @@
1
- ---
2
- name: challenge
3
- description: Use when requirements or consequences leave a real decision open
4
- ---
5
-
6
- # Challenge open choices with evidence
7
-
8
- Skip this method for a precise, settled request.
9
-
10
- ## Method
11
-
12
- 1. Ground the outcome, constraints, evidence, and important unknowns. Inspect
13
- code and project docs before asking about facts they can answer.
14
- 2. Separate observations, inferences, and proposals. Test assumptions that
15
- could change the recommendation.
16
- 3. When viable alternatives exist, present two or three genuinely different
17
- approaches. Recommend one; for each, give its benefit, cost or risk, when it
18
- fits, and the smallest useful validation. Do not invent a weak alternative.
19
- 4. Ask only consequential questions, in dependency order. Group independent
20
- choices into one small round. If an authorized reversible default works,
21
- state it and proceed.
22
- 5. Record a settled choice once in the existing conversation or task context.
23
- A discussion decision is knowledge; never fabricate a native permission
24
- receipt or reopen it without new evidence.
25
- 6. Create durable documentation only when requested or when the decision needs
26
- a future reader. Brainstorming alone does not require a spec.
27
-
28
- Say directly when evidence shows a proposal is weak, overcomplicated, or solves
29
- the wrong problem. Raise a counter-case only when it adds a real constraint.
30
-
31
- ## Guardrails
32
-
33
- - Do not create a universal spec, plan, approval chain, or second lifecycle.
34
- - Unknowns that affect an action remain unresolved until evidence or a user
35
- decision closes them.
36
- - Use shared operations for tracked state and provenance; never write task
37
- metadata directly.
38
- - An in-session counter-case is not fresh-context review.
39
-
40
- ## Common mistakes
41
-
42
- | Mistake | Correction |
43
- | --- | --- |
44
- | Asking what code or docs already answer | Ground first; the user owns choices, not lookups. |
45
- | Listing every hypothetical objection | Raise only objections that can change the decision. |
46
- | Asking for agreement without a recommendation | Recommend an option and explain the tradeoff. |
47
- | Treating a discussion decision as permission | Record it as knowledge; preserve native host authority. |
48
-
49
-
50
- ## In Claude Code
51
-
52
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
53
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
54
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
55
- and `implementer` agents take independent verification, fresh-context
56
- review and isolated implementation.
@@ -1,36 +0,0 @@
1
- ---
2
- name: diagram
3
- description: Use when a spec or plan needs a flow or architecture diagram
4
- ---
5
-
6
- # Mermaid when needed, never by default
7
-
8
- Tables first, ASCII trees second, mermaid only when a flow or architecture
9
- needs it. Flowchart, sequence, state, or ER only. No renderer, no network.
10
-
11
-
12
- ## Syntax rules (mermaid v11)
13
-
14
- - Fence as ` ```mermaid `, no surrounding prose inside the fence.
15
- - Quote node labels containing punctuation: `A["input (x, y)"]`.
16
- - One direction per diagram (`TD` or `LR`); keep nodes under twelve.
17
- - Name actors exactly as the codebase names them (real symbols only).
18
-
19
- ## Verify
20
-
21
- Re-read the fence before commit: balanced quotes/brackets, every node
22
- reachable, labels match spec terms. If it cannot be verified by reading,
23
- delete it.
24
-
25
- ## Completion
26
-
27
- One diagram that argues a decision, or nothing. Never a diagram suite.
28
-
29
-
30
- ## In Claude Code
31
-
32
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
33
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
34
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
35
- and `implementer` agents take independent verification, fresh-context
36
- review and isolated implementation.
@@ -1,33 +0,0 @@
1
- ---
2
- name: green-run
3
- description: Use to drive a red CI pipeline back to green, usually inside babysit
4
- ---
5
-
6
- # The CI loop
7
-
8
- Watch, classify, fix, push once, re-verify. Host-native (`gh` / GitLab);
9
- never invent CI APIs.
10
-
11
-
12
- ## Method
13
-
14
- 1. Read the failing checks, not the summary. Quote the failing log lines.
15
- 2. Classify each: flake (rerun once, note it) / stale base (update after
16
- merge-base check) / real failure (reproduce locally, then fix).
17
- 3. Fix at root cause with a regression test; push one wave.
18
- 4. Re-verify the same checks green on the new head. A fix without a
19
- green re-run is not a fix.
20
-
21
- ## Completion
22
-
23
- Green pipeline on the merge head, or an escalated finding with the exact
24
- failing logs when the fix needs the human.
25
-
26
-
27
- ## In Claude Code
28
-
29
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
30
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
31
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
32
- and `implementer` agents take independent verification, fresh-context
33
- review and isolated implementation.
@@ -1,46 +0,0 @@
1
- ---
2
- name: handoff
3
- description: Use when work must continue in another session, host, or agent after interruption, transfer, or compaction
4
- ---
5
-
6
- # Handoff durable task state
7
-
8
- Transfer continuity, not live authority. Use this method when a tracked item
9
- must continue in another session, host, or agent.
10
-
11
- ## Method
12
-
13
- 1. Inspect the current task, scope, decisions, policy requirements, candidate,
14
- evidence, findings, progress, and worker states with shared `task` and `state`
15
- operations.
16
- 2. Export the compact state through `state.export`. Preserve the objective,
17
- exclusions, accepted decisions and reasons, evidence references, open gaps,
18
- findings, candidate identity, blockers, and next action. Do not include
19
- credentials, live writer ownership, or host authority.
20
- 3. Import only through the destination `state.import` operation and its expected
21
- workspace revision. The destination starts paused or otherwise unauthorised
22
- until it observes its own host/session and reconciles stale evidence and
23
- uncertain workers.
24
- 4. Resume or continue through shared `task`, `policy`, `worker`, `writer`, and
25
- `evidence` operations. Record what changed instead of copying a transcript.
26
-
27
- A handoff does not require a formal spec or plan unless those are separate
28
- selected requirements. Never grant destination authority from imported prose,
29
- create a second lifecycle, or edit task metadata directly.
30
-
31
- ## Common mistakes
32
-
33
- | Mistake | Correction |
34
- | --- | --- |
35
- | Sending the whole transcript | Export compact decisions, gaps, evidence, and next action. |
36
- | Restoring the old writer or credentials | Re-observe authority in the destination. |
37
- | Calling a handoff complete without reconciliation | Recheck stale files and uncertain workers first. |
38
-
39
-
40
- ## In Claude Code
41
-
42
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
43
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
44
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
45
- and `implementer` agents take independent verification, fresh-context
46
- review and isolated implementation.
@@ -1,32 +0,0 @@
1
- ---
2
- name: mockup
3
- description: Use when a UI decision needs sketching before implementation
4
- ---
5
-
6
- # ASCII mockups before UI code
7
-
8
- Sketch, don't build. Three genuinely different layout hypotheses maximum,
9
- ASCII only, no code output.
10
-
11
-
12
- ## Method
13
-
14
- 1. Fix a legend (`┌─┐ │ └─┘ ░ ≈ [ ] ( )`) and keep sketches 60-80 cols,
15
- 8-20 rows.
16
- 2. Per hypothesis: regions, component reuse vs new (named against the
17
- existing codebase), empty/loading/populated/error states, nav flow.
18
- 3. Ask at most one clarifying question, then recommend. Flag hi-fi
19
- escalation when ASCII cannot settle it (density, motion, brand).
20
-
21
- ## Completion
22
-
23
- The sketch plus the decision lands in the spec dir. Throwaway by design.
24
-
25
-
26
- ## In Claude Code
27
-
28
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
29
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
30
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
31
- and `implementer` agents take independent verification, fresh-context
32
- review and isolated implementation.
@@ -1,54 +0,0 @@
1
- ---
2
- name: plan
3
- description: Use when dependencies, sequencing, coordination, or resumption make durable next actions useful
4
- ---
5
-
6
- # Plan useful coordination
7
-
8
- A plan is useful when work has dependent steps, a handoff, concurrent actors, or
9
- meaningful unresolved choices. It is never a prerequisite for implementation.
10
-
11
- ## Method
12
-
13
- 1. Establish the requested outcome, constraints, dependencies, and important
14
- unknowns from the available code, docs, and configuration before asking.
15
- 2. Record only the useful sequence: objective, dependency, bounded outcome,
16
- evidence needed, and next action. Reuse the project's existing format.
17
- 3. Start a tracked record only when handoff, dependent steps, coordination, or
18
- durable decisions need continuity. Infer observable facts instead of asking
19
- the user to fill redundant protocol fields.
20
- 4. Create a spec when a durable behavior contract or interface is requested or
21
- will help a future reader. Use an ADR for a consequential trade-off and a
22
- glossary for stable terms. A small fix needs no document.
23
- 5. If the user authorized implementation, proceed through the agreed endpoint
24
- and applicable checks. Do not ask for a separate plan approval or repeat
25
- "continue?". Stop for a new consequential choice, host denial, conflict, or
26
- blocker that cannot be resolved safely.
27
- 6. Update a checkpoint at meaningful boundaries when another session may need
28
- to resume. Keep settled decisions, changed files, actual checks, blockers,
29
- and the next action; omit transcript and process trivia.
30
-
31
- ## Shape
32
-
33
- Plan slices as small end-to-end outcomes with explicit dependencies and checks.
34
- Keep repo policy and native host authority separate. Preserve the user's branch
35
- and commit conventions. Do not prescribe one commit per step or create a second
36
- approval chain unless the user asked for that delivery format.
37
-
38
- ## Common mistakes
39
-
40
- | Mistake | Correction |
41
- | --- | --- |
42
- | Writing a packet for a bounded reversible change | Leave the change undocumented. |
43
- | Asking what repository state or code can answer | Inspect it first. |
44
- | Treating the plan as authority | Follow the user's scope and native host permissions. |
45
- | Copying a transcript into the plan | Keep decisions, evidence, gaps, and next action only. |
46
-
47
-
48
- ## In Claude Code
49
-
50
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
51
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
52
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
53
- and `implementer` agents take independent verification, fresh-context
54
- review and isolated implementation.
@@ -1,48 +0,0 @@
1
- ---
2
- name: steer
3
- description: Use when new instructions, interruptions, or forgotten items change substantial ongoing work
4
- ---
5
-
6
- # Steer without forced lifecycle
7
-
8
- Classify new input as a quick question, a same-task adjustment, or a separate
9
- request. Preserve continuity when it helps; do not manufacture task mutations.
10
-
11
- ## Method
12
-
13
- 1. Answer a quick question from current context. Do not pause, resume, create,
14
- assess, or close a task just to answer it.
15
- 2. For a same-task adjustment, update only affected constraints and next actions.
16
- Keep an existing checkpoint when it helps; reassess policy only when evidence
17
- or constraints changed enough to affect a rule.
18
- 3. For a separate request, do not silently resume an old objective. Park a
19
- concise checkpoint only when substantial work needs to continue later. Start
20
- a distinct tracked record only if the new work benefits from continuity,
21
- dependencies, coordination, or durable decisions. For work spanning repos,
22
- checkpoint each unfinished item with its checkout, branch, requested
23
- deliverables, and delivery endpoint. Keep explicitly held items parked with
24
- their resume condition until the user resumes them. A conversational
25
- checkpoint is sufficient when no task record is needed.
26
- 4. Before resuming a named tracked task, reconcile its checkout, branch, dirty
27
- state, current policy, stale evidence, uncertain effects, and ownership. Do
28
- not change branches, stash, fetch large histories, or seize ownership just
29
- to display a history choice.
30
- 5. Continue to the user's authorized endpoint with applicable checks. Ask only
31
- about consequential choices the code and available context cannot resolve.
32
-
33
- ## Common mistakes
34
-
35
- | Mistake | Correction |
36
- | --- | --- |
37
- | Treating a quick question as interruption | Answer without changing task state. |
38
- | Auto-resuming an old objective | Wait for the user's direction to resume it. |
39
- | Rebuilding continuity from a transcript | Keep a compact checkpoint with evidence and next action. |
40
-
41
-
42
- ## In Claude Code
43
-
44
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
45
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
46
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
47
- and `implementer` agents take independent verification, fresh-context
48
- review and isolated implementation.