@brainervirus/workit-claude-code 4.0.0 → 6.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (40) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/README.md +1 -1
  3. package/agents/implementer.md +25 -15
  4. package/agents/reviewer.md +21 -15
  5. package/agents/verifier.md +27 -18
  6. package/assets/templates/plan-template.md +17 -18
  7. package/assets/templates/spec-template.md +4 -3
  8. package/assets/templates/workit-contract.md +3 -2
  9. package/dist/workit-hook.js +111 -139
  10. package/dist/workit.js +9268 -14584
  11. package/package.json +3 -3
  12. package/skills/bdd/SKILL.md +35 -38
  13. package/skills/continue/SKILL.md +53 -0
  14. package/skills/debug/SKILL.md +41 -45
  15. package/skills/deslop/SKILL.md +35 -34
  16. package/skills/fanout/SKILL.md +62 -0
  17. package/skills/fanout/references/brief.md +56 -0
  18. package/skills/implement/SKILL.md +48 -53
  19. package/skills/review/SKILL.md +42 -60
  20. package/skills/review/references/impact.md +24 -0
  21. package/skills/shape/SKILL.md +71 -0
  22. package/skills/shape/references/diagrams.md +17 -0
  23. package/skills/shape/references/knowledge.md +58 -0
  24. package/skills/shape/references/mockups.md +15 -0
  25. package/skills/shape/references/slicing.md +42 -0
  26. package/skills/ship/SKILL.md +52 -0
  27. package/skills/test-audit/SKILL.md +10 -11
  28. package/skills/verify-app/SKILL.md +63 -0
  29. package/skills/verify-app/references/template.md +49 -0
  30. package/assets/templates/execution-contract.md +0 -40
  31. package/skills/babysit/SKILL.md +0 -46
  32. package/skills/behavioral-tdd/SKILL.md +0 -65
  33. package/skills/blast-radius/SKILL.md +0 -35
  34. package/skills/challenge/SKILL.md +0 -56
  35. package/skills/diagram/SKILL.md +0 -36
  36. package/skills/green-run/SKILL.md +0 -33
  37. package/skills/handoff/SKILL.md +0 -46
  38. package/skills/mockup/SKILL.md +0 -32
  39. package/skills/plan/SKILL.md +0 -54
  40. package/skills/steer/SKILL.md +0 -48
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@brainervirus/workit-claude-code",
3
- "version": "4.0.0",
3
+ "version": "6.0.0",
4
4
  "private": false,
5
5
  "description": "Workit Claude Code plugin — session and per-turn task context, branch policy on git shell commands, workit method skills, and verifier/reviewer/implementer agents",
6
6
  "keywords": [
@@ -39,8 +39,8 @@
39
39
  "build": "bun scripts/build.ts"
40
40
  },
41
41
  "devDependencies": {
42
- "@brainervirus/workit-cli": "^4.0.0",
43
- "@brainervirus/workit-core": "^4.0.0"
42
+ "@brainervirus/workit-cli": "^6.0.0",
43
+ "@brainervirus/workit-core": "^6.0.0"
44
44
  },
45
45
  "engines": {
46
46
  "node": ">=24"
@@ -1,51 +1,48 @@
1
1
  ---
2
2
  name: bdd
3
- description: Use when turning a requirement, issue or acceptance criterion into tests, or when tests should read as behavior (Given/When/Then, BDD, scenarios, acceptance tests, test names and seams)
3
+ description: Turn requirements into Given/When/Then scenarios, agree the test seam, and work test-first in vertical RED/GREEN slices. Use for BDD, TDD, acceptance criteria, scenarios, Given/When/Then, write a test first.
4
4
  ---
5
5
 
6
6
  # Behavior first: Given/When/Then
7
7
 
8
- Write each acceptance criterion as Given/When/Then before any code, then let
9
- it name the test and pick the seam. This skill shapes the scenarios; the
10
- RED/GREEN loop itself is workit-behavioral-tdd.
11
-
12
- ## Method
13
-
14
- 1. Write the scenarios. One behavior per scenario, in the user's or caller's
15
- words: `Given <state>, When <action>, Then <observable result>`. Include the
16
- unhappy paths a caller depends on (denied, empty, invalid, timeout).
17
- 2. Agree the seams. Pick the highest stable interface the scenario can be
18
- observed through: a CLI verb, a public function, an HTTP route. Ideally one
19
- seam per feature. Write the seams down; do not test at an unagreed seam.
20
- 3. Name the tests after the scenarios. The test name is the Given/When/Then
21
- sentence; the body is arrange (Given), act (When), assert (Then). Expected
22
- values come from the scenario (a literal from a worked example or the
23
- spec), never from the code.
24
- 4. Use Gherkin only where the repo already does (`.feature` files with
25
- playwright-bdd, cucumber, jest-cucumber). Otherwise plain test names carry
26
- the scenario; do not add a BDD framework.
27
- 5. Build in vertical slices, one scenario at a time, with
28
- workit-behavioral-tdd: run `workit check test` RED for the new scenario,
29
- make the smallest change, run `workit check test` GREEN, then the next.
30
- 6. Mock only at system boundaries: network, clock, randomness, other
31
- processes, sometimes the filesystem. Never the unit or its internal
8
+ 1. **Write the scenarios.** One behavior each, in the caller's words:
9
+ `Given <state>, When <action>, Then <observable result>`. Include the
10
+ unhappy paths callers depend on (denied, empty, invalid, timeout).
11
+ 2. **Agree the seam.** The highest stable interface the scenario can be
12
+ observed through: a CLI verb, a public function, an HTTP route; ideally one
13
+ per feature. Do not test at a seam nobody agreed to.
14
+ 3. **Name tests after scenarios.** The name is the Given/When/Then sentence;
15
+ the body is arrange, act, assert. Expected values come from the scenario (a
16
+ literal from a worked example, the spec, an external contract), never from
17
+ the code under test.
18
+ 4. **Build in vertical slices.** Write one vertical RED slice that fails for
19
+ the missing behavior and run it through the CLI so the failure is observed:
20
+ `workit check test`. Make the smallest change, run the same check GREEN,
21
+ then take the next scenario. Any edit makes the observation stale; re-run
22
+ before you claim it. A recorded "tests pass" is a note, and an ad-hoc
23
+ `workit check -- <cmd>` never satisfies the gate: only the configured
24
+ `test` check does. No `test` detected? Create `workit.checks.json`, copying
25
+ in every check the repo already runs: once it exists it replaces the
26
+ detected defaults.
27
+ 5. **Mock only at system boundaries:** network, clock, randomness, other
28
+ processes, sometimes the filesystem. Never the unit or its own
32
29
  collaborators; use the real thing or an in-memory adapter behind a port.
30
+ 6. **Gherkin only where the repo already uses it** (`.feature` files with
31
+ playwright-bdd, cucumber, jest-cucumber). Otherwise test names carry it.
33
32
 
34
- ## Completion
33
+ Reject noise: version-pin assertions, tests that mirror private structure,
34
+ assertions inside a possibly-empty loop, smoke-only renders, duplicates. If a
35
+ test still passes when every imported function returns `undefined`, rewrite it
36
+ (workit-test-audit finds these).
35
37
 
36
- Every acceptance criterion maps to a named test at an agreed seam, each was
37
- seen RED then GREEN through `workit check test`, and the new tests have no
38
- tautologies:
38
+ ## Example
39
39
 
40
- ```sh
41
- workit test-audit --diff && workit check test
42
- ```
40
+ Bad: `test("calculateTotal works", () => expect(calculateTotal(items)).toBe(items.reduce((s, i) => s + i.price, 0)))`
43
41
 
42
+ Good: `test("Given two items of 5 and 10, When totalled, Then the total is 15", () => expect(calculateTotal([{ price: 5 }, { price: 10 }])).toBe(15))`
44
43
 
45
- ## In Claude Code
44
+ ## Check
46
45
 
47
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
48
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
49
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
50
- and `implementer` agents take independent verification, fresh-context
51
- review and isolated implementation.
46
+ ```sh
47
+ workit test-audit --diff && workit check test
48
+ ```
@@ -0,0 +1,53 @@
1
+ ---
2
+ name: continue
3
+ description: Keep work on track across interruptions and sessions - sort new input, checkpoint, hand off with a resume brief, and pick up by verifying inherited claims. Use for resume, pick up, handoff, new session, interruption, change of direction.
4
+ ---
5
+
6
+ # Continue without losing the thread
7
+
8
+ ## New input mid-task
9
+
10
+ - **Quick question:** answer it; change nothing else.
11
+ - **Same-task adjustment:** update the affected constraint and next step, then
12
+ keep going.
13
+ - **Separate request:** do not silently resume an old objective, and do not
14
+ drop the current one. Checkpoint it (below) if it must continue later, then
15
+ start the new work. Held items stay parked with their resume condition until
16
+ the user resumes them.
17
+
18
+ ## Checkpoint and hand off
19
+
20
+ ```sh
21
+ workit git commit -m "wip: <state>" --all # nothing lives only in your context
22
+ workit handoff --note "<state in one line>" --next "<next command>" --record
23
+ ```
24
+
25
+ The brief carries the branch, HEAD, dirty state, check freshness, verdict,
26
+ rulings and the next command. Add only what it cannot know: choices still
27
+ open, approaches that failed and why. Work spanning repos gets one brief per
28
+ checkout, each with its branch and delivery endpoint.
29
+
30
+ ## Pick up
31
+
32
+ 1. In the checkout: `workit handoff`, then `workit ledger list` and
33
+ `git log --oneline -10`.
34
+ 2. Trust the trail, verify the claims: re-run the checks the brief calls
35
+ stale, and confirm each "done" item against the goal on the real artifact
36
+ (a pushed SHA, a PR state, a running feature). Do not re-derive settled
37
+ decisions.
38
+ 3. Continue to the recorded endpoint with the brief's next command.
39
+
40
+ ## Example
41
+
42
+ Bad: a new session re-reads the whole codebase, re-asks the user which
43
+ approach to take, and redoes a finished slice.
44
+
45
+ Good: "`workit handoff`: feature/usage at 4be1, `test` stale, verdict none,
46
+ next `workit check test`. Re-ran it: exit 0. The brief says PR #42 is open:
47
+ `workit pr status` confirms, CI pending. Continuing with workit-ship."
48
+
49
+ ## Check
50
+
51
+ ```sh
52
+ workit handoff # read: "next command" is set and no check is listed as stale
53
+ ```
@@ -1,49 +1,45 @@
1
1
  ---
2
2
  name: debug
3
- description: Use when behavior is failing, surprising, contradictory, or regressed and the root cause is not established
3
+ description: Find a root cause before patching - start from a red-capable deterministic repro, rank hypotheses, bisect regressions, fix at the root with a regression test. Use for bug, broken, failing, flaky, regression, error, why does.
4
4
  ---
5
5
 
6
- # Debug the root cause
7
-
8
- Debugging is investigation, not a fast symptom patch. Use this method when
9
- assessment selects `root-cause-investigation` or behavior is failing without an
10
- established root cause.
11
-
12
-
13
- ## Method
14
-
15
- 1. Inspect task scope, caller authority, candidate identity, existing evidence,
16
- findings, and worker/writer state with shared `task`, `policy`, `evidence`, and
17
- `finding` operations.
18
- 2. Reproduce the failure at a stable behavioral boundary. Record observed facts,
19
- inferences, and unknowns with references; trace the failing value and all
20
- relevant callers before editing.
21
- 3. State the root-cause hypothesis and the smallest in-scope fix. Write a focused
22
- regression at the boundary when practical, then run RED and GREEN through
23
- `workit check <name>` so the results are observed, not reported.
24
- 4. Acquire writer authority through `writer` before mutation. Reconcile the
25
- candidate, evidence, and findings after the change; investigate sibling paths
26
- and stale conclusions rather than assuming the first patch worked.
27
-
28
- Respect the user's scope and native authority. For a deterministic failure, make
29
- one focused reproduction that exercises the affected boundary and add a
30
- regression check when practical. If no direct reproduction exists, gather the
31
- available evidence and state what remains uncertain instead of inventing a red
32
- loop or blocking unrelated work.
33
-
34
- ## Common mistakes
35
-
36
- | Mistake | Correction |
37
- | ------------------------------------- | ------------------------------------------------------ |
38
- | Patching the nearest stack frame | Trace the input, callers, and shared cause. |
39
- | Reproducing only after editing | Capture the failure before mutation. |
40
- | Treating one passing command as proof | Verify the affected behavior and record real evidence. |
41
-
42
-
43
- ## In Claude Code
44
-
45
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
46
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
47
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
48
- and `implementer` agents take independent verification, fresh-context
49
- review and isolated implementation.
6
+ # Debug from a red loop
7
+
8
+ ## Steps
9
+
10
+ 1. **Build the loop before any hypothesis.** One command that is red for the
11
+ user's exact symptom, deterministic, fast and runnable by you:
12
+ `workit check -- <repro>`. Shrink it until it fails in seconds. If the
13
+ symptom only shows on the running app, drive it with the project's
14
+ verify-<app> skill. No loop yet? Building it is the task; do not guess-patch.
15
+ 2. **Read the failure, not the summary:** the full error, the failing value,
16
+ and every caller on its path.
17
+ 3. **Rank three to five falsifiable hypotheses**, likeliest first. Test one at
18
+ a time with the loop or a tagged log line (`[DEBUG-<id>]`, removed at the
19
+ end with one grep). Keep a short hypothesis log so a dead idea stays dead.
20
+ 4. **Regression? Bisect it:** `git bisect start <bad> <good>` then
21
+ `git bisect run <repro>`. The first bad commit names the cause.
22
+ 5. **Fix at the root**, the one place every failing caller passes through.
23
+ Add a regression test at a seam that exercises the real bug pattern; if no
24
+ such seam exists, report that as a finding.
25
+ 6. **Stop rule:** after three dead hypotheses, write down what is measured and
26
+ what is inferred, widen the loop, or ask for the one fact only the user has.
27
+
28
+ Every shipped line traces to evidence from the loop. A "might help" retry or
29
+ guard is a hypothesis, not a fix.
30
+
31
+ ## Example
32
+
33
+ Bad: "Probably a race; added a retry." (no repro, nothing measured)
34
+
35
+ Good: "Repro: `workit check -- bun test lock.test.ts -t stale` red 10/10.
36
+ H1 dead pid not reclaimed - confirmed: `kill(pid, 0)` throws EPERM for another
37
+ user's pid and we treated it as dead. Fix: EPERM means alive. Loop green 10/10;
38
+ regression test pins the EPERM case."
39
+
40
+ ## Check
41
+
42
+ ```sh
43
+ workit check -- <repro> # red before the fix, green after
44
+ workit check test
45
+ ```
@@ -1,40 +1,41 @@
1
1
  ---
2
2
  name: deslop
3
- description: Use before opening a PR or after implementation to remove AI slop from code and prose
3
+ description: Remove AI slop before a PR - dead code, comments that restate the code, filler prose - with a minimal diff and identical behavior. Use for deslop, clean up, slop, tidy before PR, remove dead code, trim the PR body.
4
4
  ---
5
5
 
6
6
  # Deslop code and prose
7
7
 
8
- Throughput without quality is slop. Clean it with a minimal diff — deslop
9
- never refactors behavior.
10
-
11
-
12
- ## Method
13
-
14
- 1. Code: delete dead helpers, redundant validators, stub references, and
15
- comments that restate the code. Comments die by default; keep one only
16
- with proof of an unchangeable constraint, encoded structurally if cheap.
17
- 2. Prose (PR body, spec, docs): cut filler, keep real symbol names and
18
- before→after numbers. One doc, one purpose.
19
- 3. Keep the diff minimal: deslop removes lines, never moves logic. If a
20
- cleanup wants behavior change, it becomes its own tasked change.
21
-
22
- ## Completion
23
-
24
- A smaller diff with identical behavior and green checks. Report lines
25
- removed, not lines written.
26
-
27
- Record passing check evidence linked to the `pre-pr-cleanup` requirement id
28
- from the current policy (`kind: check`, `result: passed`, summary naming what
29
- was removed). That requirement gates `hosting.pull_request` and close. If the
30
- change genuinely has nothing to clean, ask for an approved limitation
31
- decision instead of recording evidence that did not happen.
32
-
33
-
34
- ## In Claude Code
35
-
36
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
37
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
38
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
39
- and `implementer` agents take independent verification, fresh-context
40
- review and isolated implementation.
8
+ Throughput without quality is slop. Deslop only removes; it never changes
9
+ behavior. A change that wants new behavior is its own change.
10
+
11
+ 1. **Find it with tools first.** The repo's dead-code and lint tools on the
12
+ branch diff (for example `knip`, `ts-prune`, `vulture`, `cargo udeps`, or
13
+ the linter's unused rules), then read the diff:
14
+ `git diff <base>...HEAD`.
15
+ 2. **Code.** Delete unused helpers and exports, stub references, debug
16
+ leftovers (`[DEBUG-` tags, stray logs), and comments that restate the next
17
+ line. Keep comments that say *why* (a constraint, a workaround with its
18
+ link), license headers and tool directives.
19
+ 3. **Prose** (PR body, spec, docs): cut filler and hedging, keep real symbol
20
+ names and before-to-after numbers. One doc, one purpose.
21
+ 4. **Minimal diff.** Deslop removes lines; it never moves logic. A removed
22
+ validator that changes behavior is not deslop.
23
+ 5. **Re-run the checks** and report lines removed, not lines written. Nothing
24
+ to clean is a valid result: say what you checked ("0 removals; ran knip and
25
+ read the diff"). When a tracked task lists a `pre-pr-cleanup` requirement,
26
+ record this result as its evidence.
27
+
28
+ ## Example
29
+
30
+ Bad: deleting `// retry: the gateway drops the first request after idle (#412)`
31
+ because "comments die".
32
+
33
+ Good: deleting `// increment the counter` above `count += 1`, an unused
34
+ `formatLegacyDate` export reported by knip, and two hedging paragraphs from the
35
+ PR body: "-34 lines, behavior unchanged, `workit check test` exit 0".
36
+
37
+ ## Check
38
+
39
+ ```sh
40
+ workit check test && git diff --stat <base>...HEAD
41
+ ```
@@ -0,0 +1,62 @@
1
+ ---
2
+ name: fanout
3
+ description: Run independent slices in parallel - one worker per isolated worktree with a fixed brief and file-scope manifest, a non-author verifier per slice, results in the ledger. Use for fan out, parallelize, parallel agents, split the work, delegate, swarm.
4
+ ---
5
+
6
+ # Fan out parallel workers
7
+
8
+ Fan out only independent slices: disjoint files, no shared mutable state, each
9
+ verifiable alone. Code-coupled work stays with one owner, who fans out after
10
+ the blocking part lands. A worker whose whole job is re-running one command is
11
+ ceremony; do it yourself.
12
+
13
+ 1. **Slice** (workit-shape): each slice gets a branch, a file-scope manifest
14
+ (the globs it may write) and, if it depends on another, its stack parent.
15
+ 2. **Check disjointness.** No two manifests overlap. Shared files (lockfile,
16
+ registry, barrel exports) belong to one slice, or to you after fan-in.
17
+ 3. **Brief each worker** with the fixed template and refuse to spawn while a
18
+ field is empty: GOAL, SCOPE (the manifest), CONTEXT (pointers, not pasted
19
+ text), ACCEPTANCE (Given/When/Then), VERIFY (exact commands), TIMEBOX,
20
+ FORBIDDEN, REPORT, STANDING. STANDING is every standing order and user
21
+ directive so far, pasted verbatim into each spawn and respawn, because
22
+ directives decay across resumes. Template: `references/brief.md`.
23
+ 4. **Spawn all workers in one message**, in the background, each in its own
24
+ worktree (Claude Code: the `implementer` agent; elsewhere
25
+ `git worktree add --detach ../<repo>-wt/<slug> origin/<base>`). The first
26
+ command a worker runs is `workit git branch <branch> --base <base>`.
27
+ 5. **Judge liveness by side effects only:** new commits and pushes
28
+ (`git log <branch>`), PR and check changes (`workit pr status --branch <b>`).
29
+ No progress past the timebox means stuck. Stop the old worker and observe
30
+ that it exited (a timeout is not proof). `git worktree remove --force`
31
+ drops its uncommitted changes, so first record `git -C <wt> status --short`
32
+ in the ledger or your report; only then remove the worktree. Respawn with the brief in
33
+ `MODE: resume` (consolidated: original, later directives, its last report):
34
+ the new worker runs `git switch <branch>` in its fresh worktree instead of
35
+ `workit git branch`. Never two live workers on one branch. Replace at most
36
+ twice, then re-slice or report the gap. Never chain resumes.
37
+ 6. **Verify each slice independently.** A fresh agent that did not write it
38
+ (Claude Code: the `verifier` agent) runs VERIFY and verify-<app>, then
39
+ `workit ledger verdict <result> --branch <b> --how "<evidence>"` under the
40
+ session you started it with (`WORKIT_SESSION_ID=<lead>-v<n>`, set by you,
41
+ never chosen by the author; Claude Code: the hook names one).
42
+ A worker's report is a pointer, never evidence.
43
+ 7. **Fan in.** Compare `git diff --name-only <base>...<b>` with the slice's
44
+ SCOPE: any file outside it stops the fan-in with a report. Then
45
+ `workit ledger check --branch <b>` for each slice; restack
46
+ stacked slices with `workit stack sync`. Only you touch topology: workers
47
+ never rebase, retarget or merge. Then workit-ship.
48
+
49
+ ## Example
50
+
51
+ Bad brief: "Do the API part and add tests." (no scope, no acceptance, no
52
+ verify command, so nobody can tell when it is done)
53
+
54
+ Good brief: `references/brief.md` (GOAL: `GET /v1/usage` returns daily run
55
+ counts; SCOPE: `src/routes/usage.ts`, `test/usage.test.ts`; VERIFY:
56
+ `workit check test`; ...).
57
+
58
+ ## Check
59
+
60
+ ```sh
61
+ workit ledger check --branch <b> # per slice: accepted (current, passing, independent)
62
+ ```
@@ -0,0 +1,56 @@
1
+ # Worker brief template
2
+
3
+ Every field is required. A brief with an empty field is not spawned. Point to
4
+ files and ledger rows instead of pasting their content.
5
+
6
+ ```md
7
+ MODE: <new | resume (a replacement continuing an existing branch)>
8
+ GOAL: <one observable outcome, in the user's terms>
9
+ SCOPE: <file-scope manifest: the globs this worker may write; everything else is read-only>
10
+ branch: <type>/<slug> base: <trunk or parent branch>
11
+ CONTEXT: <pointers: spec section, ledger decisions, the neighbour file to imitate>
12
+ ACCEPTANCE:
13
+ - Given <state>, When <action>, Then <observable result>
14
+ VERIFY: <exact commands, e.g. `workit check test`, the verify-<app> feature to drive>
15
+ TIMEBOX: <wall clock or turn budget; past it without a new commit you will be replaced>
16
+ FORBIDDEN: <no edits outside SCOPE; no rebase, retarget, merge or force-push; no new dependencies; ...>
17
+ REPORT: branch, head SHA, files changed, each VERIFY command with its exit code,
18
+ each ACCEPTANCE line met / not met, rulings you made (`workit ledger ruling`),
19
+ anything out of scope as a follow-up, not a diff.
20
+ STANDING: <the standing orders, verbatim: user preferences and every directive given so far>
21
+ export WORKIT_SESSION_ID=<lead>-w<n> (set by the lead; a verifier brief gets <lead>-v<n>)
22
+ ```
23
+
24
+ ## Worked example
25
+
26
+ ```md
27
+ MODE: new
28
+ GOAL: `GET /v1/usage` returns the workspace's run count per UTC day for the last 7 days.
29
+ SCOPE: src/routes/usage.ts, src/queries/usage.ts, test/usage.test.ts
30
+ branch: feature/usage-endpoint base: main
31
+ CONTEXT: spec docs/usage/spec.md "Behavior"; ledger decision "counts are per UTC day";
32
+ imitate src/routes/runs.ts for auth and error shape.
33
+ ACCEPTANCE:
34
+ - Given 3 runs today and 1 yesterday, When GET /v1/usage, Then the last two entries are {"runs":1} and {"runs":3}
35
+ - Given no auth header, When GET /v1/usage, Then the status is 401
36
+ VERIFY: `workit check test`; verify-api feature "usage"
37
+ TIMEBOX: 45 minutes
38
+ FORBIDDEN: no edits outside SCOPE; no schema migration; no rebase or force-push; no new packages
39
+ REPORT: as in the template
40
+ STANDING: conventional commits; no comments that restate code; ask nothing, record rulings instead
41
+ ```
42
+
43
+ ## Worker rules (paste into the brief when the host has no implementer agent)
44
+
45
+ 1. First command. `MODE: new`: `workit git branch <branch> --base <base>` (the
46
+ worktree may start on a name that breaks branch policy). `MODE: resume`:
47
+ `git switch <branch>` (the branch exists; the lead removed the dead
48
+ worker's worktree after its exit, so the switch succeeds), then continue
49
+ from its head.
50
+ 2. Decide ambiguities yourself and record them:
51
+ `workit ledger ruling "<what>" --why "<why>" --cost-if-wrong "<cost>"`.
52
+ Stop only for an irreversible action, a security-sensitive one, or a side
53
+ effect outside the worktree.
54
+ 3. Commit with `workit git commit -m "<msg>" -- <paths in SCOPE>`; push only if
55
+ the brief says so.
56
+ 4. Never record a verdict on your own work.
@@ -1,57 +1,52 @@
1
1
  ---
2
2
  name: implement
3
- description: Use when implementing requested code changes in a repository. Follow its rules and host permissions; use Workit tracking and delegation only when continuity or coordination helps.
3
+ description: Build a requested change in small verified steps - follow local patterns, run real checks with workit check, prove it on the running app, hand verification to a non-author. Use for implement, build, add a feature, make the change, code it.
4
4
  ---
5
5
 
6
- # Implement within authority
7
-
8
- Implement within the user request and native host permissions. A task record and
9
- writer ownership are optional coordination tools. Assignment never expands the
10
- request, and a timeout is not proof that a worker stopped.
11
-
12
- ## Method
13
-
14
- 1. Inspect the repository, relevant rules, host capabilities, and current work.
15
- Inspect task state only when this work is already tracked.
16
- 2. If a helper is useful, assign one bounded objective with allowed paths,
17
- applicable requirements, evidence needed, and a stopping condition. Helpers
18
- cannot change scope, record binding decisions, close or pause the task, assign
19
- helpers, or resolve blockers for the lead.
20
- 3. Use writer ownership only when another Workit actor may mutate the same
21
- checkout; release it when coordination ends. Cancellation remains uncertain
22
- until process exit or explicit recovery.
23
- 4. Reconcile helper reports and run the checks appropriate to the requested
24
- outcome. If delegation is unavailable, continue inline when useful.
25
- 5. Before a cross-repo mutation, resolve the actual checkout, branch and remote
26
- from the request and current context. Ask only if competing plausible targets
27
- remain unresolved. Branch, commit, direct push, PR-ready, merge and release
28
- are distinct endpoints; perform only the authorized one under target rules.
29
- Before saying done, reconcile every named deliverable against that checkout
30
- and verify the requested result. For a push, observe the destination remote
31
- ref and confirm it contains the delivered commit; report drift or missing
32
- items as blockers instead of treating local success as remote delivery.
33
-
34
- Do not edit Workit metadata directly, create nested helper trees, widen paths, or
35
- create a second lifecycle. Read-only investigation and bounded reports do not
36
- grant product-write ownership.
37
-
38
- Do not request writer ownership for a solo edit. If a concurrent Workit writer
39
- cannot be fenced, stop only the conflicting managed writes; ordinary writes still
40
- follow the native host permission and sandbox.
41
-
42
- ## Common mistakes
43
-
44
- | Mistake | Correction |
45
- | ---------------------------------------------- | --------------------------------------------------- |
46
- | "The helper timed out, so the writer is free" | Observe exit or perform explicit recovery. |
47
- | Letting a helper approve its own exception | Return the decision to the lead/user. |
48
- | Running a build while another writer is active | Treat builds and tests that mutate state as writes. |
49
-
50
-
51
- ## In Claude Code
52
-
53
- Workit operations (`task`, `evidence`, `policy`, `decision`, …) run through
54
- the `workit` CLI on the Bash tool: `workit <family> <action> --json`
55
- (`workit --help` lists the verbs). The plugin's `verifier`, `reviewer`
56
- and `implementer` agents take independent verification, fresh-context
57
- review and isolated implementation.
6
+ # Implement and prove it works
7
+
8
+ ## Steps
9
+
10
+ 1. Read before writing: the files you will touch, their callers, and one
11
+ neighbour that already does something similar. Copy its patterns, names and
12
+ error handling. Repo rules (AGENTS.md, CLAUDE.md, lint config) win.
13
+ 2. On the default branch? Branch first: `workit git branch --kind feature --slug <s>`.
14
+ 3. Small steps that each leave the tree green. Behavior change: write the
15
+ acceptance as Given/When/Then and see a test fail first (workit-bdd).
16
+ Mechanical change: the existing checks are enough.
17
+ 4. Run the real checks: `workit check test` (and `lint`, `typecheck` when the
18
+ repo has them). A recorded "tests pass" is a note; an observed run counts.
19
+ 5. Prove the feature on its real surface with the project's `verify-<app>`
20
+ skill (none yet? workit-verify-app writes one). Tests show branch behavior,
21
+ not that the feature works.
22
+ 6. Commit: `workit git commit -m "<type>: <what>" -- <paths>` (or `--all`).
23
+ No endpoint named? Stop here and state the next command. Push and open a
24
+ PR (`workit git push`, `workit pr create --fill`, then workit-ship) only when
25
+ that was requested, or the request implies delivery and `workit grant show`
26
+ reports `defaultEndpoint` `pr`; otherwise the endpoint is `commit`.
27
+ 7. Hand off verification. Never record a passing verdict on your own work.
28
+ Start a fresh verifier with its own session (`WORKIT_SESSION_ID=<yours>-v1`;
29
+ Claude Code: the `verifier` agent, which the hook gives one); it runs
30
+ verify-<app> and `workit ledger verdict`. Before saying done, reconcile
31
+ every named deliverable against the target checkout and observe it (for a push:
32
+ `workit verify-delivery push`).
33
+
34
+ Independent slices that could run in parallel go to workit-fanout. When a step
35
+ stalls on a fact, find it (read, run, prototype); ask only for a product or
36
+ preference choice, with your recommended answer.
37
+
38
+ ## Example
39
+
40
+ Bad: "Done - added the --since flag, tests pass." (no check ran this turn; the
41
+ flag was never invoked)
42
+
43
+ Good: "Added `--since`. measured: `workit check test` exit 0;
44
+ `mytool log --since 2d` printed 3 entries against the fixture repo. inferred:
45
+ the GitLab path behaves the same (shared parser, not run). Verification handed
46
+ to the verifier agent."
47
+
48
+ ## Check
49
+
50
+ ```sh
51
+ workit check test # then, when the endpoint was a push or beyond: workit verify-delivery push
52
+ ```