superwiki 0.1.2 → 0.1.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "sw",
3
3
  "description": "An LLM-maintained wiki and task tracker in docs/ for coding agents. Obsidian friendly, with a static viewer.",
4
- "version": "0.1.2",
4
+ "version": "0.1.4",
5
5
  "license": "MIT",
6
6
  "keywords": [
7
7
  "wiki",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "sw",
3
- "version": "0.1.2",
3
+ "version": "0.1.4",
4
4
  "description": "An LLM-maintained wiki and task tracker in docs/ for coding agents. Obsidian friendly, with a static viewer.",
5
5
  "license": "MIT",
6
6
  "skills": "./skills/",
package/README.md CHANGED
@@ -17,9 +17,13 @@ Then, in a project: `/sw-init`.
17
17
 
18
18
  ## Why
19
19
 
20
- Superwiki follows the LLM Wiki pattern described by Andrej Karpathy: raw sources you curate, a wiki the agent owns, and a short schema that tells the agent how to maintain it. On top of that it adds what a software project needs: tasks with dependencies, plans, and a record of decisions and lessons.
20
+ Superwiki follows the LLM Wiki pattern described by Andrej Karpathy: raw sources you curate, a wiki the agent owns, and a short schema that tells the agent how to maintain it. On top of that it adds what a software project needs: tasks with dependencies, plans, reviews, and a record of decisions and lessons.
21
21
 
22
- It is built to be cheap for the agent. One small index to read, one file per task, and a script that answers "what is ready?", "what blocks this?" or "is anything broken?" without the agent reading the vault.
22
+ It is built to be cheap for the agent:
23
+
24
+ - **Little to read.** One small index, one file per task, and a script that answers "what is ready?", "what blocks this?" or "is anything broken?" without the agent reading the vault.
25
+ - **Work in clean contexts.** Planning, implementing and reviewing run in subagents, each on the model you choose for it. The main session only keeps the task's status true, so it stays small.
26
+ - **Cost you can see.** `sw-stats` shows what a session used, per agent.
23
27
 
24
28
  Measured on a real project with 165 tasks, converted from a single markdown index:
25
29
 
@@ -29,7 +33,7 @@ Measured on a real project with 165 tasks, converted from a single markdown inde
29
33
  | Read to start one task | the index, then the task's section | one file, 2 KB at the median |
30
34
  | Marking a task done | a status cell, plus a ✅ at every reference to it (median 12 places) | one frontmatter line |
31
35
 
32
- > Status: early. The CLI, the viewer, `sw-init` and the migration script are tested, and the planning flow has been run in all three agents; some skills have only been exercised once. [DESIGN.md](DESIGN.md) lists what has and has not been proven.
36
+ > Status: early. The CLI, the viewer, `sw-init` and the migration script are tested. Planning, implementing and reviewing have been run end to end on real projects in Claude Code, and the planning flow in Codex and Copilot CLI; some skills have only been exercised once. [DESIGN.md](DESIGN.md) lists what has and has not been proven.
33
37
 
34
38
  ## What you get
35
39
 
@@ -84,9 +88,13 @@ git clone https://github.com/mhmtsrfglu/superwiki ~/.superwiki
84
88
 
85
89
  A linked install follows the clone: `git pull` updates every agent. `install.sh --copy` copies instead, `--uninstall` removes.
86
90
 
87
- All three agents below were checked the same way: the agent found the skills, refused to start a task with an unfinished dependency, and ran `sw-plan` end to end with the planner subagent.
91
+ ### Per agent
92
+
93
+ All three agents were checked the same way: the agent found the skills, refused to start a task with an unfinished dependency, and ran `sw-plan` end to end with the planner subagent.
88
94
 
89
- ### Claude Code
95
+ The model for each role (planning, implementing, reviewing) is set with `sw-config`, which writes one agent file per role into the project. The skills dispatch those agents by name.
96
+
97
+ #### Claude Code
90
98
 
91
99
  ```bash
92
100
  npx superwiki install claude
@@ -94,11 +102,11 @@ npx superwiki install claude
94
102
 
95
103
  Invoke with a slash: `/sw-init`, `/sw-plan T-01`.
96
104
 
97
- Claude Code reads `~/.claude/skills/` (and a project's `.claude/skills/`); it does not read `~/.agents/skills/`, so `global` is not enough for it.
98
-
99
- `sw-plan` enters plan mode when the session offers it. A model set with `sw-config` for `claude` applies to the planner and implementer subagents, written to `.claude/agents/`; if those agents are not loaded, the skills fall back to built-in agents with the same model.
105
+ - Claude Code reads `~/.claude/skills/` (and a project's `.claude/skills/`). It does not read `~/.agents/skills/`, so `global` is not enough for it.
106
+ - `sw-plan` enters plan mode when the session offers it.
107
+ - Agent files: `.claude/agents/sw-planner.md`, `sw-implementer.md` and `sw-reviewer.md`. They load when a session starts; in a session that began before they existed, the skills fall back to built-in agents with the same model.
100
108
 
101
- ### Codex CLI
109
+ #### Codex CLI
102
110
 
103
111
  ```bash
104
112
  npx superwiki install codex
@@ -106,9 +114,10 @@ npx superwiki install codex
106
114
 
107
115
  Invoke with a dollar sign, or by name in a sentence: `$sw-init`, `$sw-plan T-01`, "use the sw-plan skill for T-01". Checked with CLI 0.153.
108
116
 
109
- A skill cannot switch Codex into plan mode; start planning yourself with `/plan` if you want the mode, or let `sw-plan` proceed without it (it changes no file before you approve). A model set with `sw-config` for `codex` is written to `.codex/agents/sw-planner.toml` and `sw-implementer.toml`, and the skills spawn those agents by name. Subagents must be enabled (they are by default in current releases).
117
+ - A skill cannot switch Codex into plan mode. Start planning yourself with `/plan` if you want the mode, or let `sw-plan` proceed without it: it changes no file before you approve.
118
+ - Agent files: `.codex/agents/sw-planner.toml`, `sw-implementer.toml` and `sw-reviewer.toml`. Subagents must be enabled; they are by default in current releases.
110
119
 
111
- ### GitHub Copilot CLI
120
+ #### GitHub Copilot CLI
112
121
 
113
122
  ```bash
114
123
  npx superwiki install copilot
@@ -116,11 +125,11 @@ npx superwiki install copilot
116
125
 
117
126
  Invoke with a slash, or by name in a sentence: `/sw-init`, "use the sw-plan skill for T-01". Checked with CLI 1.0.31.
118
127
 
119
- Copilot CLI also reads `~/.agents/skills/`, so if you installed `codex` or `global` it already has the skills.
128
+ - Copilot CLI also reads `~/.agents/skills/`, so if you installed `codex` or `global` it already has the skills.
129
+ - A skill cannot switch Copilot into plan mode. Start with `copilot --mode plan` or `/plan` if you want it.
130
+ - Agent files: `.github/agents/sw-planner.agent.md`, `sw-implementer.agent.md` and `sw-reviewer.agent.md`, dispatched with the `task` tool. Whether Copilot honours the `model:` field of those files has not been checked.
120
131
 
121
- A skill cannot switch Copilot into plan mode; start with `copilot --mode plan` or `/plan` if you want it. A model set with `sw-config` for `copilot` is written to `.github/agents/sw-planner.agent.md` and `sw-implementer.agent.md`; the skills dispatch them with the `task` tool. Whether Copilot honours the `model:` field of those files has not been checked.
122
-
123
- ### Other agents
132
+ #### Other agents
124
133
 
125
134
  ```bash
126
135
  npx superwiki install global
@@ -128,8 +137,9 @@ npx superwiki install global
128
137
 
129
138
  Agents that load `SKILL.md` folders from `~/.agents/skills` pick the skills up from there. For an agent with its own skills folder (Cursor, Gemini CLI, OpenCode and others), copy the `skills/sw-*` folders from a clone into it by hand. Nothing has been run in these agents. What will differ:
130
139
 
131
- - the skills name Claude Code, Codex and Copilot tools when they dispatch subagents; elsewhere they fall back to doing the planning or implementing in the main session, and say so;
132
- - `sw-config` writes agent files only for `claude`, `codex` and `copilot`, so a per-role model cannot be set.
140
+ - the skills name Claude Code, Codex and Copilot tools when they dispatch subagents; elsewhere they fall back to doing the work in the main session, and say so;
141
+ - `sw-config` writes agent files only for `claude`, `codex` and `copilot`, so a per-role model cannot be set;
142
+ - `sw-stats` reads the session records of those three tools only.
133
143
 
134
144
  Everything else (the vault, the CLI, the viewer, ingest, lint, explain, triage) depends only on Node and on the agent following the skill text.
135
145
 
@@ -156,15 +166,16 @@ A project's `docs/` folder is plain markdown and keeps working as an Obsidian va
156
166
  | Skill | What it does |
157
167
  | --- | --- |
158
168
  | `sw-init` | set up `docs/` in the current project, or upgrade it |
159
- | `sw-migrate` | convert an existing table-based task index, on a new git branch |
169
+ | `sw-migrate` | convert an existing table-based task index, on a git branch of its own |
160
170
  | `sw-ingest` | file a source into the wiki |
161
171
  | `sw-plan` | plan a task with the planner subagent and get your approval |
162
- | `sw-implement` | run a task with the implementer subagent and record the result |
172
+ | `sw-implement` | run a task with the implementer subagent, have it reviewed if the task asks for that, and record the result |
163
173
  | `sw-explain` | explain a task: what, why, dependencies, what it unblocks |
164
174
  | `sw-triage` | for a problem: seen before? lessons, likely causes |
165
175
  | `sw-lint` | structural checks by script, semantic review on request |
166
176
  | `sw-visualize` | open the viewer |
167
- | `sw-config` | the model each tool uses for planning and implementing; task areas |
177
+ | `sw-stats` | what the current session has cost: tokens, context, steps and tool calls, per agent |
178
+ | `sw-config` | the model each tool uses for planning, implementing and reviewing; task areas |
168
179
 
169
180
  ### Examples
170
181
 
@@ -188,9 +199,10 @@ Shown as typed in Claude Code. In Codex, write `$sw-plan` instead of `/sw-plan`.
188
199
  /sw-ingest ~/Downloads/interview-notes.md
189
200
  file a source and summarise it into the wiki
190
201
 
191
- /sw-config plan with opus, implement with sonnet
202
+ /sw-config plan with opus, implement with sonnet, review with opus
192
203
  /sw-lint check links, frontmatter and task dependencies
193
204
  /sw-visualize open the task board and the wiki in the browser
205
+ /sw-stats what this session has cost so far, per agent
194
206
  ```
195
207
 
196
208
  You do not have to type a command. The rules `sw-init` adds to `AGENTS.md` tell the agent which skill fits, so a plain request should reach the same skill:
@@ -204,6 +216,22 @@ Users get the magic-link email twice. Have we seen this before?
204
216
 
205
217
  A filled-in example vault is in [examples/demo/docs](examples/demo/docs).
206
218
 
219
+ ### What a session cost
220
+
221
+ `sw-stats` reads the record your agent keeps of the session and prints one row for the main session and one for each subagent. It writes nothing. `sw-implement` ends its report with the same table. This one is a real task: planned, implemented and reviewed in 37 minutes.
222
+
223
+ ```text
224
+ session claude fe4cbd6c-8abf-4db6-9755-469dc7321dc7 2026-10-05 10:27 to 11:04, 37 min
225
+ agent model steps first peak sent cached output tools min
226
+ main claude-opus-5-5 27 79k 126k 2.8M 96% 18k 23 37
227
+ sw-planner claude-opus-5-5 31 59k 154k 3.5M 96% 8k 32 6
228
+ sw-implementer claude-sonnet-5-5 68 59k 252k 12.2M 96% 16k 76 26
229
+ sw-reviewer claude-opus-5-5 23 60k 137k 2.4M 89% 372 24 12
230
+ total 149 - - 20.9M 95% 43k 155
231
+ ```
232
+
233
+ `first` and `peak` are the tokens sent with one request. `sent` is that, summed over every step: each step sends the whole context again, which is why a long session in one context is expensive. In Copilot CLI the token columns fill in once the session has closed.
234
+
207
235
  ### The CLI
208
236
 
209
237
  The skills call a small script that answers questions without the agent reading the vault. You can run it yourself, from the project root:
@@ -211,11 +239,12 @@ The skills call a small script that answers questions without the agent reading
211
239
  ```bash
212
240
  node docs/.sw/sw.mjs status # counts per area
213
241
  node docs/.sw/sw.mjs ready # tasks that can start now
214
- node docs/.sw/sw.mjs check P-15 # can it start or finish, and what is open
242
+ node docs/.sw/sw.mjs check P-15 # can it start or finish, what is open, is a review required
215
243
  node docs/.sw/sw.mjs explain P-15 # dependencies, what it unblocks, plan, area guide
216
244
  node docs/.sw/sw.mjs search sync timeout # where something is mentioned
217
245
  node docs/.sw/sw.mjs next-id P # next free id in an area
218
246
  node docs/.sw/sw.mjs lint # broken links, bad frontmatter, dependency errors
247
+ node docs/.sw/sw.mjs stats # tokens, context and steps of the agent session here
219
248
  node docs/.sw/sw.mjs serve --open # the viewer, reading files live
220
249
  node docs/.sw/sw.mjs snapshot # or: freeze the vault into docs/viewer.html, no server
221
250
  ```
@@ -223,14 +252,23 @@ node docs/.sw/sw.mjs snapshot # or: freeze the vault into docs/vie
223
252
  ## Develop
224
253
 
225
254
  ```bash
226
- npm test # builds skills/sw-init/assets/sw.mjs, then runs the tests
255
+ npm test # builds skills/sw-init/assets/sw.mjs and viewer.html, then runs the tests
227
256
  ```
228
257
 
229
- Releases are cut by the `Release` workflow (Actions → Release → Run workflow): it tests, bumps the version in `package.json` and the plugin manifests, publishes to npm, tags, and creates a GitHub release. It publishes through npm trusted publishing, so no token is stored: the package's settings on npmjs.com name this repository and `release.yml` as its trusted publisher.
258
+ Edit sources in `src/`:
259
+
260
+ | File | What it is |
261
+ | --- | --- |
262
+ | `src/core.js` | the vault model, derived task state, lint and search; shared by the CLI and the viewer |
263
+ | `src/stats.js` | reads the agents' session records |
264
+ | `src/cli.js` | the commands |
265
+ | `src/viewer.html` | the viewer |
266
+
267
+ `scripts/build.mjs` bundles them into `skills/sw-init/assets/sw.mjs` and `viewer.html`. Those two files are generated: do not edit them.
230
268
 
231
269
  `node scripts/build-demo.mjs` builds the public demo into `site/` (the viewer with the example vault baked in); the Pages workflow deploys it on every push to `main`.
232
270
 
233
- `src/core.js` is shared by the CLI and the viewer. Edit sources in `src/`; the files in `skills/sw-init/assets/` named `sw.mjs` and `viewer.html` are generated.
271
+ Releases are cut by the `Release` workflow (Actions → Release → Run workflow): it tests, bumps the version in `package.json` and the plugin manifests, publishes to npm, tags, and creates a GitHub release. It publishes with the repository secret `NPM_TOKEN`, an npm access token allowed to publish `superwiki`.
234
272
 
235
273
  ## License
236
274
 
@@ -0,0 +1,5 @@
1
+ ---
2
+ description: Show what this session has cost so far: tokens, context, steps and tool calls per agent
3
+ ---
4
+
5
+ Use the sw-stats skill. User arguments: $ARGUMENTS
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "superwiki",
3
- "version": "0.1.2",
3
+ "version": "0.1.4",
4
4
  "description": "Agent skills that turn docs/ into an LLM-maintained wiki and task tracker. Obsidian-friendly. Works with Claude Code, Codex and Copilot CLI.",
5
5
  "keywords": [
6
6
  "claude-code",
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: sw-config
3
- description: Use when the user wants to choose or change which model plans or implements Superwiki tasks (opus, sonnet, gpt and so on), add task areas, see the Superwiki configuration, or invokes sw-config or sw:config.
3
+ description: Use when the user wants to choose or change which model plans, implements or reviews Superwiki tasks (opus, sonnet, gpt and so on), add task areas, see the Superwiki configuration, or invokes sw-config or sw:config.
4
4
  ---
5
5
 
6
6
  # sw-config
@@ -10,7 +10,8 @@ Settings live in `docs/.sw/config.json`. Change them with the script, from the p
10
10
  ```bash
11
11
  node <skill-dir>/scripts/config.mjs show
12
12
  node <skill-dir>/scripts/config.mjs model plan claude opus
13
- node <skill-dir>/scripts/config.mjs model implement codex gpt-6
13
+ node <skill-dir>/scripts/config.mjs model implement claude sonnet
14
+ node <skill-dir>/scripts/config.mjs model review codex gpt-6
14
15
  node <skill-dir>/scripts/config.mjs model plan copilot --unset
15
16
  node <skill-dir>/scripts/config.mjs areas "M=Mobile,B=Backend"
16
17
  node <skill-dir>/scripts/config.mjs sync --tools claude,codex,copilot
@@ -18,13 +19,21 @@ node <skill-dir>/scripts/config.mjs sync --tools claude,codex,copilot
18
19
 
19
20
  ## Models
20
21
 
21
- A model is chosen per role (`plan`, `implement`) and per tool (`claude`, `codex`, `copilot`), because each tool can only run its own models. No tool lets a skill change the model of the running session, so sw-plan and sw-implement hand the work to a subagent, and the subagent's file carries the model.
22
+ A model is chosen per role and per tool, because each tool can only run its own models.
22
23
 
23
- | Tool | Files the script writes |
24
- |---|---|
25
- | Claude Code | `.claude/agents/sw-planner.md`, `sw-implementer.md` |
26
- | Codex | `.codex/agents/sw-planner.toml`, `sw-implementer.toml` |
27
- | Copilot CLI | `.github/agents/sw-planner.agent.md`, `sw-implementer.agent.md` |
24
+ | Role | Does | Used by |
25
+ | --- | --- | --- |
26
+ | `plan` | writes the plan file for a task | sw-plan |
27
+ | `implement` | does the work of a task | sw-implement |
28
+ | `review` | reviews the implementation from a clean context, for tasks whose frontmatter has `review:` | sw-implement |
29
+
30
+ No tool lets a skill change the model of the running session, so the skills hand the work to a subagent, and the subagent's file carries the model. The script writes one file per role:
31
+
32
+ | Tool | Folder | Files |
33
+ | --- | --- | --- |
34
+ | Claude Code | `.claude/agents/` | `sw-planner.md`, `sw-implementer.md`, `sw-reviewer.md` |
35
+ | Codex | `.codex/agents/` | `sw-planner.toml`, `sw-implementer.toml`, `sw-reviewer.toml` |
36
+ | Copilot CLI | `.github/agents/` | `sw-planner.agent.md`, `sw-implementer.agent.md`, `sw-reviewer.agent.md` |
28
37
 
29
38
  When the user asks to set a model:
30
39
 
@@ -32,6 +41,8 @@ When the user asks to set a model:
32
41
  2. Run the `model` command. Show its output.
33
42
  3. Say that a tool picks up new agent files when its next session starts.
34
43
 
44
+ After Superwiki itself is updated, run `sync` once: the roles' instructions are part of the agent files.
45
+
35
46
  Do not edit the generated agent files or `config.json` by hand; the next `sync` overwrites agent files.
36
47
 
37
48
  ## Areas
@@ -0,0 +1,41 @@
1
+ # Reviewer
2
+
3
+ You review the implementation of one task in a Superwiki vault. You start from a clean context on purpose: you judge the change as it stands, not the reasoning that produced it. You change no file in the repository.
4
+
5
+ Input: a task id and the list of files the implementation changed; possibly which checks you may run that need services or data.
6
+
7
+ What you read is what this review costs, and every extra step re-sends everything you have read so far. Read little, in few steps. Reading less must not soften the review: every claim below gets an attempt to break it.
8
+
9
+ ## Start
10
+
11
+ 1. Read `docs/tasks/<ID>.md`. If `docs/plans/<ID>-plan.md` exists, read its `## Approach` only.
12
+ 2. Project rules (`AGENTS.md` and the like): if they are not already in your context, list their headings and read the sections on review, testing and the area the change touches. Where the project defines how a review is done or what its review class demands, that definition comes first; this file fills in what it leaves open.
13
+ 3. Read the change: the listed files, at the places that changed. Use the version control diff if you may run it; otherwise read the files.
14
+
15
+ ## Review
16
+
17
+ 1. **List the claims.** Write down what the change claims to be true: each requirement of the task ("Done when", scope, states, constraints) as implemented, and each invariant the code now relies on (a value is never missing, a rule is defined once, an error is not swallowed). Aim for the claims whose failure would be silent.
18
+ 2. **Try to break each claim.** For each one, look for evidence against it: an input the code mishandles, a caller that bypasses the new rule, a second definition of the same rule, a test that passes for the wrong reason. Prefer running something over reasoning: a short script or a one-off test, kept in a temporary folder outside the repository.
19
+ 3. **Check the tests.** Does a test fail if the claim is false? Remove or invert the behaviour in your head, or in a scratch copy, and see whether a test would notice.
20
+ 4. **Classify what you find.**
21
+ - `blocking`: the task's requirement is not met, or the change can produce a wrong result without anyone noticing.
22
+ - `important`: a real defect or gap that does not make the result wrong today.
23
+ - `minor`: clarity, naming, small cleanups.
24
+
25
+ A check marked `needs: ...` in the plan runs only if your input says it may.
26
+
27
+ ## How to read
28
+
29
+ - **Locate, then open.** Search for the symbol or string first; open the range the search points at.
30
+ - **Follow the change outward only as far as a claim needs.** A caller matters when a claim depends on how it calls.
31
+ - **Batch lookups.** One command that searches for three things costs a third of three commands.
32
+ - **Never read twice.**
33
+
34
+ ## Report
35
+
36
+ About 30 lines:
37
+
38
+ - `Verdict:` `pass` when nothing is blocking, otherwise `changes needed`;
39
+ - `Claims:` each claim, one line, with what you tried against it and the result;
40
+ - `Findings:` each finding with its class, the file and line, and the evidence (the input, command or reading that shows it). No finding without evidence;
41
+ - `Not checked:` what you could not verify, and what it would need.
@@ -21,6 +21,11 @@ const ROLES = {
21
21
  instructions: 'implementer.md',
22
22
  description: 'Implements one Superwiki task from its task and plan files. Use from sw-implement.',
23
23
  },
24
+ review: {
25
+ name: 'sw-reviewer',
26
+ instructions: 'reviewer.md',
27
+ description: 'Reviews the implementation of one Superwiki task from a clean context. Use from sw-implement.',
28
+ },
24
29
  };
25
30
 
26
31
  // How each tool wants an agent defined: where the file goes and what it looks like.
@@ -81,7 +86,7 @@ function loadConfig(root) {
81
86
  const path = join(root, 'docs/.sw/config.json');
82
87
  if (!existsSync(path)) throw new UsageError('no docs/.sw/config.json here; run sw-init first');
83
88
  const config = JSON.parse(readFileSync(path, 'utf8'));
84
- config.models = { plan: {}, implement: {}, ...config.models };
89
+ config.models = { ...Object.fromEntries(ROLE_NAMES.map(role => [role, {}])), ...config.models };
85
90
  config.tools ??= [];
86
91
  return { config, save: () => writeFileSync(path, JSON.stringify(config, null, 2) + '\n') };
87
92
  }
@@ -5,7 +5,9 @@ description: Use when the user wants to implement, build, execute, start or cont
5
5
 
6
6
  # sw-implement
7
7
 
8
- Runs one task. You keep the task's status true and judge the result; an implementer subagent, running the model set in sw-config, does the work from the task and plan files. You do not read the code or the plan: the report is your input.
8
+ Runs one task. You keep the task's status true and judge the result. The work is done by subagents that start from a clean context, on the models set in sw-config: an implementer, and a reviewer when the task asks for one. You do not read the code or the plan: their reports are your input.
9
+
10
+ That split is what keeps a task cheap. A long session sends its whole context again on every step; work done in a fresh context does not carry yours, and yours stays small because the work never enters it.
9
11
 
10
12
  Run commands from the project root. `<skill-dir>` is the directory this SKILL.md is in.
11
13
 
@@ -13,44 +15,68 @@ Run commands from the project root. `<skill-dir>` is the directory this SKILL.md
13
15
 
14
16
  1. **Pick the task.** Id given: use it. Otherwise `node docs/.sw/sw.mjs ready` and let the user choose; tasks already in progress come first.
15
17
  2. **Gate**: `node docs/.sw/sw.mjs check <ID>`.
16
- - `can start: no open deps: ...`: stop. Tell the user which tasks block it and offer to run the first blocker instead; do not run it unasked. Do not start the task anyway, and do not edit `deps` to get past this.
17
- - `can start: n/a, status is in-progress`: this is a continuation; skip step 3.
18
- - `can start: n/a, status is done` or `cancelled`: stop and ask what the user wants.
19
- - `plan: ... (draft, not approved)`: stop; the plan needs the user's approval (sw-plan).
20
- - `plan: none`: fine for a small task (one area, three "Done when" items or fewer, nothing open in its notes, a few files). For anything larger, recommend sw-plan first and let the user choose.
21
- 3. **Mark it started** before any work: in the frontmatter of `docs/tasks/<ID>.md` set `status: in-progress` and `started:` today. Append `## [date] task | <ID> started` to `docs/log.md`, in the layout its last entries use.
22
- 4. **Checks that need the environment.** If the task has a plan, look for `needs:` in it: `grep -n 'needs:' docs/plans/<ID>-plan.md`. Each hit is a check that starts services or changes data. Ask the user which of them the implementer may run; without a yes, none.
23
- 5. **Dispatch the implementer.** Its prompt is the task id, the project root if it is not your working directory, and which `needs:` checks it may run. Do not paste the plan into the prompt; it reads the files.
24
18
 
25
- | Tool | How |
19
+ | `check` says | Do |
26
20
  | --- | --- |
27
- | Claude Code | agent `sw-implementer`. If it is not among your agent types, use a general-purpose agent, tell it to read `<skill-dir>/../sw-config/assets/implementer.md` first and follow it, and pass the model from `models.implement.claude` in `docs/.sw/config.json` if set |
28
- | Codex | spawn the custom agent `sw_implementer` |
29
- | Copilot CLI | `task` tool with agent `sw-implementer` |
30
- | No subagents available, or the agent is not defined | follow `implementer.md` yourself, in this session, and tell the user the configured model was not used |
21
+ | `can start: no open deps: ...` | stop. Tell the user which tasks block it and offer to run the first blocker instead; do not run it unasked. Do not start the task anyway, and do not edit `deps` to get past this |
22
+ | `can start: n/a, status is in-progress` | this is a continuation; skip step 3 |
23
+ | `can start: n/a, status is done` or `cancelled` | stop and ask what the user wants |
24
+ | `plan: ... (draft, not approved)` | stop; the plan needs the user's approval (sw-plan) |
25
+ | `plan: none` | fine for a small task: one area, three "Done when" items or fewer, nothing open in its notes, a few files. For anything larger, recommend sw-plan first and let the user choose |
26
+ | `review: required (...)` | remember it for step 7 |
31
27
 
28
+ 3. **Mark it started** before any work: in the frontmatter of `docs/tasks/<ID>.md` set `status: in-progress` and `started:` today. Append `## [date] task | <ID> started` to `docs/log.md`, in the layout its last entries use.
29
+ 4. **Checks that need the environment.** If the task has a plan, look for `needs:` in it: `grep -n 'needs:' docs/plans/<ID>-plan.md`. Each hit is a check that starts services or changes data. Ask the user which of them may run; without a yes, none.
30
+ 5. **Dispatch the implementer** (how: "Dispatching" below). Its prompt is the task id, the project root if it is not your working directory, and which `needs:` checks it may run. Do not paste the plan into the prompt; it reads the files.
32
31
  6. **Judge the report.** Its `Requirements:` list must name every "Done when" item and every scope, state or constraint item of the task; compare it with the task file.
33
32
  - `met` needs evidence: a command or test and its result. Re-run one verification command yourself when the evidence is vague.
34
33
  - `not met`, or missing from the list: the task is not done.
35
34
  - `differs`: the implementer built something other than what the task says. That is the user's call: show it and ask. Until they accept it, the item is not met.
36
- 7. **Record the outcome.**
35
+ 7. **Review, if the task requires it.** Only when every requirement is met or accepted: dispatch the reviewer with the task id, the files the implementer changed, and which `needs:` checks it may run.
36
+ - `Verdict: pass`: go on. Pass `important` and `minor` findings to the user in your report; they do not block.
37
+ - `Verdict: changes needed`: dispatch the implementer again with the blocking findings, word for word, then the reviewer again with the files changed since. After two rounds that still end in `changes needed`, stop and put the findings to the user.
38
+ - Do not review the change yourself in place of the reviewer, and do not argue a blocking finding away. If you think a finding is wrong, say so to the user and let them decide.
39
+ 8. **Record the outcome.**
37
40
 
38
41
  | Outcome | Task file | Log entry |
39
42
  | --- | --- | --- |
40
- | Every requirement met or accepted, and `check <ID>` says `can finish: yes` | `status: done`, `finished:` today | `task \| <ID> done`, then one body line on what was verified |
43
+ | Every requirement met or accepted, review passed where required, and `check <ID>` says `can finish: yes` | `status: done`, `finished:` today | `task \| <ID> done`, then one body line on what was verified and, where it ran, the review verdict |
41
44
  | Requirements met but soft deps open | stays `in-progress` | `task \| <ID> waiting on <ids>` |
42
- | Anything not met, unverified or awaiting the user's call | stays `in-progress`; add what is left to "Notes" | `task \| <ID> blocked: <reason>` |
45
+ | Anything not met, unverified, not reviewed or awaiting the user's call | stays `in-progress`; add what is left to "Notes" | `task \| <ID> blocked: <reason>` |
43
46
 
44
- 8. **Keep the area guide, if the area has one.** `node docs/.sw/sw.mjs explain <ID>` prints `area guide:` with a path or `none`.
45
- - A guide exists: add the report's `Guide:` lines to it, one line per fact under Layout, Patterns, Verify or Gotchas. Replace a line the new fact corrects, and keep the page under 60 lines.
47
+ 9. **Keep the area guide, if the area has one.** `node docs/.sw/sw.mjs explain <ID>` prints `area guide:` with a path or `none`.
48
+ - A guide exists: add the reports' `Guide:` lines to it, one line per fact under Layout, Patterns, Verify or Gotchas. Replace a line the new fact corrects, and keep the page under 60 lines.
46
49
  - No guide: do nothing. A guide is worth starting once several tasks in an area have needed the same facts; if the user asks for one, create `docs/wiki/guide-<area, lowercase>.md` from `docs/.sw/templates/guide.md` and list it in `index.md`.
47
- 9. **File what else was learned.** If the implementer reported a decision or constraint the wiki should hold, offer to save it as a wiki page (`type: decision` or `concept`) and add it to `index.md`. If the task fixed a problem whose cause is now known, offer a `type: lesson` page (Symptom, Cause, Fix, How to notice it earlier); sw-triage finds these later. If it reported follow-up work, offer to create the tasks. These are separate offers: act on each only when the user says yes to that one.
48
- 10. **Report** to the user: outcome, each requirement with its evidence, anything that differs from the task, files changed, and which tasks this unblocked (`node docs/.sw/sw.mjs ready`). Commit only if the user asks.
50
+ 10. **File what else was learned.** These are separate offers: act on each only when the user says yes to that one.
51
+ - A report held a decision or constraint the wiki should keep: offer a wiki page (`type: decision` or `concept`), added to `index.md`.
52
+ - The task fixed a problem whose cause is now known, or the review caught a defect worth remembering: offer a `type: lesson` page (Symptom, Cause, Fix, How to notice it earlier); sw-triage finds these later.
53
+ - A report named follow-up work: offer to create the tasks.
54
+ 11. **Report** to the user, in this order. Commit only if the user asks.
55
+ - the outcome;
56
+ - each requirement with its evidence, and anything that differs from the task;
57
+ - the review verdict and its findings;
58
+ - the files changed;
59
+ - the tasks this unblocked (`node docs/.sw/sw.mjs ready`);
60
+ - what the task cost: run `node docs/.sw/sw.mjs stats` and show its table as printed;
61
+ - one last line: the task is recorded, so the next task is cheapest in a new session.
62
+
63
+ ## Dispatching
64
+
65
+ The same table serves both roles: `sw-implementer` with `implementer.md`, `sw-reviewer` with `reviewer.md`, model from `models.implement` or `models.review`.
66
+
67
+ | Tool | How |
68
+ | --- | --- |
69
+ | Claude Code | the agent by name. If it is not among your agent types, use a general-purpose agent, tell it to read `<skill-dir>/../sw-config/assets/<role file>` first and follow it, and pass the model from `docs/.sw/config.json` (`models.<role>.claude`) if set |
70
+ | Codex | spawn the custom agent `sw_implementer` or `sw_reviewer` |
71
+ | Copilot CLI | `task` tool with the agent name |
72
+ | No subagents available, the agent is not defined, or the project's rules forbid subagents | follow the role file yourself, in this session, and tell the user the configured model and the clean context were not used. A review done this way is weaker: say so |
49
73
 
50
74
  ## Common mistakes
51
75
 
52
76
  - Marking `done` because the implementer said so. Done means every requirement has evidence.
53
77
  - Accepting a `differs` item on the user's behalf. A sensible alternative is still not what the task asked for.
78
+ - Skipping the review on a task that requires it, or doing it yourself in the same context that judged the implementation.
54
79
  - Starting work before the task file says `in-progress`. If the session dies, nobody knows the task was touched.
55
- - Letting the implementer edit the task file or the log. One writer for status: you.
56
- - Reading the plan or the code "to follow along". The implementer already paid for that.
80
+ - Letting a subagent edit the task file or the log. One writer for status: you.
81
+ - Reading the plan or the code "to follow along". The subagents already paid for that.
82
+ - Running the next task in the same session out of momentum.