superwiki 0.1.3 → 0.1.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "sw",
3
3
  "description": "An LLM-maintained wiki and task tracker in docs/ for coding agents. Obsidian friendly, with a static viewer.",
4
- "version": "0.1.3",
4
+ "version": "0.1.4",
5
5
  "license": "MIT",
6
6
  "keywords": [
7
7
  "wiki",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "sw",
3
- "version": "0.1.3",
3
+ "version": "0.1.4",
4
4
  "description": "An LLM-maintained wiki and task tracker in docs/ for coding agents. Obsidian friendly, with a static viewer.",
5
5
  "license": "MIT",
6
6
  "skills": "./skills/",
package/README.md CHANGED
@@ -17,9 +17,13 @@ Then, in a project: `/sw-init`.
17
17
 
18
18
  ## Why
19
19
 
20
- Superwiki follows the LLM Wiki pattern described by Andrej Karpathy: raw sources you curate, a wiki the agent owns, and a short schema that tells the agent how to maintain it. On top of that it adds what a software project needs: tasks with dependencies, plans, and a record of decisions and lessons.
20
+ Superwiki follows the LLM Wiki pattern described by Andrej Karpathy: raw sources you curate, a wiki the agent owns, and a short schema that tells the agent how to maintain it. On top of that it adds what a software project needs: tasks with dependencies, plans, reviews, and a record of decisions and lessons.
21
21
 
22
- It is built to be cheap for the agent. One small index to read, one file per task, and a script that answers "what is ready?", "what blocks this?" or "is anything broken?" without the agent reading the vault.
22
+ It is built to be cheap for the agent:
23
+
24
+ - **Little to read.** One small index, one file per task, and a script that answers "what is ready?", "what blocks this?" or "is anything broken?" without the agent reading the vault.
25
+ - **Work in clean contexts.** Planning, implementing and reviewing run in subagents, each on the model you choose for it. The main session only keeps the task's status true, so it stays small.
26
+ - **Cost you can see.** `sw-stats` shows what a session used, per agent.
23
27
 
24
28
  Measured on a real project with 165 tasks, converted from a single markdown index:
25
29
 
@@ -29,7 +33,7 @@ Measured on a real project with 165 tasks, converted from a single markdown inde
29
33
  | Read to start one task | the index, then the task's section | one file, 2 KB at the median |
30
34
  | Marking a task done | a status cell, plus a ✅ at every reference to it (median 12 places) | one frontmatter line |
31
35
 
32
- > Status: early. The CLI, the viewer, `sw-init` and the migration script are tested, and the planning flow has been run in all three agents; some skills have only been exercised once. [DESIGN.md](DESIGN.md) lists what has and has not been proven.
36
+ > Status: early. The CLI, the viewer, `sw-init` and the migration script are tested. Planning, implementing and reviewing have been run end to end on real projects in Claude Code, and the planning flow in Codex and Copilot CLI; some skills have only been exercised once. [DESIGN.md](DESIGN.md) lists what has and has not been proven.
33
37
 
34
38
  ## What you get
35
39
 
@@ -84,9 +88,13 @@ git clone https://github.com/mhmtsrfglu/superwiki ~/.superwiki
84
88
 
85
89
  A linked install follows the clone: `git pull` updates every agent. `install.sh --copy` copies instead, `--uninstall` removes.
86
90
 
87
- All three agents below were checked the same way: the agent found the skills, refused to start a task with an unfinished dependency, and ran `sw-plan` end to end with the planner subagent.
91
+ ### Per agent
92
+
93
+ All three agents were checked the same way: the agent found the skills, refused to start a task with an unfinished dependency, and ran `sw-plan` end to end with the planner subagent.
88
94
 
89
- ### Claude Code
95
+ The model for each role (planning, implementing, reviewing) is set with `sw-config`, which writes one agent file per role into the project. The skills dispatch those agents by name.
96
+
97
+ #### Claude Code
90
98
 
91
99
  ```bash
92
100
  npx superwiki install claude
@@ -94,11 +102,11 @@ npx superwiki install claude
94
102
 
95
103
  Invoke with a slash: `/sw-init`, `/sw-plan T-01`.
96
104
 
97
- Claude Code reads `~/.claude/skills/` (and a project's `.claude/skills/`); it does not read `~/.agents/skills/`, so `global` is not enough for it.
98
-
99
- `sw-plan` enters plan mode when the session offers it. A model set with `sw-config` for `claude` applies to the planner and implementer subagents, written to `.claude/agents/`; if those agents are not loaded, the skills fall back to built-in agents with the same model.
105
+ - Claude Code reads `~/.claude/skills/` (and a project's `.claude/skills/`). It does not read `~/.agents/skills/`, so `global` is not enough for it.
106
+ - `sw-plan` enters plan mode when the session offers it.
107
+ - Agent files: `.claude/agents/sw-planner.md`, `sw-implementer.md` and `sw-reviewer.md`. They load when a session starts; in a session that began before they existed, the skills fall back to built-in agents with the same model.
100
108
 
101
- ### Codex CLI
109
+ #### Codex CLI
102
110
 
103
111
  ```bash
104
112
  npx superwiki install codex
@@ -106,9 +114,10 @@ npx superwiki install codex
106
114
 
107
115
  Invoke with a dollar sign, or by name in a sentence: `$sw-init`, `$sw-plan T-01`, "use the sw-plan skill for T-01". Checked with CLI 0.153.
108
116
 
109
- A skill cannot switch Codex into plan mode; start planning yourself with `/plan` if you want the mode, or let `sw-plan` proceed without it (it changes no file before you approve). A model set with `sw-config` for `codex` is written to `.codex/agents/sw-planner.toml` and `sw-implementer.toml`, and the skills spawn those agents by name. Subagents must be enabled (they are by default in current releases).
117
+ - A skill cannot switch Codex into plan mode. Start planning yourself with `/plan` if you want the mode, or let `sw-plan` proceed without it: it changes no file before you approve.
118
+ - Agent files: `.codex/agents/sw-planner.toml`, `sw-implementer.toml` and `sw-reviewer.toml`. Subagents must be enabled; they are by default in current releases.
110
119
 
111
- ### GitHub Copilot CLI
120
+ #### GitHub Copilot CLI
112
121
 
113
122
  ```bash
114
123
  npx superwiki install copilot
@@ -116,11 +125,11 @@ npx superwiki install copilot
116
125
 
117
126
  Invoke with a slash, or by name in a sentence: `/sw-init`, "use the sw-plan skill for T-01". Checked with CLI 1.0.31.
118
127
 
119
- Copilot CLI also reads `~/.agents/skills/`, so if you installed `codex` or `global` it already has the skills.
128
+ - Copilot CLI also reads `~/.agents/skills/`, so if you installed `codex` or `global` it already has the skills.
129
+ - A skill cannot switch Copilot into plan mode. Start with `copilot --mode plan` or `/plan` if you want it.
130
+ - Agent files: `.github/agents/sw-planner.agent.md`, `sw-implementer.agent.md` and `sw-reviewer.agent.md`, dispatched with the `task` tool. Whether Copilot honours the `model:` field of those files has not been checked.
120
131
 
121
- A skill cannot switch Copilot into plan mode; start with `copilot --mode plan` or `/plan` if you want it. A model set with `sw-config` for `copilot` is written to `.github/agents/sw-planner.agent.md` and `sw-implementer.agent.md`; the skills dispatch them with the `task` tool. Whether Copilot honours the `model:` field of those files has not been checked.
122
-
123
- ### Other agents
132
+ #### Other agents
124
133
 
125
134
  ```bash
126
135
  npx superwiki install global
@@ -128,8 +137,9 @@ npx superwiki install global
128
137
 
129
138
  Agents that load `SKILL.md` folders from `~/.agents/skills` pick the skills up from there. For an agent with its own skills folder (Cursor, Gemini CLI, OpenCode and others), copy the `skills/sw-*` folders from a clone into it by hand. Nothing has been run in these agents. What will differ:
130
139
 
131
- - the skills name Claude Code, Codex and Copilot tools when they dispatch subagents; elsewhere they fall back to doing the planning or implementing in the main session, and say so;
132
- - `sw-config` writes agent files only for `claude`, `codex` and `copilot`, so a per-role model cannot be set.
140
+ - the skills name Claude Code, Codex and Copilot tools when they dispatch subagents; elsewhere they fall back to doing the work in the main session, and say so;
141
+ - `sw-config` writes agent files only for `claude`, `codex` and `copilot`, so a per-role model cannot be set;
142
+ - `sw-stats` reads the session records of those three tools only.
133
143
 
134
144
  Everything else (the vault, the CLI, the viewer, ingest, lint, explain, triage) depends only on Node and on the agent following the skill text.
135
145
 
@@ -156,7 +166,7 @@ A project's `docs/` folder is plain markdown and keeps working as an Obsidian va
156
166
  | Skill | What it does |
157
167
  | --- | --- |
158
168
  | `sw-init` | set up `docs/` in the current project, or upgrade it |
159
- | `sw-migrate` | convert an existing table-based task index, on a new git branch |
169
+ | `sw-migrate` | convert an existing table-based task index, on a git branch of its own |
160
170
  | `sw-ingest` | file a source into the wiki |
161
171
  | `sw-plan` | plan a task with the planner subagent and get your approval |
162
172
  | `sw-implement` | run a task with the implementer subagent, have it reviewed if the task asks for that, and record the result |
@@ -164,6 +174,7 @@ A project's `docs/` folder is plain markdown and keeps working as an Obsidian va
164
174
  | `sw-triage` | for a problem: seen before? lessons, likely causes |
165
175
  | `sw-lint` | structural checks by script, semantic review on request |
166
176
  | `sw-visualize` | open the viewer |
177
+ | `sw-stats` | what the current session has cost: tokens, context, steps and tool calls, per agent |
167
178
  | `sw-config` | the model each tool uses for planning, implementing and reviewing; task areas |
168
179
 
169
180
  ### Examples
@@ -191,6 +202,7 @@ Shown as typed in Claude Code. In Codex, write `$sw-plan` instead of `/sw-plan`.
191
202
  /sw-config plan with opus, implement with sonnet, review with opus
192
203
  /sw-lint check links, frontmatter and task dependencies
193
204
  /sw-visualize open the task board and the wiki in the browser
205
+ /sw-stats what this session has cost so far, per agent
194
206
  ```
195
207
 
196
208
  You do not have to type a command. The rules `sw-init` adds to `AGENTS.md` tell the agent which skill fits, so a plain request should reach the same skill:
@@ -204,6 +216,22 @@ Users get the magic-link email twice. Have we seen this before?
204
216
 
205
217
  A filled-in example vault is in [examples/demo/docs](examples/demo/docs).
206
218
 
219
+ ### What a session cost
220
+
221
+ `sw-stats` reads the record your agent keeps of the session and prints one row for the main session and one for each subagent. It writes nothing. `sw-implement` ends its report with the same table. This one is a real task: planned, implemented and reviewed in 37 minutes.
222
+
223
+ ```text
224
+ session claude fe4cbd6c-8abf-4db6-9755-469dc7321dc7 2026-10-05 10:27 to 11:04, 37 min
225
+ agent model steps first peak sent cached output tools min
226
+ main claude-opus-5-5 27 79k 126k 2.8M 96% 18k 23 37
227
+ sw-planner claude-opus-5-5 31 59k 154k 3.5M 96% 8k 32 6
228
+ sw-implementer claude-sonnet-5-5 68 59k 252k 12.2M 96% 16k 76 26
229
+ sw-reviewer claude-opus-5-5 23 60k 137k 2.4M 89% 372 24 12
230
+ total 149 - - 20.9M 95% 43k 155
231
+ ```
232
+
233
+ `first` and `peak` are the tokens sent with one request. `sent` is that, summed over every step: each step sends the whole context again, which is why a long session in one context is expensive. In Copilot CLI the token columns fill in once the session has closed.
234
+
207
235
  ### The CLI
208
236
 
209
237
  The skills call a small script that answers questions without the agent reading the vault. You can run it yourself, from the project root:
@@ -211,11 +239,12 @@ The skills call a small script that answers questions without the agent reading
211
239
  ```bash
212
240
  node docs/.sw/sw.mjs status # counts per area
213
241
  node docs/.sw/sw.mjs ready # tasks that can start now
214
- node docs/.sw/sw.mjs check P-15 # can it start or finish, and what is open
242
+ node docs/.sw/sw.mjs check P-15 # can it start or finish, what is open, is a review required
215
243
  node docs/.sw/sw.mjs explain P-15 # dependencies, what it unblocks, plan, area guide
216
244
  node docs/.sw/sw.mjs search sync timeout # where something is mentioned
217
245
  node docs/.sw/sw.mjs next-id P # next free id in an area
218
246
  node docs/.sw/sw.mjs lint # broken links, bad frontmatter, dependency errors
247
+ node docs/.sw/sw.mjs stats # tokens, context and steps of the agent session here
219
248
  node docs/.sw/sw.mjs serve --open # the viewer, reading files live
220
249
  node docs/.sw/sw.mjs snapshot # or: freeze the vault into docs/viewer.html, no server
221
250
  ```
@@ -223,14 +252,23 @@ node docs/.sw/sw.mjs snapshot # or: freeze the vault into docs/vie
223
252
  ## Develop
224
253
 
225
254
  ```bash
226
- npm test # builds skills/sw-init/assets/sw.mjs, then runs the tests
255
+ npm test # builds skills/sw-init/assets/sw.mjs and viewer.html, then runs the tests
227
256
  ```
228
257
 
229
- Releases are cut by the `Release` workflow (Actions → Release → Run workflow): it tests, bumps the version in `package.json` and the plugin manifests, publishes to npm, tags, and creates a GitHub release. It publishes through npm trusted publishing, so no token is stored: the package's settings on npmjs.com name this repository and `release.yml` as its trusted publisher.
258
+ Edit sources in `src/`:
259
+
260
+ | File | What it is |
261
+ | --- | --- |
262
+ | `src/core.js` | the vault model, derived task state, lint and search; shared by the CLI and the viewer |
263
+ | `src/stats.js` | reads the agents' session records |
264
+ | `src/cli.js` | the commands |
265
+ | `src/viewer.html` | the viewer |
266
+
267
+ `scripts/build.mjs` bundles them into `skills/sw-init/assets/sw.mjs` and `viewer.html`. Those two files are generated: do not edit them.
230
268
 
231
269
  `node scripts/build-demo.mjs` builds the public demo into `site/` (the viewer with the example vault baked in); the Pages workflow deploys it on every push to `main`.
232
270
 
233
- `src/core.js` is shared by the CLI and the viewer. Edit sources in `src/`; the files in `skills/sw-init/assets/` named `sw.mjs` and `viewer.html` are generated.
271
+ Releases are cut by the `Release` workflow (Actions → Release → Run workflow): it tests, bumps the version in `package.json` and the plugin manifests, publishes to npm, tags, and creates a GitHub release. It publishes with the repository secret `NPM_TOKEN`, an npm access token allowed to publish `superwiki`.
234
272
 
235
273
  ## License
236
274
 
@@ -0,0 +1,5 @@
1
+ ---
2
+ description: Show what this session has cost so far: tokens, context, steps and tool calls per agent
3
+ ---
4
+
5
+ Use the sw-stats skill. User arguments: $ARGUMENTS
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "superwiki",
3
- "version": "0.1.3",
3
+ "version": "0.1.4",
4
4
  "description": "Agent skills that turn docs/ into an LLM-maintained wiki and task tracker. Obsidian-friendly. Works with Claude Code, Codex and Copilot CLI.",
5
5
  "keywords": [
6
6
  "claude-code",
@@ -7,7 +7,7 @@ description: Use when the user wants to implement, build, execute, start or cont
7
7
 
8
8
  Runs one task. You keep the task's status true and judge the result. The work is done by subagents that start from a clean context, on the models set in sw-config: an implementer, and a reviewer when the task asks for one. You do not read the code or the plan: their reports are your input.
9
9
 
10
- That split is what keeps a task cheap. A long session re-sends its whole context on every step; work done in a fresh context does not carry yours, and yours stays small because the work never enters it.
10
+ That split is what keeps a task cheap. A long session sends its whole context again on every step; work done in a fresh context does not carry yours, and yours stays small because the work never enters it.
11
11
 
12
12
  Run commands from the project root. `<skill-dir>` is the directory this SKILL.md is in.
13
13
 
@@ -15,12 +15,16 @@ Run commands from the project root. `<skill-dir>` is the directory this SKILL.md
15
15
 
16
16
  1. **Pick the task.** Id given: use it. Otherwise `node docs/.sw/sw.mjs ready` and let the user choose; tasks already in progress come first.
17
17
  2. **Gate**: `node docs/.sw/sw.mjs check <ID>`.
18
- - `can start: no open deps: ...`: stop. Tell the user which tasks block it and offer to run the first blocker instead; do not run it unasked. Do not start the task anyway, and do not edit `deps` to get past this.
19
- - `can start: n/a, status is in-progress`: this is a continuation; skip step 3.
20
- - `can start: n/a, status is done` or `cancelled`: stop and ask what the user wants.
21
- - `plan: ... (draft, not approved)`: stop; the plan needs the user's approval (sw-plan).
22
- - `plan: none`: fine for a small task (one area, three "Done when" items or fewer, nothing open in its notes, a few files). For anything larger, recommend sw-plan first and let the user choose.
23
- - `review: required (...)`: remember it for step 7.
18
+
19
+ | `check` says | Do |
20
+ | --- | --- |
21
+ | `can start: no open deps: ...` | stop. Tell the user which tasks block it and offer to run the first blocker instead; do not run it unasked. Do not start the task anyway, and do not edit `deps` to get past this |
22
+ | `can start: n/a, status is in-progress` | this is a continuation; skip step 3 |
23
+ | `can start: n/a, status is done` or `cancelled` | stop and ask what the user wants |
24
+ | `plan: ... (draft, not approved)` | stop; the plan needs the user's approval (sw-plan) |
25
+ | `plan: none` | fine for a small task: one area, three "Done when" items or fewer, nothing open in its notes, a few files. For anything larger, recommend sw-plan first and let the user choose |
26
+ | `review: required (...)` | remember it for step 7 |
27
+
24
28
  3. **Mark it started** before any work: in the frontmatter of `docs/tasks/<ID>.md` set `status: in-progress` and `started:` today. Append `## [date] task | <ID> started` to `docs/log.md`, in the layout its last entries use.
25
29
  4. **Checks that need the environment.** If the task has a plan, look for `needs:` in it: `grep -n 'needs:' docs/plans/<ID>-plan.md`. Each hit is a check that starts services or changes data. Ask the user which of them may run; without a yes, none.
26
30
  5. **Dispatch the implementer** (how: "Dispatching" below). Its prompt is the task id, the project root if it is not your working directory, and which `needs:` checks it may run. Do not paste the plan into the prompt; it reads the files.
@@ -43,8 +47,18 @@ Run commands from the project root. `<skill-dir>` is the directory this SKILL.md
43
47
  9. **Keep the area guide, if the area has one.** `node docs/.sw/sw.mjs explain <ID>` prints `area guide:` with a path or `none`.
44
48
  - A guide exists: add the reports' `Guide:` lines to it, one line per fact under Layout, Patterns, Verify or Gotchas. Replace a line the new fact corrects, and keep the page under 60 lines.
45
49
  - No guide: do nothing. A guide is worth starting once several tasks in an area have needed the same facts; if the user asks for one, create `docs/wiki/guide-<area, lowercase>.md` from `docs/.sw/templates/guide.md` and list it in `index.md`.
46
- 10. **File what else was learned.** If a report held a decision or constraint the wiki should keep, offer to save it as a wiki page (`type: decision` or `concept`) and add it to `index.md`. If the task fixed a problem whose cause is now known, or the review caught a defect worth remembering, offer a `type: lesson` page (Symptom, Cause, Fix, How to notice it earlier); sw-triage finds these later. If a report named follow-up work, offer to create the tasks. These are separate offers: act on each only when the user says yes to that one.
47
- 11. **Report** to the user: outcome, each requirement with its evidence, the review verdict and findings, anything that differs from the task, files changed, and which tasks this unblocked (`node docs/.sw/sw.mjs ready`). Commit only if the user asks. End with one line: the task is recorded, so the next task is cheapest in a new session.
50
+ 10. **File what else was learned.** These are separate offers: act on each only when the user says yes to that one.
51
+ - A report held a decision or constraint the wiki should keep: offer a wiki page (`type: decision` or `concept`), added to `index.md`.
52
+ - The task fixed a problem whose cause is now known, or the review caught a defect worth remembering: offer a `type: lesson` page (Symptom, Cause, Fix, How to notice it earlier); sw-triage finds these later.
53
+ - A report named follow-up work: offer to create the tasks.
54
+ 11. **Report** to the user, in this order. Commit only if the user asks.
55
+ - the outcome;
56
+ - each requirement with its evidence, and anything that differs from the task;
57
+ - the review verdict and its findings;
58
+ - the files changed;
59
+ - the tasks this unblocked (`node docs/.sw/sw.mjs ready`);
60
+ - what the task cost: run `node docs/.sw/sw.mjs stats` and show its table as printed;
61
+ - one last line: the task is recorded, so the next task is cheapest in a new session.
48
62
 
49
63
  ## Dispatching
50
64
 
@@ -1,5 +1,11 @@
1
1
  #!/usr/bin/env node
2
- // Generated by scripts/build.mjs from src/core.js and src/cli.js. Do not edit.
2
+ // Generated by scripts/build.mjs from src/core.js, src/stats.js, src/cli.js. Do not edit.
3
+ import { existsSync, mkdirSync, readFileSync, readdirSync, realpathSync, statSync, writeFileSync } from 'node:fs';
4
+ import { homedir } from 'node:os';
5
+ import { basename, dirname, join, resolve as resolvePath, sep } from 'node:path';
6
+ import { spawn } from 'node:child_process';
7
+ import { createServer } from 'node:http';
8
+ import { fileURLToPath } from 'node:url';
3
9
  // Superwiki core: vault model, derived task state and lint. Pure: no fs, no DOM.
4
10
  // Runs in Node (docs/.sw/sw.mjs) and inlined in the viewer, so both report the same findings.
5
11
 
@@ -316,13 +322,368 @@ export function guideFor(vault, area) {
316
322
  return vault.pages.find(p => p.folder === 'wiki' && p.data.type === 'guide' && key(p.data.area ?? '') === key(area)) || null;
317
323
  }
318
324
 
325
+ // Session statistics: what the agent session working in this project has cost so far.
326
+ // Claude Code, Codex and Copilot CLI each keep a record of every session on disk. This finds the
327
+ // record of the project's session and reduces it to one row per agent: the main session and each
328
+ // subagent it started.
329
+
330
+ export const STATS_TOOLS = ['claude', 'codex', 'copilot'];
331
+
332
+ const MAIN = 'main';
333
+ // Codex keeps every project's sessions in one tree; only the most recent files are opened.
334
+ const CODEX_FILES_SCANNED = 100;
335
+
336
+ // ---------- Reading records ----------
337
+
338
+ // A record still being written can end in half a line; lines that do not parse are skipped.
339
+ function jsonLines(path) {
340
+ const records = [];
341
+ for (const line of readFileSync(path, 'utf8').split('\n')) {
342
+ if (!line) continue;
343
+ try {
344
+ records.push(JSON.parse(line));
345
+ } catch {}
346
+ }
347
+ return records;
348
+ }
349
+
350
+ function jsonFile(path) {
351
+ try {
352
+ return JSON.parse(readFileSync(path, 'utf8'));
353
+ } catch {
354
+ return {};
355
+ }
356
+ }
357
+
358
+ const modifiedAt = path => statSync(path).mtimeMs;
359
+
360
+ // The project root as typed and as resolved: tools record the working directory either way.
361
+ function rootForms(root) {
362
+ const forms = new Set([root]);
363
+ try {
364
+ forms.add(realpathSync(root));
365
+ } catch {}
366
+ return [...forms];
367
+ }
368
+
369
+ const isWithin = (roots, dir) => Boolean(dir) && roots.some(root => dir === root || dir.startsWith(root + sep));
370
+
371
+ // ---------- The summary being built ----------
372
+
373
+ const newSession = (tool, id) => ({ tool, id, agents: [], tools: {}, note: '' });
374
+
375
+ // Token fields stay null when the record does not hold them, so "unknown" never prints as 0.
376
+ const newAgent = name => ({
377
+ name, model: '', steps: 0, first: null, peak: null, sent: null, cached: null, output: null, toolCalls: 0, start: null, end: null,
378
+ });
379
+
380
+ const plus = (sum, n) => (sum ?? 0) + (n || 0);
381
+
382
+ // One model request. `context` is everything sent with it, `cached` the part read from the cache.
383
+ function addStep(agent, { context, cached, output }) {
384
+ agent.steps++;
385
+ agent.first ??= context;
386
+ agent.peak = Math.max(agent.peak ?? 0, context);
387
+ agent.sent = plus(agent.sent, context);
388
+ agent.cached = plus(agent.cached, cached);
389
+ agent.output = plus(agent.output, output);
390
+ }
391
+
392
+ function addToolCall(session, agent, name) {
393
+ agent.toolCalls++;
394
+ session.tools[name] = (session.tools[name] || 0) + 1;
395
+ }
396
+
397
+ // Widens the agent's working period to include this record.
398
+ function touch(agent, timestamp) {
399
+ const time = Date.parse(timestamp);
400
+ if (Number.isNaN(time)) return;
401
+ agent.start = Math.min(agent.start ?? time, time);
402
+ agent.end = Math.max(agent.end ?? time, time);
403
+ }
404
+
405
+ // Subagents are listed in the order they were started.
406
+ const byStart = (a, b) => (a.start ?? 0) - (b.start ?? 0);
407
+
408
+ // ---------- Claude Code ----------
409
+ // ~/.claude/projects/<working directory, non-alphanumerics as dashes>/<session>.jsonl, and next to
410
+ // it <session>/subagents/agent-<id>.jsonl with a .meta.json naming the agent type.
411
+
412
+ const claudeHome = () => process.env.CLAUDE_CONFIG_DIR || join(homedir(), '.claude');
413
+
414
+ function claudeSessions(root) {
415
+ const sessions = [];
416
+ for (const form of rootForms(root)) {
417
+ const dir = join(claudeHome(), 'projects', form.replace(/[^A-Za-z0-9]/g, '-'));
418
+ if (!existsSync(dir)) continue;
419
+ for (const name of readdirSync(dir)) {
420
+ if (!name.endsWith('.jsonl')) continue;
421
+ const id = name.slice(0, -'.jsonl'.length);
422
+ const path = join(dir, name);
423
+ sessions.push({
424
+ tool: 'claude',
425
+ id,
426
+ modified: modifiedAt(path),
427
+ current: id === process.env.CLAUDE_CODE_SESSION_ID,
428
+ read: () => readClaude(id, path, join(dir, id, 'subagents')),
429
+ });
430
+ }
431
+ }
432
+ return sessions;
433
+ }
434
+
435
+ function readClaudeAgent(session, name, path) {
436
+ const agent = newAgent(name);
437
+ // A reply is written as one line per content block. Each line repeats the reply's usage, and
438
+ // only the last one has the final output count, so the last line of a reply is the one kept.
439
+ const replies = new Map();
440
+ for (const record of jsonLines(path)) {
441
+ touch(agent, record.timestamp);
442
+ const message = record.type === 'assistant' && record.message;
443
+ if (!message) continue;
444
+ for (const block of message.content || []) {
445
+ if (block.type === 'tool_use') addToolCall(session, agent, block.name);
446
+ }
447
+ if (message.usage) replies.set(message.id, message);
448
+ }
449
+ for (const { model, usage } of replies.values()) {
450
+ agent.model = model || agent.model;
451
+ const cached = usage.cache_read_input_tokens || 0;
452
+ addStep(agent, {
453
+ context: (usage.input_tokens || 0) + cached + (usage.cache_creation_input_tokens || 0),
454
+ cached,
455
+ output: usage.output_tokens,
456
+ });
457
+ }
458
+ return agent;
459
+ }
460
+
461
+ function readClaude(id, path, subagentDir) {
462
+ const session = newSession('claude', id);
463
+ session.agents.push(readClaudeAgent(session, MAIN, path));
464
+ if (!existsSync(subagentDir)) return session;
465
+ const subagents = readdirSync(subagentDir)
466
+ .filter(name => name.endsWith('.jsonl'))
467
+ .map(name => {
468
+ const meta = jsonFile(join(subagentDir, name.replace(/\.jsonl$/, '.meta.json')));
469
+ return readClaudeAgent(session, meta.agentType || 'subagent', join(subagentDir, name));
470
+ });
471
+ session.agents.push(...subagents.sort(byStart));
472
+ return session;
473
+ }
474
+
475
+ // ---------- Codex ----------
476
+ // ~/.codex/sessions/<year>/<month>/<day>/rollout-*.jsonl. The first line is the session's meta:
477
+ // its working directory and, for a subagent, the session that spawned it and its role.
478
+
479
+ const codexHome = () => process.env.CODEX_HOME || join(homedir(), '.codex');
480
+
481
+ function rolloutFiles(dir, found = []) {
482
+ if (!existsSync(dir)) return found;
483
+ for (const entry of readdirSync(dir, { withFileTypes: true })) {
484
+ const path = join(dir, entry.name);
485
+ if (entry.isDirectory()) rolloutFiles(path, found);
486
+ else if (entry.name.endsWith('.jsonl')) found.push({ path, modified: modifiedAt(path) });
487
+ }
488
+ return found;
489
+ }
490
+
491
+ function codexSessions(root) {
492
+ const roots = rootForms(root);
493
+ const recent = rolloutFiles(join(codexHome(), 'sessions')).sort((a, b) => b.modified - a.modified).slice(0, CODEX_FILES_SCANNED);
494
+ const byId = new Map();
495
+ for (const { path, modified } of recent) {
496
+ const records = jsonLines(path);
497
+ const meta = records[0]?.type === 'session_meta' ? records[0].payload : null;
498
+ if (!meta || !isWithin(roots, meta.cwd)) continue;
499
+ const id = meta.session_id || meta.id;
500
+ if (!byId.has(id)) byId.set(id, { modified: 0, files: [] });
501
+ const group = byId.get(id);
502
+ group.modified = Math.max(group.modified, modified);
503
+ group.files.push({ meta, records });
504
+ }
505
+ return [...byId].map(([id, { modified, files }]) => ({ tool: 'codex', id, modified, current: false, read: () => readCodex(id, files) }));
506
+ }
507
+
508
+ function readCodexAgent(session, name, records) {
509
+ const agent = newAgent(name);
510
+ let lastTotal = null;
511
+ for (const record of records) {
512
+ touch(agent, record.timestamp);
513
+ const payload = record.payload || {};
514
+ if (record.type === 'turn_context') agent.model = payload.model || agent.model;
515
+ if (record.type === 'response_item' && /_call$/.test(payload.type || '')) addToolCall(session, agent, payload.name || payload.type);
516
+ if (record.type !== 'event_msg' || payload.type !== 'token_count' || !payload.info) continue;
517
+ // The count is also repeated when only the rate limits change; a new request moves the total.
518
+ const total = payload.info.total_token_usage?.total_tokens;
519
+ if (total === lastTotal) continue;
520
+ lastTotal = total;
521
+ const last = payload.info.last_token_usage || {};
522
+ addStep(agent, { context: last.input_tokens || 0, cached: last.cached_input_tokens, output: last.output_tokens });
523
+ }
524
+ return agent;
525
+ }
526
+
527
+ function readCodex(id, files) {
528
+ const session = newSession('codex', id);
529
+ const subagents = [];
530
+ for (const { meta, records } of files) {
531
+ if (meta.thread_source === 'subagent') subagents.push(readCodexAgent(session, meta.agent_role || meta.agent_nickname || 'subagent', records));
532
+ else session.agents.push(readCodexAgent(session, MAIN, records));
533
+ }
534
+ session.agents.push(...subagents.sort(byStart));
535
+ return session;
536
+ }
537
+
538
+ // ---------- Copilot CLI ----------
539
+ // ~/.copilot/session-state/<session>/events.jsonl, with the working directory in workspace.yaml.
540
+ // Events of a subagent carry its agentId. Token counts are written only when the session closes.
541
+
542
+ const copilotHome = () => join(homedir(), '.copilot');
543
+
544
+ function copilotSessions(root) {
545
+ const roots = rootForms(root);
546
+ const base = join(copilotHome(), 'session-state');
547
+ if (!existsSync(base)) return [];
548
+ const sessions = [];
549
+ for (const id of readdirSync(base)) {
550
+ const events = join(base, id, 'events.jsonl');
551
+ const workspace = join(base, id, 'workspace.yaml');
552
+ if (!existsSync(events) || !existsSync(workspace)) continue;
553
+ const cwd = (readFileSync(workspace, 'utf8').match(/^cwd: (.*)$/m) || [])[1];
554
+ if (!isWithin(roots, cwd)) continue;
555
+ sessions.push({ tool: 'copilot', id, modified: modifiedAt(events), current: false, read: () => readCopilot(id, events) });
556
+ }
557
+ return sessions;
558
+ }
559
+
560
+ function readCopilot(id, path) {
561
+ const session = newSession('copilot', id);
562
+ const agents = new Map();
563
+ const agentOf = key => {
564
+ if (!agents.has(key)) agents.set(key, newAgent(key));
565
+ return agents.get(key);
566
+ };
567
+ agentOf(MAIN);
568
+ let closed = false;
569
+ for (const event of jsonLines(path)) {
570
+ const data = event.data || {};
571
+ const agent = agentOf(event.agentId || MAIN);
572
+ touch(agent, event.timestamp);
573
+ if (event.type === 'subagent.started') {
574
+ agent.name = data.agentName || agent.name;
575
+ } else if (event.type === 'assistant.message') {
576
+ agent.steps++;
577
+ agent.model = data.model || agent.model;
578
+ } else if (event.type === 'tool.execution_start') {
579
+ addToolCall(session, agent, data.toolName);
580
+ } else if (event.type === 'session.shutdown' && data.agentMetrics) {
581
+ // A resumed session closes more than once; each close reports the run that ended with it.
582
+ closed = true;
583
+ for (const [key, metrics] of Object.entries(data.agentMetrics)) {
584
+ const reported = agentOf(key);
585
+ for (const { usage = {} } of Object.values(metrics.modelMetrics || {})) {
586
+ reported.sent = plus(reported.sent, usage.inputTokens);
587
+ reported.cached = plus(reported.cached, usage.cacheReadTokens);
588
+ reported.output = plus(reported.output, usage.outputTokens);
589
+ }
590
+ }
591
+ }
592
+ }
593
+ session.agents = [...agents.values()];
594
+ if (!closed) session.note = 'Copilot writes token counts when the session closes; until then /usage shows them.';
595
+ return session;
596
+ }
597
+
598
+ // ---------- Choosing the session ----------
599
+
600
+ const LISTERS = { claude: claudeSessions, codex: codexSessions, copilot: copilotSessions };
601
+
602
+ // The session to report: the one named by `id` (a prefix is enough), else the session this command
603
+ // runs in when the tool says which one that is, else the most recently written one.
604
+ export function sessionStats(root, { tool, id } = {}) {
605
+ let sessions = (tool ? [tool] : STATS_TOOLS).flatMap(name => LISTERS[name](root));
606
+ if (id) sessions = sessions.filter(session => session.id.startsWith(id));
607
+ sessions.sort((a, b) => b.modified - a.modified);
608
+ const chosen = (!id && sessions.find(session => session.current)) || sessions[0];
609
+ return chosen ? chosen.read() : null;
610
+ }
611
+
612
+ // ---------- Printing ----------
613
+
614
+ const COLUMNS = ['agent', 'model', 'steps', 'first', 'peak', 'sent', 'cached', 'output', 'tools', 'min'];
615
+ const TEXT_COLUMNS = 2; // agent and model align left; the numbers after them align right
616
+ const LEGEND = 'first, peak: tokens sent with one request. sent: that, summed over every step. cached: the share of sent read from the cache.';
617
+
618
+ function tokens(n) {
619
+ if (n == null) return '-';
620
+ if (n < 1000) return String(n);
621
+ if (n < 1e6) return `${Math.round(n / 1000)}k`;
622
+ return `${(n / 1e6).toFixed(1)}M`;
623
+ }
624
+
625
+ const cachedShare = agent => (agent.sent ? `${Math.round((100 * (agent.cached || 0)) / agent.sent)}%` : '-');
626
+ const minutes = (start, end) => Math.round((end - start) / 60000);
627
+
628
+ function clock(time) {
629
+ const d = new Date(time);
630
+ const two = n => String(n).padStart(2, '0');
631
+ return `${d.getFullYear()}-${two(d.getMonth() + 1)}-${two(d.getDate())} ${two(d.getHours())}:${two(d.getMinutes())}`;
632
+ }
633
+
634
+ // "2026-10-05 10:27 to 11:04, 37 min"; the end repeats the date only when it is another day.
635
+ function period(start, end) {
636
+ const [from, to] = [clock(start), clock(end)];
637
+ const sameDay = from.slice(0, 10) === to.slice(0, 10);
638
+ return `${from} to ${sameDay ? to.slice(11) : to}, ${minutes(start, end)} min`;
639
+ }
640
+
641
+ function sessionPeriod(agents) {
642
+ const timed = agents.filter(agent => agent.start != null);
643
+ if (!timed.length) return '';
644
+ return period(Math.min(...timed.map(agent => agent.start)), Math.max(...timed.map(agent => agent.end)));
645
+ }
646
+
647
+ // The sum of the agents. It has no context of its own and no period, so those cells stay empty.
648
+ function totalOf(agents) {
649
+ const total = newAgent('total');
650
+ for (const agent of agents) {
651
+ total.steps += agent.steps;
652
+ total.toolCalls += agent.toolCalls;
653
+ for (const field of ['sent', 'cached', 'output']) {
654
+ if (agent[field] != null) total[field] = plus(total[field], agent[field]);
655
+ }
656
+ }
657
+ return total;
658
+ }
659
+
660
+ const cells = agent => [
661
+ agent.name, agent.model, String(agent.steps), tokens(agent.first), tokens(agent.peak), tokens(agent.sent), cachedShare(agent),
662
+ tokens(agent.output), String(agent.toolCalls), agent.start == null ? '' : String(minutes(agent.start, agent.end)),
663
+ ];
664
+
665
+ function table(rows) {
666
+ const widths = COLUMNS.map((_, column) => Math.max(...rows.map(row => row[column].length)));
667
+ const pad = (cell, column) => (column < TEXT_COLUMNS ? cell.padEnd(widths[column]) : cell.padStart(widths[column]));
668
+ return rows.map(row => row.map(pad).join(' ').trimEnd());
669
+ }
670
+
671
+ export function formatStats(session) {
672
+ const { agents } = session;
673
+ const rows = [COLUMNS, ...agents.map(cells)];
674
+ if (agents.length > 1) rows.push(cells(totalOf(agents)));
675
+ const calls = Object.entries(session.tools).sort((a, b) => b[1] - a[1]).map(([name, count]) => `${name} ${count}`);
676
+ return [
677
+ ['session', session.tool, session.id, sessionPeriod(agents)].filter(Boolean).join(' '),
678
+ ...table(rows),
679
+ `tool calls ${calls.join(' ') || 'none'}`,
680
+ LEGEND,
681
+ ...(session.note ? [session.note] : []),
682
+ ].join('\n');
683
+ }
684
+
319
685
  // Superwiki CLI. Lives in a project at docs/.sw/sw.mjs and prints short answers,
320
686
  // so agents do not have to read the vault to get them.
321
- import { spawn } from 'node:child_process';
322
- import { existsSync, mkdirSync, readFileSync, readdirSync, realpathSync, writeFileSync } from 'node:fs';
323
- import { createServer } from 'node:http';
324
- import { basename, dirname, join, resolve as resolvePath } from 'node:path';
325
- import { fileURLToPath } from 'node:url';
326
687
 
327
688
  const HELP = `sw <command> [--docs <dir>] [--json]
328
689
 
@@ -333,6 +694,8 @@ const HELP = `sw <command> [--docs <dir>] [--json]
333
694
  search <words> pages and log entries that mention the words, best match first
334
695
  next-id <AREA> next free task id for an area (numbers are never reused)
335
696
  lint structural checks; exit code 1 on errors
697
+ stats what the agent session here has cost so far: steps, context and tokens per agent
698
+ [--session <id>] another session of this project [--tool ${STATS_TOOLS.join('|')}]
336
699
  serve [--open] start (or reuse) a local viewer at http://127.0.0.1:<port>/ that reads the files live
337
700
  snapshot write docs/.sw/data.js so docs/viewer.html opens as a file, frozen at this moment`;
338
701
 
@@ -394,8 +757,8 @@ export function vaultData(docs) {
394
757
  export const dataScript = data => `window.SW_DATA = ${JSON.stringify(data).replace(/</g, '\\u003c')};\n`;
395
758
 
396
759
  // ---------- Commands ----------
397
- // Each command gets { docs, vault, args, flags } and returns { data, text, code? }.
398
- // `data` is what --json prints; `text` is the default output.
760
+ // Each command gets { docs, vault, args, flags } and returns { data, text, code? } or
761
+ // { error, code? }. `data` is what --json prints; `text` is the default output.
399
762
 
400
763
  function status({ vault }) {
401
764
  const s = summary(vault);
@@ -557,6 +920,17 @@ function lintCommand({ vault }) {
557
920
  };
558
921
  }
559
922
 
923
+ // Agent sessions are recorded by the folder they ran in: the project root, which holds docs/.
924
+ function stats({ docs, flags }) {
925
+ if (flags.tool && !STATS_TOOLS.includes(flags.tool)) {
926
+ return { error: `usage: sw stats [--session <id>] [--tool ${STATS_TOOLS.join('|')}]` };
927
+ }
928
+ const root = dirname(docs);
929
+ const session = sessionStats(root, { tool: flags.tool, id: flags.session });
930
+ if (!session) return { error: `no ${flags.tool || 'agent'} session record found for ${root}`, code: 1 };
931
+ return { data: session, text: formatStats(session) };
932
+ }
933
+
560
934
  function snapshot({ docs }) {
561
935
  const data = vaultData(docs);
562
936
  mkdirSync(join(docs, '.sw'), { recursive: true });
@@ -664,20 +1038,24 @@ const COMMANDS = {
664
1038
  search: { run: searchCommand, needsVault: true },
665
1039
  'next-id': { run: nextIdCommand, needsVault: true },
666
1040
  lint: { run: lintCommand, needsVault: true },
1041
+ stats: { run: stats, needsVault: false },
667
1042
  snapshot: { run: snapshot, needsVault: false },
668
1043
  serve: { run: serve, needsVault: false },
669
1044
  };
670
1045
 
1046
+ // Flags that are on or off, and flags that take the next argument as their value.
1047
+ const SWITCHES = ['json', 'open', 'foreground'];
1048
+ const OPTIONS = ['docs', 'tool', 'session'];
1049
+
1050
+ // Anything that is not a known flag is positional, so search words may start with dashes.
671
1051
  function parseArgs(argv) {
672
- const flags = { json: false, open: false, foreground: false, docs: null };
1052
+ const flags = Object.fromEntries([...SWITCHES.map(name => [name, false]), ...OPTIONS.map(name => [name, null])]);
673
1053
  const positional = [];
674
1054
  for (let i = 0; i < argv.length; i++) {
675
- const arg = argv[i];
676
- if (arg === '--docs') flags.docs = argv[++i] || '.';
677
- else if (arg === '--json') flags.json = true;
678
- else if (arg === '--open') flags.open = true;
679
- else if (arg === '--foreground') flags.foreground = true;
680
- else positional.push(arg);
1055
+ const name = argv[i].startsWith('--') ? argv[i].slice(2) : null;
1056
+ if (SWITCHES.includes(name)) flags[name] = true;
1057
+ else if (OPTIONS.includes(name)) flags[name] = argv[++i] ?? null;
1058
+ else positional.push(argv[i]);
681
1059
  }
682
1060
  return { command: positional[0], args: positional.slice(1), flags };
683
1061
  }
@@ -0,0 +1,50 @@
1
+ ---
2
+ name: sw-stats
3
+ description: Use when the user asks what the current agent session has cost or used (tokens, context size, steps, tool calls, time, per subagent), wants a session summary or statistics, or invokes sw-stats or sw:stats.
4
+ ---
5
+
6
+ # sw-stats
7
+
8
+ Shows what this session has cost so far: one row for the main session and one for each subagent it started. A script reads the record your tool keeps of the session; you read nothing yourself and write no file.
9
+
10
+ Run from the project root:
11
+
12
+ ```bash
13
+ node docs/.sw/sw.mjs stats
14
+ ```
15
+
16
+ It reports the session it runs in. For another session of this project, add `--session <id or its first characters>`; to look only at one tool's sessions, `--tool claude|codex|copilot`.
17
+
18
+ ## Answer
19
+
20
+ 1. **Show the table as printed**, in a code block. Do not round, reorder or translate it.
21
+ 2. **Add at most three observations**, in the user's language, each one a number from the table and what it means. Pick the ones that apply:
22
+
23
+ | In the table | Say |
24
+ | --- | --- |
25
+ | one agent's `sent` is most of the total | which agent, and its share. That is where the session's cost is |
26
+ | an agent's `peak` is several times its `first` | its context grew during the work; `steps` times a large context is what makes `sent` large |
27
+ | `first` is large for `main` | the session starts heavy before it reads anything: rules files, plugins and tool definitions. The tool's own context command shows what fills it |
28
+ | `cached` is well below the others for one agent | much of what it sent was not served from the cache, which costs more per token. Long pauses do that |
29
+ | `main` has many `steps` or a `peak` far above its `first` | work is running in the main session that a subagent could do in a clean context |
30
+
31
+ 3. Nothing else: no advice the table does not support, and no price. The table counts tokens; what a token costs depends on the user's plan.
32
+
33
+ ## Columns
34
+
35
+ So that you can answer a question about them:
36
+
37
+ - `steps`: model requests. `tools`: tool calls. `min`: minutes between the agent's first and last record.
38
+ - `first`, `peak`: tokens sent with the first request and with the largest one.
39
+ - `sent`: tokens sent, summed over every step. A step re-sends the whole context, so this is far larger than `peak`.
40
+ - `cached`: the share of `sent` that was read from the cache.
41
+ - `output`: tokens the model wrote.
42
+ - `-`: the tool's record does not hold that number.
43
+
44
+ ## If it fails
45
+
46
+ | Output | Do |
47
+ | --- | --- |
48
+ | `unknown command stats` | the project's `docs/.sw/sw.mjs` is older than this skill; offer to run sw-init, which updates it |
49
+ | `no agent session record found` | say so, and name the tool's own command instead: `/cost` or `/context` (Claude Code), `/status` (Codex), `/usage` (Copilot CLI). A session started in a subfolder of the project is recorded under that folder and is not found |
50
+ | a note that Copilot has not written token counts yet | pass the note on; steps and tool calls are still valid |