planrails 0.7.0 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,58 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.8.0 — 2026-10-04
4
+
5
+ Plans that run unattended, compact between tasks, and end instead of lingering.
6
+ Every change below came from a test run or a real project's plans.
7
+
8
+ - **An optional Claude Code plugin lets the model compact its own session
9
+ between tasks.** `claude plugin marketplace add vivmagarwal/planrails`, then
10
+ `claude plugin install planrails@planrails`. It gives the model two tools,
11
+ `context_usage` and `compact_after_turn`, and decides nothing: when a task is
12
+ recorded as done it puts the context reading (and the growth since the session
13
+ started or last compacted) beside the result, and step 6 of the plan's loop has
14
+ the model decide. The compaction runs once the turn ends, waits while a
15
+ sub-agent is still going (ask again later), keeps the session id, then resumes with "Continue the plan.";
16
+ a failure is reported to the model as text. Measured in CourseGen Lab on an
17
+ 11-task plan run unattended on Opus, two pairs of runs: compactions only at
18
+ recorded boundaries; inside each plugin run its compactions saved about 30% and
19
+ 50% of its bill; across pairs the plugin runs read 20% and 50% fewer tokens and
20
+ cost 1–32% less (runs vary: the two controls differed by 45%); blind reviews of
21
+ both pairs found no quality loss attributable to it.
22
+ Without the plugin, or headless, nothing changes. Installed per user (or per
23
+ project) from this repo's marketplace into `~/.claude/plugins/`; it adds
24
+ nothing measurable to a normal chat's start (3.3 s against 3.0 s). If an
25
+ `allowedMcpServers` list hides its tools, it stays idle (an earlier build
26
+ stalled every session start by ~8 s), and `npx planrails init` and the planner
27
+ warn and offer three choices: one entry in `~/.claude/settings.json`, the same
28
+ in a project's `.claude/settings.local.json`, or removing the list, with the
29
+ count of MCP servers that would also turn on. The plugin uses Claude Code's mods
30
+ API, which its own types call early access.
31
+ - **Go on until the plan stops you.** Nothing told an unattended session to start
32
+ the next task. The loop now does, works around a blocked task, and lists when
33
+ to stop.
34
+ - **Plans that only wait on the owner stop reloading.** A real project's three
35
+ active plans were 97–99% done, held open by one or two owner rows each, and
36
+ reloaded ~80,000 tokens into every call for days. New status `waiting`: §5
37
+ hands such a plan over (its reload line out, one line under "Waiting on you"
38
+ in `CLAUDE.local.md`); the checker notes an active plan in that state.
39
+ - **Only the plans being worked reload.** The planner reports what reloads and
40
+ its size, and clears it before adding a line, with the person's yes.
41
+ - **A plan is one feature of ~10–30 tasks**, ended at a phase boundary and
42
+ continued in a new plan that carries every Rule, Decision and Learning
43
+ forward. The size note says so when finished rows are most of the plan (the
44
+ 0.6 note asked for a trim the method forbade, and was ignored for weeks), and
45
+ no longer counts the loop block every plan carries.
46
+ - **Slices a stranger can pick up cold**: each row names its files and `after
47
+ Tn`; a proof fails until its task is done (a project-wide check goes after the
48
+ task's own test); irreversible steps are owner rows; Context is the living
49
+ "what exists"; before compacting, what the next task needs goes in the plan.
50
+ - **The stamp is taken in the same command as the proof.** Typed timestamps were
51
+ the slip real plans recorded most.
52
+
53
+ To update: `npx planrails@latest init`. Active plans keep working; to give one
54
+ the new loop, replace its "How to work this plan" block with the template's.
55
+
3
56
  ## 0.7.0 — 2026-10-03
4
57
 
5
58
  The reload line moves from `CLAUDE.md` to `CLAUDE.local.md`, which git ignores.
package/PLANNER.md CHANGED
@@ -1,4 +1,4 @@
1
- <!-- planrails 0.7.0 -->
1
+ <!-- planrails 0.8.0 -->
2
2
  # The Planner
3
3
 
4
4
  You are about to plan a piece of work with a person, then help execute it so the
@@ -11,7 +11,7 @@ project runs Node — copy in one small checker.
11
11
  **How a person starts you.** They say something like *"Follow
12
12
  `.project-management/planrails/PLANNER.md` and tell me when you are ready to plan
13
13
  the next feature with me."* When they do, run **§1 Get ready**, give the
14
- eight-line report, and stop.
14
+ nine-line report, and stop.
15
15
  Do not start planning until they answer.
16
16
 
17
17
  ---
@@ -63,14 +63,17 @@ tell you what the repo already says.
63
63
  and one line of what it holds.
64
64
  3. **Read `.project-management/`.** Read in full the plan this session is
65
65
  assigned — the person names it; if one plan is active, that one — and list
66
- the others by id, with who holds each (NOW's `session:` line).
66
+ the others by id, with who holds each (NOW's `session:` line). Note what
67
+ reloads: each bare `@` plan line in `CLAUDE.local.md` and `CLAUDE.md`, and
68
+ its size (`wc -w`); and what waits on the person (the "Waiting on you" lines),
69
+ to remind them.
67
70
  4. **Read the last ~20 commits** (`git log --oneline -20`) for how the code moves.
68
71
 
69
72
  **A sub-agent's finding is a lead, not a fact.** If it names a file and line, you
70
73
  can check it. If it is only a number, re-derive it yourself or drop it. A wrong
71
74
  number in a plan becomes a wrong decision later.
72
75
 
73
- Then **report ready in eight lines** and stop:
76
+ Then **report ready in nine lines** and stop:
74
77
 
75
78
  ```
76
79
  READY — <project name>
@@ -78,6 +81,7 @@ Stack: <language, framework, package manager>
78
81
  Always-on: <CLAUDE.md | AGENTS.md | README.md>
79
82
  Check: <the check command> Test: <the test command>
80
83
  Plans: <existing plan ids; which are active and who holds each, or "none">
84
+ Reloading: <the plans on @ lines, with their words; or "none">
81
85
  Docs: <relevant doc paths, or "none">
82
86
  Touches: <the folders this feature will change>
83
87
  Questions: <up to 4 things the repo did not answer, or "none">
@@ -107,12 +111,21 @@ templates at the end of this file (§ Template — PLAN.md, § Template — LOG.
107
111
  Pick a short kebab-case `<id>` (`weekly-digest`). Then:
108
112
 
109
113
  - **Write for a senior engineer.** Decisions and context, not obvious steps.
110
- - **Name the proof for every task before any work.** The proof is the command
111
- that shows the task is done: a test, a build, a check. A task that no command
114
+ - **Name the proof for every task before any work.** The proof is a command
115
+ that would fail if this task were not done. A project-wide check (`npm run
116
+ check`) shows only that nothing broke; when you use it, put the task's own test
117
+ first in the same cell: `npx vitest run tests/digest.test.ts && npm run check`.
118
+ Anything that cannot be undone or reaches beyond the repo (a deploy, a publish,
119
+ a push, a migration on shared data, a message to people) is its own task with
120
+ proof `owner`. A task that no command
112
121
  can prove is proven by the person's word — write `owner` in the proof column
113
122
  and record their words and the date in evidence when they give it (the checker
114
123
  wants the date there). The proof cell is exactly one `command`, or `owner`.
115
- - **A task is one sitting's work with one proof.** Longer than that is two tasks.
124
+ - **A task is a slice a stranger can pick up cold.** Its row says what changes,
125
+ the files, and `after T3` when it needs an earlier task; the why is in
126
+ Decisions, the facts in Context. It fits in one stretch between compactions:
127
+ if it means reading more than a handful of files or changing two areas, it is
128
+ two tasks. One task, one proof.
116
129
  Task ids are permanent: never renumber, and a task you will not do keeps its
117
130
  row with status `dropped`.
118
131
  - **A plan with phases ends each phase with a task** whose proof is the Done-when
@@ -123,6 +136,26 @@ Pick a short kebab-case `<id>` (`weekly-digest`). Then:
123
136
  word is paid for again and again. Trim prose before Learnings or Decisions.
124
137
  History goes in LOG.md, not here. Past the line, the checker prints a `note:` —
125
138
  advice, not a failure: the run still exits 0, and you decide what to trim.
139
+ - **A plan is one feature: about 10–30 tasks.** Finished rows stay in the plan
140
+ and reload with it, so a plan that keeps growing keeps costing. When the work
141
+ outgrows that, end the plan at a phase boundary (§5) and go on in a new plan.
142
+ Ending a plan must not lose what it knows: the new plan carries forward every
143
+ Rule, Decision and Learning that still holds, and its Context gets a short
144
+ "What exists" (what the last phase built, where) and a link to the old plan,
145
+ whose rows stay on disk to read when a task needs the detail. What stops
146
+ reloading is the finished rows' narrative, which is most of a big plan's
147
+ words and little of its guidance. An evidence cell holds the stamp, the exit
148
+ code and the last line; the story goes in LOG.md.
149
+ - **Clear what reloads before you add to it.** Every plan on a bare `@` line
150
+ in `CLAUDE.local.md` or `CLAUDE.md` rides on every call of every session in
151
+ this checkout, whether that session works it or not. A session needs only the
152
+ plan it works: the hold check reads the others from disk. So, for each plan
153
+ that reloads: finished, close it (§5); only the owner can move it, hand it over
154
+ (§5); nobody works it now, pause it (`status: paused`, its line out); held by
155
+ a live session (§4's check), leave it. A line for someone else's plan belongs
156
+ in their own `CLAUDE.local.md`; one in the shared `CLAUDE.md` predates 0.7.
157
+ Say what you found, and change a plan only with the person's yes. The aim:
158
+ only the plans being worked right now reload.
126
159
  - **Add the reload line.** In the project-root `CLAUDE.local.md` (create it if
127
160
  needed), under a short "Active plans" heading, add:
128
161
  ```
@@ -146,7 +179,7 @@ Pick a short kebab-case `<id>` (`weekly-digest`). Then:
146
179
  - **One plan per session, one session per plan.** A repo may hold several active
147
180
  plans, each with its own reload line and each held by one session, named on its
148
181
  NOW `session:` line. Every session in this checkout reloads all of them, so keep them few; a plan
149
- nobody is working is paused (`status: paused`, its line in backticks). A session
182
+ nobody is working is paused (`status: paused`, its line out). A session
150
183
  works only the plan the person in that session assigned; the other reloaded
151
184
  plans are context, not work orders. Naming the session after the plan
152
185
  (`claude -n <id>`) shows the holder in every listing and in the terminal title;
@@ -167,8 +200,21 @@ Pick a short kebab-case `<id>` (`weekly-digest`). Then:
167
200
  holds has no reload line, if its NOW lost its RESUME line, or if NOW points at
168
201
  a finished task. If the project is not Node,
169
202
  skip this; the plan still works, and you enforce the gate yourself.
203
+ - **Check the compaction tools, if the planrails plugin is installed.** If you
204
+ lack `context_usage` although `claude plugin list` shows `planrails`, an
205
+ `allowedMcpServers` list in the settings hides them (the plugin then stays
206
+ idle and costs nothing). Tell the person, and offer the choices
207
+ (`npx planrails init` prints them too): `{ "serverName": "planrails" }` in
208
+ `allowedMcpServers` in `~/.claude/settings.json` (every project, every other
209
+ server still off); the same entry in `.claude/settings.local.json` (this
210
+ project only; make sure git ignores it); or removing the list, which also turns
211
+ on every MCP server it keeps off (count them in `~/.claude.json` and
212
+ `.mcp.json` first). Do what they pick; it applies from the next session. A
213
+ managed policy that locks the list (`allowManagedMcpServersOnly`) cannot be
214
+ widened: say who can, and carry on without.
170
215
  - **Check the plan before the first task.** You, or a fresh sub-agent with no
171
- chat context: open every path the plan names, start every proof command,
216
+ chat context: open every path the plan names, start every proof command and
217
+ see it fail on work not yet built,
172
218
  confirm the reload line loads and the checker passes, and log what you changed.
173
219
  A plan can name a file that does not exist, or a proof that proves nothing;
174
220
  ten minutes here saves an hour later.
@@ -238,8 +284,10 @@ The loop for each task:
238
284
  2. **Do the work.** Fix the cause, not the symptom. The simplest change that
239
285
  works, end to end.
240
286
  3. **Run the proof.** Right now, not from memory. Copy the exit code and the last
241
- line of output. Take the time from the shell clock (`date`; on Windows
242
- `Get-Date`), never from memory.
287
+ line of output. Take the stamp in the same command, so it cannot be typed:
288
+ `npx vitest run x > /tmp/proof.out 2>&1; echo "exit $?"; date "+%Y-%m-%d %H:%M"; tail -1 /tmp/proof.out`
289
+ (on Windows, `Get-Date -Format "yyyy-MM-dd HH:mm"`). A stamp typed from memory
290
+ is the slip real plans recorded most.
243
291
  4. **Paste the evidence.** Into the task's evidence cell:
244
292
  `2026-09-12 14:20 · exit 0 · "6 passed"`. Exit 0, or the task is not done.
245
293
  5. **Set it done.** Only now. A proof of `owner` is closed only by the person's
@@ -254,12 +302,48 @@ The loop for each task:
254
302
  "trap → rule", with the real case. LOG.md holds what happened; Learnings holds
255
303
  the rule, because Learnings reloads every session and the log does not.
256
304
 
305
+ **Go on until the plan stops you.** When the person asked you to work the plan,
306
+ start the next task without ending your turn to report. A blocked task does not
307
+ end the run: take the next one that does not wait on it. Stop only when every
308
+ row is done or dropped (close, §5), only owner rows and the close are open (hand
309
+ over, §5), every open task is blocked or waits on one, the hold check says stop,
310
+ or something breaks that you cannot diagnose. Write NOW first, then report.
311
+
257
312
  **Update NOW before you end any turn.** NOW is the first thing a fresh session
258
313
  reads, so write it for a stranger. On a long task, note the sub-step in NOW at
259
314
  each checkpoint. If a task cannot proceed, set it `blocked`, put the reason in its
260
315
  evidence cell, and say so in NOW. If NOW is stale, the next session repeats your
261
316
  work or starts in the wrong place.
262
317
 
318
+ **Compact between tasks, if you have the tools.** The planrails plugin gives you
319
+ `context_usage` and `compact_after_turn` (`mcp__planrails__…`; Claude Code may
320
+ list them as deferred: load them with ToolSearch); without them, skip this, and Claude
321
+ Code compacts on its own when the context fills. Every call re-reads the whole
322
+ context, so a big one costs on every step that follows. When a task is done and
323
+ recorded (evidence pasted, NOW updated, LOG appended, and in the plan whatever
324
+ the next task needs: a person's words in Decisions with the date, a trap in
325
+ Learnings, what you built in Context), call `context_usage` and decide:
326
+
327
+ - Compact when the context has grown past about 50,000 tokens since the session
328
+ started or last compacted (`context_usage` reports it; a real project starts a
329
+ session at tens of thousands, which no compaction can drop), or past the line
330
+ the plan's Rules set, and tasks remain. The saving is roughly the tokens dropped
331
+ times the calls still to come.
332
+ - Do not compact mid-task, near the end of the plan, or when the next task works
333
+ in the files you just read: re-reading them costs more than keeping them.
334
+
335
+ To compact, first write into the plan what the next task needs: what you built,
336
+ in Context; a trap, in Learnings; a person's words, in Decisions. The summary is
337
+ rewritten at every compaction, so a fact that lives only there is lost at the
338
+ next one; in a real run, facts left in the instructions ("not yet written down
339
+ anywhere") were exactly that. Then call `compact_after_turn` with
340
+ `instructions` naming the plan and the next task, and `resume: "Continue the
341
+ plan."`, and end your turn. It compacts once the turn
342
+ ends (it refuses while a sub-agent is still going: ask again once it is done),
343
+ and the plan reloads with your session's
344
+ claim intact. If `context_usage` reports the last compaction failed, carry on
345
+ without it.
346
+
263
347
  **A red test starts an investigation**, not an edit: is the product wrong, the
264
348
  test stale, or the environment wrong? Decide which before changing anything, and
265
349
  never weaken an assertion to get green.
@@ -293,13 +377,25 @@ cannot see the screen. You can.
293
377
  - Gate green, work wrong: refuse to close it. Say what the gate cannot see.
294
378
  - A green check is not proof the work is good. It is proof one command exited 0.
295
379
 
296
- **Never edit a plan the way a script would.** If something crashes, log it, set
297
- the task `blocked` with the reason, and stop. Do not hand-fix state to look done.
380
+ **Never edit a plan the way a script would.** If something crashes, log it and set
381
+ the task `blocked` with the reason; go on only with tasks that do not wait on it. Do not hand-fix state to look done.
298
382
 
299
383
  ---
300
384
 
301
385
  ## §5 Close
302
386
 
387
+ **When only the owner can move the plan.** If every open row is an `owner` row
388
+ or the close itself, the work is done and only the person's word is missing. Do
389
+ not leave the plan `active`: it reloads into every session for nothing, and real
390
+ plans have sat that way for days. Run steps 1–3 below now, then hand it over: set
391
+ `status: waiting`, NOW's `session:` to `none` and RESUME to the owner rows; take
392
+ its reload line out of `CLAUDE.local.md`, and add one line there under a
393
+ "Waiting on you" heading: `<id> — T7: the owner tries the export and says it
394
+ works (.project-management/plans/<id>/PLAN.md)`. Tell the person. When they give
395
+ their word, a session sets the plan `active`, puts its reload line back, records
396
+ their words and the date, re-runs every proof, and closes by steps 4–5. The checker notes an active plan
397
+ in this state.
398
+
303
399
  1. **Re-run every proof.** A plan closes on what the checks say now, not on their
304
400
  last recorded run. Paste fresh evidence.
305
401
  2. **Fresh-context review.** A sub-agent, or a new session, with the plan and
@@ -352,8 +448,11 @@ Each line here was paid for by a real failure in earlier planning systems:
352
448
  stamped two hours behind its own log. NOW is four lines; LOG.md is the history.
353
449
  - **The checker's notes advise; they never block.** A script can verify an exit
354
450
  code, so that is a problem and fails the build. It can only suspect that a plan
355
- is too long, so that is a note, and you judge. Four of four real plans had
356
- outgrown the word line with nothing saying so. The check command already runs
451
+ is too long, that only its owner can move it, or that a finished plan still
452
+ reloads, so those are notes, and you judge. Four of four real plans had
453
+ outgrown the word line with nothing saying so; three of three active plans in
454
+ one project were held open by owner rows. A note names an action the method
455
+ allows, or it is ignored. The check command already runs
357
456
  before every commit and its output is already read, so a reminder there needs
358
457
  no hook. A note needs a real failure behind it, or it is noise.
359
458
  - **Sub-agent findings are leads** because four spot-checked findings were each
@@ -401,14 +500,15 @@ updated: <YYYY-MM-DD HH:MM, from the shell clock>
401
500
  session: <none, or the session working this plan: name · the first 8 characters of its session id · since YYYY-MM-DD HH:MM>
402
501
 
403
502
  ## How to work this plan
404
- Read NOW, then Rules and Learnings; do not re-read Context. One task at a time:
503
+ Read NOW, then Rules and Learnings; open a file Context lists only when your task touches it. One task at a time:
405
504
  1. Check who holds the plan (below); set it `doing`; point NOW at it; put your session on its `session:` line.
406
505
  2. Do the work: fix causes, not symptoms; the simplest change that works end to end.
407
- 3. Run the proof now. Paste `YYYY-MM-DD HH:MM · exit N · "last line"` into evidence. Every stamp comes from the shell clock (`date`; on Windows `Get-Date`), never typed from memory.
506
+ 3. Run the proof now, with the stamp taken in the same command (`cmd > /tmp/proof.out 2>&1; echo "exit $?"; date "+%Y-%m-%d %H:%M"; tail -1 /tmp/proof.out`), never typed: `date`; in PowerShell `$LASTEXITCODE` and `Get-Date`. Paste `YYYY-MM-DD HH:MM · exit N · "last line"` into evidence, and only that; the story goes in LOG.md.
408
507
  4. `done` only if N is 0. A proof of `owner` is closed only by the person's words with the date, never by you: until then the row is `blocked`, the reason in its evidence cell, and RESUME says so. Point NOW at the next task; append an entry to LOG.md: did, files, proof, next, learned.
409
- 5. If it fought back, add a Learning: the trap, then the rule. A verified fact goes in Context, a choice in Decisions.
508
+ 5. If it fought back, add a Learning: the trap, then the rule. A verified fact, or what this task built that a later one uses, goes in Context (path — what it is); a choice, or a person's words with the date, in Decisions.
509
+ 6. If you have the planrails plugin's tools (`mcp__planrails__context_usage`; load it with ToolSearch if it is listed as deferred), call `context_usage` now: grown past ~50,000 tokens since the session started or last compacted, with tasks left and the next one in other files, first write into the plan what the next task needs (what you built, in Context; a trap, in Learnings), then call `compact_after_turn` (instructions: the plan and the next task; resume: "Continue the plan.") and end your turn. Otherwise go on to the next task without stopping to report; a blocked task does not end the run: take the next one that does not wait on it. Stop when no task can go ahead, the plan is ready to close or hand over, or the hold check says stop; write NOW first.
410
510
 
411
- NOW is exactly these four lines, RESUME, NEXT, updated and session, written for a stranger; on a long task note the sub-step in RESUME, and update NOW before any turn ends. Blocked: say so in RESUME, reason in the evidence cell. A task you will not do is `dropped`; its row stays. After a compaction, `git status --short` shows the in-flight work. Edit this file with the editor tool; an unquoted shell string eats backticks. A self-contained task may go to a sub-agent briefed with its row, Rules, Decisions, Learnings and Context; it writes to a named file and reports a few lines, which are leads; you run the proof before pasting evidence. If the project runs Node, `node .project-management/planrails/check-plans.mjs` must pass. When no row is left open, done or dropped, close by PLANNER.md §5: proofs re-run, a fresh-context review, `status: done`, `session: none`, the reload line deleted from CLAUDE.local.md. Full method: `.project-management/planrails/PLANNER.md` §4.
511
+ NOW is exactly these four lines, RESUME, NEXT, updated and session, written for a stranger; on a long task note the sub-step in RESUME, and update NOW before any turn ends. Blocked: say so in RESUME, reason in the evidence cell. A task you will not do is `dropped`; its row stays. After a compaction, `git status --short` and the last LOG.md entry show where you are. Edit this file with the editor tool; an unquoted shell string eats backticks. A self-contained task may go to a sub-agent briefed with its row, Rules, Decisions, Learnings and Context; it writes to a named file and reports a few lines, which are leads; you run the proof before pasting evidence. If the project runs Node, `node .project-management/planrails/check-plans.mjs` must pass. When no row is left open, done or dropped, close by PLANNER.md §5: proofs re-run, a fresh-context review, `status: done`, `session: none`, the reload line deleted from CLAUDE.local.md. When only owner rows and the close are open, hand it over by §5: `status: waiting`, `session: none`, its reload line out, one line under "Waiting on you" in CLAUDE.local.md. Full method: `.project-management/planrails/PLANNER.md` §4.
412
512
 
413
513
  Who holds the plan. `session:` names the one session working it: `name · first 8 characters of its session id · since YYYY-MM-DD HH:MM`. Match by id, never by name: the line's 8 characters start a `sessionId` in the listing. Before any task goes `doing`, even when asked to continue: re-read NOW from disk, run `claude agents --json` (your own id is `$CLAUDE_CODE_SESSION_ID`) and `git status --short`, then:
414
514
  - the id is yours: go on.
@@ -430,9 +530,10 @@ Work only the plan assigned in this session; other reloaded plans are context. W
430
530
  ## Tasks
431
531
  | id | task | status | proof | evidence |
432
532
  |----|------|--------|-------|----------|
433
- | T1 | <what to do> (<path>) | todo | `<command that proves it>` | |
434
- | T2 | <what to do> (<path>) | todo | `<command>` | |
435
- | T3 | update the docs this work changed | todo | owner | |
533
+ | T1 | <what changes> (<paths>) | todo | `<a command that fails until T1 is done>` | |
534
+ | T2 | <what changes> (<paths>; after T1) | todo | `<its own test> && <the project check>` | |
535
+ | T3 | the owner tries it on staging and says it works | todo | owner | |
536
+ | T4 | close the plan by PLANNER.md §5 (after T1–T3) | todo | `<the Done-when checks>` | |
436
537
 
437
538
  ## Rules for this plan
438
539
  - <a rule that governs this area — e.g. "every email goes through lib/email, never a bare send">
@@ -445,8 +546,9 @@ Work only the plan assigned in this session; other reloaded plans are context. W
445
546
  ## Learnings
446
547
  - <the trap you hit> → <the rule that avoids it> (<the real case, one line>)
447
548
 
448
- ## Context (read during planning — do not re-read)
549
+ ## Context (facts and what exists; open a listed file only when your task touches it)
449
550
  - <path> — <one line of what it holds; mark a sub-agent's unverified finding as a lead>
551
+ - <path> — <what a task built here that a later task uses: the function, the table, the route>
450
552
  ````
451
553
 
452
554
  ## Template — LOG.md
package/README.md CHANGED
@@ -58,7 +58,7 @@ Then, two steps:
58
58
  > Follow `.project-management/planrails/PLANNER.md` and tell me when you are
59
59
  > ready to plan the next feature with me.
60
60
 
61
- It reads your repo, reports what it found in eight lines, and waits. Then you
61
+ It reads your repo, reports what it found in nine lines, and waits. Then you
62
62
  plan together, and it writes the plan and wires the reload line.
63
63
 
64
64
  2. **Wire the checker into your build.** Add this to the command you run before
@@ -103,6 +103,73 @@ the same flow in any project that has run `npx planrails init` — the skill rea
103
103
  the project's own copy of `PLANNER.md`, so every project follows the version it
104
104
  has. (It is not called `/plan`, because Claude Code has a `/plan` of its own.)
105
105
 
106
+ ### Optional: let the model compact between tasks (Claude Code plugin)
107
+
108
+ Every model call re-reads the whole conversation, so a long plan pays for its
109
+ early tasks again on every later step. Claude Code compacts on its own only when
110
+ the context is nearly full, wherever that falls, mid-task included. The
111
+ planrails plugin lets the model do it at a better moment: right after a task is
112
+ done and recorded, when the plan already holds everything that matters.
113
+
114
+ ```bash
115
+ claude plugin marketplace add vivmagarwal/planrails
116
+ claude plugin install planrails@planrails
117
+ ```
118
+
119
+ It adds two tools and decides nothing: `context_usage` (how full the context is,
120
+ how much it has grown since the session started or last compacted, how many
121
+ sub-agents run, how the last compaction went) and `compact_after_turn` (compact
122
+ once this turn ends, then go on with "Continue the plan."). When a task is
123
+ recorded as done in the plan, by whatever tool, the plugin puts that reading
124
+ beside the result, and step 6 of the plan's loop has the model decide: after
125
+ growth of about 50,000 tokens, with tasks left and the next one in other files,
126
+ it writes what the next task needs into the plan and compacts. The compaction
127
+ refuses while a sub-agent is still going, keeps the session's id (so the plan's `session:`
128
+ claim holds), and the plan reloads afterwards. Without the plugin, or in a
129
+ headless `claude -p` run, nothing changes.
130
+
131
+ Measured on a real app (CourseGen Lab, an 11-task plan, unattended Opus
132
+ sessions, two pairs of runs with and without the plugin): the sessions with it
133
+ compacted only at recorded task boundaries; each run's own compactions saved
134
+ about 30% and 50% of its bill, against the same run replayed without them. Across
135
+ the pairs the plugin runs read 20% and 50% fewer tokens and cost 1–32% less, a
136
+ wide range because runs vary: the two runs without the plugin differed by 45%.
137
+ Blind reviews of both pairs found no quality loss attributable to compaction
138
+ (one pair favoured each side, on the same bug). The runs, the method and the
139
+ numbers are in `.project-management/research/auto-compact/`; the final wording
140
+ of step 6 came after them.
141
+
142
+ **Where it lives.** Not in the npm package: `claude plugin install` copies it
143
+ into `~/.claude/plugins/cache/planrails/` and switches it on in the scope you
144
+ choose (`--scope user`, the default: every project, for you; `project`: the
145
+ shared `.claude/settings.json`, for the whole team; `local`: this project, for
146
+ you). `claude plugin uninstall planrails@planrails` removes it.
147
+
148
+ **What it costs a normal chat.** Nothing measurable. Sessions with it started as
149
+ fast as without (median 3.3 s against 3.0 s, three runs each). Its two tools
150
+ are listed by name only (Claude Code defers MCP tools until a model loads one),
151
+ and it does nothing unless a plan under `.project-management/plans/` gains a
152
+ done row or the model calls one of its tools. It never compacts on its own.
153
+
154
+ **If an `allowedMcpServers` list hides it.** Such a list keeps every MCP server
155
+ it does not name off, the plugin's included; the plugin notices, stays idle and
156
+ costs nothing (an earlier build stalled every session start by ~8 s here).
157
+ `npx planrails init` and the planner both say so, and offer three choices,
158
+ each applied from the next session:
159
+
160
+ | Choice | Effect |
161
+ |---|---|
162
+ | add `{ "serverName": "planrails" }` to `allowedMcpServers` in `~/.claude/settings.json` | every project; every other server stays off |
163
+ | the same entry in a project's `.claude/settings.local.json` | that project only, for you |
164
+ | remove the `allowedMcpServers` line | also turns on every MCP server it was keeping off (`init` counts them) |
165
+
166
+ If managed settings set `allowManagedMcpServersOnly`, only whoever manages
167
+ Claude Code can add the entry. Without the tools, planrails works as before.
168
+
169
+ The plugin uses Claude Code's mods API, which its own types call early access:
170
+ if a Claude Code update changes it, the tools go missing or report an error,
171
+ and plans run as they do without the plugin.
172
+
106
173
  ## How a plan works
107
174
 
108
175
  Each plan is two files: `PLAN.md` (the map and tracker) and `LOG.md` (append-only
@@ -172,13 +239,24 @@ structural check, biased toward catching a faked "done":
172
239
 
173
240
  The checker also **advises without blocking**. What a script can only suspect, it
174
241
  prints as a `note:` before its final line; the run still exits 0, and the agent
175
- that ran the check decides. There is one today: an active `PLAN.md` over ~3,000
176
- words, because the plan reloads into every session and nothing else says when it
177
- has grown. Reminders ride on the check your agent already runs before every
178
- commit, so they need no hook.
242
+ that ran the check decides. Each one names an action the method allows:
243
+
244
+ - an active `PLAN.md` over ~3,000 words. When most of the words are finished
245
+ task rows, trimming prose cannot help, so the note says to end the plan at a
246
+ phase boundary and go on in a new one that carries the Rules, Decisions and
247
+ Learnings forward;
248
+ - an active plan that only the owner can move: every open row is an `owner`
249
+ row or the close. It says to hand the plan over (`status: waiting`, its reload
250
+ line out, one line under "Waiting on you" in `CLAUDE.local.md`), so it stops
251
+ reloading until the owner's word reopens it.
252
+
253
+ Both come from a real project's plans: three active plans there were 97–99%
254
+ done, held open by one or two owner rows each, and reloaded about 59,000 words
255
+ into every session. Reminders ride on the check your agent already runs before
256
+ every commit, so they need no hook.
179
257
 
180
258
  ```
181
- check-plans: note: weekly-digest: PLAN.md is 3,588 words, over the ~3,000 line — it reloads into every session; move history to LOG.md and trim prose before Learnings or Decisions
259
+ check-plans: note: weekly-digest: PLAN.md is 3,588 words besides its loop block, over the ~3,000 line — it reloads into every session; move history to LOG.md and trim prose before Learnings or Decisions
182
260
  check-plans: 1 plan(s) ok — every completion claim has a proof and exit 0 evidence, this session's plan reloads, NOW is current
183
261
  ```
184
262
 
@@ -256,7 +334,10 @@ session working a plan and has every session check for a live holder before it
256
334
  touches the plan. 0.6.0 lets the checker advise without blocking: a `note:` for
257
335
  what a script can only suspect, starting with a plan that has outgrown its word
258
336
  line. 0.7.0 moves the reload line to `CLAUDE.local.md`, so a plan reloads
259
- only for whoever works it and a team's `CLAUDE.md` is left alone. See
337
+ only for whoever works it and a team's `CLAUDE.md` is left alone. 0.8.0 adds an
338
+ optional Claude Code plugin that lets the model compact its session between
339
+ tasks, a `waiting` state for plans only the owner can move, and plans that end
340
+ at phase boundaries without losing what they know. See
260
341
  [`CHANGELOG.md`](CHANGELOG.md).
261
342
 
262
343
  MIT.
package/bin/planrails.mjs CHANGED
@@ -18,6 +18,7 @@
18
18
  */
19
19
  import { readFileSync, copyFileSync, mkdirSync, existsSync, writeFileSync, realpathSync, readdirSync } from "node:fs";
20
20
  import { join, dirname, resolve } from "node:path";
21
+ import { homedir } from "node:os";
21
22
  import { fileURLToPath } from "node:url";
22
23
  import { checkPlans, isActive, formatNotes, reloadLines } from "../tools/check-plans.mjs";
23
24
 
@@ -49,6 +50,41 @@ export function stampOf(text) {
49
50
  }
50
51
  const newer = (a, b) => { const [x, y] = [a, b].map((v) => v.split(".").map(Number)); for (let i = 0; i < 3; i++) if (x[i] !== y[i]) return x[i] > y[i]; return false; };
51
52
 
53
+ /**
54
+ * Whether Claude Code's settings hide the planrails plugin's tools. Reads the
55
+ * settings files only (user, project, local, managed); writes nothing. Returns
56
+ * null when the plugin is not installed or the tools are allowed, else what hides
57
+ * them and the choices, as lines for init to print.
58
+ */
59
+ export function pluginToolsCheck(project, { configDir = process.env.CLAUDE_CONFIG_DIR || join(homedir(), ".claude"), platform = process.platform, managedPath } = {}) {
60
+ const read = (p) => { try { return JSON.parse(readFileSync(p, "utf8")); } catch { return null; } };
61
+ const managed = managedPath ?? ({ darwin: "/Library/Application Support/ClaudeCode/managed-settings.json", win32: "C:\\Program Files\\ClaudeCode\\managed-settings.json" }[platform] || "/etc/claude-code/managed-settings.json");
62
+ const sources = [
63
+ ["~/.claude/settings.json", join(configDir, "settings.json")],
64
+ [".claude/settings.json", join(project, ".claude", "settings.json")],
65
+ [".claude/settings.local.json", join(project, ".claude", "settings.local.json")],
66
+ ["managed settings", managed],
67
+ ].map(([label, path]) => ({ label, s: read(path) })).filter((x) => x.s);
68
+ const installed = sources.some(({ s }) => Object.entries(s.enabledPlugins || {}).some(([k, v]) => k.startsWith("planrails@") && v));
69
+ if (!installed) return null;
70
+ const lists = sources.filter(({ s }) => Array.isArray(s.allowedMcpServers));
71
+ if (!lists.length) return null;
72
+ const opens = (x) => x?.serverName === "planrails" || x?.serverUrl !== undefined || x?.serverCommand !== undefined;
73
+ const policy = sources.find((x) => x.label === "managed settings")?.s;
74
+ const locked = policy?.allowManagedMcpServersOnly === true;
75
+ const counted = locked ? [policy.allowedMcpServers || []] : lists.map(({ s }) => s.allowedMcpServers);
76
+ if (counted.some((l) => l.some(opens))) return null;
77
+ const where = lists.map((x) => x.label).join(" and ");
78
+ if (locked) return [`the planrails plugin is installed, but your organisation's managed settings lock allowedMcpServers (allowManagedMcpServersOnly), so its tools stay hidden and the model cannot compact between tasks. Ask whoever manages Claude Code to add { "serverName": "planrails" }. planrails works without it; the plugin stays idle.`];
79
+ const servers = Object.keys(read(existsSync(join(configDir, ".claude.json")) ? join(configDir, ".claude.json") : join(dirname(configDir), ".claude.json"))?.mcpServers || {}).length + Object.keys(read(join(project, ".mcp.json"))?.mcpServers || {}).length;
80
+ return [
81
+ `the planrails plugin is installed, but allowedMcpServers in ${where} hides its tools, so the model cannot compact between tasks. planrails works without them; the plugin stays idle and costs nothing. To turn them on, pick one; it applies from the next session:`,
82
+ ` a) every project, every other server still off: add { "serverName": "planrails" } to allowedMcpServers in ~/.claude/settings.json`,
83
+ ` b) this project only: add the same entry to .claude/settings.local.json (keep it out of git)`,
84
+ ` c) remove the allowedMcpServers line${servers ? `: it also turns on the ${servers} MCP server(s) configured in ~/.claude.json and .mcp.json that it now keeps off` : ""}`,
85
+ ];
86
+ }
87
+
52
88
  function init(args) {
53
89
  const target = dirArg(args);
54
90
  const force = args.includes("--force");
@@ -91,6 +127,8 @@ function init(args) {
91
127
  const claudeMd = join(target, "CLAUDE.md");
92
128
  const shared = existsSync(claudeMd) ? [...reloadLines(readFileSync(claudeMd, "utf8"))] : [];
93
129
  if (shared.length) console.log(` ! CLAUDE.md reloads ${shared.join(", ")}: since 0.7 the reload line goes in CLAUDE.local.md, which git ignores, so a plan reloads only for whoever works it and CLAUDE.md stays the team's. The line still works where it is; to move it, cut it from CLAUDE.md, paste it into CLAUDE.local.md, and add CLAUDE.local.md to .gitignore`);
130
+ const tools = pluginToolsCheck(target, process.env.PLANRAILS_MANAGED_SETTINGS ? { managedPath: process.env.PLANRAILS_MANAGED_SETTINGS } : {});
131
+ if (tools) console.log(` ! ${tools.join("\n ")}`);
94
132
  const loose = ["PLANNER.md", "check-plans.mjs"].filter((f) => existsSync(join(pm, f)));
95
133
  if (loose.length) console.log(` ! 0.2.x files at .project-management/ root: ${loose.join(", ")} — the copies now live in planrails/; delete the loose ones and point your check command at .project-management/planrails/check-plans.mjs`);
96
134
 
@@ -102,7 +140,10 @@ Next:
102
140
  node .project-management/planrails/check-plans.mjs
103
141
  3. When the agent writes a plan, it adds one line to CLAUDE.local.md, and that file
104
142
  to .gitignore, so the plan reloads after every compaction, for you alone:
105
- @.project-management/plans/<id>/PLAN.md`);
143
+ @.project-management/plans/<id>/PLAN.md
144
+ 4. Optional, in Claude Code: let the model compact its session between tasks.
145
+ claude plugin marketplace add vivmagarwal/planrails
146
+ claude plugin install planrails@planrails`);
106
147
  return 0;
107
148
  }
108
149
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "planrails",
3
- "version": "0.7.0",
3
+ "version": "0.8.0",
4
4
  "description": "Plans that survive a lost session, and \"done\" that means done. A planner prompt and a tiny checker.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -1,5 +1,5 @@
1
1
  #!/usr/bin/env node
2
- // planrails 0.7.0
2
+ // planrails 0.8.0
3
3
  /**
4
4
  * check-plans — the machine-checked rails of the planner.
5
5
  *
@@ -16,12 +16,15 @@
16
16
  * may not point at a plan that does not exist; a CLAUDE.local.md line is not
17
17
  * judged, because that file is untracked and outlives a branch switch.
18
18
  * - NOW must be current: an active plan with a NOW section keeps its RESUME
19
- * line, and that line may not name only finished tasks. A retired plan says `status: done` (or paused) and is exempt
20
- * from both; a plan with no status line counts as active.
19
+ * line, and that line may not name only finished tasks. A plan that says
20
+ * `status: done`, `paused` or `waiting` is exempt from both; a plan with no
21
+ * status line counts as active.
21
22
  *
22
- * It also prints notes: what a script can only suspect (an active PLAN.md over
23
- * ~3,000 words, which reloads into every session). A note goes to stdout before
24
- * the final line and never changes the exit code; whoever ran the check decides.
23
+ * It also prints notes: what a script can only suspect. An active PLAN.md over
24
+ * ~3,000 words of its own (the loop block not counted); an active plan only the
25
+ * owner can move; a plan that is no longer active but still reloads. A note goes
26
+ * to stdout before the final line and never changes the exit code; whoever ran
27
+ * the check decides.
25
28
  *
26
29
  * The rule is biased toward catching a faked "done": a task counts as a
27
30
  * completion claim UNLESS its status is blank or an explicit not-done word
@@ -254,7 +257,7 @@ export function checkPlan({ id, text, verify = false, run = null }) {
254
257
  * finished word exempts a plan from the reload and NOW rules. A plan with no
255
258
  * status line counts as active — the safe direction for a gate.
256
259
  */
257
- const NOT_ACTIVE = new Set(["done", "paused", "retired", "closed", "shipped", "finished", "complete", "completed", "archived", "dropped", "cancelled", "canceled", "abandoned", "superseded", "onhold", "hold"]);
260
+ const NOT_ACTIVE = new Set(["done", "paused", "retired", "closed", "shipped", "finished", "complete", "completed", "archived", "dropped", "cancelled", "canceled", "abandoned", "superseded", "onhold", "hold", "waiting"]);
258
261
  export function isActive(text) {
259
262
  const m = text.match(/^status:\s*([^\s·|]+)/im);
260
263
  return !(m && NOT_ACTIVE.has(norm(m[1])));
@@ -323,11 +326,33 @@ export function runProof(cmd, root, timeoutMs) {
323
326
  */
324
327
  export const WORD_LINE = 3000;
325
328
  const commas = (n) => String(n).replace(/\B(?=(\d{3})+$)/g, ","); // no Intl: a Node built without it would drop the comma
329
+ const countWords = (s) => s.split(/\s+/).filter(Boolean).length;
330
+ // A row no one will finish: not a completion claim, and not waiting to be done.
331
+ const SET_ASIDE = new Set(["dropped", "cancelled", "canceled", "wontfix", "wontdo", "abandoned", "skip", "skipped", "na"]);
332
+ const isOpen = (t) => (EMPTY.test(t.status) || NOT_DONE.has(norm(t.status))) && !SET_ASIDE.has(norm(t.status));
333
+ const isOwnerRow = (t) => norm(t.proof) === "owner";
334
+ // The plan's own close: "close by PLANNER.md §5", "close the plan", "retire the plan" — not "close the dialog".
335
+ const isCloseRow = (t) => /§\s*5|\bclose (?:by|per) planner\b|\b(?:close|retire) (?:the|this) plan\b/i.test(t.task);
326
336
  export function planNotes({ id, text }) {
327
337
  const notes = [];
328
338
  if (!isActive(text)) return notes;
329
- const words = text.split(/\s+/).filter(Boolean).length;
330
- if (words > WORD_LINE) notes.push(`${id}: PLAN.md is ${commas(words)} words, over the ~${commas(WORD_LINE)} line — it reloads into every session; move history to LOG.md and trim prose before Learnings or Decisions`);
339
+ // The "How to work this plan" block is the method's, the same in every plan:
340
+ // count the plan's own words, which its author can trim.
341
+ const block = text.match(/^## How to work this plan[\s\S]*?(?=^## |(?![\s\S]))/m);
342
+ const words = countWords((block ? text.replace(block[0], "") : text).replace(/\|/g, " "));
343
+ const { tasks } = parseTasks(text);
344
+ if (words > WORD_LINE) {
345
+ // When the task rows are most of the plan, trimming prose cannot help, and the
346
+ // method keeps every row: ending the plan is the way to stop reloading them.
347
+ const rows = tasks.reduce((n, t) => n + countWords(`${t.id} ${t.task} ${t.status} ${t.proof} ${t.evidence}`), 0);
348
+ notes.push(rows * 2 > words
349
+ ? `${id}: PLAN.md is ${commas(words)} words besides its loop block, over the ~${commas(WORD_LINE)} line, and ${commas(rows)} are its task rows — it reloads into every session; end it at a phase boundary and go on in a new plan (its rows stay on disk as the record), and keep each evidence cell to the stamp, the exit code and the last line`
350
+ : `${id}: PLAN.md is ${commas(words)} words besides its loop block, over the ~${commas(WORD_LINE)} line — it reloads into every session; move history to LOG.md and trim prose before Learnings or Decisions`);
351
+ }
352
+ const open = tasks.filter(isOpen);
353
+ const owner = open.filter(isOwnerRow);
354
+ if (owner.length && open.every((t) => isOwnerRow(t) || isCloseRow(t)))
355
+ notes.push(`${id}: only the owner can move it now (${owner.map((t) => t.id).join(", ")}) — hand it over by PLANNER.md §5: status: waiting, session: none, its reload line out, and one line naming those rows under "Waiting on you" in CLAUDE.local.md; it stops reloading until their word reopens it`);
331
356
  return notes;
332
357
  }
333
358
  /** The notes as printed: one "note:" line each, newline-terminated, or "" when there are none. */
@@ -356,6 +381,8 @@ export function checkPlans({ root = ".", verify = false, verifyTimeoutMs = 10 *
356
381
  try { text = readFileSync(path, "utf8"); } catch (e) { problems.push(`${id}: cannot read ${path} (${e.code || e.message})`); continue; }
357
382
  problems.push(...checkPlan({ id, text, verify, run }));
358
383
  notes.push(...planNotes({ id, text }));
384
+ if (!isActive(text) && loaded.has(id))
385
+ notes.push(`${id}: it is no longer active but still reloads (a bare @ line in CLAUDE.local.md or CLAUDE.md) — take the line out, or wrap it in backticks; it costs every call of every session here`);
359
386
  const holder = holderId(text);
360
387
  if (me && holder && me.startsWith(holder) && isActive(text) && !loaded.has(id))
361
388
  problems.push(`${id}: this session holds the plan but nothing reloads it here — add "@.project-management/plans/${id}/PLAN.md" on its own line, outside backticks, to CLAUDE.local.md, or the plan will not survive a compaction`);