tiny-spec 0.4.0__tar.gz → 0.5.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
- Metadata-Version: 2.4
1
+ Metadata-Version: 2.5
2
2
  Name: tiny-spec
3
- Version: 0.4.0
3
+ Version: 0.5.0
4
4
  Summary: A tiny, opinionated take on spec-driven development.
5
5
  Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
6
6
  Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
@@ -44,10 +44,24 @@ Description-Content-Type: text/markdown
44
44
  <p align="center">A tiny, opinionated take on spec-driven development.</p>
45
45
 
46
46
  <p align="center">
47
+ <a href="https://pypi.org/project/tiny-spec/"><img src="https://img.shields.io/pypi/v/tiny-spec.svg?color=d97757" alt="PyPI version"></a>
48
+ <a href="https://pypi.org/project/tiny-spec/"><img src="https://img.shields.io/pypi/dm/tiny-spec.svg?color=d97757" alt="PyPI downloads"></a>
47
49
  <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue.svg" alt="License: MIT"></a>
48
50
  <a href="https://docs.claude.com/en/docs/claude-code/overview"><img src="https://img.shields.io/badge/Claude%20Code-skills-d97757.svg" alt="Claude Code"></a>
51
+ <a href="https://github.com/GrayMa77er/tiny-spec/stargazers"><img src="https://img.shields.io/github/stars/GrayMa77er/tiny-spec?color=d97757" alt="Stars"></a>
49
52
  </p>
50
53
 
54
+ <p align="center">
55
+ <code>uvx tiny-spec install</code>
56
+ </p>
57
+
58
+ <!-- DEMO GIF SLOT — images/demo.gif
59
+ Record the FAIL -> fix -> PASS loop, then uncomment:
60
+ <p align="center">
61
+ <img src="images/demo.gif" alt="tiny-spec build loop: the reviewer fails a task, the executor fixes it, the task commits" width="800">
62
+ </p>
63
+ -->
64
+
51
65
  tiny-spec is a four-step workflow for Claude Code that turns a ticket into shipped,
52
66
  reviewed code. You write the intent, it produces a design, a task list, and then
53
67
  builds the work one task at a time. Every task is implemented by one agent and
@@ -57,7 +71,9 @@ committed.
57
71
  That core is **four skills and two agents**. In front of it sit **two optional
58
72
  planning on-ramps** — `tiny-spec-prd` (idea → PRD) and `tiny-spec-breakdown`
59
73
  (PRD → stories) — for when you're starting from an idea rather than a ready ticket.
60
- `tiny-spec-run` walks the three planning steps in one command. No config file, no
74
+ Two optional routers sit over the top: `tiny-spec-run` walks the three planning steps
75
+ in one command, and `tiny-spec-loop` works a whole list of stories — branch, plan,
76
+ build, merge, next — until they're built or something stops it. No config file, no
61
77
  build step.
62
78
 
63
79
  ```
@@ -69,6 +85,10 @@ PLANNING (optional on-ramps) EXECUTION (the core loop, one story at
69
85
  idea → PRD PRD → stories intent design tasks per-task loop
70
86
  PRD.md BREAKDOWN.md SPEC.md PLAN.md + tasks.md plan → implement → review → commit
71
87
  constitution
88
+ ▲ ▲
89
+ └──────────────── tiny-spec-loop ─────────────────┘
90
+ optional: per story — branch, run, build, merge,
91
+ then the next story. Halts on a wall.
72
92
  ```
73
93
 
74
94
  The two on-ramps are **optional** and stack. Have nothing written down? Run
@@ -79,6 +99,97 @@ work? Skip both and start at `tiny-spec-create`. Both on-ramps write a regenerab
79
99
  file at your project root (not under `.spec/`); `tiny-spec-create` then reads the
80
100
  breakdown one story at a time.
81
101
 
102
+ ## How small
103
+
104
+ Every other kit in this space is bigger. That is the whole pitch, so here is the
105
+ receipt rather than the adjective:
106
+
107
+ | | skills / commands | agents | config | artifacts per feature |
108
+ |---|---|---|---|---|
109
+ | **tiny-spec** | **8** (4 core + 4 optional) | **2** | **none** | **`SPEC` `PLAN` `tasks`** |
110
+ | [GitHub Spec Kit](https://github.com/github/spec-kit) | 10 | — | `specify init` | `spec` `plan` `tasks` `checklist` `constitution` `research` `data-model` `contracts/` `quickstart` |
111
+ | [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
112
+ | [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
113
+
114
+ The whole of tiny-spec is **2,079 lines** of markdown across 8 skills and 2 agents.
115
+ There is no config file, no build step, no orchestrator, and no CLI to initialize a
116
+ project — the skills are the product.
117
+
118
+ <sub>Counts taken from each project's own repository on 2026-08-03 by listing its
119
+ command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
120
+ `templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
121
+ OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
122
+ personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
123
+ `customize.toml` files (35). These projects are all substantially more popular than
124
+ tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
125
+ bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
126
+
127
+ ## Why it's small
128
+
129
+ Most spec frameworks are generous by default:
130
+ many phases, many agents, many generated documents. tiny-spec makes the opposite
131
+ bet. Keep one safeguard, drop the rest.
132
+
133
+ A green unit test suite is not the same as working software, so the reviewer
134
+ exercises acceptance criteria end to end and a final smoke test confirms the whole
135
+ spec. That independent review is the safeguard — not the volume of planning
136
+ artifacts. One task, one commit, an external reviewer. Nothing gets added unless
137
+ it earns its place.
138
+
139
+ The case for staying small:
140
+
141
+ - **Documents are context, and context isn't free.** Generating large `spec.md`,
142
+ `plan.md`, `research.md`, and `data-model.md` files costs tokens to write, then
143
+ costs context to carry. Every paragraph the agent has to hold is room it no
144
+ longer has for your actual code. tiny-spec keeps the spine small — a
145
+ constitution and a short memory — and injects only what each task needs.
146
+ - **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
147
+ assume you're bootstrapping a project from a blank page. Day to day, you pick up
148
+ a ticket and change part of a system that already exists. tiny-spec binds to a
149
+ ticket, works one at a time, and references your task platform instead of
150
+ re-describing the world.
151
+ - **Rigid pipelines fight the user.** Mandatory phases and required sections
152
+ impose ceremony on work that doesn't need it. tiny-spec's extra structure is
153
+ optional by design — add shape where it pays, skip it where it doesn't.
154
+ - **More moving parts is more to maintain.** Orchestrators, ownership contracts,
155
+ checkpoint matrices, and config files are themselves a system you have to learn
156
+ and keep in sync. A few small skills and two agents are not.
157
+ - **Generated docs can fake rigor.** A folder of polished planning artifacts looks
158
+ like progress, but it isn't proof. The proof is the reviewer running your real
159
+ tests before each commit.
160
+
161
+ That's the whole trade: where larger kits add machinery, tiny-spec adds one
162
+ independent reviewer and stops.
163
+
164
+ ### Does the trade hold up?
165
+
166
+ Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
167
+ runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
168
+ code with held-out tests the suite never sees, and checks the reviewer's own verdict
169
+ against that ground truth — the number that matters being **false-PASS rate**, how
170
+ often the gate blesses code that is actually broken.
171
+
172
+ Most recent run — 5 benchmark tasks, v0.4.0 (`7810074`), 2026-08-03:
173
+
174
+ | | |
175
+ |---|---|
176
+ | held-out pass rate | **100%** (5/5) — the code works, judged by tests it never saw |
177
+ | suite/truth agreement | **100%** — the reviewer's verdict matched the held-out grader every time |
178
+ | **false-PASS rate** | **0%** — it never called broken code done |
179
+ | blocker rate | 0% |
180
+
181
+ The same run a month earlier on `f973c85` scored identically, which is the more
182
+ interesting result: the loop is at least stable across versions rather than tuned to a
183
+ single commit.
184
+
185
+ See [`docs/eval/README.md`](docs/eval/README.md) for the method and
186
+ [`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the scoring rubric
187
+ this repo grades itself against. **Read both skeptically: it is my benchmark, of my own
188
+ tool, scored by me.** The rubric says so itself — *"treat the tiny-spec column as
189
+ self-assessment to pressure-test, not gospel."* The tasks are small and self-contained,
190
+ so the harness measures the loop on well-specified work, not large-codebase performance.
191
+ The harness is in the repo precisely so you don't have to take my word for it.
192
+
82
193
  ## New to spec-driven development?
83
194
 
84
195
  Spec-driven development (SDD) means writing down *what* you want and *why* before
@@ -121,6 +232,17 @@ Or collapse the three planning steps into one and go straight to building:
121
232
  **stops before `tiny-spec-build`** — that's where you actually review the work — and
122
233
  it writes nothing itself, it only delegates.
123
234
 
235
+ Or hand it a whole list of stories:
236
+
237
+ ```
238
+ /tiny-spec-loop # per story: branch → plan → build → merge → next
239
+ ```
240
+
241
+ `tiny-spec-loop` reads your `BREAKDOWN.md` (or a list you paste) and works the stories
242
+ one after another, merging each finished branch into `main` locally before starting the
243
+ next — until it's done or reaches a **terminal state** it names out loud. See
244
+ [Working a whole list](#working-a-whole-list--tiny-spec-loop).
245
+
124
246
  **Building from a mockup?** There is no design flag to pass. Commit your exports to
125
247
  `design/` before the run and `tiny-spec-create` — whether you invoke it directly or
126
248
  reach it through `tiny-spec-run` — reads what's there and asks which screens this
@@ -141,7 +263,7 @@ git clone https://github.com/GrayMa77er/tiny-spec.git
141
263
  cd tiny-spec
142
264
 
143
265
  mkdir -p "$HOME/.claude/skills" "$HOME/.claude/agents"
144
- for s in tiny-spec-prd tiny-spec-breakdown tiny-spec-run tiny-spec-create tiny-spec-plan tiny-spec-tasks tiny-spec-build; do
266
+ for s in tiny-spec-prd tiny-spec-breakdown tiny-spec-run tiny-spec-loop tiny-spec-create tiny-spec-plan tiny-spec-tasks tiny-spec-build; do
145
267
  cp -R "$s" "$HOME/.claude/skills/$s"
146
268
  done
147
269
  cp agents/*.md "$HOME/.claude/agents/"
@@ -152,6 +274,13 @@ or install one set at a time.
152
274
 
153
275
  </details>
154
276
 
277
+ ### See a finished run first
278
+
279
+ [`examples/todo-cli/`](examples/todo-cli/) is a real run of the flow on one small
280
+ ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`,
281
+ `tasks.md` and `constitution.md` the suite wrote, and the code and tests it produced.
282
+ The tests pass; you can clone it and run the gate yourself.
283
+
155
284
  ## How it works
156
285
 
157
286
  The constitution (`constitution.md`) is the spine. `tiny-spec-create` seeds it from a
@@ -332,19 +461,21 @@ unnoticed.
332
461
  3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
333
462
  end to end and grades against the constitution and the task's acceptance.
334
463
  4. On pass, commit the code plus a checklist tick. On fail, loop back to the
335
- executor with the findings. After two failed attempts it becomes a blocker.
464
+ executor with the findings. After two failed attempts the run halts `exhausted`.
336
465
 
337
466
  ```mermaid
338
467
  flowchart TB
339
468
  SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
340
469
 
470
+ TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
341
471
  TASKS --> P[Plan task]
342
472
  P --> I[Implement<br/>executor]
343
473
  I --> R[Review + run gate<br/>reviewer]
474
+ I -->|blocker| B
344
475
  R -->|pass| C[Commit + tick]
345
476
  C --> TASKS
346
477
  R -->|fail| I
347
- R -->|fail twice| B[Blocker logged to decisions.md]
478
+ R -->|fail twice| B[Halt — blocked / exhausted<br/>logged to decisions.md]
348
479
 
349
480
  CON([constitution.md]) -.-> P & I & R
350
481
  MEM([memory.md]) -.-> I & R
@@ -365,42 +496,74 @@ and logs a blocker instead of hacking around it. You fix the gap upstream in
365
496
  `tiny-spec-plan` or `tiny-spec-create`, then resume. Work runs one ticket at a time and
366
497
  resumes from the checklist state.
367
498
 
368
- ## Why it's small
499
+ ### Working a whole list — `tiny-spec-loop`
369
500
 
370
- Most spec frameworks are generous by default:
371
- many phases, many agents, many generated documents. tiny-spec makes the opposite
372
- bet. Keep one safeguard, drop the rest.
501
+ `/tiny-spec-loop` takes a list of stories and works them one after another. Per story
502
+ it does the same four moves:
373
503
 
374
- A green unit test suite is not the same as working software, so the reviewer
375
- exercises acceptance criteria end to end and a final smoke test confirms the whole
376
- spec. That independent review is the safeguard — not the volume of planning
377
- artifacts. One task, one commit, an external reviewer. Nothing gets added unless
378
- it earns its place.
504
+ ```
505
+ cut a branch from main → tiny-spec-run → tiny-spec-build → merge back to main
506
+ ```
379
507
 
380
- The case for staying small:
508
+ Then the next story. Each branch is cut **fresh from main**, so story 3 sees stories 1
509
+ and 2 already merged — which is what makes an ordered list build correctly.
381
510
 
382
- - **Documents are context, and context isn't free.** Generating large `spec.md`,
383
- `plan.md`, `research.md`, and `data-model.md` files costs tokens to write, then
384
- costs context to carry. Every paragraph the agent has to hold is room it no
385
- longer has for your actual code. tiny-spec keeps the spine small — a
386
- constitution and a short memory — and injects only what each task needs.
387
- - **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
388
- assume you're bootstrapping a project from a blank page. Day to day, you pick up
389
- a ticket and change part of a system that already exists. tiny-spec binds to a
390
- ticket, works one at a time, and references your task platform instead of
391
- re-describing the world.
392
- - **Rigid pipelines fight the user.** Mandatory phases and required sections
393
- impose ceremony on work that doesn't need it. tiny-spec's extra structure is
394
- optional by design — add shape where it pays, skip it where it doesn't.
395
- - **More moving parts is more to maintain.** Orchestrators, ownership contracts,
396
- checkpoint matrices, and config files are themselves a system you have to learn
397
- and keep in sync. A few small skills and two agents are not.
398
- - **Generated docs can fake rigor.** A folder of polished planning artifacts looks
399
- like progress, but it isn't proof. The proof is the reviewer running your real
400
- tests before each commit.
511
+ **The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
512
+ each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
513
+ a list at invocation instead and that wins; but a bare feature name has no acceptance
514
+ criteria, so `tiny-spec-create` will interview you when it reaches it. That's the
515
+ honest trade: a breakdown runs unattended, a pasted list is supervised.
401
516
 
402
- That's the whole trade: where larger kits add machinery, tiny-spec adds one
403
- independent reviewer and stops.
517
+ **Every run ends in exactly one of six states, and it says which:**
518
+
519
+ | | |
520
+ |---|---|
521
+ | `done` | every story built **and merged** |
522
+ | `blocked` | an upstream document is wrong — go fix the spec or the plan |
523
+ | `exhausted` | a task stayed red past two fix attempts |
524
+ | `paused` | it reached a `pause:` point |
525
+ | `fork` | a real either/or the plan doesn't answer |
526
+ | `conflict` | a story's branch wouldn't merge cleanly |
527
+
528
+ **Only `done` means the work is built** — and in a story loop that means *all* of them.
529
+ A run that stopped at story 2 of 7 reporting "done" is the single thing autonomous loops
530
+ get wrong most often, so the state is always named, along with what got merged and how
531
+ many stories are still untouched.
532
+
533
+ **A halt stops the whole run, not just that story.** Later stories in a list you wrote
534
+ top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
535
+ just produces a second, more confusing failure downstream.
536
+
537
+ **Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
538
+ the build halts *before* running it:
539
+
540
+ ```
541
+ - [ ] T7 — run the schema migration against the restored snapshot
542
+ - acceptance: every row in orders has a non-null tenant_id
543
+ - pause: confirm the snapshot is current before this touches real rows
544
+ ```
545
+
546
+ `tiny-spec-tasks` proposes these for genuinely irreversible work — migrations,
547
+ destructive file operations, a new dependency, an auth boundary, a public API contract.
548
+ You can also give the run a standing policy up front ("halt before anything that touches
549
+ auth") and it gets applied as each story's tasks are sliced.
550
+
551
+ **What it will not do to your repo.** It runs exactly five git commands — `switch`,
552
+ `switch -c`, `merge --no-ff`, `merge --abort`, and reads. It refuses to start on a dirty
553
+ tree. It **never pushes**, never rebases, never resets, never deletes a branch, and
554
+ never opens a PR. Merges are local, so a bad run is one `git reset` away; publishing
555
+ stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
556
+ the undo command rather than running it.
557
+
558
+ **Walk away and come back.** Nothing is written down to track progress — it's derived:
559
+ a story whose ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked
560
+ task is in progress, no directory means not started. So re-running `/tiny-spec-loop`
561
+ tomorrow in a fresh session picks up exactly where it stopped. No run-state file, no
562
+ lock, and no budget to configure: the story list *is* the budget.
563
+
564
+ **It never fixes a blocker for you.** A blocker means one of your documents is wrong,
565
+ and a loop allowed to rewrite the requirement its own task just failed would be grading
566
+ its own homework. It reports and stops; the upstream edit is yours.
404
567
 
405
568
  ## Project layout
406
569
 
@@ -7,10 +7,24 @@
7
7
  <p align="center">A tiny, opinionated take on spec-driven development.</p>
8
8
 
9
9
  <p align="center">
10
+ <a href="https://pypi.org/project/tiny-spec/"><img src="https://img.shields.io/pypi/v/tiny-spec.svg?color=d97757" alt="PyPI version"></a>
11
+ <a href="https://pypi.org/project/tiny-spec/"><img src="https://img.shields.io/pypi/dm/tiny-spec.svg?color=d97757" alt="PyPI downloads"></a>
10
12
  <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue.svg" alt="License: MIT"></a>
11
13
  <a href="https://docs.claude.com/en/docs/claude-code/overview"><img src="https://img.shields.io/badge/Claude%20Code-skills-d97757.svg" alt="Claude Code"></a>
14
+ <a href="https://github.com/GrayMa77er/tiny-spec/stargazers"><img src="https://img.shields.io/github/stars/GrayMa77er/tiny-spec?color=d97757" alt="Stars"></a>
12
15
  </p>
13
16
 
17
+ <p align="center">
18
+ <code>uvx tiny-spec install</code>
19
+ </p>
20
+
21
+ <!-- DEMO GIF SLOT — images/demo.gif
22
+ Record the FAIL -> fix -> PASS loop, then uncomment:
23
+ <p align="center">
24
+ <img src="images/demo.gif" alt="tiny-spec build loop: the reviewer fails a task, the executor fixes it, the task commits" width="800">
25
+ </p>
26
+ -->
27
+
14
28
  tiny-spec is a four-step workflow for Claude Code that turns a ticket into shipped,
15
29
  reviewed code. You write the intent, it produces a design, a task list, and then
16
30
  builds the work one task at a time. Every task is implemented by one agent and
@@ -20,7 +34,9 @@ committed.
20
34
  That core is **four skills and two agents**. In front of it sit **two optional
21
35
  planning on-ramps** — `tiny-spec-prd` (idea → PRD) and `tiny-spec-breakdown`
22
36
  (PRD → stories) — for when you're starting from an idea rather than a ready ticket.
23
- `tiny-spec-run` walks the three planning steps in one command. No config file, no
37
+ Two optional routers sit over the top: `tiny-spec-run` walks the three planning steps
38
+ in one command, and `tiny-spec-loop` works a whole list of stories — branch, plan,
39
+ build, merge, next — until they're built or something stops it. No config file, no
24
40
  build step.
25
41
 
26
42
  ```
@@ -32,6 +48,10 @@ PLANNING (optional on-ramps) EXECUTION (the core loop, one story at
32
48
  idea → PRD PRD → stories intent design tasks per-task loop
33
49
  PRD.md BREAKDOWN.md SPEC.md PLAN.md + tasks.md plan → implement → review → commit
34
50
  constitution
51
+ ▲ ▲
52
+ └──────────────── tiny-spec-loop ─────────────────┘
53
+ optional: per story — branch, run, build, merge,
54
+ then the next story. Halts on a wall.
35
55
  ```
36
56
 
37
57
  The two on-ramps are **optional** and stack. Have nothing written down? Run
@@ -42,6 +62,97 @@ work? Skip both and start at `tiny-spec-create`. Both on-ramps write a regenerab
42
62
  file at your project root (not under `.spec/`); `tiny-spec-create` then reads the
43
63
  breakdown one story at a time.
44
64
 
65
+ ## How small
66
+
67
+ Every other kit in this space is bigger. That is the whole pitch, so here is the
68
+ receipt rather than the adjective:
69
+
70
+ | | skills / commands | agents | config | artifacts per feature |
71
+ |---|---|---|---|---|
72
+ | **tiny-spec** | **8** (4 core + 4 optional) | **2** | **none** | **`SPEC` `PLAN` `tasks`** |
73
+ | [GitHub Spec Kit](https://github.com/github/spec-kit) | 10 | — | `specify init` | `spec` `plan` `tasks` `checklist` `constitution` `research` `data-model` `contracts/` `quickstart` |
74
+ | [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
75
+ | [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
76
+
77
+ The whole of tiny-spec is **2,079 lines** of markdown across 8 skills and 2 agents.
78
+ There is no config file, no build step, no orchestrator, and no CLI to initialize a
79
+ project — the skills are the product.
80
+
81
+ <sub>Counts taken from each project's own repository on 2026-08-03 by listing its
82
+ command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
83
+ `templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
84
+ OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
85
+ personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
86
+ `customize.toml` files (35). These projects are all substantially more popular than
87
+ tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
88
+ bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
89
+
90
+ ## Why it's small
91
+
92
+ Most spec frameworks are generous by default:
93
+ many phases, many agents, many generated documents. tiny-spec makes the opposite
94
+ bet. Keep one safeguard, drop the rest.
95
+
96
+ A green unit test suite is not the same as working software, so the reviewer
97
+ exercises acceptance criteria end to end and a final smoke test confirms the whole
98
+ spec. That independent review is the safeguard — not the volume of planning
99
+ artifacts. One task, one commit, an external reviewer. Nothing gets added unless
100
+ it earns its place.
101
+
102
+ The case for staying small:
103
+
104
+ - **Documents are context, and context isn't free.** Generating large `spec.md`,
105
+ `plan.md`, `research.md`, and `data-model.md` files costs tokens to write, then
106
+ costs context to carry. Every paragraph the agent has to hold is room it no
107
+ longer has for your actual code. tiny-spec keeps the spine small — a
108
+ constitution and a short memory — and injects only what each task needs.
109
+ - **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
110
+ assume you're bootstrapping a project from a blank page. Day to day, you pick up
111
+ a ticket and change part of a system that already exists. tiny-spec binds to a
112
+ ticket, works one at a time, and references your task platform instead of
113
+ re-describing the world.
114
+ - **Rigid pipelines fight the user.** Mandatory phases and required sections
115
+ impose ceremony on work that doesn't need it. tiny-spec's extra structure is
116
+ optional by design — add shape where it pays, skip it where it doesn't.
117
+ - **More moving parts is more to maintain.** Orchestrators, ownership contracts,
118
+ checkpoint matrices, and config files are themselves a system you have to learn
119
+ and keep in sync. A few small skills and two agents are not.
120
+ - **Generated docs can fake rigor.** A folder of polished planning artifacts looks
121
+ like progress, but it isn't proof. The proof is the reviewer running your real
122
+ tests before each commit.
123
+
124
+ That's the whole trade: where larger kits add machinery, tiny-spec adds one
125
+ independent reviewer and stops.
126
+
127
+ ### Does the trade hold up?
128
+
129
+ Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
130
+ runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
131
+ code with held-out tests the suite never sees, and checks the reviewer's own verdict
132
+ against that ground truth — the number that matters being **false-PASS rate**, how
133
+ often the gate blesses code that is actually broken.
134
+
135
+ Most recent run — 5 benchmark tasks, v0.4.0 (`7810074`), 2026-08-03:
136
+
137
+ | | |
138
+ |---|---|
139
+ | held-out pass rate | **100%** (5/5) — the code works, judged by tests it never saw |
140
+ | suite/truth agreement | **100%** — the reviewer's verdict matched the held-out grader every time |
141
+ | **false-PASS rate** | **0%** — it never called broken code done |
142
+ | blocker rate | 0% |
143
+
144
+ The same run a month earlier on `f973c85` scored identically, which is the more
145
+ interesting result: the loop is at least stable across versions rather than tuned to a
146
+ single commit.
147
+
148
+ See [`docs/eval/README.md`](docs/eval/README.md) for the method and
149
+ [`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the scoring rubric
150
+ this repo grades itself against. **Read both skeptically: it is my benchmark, of my own
151
+ tool, scored by me.** The rubric says so itself — *"treat the tiny-spec column as
152
+ self-assessment to pressure-test, not gospel."* The tasks are small and self-contained,
153
+ so the harness measures the loop on well-specified work, not large-codebase performance.
154
+ The harness is in the repo precisely so you don't have to take my word for it.
155
+
45
156
  ## New to spec-driven development?
46
157
 
47
158
  Spec-driven development (SDD) means writing down *what* you want and *why* before
@@ -84,6 +195,17 @@ Or collapse the three planning steps into one and go straight to building:
84
195
  **stops before `tiny-spec-build`** — that's where you actually review the work — and
85
196
  it writes nothing itself, it only delegates.
86
197
 
198
+ Or hand it a whole list of stories:
199
+
200
+ ```
201
+ /tiny-spec-loop # per story: branch → plan → build → merge → next
202
+ ```
203
+
204
+ `tiny-spec-loop` reads your `BREAKDOWN.md` (or a list you paste) and works the stories
205
+ one after another, merging each finished branch into `main` locally before starting the
206
+ next — until it's done or reaches a **terminal state** it names out loud. See
207
+ [Working a whole list](#working-a-whole-list--tiny-spec-loop).
208
+
87
209
  **Building from a mockup?** There is no design flag to pass. Commit your exports to
88
210
  `design/` before the run and `tiny-spec-create` — whether you invoke it directly or
89
211
  reach it through `tiny-spec-run` — reads what's there and asks which screens this
@@ -104,7 +226,7 @@ git clone https://github.com/GrayMa77er/tiny-spec.git
104
226
  cd tiny-spec
105
227
 
106
228
  mkdir -p "$HOME/.claude/skills" "$HOME/.claude/agents"
107
- for s in tiny-spec-prd tiny-spec-breakdown tiny-spec-run tiny-spec-create tiny-spec-plan tiny-spec-tasks tiny-spec-build; do
229
+ for s in tiny-spec-prd tiny-spec-breakdown tiny-spec-run tiny-spec-loop tiny-spec-create tiny-spec-plan tiny-spec-tasks tiny-spec-build; do
108
230
  cp -R "$s" "$HOME/.claude/skills/$s"
109
231
  done
110
232
  cp agents/*.md "$HOME/.claude/agents/"
@@ -115,6 +237,13 @@ or install one set at a time.
115
237
 
116
238
  </details>
117
239
 
240
+ ### See a finished run first
241
+
242
+ [`examples/todo-cli/`](examples/todo-cli/) is a real run of the flow on one small
243
+ ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`,
244
+ `tasks.md` and `constitution.md` the suite wrote, and the code and tests it produced.
245
+ The tests pass; you can clone it and run the gate yourself.
246
+
118
247
  ## How it works
119
248
 
120
249
  The constitution (`constitution.md`) is the spine. `tiny-spec-create` seeds it from a
@@ -295,19 +424,21 @@ unnoticed.
295
424
  3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
296
425
  end to end and grades against the constitution and the task's acceptance.
297
426
  4. On pass, commit the code plus a checklist tick. On fail, loop back to the
298
- executor with the findings. After two failed attempts it becomes a blocker.
427
+ executor with the findings. After two failed attempts the run halts `exhausted`.
299
428
 
300
429
  ```mermaid
301
430
  flowchart TB
302
431
  SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
303
432
 
433
+ TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
304
434
  TASKS --> P[Plan task]
305
435
  P --> I[Implement<br/>executor]
306
436
  I --> R[Review + run gate<br/>reviewer]
437
+ I -->|blocker| B
307
438
  R -->|pass| C[Commit + tick]
308
439
  C --> TASKS
309
440
  R -->|fail| I
310
- R -->|fail twice| B[Blocker logged to decisions.md]
441
+ R -->|fail twice| B[Halt — blocked / exhausted<br/>logged to decisions.md]
311
442
 
312
443
  CON([constitution.md]) -.-> P & I & R
313
444
  MEM([memory.md]) -.-> I & R
@@ -328,42 +459,74 @@ and logs a blocker instead of hacking around it. You fix the gap upstream in
328
459
  `tiny-spec-plan` or `tiny-spec-create`, then resume. Work runs one ticket at a time and
329
460
  resumes from the checklist state.
330
461
 
331
- ## Why it's small
462
+ ### Working a whole list — `tiny-spec-loop`
332
463
 
333
- Most spec frameworks are generous by default:
334
- many phases, many agents, many generated documents. tiny-spec makes the opposite
335
- bet. Keep one safeguard, drop the rest.
464
+ `/tiny-spec-loop` takes a list of stories and works them one after another. Per story
465
+ it does the same four moves:
336
466
 
337
- A green unit test suite is not the same as working software, so the reviewer
338
- exercises acceptance criteria end to end and a final smoke test confirms the whole
339
- spec. That independent review is the safeguard — not the volume of planning
340
- artifacts. One task, one commit, an external reviewer. Nothing gets added unless
341
- it earns its place.
467
+ ```
468
+ cut a branch from main → tiny-spec-run → tiny-spec-build → merge back to main
469
+ ```
342
470
 
343
- The case for staying small:
471
+ Then the next story. Each branch is cut **fresh from main**, so story 3 sees stories 1
472
+ and 2 already merged — which is what makes an ordered list build correctly.
344
473
 
345
- - **Documents are context, and context isn't free.** Generating large `spec.md`,
346
- `plan.md`, `research.md`, and `data-model.md` files costs tokens to write, then
347
- costs context to carry. Every paragraph the agent has to hold is room it no
348
- longer has for your actual code. tiny-spec keeps the spine small — a
349
- constitution and a short memory — and injects only what each task needs.
350
- - **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
351
- assume you're bootstrapping a project from a blank page. Day to day, you pick up
352
- a ticket and change part of a system that already exists. tiny-spec binds to a
353
- ticket, works one at a time, and references your task platform instead of
354
- re-describing the world.
355
- - **Rigid pipelines fight the user.** Mandatory phases and required sections
356
- impose ceremony on work that doesn't need it. tiny-spec's extra structure is
357
- optional by design — add shape where it pays, skip it where it doesn't.
358
- - **More moving parts is more to maintain.** Orchestrators, ownership contracts,
359
- checkpoint matrices, and config files are themselves a system you have to learn
360
- and keep in sync. A few small skills and two agents are not.
361
- - **Generated docs can fake rigor.** A folder of polished planning artifacts looks
362
- like progress, but it isn't proof. The proof is the reviewer running your real
363
- tests before each commit.
474
+ **The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
475
+ each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
476
+ a list at invocation instead and that wins; but a bare feature name has no acceptance
477
+ criteria, so `tiny-spec-create` will interview you when it reaches it. That's the
478
+ honest trade: a breakdown runs unattended, a pasted list is supervised.
364
479
 
365
- That's the whole trade: where larger kits add machinery, tiny-spec adds one
366
- independent reviewer and stops.
480
+ **Every run ends in exactly one of six states, and it says which:**
481
+
482
+ | | |
483
+ |---|---|
484
+ | `done` | every story built **and merged** |
485
+ | `blocked` | an upstream document is wrong — go fix the spec or the plan |
486
+ | `exhausted` | a task stayed red past two fix attempts |
487
+ | `paused` | it reached a `pause:` point |
488
+ | `fork` | a real either/or the plan doesn't answer |
489
+ | `conflict` | a story's branch wouldn't merge cleanly |
490
+
491
+ **Only `done` means the work is built** — and in a story loop that means *all* of them.
492
+ A run that stopped at story 2 of 7 reporting "done" is the single thing autonomous loops
493
+ get wrong most often, so the state is always named, along with what got merged and how
494
+ many stories are still untouched.
495
+
496
+ **A halt stops the whole run, not just that story.** Later stories in a list you wrote
497
+ top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
498
+ just produces a second, more confusing failure downstream.
499
+
500
+ **Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
501
+ the build halts *before* running it:
502
+
503
+ ```
504
+ - [ ] T7 — run the schema migration against the restored snapshot
505
+ - acceptance: every row in orders has a non-null tenant_id
506
+ - pause: confirm the snapshot is current before this touches real rows
507
+ ```
508
+
509
+ `tiny-spec-tasks` proposes these for genuinely irreversible work — migrations,
510
+ destructive file operations, a new dependency, an auth boundary, a public API contract.
511
+ You can also give the run a standing policy up front ("halt before anything that touches
512
+ auth") and it gets applied as each story's tasks are sliced.
513
+
514
+ **What it will not do to your repo.** It runs exactly five git commands — `switch`,
515
+ `switch -c`, `merge --no-ff`, `merge --abort`, and reads. It refuses to start on a dirty
516
+ tree. It **never pushes**, never rebases, never resets, never deletes a branch, and
517
+ never opens a PR. Merges are local, so a bad run is one `git reset` away; publishing
518
+ stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
519
+ the undo command rather than running it.
520
+
521
+ **Walk away and come back.** Nothing is written down to track progress — it's derived:
522
+ a story whose ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked
523
+ task is in progress, no directory means not started. So re-running `/tiny-spec-loop`
524
+ tomorrow in a fresh session picks up exactly where it stopped. No run-state file, no
525
+ lock, and no budget to configure: the story list *is* the budget.
526
+
527
+ **It never fixes a blocker for you.** A blocker means one of your documents is wrong,
528
+ and a loop allowed to rewrite the requirement its own task just failed would be grading
529
+ its own homework. It reports and stops; the upstream edit is yours.
367
530
 
368
531
  ## Project layout
369
532
 
@@ -0,0 +1,69 @@
1
+ # Example: todo CLI
2
+
3
+ A **real run** of the tiny-spec suite on one small ticket — not a hand-written ideal.
4
+ Everything in this folder (the `.spec/` artifacts and the code) was produced by
5
+ actually driving `tiny-spec-create → plan → tasks → build` autonomously on
6
+ [`TICKET.md`](TICKET.md), then committed verbatim. Browse it to see exactly what the
7
+ workflow generates before you run it on your own work.
8
+
9
+ > Produced on 2026-06-29 with tiny-spec **v0.2.1**, and kept as-run rather than
10
+ > retouched to match the current artifact formats — a doctored example would defeat
11
+ > the point. Later versions add a `## Design` section to `SPEC.md` (unused here: this
12
+ > is a CLI with no design surface) and richer status flags. The shape is the same.
13
+
14
+ ## What's here
15
+
16
+ ```
17
+ TICKET.md the input — a small "todo CLI" ask
18
+ .spec/
19
+ constitution.md the shared spine: style, standards, invariants,
20
+ definition of done, and the verification gate
21
+ todo-cli/
22
+ SPEC.md intent + REQ-1..REQ-9 (what "done" means)
23
+ PLAN.md the design and how each requirement is covered
24
+ tasks.md the ordered checklist, all ticked [x]
25
+ todo.py the produced CLI (stdlib only, 117 lines)
26
+ test_todo.py the produced end-to-end test (129 lines)
27
+ ```
28
+
29
+ Read them in flow order: `TICKET.md` → `SPEC.md` → `PLAN.md` → `tasks.md` →
30
+ `todo.py`. The `constitution.md` is the persistent context injected into every task.
31
+
32
+ ## It passes its own gate
33
+
34
+ The committed code clears the gate defined in its own
35
+ [`constitution.md`](.spec/constitution.md) — the same check `tiny-spec-build-reviewer`
36
+ runs. From this directory:
37
+
38
+ ```sh
39
+ python3 -m pytest test_todo.py # 8 passed
40
+ ```
41
+
42
+ And the ticket's example session reproduces exactly:
43
+
44
+ ```sh
45
+ python3 todo.py add "buy milk" # -> 1
46
+ python3 todo.py add "write tests" # -> 2
47
+ python3 todo.py list # two open lines
48
+ python3 todo.py done 1
49
+ python3 todo.py list # 1 now shows [x]
50
+ ```
51
+
52
+ > Running these creates a `todos.json` in your working directory — delete it
53
+ > afterward, and don't commit it.
54
+
55
+ ## How it was generated
56
+
57
+ Driven headlessly the same way the eval harness drives its benchmark tasks: a
58
+ hermetic sandbox with the suite vendored into a local `.claude/`, seeded with
59
+ `TICKET.md`, and run via `claude -p` with instructions to complete the flow
60
+ autonomously. See [`docs/eval/harness/run.sh`](../../docs/eval/harness/run.sh) for the
61
+ exact recipe — the `DRIVER` prompt near the top is what was handed to the model.
62
+
63
+ This run finished in 55 turns at roughly $3.49 with no blockers; all five tasks
64
+ passed review on the first or second attempt.
65
+
66
+ To produce your own equivalent, point that harness recipe at this `TICKET.md`, or
67
+ just run the flow interactively in a scratch directory. Output won't be
68
+ byte-identical — the model is nondeterministic — but it will be an equivalent,
69
+ gate-passing run.
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
4
4
 
5
5
  [project]
6
6
  name = "tiny-spec"
7
- version = "0.4.0"
7
+ version = "0.5.0"
8
8
  description = "A tiny, opinionated take on spec-driven development."
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.9"
@@ -37,6 +37,7 @@ packages = ["tiny_spec"]
37
37
  "tiny-spec-prd" = "tiny_spec/_bundle/tiny-spec-prd"
38
38
  "tiny-spec-breakdown" = "tiny_spec/_bundle/tiny-spec-breakdown"
39
39
  "tiny-spec-run" = "tiny_spec/_bundle/tiny-spec-run"
40
+ "tiny-spec-loop" = "tiny_spec/_bundle/tiny-spec-loop"
40
41
  "tiny-spec-create" = "tiny_spec/_bundle/tiny-spec-create"
41
42
  "tiny-spec-plan" = "tiny_spec/_bundle/tiny-spec-plan"
42
43
  "tiny-spec-tasks" = "tiny_spec/_bundle/tiny-spec-tasks"
@@ -48,6 +49,7 @@ include = [
48
49
  "tiny_spec",
49
50
  "tiny-spec-prd",
50
51
  "tiny-spec-run",
52
+ "tiny-spec-loop",
51
53
  "tiny-spec-breakdown",
52
54
  "tiny-spec-create",
53
55
  "tiny-spec-plan",
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: tiny-spec-build
3
- description: Build the spec — run the per-task loop plan→implement→review→commit, one task at a time. Implements with a fresh executor, grades with an independent reviewer running the real gate, commits per passed task, keeps a lean memory. Resumes from the checkbox state.
3
+ description: Build the spec — run the per-task loop plan→implement→review→commit, one task at a time. Implements with a fresh executor, grades with an independent reviewer running the real gate, commits per passed task, keeps a lean memory. Halts on a blocker, a spent convergence budget, a task's pause: point, or a genuine fork, and names which. Resumes from the checkbox state.
4
4
  ---
5
5
 
6
6
  # tiny-spec-build
@@ -42,6 +42,15 @@ alongside this skill (see the suite README).
42
42
  For the selected task, run these steps in order. **Do not tick a task until its
43
43
  reviewer passes.**
44
44
 
45
+ **Before step 1, check the task's `pause:` field.** If it has one and the user hasn't
46
+ waived it for this run, **halt `paused`** (see **Halting** below): leave the task
47
+ `[ ]`, dispatch nothing, and don't touch `tasks.md`. A pause stops *before* the work so
48
+ the approach gets reviewed while redirecting it is still cheap.
49
+
50
+ A pause can be waived or added for the current run only, conversationally — "skip the
51
+ pause on T5", "also stop before T9". Honor it for this run and **never write it back
52
+ into `tasks.md`**: the file records the standing intent, not one run's override.
53
+
45
54
  ### 1. PLAN (inline, brief)
46
55
  Restate the task as a 2–4 step micro-plan against the constitution: which
47
56
  **invariants** apply, which files it touches, which **Definition of Done** items
@@ -90,8 +99,12 @@ It fails on what it measured or saw, and merely flags what is taste.
90
99
 
91
100
  ### 4. CONVERGE (on FAIL)
92
101
  Re-dispatch the **executor** with the reviewer's findings appended to its brief.
93
- Bound this to **2 fix attempts**. If it still fails after that, stop and treat it
94
- as a **blocker** (below) — don't keep grinding or hand-fix past the loop silently.
102
+ Bound this to **2 fix attempts**. If it still fails after that, stop and **halt
103
+ `exhausted`** (see **Halting** below) — don't keep grinding or hand-fix past the loop
104
+ silently. `exhausted` is recorded as a blocker, because a task that three attempts
105
+ couldn't turn green is a signal the upstream document is wrong — but name it
106
+ `exhausted`, not `blocked`: the executor never said it couldn't proceed, it just
107
+ never got there.
95
108
 
96
109
  ### 5. COMMIT + TICK (on PASS)
97
110
  Two commits, in order (keeps code history clean of planning churn), both in
@@ -156,12 +169,62 @@ On first use, create the file with this structure:
156
169
  Report the task outcome (built, reviewed, committed). Then:
157
170
  - **Interactive default:** continue to the next unchecked task. Pausing for the
158
171
  user between tasks is fine and expected.
159
- - If the user asked to **run it through** ("do it all", "build everything"), keep
160
- looping until done or a blocker stops you — committing per passed task as you go.
172
+ - If the user asked to **run it through** ("do it all", "build everything"), or
173
+ `tiny-spec-loop` invoked you, keep looping until one of the five terminal states
174
+ below is reached — committing per passed task as you go.
175
+
176
+ **The task list is the budget.** The loop ends when the tasks end. There is no turn
177
+ ceiling, no token cap, and no max-tasks knob to set — adding one would put a second,
178
+ weaker stopping rule next to the real one.
161
179
 
162
180
  There is **no** separate autonomous mode and **no** checkpoint config: one commit
163
181
  per passed task, always, on the current branch. (If the user wants a feature
164
- branch, create it once up front — that's their call, not a knob here.)
182
+ branch, create it once up front — that's their call, not a knob here.) Running
183
+ through is a caller's instruction, not a mode — every step above is identical either
184
+ way, which is why a loop and a hand-driven build can't drift apart.
185
+
186
+ ## Halting
187
+
188
+ Every build run ends in exactly **one** of five terminal states. Name it out loud in
189
+ the closing report:
190
+
191
+ | State | Cause |
192
+ |---|---|
193
+ | `done` | every task `[x]` and the final smoke passed |
194
+ | `blocked` | the executor reported a `BLOCKER`, or the reviewer found a missing `visual:` command |
195
+ | `exhausted` | convergence (step 4) spent its 2 fix attempts and the task is still red |
196
+ | `paused` | a task's `pause:` point was reached |
197
+ | `fork` | a genuine either/or the plan doesn't pin down |
198
+
199
+ **Never round up.** Only `done` may report the work as built. A run that ended at
200
+ `blocked`, `exhausted`, `paused`, or `fork` has unbuilt work in it, and a closing
201
+ report that calls it done is worse than the halt itself — it converts a stop the user
202
+ could act on into a false completion they won't check. Say the state, then say what's
203
+ left and which command resolves it.
204
+
205
+ **`blocked` and `exhausted`** log **one** entry with `type: blocker` — an upstream doc
206
+ must change; see **Blockers** below. **`fork`** logs `type: halt`:
207
+
208
+ ```
209
+ ## D-NNN — <short title>
210
+ - type: halt
211
+ - state: fork
212
+ - date: <ISO date>
213
+ - affects: REQ-N, T<n>
214
+ - note: <the options, your recommendation, and what the user has to decide>
215
+ ```
216
+
217
+ Write one entry per halt, not one per cause — a blocker that halts the loop is a single
218
+ event and reads as two records if you log it twice.
219
+
220
+ **`done` and `paused` log nothing.** A green tree with every box ticked already says
221
+ `done`. And a `paused` run stopped at a task whose `pause:` line is sitting right there
222
+ in `tasks.md` saying why — an entry would restate what the file already tells you, and
223
+ the log is for things that aren't otherwise visible. Report it, don't record it.
224
+
225
+ The task the loop stopped at stays `[ ]`. That is the whole resume mechanism:
226
+ re-running `tiny-spec-build` (or `tiny-spec-loop`) picks up at exactly that task with no
227
+ conversational context needed.
165
228
 
166
229
  ## Blockers (never hack around)
167
230
 
@@ -170,16 +233,22 @@ command** for a `design:` task, or when convergence (step 4) exhausts its attemp
170
233
 
171
234
  1. Leave the task `[ ]`.
172
235
  2. Log it to `.spec/<active>/decisions.md` using the fixed skeleton (`type: blocker`,
173
- naming the upstream doc to fix in `note:`, and the affected `REQ-N`/`T<n>` in
174
- `affects:`). Create the file if absent:
236
+ the terminal state in `state:`, the upstream doc to fix named in `note:`, and the
237
+ affected `REQ-N`/`T<n>` in `affects:`). Create the file if absent:
175
238
 
176
239
  ```
177
240
  ## D-NNN — <short title>
178
241
  - type: blocker
242
+ - state: blocked
179
243
  - date: <ISO date>
180
244
  - affects: REQ-N, T<n>
181
245
  - note: <what stopped you; which upstream doc must change>
182
246
  ```
247
+
248
+ `state:` is `blocked` for an executor blocker or a missing `visual:`, and
249
+ `exhausted` when convergence ran out. Both route upstream the same way; the
250
+ distinction is what the user reads to know whether the task was impossible or just
251
+ never converged.
183
252
  3. **Surface and route upstream:** `tiny-spec-plan` (a design gap, a missing `visual:`
184
253
  command, or a token the Design system doesn't define) or `tiny-spec-create` (a
185
254
  requirement is wrong/impossible), in update mode. Never resolve a missing `visual:`
@@ -190,6 +259,12 @@ command** for a `design:` task, or when convergence (step 4) exhausts its attemp
190
259
  A genuine fork the plan doesn't pin down → don't guess: present the options + your
191
260
  recommendation, get the user's call, record it in `decisions.md`, then continue.
192
261
 
262
+ **When running through, a fork is a halt** (`type: halt`, `state: fork`) — record the
263
+ options and your recommendation in the `note:` and stop, rather than blocking on an
264
+ answer nobody may be present to give. Ask only when the user is in the conversation
265
+ with you; a question asked into an empty room stalls the run without ending it, and
266
+ the halt is what makes the same fork readable hours later.
267
+
193
268
  ## Completion
194
269
 
195
270
  When every task in `tasks.md` is `[x]`:
@@ -199,10 +274,11 @@ When every task in `tasks.md` is `[x]`:
199
274
  setup — install/build, not a test-runner shortcut). There is no separate
200
275
  verify skill — this final smoke confirms the requirements actually work end-to-end, not
201
276
  just that tasks are ticked.
202
- 2. **Report** — what was built, the commits made (with the branch), and any open
203
- `decisions.md` items (blockers, tasks unchecked by a reconcile). If the final
277
+ 2. **Report** — the terminal state by name (`done` here, since every task is `[x]`
278
+ and the smoke passed), what was built, the commits made (with the branch), and any
279
+ open `decisions.md` items (blockers, tasks unchecked by a reconcile). If the final
204
280
  smoke reveals a gap, it's a bug to fix now (new/edited task) or a blocker to
205
- route upstream — not a pass.
281
+ route upstream — **not** a pass, and not `done`.
206
282
  3. **Close the ticket (reference-only).** If a ticket is bound, this is where the
207
283
  **closing** keyword belongs — on the final / PR commit, not the per-task ones:
208
284
  GitHub `Closes #123`, ADO `Fixes AB#123`. Jira/Monday have no closing keyword,
@@ -111,6 +111,21 @@ Only the **project-wide** questions collapse — still confirm this story's bind
111
111
  requirements. If there is **no `BREAKDOWN.md`**, or no entry matches, run the **full
112
112
  interview** below unchanged.
113
113
 
114
+ **Exception — a loop run.** If the caller says this is a **loop run** and names the
115
+ story's slug (`tiny-spec-loop` does, through `tiny-spec-run`), skip step 6's
116
+ confirmation and write `SPEC.md` straight out. The `AC:` lines *are* the approval: the
117
+ user reviewed them when they wrote or accepted `BREAKDOWN.md`, and re-confirming them
118
+ once per story is asking the same question twice — across a seven-story run it is the
119
+ difference between walking away and being interrupted seven times. Report the captured
120
+ `REQ-N` instead of asking about them.
121
+
122
+ This waives **only** that confirmation. Everything that is a genuine question still
123
+ stops the run: a story whose `AC:` lines contradict each other or the Decisions block,
124
+ an `AC:` you cannot turn into a testable `REQ-N`, a missing ticket id the platform
125
+ needs, a design export the story names that isn't on disk. Those are not ceremony —
126
+ they are the human input the loop is supposed to stop for. Never guess past one because
127
+ a loop is running.
128
+
114
129
  ## Pick the slug (resolve the active dir)
115
130
 
116
131
  **Fresh modes only** — in reseed or update mode the active dir already exists and you
@@ -0,0 +1,188 @@
1
+ ---
2
+ name: tiny-spec-loop
3
+ description: Work a list of stories end to end, one after another — for each story cut a branch from main, plan it, build it, merge it back locally, then move to the next. Reads BREAKDOWN.md by default, or a list pasted at invocation. Halts the WHOLE run on a blocker, a spent convergence budget, a task's pause: point, a genuine fork, a merge conflict, or a red gate after a merge — and names which. Use for "build the backlog", "work through the breakdown", "build these features one after another". NOT for a single story — for that use tiny-spec-run then tiny-spec-build.
4
+ ---
5
+
6
+ # tiny-spec-loop
7
+
8
+ A **router**, not a stage. It owns no artifact, writes nothing, and holds no state
9
+ file. It walks a list of stories and, for each one, drives the existing flow to a
10
+ finished, merged branch — then moves to the next.
11
+
12
+ Per story the cycle is always the same four moves:
13
+
14
+ ```
15
+ cut a branch from main → tiny-spec-run → tiny-spec-build → merge back to main
16
+ ```
17
+
18
+ **Everything it decides comes from files on disk and from git.** Which stories are
19
+ built, which are merged, which haven't started — all derived, never recorded. That is
20
+ what makes a run abandoned overnight resume correctly tomorrow, and it is why this
21
+ skill needs no run log, pointer, or lock. Do not "optimize" this into stored state.
22
+
23
+ **It supplies one thing the flow was missing: a stopping rule.** The task list is the
24
+ goal, the independent reviewer running the real gate is the verification, `memory.md`
25
+ plus a commit per passed task is the memory. This skill adds when to stop, and nothing
26
+ else — no budget, no turn ceiling, no config.
27
+
28
+ ## Step 0 — the run brief
29
+
30
+ The one interactive moment. Everything after it runs without asking unless a stage
31
+ genuinely needs an answer.
32
+
33
+ 1. **Check the preconditions.** All three, before touching anything:
34
+ - it is a git repo (`git rev-parse --is-inside-work-tree`);
35
+ - the working tree is **clean** (`git status --porcelain` prints nothing) — a dirty
36
+ tree would get swept into the first story's commits;
37
+ - an integration branch exists — `main`, or `master` if there is no `main`.
38
+
39
+ Any of them failing → say which and stop. Do not offer to stash, commit, or create
40
+ the branch for the user.
41
+
42
+ 2. **Resolve the story list** (Step 1) and play it back in order — slug and title —
43
+ with the state you derived for each (Step 2). This is the user's chance to
44
+ reorder, drop, or narrow before anything is cut.
45
+
46
+ 3. **Take the pause policy.** Ask for, or accept, standing technical stop points for
47
+ this run — *"halt before anything that touches auth"*, *"stop before any schema
48
+ migration"*. Carry the wording verbatim into every story's `tiny-spec-tasks` stage
49
+ (Step 3), where it becomes a real `pause:` line on the matching task. Pause points
50
+ are **technical**, not per-story: there is no way to mark a whole story for review,
51
+ because the thing worth looking at is a migration or an auth boundary, not a
52
+ feature heading.
53
+
54
+ Step 0 is a **once-per-run** check. Do not re-run it between stories.
55
+
56
+ ## Step 1 — the story list
57
+
58
+ **Default: `BREAKDOWN.md` at the project root**, in file order. Each `- Story:` under
59
+ each `## Feature:` heading is one item; take its **`slug:`** — that names both the
60
+ branch and the `.spec/<slug>/` dir. A `## Feature:` heading is a grouping, not an item:
61
+ it carries no slug, so a Feature with three stories is three branches and three merges.
62
+
63
+ **A list pasted at invocation wins** over `BREAKDOWN.md` when the user gives one. Treat
64
+ each line as a story title and derive a kebab-case slug from it. Such a story has **no
65
+ acceptance criteria**, so `tiny-spec-create` will run its full interview when it reaches
66
+ that story — which is correct, not a failure: a one-line feature name is not enough to
67
+ build from, and that interview *is* the human input the run stops for. Say so in Step 0
68
+ so the user knows a pasted list is a supervised run, not a walk-away one.
69
+
70
+ If neither exists, stop and say so. Never invent the list.
71
+
72
+ ## Step 2 — where each story stands (derive, don't record)
73
+
74
+ For each story slug, in order, ask git — **in this order**, first match wins:
75
+
76
+ 1. **`git show <integration>:.spec/<slug>/tasks.md`** succeeds and every task is `[x]`
77
+ → **built and merged.** Skip it.
78
+ 2. Otherwise, if the branch exists (`git rev-parse --verify <slug>`), read
79
+ **`git show <slug>:.spec/<slug>/tasks.md`**:
80
+ - succeeds, every task `[x]` → **built, not merged** → resume at the **merge** (4).
81
+ - succeeds, at least one `[ ]` → **in progress** → resume at the **build** (3).
82
+ - fails (no task list yet) → **planning incomplete** → resume at the **plan** (2).
83
+ 3. Otherwise → **not started.** Run the full cycle.
84
+
85
+ Ask **git**, not the working tree. A story that isn't merged yet has no `.spec/<slug>/`
86
+ on the integration branch at all, so "does the directory exist" can't tell "not started"
87
+ apart from "built on a branch you haven't merged" — reading each ref explicitly can.
88
+ Sourcing step 1 from the integration branch is also what survives a deleted branch: once
89
+ a story is merged, its ticked task list is part of `main` whether or not the branch that
90
+ built it still exists.
91
+
92
+ ## Step 3 — the per-story cycle
93
+
94
+ For the first story that isn't already built and merged:
95
+
96
+ 1. **Branch.** `git switch <slug>` if it already exists; otherwise
97
+ `git switch -c <slug> <integration>` — cut **fresh from the integration branch** so
98
+ this story sees every story merged before it. That is what makes an ordered list
99
+ build correctly: story 3 gets stories 1 and 2 already in its tree.
100
+
101
+ 2. **Plan.** Invoke **`tiny-spec-run`**, once. It resolves `.spec/<slug>/` from the
102
+ branch name automatically, so you don't name the ticket dir. Hand it two things:
103
+ - a **create-stage brief**: *this is a loop run; seed from the `BREAKDOWN.md` story
104
+ with slug `<slug>` and do not stop to confirm the requirements — its `AC:` lines
105
+ are already approved.* Omit the seeding clause for a pasted-list story; there is
106
+ nothing to seed from and the interview is correct.
107
+ - a **tasks-stage brief**: the pause policy from Step 0, verbatim, so the technical
108
+ stop points land as `pause:` lines.
109
+
110
+ If `tiny-spec-run` stops anywhere other than **L9** (tasks ready) or **L10** (already
111
+ built), that is a halt — report and stop the whole run. **Never invoke it twice to
112
+ push past its own stop:** each of those is a human decision it deliberately declined
113
+ to make, and running it again declines again.
114
+
115
+ 3. **Build.** Invoke **`tiny-spec-build`**, once, briefed to run it through. It owns
116
+ the per-task loop and writes its own halt record. Anything other than `done` halts
117
+ the whole run — **do not merge a story that didn't finish**, and do not invoke build
118
+ a second time: it resumes from the checkbox state, so it lands on the very task that
119
+ just halted and halts there again.
120
+
121
+ 4. **Merge — only on `done`.** In order:
122
+ - `git switch <integration>`
123
+ - `git merge --no-ff <slug>` — the merge commit keeps each story legible in history.
124
+ - **Conflict** → `git merge --abort`, then halt `conflict` (below).
125
+ - **Run the constitution's Verification commands on the merged result**, exercised
126
+ the way a user would. A story that was green alone can still break against work
127
+ merged before it, and that is exactly what this catches. Red → halt `blocked`.
128
+ - **Never push.** Merging locally keeps a bad run one `git reset` away; sending it
129
+ to a remote is the user's call, and this suite makes no network calls.
130
+
131
+ 5. **Next story.** Return to Step 2. Do not re-run Step 0.
132
+
133
+ ## Step 4 — report
134
+
135
+ Name, in this order:
136
+
137
+ 1. **The terminal state** — exactly one of `done`, `blocked`, `exhausted`, `paused`,
138
+ `fork`, `conflict`. Use the word.
139
+ 2. **Stories built and merged**, in order, with their merge commits.
140
+ 3. **The story it stopped on**, the task within it, and why in one line.
141
+ 4. **Stories never started** — say how many are left, by name. A run that stopped at
142
+ story 2 of 7 must not read like a finished backlog.
143
+ 5. **The one command that resolves it** — `tiny-spec-create`/`tiny-spec-plan` in update
144
+ mode for `blocked`/`exhausted`, `tiny-spec-loop` again for `paused`, the decision
145
+ the user owes you for a `fork`, or the conflicted paths for a `conflict`.
146
+
147
+ **Only `done` — every story merged — may report the work as built.** A run that halted
148
+ has unbuilt stories in it, and a report that rounds `blocked`, `exhausted`, `paused`,
149
+ `fork`, or `conflict` up to done converts a stop the user could act on into a false
150
+ completion they won't check. Say the state, then say what's left.
151
+
152
+ ### Halting
153
+
154
+ A halt stops **the whole run**, not just the current story. Later stories in a list you
155
+ wrote top to bottom usually assume the earlier ones landed, so skipping ahead past a
156
+ failure produces a second, more confusing failure downstream.
157
+
158
+ `tiny-spec-build` records its own halts (`blocked`, `exhausted`, `fork`) in the story's
159
+ `decisions.md`. **You record nothing** — the merge-stage halts are already legible
160
+ without a log: a `conflict` leaves the story's branch unmerged and git itself reports
161
+ the conflicted paths, and a red gate after a merge is reported by the gate. A log entry
162
+ restating what git already shows is a second source of truth with extra steps.
163
+
164
+ **On a red gate after a merge, leave the merge in place.** Report it, name the story,
165
+ and tell the user that `git reset --hard HEAD~1` on the integration branch undoes it.
166
+ Do not undo it yourself: fixing forward and rolling back are both reasonable, the
167
+ choice is theirs, and discarding a real merge is not a call a router gets to make.
168
+
169
+ ### Hard rules
170
+
171
+ - **Never write, edit, or flip anything.** No `status:`, no checkbox, no
172
+ `decisions.md` entry, no code, no `BREAKDOWN.md` edit. Delegate or stop.
173
+ - **Never push, force, rebase, reset, or delete a branch.** The only git commands this
174
+ skill issues are `switch`, `switch -c`, `merge --no-ff`, `merge --abort`, and
175
+ read-only queries. Everything outward-facing or destructive is the user's.
176
+ - **Never merge a story whose build didn't return `done`.**
177
+ - **At most one `tiny-spec-run` and one `tiny-spec-build` per story.**
178
+ - **Never invoke `tiny-spec-loop`.** Re-entering means re-reading these steps, not
179
+ calling yourself. Self-invocation compounds context and does not terminate.
180
+ - **Never invoke `tiny-spec-create`, `tiny-spec-plan`, or `tiny-spec-tasks` directly.**
181
+ That ladder is `tiny-spec-run`'s; walking it here would duplicate it, and two ladders
182
+ drift. Briefs are passed *through* run, not around it.
183
+ - **Never resolve a halt yourself.** A blocker means an upstream document is wrong,
184
+ which is the user's call. Routing to `plan`/`create` in update mode automatically
185
+ would let the loop rewrite the requirement its own task just failed to satisfy — the
186
+ agent grading its own homework, one level up.
187
+ - **There is no budget to set.** The story list is the budget: the run ends when the
188
+ stories end. No turn ceiling, no token cap, no max-stories knob.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: tiny-spec-run
3
- description: Run the spec flow end to end — read each artifact's status flag to work out where the active ticket stands, then invoke tiny-spec-create → tiny-spec-plan → tiny-spec-tasks in order, taking anything marked stale before anything missing. STOPS before tiny-spec-build; it never builds, never commits, and never writes an artifact itself. Use for "run the whole spec flow", "spec this out end to end", or "pick the chain back up after a change". NOT for a single stage — for that, invoke tiny-spec-create, tiny-spec-plan, or tiny-spec-tasks directly.
3
+ description: Run the spec flow end to end — read each artifact's status flag to work out where the active ticket stands, then invoke tiny-spec-create → tiny-spec-plan → tiny-spec-tasks in order, taking anything marked stale before anything missing. STOPS before tiny-spec-build; it never builds, never commits, and never writes an artifact itself. Use for "run the whole spec flow", "spec this out end to end", or "pick the chain back up after a change". NOT for a single stage — for that, invoke tiny-spec-create, tiny-spec-plan, or tiny-spec-tasks directly. NOT for building — for "spec it out and build it", use tiny-spec-loop.
4
4
  ---
5
5
 
6
6
  # tiny-spec-run
@@ -29,9 +29,11 @@ to resolve. `tiny-spec-create` picks `BREAKDOWN.md` up on its own when it exists
29
29
  If the user named a **single stage** ("update the spec", "redo the tasks"), hand off
30
30
  to that skill and stop. `run` is for walking the chain, not for wrapping one stage.
31
31
 
32
- If they asked for the flow **and** the build ("spec this out and build it"), run the
33
- chain as normal and stop at the end anyway — then tell them `tiny-spec-build` is the
34
- next command. The stop is not negotiable; see the hard rules.
32
+ If they asked for the flow **and** the build ("spec this out and build it"), that is
33
+ `tiny-spec-loop` — hand off to it and stop. If you are *already* inside a
34
+ `tiny-spec-loop` run (it invoked you), run the chain as normal and stop at the end
35
+ anyway; the loop takes it from there. Either way **you** never enter the build. The
36
+ stop is not negotiable; see the hard rules.
35
37
 
36
38
  Step 0 is a **once-per-run** check on the user's opening request. Do not re-run it
37
39
  when you return to Step 1 after a stage.
@@ -181,6 +183,13 @@ something new:
181
183
  fall back to an existing ticket dir — the user confirmed this is new work.*
182
184
  (`tiny-spec-create` honors that phrase by skipping its sole-dir fallback.)
183
185
 
186
+ **Caller briefs — pass them through verbatim.** A caller (today, `tiny-spec-loop`) may
187
+ hand you a brief addressed to a specific stage: *"for the create stage: …"*, *"for the
188
+ tasks stage: …"*. Append it, **word for word**, to the scope you send that stage, and
189
+ send nothing to the others. Do not summarize it, act on it yourself, or let it change
190
+ which rung fires — the ladder is still yours, and a brief is cargo, not an instruction
191
+ to you. A brief for a stage this run never reaches is simply never delivered.
192
+
184
193
  **A stage's closing "point the user at X" is not a terminus.** Each stage ends by
185
194
  naming the next skill ("point the user at `tiny-spec-plan`"). Inside a run that
186
195
  sentence is a *report*, not a stop — when a stage finishes, return to Step 1 and keep
@@ -224,5 +233,9 @@ invent a `## Design` section and the condition would still hold on the next pass
224
233
  Telling the user once and letting them decide is the version that terminates.
225
234
 
226
235
  Then hand off explicitly: **run `tiny-spec-build` when you're ready to build** (it
227
- starts at the first unchecked task). If you stopped at L4 or on a bound, say exactly
228
- what stopped you and what the user needs to decide.
236
+ starts at the first unchecked task, reviewing as it goes), or **`tiny-spec-loop`** to
237
+ run it through until the work is built or something stops it. If you stopped at L4 or
238
+ on a bound, say exactly what stopped you and what the user needs to decide.
239
+
240
+ (If `tiny-spec-loop` invoked you, this report is what it reads to decide whether to
241
+ build — so say which rung you stopped on either way.)
@@ -45,6 +45,7 @@ For each task, write:
45
45
  - type: feat # optional; Conventional Commit type (defaults to feat)
46
46
  - req: REQ-n # optional; the REQ-N this task delivers
47
47
  - design: D-n # optional; the SPEC.md D<n> screen this task builds — arms the visual gate
48
+ - pause: <why> # optional; halt the loop before this task so a human looks first
48
49
  - files: <comma-separated hint of files it will touch>
49
50
  ```
50
51
 
@@ -70,6 +71,32 @@ other. Two rules:
70
71
  - Only set it if the constitution has a `visual:` verification command. Without one
71
72
  the reviewer cannot render anything and will raise a blocker instead of a verdict.
72
73
 
74
+ **pause** is the other field that changes what happens rather than what gets written —
75
+ but where `design:` changes how a task is **graded**, `pause:` changes whether the loop
76
+ **continues**. Set it and `tiny-spec-build` halts *before* that task, leaving it
77
+ unchecked, so a human reviews the approach while redirecting it is still cheap. The
78
+ value is one line saying what to look at.
79
+
80
+ Propose it only on work that is genuinely **irreversible or wide-blast-radius**:
81
+
82
+ - a data migration, or anything that writes to real rows;
83
+ - a destructive or bulk file operation;
84
+ - pulling in a new third-party dependency;
85
+ - an auth, permissions, or trust boundary;
86
+ - a public API or schema contract other people's code depends on.
87
+
88
+ **When in doubt, leave it out.** A pause the user didn't want is worse than no pause at
89
+ all — it trains them to wave past the ones they did want, which is exactly the reflex
90
+ that makes the mechanism useless the one time it matters. Most task lists should carry
91
+ zero or one.
92
+
93
+ **A caller may hand you a standing pause policy** — *"halt before anything that touches
94
+ auth"*, *"stop before any schema migration"* (`tiny-spec-loop` passes one through
95
+ `tiny-spec-run`). Apply it to **this** task list: any task matching the description gets
96
+ a `pause:` naming the policy that put it there, on top of whatever you'd have set
97
+ anyway. A policy that matches nothing here is not an error — say so and move on, rather
98
+ than stretching a task to fit it.
99
+
73
100
  Cover **every** part of the approach — together the tasks must deliver all
74
101
  `REQ-N`. Don't leave a requirement with no task. Likewise, if `SPEC.md` has a
75
102
  `## Design` section, every `D<n>` in it needs at least one task carrying that
@@ -91,9 +118,10 @@ updated: <ISO date>
91
118
  # Tasks — <project / feature name>
92
119
 
93
120
  > Executed top to bottom, one at a time. A checked `[x]` task is implemented AND
94
- > reviewed. `type:`, `req:`, and `design:` are optional; `files:` is a hint, not an
95
- > ownership contract. A task with `design:` is also graded against that screen's
96
- > `D<n>` entry and the constitution's Design system.
121
+ > reviewed. `type:`, `req:`, `design:`, and `pause:` are optional; `files:` is a hint,
122
+ > not an ownership contract. A task with `design:` is also graded against that screen's
123
+ > `D<n>` entry and the constitution's Design system. A task with `pause:` halts the
124
+ > build before it runs, so a human looks first.
97
125
 
98
126
  ## Tasks
99
127
 
@@ -112,6 +140,7 @@ updated: <ISO date>
112
140
 
113
141
  - [ ] T3 — <…>
114
142
  - acceptance: <observable outcome>
143
+ - pause: <optional; what to check before this runs — irreversible work only>
115
144
  - files: <path, path>
116
145
  ```
117
146
 
@@ -138,5 +167,7 @@ When `tasks.md` is `status: stale`:
138
167
 
139
168
  ## When done
140
169
 
141
- Report the task count and point the user at `tiny-spec-build` (one task at a time) or
142
- note they can run it straight through.
170
+ Report the task count — and any `pause:` points you set, with their reason, so the user
171
+ can drop one before it fires. Then point them at `tiny-spec-build` (one task at a time,
172
+ reviewing as it goes) or `tiny-spec-loop` (run it through until it's built or something
173
+ stops it).
@@ -5,4 +5,4 @@ into your Claude Code config directory (``~/.claude/`` by default). It adds no
5
5
  runtime behavior to the suite itself.
6
6
  """
7
7
 
8
- __version__ = "0.4.0"
8
+ __version__ = "0.5.0"
@@ -3,6 +3,7 @@
3
3
  "tiny-spec-prd",
4
4
  "tiny-spec-breakdown",
5
5
  "tiny-spec-run",
6
+ "tiny-spec-loop",
6
7
  "tiny-spec-create",
7
8
  "tiny-spec-plan",
8
9
  "tiny-spec-tasks",
File without changes
File without changes
File without changes
File without changes