tiny-spec 1.0.0__tar.gz → 1.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: tiny-spec
3
- Version: 1.0.0
3
+ Version: 1.1.0
4
4
  Summary: A tiny, opinionated take on spec-driven development.
5
5
  Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
6
6
  Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
@@ -115,54 +115,36 @@ receipt rather than the adjective:
115
115
  | [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
116
116
  | [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
117
117
 
118
- The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents.
119
- There is no config file, no build step, no orchestrator, and no CLI to initialize a
120
- project — the skills are the product. (That line count went *up* slightly in 1.0 while
121
- the skill count went down: `tiny-spec-adopt` is genuinely new capability, and merging
122
- four skills into two removed commands to learn rather than words to read.)
123
-
124
- <sub>Counts taken from each project's own repository on 2026-08-03 by listing its
125
- command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
126
- `templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
127
- OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
128
- personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
129
- `customize.toml` files (35). These projects are all substantially more popular than
130
- tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
131
- bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
118
+ The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents. No
119
+ config file, no build step, no orchestrator, no CLI to initialize a project — the skills
120
+ are the product.
132
121
 
133
- ## Why it's small
122
+ <sub>Counted on 2026-08-03 from each project's own repository, not its marketing copy —
123
+ Spec Kit's `templates/commands/`, OpenSpec's `skills/`, BMAD's `SKILL.md` and
124
+ `customize.toml` files on `main`. All three are far more popular than tiny-spec, and
125
+ being smaller is a bet, not a proof of quality — re-run the counts yourself before
126
+ believing them.</sub>
134
127
 
135
- Most spec frameworks are generous by default:
136
- many phases, many agents, many generated documents. tiny-spec makes the opposite
137
- bet. Keep one safeguard, drop the rest.
128
+ ## Why it's small
138
129
 
139
- A green unit test suite is not the same as working software, so the reviewer
140
- exercises acceptance criteria end to end and a final smoke test confirms the whole
141
- spec. That independent review is the safeguard — not the volume of planning
142
- artifacts. One task, one commit, an external reviewer. Nothing gets added unless
143
- it earns its place.
130
+ Most spec frameworks are generous by default: many phases, many agents, many generated
131
+ documents. tiny-spec makes the opposite bet — keep one safeguard, drop the rest. That
132
+ safeguard is the independent reviewer, which exercises acceptance criteria end to end,
133
+ because a green unit test suite is not the same as working software.
144
134
 
145
135
  The case for staying small:
146
136
 
147
- - **Documents are context, and context isn't free.** Generating large `spec.md`,
148
- `plan.md`, `research.md`, and `data-model.md` files costs tokens to write, then
149
- costs context to carry. Every paragraph the agent has to hold is room it no
150
- longer has for your actual code. tiny-spec keeps the spine small — a
151
- constitution and a short memory — and injects only what each task needs.
152
- - **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
153
- assume you're bootstrapping a project from a blank page. Day to day, you pick up
154
- a ticket and change part of a system that already exists. tiny-spec binds to a
155
- ticket, works one at a time, and references your task platform instead of
156
- re-describing the world.
157
- - **Rigid pipelines fight the user.** Mandatory phases and required sections
158
- impose ceremony on work that doesn't need it. tiny-spec's extra structure is
159
- optional by design — add shape where it pays, skip it where it doesn't.
160
- - **More moving parts is more to maintain.** Orchestrators, ownership contracts,
161
- checkpoint matrices, and config files are themselves a system you have to learn
162
- and keep in sync. A few small skills and two agents are not.
163
- - **Generated docs can fake rigor.** A folder of polished planning artifacts looks
164
- like progress, but it isn't proof. The proof is the reviewer running your real
165
- tests before each commit.
137
+ - **Documents are context, and context isn't free.** Every paragraph the agent has to
138
+ hold is room it no longer has for your actual code. tiny-spec keeps the spine small —
139
+ a constitution and a short memory — and injects only what each task needs.
140
+ - **Real work is a ticket inside a system, not a greenfield repo.** tiny-spec binds to a
141
+ ticket and references your task platform instead of re-describing the world.
142
+ - **Rigid pipelines fight the user.** Mandatory phases impose ceremony on work that
143
+ doesn't need it. The extra structure here is optional — add shape where it pays.
144
+ - **More moving parts is more to maintain.** Orchestrators, ownership contracts, and
145
+ config files are themselves a system you have to keep in sync. Two agents are not.
146
+ - **Generated docs can fake rigor.** A folder of polished planning artifacts looks like
147
+ progress, but the proof is the reviewer running your real tests before each commit.
166
148
 
167
149
  That's the whole trade: where larger kits add machinery, tiny-spec adds one
168
150
  independent reviewer and stops.
@@ -172,10 +154,10 @@ independent reviewer and stops.
172
154
  Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
173
155
  runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
174
156
  code with held-out tests the suite never sees, and checks the reviewer's own verdict
175
- against that ground truth — the number that matters being **false-PASS rate**, how
176
- often the gate blesses code that is actually broken.
157
+ against that ground truth. The number that matters is **false-PASS rate** — how often
158
+ the gate blesses code that is actually broken.
177
159
 
178
- Most recent run — 5 benchmark tasks, v1.0.0 (`b18e32a`), 2026-09-07:
160
+ Most recent run — 5 benchmark tasks on v1.0.0:
179
161
 
180
162
  | | |
181
163
  |---|---|
@@ -184,29 +166,24 @@ Most recent run — 5 benchmark tasks, v1.0.0 (`b18e32a`), 2026-09-07:
184
166
  | **false-PASS rate** | **0%** — it never called broken code done |
185
167
  | blocker rate | 0% |
186
168
 
187
- Two earlier runs — `7810074` (2026-08-03) and `f973c85` a month before it — scored
188
- identically, which is the more interesting result: the loop is stable across versions
189
- rather than tuned to a single commit. The 1.0 run matters most of the three, because
190
- 1.0 rebuilt the planning chain around it — `tiny-spec-plan` now writes the task list
191
- that `tiny-spec-build` executes — and the gate came through unchanged.
169
+ Two earlier versions scored identically, which is the more interesting result: the loop
170
+ is stable across releases rather than tuned to one commit — including through 1.0, which
171
+ rebuilt the planning chain underneath it.
192
172
 
193
173
  See [`docs/eval/README.md`](docs/eval/README.md) for the method and
194
- [`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the scoring rubric
195
- this repo grades itself against. **Read both skeptically: it is my benchmark, of my own
196
- tool, scored by me.** The rubric says so itself — *"treat the tiny-spec column as
197
- self-assessment to pressure-test, not gospel."* The tasks are small and self-contained,
198
- so the harness measures the loop on well-specified work, not large-codebase performance.
199
- The harness is in the repo precisely so you don't have to take my word for it.
174
+ [`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the rubric. **Read
175
+ both skeptically: it is my benchmark, of my own tool, scored by me**, on small
176
+ self-contained tasks — so it measures the loop on well-specified work, not
177
+ large-codebase performance. The harness is in the repo so you don't have to take my
178
+ word for it.
200
179
 
201
180
  ## New to spec-driven development?
202
181
 
203
- Spec-driven development (SDD) means writing down *what* you want and *why* before
204
- any code exists, then letting that spec drive the build. Instead of prompting an
205
- agent and hoping, you hand it a small, explicit contract — the intent, a design,
206
- and an ordered list of tasks — and it implements against that. The payoff: the
207
- agent stops guessing. It knows what "done" looks like, you can review the plan
208
- before a single line is written, and the result is checked against the spec
209
- rather than vibes. tiny-spec is one small take on that idea.
182
+ Spec-driven development (SDD) means writing down *what* you want and *why* before any
183
+ code exists, then letting that spec drive the build. Instead of prompting an agent and
184
+ hoping, you hand it a small, explicit contract — the intent, a design, and an ordered
185
+ list of tasks. The payoff: the agent stops guessing, you review the plan before a line
186
+ is written, and the result is checked against the spec rather than vibes.
210
187
 
211
188
  ## Quickstart
212
189
 
@@ -297,13 +274,12 @@ hardens it with concrete engineering rules, and `tiny-spec-build` injects it who
297
274
  every task. It holds your style, standards, invariants, definition of done, and
298
275
  verification commands.
299
276
 
300
- **On an existing codebase, deriving beats asking.** Your test command, your layout, and
301
- your conventions are already written down — in `package.json`, in CI, in your linter
302
- config, in the shape of the tree. `tiny-spec-adopt` reads them and marks each section
303
- as declared or inferred, so you know which parts to distrust. It then *runs* the
304
- verification commands it derived and reports which went green: a gate that has never
305
- been run is the most dangerous thing in the file, because it turns every future review
306
- into theatre.
277
+ **On an existing codebase, deriving beats asking.** Your test command, layout, and
278
+ conventions are already written down — in `package.json`, in CI, in your linter config,
279
+ in the shape of the tree. `tiny-spec-adopt` reads them, marks each section as declared
280
+ or inferred so you know which parts to distrust, then *runs* the verification commands
281
+ it derived and reports which went green. A gate that has never been run turns every
282
+ future review into theatre.
307
283
 
308
284
  Because it is project-wide the constitution can also go missing — deleted, or never
309
285
  committed — while your specs survive. Re-running `tiny-spec-adopt` (or
@@ -404,13 +380,12 @@ FINDINGS:
404
380
 
405
381
  Numbers first, and never a screenshot diff — pixel comparison goes flaky on font
406
382
  antialiasing and teams end up muting it. But no measurement catches an element that is
407
- present, on-token, and still not on screen — occluded, clipped, truncated, or the same
383
+ present, on-token, and still not on screen: occluded, clipped, truncated, or the same
408
384
  color as what's behind it. So the last step reads a screenshot of each state next to
409
385
  your export and grades presence, legibility, and correspondence. **Where a number
410
- already settled the question the eye may only flag** — on-scale-but-cramped is never a
411
- fail — which keeps the two halves from contradicting each other. Measurable or visible
412
- violations fail; taste comes back as `flag:` notes so a bounded fix loop can't thrash. A
413
- task with no `design:` tag is graded exactly as before.
386
+ already settled the question the eye may only flag**, which keeps the two halves from
387
+ contradicting each other — measurable or visible violations fail, taste comes back as
388
+ `flag:` notes. A task with no `design:` tag is graded exactly as before.
414
389
 
415
390
  <details>
416
391
  <summary>The <code>visual:</code> command (you write this once)</summary>
@@ -454,14 +429,12 @@ await b.close();
454
429
  ```
455
430
 
456
431
  The missing-selector branch **must print something** — that is what turns a renamed
457
- element into a failure instead of a silent skip. `top` is what the layout-order check
458
- reads, and `opacity`/`visibility` are worth printing because they turn the cheapest kind
459
- of invisible element into a numeric failure. The `SCREENSHOT` line is what the reviewer
460
- reads back to look at the render and catch the rest — occlusion, clipping, truncation,
461
- same-color-on-same-color — none of which any single property reports. Drop it and the
462
- numeric half still gates exactly as before, with the reviewer reporting `judge: not run`
463
- rather than quietly skipping it. Without a `visual:` command at all, a task tagged
464
- `design:` raises a blocker rather than passing quietly.
432
+ element into a failure instead of a silent skip. `top` feeds the layout-order check, and
433
+ `opacity`/`visibility` turn the cheapest kind of invisible element into a numeric
434
+ failure. The `SCREENSHOT` line arms the judge; drop it and the numeric half still gates
435
+ exactly as before, with the reviewer reporting `judge: not run` rather than quietly
436
+ skipping it. Without a `visual:` command at all, a task tagged `design:` raises a
437
+ blocker rather than passing quietly.
465
438
 
466
439
  </details>
467
440
 
@@ -476,10 +449,19 @@ unnoticed.
476
449
  1. Plan the task against the constitution (inline, brief).
477
450
  2. Implement it with a fresh `tiny-spec-build-executor` agent.
478
451
  3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
479
- end to end and grades against the constitution and the task's acceptance.
452
+ and grades against the constitution and the task's acceptance.
480
453
  4. On pass, commit the code plus a checklist tick. On fail, loop back to the
481
454
  executor with the findings. After two failed attempts the run halts `exhausted`.
482
455
 
456
+ Every task's acceptance is exercised end to end, black-box, by the reviewer — that part
457
+ never varies. What does vary is how much of the toolchain gets re-proven around it: a
458
+ clean `install → build` runs on the first task, the last one, anything touching
459
+ dependency or build config, and anything carrying `design:`; in between the reviewer runs
460
+ lint and tests against the existing build. It escalates back to the full gate whenever the
461
+ narrower one can't carry the verdict, and the authoritative clean run happens once more
462
+ over the whole project as the final smoke. Re-installing from scratch on every task
463
+ proved the same thing repeatedly and was the slowest part of the loop.
464
+
483
465
  ```mermaid
484
466
  flowchart TB
485
467
  SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
@@ -487,7 +469,7 @@ flowchart TB
487
469
  TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
488
470
  TASKS --> P[Plan task]
489
471
  P --> I[Implement<br/>executor]
490
- I --> R[Review + run gate<br/>reviewer]
472
+ I --> R[Review + run gate<br/>full or scoped<br/>reviewer]
491
473
  I -->|blocker| B
492
474
  R -->|pass| C[Commit + tick]
493
475
  C --> TASKS
@@ -543,10 +525,9 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
543
525
  | `fork` | a real either/or the plan doesn't answer |
544
526
  | `conflict` | a story's branch wouldn't merge cleanly |
545
527
 
546
- **Only `done` means the work is built** — and in a story run that means *all* of them.
547
- A run that stopped at story 2 of 7 reporting "done" is the single thing autonomous loops
548
- get wrong most often, so the state is always named, along with what got merged and how
549
- many stories are still untouched.
528
+ **Only `done` means the work is built** — and in a story run, that means *all* of them.
529
+ Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
530
+ often, so the state is always named alongside what merged and what's still untouched.
550
531
 
551
532
  **A halt stops the whole run, not just that story.** Later stories in a list you wrote
552
533
  top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
@@ -573,27 +554,26 @@ never opens a PR. Merges are local, so a bad run is one `git reset` away; publis
573
554
  stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
574
555
  the undo command rather than running it.
575
556
 
576
- **Walk away and come back.** Nothing is written down to track progress — it's derived:
577
- a story whose ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked
578
- task is in progress, no directory means not started. So asking `/tiny-spec-run` to build
579
- the backlog again tomorrow in a fresh session picks up exactly where it stopped. No
580
- run-state file, no lock, and no budget to configure: the story list *is* the budget.
557
+ **Walk away and come back.** Progress isn't written down, it's derived: a story whose
558
+ ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
559
+ progress, no directory means not started. Ask again tomorrow in a fresh session and it
560
+ picks up where it stopped. No run-state file, no lock, no budget to configure — the
561
+ story list *is* the budget.
581
562
 
582
563
  **It never fixes a blocker for you.** A blocker means one of your documents is wrong,
583
564
  and a run allowed to rewrite the requirement its own task just failed would be grading
584
565
  its own homework. It reports and stops; the upstream edit is yours.
585
566
 
586
567
  **And it never decides to build on its own.** Whether a run stops before the build or
587
- goes all the way is fixed from your opening request and never revised — not by a stage's
588
- closing line, not by a follow-up message. Build is your review gate, so promoting
589
- yourself into it is the one thing a router doesn't get to do.
568
+ goes all the way is fixed from your opening request and never revised. Build is your
569
+ review gate, so promoting itself into it is the one thing a router doesn't get to do.
590
570
 
591
571
  ## Project layout
592
572
 
593
573
  Each skill is one self-contained `SKILL.md`, with every document skeleton inline in
594
- it — no companion template files, no absolute paths, and no shared parent required
595
- at runtime, so a skill works wherever you drop it. (It also means a run never stops
596
- to ask permission to read a template out of your Claude config directory.)
574
+ it — no companion template files, no absolute paths, no shared parent required at
575
+ runtime. A skill works wherever you drop it, and a run never stops to ask permission to
576
+ read a template out of your Claude config directory.
597
577
 
598
578
  tiny-spec creates a `.spec/` directory in your project root, never inside a skill.
599
579
  It is namespaced per ticket, with a shared spine at the root:
@@ -78,54 +78,36 @@ receipt rather than the adjective:
78
78
  | [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
79
79
  | [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
80
80
 
81
- The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents.
82
- There is no config file, no build step, no orchestrator, and no CLI to initialize a
83
- project — the skills are the product. (That line count went *up* slightly in 1.0 while
84
- the skill count went down: `tiny-spec-adopt` is genuinely new capability, and merging
85
- four skills into two removed commands to learn rather than words to read.)
86
-
87
- <sub>Counts taken from each project's own repository on 2026-08-03 by listing its
88
- command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
89
- `templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
90
- OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
91
- personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
92
- `customize.toml` files (35). These projects are all substantially more popular than
93
- tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
94
- bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
81
+ The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents. No
82
+ config file, no build step, no orchestrator, no CLI to initialize a project — the skills
83
+ are the product.
95
84
 
96
- ## Why it's small
85
+ <sub>Counted on 2026-08-03 from each project's own repository, not its marketing copy —
86
+ Spec Kit's `templates/commands/`, OpenSpec's `skills/`, BMAD's `SKILL.md` and
87
+ `customize.toml` files on `main`. All three are far more popular than tiny-spec, and
88
+ being smaller is a bet, not a proof of quality — re-run the counts yourself before
89
+ believing them.</sub>
97
90
 
98
- Most spec frameworks are generous by default:
99
- many phases, many agents, many generated documents. tiny-spec makes the opposite
100
- bet. Keep one safeguard, drop the rest.
91
+ ## Why it's small
101
92
 
102
- A green unit test suite is not the same as working software, so the reviewer
103
- exercises acceptance criteria end to end and a final smoke test confirms the whole
104
- spec. That independent review is the safeguard — not the volume of planning
105
- artifacts. One task, one commit, an external reviewer. Nothing gets added unless
106
- it earns its place.
93
+ Most spec frameworks are generous by default: many phases, many agents, many generated
94
+ documents. tiny-spec makes the opposite bet — keep one safeguard, drop the rest. That
95
+ safeguard is the independent reviewer, which exercises acceptance criteria end to end,
96
+ because a green unit test suite is not the same as working software.
107
97
 
108
98
  The case for staying small:
109
99
 
110
- - **Documents are context, and context isn't free.** Generating large `spec.md`,
111
- `plan.md`, `research.md`, and `data-model.md` files costs tokens to write, then
112
- costs context to carry. Every paragraph the agent has to hold is room it no
113
- longer has for your actual code. tiny-spec keeps the spine small — a
114
- constitution and a short memory — and injects only what each task needs.
115
- - **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
116
- assume you're bootstrapping a project from a blank page. Day to day, you pick up
117
- a ticket and change part of a system that already exists. tiny-spec binds to a
118
- ticket, works one at a time, and references your task platform instead of
119
- re-describing the world.
120
- - **Rigid pipelines fight the user.** Mandatory phases and required sections
121
- impose ceremony on work that doesn't need it. tiny-spec's extra structure is
122
- optional by design — add shape where it pays, skip it where it doesn't.
123
- - **More moving parts is more to maintain.** Orchestrators, ownership contracts,
124
- checkpoint matrices, and config files are themselves a system you have to learn
125
- and keep in sync. A few small skills and two agents are not.
126
- - **Generated docs can fake rigor.** A folder of polished planning artifacts looks
127
- like progress, but it isn't proof. The proof is the reviewer running your real
128
- tests before each commit.
100
+ - **Documents are context, and context isn't free.** Every paragraph the agent has to
101
+ hold is room it no longer has for your actual code. tiny-spec keeps the spine small —
102
+ a constitution and a short memory — and injects only what each task needs.
103
+ - **Real work is a ticket inside a system, not a greenfield repo.** tiny-spec binds to a
104
+ ticket and references your task platform instead of re-describing the world.
105
+ - **Rigid pipelines fight the user.** Mandatory phases impose ceremony on work that
106
+ doesn't need it. The extra structure here is optional — add shape where it pays.
107
+ - **More moving parts is more to maintain.** Orchestrators, ownership contracts, and
108
+ config files are themselves a system you have to keep in sync. Two agents are not.
109
+ - **Generated docs can fake rigor.** A folder of polished planning artifacts looks like
110
+ progress, but the proof is the reviewer running your real tests before each commit.
129
111
 
130
112
  That's the whole trade: where larger kits add machinery, tiny-spec adds one
131
113
  independent reviewer and stops.
@@ -135,10 +117,10 @@ independent reviewer and stops.
135
117
  Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
136
118
  runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
137
119
  code with held-out tests the suite never sees, and checks the reviewer's own verdict
138
- against that ground truth — the number that matters being **false-PASS rate**, how
139
- often the gate blesses code that is actually broken.
120
+ against that ground truth. The number that matters is **false-PASS rate** — how often
121
+ the gate blesses code that is actually broken.
140
122
 
141
- Most recent run — 5 benchmark tasks, v1.0.0 (`b18e32a`), 2026-09-07:
123
+ Most recent run — 5 benchmark tasks on v1.0.0:
142
124
 
143
125
  | | |
144
126
  |---|---|
@@ -147,29 +129,24 @@ Most recent run — 5 benchmark tasks, v1.0.0 (`b18e32a`), 2026-09-07:
147
129
  | **false-PASS rate** | **0%** — it never called broken code done |
148
130
  | blocker rate | 0% |
149
131
 
150
- Two earlier runs — `7810074` (2026-08-03) and `f973c85` a month before it — scored
151
- identically, which is the more interesting result: the loop is stable across versions
152
- rather than tuned to a single commit. The 1.0 run matters most of the three, because
153
- 1.0 rebuilt the planning chain around it — `tiny-spec-plan` now writes the task list
154
- that `tiny-spec-build` executes — and the gate came through unchanged.
132
+ Two earlier versions scored identically, which is the more interesting result: the loop
133
+ is stable across releases rather than tuned to one commit — including through 1.0, which
134
+ rebuilt the planning chain underneath it.
155
135
 
156
136
  See [`docs/eval/README.md`](docs/eval/README.md) for the method and
157
- [`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the scoring rubric
158
- this repo grades itself against. **Read both skeptically: it is my benchmark, of my own
159
- tool, scored by me.** The rubric says so itself — *"treat the tiny-spec column as
160
- self-assessment to pressure-test, not gospel."* The tasks are small and self-contained,
161
- so the harness measures the loop on well-specified work, not large-codebase performance.
162
- The harness is in the repo precisely so you don't have to take my word for it.
137
+ [`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the rubric. **Read
138
+ both skeptically: it is my benchmark, of my own tool, scored by me**, on small
139
+ self-contained tasks — so it measures the loop on well-specified work, not
140
+ large-codebase performance. The harness is in the repo so you don't have to take my
141
+ word for it.
163
142
 
164
143
  ## New to spec-driven development?
165
144
 
166
- Spec-driven development (SDD) means writing down *what* you want and *why* before
167
- any code exists, then letting that spec drive the build. Instead of prompting an
168
- agent and hoping, you hand it a small, explicit contract — the intent, a design,
169
- and an ordered list of tasks — and it implements against that. The payoff: the
170
- agent stops guessing. It knows what "done" looks like, you can review the plan
171
- before a single line is written, and the result is checked against the spec
172
- rather than vibes. tiny-spec is one small take on that idea.
145
+ Spec-driven development (SDD) means writing down *what* you want and *why* before any
146
+ code exists, then letting that spec drive the build. Instead of prompting an agent and
147
+ hoping, you hand it a small, explicit contract — the intent, a design, and an ordered
148
+ list of tasks. The payoff: the agent stops guessing, you review the plan before a line
149
+ is written, and the result is checked against the spec rather than vibes.
173
150
 
174
151
  ## Quickstart
175
152
 
@@ -260,13 +237,12 @@ hardens it with concrete engineering rules, and `tiny-spec-build` injects it who
260
237
  every task. It holds your style, standards, invariants, definition of done, and
261
238
  verification commands.
262
239
 
263
- **On an existing codebase, deriving beats asking.** Your test command, your layout, and
264
- your conventions are already written down — in `package.json`, in CI, in your linter
265
- config, in the shape of the tree. `tiny-spec-adopt` reads them and marks each section
266
- as declared or inferred, so you know which parts to distrust. It then *runs* the
267
- verification commands it derived and reports which went green: a gate that has never
268
- been run is the most dangerous thing in the file, because it turns every future review
269
- into theatre.
240
+ **On an existing codebase, deriving beats asking.** Your test command, layout, and
241
+ conventions are already written down — in `package.json`, in CI, in your linter config,
242
+ in the shape of the tree. `tiny-spec-adopt` reads them, marks each section as declared
243
+ or inferred so you know which parts to distrust, then *runs* the verification commands
244
+ it derived and reports which went green. A gate that has never been run turns every
245
+ future review into theatre.
270
246
 
271
247
  Because it is project-wide the constitution can also go missing — deleted, or never
272
248
  committed — while your specs survive. Re-running `tiny-spec-adopt` (or
@@ -367,13 +343,12 @@ FINDINGS:
367
343
 
368
344
  Numbers first, and never a screenshot diff — pixel comparison goes flaky on font
369
345
  antialiasing and teams end up muting it. But no measurement catches an element that is
370
- present, on-token, and still not on screen — occluded, clipped, truncated, or the same
346
+ present, on-token, and still not on screen: occluded, clipped, truncated, or the same
371
347
  color as what's behind it. So the last step reads a screenshot of each state next to
372
348
  your export and grades presence, legibility, and correspondence. **Where a number
373
- already settled the question the eye may only flag** — on-scale-but-cramped is never a
374
- fail — which keeps the two halves from contradicting each other. Measurable or visible
375
- violations fail; taste comes back as `flag:` notes so a bounded fix loop can't thrash. A
376
- task with no `design:` tag is graded exactly as before.
349
+ already settled the question the eye may only flag**, which keeps the two halves from
350
+ contradicting each other — measurable or visible violations fail, taste comes back as
351
+ `flag:` notes. A task with no `design:` tag is graded exactly as before.
377
352
 
378
353
  <details>
379
354
  <summary>The <code>visual:</code> command (you write this once)</summary>
@@ -417,14 +392,12 @@ await b.close();
417
392
  ```
418
393
 
419
394
  The missing-selector branch **must print something** — that is what turns a renamed
420
- element into a failure instead of a silent skip. `top` is what the layout-order check
421
- reads, and `opacity`/`visibility` are worth printing because they turn the cheapest kind
422
- of invisible element into a numeric failure. The `SCREENSHOT` line is what the reviewer
423
- reads back to look at the render and catch the rest — occlusion, clipping, truncation,
424
- same-color-on-same-color — none of which any single property reports. Drop it and the
425
- numeric half still gates exactly as before, with the reviewer reporting `judge: not run`
426
- rather than quietly skipping it. Without a `visual:` command at all, a task tagged
427
- `design:` raises a blocker rather than passing quietly.
395
+ element into a failure instead of a silent skip. `top` feeds the layout-order check, and
396
+ `opacity`/`visibility` turn the cheapest kind of invisible element into a numeric
397
+ failure. The `SCREENSHOT` line arms the judge; drop it and the numeric half still gates
398
+ exactly as before, with the reviewer reporting `judge: not run` rather than quietly
399
+ skipping it. Without a `visual:` command at all, a task tagged `design:` raises a
400
+ blocker rather than passing quietly.
428
401
 
429
402
  </details>
430
403
 
@@ -439,10 +412,19 @@ unnoticed.
439
412
  1. Plan the task against the constitution (inline, brief).
440
413
  2. Implement it with a fresh `tiny-spec-build-executor` agent.
441
414
  3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
442
- end to end and grades against the constitution and the task's acceptance.
415
+ and grades against the constitution and the task's acceptance.
443
416
  4. On pass, commit the code plus a checklist tick. On fail, loop back to the
444
417
  executor with the findings. After two failed attempts the run halts `exhausted`.
445
418
 
419
+ Every task's acceptance is exercised end to end, black-box, by the reviewer — that part
420
+ never varies. What does vary is how much of the toolchain gets re-proven around it: a
421
+ clean `install → build` runs on the first task, the last one, anything touching
422
+ dependency or build config, and anything carrying `design:`; in between the reviewer runs
423
+ lint and tests against the existing build. It escalates back to the full gate whenever the
424
+ narrower one can't carry the verdict, and the authoritative clean run happens once more
425
+ over the whole project as the final smoke. Re-installing from scratch on every task
426
+ proved the same thing repeatedly and was the slowest part of the loop.
427
+
446
428
  ```mermaid
447
429
  flowchart TB
448
430
  SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
@@ -450,7 +432,7 @@ flowchart TB
450
432
  TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
451
433
  TASKS --> P[Plan task]
452
434
  P --> I[Implement<br/>executor]
453
- I --> R[Review + run gate<br/>reviewer]
435
+ I --> R[Review + run gate<br/>full or scoped<br/>reviewer]
454
436
  I -->|blocker| B
455
437
  R -->|pass| C[Commit + tick]
456
438
  C --> TASKS
@@ -506,10 +488,9 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
506
488
  | `fork` | a real either/or the plan doesn't answer |
507
489
  | `conflict` | a story's branch wouldn't merge cleanly |
508
490
 
509
- **Only `done` means the work is built** — and in a story run that means *all* of them.
510
- A run that stopped at story 2 of 7 reporting "done" is the single thing autonomous loops
511
- get wrong most often, so the state is always named, along with what got merged and how
512
- many stories are still untouched.
491
+ **Only `done` means the work is built** — and in a story run, that means *all* of them.
492
+ Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
493
+ often, so the state is always named alongside what merged and what's still untouched.
513
494
 
514
495
  **A halt stops the whole run, not just that story.** Later stories in a list you wrote
515
496
  top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
@@ -536,27 +517,26 @@ never opens a PR. Merges are local, so a bad run is one `git reset` away; publis
536
517
  stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
537
518
  the undo command rather than running it.
538
519
 
539
- **Walk away and come back.** Nothing is written down to track progress — it's derived:
540
- a story whose ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked
541
- task is in progress, no directory means not started. So asking `/tiny-spec-run` to build
542
- the backlog again tomorrow in a fresh session picks up exactly where it stopped. No
543
- run-state file, no lock, and no budget to configure: the story list *is* the budget.
520
+ **Walk away and come back.** Progress isn't written down, it's derived: a story whose
521
+ ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
522
+ progress, no directory means not started. Ask again tomorrow in a fresh session and it
523
+ picks up where it stopped. No run-state file, no lock, no budget to configure — the
524
+ story list *is* the budget.
544
525
 
545
526
  **It never fixes a blocker for you.** A blocker means one of your documents is wrong,
546
527
  and a run allowed to rewrite the requirement its own task just failed would be grading
547
528
  its own homework. It reports and stops; the upstream edit is yours.
548
529
 
549
530
  **And it never decides to build on its own.** Whether a run stops before the build or
550
- goes all the way is fixed from your opening request and never revised — not by a stage's
551
- closing line, not by a follow-up message. Build is your review gate, so promoting
552
- yourself into it is the one thing a router doesn't get to do.
531
+ goes all the way is fixed from your opening request and never revised. Build is your
532
+ review gate, so promoting itself into it is the one thing a router doesn't get to do.
553
533
 
554
534
  ## Project layout
555
535
 
556
536
  Each skill is one self-contained `SKILL.md`, with every document skeleton inline in
557
- it — no companion template files, no absolute paths, and no shared parent required
558
- at runtime, so a skill works wherever you drop it. (It also means a run never stops
559
- to ask permission to read a template out of your Claude config directory.)
537
+ it — no companion template files, no absolute paths, no shared parent required at
538
+ runtime. A skill works wherever you drop it, and a run never stops to ask permission to
539
+ read a template out of your Claude config directory.
560
540
 
561
541
  tiny-spec creates a `.spec/` directory in your project root, never inside a skill.
562
542
  It is namespaced per ticket, with a shared spine at the root:
@@ -24,7 +24,10 @@ Everything you need and nothing you don't:
24
24
  so you don't re-learn a pitfall a past run already paid for;
25
25
  - **if the task carries `design:`** — that screen's `D<n>` entry from `SPEC.md` and
26
26
  the path to its committed export;
27
- - the specific existing files that are your starting point, named explicitly.
27
+ - the specific existing files that are your starting point, named explicitly;
28
+ - sometimes, **the files the previous task left behind** — paths only. Start there when
29
+ orienting: it is the most recently moved ground in the codebase and the most likely
30
+ place your task connects to.
28
31
 
29
32
  You are **blind to the workflow, not to the codebase.** You don't get the plan,
30
33
  other tasks, or shared state. But the named files are a launch point, not the whole
@@ -64,7 +67,11 @@ violated.
64
67
  6. You MAY run a **narrow self-check** of your own work (the one test file you
65
68
  wrote, a syntax/import check). You do **not** need to run the full gate — the
66
69
  independent **reviewer** runs the authoritative Verification commands next.
67
- Leave the tree in a clean, buildable state for it.
70
+ **Leave the tree in a clean, buildable, installed state for it** — the reviewer may
71
+ run against the existing build rather than rebuilding from scratch, so a tree you
72
+ left half-migrated or with a stale artifact costs the task a review cycle. If you
73
+ changed dependencies or build configuration, apply them (install, regenerate the
74
+ lockfile) rather than leaving it for someone else, and say so in `CHANGES`.
68
75
 
69
76
  ## Hard constraints
70
77
 
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: tiny-spec-build-reviewer
3
- description: Independently reviews a single finished task — runs the project's real gate end-to-end and checks the code against the constitution and the task's acceptance. Blind to how the code was written. Returns PASS/FAIL plus findings. Spawned (one per task) by tiny-spec-build. Does not fix code, plan, spawn agents, or invoke skills.
3
+ description: Independently reviews a single finished task — runs the project's real gate at the scope it was given (full or scoped, escalating when scoped can't carry the verdict), exercises the acceptance end-to-end, and checks the code against the constitution. Blind to how the code was written. Returns PASS/FAIL plus findings. Spawned (one per task) by tiny-spec-build. Does not fix code, plan, spawn agents, or invoke skills.
4
4
  tools: Read, Write, Edit, Bash, Grep, Glob
5
5
  ---
6
6
 
@@ -17,6 +17,9 @@ back to `tiny-spec-build`; return data, not pleasantries.
17
17
  - the full **constitution** (`constitution.md`) — especially **Guiding invariants**,
18
18
  **Definition of Done**, and **Verification commands**, plus the **Design system**
19
19
  token table if the project has one;
20
+ - the **gate scope** for this task — `full` or `scoped`. `tiny-spec-build` sets it; it
21
+ knows where this task sits in the run and you don't. If the brief names no scope,
22
+ treat it as `full`;
20
23
  - the project's **memory** if any (`memory.md`) — operational lessons (e.g. the
21
24
  gate needs the package installed first); honor them so you don't false-fail on a
22
25
  known precondition;
@@ -32,10 +35,26 @@ acceptance and the constitution — verified, not inferred?**
32
35
  1. **Read the changed code.** Check it against the constitution: does it honor the
33
36
  **Guiding invariants**, match the **Style** and **Layout**, meet the
34
37
  **Definition of Done**? Note any violation as a finding.
35
- 2. **Run the real gate.** Execute the constitution's **Verification commands**
36
- end-to-end (install → lint → test → build → run, as applicable) from a clean
37
- state, after the documented setup — not a test-runner shortcut. Capture the
38
- real output.
38
+ 2. **Run the real gate, at the scope you were given.** Capture the real output either
39
+ way — never a test-runner shortcut, never an inferred result.
40
+
41
+ - **`full`** — execute the constitution's **Verification commands** end-to-end
42
+ (install → lint → test → build → run, as applicable) from a clean state, after the
43
+ documented setup.
44
+ - **`scoped`** — run `lint:` and `test:` against the working tree, and rely on the
45
+ existing install/build state rather than rebuilding it from clean. Step 3 below is
46
+ **not** scoped: you still exercise the acceptance end-to-end, black-box, with real
47
+ input. A scoped gate skips re-proving the toolchain; it never skips proving the task.
48
+
49
+ **Escalate rather than guess.** Move a `scoped` gate up to `full` the moment the
50
+ narrower one can't carry the verdict — you need a build artifact that isn't there or
51
+ looks stale, `install`/`build` state seems inconsistent with the changed files, a
52
+ command fails in a way that might be environmental, or you simply cannot exercise the
53
+ acceptance without building. Escalation is always allowed and always safe. **Passing a
54
+ task on evidence you found insufficient is not** — that is the one thing this whole
55
+ role exists to prevent. Never narrow a `full` gate to `scoped`.
56
+
57
+ Say which scope you actually ran in `GATE:`, including that you escalated and why.
39
58
  3. **Exercise the acceptance.** Trigger the task's stated outcome the most
40
59
  black-box way available (CLI > HTTP > public API) with realistic input,
41
60
  including a negative case if the acceptance implies a boundary or rejection.
@@ -158,7 +177,8 @@ Never spawn subagents or invoke skills.
158
177
  ```
159
178
  TASK: <task id>
160
179
  VERDICT: PASS | FAIL
161
- GATE: <the Verification commands you ran + the real result (pass/fail + key output)>
180
+ GATE: <scope: full | scoped (+ "escalated from scoped: <why>" if you moved it up);
181
+ the Verification commands you ran + the real result (pass/fail + key output)>
162
182
  ACCEPTANCE: <how you exercised it + the observed effect, or why you couldn't>
163
183
  DESIGN: <omit unless the task carried `design:`. The D<n> checked, the measurements
164
184
  you read back vs the tokens they should match, and which states you exercised —
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
4
4
 
5
5
  [project]
6
6
  name = "tiny-spec"
7
- version = "1.0.0"
7
+ version = "1.1.0"
8
8
  description = "A tiny, opinionated take on spec-driven development."
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.9"
@@ -68,7 +68,11 @@ Spawn one **`tiny-spec-build-executor`** with a fresh, self-contained prompt:
68
68
  `.spec/<active>/SPEC.md` plus its `export:` path, so the executor can look at the
69
69
  design instead of guessing at it;
70
70
  - only the specific existing files the task starts from, named explicitly (so it
71
- edits with the real current contents, not blind).
71
+ edits with the real current contents, not blind);
72
+ - the **previous passed task's `CHANGES` list**, if there is one, labelled as what the
73
+ last task left behind. Every executor starts cold and re-derives the codebase; naming
74
+ the files that just moved is the cheapest way to cut that, and you already have the
75
+ list in hand. Paths only — not the previous task's description, findings, or report.
72
76
 
73
77
  Do **not** pass the plan, sibling tasks, or other chatter. It returns a structured
74
78
  report (`STATUS`, `CHANGES`, `DECISIONS`, `BLOCKER`). A `STATUS: blocked` →
@@ -80,11 +84,29 @@ Spawn one **`tiny-spec-build-reviewer`**, **blind to step 2**, with:
80
84
  - the task id, description, and **acceptance**;
81
85
  - the **whole** `.spec/constitution.md`;
82
86
  - the list of changed files (from the executor's `CHANGES`) to read;
83
- - the **Verification commands** from the constitution to run;
87
+ - the **Verification commands** from the constitution to run, and the **gate scope**
88
+ for this task — `full` or `scoped` (decide it as below; you have the run context,
89
+ the reviewer does not);
84
90
  - **if the task has a `design:` field** — the same `D<n>` entry and `export:` path
85
91
  you gave the executor, so it grades against the contract rather than its taste.
86
92
 
87
- It runs the real gate end-to-end, checks the code against the constitution's
93
+ **Gate scope.** A clean `install → build` proves the same thing on task 7 that it proved
94
+ on task 1 unless a task changed what it installs or builds, so running it every time buys
95
+ little and costs the most wall-clock in the loop. Name the scope in the brief:
96
+
97
+ - **`full`** — the whole gate from a clean state (install → lint → test → build → run).
98
+ Use it on: the **first** task of this run; any task whose `CHANGES` touch dependency or
99
+ build configuration (`package.json`/lockfiles, `pyproject.toml`, `go.mod`, `Cargo.toml`,
100
+ `Dockerfile`, CI config, build scripts); any task carrying `design:`; and the **last**
101
+ task in `tasks.md`.
102
+ - **`scoped`** — otherwise. `lint:` + `test:` plus the acceptance exercised end-to-end,
103
+ black-box, exactly as on a full gate. The acceptance exercise is never scoped away; it
104
+ is the whole reason the reviewer exists.
105
+
106
+ The reviewer may **escalate** `scoped` to `full` on its own and will say so — that
107
+ direction is always safe. It may never go the other way.
108
+
109
+ It runs the gate at the scope you named, checks the code against the constitution's
88
110
  **Definition of Done** and **invariants**, confirms the **acceptance** actually
89
111
  holds (exercised, not inferred), and returns `VERDICT: PASS | FAIL` + findings.
90
112
  On a `design:` task it also runs the constitution's `visual:` command, measures the
@@ -94,11 +116,15 @@ numbers can't reach, like an element that measures perfectly and shows up invisi
94
116
  It fails on what it measured or saw, and merely flags what is taste.
95
117
 
96
118
  > Why independent: unit-green ≠ working, and the author is the worst judge of its
97
- > own blind spots. The reviewer running the gate from a clean state is the
98
- > safeguard that keeps scope and quality honest without an ownership contract.
119
+ > own blind spots. The reviewer exercising the acceptance black-box — on every task,
120
+ > at either scope — is the safeguard that keeps scope and quality honest without an
121
+ > ownership contract.
99
122
 
100
123
  ### 4. CONVERGE (on FAIL)
101
- Re-dispatch the **executor** with the reviewer's findings appended to its brief.
124
+ Re-dispatch the **executor** with the reviewer's findings appended to its brief, then
125
+ review again — recomputing the **gate scope** from the fix's `CHANGES`, and using `full`
126
+ if the failure was in the gate itself or the reviewer escalated last time. A task that
127
+ already failed once has earned the wider check.
102
128
  Bound this to **2 fix attempts**. If it still fails after that, stop and **halt
103
129
  `exhausted`** (see **Halting** below) — don't keep grinding or hand-fix past the loop
104
130
  silently. `exhausted` is recorded as a blocker, because a task that three attempts
@@ -274,6 +300,12 @@ When every task in `tasks.md` is `[x]`:
274
300
  setup — install/build, not a test-runner shortcut). There is no separate
275
301
  verify skill — this final smoke confirms the requirements actually work end-to-end, not
276
302
  just that tasks are ticked.
303
+
304
+ **This is the authoritative clean-state run for the whole build**, and it is never
305
+ scoped or skipped. Per-task gates run at the scope step 3 named; this one always runs
306
+ everything from clean. A build or install regression that a `scoped` task gate did not
307
+ re-prove surfaces here — which is why a red final smoke is a gap to fix or route
308
+ upstream, never a pass and never `done`.
277
309
  2. **Report** — the terminal state by name (`done` here, since every task is `[x]`
278
310
  and the smoke passed), what was built, the commits made (with the branch), and any
279
311
  open `decisions.md` items (blockers, tasks unchecked by a reconcile). If the final
@@ -138,25 +138,59 @@ past it into a task list — say what's unresolved and route back to `tiny-spec-
138
138
  ## Step 3 — slice the approach into tasks
139
139
 
140
140
  Walk the `## Approach` you just wrote and break it into tasks. No waves, no
141
- parallelism, no `owns:` contracts — tasks run one at a time, top to bottom. Each task is:
142
-
143
- - **Small and independently checkable** — one slice a single executor can finish and a
144
- reviewer can grade in one pass. If you can't write a one-line acceptance for it, it's
145
- too big — split it.
141
+ parallelism, no `owns:` contracts — tasks run one at a time, top to bottom.
142
+
143
+ **The unit is one coherent commit** — what a competent engineer does in one focused
144
+ sitting and commits as a single unit of work. That typically touches several files and
145
+ often satisfies several `REQ-N` at once. It is the size you would open as one reviewable
146
+ change, not the smallest thing you could name.
147
+
148
+ **Know what a task costs, so size has something to trade against.** Every task you write
149
+ spends two cold-start agents that must re-orient in the codebase from scratch, a reviewer
150
+ that runs the gate and exercises the acceptance end-to-end, and two commits. That
151
+ overhead is **fixed** — it does not shrink for a small task. A task the executor
152
+ finishes in thirty seconds still pays all of it. Size each task so the work inside it
153
+ clearly outweighs the machinery around it.
154
+
155
+ So:
156
+
157
+ - **Split on independent failure, not on sentence length.** Split a task when it carries
158
+ two **unrelated** observable outcomes that could fail independently of each other. Do
159
+ **not** split because the acceptance got long — a good acceptance is usually several
160
+ clauses covering the happy path *and* its negatives. See `T4` in the `examples/todo-cli`
161
+ task list: two CLI commands, four requirements, and three negative cases, in one task,
162
+ with one acceptance. That is the calibration point, not the exception.
163
+ - **One task may satisfy several `REQ-N`** — `req:` takes a list. Coverage means every
164
+ requirement has a *home*, not that every requirement gets its *own* task. **A 1:1
165
+ REQ→task mapping is the single most common way this list comes out too granular.**
166
+ Group the requirements that one coherent change delivers together.
146
167
  - **Ordered so each builds on the last.** Tasks run sequentially, so a later task may
147
168
  freely assume an earlier task's code already exists. Put foundational work (types,
148
169
  schema, scaffolding) first. Order by dependency, not by guesswork.
149
- - **Right-sized, not fragmented.** Don't split a cohesive change into five files' worth
150
- of micro-tasks just to look granular. Earned ceremony: fewer, meaningful tasks beat
151
- many trivial ones.
170
+
171
+ **Smells that mean you sliced below the commit line** — fold each of these back into the
172
+ task it belongs to:
173
+
174
+ - a task that only defines types, interfaces, or schema with no behavior behind them;
175
+ - a task that only adds tests for the task before it (the constitution's **Definition of
176
+ Done** already requires the tests to ship with the code);
177
+ - one task per file, or one task per function;
178
+ - a "wire it up" / "integrate the pieces" task trailing the pieces it wires.
179
+
180
+ **Count is a smell, not a cap.** A story sized the way `tiny-spec-scope` describes
181
+ usually lands in **3–6 tasks**. If you are past about eight, re-read the list: you have
182
+ either sliced below the commit line, or the story itself was too big and should have been
183
+ split upstream. Check the list against that; do **not** enforce a number, and never drop
184
+ or merge coverage just to hit one.
152
185
 
153
186
  For each task, write:
154
187
 
155
188
  ```
156
189
  - [ ] T<n> — <imperative description>
157
- - acceptance: <one user-observable outcome that proves it's done>
190
+ - acceptance: <one user-observable outcome that proves it's done — happy path and the
191
+ negatives that bound it, in one entry>
158
192
  - type: feat # optional; Conventional Commit type (defaults to feat)
159
- - req: REQ-n # optional; the REQ-N this task delivers
193
+ - req: REQ-n, REQ-n # optional; the REQ-N this task delivers — a list, not one
160
194
  - design: D-n # optional; the SPEC.md D<n> screen this task builds — arms the visual gate
161
195
  - pause: <why> # optional; halt the build before this task so a human looks first
162
196
  - files: <comma-separated hint of files it will touch>
@@ -164,7 +198,11 @@ For each task, write:
164
198
 
165
199
  The **acceptance** is what the reviewer checks against — make it observable
166
200
  ("`spec --version` prints the version and exits 0"), not internal ("version logic
167
- added"). **type** picks the Conventional Commit type `tiny-spec-build` uses for this
201
+ added"). It states **one outcome**, but one outcome is not one clause: spell out the
202
+ happy path and the negatives that bound it in the same acceptance, separated by
203
+ semicolons. A long acceptance is a well-specified task, not an oversized one — it is the
204
+ *number of unrelated things that could fail* that decides whether to split, not the
205
+ length of the line. **type** picks the Conventional Commit type `tiny-spec-build` uses for this
168
206
  task's code commit (`feat | fix | docs | refactor | test | chore | build | ci | perf | style`);
169
207
  set it when the task is clearly not a feature, otherwise omit and it defaults to `feat`.
170
208
  **req** ties the task to the requirement it satisfies (traceability). The **files** line
@@ -211,7 +249,9 @@ anyway. A policy that matches nothing here is not an error — say so and move o
211
249
  than stretching a task to fit it.
212
250
 
213
251
  Cover **every** part of the approach — together the tasks must deliver all `REQ-N`.
214
- Don't leave a requirement with no task. Likewise, if `SPEC.md` has a `## Design`
252
+ Don't leave a requirement with no task. Coverage is about requirements having a home,
253
+ not about the shape of the mapping: several `REQ-N` on one task is the normal case, and
254
+ a task per requirement is the anti-pattern. Likewise, if `SPEC.md` has a `## Design`
215
255
  section, every `D<n>` in it needs at least one task carrying that `design:` reference —
216
256
  a screen nobody is graded against is a screen that will be built wrong.
217
257
 
@@ -236,25 +276,31 @@ updated: <ISO date>
236
276
 
237
277
  ## Tasks
238
278
 
239
- - [ ] T1 — <one small, independently-checkable slice of work>
240
- - acceptance: <one user-observable outcome that proves T1 is done>
279
+ - [ ] T1 — <one coherent commit's worth of foundational work; usually several files>
280
+ - acceptance: <the observable outcome; happy path; and the negative case that bounds it>
241
281
  - type: feat # optional; Conventional Commit type for this task's commit (defaults to feat)
242
- - req: REQ-1 # optional; the REQ-N this task delivers
243
- - files: <path, path>
282
+ - req: REQ-1, REQ-2 # optional; the REQ-N this task delivers — several is normal
283
+ - files: <path, path, path>
244
284
 
245
- - [ ] T2 — <next slice; assume T1's code exists>
246
- - acceptance: <observable outcome>
285
+ - [ ] T2 — <next coherent change; assume T1's code exists>
286
+ - acceptance: <observable outcome; plus what it rejects and how it fails>
247
287
  - type: feat
248
- - req: REQ-2
288
+ - req: REQ-3, REQ-4, REQ-5
249
289
  - design: D1 # optional; only on tasks that build the visible surface
250
290
  - files: <path, path>
251
291
 
252
292
  - [ ] T3 — <…>
253
293
  - acceptance: <observable outcome>
294
+ - req: REQ-6
254
295
  - pause: <optional; what to check before this runs — irreversible work only>
255
296
  - files: <path, path>
256
297
  ```
257
298
 
299
+ The `req:` lists above are the shape to aim for, not filler: a handful of tasks each
300
+ carrying the requirements one coherent change delivers. A skeleton filled in as
301
+ `REQ-1`, `REQ-2`, `REQ-3` down a column of single-requirement tasks is the granularity
302
+ failure described in Step 3.
303
+
258
304
  ## Update mode (SPEC changed → PLAN and tasks are stale)
259
305
 
260
306
  When `PLAN.md` or `tasks.md` is `status: stale`, reconcile **both in one pass** — they
@@ -5,4 +5,4 @@ into your Claude Code config directory (``~/.claude/`` by default). It adds no
5
5
  runtime behavior to the suite itself.
6
6
  """
7
7
 
8
- __version__ = "0.5.0"
8
+ __version__ = "1.1.0"
File without changes
File without changes
File without changes
File without changes