tiny-spec 1.0.1__tar.gz → 1.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: tiny-spec
3
- Version: 1.0.1
3
+ Version: 1.2.0
4
4
  Summary: A tiny, opinionated take on spec-driven development.
5
5
  Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
6
6
  Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
@@ -449,10 +449,19 @@ unnoticed.
449
449
  1. Plan the task against the constitution (inline, brief).
450
450
  2. Implement it with a fresh `tiny-spec-build-executor` agent.
451
451
  3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
452
- end to end and grades against the constitution and the task's acceptance.
452
+ and grades against the constitution and the task's acceptance.
453
453
  4. On pass, commit the code plus a checklist tick. On fail, loop back to the
454
454
  executor with the findings. After two failed attempts the run halts `exhausted`.
455
455
 
456
+ Every task's acceptance is exercised end to end, black-box, by the reviewer — that part
457
+ never varies. What does vary is how much of the toolchain gets re-proven around it: a
458
+ clean `install → build` runs on the first task, the last one, anything touching
459
+ dependency or build config, and anything carrying `design:`; in between the reviewer runs
460
+ lint and tests against the existing build. It escalates back to the full gate whenever the
461
+ narrower one can't carry the verdict, and the authoritative clean run happens once more
462
+ over the whole project as the final smoke. Re-installing from scratch on every task
463
+ proved the same thing repeatedly and was the slowest part of the loop.
464
+
456
465
  ```mermaid
457
466
  flowchart TB
458
467
  SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
@@ -460,7 +469,7 @@ flowchart TB
460
469
  TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
461
470
  TASKS --> P[Plan task]
462
471
  P --> I[Implement<br/>executor]
463
- I --> R[Review + run gate<br/>reviewer]
472
+ I --> R[Review + run gate<br/>full or scoped<br/>reviewer]
464
473
  I -->|blocker| B
465
474
  R -->|pass| C[Commit + tick]
466
475
  C --> TASKS
@@ -489,15 +498,38 @@ resumes from the checklist state.
489
498
  ### Working a whole list
490
499
 
491
500
  Ask `/tiny-spec-run` to build — "build the backlog", "work through the breakdown",
492
- "spec it out and build it" — and it takes a list of stories and works them one after
493
- another. Per story it does the same four moves:
501
+ "spec it out and build it" — and it takes a list of stories and works them in batches.
502
+ Per story it does the same four moves:
494
503
 
495
504
  ```
496
505
  cut a branch from main → walk the chain → tiny-spec-build → merge back to main
497
506
  ```
498
507
 
499
- Then the next story. Each branch is cut **fresh from main**, so story 3 sees stories 1
500
- and 2 already merged — which is what makes an ordered list build correctly.
508
+ Each branch is cut **fresh from main**, so a later story sees the earlier ones already
509
+ merged — which is what makes an ordered list build correctly.
510
+
511
+ **Independent stories build at the same time.** A story can declare what it must follow
512
+ with a `needs:` line in `BREAKDOWN.md`; everything with no unmet `needs:` forms a batch
513
+ and builds **concurrently, one git worktree per story**, three at a time by default.
514
+ The batch merges, then the next one starts.
515
+
516
+ ```
517
+ ## Feature: interface
518
+
519
+ - Story: expose both helpers on a CLI slug: cli
520
+ - AC: `textkit slugify "Hi There"` prints "hi-there"
521
+ - needs: slugify, wordwrap
522
+ ```
523
+
524
+ Omit `needs:` when a story stands alone — that's the common case, and the field is meant
525
+ to be rare. A `needs:` you didn't need costs you parallelism forever; one you missed
526
+ costs a single merge conflict, which the run already catches and halts on. You can also
527
+ just name the set yourself at invocation ("build these three at once"), which overrides
528
+ the graph. A cycle, or a `needs:` naming a story that isn't there, stops the run rather
529
+ than being guessed past.
530
+
531
+ **Tasks *inside* a story never run in parallel.** They share files and each one assumes
532
+ the last landed, so they stay strictly sequential. Parallelism is across stories only.
501
533
 
502
534
  **The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
503
535
  each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
@@ -520,9 +552,13 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
520
552
  Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
521
553
  often, so the state is always named alongside what merged and what's still untouched.
522
554
 
523
- **A halt stops the whole run, not just that story.** Later stories in a list you wrote
524
- top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
525
- just produces a second, more confusing failure downstream.
555
+ **A halt stops the lane it happened in, and ends the run after that batch.** Its siblings
556
+ were declared independent, so they finish and merge — killing working lanes because one
557
+ failed throws away good work. But the run does not start the next batch: later stories
558
+ usually assume the earlier ones landed, so skipping ahead past a failure just produces a
559
+ second, more confusing failure downstream. With more than one lane you get each story's
560
+ own state, and the run's state is the worst of them — four green lanes and one `blocked`
561
+ is a `blocked` run.
526
562
 
527
563
  **Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
528
564
  the build halts *before* running it:
@@ -538,12 +574,15 @@ destructive file operations, a new dependency, an auth boundary, a public API co
538
574
  You can also give the run a standing policy up front ("halt before anything that touches
539
575
  auth") and it gets applied as each story's tasks are sliced.
540
576
 
541
- **What it will not do to your repo.** It runs exactly five git commands — `switch`,
542
- `switch -c`, `merge --no-ff`, `merge --abort`, and reads. It refuses to start on a dirty
543
- tree. It **never pushes**, never rebases, never resets, never deletes a branch, and
544
- never opens a PR. Merges are local, so a bad run is one `git reset` away; publishing
545
- stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
546
- the undo command rather than running it.
577
+ **What it will not do to your repo.** It runs exactly seven git commands — `switch`,
578
+ `switch -c`, `merge --no-ff`, `merge --abort`, `worktree add`, `worktree list`, and
579
+ reads. It refuses to start on a dirty tree, or if a worktree path it needs already
580
+ exists. It **never pushes**, never rebases, never resets, never deletes a branch, never
581
+ removes a worktree, and never opens a PR. Merges are local, so a bad run is one
582
+ `git reset` away; publishing stays yours. If the gate goes red after a merge it leaves
583
+ the merge alone and tells you the undo command rather than running it — and it hands
584
+ back the `git worktree remove` commands for the lanes instead of running those either,
585
+ since a halted lane's worktree is the tree you need to look at.
547
586
 
548
587
  **Walk away and come back.** Progress isn't written down, it's derived: a story whose
549
588
  ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
@@ -412,10 +412,19 @@ unnoticed.
412
412
  1. Plan the task against the constitution (inline, brief).
413
413
  2. Implement it with a fresh `tiny-spec-build-executor` agent.
414
414
  3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
415
- end to end and grades against the constitution and the task's acceptance.
415
+ and grades against the constitution and the task's acceptance.
416
416
  4. On pass, commit the code plus a checklist tick. On fail, loop back to the
417
417
  executor with the findings. After two failed attempts the run halts `exhausted`.
418
418
 
419
+ Every task's acceptance is exercised end to end, black-box, by the reviewer — that part
420
+ never varies. What does vary is how much of the toolchain gets re-proven around it: a
421
+ clean `install → build` runs on the first task, the last one, anything touching
422
+ dependency or build config, and anything carrying `design:`; in between the reviewer runs
423
+ lint and tests against the existing build. It escalates back to the full gate whenever the
424
+ narrower one can't carry the verdict, and the authoritative clean run happens once more
425
+ over the whole project as the final smoke. Re-installing from scratch on every task
426
+ proved the same thing repeatedly and was the slowest part of the loop.
427
+
419
428
  ```mermaid
420
429
  flowchart TB
421
430
  SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
@@ -423,7 +432,7 @@ flowchart TB
423
432
  TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
424
433
  TASKS --> P[Plan task]
425
434
  P --> I[Implement<br/>executor]
426
- I --> R[Review + run gate<br/>reviewer]
435
+ I --> R[Review + run gate<br/>full or scoped<br/>reviewer]
427
436
  I -->|blocker| B
428
437
  R -->|pass| C[Commit + tick]
429
438
  C --> TASKS
@@ -452,15 +461,38 @@ resumes from the checklist state.
452
461
  ### Working a whole list
453
462
 
454
463
  Ask `/tiny-spec-run` to build — "build the backlog", "work through the breakdown",
455
- "spec it out and build it" — and it takes a list of stories and works them one after
456
- another. Per story it does the same four moves:
464
+ "spec it out and build it" — and it takes a list of stories and works them in batches.
465
+ Per story it does the same four moves:
457
466
 
458
467
  ```
459
468
  cut a branch from main → walk the chain → tiny-spec-build → merge back to main
460
469
  ```
461
470
 
462
- Then the next story. Each branch is cut **fresh from main**, so story 3 sees stories 1
463
- and 2 already merged — which is what makes an ordered list build correctly.
471
+ Each branch is cut **fresh from main**, so a later story sees the earlier ones already
472
+ merged — which is what makes an ordered list build correctly.
473
+
474
+ **Independent stories build at the same time.** A story can declare what it must follow
475
+ with a `needs:` line in `BREAKDOWN.md`; everything with no unmet `needs:` forms a batch
476
+ and builds **concurrently, one git worktree per story**, three at a time by default.
477
+ The batch merges, then the next one starts.
478
+
479
+ ```
480
+ ## Feature: interface
481
+
482
+ - Story: expose both helpers on a CLI slug: cli
483
+ - AC: `textkit slugify "Hi There"` prints "hi-there"
484
+ - needs: slugify, wordwrap
485
+ ```
486
+
487
+ Omit `needs:` when a story stands alone — that's the common case, and the field is meant
488
+ to be rare. A `needs:` you didn't need costs you parallelism forever; one you missed
489
+ costs a single merge conflict, which the run already catches and halts on. You can also
490
+ just name the set yourself at invocation ("build these three at once"), which overrides
491
+ the graph. A cycle, or a `needs:` naming a story that isn't there, stops the run rather
492
+ than being guessed past.
493
+
494
+ **Tasks *inside* a story never run in parallel.** They share files and each one assumes
495
+ the last landed, so they stay strictly sequential. Parallelism is across stories only.
464
496
 
465
497
  **The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
466
498
  each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
@@ -483,9 +515,13 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
483
515
  Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
484
516
  often, so the state is always named alongside what merged and what's still untouched.
485
517
 
486
- **A halt stops the whole run, not just that story.** Later stories in a list you wrote
487
- top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
488
- just produces a second, more confusing failure downstream.
518
+ **A halt stops the lane it happened in, and ends the run after that batch.** Its siblings
519
+ were declared independent, so they finish and merge — killing working lanes because one
520
+ failed throws away good work. But the run does not start the next batch: later stories
521
+ usually assume the earlier ones landed, so skipping ahead past a failure just produces a
522
+ second, more confusing failure downstream. With more than one lane you get each story's
523
+ own state, and the run's state is the worst of them — four green lanes and one `blocked`
524
+ is a `blocked` run.
489
525
 
490
526
  **Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
491
527
  the build halts *before* running it:
@@ -501,12 +537,15 @@ destructive file operations, a new dependency, an auth boundary, a public API co
501
537
  You can also give the run a standing policy up front ("halt before anything that touches
502
538
  auth") and it gets applied as each story's tasks are sliced.
503
539
 
504
- **What it will not do to your repo.** It runs exactly five git commands — `switch`,
505
- `switch -c`, `merge --no-ff`, `merge --abort`, and reads. It refuses to start on a dirty
506
- tree. It **never pushes**, never rebases, never resets, never deletes a branch, and
507
- never opens a PR. Merges are local, so a bad run is one `git reset` away; publishing
508
- stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
509
- the undo command rather than running it.
540
+ **What it will not do to your repo.** It runs exactly seven git commands — `switch`,
541
+ `switch -c`, `merge --no-ff`, `merge --abort`, `worktree add`, `worktree list`, and
542
+ reads. It refuses to start on a dirty tree, or if a worktree path it needs already
543
+ exists. It **never pushes**, never rebases, never resets, never deletes a branch, never
544
+ removes a worktree, and never opens a PR. Merges are local, so a bad run is one
545
+ `git reset` away; publishing stays yours. If the gate goes red after a merge it leaves
546
+ the merge alone and tells you the undo command rather than running it — and it hands
547
+ back the `git worktree remove` commands for the lanes instead of running those either,
548
+ since a halted lane's worktree is the tree you need to look at.
510
549
 
511
550
  **Walk away and come back.** Progress isn't written down, it's derived: a story whose
512
551
  ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
@@ -15,6 +15,12 @@ pleasantries.
15
15
 
16
16
  Everything you need and nothing you don't:
17
17
 
18
+ - **the working directory** to operate in. Every path you read, write, or run a command
19
+ against resolves against it. It may be a **git worktree** rather than the main checkout
20
+ — a build can run several stories at once, each in its own worktree. **Never read or
21
+ write outside the directory you were given**, and never `cd` to a sibling worktree to
22
+ "check something": another story is being built there right now, and what you find will
23
+ be wrong by the time you act on it;
18
24
  - the **task id**, **description**, and **acceptance** (the outcome that proves it done);
19
25
  - a **`files:` hint** — likely paths to touch (guidance, not a hard boundary);
20
26
  - the full **constitution** (`constitution.md`): Style, Engineering standards,
@@ -24,7 +30,10 @@ Everything you need and nothing you don't:
24
30
  so you don't re-learn a pitfall a past run already paid for;
25
31
  - **if the task carries `design:`** — that screen's `D<n>` entry from `SPEC.md` and
26
32
  the path to its committed export;
27
- - the specific existing files that are your starting point, named explicitly.
33
+ - the specific existing files that are your starting point, named explicitly;
34
+ - sometimes, **the files the previous task left behind** — paths only. Start there when
35
+ orienting: it is the most recently moved ground in the codebase and the most likely
36
+ place your task connects to.
28
37
 
29
38
  You are **blind to the workflow, not to the codebase.** You don't get the plan,
30
39
  other tasks, or shared state. But the named files are a launch point, not the whole
@@ -64,7 +73,11 @@ violated.
64
73
  6. You MAY run a **narrow self-check** of your own work (the one test file you
65
74
  wrote, a syntax/import check). You do **not** need to run the full gate — the
66
75
  independent **reviewer** runs the authoritative Verification commands next.
67
- Leave the tree in a clean, buildable state for it.
76
+ **Leave the tree in a clean, buildable, installed state for it** — the reviewer may
77
+ run against the existing build rather than rebuilding from scratch, so a tree you
78
+ left half-migrated or with a stale artifact costs the task a review cycle. If you
79
+ changed dependencies or build configuration, apply them (install, regenerate the
80
+ lockfile) rather than leaving it for someone else, and say so in `CHANGES`.
68
81
 
69
82
  ## Hard constraints
70
83
 
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: tiny-spec-build-reviewer
3
- description: Independently reviews a single finished task — runs the project's real gate end-to-end and checks the code against the constitution and the task's acceptance. Blind to how the code was written. Returns PASS/FAIL plus findings. Spawned (one per task) by tiny-spec-build. Does not fix code, plan, spawn agents, or invoke skills.
3
+ description: Independently reviews a single finished task — runs the project's real gate at the scope it was given (full or scoped, escalating when scoped can't carry the verdict), exercises the acceptance end-to-end, and checks the code against the constitution. Blind to how the code was written. Returns PASS/FAIL plus findings. Spawned (one per task) by tiny-spec-build. Does not fix code, plan, spawn agents, or invoke skills.
4
4
  tools: Read, Write, Edit, Bash, Grep, Glob
5
5
  ---
6
6
 
@@ -13,10 +13,18 @@ back to `tiny-spec-build`; return data, not pleasantries.
13
13
 
14
14
  ## What you receive (the context contract)
15
15
 
16
+ - **the working directory** to operate in — every path and every gate command resolves
17
+ against it. It may be a **git worktree** rather than the main checkout, since a build
18
+ can run several stories at once. **Never read, write, or run a gate outside the
19
+ directory you were given**: a sibling worktree holds a different story mid-build, and
20
+ measuring it would make your verdict meaningless;
16
21
  - the **task id**, **description**, and **acceptance** (the outcome that must hold);
17
22
  - the full **constitution** (`constitution.md`) — especially **Guiding invariants**,
18
23
  **Definition of Done**, and **Verification commands**, plus the **Design system**
19
24
  token table if the project has one;
25
+ - the **gate scope** for this task — `full` or `scoped`. `tiny-spec-build` sets it; it
26
+ knows where this task sits in the run and you don't. If the brief names no scope,
27
+ treat it as `full`;
20
28
  - the project's **memory** if any (`memory.md`) — operational lessons (e.g. the
21
29
  gate needs the package installed first); honor them so you don't false-fail on a
22
30
  known precondition;
@@ -32,10 +40,38 @@ acceptance and the constitution — verified, not inferred?**
32
40
  1. **Read the changed code.** Check it against the constitution: does it honor the
33
41
  **Guiding invariants**, match the **Style** and **Layout**, meet the
34
42
  **Definition of Done**? Note any violation as a finding.
35
- 2. **Run the real gate.** Execute the constitution's **Verification commands**
36
- end-to-end (install → lint → test → build → run, as applicable) from a clean
37
- state, after the documented setup — not a test-runner shortcut. Capture the
38
- real output.
43
+ 2. **Run the real gate, at the scope you were given.** Capture the real output either
44
+ way — never a test-runner shortcut, never an inferred result.
45
+
46
+ - **`full`** — execute the constitution's **Verification commands** end-to-end
47
+ (install → lint → test → build → run, as applicable) from a clean state, after the
48
+ documented setup.
49
+ - **`scoped`** — run `lint:` and `test:` against the working tree, and rely on the
50
+ existing install/build state rather than rebuilding it from clean. Step 3 below is
51
+ **not** scoped: you still exercise the acceptance end-to-end, black-box, with real
52
+ input. A scoped gate skips re-proving the toolchain; it never skips proving the task.
53
+
54
+ **Escalate rather than guess.** Move a `scoped` gate up to `full` the moment the
55
+ narrower one can't carry the verdict — you need a build artifact that isn't there or
56
+ looks stale, `install`/`build` state seems inconsistent with the changed files, a
57
+ command fails in a way that might be environmental, or you simply cannot exercise the
58
+ acceptance without building. Escalation is always allowed and always safe. **Passing a
59
+ task on evidence you found insufficient is not** — that is the one thing this whole
60
+ role exists to prevent. Never narrow a `full` gate to `scoped`.
61
+
62
+ **A Verification command that cannot run in this environment is a `FAIL` and a
63
+ blocker — never a `PASS`.** If a documented command is red for a reason the task's
64
+ code cannot fix (the interpreter is too old, a tool isn't installed, the command was
65
+ never runnable as written), say exactly that in `GATE:`, return `FAIL`, and name the
66
+ **constitution** as the upstream document to fix. Proving the task green *somewhere
67
+ else* — a fresh venv, a container, an install you fixed by hand — is **not** a pass:
68
+ you verified a different environment than the one the gate names, and the next task
69
+ will hit the identical red. Diagnosing the cause and demonstrating the task's own code
70
+ is sound is genuinely useful; put it in `FINDINGS` so the fix is cheap. It does not
71
+ change the verdict. A constitution whose gate cannot run is the single most dangerous
72
+ file in the project, because every later task inherits the same false signal.
73
+
74
+ Say which scope you actually ran in `GATE:`, including that you escalated and why.
39
75
  3. **Exercise the acceptance.** Trigger the task's stated outcome the most
40
76
  black-box way available (CLI > HTTP > public API) with realistic input,
41
77
  including a negative case if the acceptance implies a boundary or rejection.
@@ -59,25 +95,23 @@ acceptance and the constitution — verified, not inferred?**
59
95
  - **A selector that matches nothing is a `FAIL`, never a skip.** Either the code
60
96
  didn't build the element or it named it something else — both are real, and both
61
97
  are invisible if you quietly move on. Report the selector and that it was absent.
62
- - Compare each element's numbers to the tokens its row names. Report concrete
63
- deltas ("heading is 28px, `type.heading.lg` is 24px"; "padding 19px is not on
64
- the `space.*` scale"). **Do not pixel-diff the screenshot** — font antialiasing
65
- makes image comparison flaky enough that the check gets ignored, which is
66
- exactly how visual gates die.
67
- - **Check `layout:`** — the arrangement, max width, and the **order** it names.
68
- Use the bounding rectangles: elements listed in order should appear in that
69
- order down the page (or across it, for a row). Every token can be correct on an
70
- element that is in the wrong place.
71
- - **Exercise every state the entry names** — empty, loading, error, success.
72
- Drive the UI into each one and observe what changes; **finding the word in the
73
- source is not evidence**, and it false-passes routinely (a comment saying the
74
- loading state is missing contains "loading"). A surface that renders its happy
75
- path and nothing else is a fail, not a nit.
76
- - Finally, **judge the render against the export.** Everything above proves the
77
- numbers are right. None of it can see an element that is present, on-token, and
78
- invisible — so now look. `Read` each `SCREENSHOT <state> <path>` the `visual:`
79
- command printed, `Read` the `D<n>`'s `export:` image, and grade **every state you
80
- have a screenshot for**, not just the happy path, on four lines:
98
+ - Compare each element's numbers to the tokens its row names, reporting concrete
99
+ deltas ("heading is 28px, `type.heading.lg` is 24px"; "padding 19px is not on the
100
+ `space.*` scale"). **Do not pixel-diff the screenshot** — font antialiasing makes
101
+ image comparison flaky enough that the check gets ignored, which is how visual
102
+ gates die.
103
+ - **Check `layout:`** — the arrangement, max width, and the **order** it names, from
104
+ the bounding rectangles. Every token can be correct on an element in the wrong place.
105
+ - **Exercise every state the entry names** — empty, loading, error, success. Drive the
106
+ UI into each one and observe what changes; **finding the word in the source is not
107
+ evidence**, and it false-passes routinely (a comment saying the loading state is
108
+ missing contains "loading"). A surface that renders its happy path and nothing else
109
+ is a fail, not a nit.
110
+ - Finally, **judge the render against the export.** Everything above proves the numbers
111
+ are right; none of it can see an element that is present, on-token, and invisible —
112
+ so now look. `Read` each `SCREENSHOT <state> <path>` the `visual:` command printed
113
+ and the `D<n>`'s `export:` image, and grade **every state you have a screenshot
114
+ for**, not just the happy path, on four lines:
81
115
 
82
116
  1. **Presence** — is every `elements:` row actually *visible* in the render? At
83
117
  `opacity: 0`, zero height, clipped out of view, hidden behind a sibling, or the
@@ -85,24 +119,21 @@ acceptance and the constitution — verified, not inferred?**
85
119
  not there. **FAIL** — this is why the step exists.
86
120
  2. **Legibility & occlusion** — text clipped, truncated mid-word, overlapping
87
121
  another element, or on a background it can't be read against. **FAIL.**
88
- 3. **Correspondence** — does the render show the same screen as the export: the
89
- same regions, in the reading order `layout:` names? A whole region missing is a
90
- **FAIL**; a stylistic difference is a **flag**.
122
+ 3. **Correspondence** — the same regions as the export, in the reading order
123
+ `layout:` names? A whole region missing is a **FAIL**; a stylistic difference is
124
+ a **flag**.
91
125
  4. **Hierarchy & polish** — emphasis, balance, crowding, alignment. **Always a
92
126
  flag**, never a fail.
93
127
 
94
- Three rules bound it:
95
-
96
- - **The numbers beat your eye on anything they already measured.** Padding that is
97
- on the `space.*` scale but looks cramped is a `flag:`. A color that is exactly
98
- its token but looks washed out is a `flag:`. You may fail only on what
99
- measurement *cannot* see. Contradicting your own numbers sends the executor a
100
- task it cannot fix, and the loop is bounded at two attempts.
101
- - **This is still not a pixel diff.** The export is usually a wireframe — judge
102
- structure and legibility, never visual identity.
103
- - **Cite what you saw.** Name the state whose screenshot the finding came from and
104
- what was in it ("state `error`: caption present in DOM but renders at opacity 0").
105
- An uncited visual claim reads as an opinion and gets ignored.
128
+ Three rules bound it. **The numbers beat your eye on anything they already
129
+ measured** — on-scale padding that looks cramped, or an exact token that looks
130
+ washed out, is a `flag:`; you may fail only on what measurement *cannot* see, since
131
+ contradicting your own numbers sends the executor a task it cannot fix and the loop
132
+ is bounded at two attempts. **This is still not a pixel diff** — the export is
133
+ usually a wireframe, so judge structure and legibility, never visual identity. And
134
+ **cite what you saw**, naming the state the finding came from ("state `error`:
135
+ caption present in DOM but renders at opacity 0") — an uncited visual claim reads as
136
+ an opinion and gets ignored.
106
137
 
107
138
  **If the command printed no `SCREENSHOT` line**, do not run this sub-step and do
108
139
  not eyeball a substitute. Grade on steps 1–3 above, write `judge: not run — visual:
@@ -158,7 +189,8 @@ Never spawn subagents or invoke skills.
158
189
  ```
159
190
  TASK: <task id>
160
191
  VERDICT: PASS | FAIL
161
- GATE: <the Verification commands you ran + the real result (pass/fail + key output)>
192
+ GATE: <scope: full | scoped (+ "escalated from scoped: <why>" if you moved it up);
193
+ the Verification commands you ran + the real result (pass/fail + key output)>
162
194
  ACCEPTANCE: <how you exercised it + the observed effect, or why you couldn't>
163
195
  DESIGN: <omit unless the task carried `design:`. The D<n> checked, the measurements
164
196
  you read back vs the tokens they should match, and which states you exercised —
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
4
4
 
5
5
  [project]
6
6
  name = "tiny-spec"
7
- version = "1.0.1"
7
+ version = "1.2.0"
8
8
  description = "A tiny, opinionated take on spec-driven development."
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.9"
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: tiny-spec-build
3
- description: Build the spec — run the per-task loop plan→implement→review→commit, one task at a time. Implements with a fresh executor, grades with an independent reviewer running the real gate, commits per passed task, keeps a lean memory. Halts on a blocker, a spent convergence budget, a task's pause: point, or a genuine fork, and names which. Resumes from the checkbox state.
3
+ description: Build the spec — run the per-task loop plan→implement→review→commit, one task at a time. Implements with a fresh executor, grades with an independent reviewer running the real gate, commits per passed task, keeps a lean memory. Can be handed several lanes (one worktree + slug per independent story) and runs them concurrently, one task at a time within each. Halts on a blocker, a spent convergence budget, a task's pause: point, or a genuine fork, and names which. Resumes from the checkbox state.
4
4
  ---
5
5
 
6
6
  # tiny-spec-build
@@ -19,16 +19,14 @@ alongside this skill (see the suite README).
19
19
 
20
20
  ## Inputs
21
21
 
22
- 1. **Resolve the active ticket dir** from the current git branch: the `.spec/<slug>/`
23
- whose slug matches the branch name (one branch per ticket). If none matches, use
24
- the sole ticket dir if there's exactly one; else ask which. Call it `<active>`.
25
- **Two cases pre-empt that order — ask instead of applying it:** more than one dir
26
- matches the branch (there is no defined tie-break, and inventing one here would
27
- silently disagree with every other skill), or ticket dirs exist while you are on
28
- `main`/`master` with no name match (the usual cause is a forgotten `git switch`, and
29
- the sole-dir fallback would otherwise swallow it). Detached HEAD or no git repo is a
30
- **degraded** case, not an ask case — branch match is simply unavailable, so fall
31
- through to sole-dir and ask as written.
22
+ 1. **Resolve the active ticket dir** (call it `<active>`), in order: the `.spec/<slug>/`
23
+ whose slug matches the current git branch (one branch per ticket); else the sole
24
+ ticket dir if exactly one exists; else ask. **Ask instead** when more than one dir
25
+ matches the branch, or when ticket dirs exist while you are on `main`/`master` with no
26
+ name match — neither has a safe tie-break. Detached HEAD or no git repo is
27
+ **degraded**, not an ask: fall through to sole-dir and ask as written.
28
+ In **multi-lane mode** (below) this step does not run: each lane is handed its
29
+ worktree and slug explicitly.
32
30
  2. Read `.spec/constitution.md` (**the shared constitution**), `.spec/memory.md` if
33
31
  it exists (**shared**), and `.spec/<active>/tasks.md`. The constitution + memory
34
32
  get injected **whole** into every executor and reviewer. Also note the `ticket`
@@ -37,6 +35,52 @@ alongside this skill (see the suite README).
37
35
  `tiny-spec-plan` to reconcile first.
38
36
  4. Pick the **first unchecked `[ ]`** task. If all are `[x]`, jump to **Completion**.
39
37
 
38
+ ## Multi-lane mode (optional — several stories at once)
39
+
40
+ A caller may hand you **lanes** instead of a single active dir: one or more
41
+ `(worktree path, slug)` pairs, each an independent story in its own git worktree.
42
+ `tiny-spec-run` does this for a batch of stories whose `needs:` say they don't depend on
43
+ each other. **With one lane, or none, everything below is inert and this skill behaves
44
+ exactly as it always has** — that is the common case and it must not drift.
45
+
46
+ Given lanes, run them **in rounds**:
47
+
48
+ 1. For every live lane, pick its first `[ ]` task and build the executor brief exactly as
49
+ step 2 describes — resolved against **that lane's worktree**, not the main checkout.
50
+ 2. **Dispatch every lane's executor in one message**, so they run concurrently.
51
+ 3. **Dispatch every lane's reviewer in one message**, once the executors are back.
52
+ 4. Commit and tick per passing lane, in that lane's worktree.
53
+
54
+ Then start the next round with whichever lanes are still live. A fast lane waits at the
55
+ round boundary for a slow one; that barrier is deliberate, and much simpler to follow than
56
+ letting lanes free-run.
57
+
58
+ **Within a lane nothing changes.** Tasks still run one at a time, top to bottom, with the
59
+ same convergence bound, the same gate scope rules, and the same halting states. There is
60
+ no parallelism *inside* a story and no `owns:` contract — tasks in one story share files
61
+ and assume their predecessors landed, which is exactly why they stay sequential.
62
+
63
+ **A halt stops that lane only.** The other lanes run to completion; you report each lane's
64
+ terminal state separately. The stories were declared independent, so killing working lanes
65
+ because one failed throws away finished work for nothing.
66
+
67
+ **Every dispatched agent is told its working directory**, and that all paths resolve
68
+ against it. An executor or reviewer must never read or write another lane's worktree.
69
+
70
+ ### `memory.md` in multi-lane mode
71
+
72
+ `.spec/memory.md` is **shared at the `.spec/` root**, and step 6 below writes to it. Two
73
+ lanes appending to their own copy in their own worktree would conflict on *every* parallel
74
+ batch — a guaranteed merge failure that has nothing to do with the code.
75
+
76
+ So: **in multi-lane mode, lanes do not write `memory.md` at all.** Collect each lane's
77
+ distilled lessons as it goes, and write them **once, in the main checkout on the
78
+ integration branch, after the batch has merged** — pruning superseded entries there, as
79
+ step 6 describes. Single-lane runs are unaffected and write it in place as before.
80
+
81
+ The other artifacts are safe and need no special handling: `tasks.md` and `decisions.md`
82
+ are per-story, and `constitution.md` is read-only for the whole build.
83
+
40
84
  ## The per-task loop
41
85
 
42
86
  For the selected task, run these steps in order. **Do not tick a task until its
@@ -60,6 +104,9 @@ it short and concrete.
60
104
  ### 2. IMPLEMENT (dispatch `tiny-spec-build-executor`)
61
105
  Spawn one **`tiny-spec-build-executor`** with a fresh, self-contained prompt:
62
106
 
107
+ - **the working directory** it must operate in — this lane's worktree, or the project
108
+ root in a single-lane run. Every path resolves against it, and it must never read or
109
+ write another lane's worktree;
63
110
  - the task id, description, and **acceptance**;
64
111
  - the `files:` hint;
65
112
  - the **whole** `.spec/constitution.md`;
@@ -68,7 +115,16 @@ Spawn one **`tiny-spec-build-executor`** with a fresh, self-contained prompt:
68
115
  `.spec/<active>/SPEC.md` plus its `export:` path, so the executor can look at the
69
116
  design instead of guessing at it;
70
117
  - only the specific existing files the task starts from, named explicitly (so it
71
- edits with the real current contents, not blind).
118
+ edits with the real current contents, not blind);
119
+ - the **accumulated `CHANGES` paths from every passed task in this story so far**,
120
+ labelled as the ground this story has already moved. Every executor starts cold and
121
+ re-derives the codebase from scratch — that re-derivation is the **largest single cost
122
+ in the loop**, far larger than the prompt text around it — and naming the files this
123
+ story has already touched is the cheapest way to cut it. You already have the lists in
124
+ hand. **Paths only**, deduplicated, most-recently-touched first — never the earlier
125
+ tasks' descriptions, findings, or reports. (The reviewer does *not* get this list: it
126
+ is scoped to the task's own changed files on purpose, and widening it would dilute the
127
+ independence that makes its verdict worth anything.)
72
128
 
73
129
  Do **not** pass the plan, sibling tasks, or other chatter. It returns a structured
74
130
  report (`STATUS`, `CHANGES`, `DECISIONS`, `BLOCKER`). A `STATUS: blocked` →
@@ -77,14 +133,34 @@ report (`STATUS`, `CHANGES`, `DECISIONS`, `BLOCKER`). A `STATUS: blocked` →
77
133
  ### 3. REVIEW (dispatch `tiny-spec-build-reviewer` — independent)
78
134
  Spawn one **`tiny-spec-build-reviewer`**, **blind to step 2**, with:
79
135
 
136
+ - **the working directory** it must operate in — this lane's worktree, or the project
137
+ root in a single-lane run. It runs the gate there and touches no other lane;
80
138
  - the task id, description, and **acceptance**;
81
139
  - the **whole** `.spec/constitution.md`;
82
140
  - the list of changed files (from the executor's `CHANGES`) to read;
83
- - the **Verification commands** from the constitution to run;
141
+ - the **Verification commands** from the constitution to run, and the **gate scope**
142
+ for this task — `full` or `scoped` (decide it as below; you have the run context,
143
+ the reviewer does not);
84
144
  - **if the task has a `design:` field** — the same `D<n>` entry and `export:` path
85
145
  you gave the executor, so it grades against the contract rather than its taste.
86
146
 
87
- It runs the real gate end-to-end, checks the code against the constitution's
147
+ **Gate scope.** A clean `install → build` proves the same thing on task 7 that it proved
148
+ on task 1 unless a task changed what it installs or builds, so running it every time buys
149
+ little and costs the most wall-clock in the loop. Name the scope in the brief:
150
+
151
+ - **`full`** — the whole gate from a clean state (install → lint → test → build → run).
152
+ Use it on: the **first** task of this run; any task whose `CHANGES` touch dependency or
153
+ build configuration (`package.json`/lockfiles, `pyproject.toml`, `go.mod`, `Cargo.toml`,
154
+ `Dockerfile`, CI config, build scripts); any task carrying `design:`; and the **last**
155
+ task in `tasks.md`.
156
+ - **`scoped`** — otherwise. `lint:` + `test:` plus the acceptance exercised end-to-end,
157
+ black-box, exactly as on a full gate. The acceptance exercise is never scoped away; it
158
+ is the whole reason the reviewer exists.
159
+
160
+ The reviewer may **escalate** `scoped` to `full` on its own and will say so — that
161
+ direction is always safe. It may never go the other way.
162
+
163
+ It runs the gate at the scope you named, checks the code against the constitution's
88
164
  **Definition of Done** and **invariants**, confirms the **acceptance** actually
89
165
  holds (exercised, not inferred), and returns `VERDICT: PASS | FAIL` + findings.
90
166
  On a `design:` task it also runs the constitution's `visual:` command, measures the
@@ -94,11 +170,15 @@ numbers can't reach, like an element that measures perfectly and shows up invisi
94
170
  It fails on what it measured or saw, and merely flags what is taste.
95
171
 
96
172
  > Why independent: unit-green ≠ working, and the author is the worst judge of its
97
- > own blind spots. The reviewer running the gate from a clean state is the
98
- > safeguard that keeps scope and quality honest without an ownership contract.
173
+ > own blind spots. The reviewer exercising the acceptance black-box — on every task,
174
+ > at either scope — is the safeguard that keeps scope and quality honest without an
175
+ > ownership contract.
99
176
 
100
177
  ### 4. CONVERGE (on FAIL)
101
- Re-dispatch the **executor** with the reviewer's findings appended to its brief.
178
+ Re-dispatch the **executor** with the reviewer's findings appended to its brief, then
179
+ review again — recomputing the **gate scope** from the fix's `CHANGES`, and using `full`
180
+ if the failure was in the gate itself or the reviewer escalated last time. A task that
181
+ already failed once has earned the wider check.
102
182
  Bound this to **2 fix attempts**. If it still fails after that, stop and **halt
103
183
  `exhausted`** (see **Halting** below) — don't keep grinding or hand-fix past the loop
104
184
  silently. `exhausted` is recorded as a blocker, because a task that three attempts
@@ -143,6 +223,9 @@ entry to the **shared** `.spec/memory.md` (the root — lessons are project-wide
143
223
  pruning any entry the new one supersedes. Skip code-style rules (→ shared
144
224
  `constitution.md`) and one-off history (→ the ticket's `decisions.md`). Keep it lean.
145
225
 
226
+ **In multi-lane mode, do not write the file here.** Hold the entry and write it once
227
+ after the batch merges — see **Multi-lane mode** above for why.
228
+
146
229
  On first use, create the file with this structure:
147
230
 
148
231
  ```markdown
@@ -186,7 +269,8 @@ way, which is why a loop and a hand-driven build can't drift apart.
186
269
  ## Halting
187
270
 
188
271
  Every build run ends in exactly **one** of five terminal states. Name it out loud in
189
- the closing report:
272
+ the closing report — **per lane**, when there is more than one; a lane's halt never
273
+ rounds up into another lane's `done`, and the run's own state is the worst of them:
190
274
 
191
275
  | State | Cause |
192
276
  |---|---|
@@ -274,6 +358,12 @@ When every task in `tasks.md` is `[x]`:
274
358
  setup — install/build, not a test-runner shortcut). There is no separate
275
359
  verify skill — this final smoke confirms the requirements actually work end-to-end, not
276
360
  just that tasks are ticked.
361
+
362
+ **This is the authoritative clean-state run for the whole build**, and it is never
363
+ scoped or skipped. Per-task gates run at the scope step 3 named; this one always runs
364
+ everything from clean. A build or install regression that a `scoped` task gate did not
365
+ re-prove surfaces here — which is why a red final smoke is a gap to fix or route
366
+ upstream, never a pass and never `done`.
277
367
  2. **Report** — the terminal state by name (`done` here, since every task is `[x]`
278
368
  and the smoke passed), what was built, the commits made (with the branch), and any
279
369
  open `decisions.md` items (blockers, tasks unchecked by a reconcile). If the final
@@ -18,19 +18,15 @@ sets `tasks.md`'s `status:`.
18
18
 
19
19
  Artifacts live under `.spec/`: the **shared** constitution at the root
20
20
  (`.spec/constitution.md`), the per-ticket `SPEC.md`/`PLAN.md`/`tasks.md` under
21
- `.spec/<slug>/`. **Resolve the active ticket dir** from the current git branch: the
22
- `.spec/<slug>/` whose slug matches the branch name (one branch per ticket). If none
23
- matches, use the sole ticket dir if there's exactly one; else ask which. Both skeletons
24
- are inline below — write them from there, no file to read. Requires
25
- `.spec/<active>/SPEC.md`.
26
-
27
- **Two cases pre-empt that order — ask instead of applying it:** more than one dir
28
- matches the branch (there is no defined tie-break, and inventing one here would
29
- silently disagree with every other skill), or ticket dirs exist while you are on
30
- `main`/`master` with no name match (the usual cause is a forgotten `git switch`, and
31
- the sole-dir fallback would otherwise swallow it). Detached HEAD or no git repo is a
32
- **degraded** case, not an ask case — branch match is simply unavailable, so fall
33
- through to sole-dir and ask as written.
21
+ `.spec/<slug>/`. Both skeletons are inline below — write them from there, no file to
22
+ read. Requires `.spec/<active>/SPEC.md`.
23
+
24
+ **Resolve the active ticket dir**, in order: the `.spec/<slug>/` whose slug matches the
25
+ current git branch (one branch per ticket); else the sole ticket dir if exactly one
26
+ exists; else ask. **Ask instead** when more than one dir matches the branch, or when
27
+ ticket dirs exist while you are on `main`/`master` with no name match — neither has a
28
+ safe tie-break. Detached HEAD or no git repo is **degraded**, not an ask: branch match is
29
+ simply unavailable, so fall through to sole-dir and ask as written.
34
30
 
35
31
  ## Step 1 — harden the constitution (`constitution.md`)
36
32
 
@@ -138,25 +134,68 @@ past it into a task list — say what's unresolved and route back to `tiny-spec-
138
134
  ## Step 3 — slice the approach into tasks
139
135
 
140
136
  Walk the `## Approach` you just wrote and break it into tasks. No waves, no
141
- parallelism, no `owns:` contracts — tasks run one at a time, top to bottom. Each task is:
142
-
143
- - **Small and independently checkable** — one slice a single executor can finish and a
144
- reviewer can grade in one pass. If you can't write a one-line acceptance for it, it's
145
- too big — split it.
137
+ parallelism, no `owns:` contracts — tasks run one at a time, top to bottom.
138
+
139
+ **The unit is one coherent commit** — what a competent engineer does in one focused
140
+ sitting and commits as a single unit of work. That typically touches several files and
141
+ often satisfies several `REQ-N` at once. It is the size you would open as one reviewable
142
+ change, not the smallest thing you could name.
143
+
144
+ **Know what a task costs, so size has something to trade against.** Every task you write
145
+ spends two cold-start agents that must re-orient in the codebase from scratch, a reviewer
146
+ that runs the gate and exercises the acceptance end-to-end, and two commits. That
147
+ overhead is **fixed** — it does not shrink for a small task. A task the executor
148
+ finishes in thirty seconds still pays all of it. Size each task so the work inside it
149
+ clearly outweighs the machinery around it.
150
+
151
+ So:
152
+
153
+ - **Split on independent failure, not on sentence length.** Split a task when it carries
154
+ two **unrelated** observable outcomes that could fail independently of each other. Do
155
+ **not** split because the acceptance got long — a good acceptance is usually several
156
+ clauses covering the happy path *and* its negatives. See `T4` in the `examples/todo-cli`
157
+ task list: two CLI commands, four requirements, and three negative cases, in one task,
158
+ with one acceptance. That is the calibration point, not the exception.
159
+ - **One task may satisfy several `REQ-N`** — `req:` takes a list. Coverage means every
160
+ requirement has a *home*, not that every requirement gets its *own* task. **A 1:1
161
+ REQ→task mapping is the single most common way this list comes out too granular.**
162
+ Group the requirements that one coherent change delivers together.
146
163
  - **Ordered so each builds on the last.** Tasks run sequentially, so a later task may
147
164
  freely assume an earlier task's code already exists. Put foundational work (types,
148
165
  schema, scaffolding) first. Order by dependency, not by guesswork.
149
- - **Right-sized, not fragmented.** Don't split a cohesive change into five files' worth
150
- of micro-tasks just to look granular. Earned ceremony: fewer, meaningful tasks beat
151
- many trivial ones.
166
+
167
+ **Smells that mean you sliced below the commit line** — fold each of these back into the
168
+ task it belongs to:
169
+
170
+ - a task that only defines types, interfaces, or schema with no behavior behind them;
171
+ - a task that only adds tests for the task before it (the constitution's **Definition of
172
+ Done** already requires the tests to ship with the code);
173
+ - one task per file, or one task per function;
174
+ - a "wire it up" / "integrate the pieces" task trailing the pieces it wires;
175
+ - a **leading pure-scaffold task** — a skeleton, a dispatch stub, a module that imports
176
+ cleanly and does nothing. Fold it into the first task that gives it behavior;
177
+ - a **trailing end-to-end verification task**. `tiny-spec-build`'s Completion step already
178
+ runs the whole gate against the whole project from a clean state, exercised the way a
179
+ user would. A task that re-does it buys nothing and costs the full per-task overhead.
180
+
181
+ The shipped `examples/todo-cli` task list predates these two smells and shows both: its
182
+ `T1` is a pure scaffold and its `T5` is an end-to-end verification pass. Today `T1` folds
183
+ into `T2` and `T5` doesn't exist. Read that list for `T4`'s sizing, not for its edges.
184
+
185
+ **Count is a smell, not a cap.** A story sized the way `tiny-spec-scope` describes
186
+ usually lands in **2–4 tasks**. If you are past about six, re-read the list: you have
187
+ either sliced below the commit line, or the story itself was too big and should have been
188
+ split upstream. Check the list against that; do **not** enforce a number, and never drop
189
+ or merge coverage just to hit one.
152
190
 
153
191
  For each task, write:
154
192
 
155
193
  ```
156
194
  - [ ] T<n> — <imperative description>
157
- - acceptance: <one user-observable outcome that proves it's done>
195
+ - acceptance: <one user-observable outcome that proves it's done — happy path and the
196
+ negatives that bound it, in one entry>
158
197
  - type: feat # optional; Conventional Commit type (defaults to feat)
159
- - req: REQ-n # optional; the REQ-N this task delivers
198
+ - req: REQ-n, REQ-n # optional; the REQ-N this task delivers — a list, not one
160
199
  - design: D-n # optional; the SPEC.md D<n> screen this task builds — arms the visual gate
161
200
  - pause: <why> # optional; halt the build before this task so a human looks first
162
201
  - files: <comma-separated hint of files it will touch>
@@ -164,7 +203,11 @@ For each task, write:
164
203
 
165
204
  The **acceptance** is what the reviewer checks against — make it observable
166
205
  ("`spec --version` prints the version and exits 0"), not internal ("version logic
167
- added"). **type** picks the Conventional Commit type `tiny-spec-build` uses for this
206
+ added"). It states **one outcome**, but one outcome is not one clause: spell out the
207
+ happy path and the negatives that bound it in the same acceptance, separated by
208
+ semicolons. A long acceptance is a well-specified task, not an oversized one — it is the
209
+ *number of unrelated things that could fail* that decides whether to split, not the
210
+ length of the line. **type** picks the Conventional Commit type `tiny-spec-build` uses for this
168
211
  task's code commit (`feat | fix | docs | refactor | test | chore | build | ci | perf | style`);
169
212
  set it when the task is clearly not a feature, otherwise omit and it defaults to `feat`.
170
213
  **req** ties the task to the requirement it satisfies (traceability). The **files** line
@@ -211,7 +254,9 @@ anyway. A policy that matches nothing here is not an error — say so and move o
211
254
  than stretching a task to fit it.
212
255
 
213
256
  Cover **every** part of the approach — together the tasks must deliver all `REQ-N`.
214
- Don't leave a requirement with no task. Likewise, if `SPEC.md` has a `## Design`
257
+ Don't leave a requirement with no task. Coverage is about requirements having a home,
258
+ not about the shape of the mapping: several `REQ-N` on one task is the normal case, and
259
+ a task per requirement is the anti-pattern. Likewise, if `SPEC.md` has a `## Design`
215
260
  section, every `D<n>` in it needs at least one task carrying that `design:` reference —
216
261
  a screen nobody is graded against is a screen that will be built wrong.
217
262
 
@@ -236,25 +281,31 @@ updated: <ISO date>
236
281
 
237
282
  ## Tasks
238
283
 
239
- - [ ] T1 — <one small, independently-checkable slice of work>
240
- - acceptance: <one user-observable outcome that proves T1 is done>
284
+ - [ ] T1 — <one coherent commit's worth of foundational work; usually several files>
285
+ - acceptance: <the observable outcome; happy path; and the negative case that bounds it>
241
286
  - type: feat # optional; Conventional Commit type for this task's commit (defaults to feat)
242
- - req: REQ-1 # optional; the REQ-N this task delivers
243
- - files: <path, path>
287
+ - req: REQ-1, REQ-2 # optional; the REQ-N this task delivers — several is normal
288
+ - files: <path, path, path>
244
289
 
245
- - [ ] T2 — <next slice; assume T1's code exists>
246
- - acceptance: <observable outcome>
290
+ - [ ] T2 — <next coherent change; assume T1's code exists>
291
+ - acceptance: <observable outcome; plus what it rejects and how it fails>
247
292
  - type: feat
248
- - req: REQ-2
293
+ - req: REQ-3, REQ-4, REQ-5
249
294
  - design: D1 # optional; only on tasks that build the visible surface
250
295
  - files: <path, path>
251
296
 
252
297
  - [ ] T3 — <…>
253
298
  - acceptance: <observable outcome>
299
+ - req: REQ-6
254
300
  - pause: <optional; what to check before this runs — irreversible work only>
255
301
  - files: <path, path>
256
302
  ```
257
303
 
304
+ The `req:` lists above are the shape to aim for, not filler: a handful of tasks each
305
+ carrying the requirements one coherent change delivers. A skeleton filled in as
306
+ `REQ-1`, `REQ-2`, `REQ-3` down a column of single-requirement tasks is the granularity
307
+ failure described in Step 3.
308
+
258
309
  ## Update mode (SPEC changed → PLAN and tasks are stale)
259
310
 
260
311
  When `PLAN.md` or `tasks.md` is `status: stale`, reconcile **both in one pass** — they
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: tiny-spec-run
3
- description: The one router — read each artifact's status flag to work out where work stands, then invoke the skill that moves it forward. Two stop points, chosen once at the start: by default it walks tiny-spec-adopt/create → design → plan and STOPS before tiny-spec-build; asked to build ("spec it out and build it") or handed a story list, it drives each story branch → plan → build → merge until the list ends or something halts it. Use for "run the spec flow", "pick the chain back up", "build the backlog", "work through the breakdown". NOT for a single stage — for that invoke tiny-spec-create, tiny-spec-plan, or tiny-spec-design directly.
3
+ description: The one router — read each artifact's status flag to work out where work stands, then invoke the skill that moves it forward. Two stop points, chosen once at the start: by default it walks tiny-spec-adopt/create → design → plan and STOPS before tiny-spec-build; asked to build ("spec it out and build it") or handed a story list, it drives each story branch → plan → build → merge until the list ends or something halts it. Independent stories — grouped by their BREAKDOWN.md needs: field, or named by the user — build concurrently, one git worktree per lane, merging batch by batch. Use for "run the spec flow", "pick the chain back up", "build the backlog", "work through the breakdown". NOT for a single stage — for that invoke tiny-spec-create, tiny-spec-plan, or tiny-spec-design directly.
4
4
  ---
5
5
 
6
6
  # tiny-spec-run
@@ -53,8 +53,8 @@ return to Step 2 after a stage, or between stories.
53
53
  the branch for the user.
54
54
 
55
55
  2. **Resolve the story list** (below) and play it back in order — slug and title — with
56
- the state you derived for each. This is the user's chance to reorder, drop, or narrow
57
- before anything is cut.
56
+ the state you derived for each, **grouped into the batches you will run**. This is the
57
+ user's chance to reorder, drop, narrow, or correct a dependency before anything is cut.
58
58
 
59
59
  3. **Take the pause policy.** Ask for, or accept, standing technical stop points for
60
60
  this run — *"halt before anything that touches auth"*, *"stop before any schema
@@ -63,6 +63,17 @@ return to Step 2 after a stage, or between stories.
63
63
  **technical**, not per-story: the thing worth looking at is a migration or an auth
64
64
  boundary, not a feature heading.
65
65
 
66
+ 4. **Fix the lane count.** Default **3**; take a number from the user if they gave one.
67
+ This is a concurrency limit on the machine, not a budget — the story list is still the
68
+ budget, and a batch wider than the lane count simply runs in successive fills. **One
69
+ lane must stay available**: any phrasing asking for one story at a time ("one at a
70
+ time", "no parallelism", "serially") sets it to 1, which collapses Step 4 to the
71
+ sequential behavior — no worktrees, no batching.
72
+
73
+ Then check one more precondition, once you know the slugs in the first batch: **no
74
+ target worktree path already exists.** Name the collisions and stop; do not reuse or
75
+ remove a directory you found.
76
+
66
77
  ## Step 1 — resolve the scope
67
78
 
68
79
  ### Stop-before-build: the active ticket dir
@@ -126,7 +137,26 @@ on the integration branch at all, so "does the directory exist" can't tell "not
126
137
  apart from "built on a branch you haven't merged" — reading each ref explicitly can.
127
138
  Sourcing step 1 from the integration branch is also what survives a deleted branch: once
128
139
  a story is merged, its ticked task list is part of `main` whether or not the branch that
129
- built it still exists.
140
+ built it still exists. **This survives parallelism unchanged** — a worktree is just a
141
+ branch, so a story built in a worktree but not yet merged lands on rule 2 exactly as one
142
+ built in place.
143
+
144
+ **Group the remaining stories into batches.** Level 0 is every unbuilt story whose
145
+ `needs:` are empty or already built-and-merged; level 1 is every story whose `needs:` are
146
+ all in level 0; and so on. Batches run **in order**; the stories inside one batch run
147
+ **concurrently**, capped at the lane count.
148
+
149
+ - **A parallel set the user named at invocation overrides the graph** for this run —
150
+ "build these three at once", or a pasted list they called independent. Explicit beats
151
+ derived, the same way a pasted story list already wins over `BREAKDOWN.md`. Say that you
152
+ are overriding, and name what the graph would have done.
153
+ - **A cycle in `needs:` is a stop.** Print the cycle and stop. Never break it by picking
154
+ an order — `tiny-spec-scope` says two stories that need each other are really one, and
155
+ choosing for the user hides a carve that needs fixing.
156
+ - **A `needs:` naming a slug that isn't in the list is a stop**, not a shrug. It is either
157
+ a typo or a story someone dropped, and both want a human.
158
+ - **A batch of one is the sequential behavior** — Step 4 collapses to exactly what it did
159
+ before lanes existed. This is the common case for a single ticket and must not change.
130
160
 
131
161
  ## Step 2 — the ladder
132
162
 
@@ -146,67 +176,62 @@ re-evaluate against the new state on disk.
146
176
  | — | **no rung matched** | **stop** — report the exact state you found and ask; never improvise a stage |
147
177
 
148
178
  **Upstream beats downstream — that is what the table order encodes.** Always fix the
149
- earliest artifact in the chain that needs attention, whether it's stale *or* missing.
150
- Deriving a task list from a design you already know is wrong wastes the run, and then
151
- the reconcile has to uncheck completed work all over again.
179
+ earliest artifact in the chain that needs attention, stale *or* missing. Deriving a task
180
+ list from a design you already know is wrong wastes the run, and the reconcile then has
181
+ to uncheck completed work all over again.
152
182
 
153
183
  **L0/L1 carry the greenfield/brownfield fork.** "Does this repo already contain source?"
154
- means: is there real code here beyond config and docs? If yes, the constitution should
155
- be *derived from that code* (`tiny-spec-adopt`) rather than interviewed out of the user
156
- — an interviewed constitution on an existing codebase is how you get a gate that names
157
- commands the project doesn't have. If the user is starting from an idea with no code
158
- yet, `tiny-spec-scope` is theirs to run first; it is not in the ladder (see below).
159
-
160
- **L1 is a reseed, not a fresh start.** The constitution is project-wide, so it can go
161
- missing while a perfectly good `SPEC.md` sits next to it. Invoke the chosen skill in
162
- **reseed/refresh mode** — say so explicitly, and add: *do not re-interview from scratch,
163
- do not create a ticket dir, and do not touch `SPEC.md`.* If the constitution still
164
- doesn't exist afterwards, stop and tell the user; do not loop.
165
-
166
- **L1 outranks L2** — a constitution gets reseeded even when the ticket dir or `SPEC.md`
167
- is also missing. The next pass then lands on L2 and creates the spec properly — and
168
- because that's a *different rung*, the same-skill bound below does not fire.
169
-
170
- **L3 is a stop, not a stage.** Nothing in the suite ever *sets* `SPEC.md` to stale — it
171
- is the root of the chain. A stale SPEC means someone hand-edited it, so surface it and
172
- offer `tiny-spec-create` in update mode rather than assuming intent. If the user says go
173
- ahead **in the same turn**, invoke it and carry on down the ladder — the stop is there
174
- to get a human decision, not to force a second command. (Update mode *clears* the flag
175
- on its way out, so this rung resolves rather than repeating.)
176
-
177
- **L4 is the one thing `run` checks that isn't a `status:` flag** — and it is
178
- deliberately *not* a validator. For each `D<n>` in the active `SPEC.md`, run
179
- `shasum -a 256 <export>` and compare with the `sha256:` the entry declares. You are
180
- checking whether an anchor still points at what it says it points at, not inferring
181
- whether two documents agree. A design that moved under a finished spec is otherwise
182
- completely invisible — no status flips, and the build reviews against a screen that no
183
- longer exists.
184
+ means: real code beyond config and docs. If yes, the constitution must be *derived from
185
+ that code* (`tiny-spec-adopt`) rather than interviewed out of the user — an interviewed
186
+ constitution on an existing codebase is how you get a gate that names commands the
187
+ project doesn't have. Starting from an idea with no code yet, `tiny-spec-scope` is the
188
+ user's to run first; it is not in the ladder (see below).
189
+
190
+ **L1 is a reseed, not a fresh start**, and it **outranks L2**. The constitution is
191
+ project-wide, so it can go missing while a perfectly good `SPEC.md` sits next to it.
192
+ Invoke the chosen skill in **reseed/refresh mode** — say so explicitly, and add: *do not
193
+ re-interview from scratch, do not create a ticket dir, and do not touch `SPEC.md`.* If the
194
+ constitution still doesn't exist afterwards, stop and tell the user; do not loop. The next
195
+ pass lands on L2 and creates the spec properly — a *different rung*, so the same-skill
196
+ bound below does not fire.
197
+
198
+ **L3 is a stop, not a stage.** Nothing in the suite ever *sets* `SPEC.md` to stale — it is
199
+ the root of the chain — so a stale SPEC means someone hand-edited it. Surface it and offer
200
+ `tiny-spec-create` in update mode rather than assuming intent. A go-ahead **in the same
201
+ turn** is enough to invoke it and carry on down the ladder; the stop exists to get a human
202
+ decision, not to force a second command. (Update mode *clears* the flag on its way out, so
203
+ this rung resolves rather than repeating.)
204
+
205
+ **L4 is the one thing `run` checks that isn't a `status:` flag**, and it is deliberately
206
+ *not* a validator: for each `D<n>` in the active `SPEC.md`, run `shasum -a 256 <export>`
207
+ and compare with the `sha256:` the entry declares. You are checking whether an anchor
208
+ still points at what it says it points at, not inferring whether two documents agree. A
209
+ design that moved under a finished spec is otherwise completely invisible — no status
210
+ flips, and the build reviews against a screen that no longer exists.
184
211
 
185
212
  - **Mismatch** → `tiny-spec-design` in re-anchor mode.
186
213
  - **Missing file** → **stop and tell the user**, naming the entry and the path. Don't
187
- route it: a deleted export can mean a rename, a move, or a design that was withdrawn,
188
- and each wants a different answer. Never quietly drop the entry.
214
+ route it: a deleted export can mean a rename, a move, or a withdrawn design, and each
215
+ wants a different answer. Never quietly drop the entry.
189
216
  - No `## Design` section, or every hash matches → the rung doesn't fire; fall through.
190
217
 
191
218
  **Any `status:` that isn't exactly `current`** — `stale`, missing, unreadable, or an
192
219
  unrecognized value like `draft` — counts as **stale**, on `SPEC.md`, `PLAN.md`, and
193
- `tasks.md` alike. Say so out loud. For `PLAN.md`/`tasks.md` that means reconciling
194
- (update mode preserves existing ids, so it's the non-destructive way to be wrong); for
195
- `SPEC.md` it means L3 — stop and ask.
220
+ `tasks.md` alike. Say so out loud. `PLAN.md`/`tasks.md` → reconcile (update mode preserves
221
+ existing ids, the non-destructive way to be wrong); `SPEC.md` → L3, stop and ask.
196
222
 
197
- **A `tasks.md` with no tasks at all is not "built"** — that's why L6 and L7 both require
198
- at least one task, and why an empty checklist matches neither. It means the
199
- `tiny-spec-plan` run produced nothing, so L5 catches it. If it comes back empty a second
200
- time, stop and tell the user — the plan has nothing derivable in it.
223
+ **A `tasks.md` with no tasks at all is not "built"** — L6 and L7 both require at least one
224
+ task, so an empty checklist matches neither and L5 catches it. Empty a second time → stop
225
+ and tell the user; the plan has nothing derivable in it.
201
226
 
202
227
  `run` trusts the `status:` flags. It does **not** second-guess hand edits, diff
203
- timestamps, or validate the chain — there is no validator in this suite by design. (L4
204
- is not an exception: a hash is a value the spec itself declares about a file it names,
205
- so checking it is reading state, not judging consistency. Do not use it as a precedent
206
- for adding cross-document checks.) That means it inherits each stage's propagation: if
207
- `tiny-spec-plan` update mode doesn't flip `tasks.md`, `run` will walk right past it.
208
- That trust is the price of having no validator; when a run's result looks wrong, suspect
209
- the stage's propagation before the ladder.
228
+ timestamps, or validate the chain — there is no validator in this suite by design. (L4 is
229
+ not an exception: a hash is a value the spec itself declares about a file it names, so
230
+ checking it is reading state, not judging consistency. Not a precedent for adding
231
+ cross-document checks.) So it inherits each stage's propagation: if `tiny-spec-plan`
232
+ update mode doesn't flip `tasks.md`, `run` walks right past it. That trust is the price of
233
+ having no validator; when a run's result looks wrong, suspect the stage's propagation
234
+ before the ladder.
210
235
 
211
236
  **`tiny-spec-scope` is not in the ladder.** It is a pre-spec on-ramp that writes
212
237
  `BREAKDOWN.md` at the project root, carries no `status:` frontmatter, and is the user's
@@ -253,39 +278,60 @@ to enter `tiny-spec-build` in a stop-before-build run.**
253
278
  ## Step 4 — build-through only: build and merge
254
279
 
255
280
  Reached only when Step 0 fixed the stop point at build-through and the ladder is at L6
256
- (or L7 with an unmerged branch). For the first story that isn't already built and merged:
257
-
258
- 1. **Branch.** `git switch <slug>` if it already exists; otherwise
259
- `git switch -c <slug> <integration>` — cut **fresh from the integration branch** so
260
- this story sees every story merged before it. That is what makes an ordered list
261
- build correctly: story 3 gets stories 1 and 2 already in its tree.
262
-
263
- 2. **Ladder.** Walk Steps 1–3 for this story until it reaches L6. If it stops anywhere
264
- else, that is a halt — report and stop the whole run. **Never walk the ladder twice
265
- to push past its own stop:** each of those is a human decision it deliberately
266
- declined to make, and running it again declines again.
267
-
268
- 3. **Build.** Invoke **`tiny-spec-build`**, once, briefed to run it through. It owns the
269
- per-task loop and writes its own halt record. Anything other than `done` halts the
270
- whole run — **do not merge a story that didn't finish**, and do not invoke build a
271
- second time: it resumes from the checkbox state, so it lands on the very task that
272
- just halted and halts there again.
273
-
274
- 4. **Merge — only on `done`.** In order:
281
+ (or L7 with an unmerged branch). Work **one batch at a time**, in batch order. For every
282
+ story in the current batch that isn't already built and merged:
283
+
284
+ 1. **Lane.** Give each story its own **git worktree**, so lanes cannot collide on the
285
+ filesystem: `git worktree add ../<repo>-<slug> -b <slug> <integration>`, or
286
+ `git worktree add ../<repo>-<slug> <slug>` when the branch already exists. Cut **fresh
287
+ from the integration branch** so the story sees every story merged before it — that is
288
+ what makes an ordered backlog build correctly, and it is why batches merge before the
289
+ next one starts.
290
+
291
+ With a single lane you may stay in the main checkout and `git switch` as before; a
292
+ worktree is only required when a batch has more than one story.
293
+
294
+ 2. **Ladder.** Walk Steps 1–3 for each story, in its own worktree, until it reaches L6.
295
+ A story that stops anywhere else does **not** enter the build — its lane is halted;
296
+ see the halt rule below. **Never walk the ladder twice to push past its own stop:**
297
+ each of those is a human decision it deliberately declined to make, and running it
298
+ again declines again.
299
+
300
+ 3. **Build.** Invoke **`tiny-spec-build`** **once for the whole batch**, briefed to run
301
+ through, and hand it every live lane as a `(worktree path, slug)` pair. It owns the
302
+ per-task loop, runs the lanes concurrently, and writes each story's halt record.
303
+ Do not invoke it a second time for a lane that halted: it resumes from the checkbox
304
+ state, so it lands on the very task that just halted and halts there again.
305
+
306
+ 4. **Merge — only the lanes that returned `done`, one at a time, in batch order.**
307
+ For each, in the main checkout:
275
308
  - `git switch <integration>`
276
309
  - `git merge --no-ff <slug>` — the merge commit keeps each story legible in history.
277
- - **Conflict** → `git merge --abort`, then halt `conflict`.
278
- - **Run the constitution's Verification commands on the merged result**, exercised
279
- the way a user would. A story that was green alone can still break against work
280
- merged before it, and that is exactly what this catches. Red → halt `blocked`.
310
+ - **Conflict** → `git merge --abort`, halt that story `conflict`, and carry on with
311
+ the rest of the batch. Report every conflicted story at the end.
312
+ - **Run the constitution's Verification commands on the merged result**, exercised the
313
+ way a user would. Red → halt `blocked`. **This check is load-bearing under
314
+ parallelism and is never skipped:** every lane gated against a tree that did not
315
+ contain its siblings, so the merged result is the first time they meet.
281
316
  - **Never push.** Merging locally keeps a bad run one `git reset` away; sending it to
282
317
  a remote is the user's call, and this suite makes no network calls.
283
318
 
284
- 5. **Next story.** Return to Step 1. Do not re-run Step 0.
319
+ 5. **Next batch.** If **every** lane in this batch merged cleanly, return to Step 1 for
320
+ the next batch. Do not re-run Step 0. **If any lane halted, stop the run here** — a
321
+ later batch may well depend on the story that failed, and working out which stories
322
+ are still safe is exactly the machinery this suite refuses to grow. Report and let the
323
+ user decide.
324
+
325
+ 6. **Leave the worktrees.** Do **not** run `git worktree remove`. Cleanup is the user's,
326
+ for the same reason branch deletion is — a halted lane's worktree holds the tree they
327
+ need to look at. List every path you created in the closing report, with the
328
+ `git worktree remove` commands, and let them run it.
285
329
 
286
- **A halt stops the whole run**, not just the current story. Later stories in a list
287
- written top to bottom usually assume the earlier ones landed, so skipping ahead past a
288
- failure produces a second, more confusing failure downstream.
330
+ **A halt stops the lane it happened in, and ends the run at the end of that batch.**
331
+ Within a batch the other lanes run to completion: the stories were declared independent,
332
+ so stopping the ones that are working buys nothing and throws away finished work. Across
333
+ batches the old rule stands unchanged — later stories usually assume earlier ones landed,
334
+ so never skip ahead past a failure.
289
335
 
290
336
  `tiny-spec-build` records its own halts (`blocked`, `exhausted`, `fork`) in the story's
291
337
  `decisions.md`. **You record nothing** — the merge-stage halts are already legible
@@ -309,11 +355,15 @@ L3 or on a bound, say exactly what stopped you and what the user needs to decide
309
355
  **Build-through.** Name, in this order:
310
356
 
311
357
  1. **The terminal state** — exactly one of `done`, `blocked`, `exhausted`, `paused`,
312
- `fork`, `conflict`. Use the word.
313
- 2. **Stories built and merged**, in order, with their merge commits.
314
- 3. **The story it stopped on**, the task within it, and why in one line.
358
+ `fork`, `conflict`. Use the word. With more than one lane, report **each story's own
359
+ state**, and give the run's state as the **worst** of them: four green lanes and one
360
+ `blocked` is a `blocked` run, never a `done` one with a footnote.
361
+ 2. **Stories built and merged**, in batch order, with their merge commits.
362
+ 3. **The story or stories it stopped on**, the task within each, and why in one line.
315
363
  4. **Stories never started** — say how many are left, by name.
316
- 5. **The one command that resolves it** — `tiny-spec-create`/`tiny-spec-plan` in update
364
+ 5. **Worktrees left on disk**, with the `git worktree remove` command for each. You do
365
+ not run them.
366
+ 6. **The one command that resolves it** — `tiny-spec-create`/`tiny-spec-plan` in update
317
367
  mode for `blocked`/`exhausted`, this skill again for `paused`, the decision the user
318
368
  owes you for a `fork`, or the conflicted paths for a `conflict`.
319
369
 
@@ -338,9 +388,11 @@ ticket that legitimately has no visual surface.
338
388
  - **Never write, edit, or flip anything.** No `status:`, no checkbox, no `decisions.md`
339
389
  entry, no code, no `BREAKDOWN.md` edit. Delegate or stop. A router that regenerates
340
390
  documents is a second source of truth.
341
- - **Never push, force, rebase, reset, or delete a branch.** The only git commands this
342
- skill issues are `switch`, `switch -c`, `merge --no-ff`, `merge --abort`, and
343
- read-only queries. Everything outward-facing or destructive is the user's.
391
+ - **Never push, force, rebase, reset, delete a branch, or remove a worktree.** The only
392
+ git commands this skill issues are `switch`, `switch -c`, `merge --no-ff`,
393
+ `merge --abort`, `worktree add`, `worktree list`, and read-only queries. Everything
394
+ outward-facing or destructive stays the user's — `worktree remove` included, which is
395
+ why Step 4 hands those commands over instead of running them.
344
396
  - **Never merge a story whose build didn't return `done`.**
345
397
  - **Never invoke `tiny-spec-run`.** Re-entering means re-reading these steps, not
346
398
  calling yourself. Self-invocation compounds context and does not terminate.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: tiny-spec-scope
3
- description: The greenfield on-ramp — turn a rough idea, or an existing PRD, into BREAKDOWN.md at the project root - a flat list of Features → user Stories, each with draft acceptance criteria and a slug, plus a shared Decisions block. The one skill that works from a blank page. Does not scaffold .spec/ or touch the constitution — tiny-spec-create reads BREAKDOWN.md in seeded mode and does that. For an existing codebase use tiny-spec-adopt instead. The suite works without either.
3
+ description: The greenfield on-ramp — turn a rough idea, or an existing PRD, into BREAKDOWN.md at the project root - a flat list of Features → user Stories, each with draft acceptance criteria, a slug, and an optional needs: naming the stories it must follow, plus a shared Decisions block. The one skill that works from a blank page. Does not scaffold .spec/ or touch the constitution — tiny-spec-create reads BREAKDOWN.md in seeded mode and does that. For an existing codebase use tiny-spec-adopt instead. The suite works without either.
4
4
  ---
5
5
 
6
6
  # tiny-spec-scope
@@ -135,6 +135,7 @@ Write `BREAKDOWN.md` at the **project root** (the user's cwd) with the structure
135
135
  - Story: <one user-observable capability> slug: <ado-__ | kebab>
136
136
  - AC: <a single user-observable, testable outcome — becomes REQ-1 in this story's SPEC.md>
137
137
  - AC: <another — keep each atomic; no "and" hiding two capabilities>
138
+ - needs: <slug, slug> <!-- optional; stories this one must be built after. Omit when independent. -->
138
139
  - design: <path(s) to the wireframe(s) covering this story — omit if it has no visual surface>
139
140
  - Story: <one user-observable capability> slug: <ado-__ | kebab>
140
141
  - AC: <…>
@@ -171,6 +172,24 @@ story it belongs to, so `tiny-spec-design` knows which files to open for that sp
171
172
  instead of re-deriving the mapping from filenames. A wireframe covering several stories
172
173
  is listed on each of them.
173
174
 
175
+ **`needs:` is what lets stories build in parallel — so keep it honest and keep it rare.**
176
+ It names the stories this one must be built *after*. `tiny-spec-run` uses it to group the
177
+ backlog into batches: stories with no unmet `needs:` build concurrently, each in its own
178
+ git worktree, and the next batch starts once they merge. Three rules:
179
+
180
+ - **Name a story only for a real code dependency** — this story reads or edits files that
181
+ one writes, or builds on a schema, API, or type it introduces. A shared theme, a shared
182
+ `## Feature:` heading, or "it feels like it comes second" is **not** a dependency.
183
+ - **When unsure, omit it.** The two errors are not symmetric. A `needs:` you didn't need
184
+ costs parallelism on every future run of this backlog, silently and forever. A `needs:`
185
+ you missed costs one merge conflict, which `tiny-spec-run` already detects and halts on.
186
+ - **Never point across a cycle.** If two stories each need the other, they are one story —
187
+ merge them, or split the shared part out into a third that both need.
188
+
189
+ A `## Feature:` heading carries no `needs:` — it is a grouping, not a work item. Most
190
+ backlogs need the field on a minority of stories; a `needs:` on nearly every story means
191
+ the carve is really a sequence, and you should say so rather than encode it here.
192
+
174
193
  **Slugs** mirror `tiny-spec-create`: derive from the platform key when bound — ADO
175
194
  `AB#77`→`ado-77`, GitHub `#42`→`gh-42`, Monday item→`monday-<id>`, Jira `PROJ-123`
176
195
  verbatim; **ad-hoc** → a short kebab-case slug (`email-login`). The tracker id is
@@ -5,4 +5,4 @@ into your Claude Code config directory (``~/.claude/`` by default). It adds no
5
5
  runtime behavior to the suite itself.
6
6
  """
7
7
 
8
- __version__ = "1.0.1"
8
+ __version__ = "1.2.0"
File without changes
File without changes
File without changes
File without changes