tiny-spec 1.1.0__tar.gz → 2.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: tiny-spec
3
- Version: 1.1.0
3
+ Version: 2.0.0
4
4
  Summary: A tiny, opinionated take on spec-driven development.
5
5
  Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
6
6
  Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
@@ -71,13 +71,13 @@ committed.
71
71
  That core is **three skills and two agents**. In front of it sit **two front doors** —
72
72
  pick the one that matches where you're starting. Over the top sits **one router**,
73
73
  `tiny-spec-run`, which drives the chain and, when you ask it to, works a whole list of
74
- stories through to merged code. No config file, no build step.
74
+ features through to merged code. No config file, no build step.
75
75
 
76
76
  ```
77
77
  GREENFIELD BROWNFIELD
78
78
  starting from an idea starting from a codebase
79
79
  tiny-spec-scope tiny-spec-adopt
80
- idea → Features → Stories real code → constitution
80
+ idea → Features real code → constitution
81
81
  BREAKDOWN.md constitution.md
82
82
  \ /
83
83
  └───────────┬───────────┘
@@ -85,16 +85,16 @@ stories through to merged code. No config file, no build step.
85
85
  tiny-spec-create → tiny-spec-plan → tiny-spec-build
86
86
  intent design + tasks per-task loop
87
87
  SPEC.md PLAN.md plan → implement
88
- tasks.md → review → commit
88
+ (+ ## Tasks) → review → commit
89
89
 
90
90
  tiny-spec-run one router. Walks the chain and stops before build —
91
- or, asked to, builds each story and merges it.
91
+ or, asked to, builds each feature and merges it.
92
92
  tiny-spec-design optional. Wireframes → tokens + gradeable screens.
93
93
  ```
94
94
 
95
95
  **Pick one front door, once per project.** Starting from an idea with no code yet? Run
96
96
  `tiny-spec-scope` — it interviews the idea into a `BREAKDOWN.md`, a flat list of
97
- Features → Stories with draft acceptance criteria. Working in a codebase that already
97
+ well-defined Features with draft acceptance criteria. Working in a codebase that already
98
98
  exists? Run `tiny-spec-adopt` — it reads your repo and derives the constitution from
99
99
  what's actually there: your real lint and test commands, your real layout, your real
100
100
  conventions. Have a single known ticket in a project that's already set up? Skip both
@@ -110,7 +110,7 @@ receipt rather than the adjective:
110
110
 
111
111
  | | skills / commands | agents | config | artifacts per feature |
112
112
  |---|---|---|---|---|
113
- | **tiny-spec** | **7** (3 core + 1 router + 3 optional) | **2** | **none** | **`SPEC` `PLAN` `tasks`** |
113
+ | **tiny-spec** | **7** (3 core + 1 router + 3 optional) | **2** | **none** | **`SPEC` `PLAN`** |
114
114
  | [GitHub Spec Kit](https://github.com/github/spec-kit) | 10 | — | `specify init` | `spec` `plan` `tasks` `checklist` `constitution` `research` `data-model` `contracts/` `quickstart` |
115
115
  | [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
116
116
  | [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
@@ -197,7 +197,7 @@ uvx tiny-spec install
197
197
  Restart Claude Code so it picks up the new skills, then run the flow in your project:
198
198
 
199
199
  ```
200
- /tiny-spec-scope # starting from an idea: interview it into stories (BREAKDOWN.md)
200
+ /tiny-spec-scope # starting from an idea: interview it into features (BREAKDOWN.md)
201
201
  /tiny-spec-adopt # starting from a codebase: derive the constitution from real code
202
202
  /tiny-spec-create # capture intent and requirements (binds a ticket, optional)
203
203
  /tiny-spec-plan # design it, harden the constitution, slice the task list
@@ -219,10 +219,10 @@ forward. It writes nothing itself, it only delegates. By default it **stops befo
219
219
  Ask it to build and it goes all the way instead:
220
220
 
221
221
  ```
222
- /tiny-spec-run build the backlog # per story: branch → plan → build → merge → next
222
+ /tiny-spec-run build the backlog # per feature: branch → plan → build → merge → next
223
223
  ```
224
224
 
225
- It reads your `BREAKDOWN.md` (or a list you paste) and works the stories one after
225
+ It reads your `BREAKDOWN.md` (or a list you paste) and works the features one after
226
226
  another, merging each finished branch into `main` locally before starting the next —
227
227
  until it's done or reaches a **terminal state** it names out loud. See
228
228
  [Working a whole list](#working-a-whole-list). Which of the two it does is decided from
@@ -262,8 +262,8 @@ or install one set at a time.
262
262
  ### See a finished run first
263
263
 
264
264
  [`examples/todo-cli/`](examples/todo-cli/) is a real run of the flow on one small
265
- ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`,
266
- `tasks.md` and `constitution.md` the suite wrote, and the code and tests it produced.
265
+ ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`
266
+ and `constitution.md` the suite wrote, and the code and tests it produced.
267
267
  The tests pass; you can clone it and run the gate yourself.
268
268
 
269
269
  ## How it works
@@ -464,7 +464,7 @@ proved the same thing repeatedly and was the slowest part of the loop.
464
464
 
465
465
  ```mermaid
466
466
  flowchart TB
467
- SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
467
+ SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[PLAN.md ## Tasks<br/>checklist]
468
468
 
469
469
  TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
470
470
  TASKS --> P[Plan task]
@@ -498,17 +498,39 @@ resumes from the checklist state.
498
498
  ### Working a whole list
499
499
 
500
500
  Ask `/tiny-spec-run` to build — "build the backlog", "work through the breakdown",
501
- "spec it out and build it" — and it takes a list of stories and works them one after
502
- another. Per story it does the same four moves:
501
+ "spec it out and build it" — and it takes a list of features and works them in batches.
502
+ Per feature it does the same four moves:
503
503
 
504
504
  ```
505
505
  cut a branch from main → walk the chain → tiny-spec-build → merge back to main
506
506
  ```
507
507
 
508
- Then the next story. Each branch is cut **fresh from main**, so story 3 sees stories 1
509
- and 2 already merged — which is what makes an ordered list build correctly.
508
+ Each branch is cut **fresh from main**, so a later feature sees the earlier ones already
509
+ merged — which is what makes an ordered list build correctly.
510
510
 
511
- **The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
511
+ **Independent features build at the same time.** A feature can declare what it must follow
512
+ with a `needs:` line in `BREAKDOWN.md`; everything with no unmet `needs:` forms a batch
513
+ and builds **concurrently, one git worktree per feature**, three at a time by default.
514
+ The batch merges, then the next one starts.
515
+
516
+ ```
517
+ ## Feature: expose both helpers on a CLI slug: cli
518
+
519
+ - AC: `textkit slugify "Hi There"` prints "hi-there"
520
+ - needs: slugify, wordwrap
521
+ ```
522
+
523
+ Omit `needs:` when a feature stands alone — that's the common case, and the field is meant
524
+ to be rare. A `needs:` you didn't need costs you parallelism forever; one you missed
525
+ costs a single merge conflict, which the run already catches and halts on. You can also
526
+ just name the set yourself at invocation ("build these three at once"), which overrides
527
+ the graph. A cycle, or a `needs:` naming a feature that isn't there, stops the run rather
528
+ than being guessed past.
529
+
530
+ **Tasks *inside* a feature never run in parallel.** They share files and each one assumes
531
+ the last landed, so they stay strictly sequential. Parallelism is across features only.
532
+
533
+ **The list is `BREAKDOWN.md` by default** — its `## Feature:` entries, in file order,
512
534
  each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
513
535
  a list at invocation instead and that wins; but a bare feature name has no acceptance
514
536
  criteria, so `tiny-spec-create` will interview you when it reaches it. That's the
@@ -518,22 +540,26 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
518
540
 
519
541
  | | |
520
542
  |---|---|
521
- | `done` | every story built **and merged** |
543
+ | `done` | every feature built **and merged** |
522
544
  | `blocked` | an upstream document is wrong — go fix the spec or the plan |
523
545
  | `exhausted` | a task stayed red past two fix attempts |
524
546
  | `paused` | it reached a `pause:` point |
525
547
  | `fork` | a real either/or the plan doesn't answer |
526
- | `conflict` | a story's branch wouldn't merge cleanly |
548
+ | `conflict` | a feature's branch wouldn't merge cleanly |
527
549
 
528
- **Only `done` means the work is built** — and in a story run, that means *all* of them.
529
- Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
550
+ **Only `done` means the work is built** — and in a feature run, that means *all* of them.
551
+ Stopping at feature 2 of 7 and reporting "done" is what autonomous loops get wrong most
530
552
  often, so the state is always named alongside what merged and what's still untouched.
531
553
 
532
- **A halt stops the whole run, not just that story.** Later stories in a list you wrote
533
- top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
534
- just produces a second, more confusing failure downstream.
554
+ **A halt stops the lane it happened in, and ends the run after that batch.** Its siblings
555
+ were declared independent, so they finish and merge — killing working lanes because one
556
+ failed throws away good work. But the run does not start the next batch: later features
557
+ usually assume the earlier ones landed, so skipping ahead past a failure just produces a
558
+ second, more confusing failure downstream. With more than one lane you get each feature's
559
+ own state, and the run's state is the worst of them — four green lanes and one `blocked`
560
+ is a `blocked` run.
535
561
 
536
- **Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
562
+ **Pause points are technical, not per-feature.** Any task can carry a `pause:` line, and
537
563
  the build halts *before* running it:
538
564
 
539
565
  ```
@@ -545,20 +571,23 @@ the build halts *before* running it:
545
571
  `tiny-spec-plan` proposes these for genuinely irreversible work — migrations,
546
572
  destructive file operations, a new dependency, an auth boundary, a public API contract.
547
573
  You can also give the run a standing policy up front ("halt before anything that touches
548
- auth") and it gets applied as each story's tasks are sliced.
549
-
550
- **What it will not do to your repo.** It runs exactly five git commands — `switch`,
551
- `switch -c`, `merge --no-ff`, `merge --abort`, and reads. It refuses to start on a dirty
552
- tree. It **never pushes**, never rebases, never resets, never deletes a branch, and
553
- never opens a PR. Merges are local, so a bad run is one `git reset` away; publishing
554
- stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
555
- the undo command rather than running it.
556
-
557
- **Walk away and come back.** Progress isn't written down, it's derived: a story whose
558
- ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
574
+ auth") and it gets applied as each feature's tasks are sliced.
575
+
576
+ **What it will not do to your repo.** It runs exactly seven git commands — `switch`,
577
+ `switch -c`, `merge --no-ff`, `merge --abort`, `worktree add`, `worktree list`, and
578
+ reads. It refuses to start on a dirty tree, or if a worktree path it needs already
579
+ exists. It **never pushes**, never rebases, never resets, never deletes a branch, never
580
+ removes a worktree, and never opens a PR. Merges are local, so a bad run is one
581
+ `git reset` away; publishing stays yours. If the gate goes red after a merge it leaves
582
+ the merge alone and tells you the undo command rather than running it — and it hands
583
+ back the `git worktree remove` commands for the lanes instead of running those either,
584
+ since a halted lane's worktree is the tree you need to look at.
585
+
586
+ **Walk away and come back.** Progress isn't written down, it's derived: a feature whose
587
+ ticked `## Tasks` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
559
588
  progress, no directory means not started. Ask again tomorrow in a fresh session and it
560
589
  picks up where it stopped. No run-state file, no lock, no budget to configure — the
561
- story list *is* the budget.
590
+ feature list *is* the budget.
562
591
 
563
592
  **It never fixes a blocker for you.** A blocker means one of your documents is wrong,
564
593
  and a run allowed to rewrite the requirement its own task just failed would be grading
@@ -583,7 +612,7 @@ It is namespaced per ticket, with a shared spine at the root:
583
612
  constitution.md project-wide, shared across tickets
584
613
  memory.md operational lessons, shared across tickets
585
614
  <ticket-id>/ one directory per ticket (PROJ-123/, gh-42/, …)
586
- SPEC.md PLAN.md tasks.md decisions.md
615
+ SPEC.md PLAN.md decisions.md (PLAN.md ends in the ## Tasks checklist)
587
616
  ```
588
617
 
589
618
  `BREAKDOWN.md` and `design/` sit at your project root rather than inside `.spec/`,
@@ -34,13 +34,13 @@ committed.
34
34
  That core is **three skills and two agents**. In front of it sit **two front doors** —
35
35
  pick the one that matches where you're starting. Over the top sits **one router**,
36
36
  `tiny-spec-run`, which drives the chain and, when you ask it to, works a whole list of
37
- stories through to merged code. No config file, no build step.
37
+ features through to merged code. No config file, no build step.
38
38
 
39
39
  ```
40
40
  GREENFIELD BROWNFIELD
41
41
  starting from an idea starting from a codebase
42
42
  tiny-spec-scope tiny-spec-adopt
43
- idea → Features → Stories real code → constitution
43
+ idea → Features real code → constitution
44
44
  BREAKDOWN.md constitution.md
45
45
  \ /
46
46
  └───────────┬───────────┘
@@ -48,16 +48,16 @@ stories through to merged code. No config file, no build step.
48
48
  tiny-spec-create → tiny-spec-plan → tiny-spec-build
49
49
  intent design + tasks per-task loop
50
50
  SPEC.md PLAN.md plan → implement
51
- tasks.md → review → commit
51
+ (+ ## Tasks) → review → commit
52
52
 
53
53
  tiny-spec-run one router. Walks the chain and stops before build —
54
- or, asked to, builds each story and merges it.
54
+ or, asked to, builds each feature and merges it.
55
55
  tiny-spec-design optional. Wireframes → tokens + gradeable screens.
56
56
  ```
57
57
 
58
58
  **Pick one front door, once per project.** Starting from an idea with no code yet? Run
59
59
  `tiny-spec-scope` — it interviews the idea into a `BREAKDOWN.md`, a flat list of
60
- Features → Stories with draft acceptance criteria. Working in a codebase that already
60
+ well-defined Features with draft acceptance criteria. Working in a codebase that already
61
61
  exists? Run `tiny-spec-adopt` — it reads your repo and derives the constitution from
62
62
  what's actually there: your real lint and test commands, your real layout, your real
63
63
  conventions. Have a single known ticket in a project that's already set up? Skip both
@@ -73,7 +73,7 @@ receipt rather than the adjective:
73
73
 
74
74
  | | skills / commands | agents | config | artifacts per feature |
75
75
  |---|---|---|---|---|
76
- | **tiny-spec** | **7** (3 core + 1 router + 3 optional) | **2** | **none** | **`SPEC` `PLAN` `tasks`** |
76
+ | **tiny-spec** | **7** (3 core + 1 router + 3 optional) | **2** | **none** | **`SPEC` `PLAN`** |
77
77
  | [GitHub Spec Kit](https://github.com/github/spec-kit) | 10 | — | `specify init` | `spec` `plan` `tasks` `checklist` `constitution` `research` `data-model` `contracts/` `quickstart` |
78
78
  | [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
79
79
  | [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
@@ -160,7 +160,7 @@ uvx tiny-spec install
160
160
  Restart Claude Code so it picks up the new skills, then run the flow in your project:
161
161
 
162
162
  ```
163
- /tiny-spec-scope # starting from an idea: interview it into stories (BREAKDOWN.md)
163
+ /tiny-spec-scope # starting from an idea: interview it into features (BREAKDOWN.md)
164
164
  /tiny-spec-adopt # starting from a codebase: derive the constitution from real code
165
165
  /tiny-spec-create # capture intent and requirements (binds a ticket, optional)
166
166
  /tiny-spec-plan # design it, harden the constitution, slice the task list
@@ -182,10 +182,10 @@ forward. It writes nothing itself, it only delegates. By default it **stops befo
182
182
  Ask it to build and it goes all the way instead:
183
183
 
184
184
  ```
185
- /tiny-spec-run build the backlog # per story: branch → plan → build → merge → next
185
+ /tiny-spec-run build the backlog # per feature: branch → plan → build → merge → next
186
186
  ```
187
187
 
188
- It reads your `BREAKDOWN.md` (or a list you paste) and works the stories one after
188
+ It reads your `BREAKDOWN.md` (or a list you paste) and works the features one after
189
189
  another, merging each finished branch into `main` locally before starting the next —
190
190
  until it's done or reaches a **terminal state** it names out loud. See
191
191
  [Working a whole list](#working-a-whole-list). Which of the two it does is decided from
@@ -225,8 +225,8 @@ or install one set at a time.
225
225
  ### See a finished run first
226
226
 
227
227
  [`examples/todo-cli/`](examples/todo-cli/) is a real run of the flow on one small
228
- ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`,
229
- `tasks.md` and `constitution.md` the suite wrote, and the code and tests it produced.
228
+ ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`
229
+ and `constitution.md` the suite wrote, and the code and tests it produced.
230
230
  The tests pass; you can clone it and run the gate yourself.
231
231
 
232
232
  ## How it works
@@ -427,7 +427,7 @@ proved the same thing repeatedly and was the slowest part of the loop.
427
427
 
428
428
  ```mermaid
429
429
  flowchart TB
430
- SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
430
+ SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[PLAN.md ## Tasks<br/>checklist]
431
431
 
432
432
  TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
433
433
  TASKS --> P[Plan task]
@@ -461,17 +461,39 @@ resumes from the checklist state.
461
461
  ### Working a whole list
462
462
 
463
463
  Ask `/tiny-spec-run` to build — "build the backlog", "work through the breakdown",
464
- "spec it out and build it" — and it takes a list of stories and works them one after
465
- another. Per story it does the same four moves:
464
+ "spec it out and build it" — and it takes a list of features and works them in batches.
465
+ Per feature it does the same four moves:
466
466
 
467
467
  ```
468
468
  cut a branch from main → walk the chain → tiny-spec-build → merge back to main
469
469
  ```
470
470
 
471
- Then the next story. Each branch is cut **fresh from main**, so story 3 sees stories 1
472
- and 2 already merged — which is what makes an ordered list build correctly.
471
+ Each branch is cut **fresh from main**, so a later feature sees the earlier ones already
472
+ merged — which is what makes an ordered list build correctly.
473
473
 
474
- **The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
474
+ **Independent features build at the same time.** A feature can declare what it must follow
475
+ with a `needs:` line in `BREAKDOWN.md`; everything with no unmet `needs:` forms a batch
476
+ and builds **concurrently, one git worktree per feature**, three at a time by default.
477
+ The batch merges, then the next one starts.
478
+
479
+ ```
480
+ ## Feature: expose both helpers on a CLI slug: cli
481
+
482
+ - AC: `textkit slugify "Hi There"` prints "hi-there"
483
+ - needs: slugify, wordwrap
484
+ ```
485
+
486
+ Omit `needs:` when a feature stands alone — that's the common case, and the field is meant
487
+ to be rare. A `needs:` you didn't need costs you parallelism forever; one you missed
488
+ costs a single merge conflict, which the run already catches and halts on. You can also
489
+ just name the set yourself at invocation ("build these three at once"), which overrides
490
+ the graph. A cycle, or a `needs:` naming a feature that isn't there, stops the run rather
491
+ than being guessed past.
492
+
493
+ **Tasks *inside* a feature never run in parallel.** They share files and each one assumes
494
+ the last landed, so they stay strictly sequential. Parallelism is across features only.
495
+
496
+ **The list is `BREAKDOWN.md` by default** — its `## Feature:` entries, in file order,
475
497
  each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
476
498
  a list at invocation instead and that wins; but a bare feature name has no acceptance
477
499
  criteria, so `tiny-spec-create` will interview you when it reaches it. That's the
@@ -481,22 +503,26 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
481
503
 
482
504
  | | |
483
505
  |---|---|
484
- | `done` | every story built **and merged** |
506
+ | `done` | every feature built **and merged** |
485
507
  | `blocked` | an upstream document is wrong — go fix the spec or the plan |
486
508
  | `exhausted` | a task stayed red past two fix attempts |
487
509
  | `paused` | it reached a `pause:` point |
488
510
  | `fork` | a real either/or the plan doesn't answer |
489
- | `conflict` | a story's branch wouldn't merge cleanly |
511
+ | `conflict` | a feature's branch wouldn't merge cleanly |
490
512
 
491
- **Only `done` means the work is built** — and in a story run, that means *all* of them.
492
- Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
513
+ **Only `done` means the work is built** — and in a feature run, that means *all* of them.
514
+ Stopping at feature 2 of 7 and reporting "done" is what autonomous loops get wrong most
493
515
  often, so the state is always named alongside what merged and what's still untouched.
494
516
 
495
- **A halt stops the whole run, not just that story.** Later stories in a list you wrote
496
- top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
497
- just produces a second, more confusing failure downstream.
517
+ **A halt stops the lane it happened in, and ends the run after that batch.** Its siblings
518
+ were declared independent, so they finish and merge — killing working lanes because one
519
+ failed throws away good work. But the run does not start the next batch: later features
520
+ usually assume the earlier ones landed, so skipping ahead past a failure just produces a
521
+ second, more confusing failure downstream. With more than one lane you get each feature's
522
+ own state, and the run's state is the worst of them — four green lanes and one `blocked`
523
+ is a `blocked` run.
498
524
 
499
- **Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
525
+ **Pause points are technical, not per-feature.** Any task can carry a `pause:` line, and
500
526
  the build halts *before* running it:
501
527
 
502
528
  ```
@@ -508,20 +534,23 @@ the build halts *before* running it:
508
534
  `tiny-spec-plan` proposes these for genuinely irreversible work — migrations,
509
535
  destructive file operations, a new dependency, an auth boundary, a public API contract.
510
536
  You can also give the run a standing policy up front ("halt before anything that touches
511
- auth") and it gets applied as each story's tasks are sliced.
512
-
513
- **What it will not do to your repo.** It runs exactly five git commands — `switch`,
514
- `switch -c`, `merge --no-ff`, `merge --abort`, and reads. It refuses to start on a dirty
515
- tree. It **never pushes**, never rebases, never resets, never deletes a branch, and
516
- never opens a PR. Merges are local, so a bad run is one `git reset` away; publishing
517
- stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
518
- the undo command rather than running it.
519
-
520
- **Walk away and come back.** Progress isn't written down, it's derived: a story whose
521
- ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
537
+ auth") and it gets applied as each feature's tasks are sliced.
538
+
539
+ **What it will not do to your repo.** It runs exactly seven git commands — `switch`,
540
+ `switch -c`, `merge --no-ff`, `merge --abort`, `worktree add`, `worktree list`, and
541
+ reads. It refuses to start on a dirty tree, or if a worktree path it needs already
542
+ exists. It **never pushes**, never rebases, never resets, never deletes a branch, never
543
+ removes a worktree, and never opens a PR. Merges are local, so a bad run is one
544
+ `git reset` away; publishing stays yours. If the gate goes red after a merge it leaves
545
+ the merge alone and tells you the undo command rather than running it — and it hands
546
+ back the `git worktree remove` commands for the lanes instead of running those either,
547
+ since a halted lane's worktree is the tree you need to look at.
548
+
549
+ **Walk away and come back.** Progress isn't written down, it's derived: a feature whose
550
+ ticked `## Tasks` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
522
551
  progress, no directory means not started. Ask again tomorrow in a fresh session and it
523
552
  picks up where it stopped. No run-state file, no lock, no budget to configure — the
524
- story list *is* the budget.
553
+ feature list *is* the budget.
525
554
 
526
555
  **It never fixes a blocker for you.** A blocker means one of your documents is wrong,
527
556
  and a run allowed to rewrite the requirement its own task just failed would be grading
@@ -546,7 +575,7 @@ It is namespaced per ticket, with a shared spine at the root:
546
575
  constitution.md project-wide, shared across tickets
547
576
  memory.md operational lessons, shared across tickets
548
577
  <ticket-id>/ one directory per ticket (PROJ-123/, gh-42/, …)
549
- SPEC.md PLAN.md tasks.md decisions.md
578
+ SPEC.md PLAN.md decisions.md (PLAN.md ends in the ## Tasks checklist)
550
579
  ```
551
580
 
552
581
  `BREAKDOWN.md` and `design/` sit at your project root rather than inside `.spec/`,
@@ -15,6 +15,12 @@ pleasantries.
15
15
 
16
16
  Everything you need and nothing you don't:
17
17
 
18
+ - **the working directory** to operate in. Every path you read, write, or run a command
19
+ against resolves against it. It may be a **git worktree** rather than the main checkout
20
+ — a build can run several features at once, each in its own worktree. **Never read or
21
+ write outside the directory you were given**, and never `cd` to a sibling worktree to
22
+ "check something": another feature is being built there right now, and what you find will
23
+ be wrong by the time you act on it;
18
24
  - the **task id**, **description**, and **acceptance** (the outcome that proves it done);
19
25
  - a **`files:` hint** — likely paths to touch (guidance, not a hard boundary);
20
26
  - the full **constitution** (`constitution.md`): Style, Engineering standards,
@@ -13,6 +13,11 @@ back to `tiny-spec-build`; return data, not pleasantries.
13
13
 
14
14
  ## What you receive (the context contract)
15
15
 
16
+ - **the working directory** to operate in — every path and every gate command resolves
17
+ against it. It may be a **git worktree** rather than the main checkout, since a build
18
+ can run several features at once. **Never read, write, or run a gate outside the
19
+ directory you were given**: a sibling worktree holds a different feature mid-build, and
20
+ measuring it would make your verdict meaningless;
16
21
  - the **task id**, **description**, and **acceptance** (the outcome that must hold);
17
22
  - the full **constitution** (`constitution.md`) — especially **Guiding invariants**,
18
23
  **Definition of Done**, and **Verification commands**, plus the **Design system**
@@ -54,6 +59,18 @@ acceptance and the constitution — verified, not inferred?**
54
59
  task on evidence you found insufficient is not** — that is the one thing this whole
55
60
  role exists to prevent. Never narrow a `full` gate to `scoped`.
56
61
 
62
+ **A Verification command that cannot run in this environment is a `FAIL` and a
63
+ blocker — never a `PASS`.** If a documented command is red for a reason the task's
64
+ code cannot fix (the interpreter is too old, a tool isn't installed, the command was
65
+ never runnable as written), say exactly that in `GATE:`, return `FAIL`, and name the
66
+ **constitution** as the upstream document to fix. Proving the task green *somewhere
67
+ else* — a fresh venv, a container, an install you fixed by hand — is **not** a pass:
68
+ you verified a different environment than the one the gate names, and the next task
69
+ will hit the identical red. Diagnosing the cause and demonstrating the task's own code
70
+ is sound is genuinely useful; put it in `FINDINGS` so the fix is cheap. It does not
71
+ change the verdict. A constitution whose gate cannot run is the single most dangerous
72
+ file in the project, because every later task inherits the same false signal.
73
+
57
74
  Say which scope you actually ran in `GATE:`, including that you escalated and why.
58
75
  3. **Exercise the acceptance.** Trigger the task's stated outcome the most
59
76
  black-box way available (CLI > HTTP > public API) with realistic input,
@@ -78,25 +95,23 @@ acceptance and the constitution — verified, not inferred?**
78
95
  - **A selector that matches nothing is a `FAIL`, never a skip.** Either the code
79
96
  didn't build the element or it named it something else — both are real, and both
80
97
  are invisible if you quietly move on. Report the selector and that it was absent.
81
- - Compare each element's numbers to the tokens its row names. Report concrete
82
- deltas ("heading is 28px, `type.heading.lg` is 24px"; "padding 19px is not on
83
- the `space.*` scale"). **Do not pixel-diff the screenshot** — font antialiasing
84
- makes image comparison flaky enough that the check gets ignored, which is
85
- exactly how visual gates die.
86
- - **Check `layout:`** — the arrangement, max width, and the **order** it names.
87
- Use the bounding rectangles: elements listed in order should appear in that
88
- order down the page (or across it, for a row). Every token can be correct on an
89
- element that is in the wrong place.
90
- - **Exercise every state the entry names** — empty, loading, error, success.
91
- Drive the UI into each one and observe what changes; **finding the word in the
92
- source is not evidence**, and it false-passes routinely (a comment saying the
93
- loading state is missing contains "loading"). A surface that renders its happy
94
- path and nothing else is a fail, not a nit.
95
- - Finally, **judge the render against the export.** Everything above proves the
96
- numbers are right. None of it can see an element that is present, on-token, and
97
- invisible — so now look. `Read` each `SCREENSHOT <state> <path>` the `visual:`
98
- command printed, `Read` the `D<n>`'s `export:` image, and grade **every state you
99
- have a screenshot for**, not just the happy path, on four lines:
98
+ - Compare each element's numbers to the tokens its row names, reporting concrete
99
+ deltas ("heading is 28px, `type.heading.lg` is 24px"; "padding 19px is not on the
100
+ `space.*` scale"). **Do not pixel-diff the screenshot** — font antialiasing makes
101
+ image comparison flaky enough that the check gets ignored, which is how visual
102
+ gates die.
103
+ - **Check `layout:`** — the arrangement, max width, and the **order** it names, from
104
+ the bounding rectangles. Every token can be correct on an element in the wrong place.
105
+ - **Exercise every state the entry names** — empty, loading, error, success. Drive the
106
+ UI into each one and observe what changes; **finding the word in the source is not
107
+ evidence**, and it false-passes routinely (a comment saying the loading state is
108
+ missing contains "loading"). A surface that renders its happy path and nothing else
109
+ is a fail, not a nit.
110
+ - Finally, **judge the render against the export.** Everything above proves the numbers
111
+ are right; none of it can see an element that is present, on-token, and invisible —
112
+ so now look. `Read` each `SCREENSHOT <state> <path>` the `visual:` command printed
113
+ and the `D<n>`'s `export:` image, and grade **every state you have a screenshot
114
+ for**, not just the happy path, on four lines:
100
115
 
101
116
  1. **Presence** — is every `elements:` row actually *visible* in the render? At
102
117
  `opacity: 0`, zero height, clipped out of view, hidden behind a sibling, or the
@@ -104,24 +119,21 @@ acceptance and the constitution — verified, not inferred?**
104
119
  not there. **FAIL** — this is why the step exists.
105
120
  2. **Legibility & occlusion** — text clipped, truncated mid-word, overlapping
106
121
  another element, or on a background it can't be read against. **FAIL.**
107
- 3. **Correspondence** — does the render show the same screen as the export: the
108
- same regions, in the reading order `layout:` names? A whole region missing is a
109
- **FAIL**; a stylistic difference is a **flag**.
122
+ 3. **Correspondence** — the same regions as the export, in the reading order
123
+ `layout:` names? A whole region missing is a **FAIL**; a stylistic difference is
124
+ a **flag**.
110
125
  4. **Hierarchy & polish** — emphasis, balance, crowding, alignment. **Always a
111
126
  flag**, never a fail.
112
127
 
113
- Three rules bound it:
114
-
115
- - **The numbers beat your eye on anything they already measured.** Padding that is
116
- on the `space.*` scale but looks cramped is a `flag:`. A color that is exactly
117
- its token but looks washed out is a `flag:`. You may fail only on what
118
- measurement *cannot* see. Contradicting your own numbers sends the executor a
119
- task it cannot fix, and the loop is bounded at two attempts.
120
- - **This is still not a pixel diff.** The export is usually a wireframe — judge
121
- structure and legibility, never visual identity.
122
- - **Cite what you saw.** Name the state whose screenshot the finding came from and
123
- what was in it ("state `error`: caption present in DOM but renders at opacity 0").
124
- An uncited visual claim reads as an opinion and gets ignored.
128
+ Three rules bound it. **The numbers beat your eye on anything they already
129
+ measured** — on-scale padding that looks cramped, or an exact token that looks
130
+ washed out, is a `flag:`; you may fail only on what measurement *cannot* see, since
131
+ contradicting your own numbers sends the executor a task it cannot fix and the loop
132
+ is bounded at two attempts. **This is still not a pixel diff** — the export is
133
+ usually a wireframe, so judge structure and legibility, never visual identity. And
134
+ **cite what you saw**, naming the state the finding came from ("state `error`:
135
+ caption present in DOM but renders at opacity 0") — an uncited visual claim reads as
136
+ an opinion and gets ignored.
125
137
 
126
138
  **If the command printed no `SCREENSHOT` line**, do not run this sub-step and do
127
139
  not eyeball a substitute. Grade on steps 1–3 above, write `judge: not run — visual:
@@ -87,10 +87,10 @@ Each case runs in a throwaway sandbox seeded with `IDEA.md` and the vendored ski
87
87
  ### What gets measured
88
88
 
89
89
  - **structural conformance** (deterministic) — BREAKDOWN has Problem and
90
- Goal & non-goals filled, a Decisions block, ≥1 Feature, Stories with a slug and ≥1 AC; the
90
+ Goal & non-goals filled, a Decisions block, ≥1 Feature, each with a slug and ≥1 AC; the
91
91
  planning skills left no `.spec/` behind.
92
92
  - **hand-off integrity** (LLM judge) — **coverage** (every PRD capability lands in ≥1
93
- story, nothing dropped) and **no fabrication** (every story traces to a capability,
93
+ feature, nothing dropped) and **no fabrication** (every feature traces to a capability,
94
94
  nothing invented). These are the high-value signals; a case PASSes only if structural
95
95
  conformance holds *and* the judge confirms both.
96
96
  - **quality** (LLM judge, reported not gated) — atomicity / user-observable phrasing,
@@ -10,6 +10,9 @@ workflow generates before you run it on your own work.
10
10
  > retouched to match the current artifact formats — a doctored example would defeat
11
11
  > the point. Later versions add a `## Design` section to `SPEC.md` (unused here: this
12
12
  > is a CLI with no design surface) and richer status flags. The shape is the same.
13
+ > One format move was applied after the fact: in 2.0 the checklist stopped being its
14
+ > own `tasks.md` and became `PLAN.md`'s `## Tasks` section, so the original `tasks.md`
15
+ > was moved there verbatim — same tasks, same ticks, no content edited.
13
16
 
14
17
  ## What's here
15
18
 
@@ -20,14 +23,13 @@ TICKET.md the input — a small "todo CLI" ask
20
23
  definition of done, and the verification gate
21
24
  todo-cli/
22
25
  SPEC.md intent + REQ-1..REQ-9 (what "done" means)
23
- PLAN.md the design and how each requirement is covered
24
- tasks.md the ordered checklist, all ticked [x]
26
+ PLAN.md the design, how each requirement is covered, and
27
+ the ordered ## Tasks checklist, all ticked [x]
25
28
  todo.py the produced CLI (stdlib only, 117 lines)
26
29
  test_todo.py the produced end-to-end test (129 lines)
27
30
  ```
28
31
 
29
- Read them in flow order: `TICKET.md` → `SPEC.md` → `PLAN.md` → `tasks.md` →
30
- `todo.py`. The `constitution.md` is the persistent context injected into every task.
32
+ Read them in flow order: `TICKET.md` → `SPEC.md` → `PLAN.md` → `todo.py`. The `constitution.md` is the persistent context injected into every task.
31
33
 
32
34
  ## It passes its own gate
33
35
 
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
4
4
 
5
5
  [project]
6
6
  name = "tiny-spec"
7
- version = "1.1.0"
7
+ version = "2.0.0"
8
8
  description = "A tiny, opinionated take on spec-driven development."
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.9"