tiny-spec 1.1.0__tar.gz → 2.0.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/PKG-INFO +68 -39
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/README.md +67 -38
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/agents/tiny-spec-build-executor.md +6 -0
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/agents/tiny-spec-build-reviewer.md +46 -34
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/docs/eval/README.md +2 -2
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/examples/todo-cli/README.md +6 -4
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/pyproject.toml +1 -1
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/tiny-spec-adopt/SKILL.md +2 -2
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/tiny-spec-build/SKILL.md +91 -30
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/tiny-spec-create/SKILL.md +21 -21
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/tiny-spec-design/SKILL.md +3 -3
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/tiny-spec-plan/SKILL.md +52 -50
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/tiny-spec-run/SKILL.md +178 -122
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/tiny-spec-scope/SKILL.md +72 -52
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/tiny_spec/__init__.py +1 -1
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/.gitignore +0 -0
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/LICENSE +0 -0
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/tiny_spec/cli.py +0 -0
- {tiny_spec-1.1.0 → tiny_spec-2.0.0}/tiny_spec/manifest.json +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: tiny-spec
|
|
3
|
-
Version:
|
|
3
|
+
Version: 2.0.0
|
|
4
4
|
Summary: A tiny, opinionated take on spec-driven development.
|
|
5
5
|
Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
|
|
6
6
|
Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
|
|
@@ -71,13 +71,13 @@ committed.
|
|
|
71
71
|
That core is **three skills and two agents**. In front of it sit **two front doors** —
|
|
72
72
|
pick the one that matches where you're starting. Over the top sits **one router**,
|
|
73
73
|
`tiny-spec-run`, which drives the chain and, when you ask it to, works a whole list of
|
|
74
|
-
|
|
74
|
+
features through to merged code. No config file, no build step.
|
|
75
75
|
|
|
76
76
|
```
|
|
77
77
|
GREENFIELD BROWNFIELD
|
|
78
78
|
starting from an idea starting from a codebase
|
|
79
79
|
tiny-spec-scope tiny-spec-adopt
|
|
80
|
-
idea → Features
|
|
80
|
+
idea → Features real code → constitution
|
|
81
81
|
BREAKDOWN.md constitution.md
|
|
82
82
|
\ /
|
|
83
83
|
└───────────┬───────────┘
|
|
@@ -85,16 +85,16 @@ stories through to merged code. No config file, no build step.
|
|
|
85
85
|
tiny-spec-create → tiny-spec-plan → tiny-spec-build
|
|
86
86
|
intent design + tasks per-task loop
|
|
87
87
|
SPEC.md PLAN.md plan → implement
|
|
88
|
-
|
|
88
|
+
(+ ## Tasks) → review → commit
|
|
89
89
|
|
|
90
90
|
tiny-spec-run one router. Walks the chain and stops before build —
|
|
91
|
-
or, asked to, builds each
|
|
91
|
+
or, asked to, builds each feature and merges it.
|
|
92
92
|
tiny-spec-design optional. Wireframes → tokens + gradeable screens.
|
|
93
93
|
```
|
|
94
94
|
|
|
95
95
|
**Pick one front door, once per project.** Starting from an idea with no code yet? Run
|
|
96
96
|
`tiny-spec-scope` — it interviews the idea into a `BREAKDOWN.md`, a flat list of
|
|
97
|
-
Features
|
|
97
|
+
well-defined Features with draft acceptance criteria. Working in a codebase that already
|
|
98
98
|
exists? Run `tiny-spec-adopt` — it reads your repo and derives the constitution from
|
|
99
99
|
what's actually there: your real lint and test commands, your real layout, your real
|
|
100
100
|
conventions. Have a single known ticket in a project that's already set up? Skip both
|
|
@@ -110,7 +110,7 @@ receipt rather than the adjective:
|
|
|
110
110
|
|
|
111
111
|
| | skills / commands | agents | config | artifacts per feature |
|
|
112
112
|
|---|---|---|---|---|
|
|
113
|
-
| **tiny-spec** | **7** (3 core + 1 router + 3 optional) | **2** | **none** | **`SPEC` `PLAN
|
|
113
|
+
| **tiny-spec** | **7** (3 core + 1 router + 3 optional) | **2** | **none** | **`SPEC` `PLAN`** |
|
|
114
114
|
| [GitHub Spec Kit](https://github.com/github/spec-kit) | 10 | — | `specify init` | `spec` `plan` `tasks` `checklist` `constitution` `research` `data-model` `contracts/` `quickstart` |
|
|
115
115
|
| [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
|
|
116
116
|
| [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
|
|
@@ -197,7 +197,7 @@ uvx tiny-spec install
|
|
|
197
197
|
Restart Claude Code so it picks up the new skills, then run the flow in your project:
|
|
198
198
|
|
|
199
199
|
```
|
|
200
|
-
/tiny-spec-scope # starting from an idea: interview it into
|
|
200
|
+
/tiny-spec-scope # starting from an idea: interview it into features (BREAKDOWN.md)
|
|
201
201
|
/tiny-spec-adopt # starting from a codebase: derive the constitution from real code
|
|
202
202
|
/tiny-spec-create # capture intent and requirements (binds a ticket, optional)
|
|
203
203
|
/tiny-spec-plan # design it, harden the constitution, slice the task list
|
|
@@ -219,10 +219,10 @@ forward. It writes nothing itself, it only delegates. By default it **stops befo
|
|
|
219
219
|
Ask it to build and it goes all the way instead:
|
|
220
220
|
|
|
221
221
|
```
|
|
222
|
-
/tiny-spec-run build the backlog # per
|
|
222
|
+
/tiny-spec-run build the backlog # per feature: branch → plan → build → merge → next
|
|
223
223
|
```
|
|
224
224
|
|
|
225
|
-
It reads your `BREAKDOWN.md` (or a list you paste) and works the
|
|
225
|
+
It reads your `BREAKDOWN.md` (or a list you paste) and works the features one after
|
|
226
226
|
another, merging each finished branch into `main` locally before starting the next —
|
|
227
227
|
until it's done or reaches a **terminal state** it names out loud. See
|
|
228
228
|
[Working a whole list](#working-a-whole-list). Which of the two it does is decided from
|
|
@@ -262,8 +262,8 @@ or install one set at a time.
|
|
|
262
262
|
### See a finished run first
|
|
263
263
|
|
|
264
264
|
[`examples/todo-cli/`](examples/todo-cli/) is a real run of the flow on one small
|
|
265
|
-
ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md
|
|
266
|
-
|
|
265
|
+
ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`
|
|
266
|
+
and `constitution.md` the suite wrote, and the code and tests it produced.
|
|
267
267
|
The tests pass; you can clone it and run the gate yourself.
|
|
268
268
|
|
|
269
269
|
## How it works
|
|
@@ -464,7 +464,7 @@ proved the same thing repeatedly and was the slowest part of the loop.
|
|
|
464
464
|
|
|
465
465
|
```mermaid
|
|
466
466
|
flowchart TB
|
|
467
|
-
SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[
|
|
467
|
+
SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[PLAN.md ## Tasks<br/>checklist]
|
|
468
468
|
|
|
469
469
|
TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
|
|
470
470
|
TASKS --> P[Plan task]
|
|
@@ -498,17 +498,39 @@ resumes from the checklist state.
|
|
|
498
498
|
### Working a whole list
|
|
499
499
|
|
|
500
500
|
Ask `/tiny-spec-run` to build — "build the backlog", "work through the breakdown",
|
|
501
|
-
"spec it out and build it" — and it takes a list of
|
|
502
|
-
|
|
501
|
+
"spec it out and build it" — and it takes a list of features and works them in batches.
|
|
502
|
+
Per feature it does the same four moves:
|
|
503
503
|
|
|
504
504
|
```
|
|
505
505
|
cut a branch from main → walk the chain → tiny-spec-build → merge back to main
|
|
506
506
|
```
|
|
507
507
|
|
|
508
|
-
|
|
509
|
-
|
|
508
|
+
Each branch is cut **fresh from main**, so a later feature sees the earlier ones already
|
|
509
|
+
merged — which is what makes an ordered list build correctly.
|
|
510
510
|
|
|
511
|
-
**
|
|
511
|
+
**Independent features build at the same time.** A feature can declare what it must follow
|
|
512
|
+
with a `needs:` line in `BREAKDOWN.md`; everything with no unmet `needs:` forms a batch
|
|
513
|
+
and builds **concurrently, one git worktree per feature**, three at a time by default.
|
|
514
|
+
The batch merges, then the next one starts.
|
|
515
|
+
|
|
516
|
+
```
|
|
517
|
+
## Feature: expose both helpers on a CLI slug: cli
|
|
518
|
+
|
|
519
|
+
- AC: `textkit slugify "Hi There"` prints "hi-there"
|
|
520
|
+
- needs: slugify, wordwrap
|
|
521
|
+
```
|
|
522
|
+
|
|
523
|
+
Omit `needs:` when a feature stands alone — that's the common case, and the field is meant
|
|
524
|
+
to be rare. A `needs:` you didn't need costs you parallelism forever; one you missed
|
|
525
|
+
costs a single merge conflict, which the run already catches and halts on. You can also
|
|
526
|
+
just name the set yourself at invocation ("build these three at once"), which overrides
|
|
527
|
+
the graph. A cycle, or a `needs:` naming a feature that isn't there, stops the run rather
|
|
528
|
+
than being guessed past.
|
|
529
|
+
|
|
530
|
+
**Tasks *inside* a feature never run in parallel.** They share files and each one assumes
|
|
531
|
+
the last landed, so they stay strictly sequential. Parallelism is across features only.
|
|
532
|
+
|
|
533
|
+
**The list is `BREAKDOWN.md` by default** — its `## Feature:` entries, in file order,
|
|
512
534
|
each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
|
|
513
535
|
a list at invocation instead and that wins; but a bare feature name has no acceptance
|
|
514
536
|
criteria, so `tiny-spec-create` will interview you when it reaches it. That's the
|
|
@@ -518,22 +540,26 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
|
|
|
518
540
|
|
|
519
541
|
| | |
|
|
520
542
|
|---|---|
|
|
521
|
-
| `done` | every
|
|
543
|
+
| `done` | every feature built **and merged** |
|
|
522
544
|
| `blocked` | an upstream document is wrong — go fix the spec or the plan |
|
|
523
545
|
| `exhausted` | a task stayed red past two fix attempts |
|
|
524
546
|
| `paused` | it reached a `pause:` point |
|
|
525
547
|
| `fork` | a real either/or the plan doesn't answer |
|
|
526
|
-
| `conflict` | a
|
|
548
|
+
| `conflict` | a feature's branch wouldn't merge cleanly |
|
|
527
549
|
|
|
528
|
-
**Only `done` means the work is built** — and in a
|
|
529
|
-
Stopping at
|
|
550
|
+
**Only `done` means the work is built** — and in a feature run, that means *all* of them.
|
|
551
|
+
Stopping at feature 2 of 7 and reporting "done" is what autonomous loops get wrong most
|
|
530
552
|
often, so the state is always named alongside what merged and what's still untouched.
|
|
531
553
|
|
|
532
|
-
**A halt stops the
|
|
533
|
-
|
|
534
|
-
|
|
554
|
+
**A halt stops the lane it happened in, and ends the run after that batch.** Its siblings
|
|
555
|
+
were declared independent, so they finish and merge — killing working lanes because one
|
|
556
|
+
failed throws away good work. But the run does not start the next batch: later features
|
|
557
|
+
usually assume the earlier ones landed, so skipping ahead past a failure just produces a
|
|
558
|
+
second, more confusing failure downstream. With more than one lane you get each feature's
|
|
559
|
+
own state, and the run's state is the worst of them — four green lanes and one `blocked`
|
|
560
|
+
is a `blocked` run.
|
|
535
561
|
|
|
536
|
-
**Pause points are technical, not per-
|
|
562
|
+
**Pause points are technical, not per-feature.** Any task can carry a `pause:` line, and
|
|
537
563
|
the build halts *before* running it:
|
|
538
564
|
|
|
539
565
|
```
|
|
@@ -545,20 +571,23 @@ the build halts *before* running it:
|
|
|
545
571
|
`tiny-spec-plan` proposes these for genuinely irreversible work — migrations,
|
|
546
572
|
destructive file operations, a new dependency, an auth boundary, a public API contract.
|
|
547
573
|
You can also give the run a standing policy up front ("halt before anything that touches
|
|
548
|
-
auth") and it gets applied as each
|
|
549
|
-
|
|
550
|
-
**What it will not do to your repo.** It runs exactly
|
|
551
|
-
`switch -c`, `merge --no-ff`, `merge --abort`,
|
|
552
|
-
|
|
553
|
-
|
|
554
|
-
|
|
555
|
-
the
|
|
556
|
-
|
|
557
|
-
|
|
558
|
-
|
|
574
|
+
auth") and it gets applied as each feature's tasks are sliced.
|
|
575
|
+
|
|
576
|
+
**What it will not do to your repo.** It runs exactly seven git commands — `switch`,
|
|
577
|
+
`switch -c`, `merge --no-ff`, `merge --abort`, `worktree add`, `worktree list`, and
|
|
578
|
+
reads. It refuses to start on a dirty tree, or if a worktree path it needs already
|
|
579
|
+
exists. It **never pushes**, never rebases, never resets, never deletes a branch, never
|
|
580
|
+
removes a worktree, and never opens a PR. Merges are local, so a bad run is one
|
|
581
|
+
`git reset` away; publishing stays yours. If the gate goes red after a merge it leaves
|
|
582
|
+
the merge alone and tells you the undo command rather than running it — and it hands
|
|
583
|
+
back the `git worktree remove` commands for the lanes instead of running those either,
|
|
584
|
+
since a halted lane's worktree is the tree you need to look at.
|
|
585
|
+
|
|
586
|
+
**Walk away and come back.** Progress isn't written down, it's derived: a feature whose
|
|
587
|
+
ticked `## Tasks` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
|
|
559
588
|
progress, no directory means not started. Ask again tomorrow in a fresh session and it
|
|
560
589
|
picks up where it stopped. No run-state file, no lock, no budget to configure — the
|
|
561
|
-
|
|
590
|
+
feature list *is* the budget.
|
|
562
591
|
|
|
563
592
|
**It never fixes a blocker for you.** A blocker means one of your documents is wrong,
|
|
564
593
|
and a run allowed to rewrite the requirement its own task just failed would be grading
|
|
@@ -583,7 +612,7 @@ It is namespaced per ticket, with a shared spine at the root:
|
|
|
583
612
|
constitution.md project-wide, shared across tickets
|
|
584
613
|
memory.md operational lessons, shared across tickets
|
|
585
614
|
<ticket-id>/ one directory per ticket (PROJ-123/, gh-42/, …)
|
|
586
|
-
SPEC.md PLAN.md
|
|
615
|
+
SPEC.md PLAN.md decisions.md (PLAN.md ends in the ## Tasks checklist)
|
|
587
616
|
```
|
|
588
617
|
|
|
589
618
|
`BREAKDOWN.md` and `design/` sit at your project root rather than inside `.spec/`,
|
|
@@ -34,13 +34,13 @@ committed.
|
|
|
34
34
|
That core is **three skills and two agents**. In front of it sit **two front doors** —
|
|
35
35
|
pick the one that matches where you're starting. Over the top sits **one router**,
|
|
36
36
|
`tiny-spec-run`, which drives the chain and, when you ask it to, works a whole list of
|
|
37
|
-
|
|
37
|
+
features through to merged code. No config file, no build step.
|
|
38
38
|
|
|
39
39
|
```
|
|
40
40
|
GREENFIELD BROWNFIELD
|
|
41
41
|
starting from an idea starting from a codebase
|
|
42
42
|
tiny-spec-scope tiny-spec-adopt
|
|
43
|
-
idea → Features
|
|
43
|
+
idea → Features real code → constitution
|
|
44
44
|
BREAKDOWN.md constitution.md
|
|
45
45
|
\ /
|
|
46
46
|
└───────────┬───────────┘
|
|
@@ -48,16 +48,16 @@ stories through to merged code. No config file, no build step.
|
|
|
48
48
|
tiny-spec-create → tiny-spec-plan → tiny-spec-build
|
|
49
49
|
intent design + tasks per-task loop
|
|
50
50
|
SPEC.md PLAN.md plan → implement
|
|
51
|
-
|
|
51
|
+
(+ ## Tasks) → review → commit
|
|
52
52
|
|
|
53
53
|
tiny-spec-run one router. Walks the chain and stops before build —
|
|
54
|
-
or, asked to, builds each
|
|
54
|
+
or, asked to, builds each feature and merges it.
|
|
55
55
|
tiny-spec-design optional. Wireframes → tokens + gradeable screens.
|
|
56
56
|
```
|
|
57
57
|
|
|
58
58
|
**Pick one front door, once per project.** Starting from an idea with no code yet? Run
|
|
59
59
|
`tiny-spec-scope` — it interviews the idea into a `BREAKDOWN.md`, a flat list of
|
|
60
|
-
Features
|
|
60
|
+
well-defined Features with draft acceptance criteria. Working in a codebase that already
|
|
61
61
|
exists? Run `tiny-spec-adopt` — it reads your repo and derives the constitution from
|
|
62
62
|
what's actually there: your real lint and test commands, your real layout, your real
|
|
63
63
|
conventions. Have a single known ticket in a project that's already set up? Skip both
|
|
@@ -73,7 +73,7 @@ receipt rather than the adjective:
|
|
|
73
73
|
|
|
74
74
|
| | skills / commands | agents | config | artifacts per feature |
|
|
75
75
|
|---|---|---|---|---|
|
|
76
|
-
| **tiny-spec** | **7** (3 core + 1 router + 3 optional) | **2** | **none** | **`SPEC` `PLAN
|
|
76
|
+
| **tiny-spec** | **7** (3 core + 1 router + 3 optional) | **2** | **none** | **`SPEC` `PLAN`** |
|
|
77
77
|
| [GitHub Spec Kit](https://github.com/github/spec-kit) | 10 | — | `specify init` | `spec` `plan` `tasks` `checklist` `constitution` `research` `data-model` `contracts/` `quickstart` |
|
|
78
78
|
| [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
|
|
79
79
|
| [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
|
|
@@ -160,7 +160,7 @@ uvx tiny-spec install
|
|
|
160
160
|
Restart Claude Code so it picks up the new skills, then run the flow in your project:
|
|
161
161
|
|
|
162
162
|
```
|
|
163
|
-
/tiny-spec-scope # starting from an idea: interview it into
|
|
163
|
+
/tiny-spec-scope # starting from an idea: interview it into features (BREAKDOWN.md)
|
|
164
164
|
/tiny-spec-adopt # starting from a codebase: derive the constitution from real code
|
|
165
165
|
/tiny-spec-create # capture intent and requirements (binds a ticket, optional)
|
|
166
166
|
/tiny-spec-plan # design it, harden the constitution, slice the task list
|
|
@@ -182,10 +182,10 @@ forward. It writes nothing itself, it only delegates. By default it **stops befo
|
|
|
182
182
|
Ask it to build and it goes all the way instead:
|
|
183
183
|
|
|
184
184
|
```
|
|
185
|
-
/tiny-spec-run build the backlog # per
|
|
185
|
+
/tiny-spec-run build the backlog # per feature: branch → plan → build → merge → next
|
|
186
186
|
```
|
|
187
187
|
|
|
188
|
-
It reads your `BREAKDOWN.md` (or a list you paste) and works the
|
|
188
|
+
It reads your `BREAKDOWN.md` (or a list you paste) and works the features one after
|
|
189
189
|
another, merging each finished branch into `main` locally before starting the next —
|
|
190
190
|
until it's done or reaches a **terminal state** it names out loud. See
|
|
191
191
|
[Working a whole list](#working-a-whole-list). Which of the two it does is decided from
|
|
@@ -225,8 +225,8 @@ or install one set at a time.
|
|
|
225
225
|
### See a finished run first
|
|
226
226
|
|
|
227
227
|
[`examples/todo-cli/`](examples/todo-cli/) is a real run of the flow on one small
|
|
228
|
-
ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md
|
|
229
|
-
|
|
228
|
+
ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`
|
|
229
|
+
and `constitution.md` the suite wrote, and the code and tests it produced.
|
|
230
230
|
The tests pass; you can clone it and run the gate yourself.
|
|
231
231
|
|
|
232
232
|
## How it works
|
|
@@ -427,7 +427,7 @@ proved the same thing repeatedly and was the slowest part of the loop.
|
|
|
427
427
|
|
|
428
428
|
```mermaid
|
|
429
429
|
flowchart TB
|
|
430
|
-
SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[
|
|
430
|
+
SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[PLAN.md ## Tasks<br/>checklist]
|
|
431
431
|
|
|
432
432
|
TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
|
|
433
433
|
TASKS --> P[Plan task]
|
|
@@ -461,17 +461,39 @@ resumes from the checklist state.
|
|
|
461
461
|
### Working a whole list
|
|
462
462
|
|
|
463
463
|
Ask `/tiny-spec-run` to build — "build the backlog", "work through the breakdown",
|
|
464
|
-
"spec it out and build it" — and it takes a list of
|
|
465
|
-
|
|
464
|
+
"spec it out and build it" — and it takes a list of features and works them in batches.
|
|
465
|
+
Per feature it does the same four moves:
|
|
466
466
|
|
|
467
467
|
```
|
|
468
468
|
cut a branch from main → walk the chain → tiny-spec-build → merge back to main
|
|
469
469
|
```
|
|
470
470
|
|
|
471
|
-
|
|
472
|
-
|
|
471
|
+
Each branch is cut **fresh from main**, so a later feature sees the earlier ones already
|
|
472
|
+
merged — which is what makes an ordered list build correctly.
|
|
473
473
|
|
|
474
|
-
**
|
|
474
|
+
**Independent features build at the same time.** A feature can declare what it must follow
|
|
475
|
+
with a `needs:` line in `BREAKDOWN.md`; everything with no unmet `needs:` forms a batch
|
|
476
|
+
and builds **concurrently, one git worktree per feature**, three at a time by default.
|
|
477
|
+
The batch merges, then the next one starts.
|
|
478
|
+
|
|
479
|
+
```
|
|
480
|
+
## Feature: expose both helpers on a CLI slug: cli
|
|
481
|
+
|
|
482
|
+
- AC: `textkit slugify "Hi There"` prints "hi-there"
|
|
483
|
+
- needs: slugify, wordwrap
|
|
484
|
+
```
|
|
485
|
+
|
|
486
|
+
Omit `needs:` when a feature stands alone — that's the common case, and the field is meant
|
|
487
|
+
to be rare. A `needs:` you didn't need costs you parallelism forever; one you missed
|
|
488
|
+
costs a single merge conflict, which the run already catches and halts on. You can also
|
|
489
|
+
just name the set yourself at invocation ("build these three at once"), which overrides
|
|
490
|
+
the graph. A cycle, or a `needs:` naming a feature that isn't there, stops the run rather
|
|
491
|
+
than being guessed past.
|
|
492
|
+
|
|
493
|
+
**Tasks *inside* a feature never run in parallel.** They share files and each one assumes
|
|
494
|
+
the last landed, so they stay strictly sequential. Parallelism is across features only.
|
|
495
|
+
|
|
496
|
+
**The list is `BREAKDOWN.md` by default** — its `## Feature:` entries, in file order,
|
|
475
497
|
each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
|
|
476
498
|
a list at invocation instead and that wins; but a bare feature name has no acceptance
|
|
477
499
|
criteria, so `tiny-spec-create` will interview you when it reaches it. That's the
|
|
@@ -481,22 +503,26 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
|
|
|
481
503
|
|
|
482
504
|
| | |
|
|
483
505
|
|---|---|
|
|
484
|
-
| `done` | every
|
|
506
|
+
| `done` | every feature built **and merged** |
|
|
485
507
|
| `blocked` | an upstream document is wrong — go fix the spec or the plan |
|
|
486
508
|
| `exhausted` | a task stayed red past two fix attempts |
|
|
487
509
|
| `paused` | it reached a `pause:` point |
|
|
488
510
|
| `fork` | a real either/or the plan doesn't answer |
|
|
489
|
-
| `conflict` | a
|
|
511
|
+
| `conflict` | a feature's branch wouldn't merge cleanly |
|
|
490
512
|
|
|
491
|
-
**Only `done` means the work is built** — and in a
|
|
492
|
-
Stopping at
|
|
513
|
+
**Only `done` means the work is built** — and in a feature run, that means *all* of them.
|
|
514
|
+
Stopping at feature 2 of 7 and reporting "done" is what autonomous loops get wrong most
|
|
493
515
|
often, so the state is always named alongside what merged and what's still untouched.
|
|
494
516
|
|
|
495
|
-
**A halt stops the
|
|
496
|
-
|
|
497
|
-
|
|
517
|
+
**A halt stops the lane it happened in, and ends the run after that batch.** Its siblings
|
|
518
|
+
were declared independent, so they finish and merge — killing working lanes because one
|
|
519
|
+
failed throws away good work. But the run does not start the next batch: later features
|
|
520
|
+
usually assume the earlier ones landed, so skipping ahead past a failure just produces a
|
|
521
|
+
second, more confusing failure downstream. With more than one lane you get each feature's
|
|
522
|
+
own state, and the run's state is the worst of them — four green lanes and one `blocked`
|
|
523
|
+
is a `blocked` run.
|
|
498
524
|
|
|
499
|
-
**Pause points are technical, not per-
|
|
525
|
+
**Pause points are technical, not per-feature.** Any task can carry a `pause:` line, and
|
|
500
526
|
the build halts *before* running it:
|
|
501
527
|
|
|
502
528
|
```
|
|
@@ -508,20 +534,23 @@ the build halts *before* running it:
|
|
|
508
534
|
`tiny-spec-plan` proposes these for genuinely irreversible work — migrations,
|
|
509
535
|
destructive file operations, a new dependency, an auth boundary, a public API contract.
|
|
510
536
|
You can also give the run a standing policy up front ("halt before anything that touches
|
|
511
|
-
auth") and it gets applied as each
|
|
512
|
-
|
|
513
|
-
**What it will not do to your repo.** It runs exactly
|
|
514
|
-
`switch -c`, `merge --no-ff`, `merge --abort`,
|
|
515
|
-
|
|
516
|
-
|
|
517
|
-
|
|
518
|
-
the
|
|
519
|
-
|
|
520
|
-
|
|
521
|
-
|
|
537
|
+
auth") and it gets applied as each feature's tasks are sliced.
|
|
538
|
+
|
|
539
|
+
**What it will not do to your repo.** It runs exactly seven git commands — `switch`,
|
|
540
|
+
`switch -c`, `merge --no-ff`, `merge --abort`, `worktree add`, `worktree list`, and
|
|
541
|
+
reads. It refuses to start on a dirty tree, or if a worktree path it needs already
|
|
542
|
+
exists. It **never pushes**, never rebases, never resets, never deletes a branch, never
|
|
543
|
+
removes a worktree, and never opens a PR. Merges are local, so a bad run is one
|
|
544
|
+
`git reset` away; publishing stays yours. If the gate goes red after a merge it leaves
|
|
545
|
+
the merge alone and tells you the undo command rather than running it — and it hands
|
|
546
|
+
back the `git worktree remove` commands for the lanes instead of running those either,
|
|
547
|
+
since a halted lane's worktree is the tree you need to look at.
|
|
548
|
+
|
|
549
|
+
**Walk away and come back.** Progress isn't written down, it's derived: a feature whose
|
|
550
|
+
ticked `## Tasks` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
|
|
522
551
|
progress, no directory means not started. Ask again tomorrow in a fresh session and it
|
|
523
552
|
picks up where it stopped. No run-state file, no lock, no budget to configure — the
|
|
524
|
-
|
|
553
|
+
feature list *is* the budget.
|
|
525
554
|
|
|
526
555
|
**It never fixes a blocker for you.** A blocker means one of your documents is wrong,
|
|
527
556
|
and a run allowed to rewrite the requirement its own task just failed would be grading
|
|
@@ -546,7 +575,7 @@ It is namespaced per ticket, with a shared spine at the root:
|
|
|
546
575
|
constitution.md project-wide, shared across tickets
|
|
547
576
|
memory.md operational lessons, shared across tickets
|
|
548
577
|
<ticket-id>/ one directory per ticket (PROJ-123/, gh-42/, …)
|
|
549
|
-
SPEC.md PLAN.md
|
|
578
|
+
SPEC.md PLAN.md decisions.md (PLAN.md ends in the ## Tasks checklist)
|
|
550
579
|
```
|
|
551
580
|
|
|
552
581
|
`BREAKDOWN.md` and `design/` sit at your project root rather than inside `.spec/`,
|
|
@@ -15,6 +15,12 @@ pleasantries.
|
|
|
15
15
|
|
|
16
16
|
Everything you need and nothing you don't:
|
|
17
17
|
|
|
18
|
+
- **the working directory** to operate in. Every path you read, write, or run a command
|
|
19
|
+
against resolves against it. It may be a **git worktree** rather than the main checkout
|
|
20
|
+
— a build can run several features at once, each in its own worktree. **Never read or
|
|
21
|
+
write outside the directory you were given**, and never `cd` to a sibling worktree to
|
|
22
|
+
"check something": another feature is being built there right now, and what you find will
|
|
23
|
+
be wrong by the time you act on it;
|
|
18
24
|
- the **task id**, **description**, and **acceptance** (the outcome that proves it done);
|
|
19
25
|
- a **`files:` hint** — likely paths to touch (guidance, not a hard boundary);
|
|
20
26
|
- the full **constitution** (`constitution.md`): Style, Engineering standards,
|
|
@@ -13,6 +13,11 @@ back to `tiny-spec-build`; return data, not pleasantries.
|
|
|
13
13
|
|
|
14
14
|
## What you receive (the context contract)
|
|
15
15
|
|
|
16
|
+
- **the working directory** to operate in — every path and every gate command resolves
|
|
17
|
+
against it. It may be a **git worktree** rather than the main checkout, since a build
|
|
18
|
+
can run several features at once. **Never read, write, or run a gate outside the
|
|
19
|
+
directory you were given**: a sibling worktree holds a different feature mid-build, and
|
|
20
|
+
measuring it would make your verdict meaningless;
|
|
16
21
|
- the **task id**, **description**, and **acceptance** (the outcome that must hold);
|
|
17
22
|
- the full **constitution** (`constitution.md`) — especially **Guiding invariants**,
|
|
18
23
|
**Definition of Done**, and **Verification commands**, plus the **Design system**
|
|
@@ -54,6 +59,18 @@ acceptance and the constitution — verified, not inferred?**
|
|
|
54
59
|
task on evidence you found insufficient is not** — that is the one thing this whole
|
|
55
60
|
role exists to prevent. Never narrow a `full` gate to `scoped`.
|
|
56
61
|
|
|
62
|
+
**A Verification command that cannot run in this environment is a `FAIL` and a
|
|
63
|
+
blocker — never a `PASS`.** If a documented command is red for a reason the task's
|
|
64
|
+
code cannot fix (the interpreter is too old, a tool isn't installed, the command was
|
|
65
|
+
never runnable as written), say exactly that in `GATE:`, return `FAIL`, and name the
|
|
66
|
+
**constitution** as the upstream document to fix. Proving the task green *somewhere
|
|
67
|
+
else* — a fresh venv, a container, an install you fixed by hand — is **not** a pass:
|
|
68
|
+
you verified a different environment than the one the gate names, and the next task
|
|
69
|
+
will hit the identical red. Diagnosing the cause and demonstrating the task's own code
|
|
70
|
+
is sound is genuinely useful; put it in `FINDINGS` so the fix is cheap. It does not
|
|
71
|
+
change the verdict. A constitution whose gate cannot run is the single most dangerous
|
|
72
|
+
file in the project, because every later task inherits the same false signal.
|
|
73
|
+
|
|
57
74
|
Say which scope you actually ran in `GATE:`, including that you escalated and why.
|
|
58
75
|
3. **Exercise the acceptance.** Trigger the task's stated outcome the most
|
|
59
76
|
black-box way available (CLI > HTTP > public API) with realistic input,
|
|
@@ -78,25 +95,23 @@ acceptance and the constitution — verified, not inferred?**
|
|
|
78
95
|
- **A selector that matches nothing is a `FAIL`, never a skip.** Either the code
|
|
79
96
|
didn't build the element or it named it something else — both are real, and both
|
|
80
97
|
are invisible if you quietly move on. Report the selector and that it was absent.
|
|
81
|
-
- Compare each element's numbers to the tokens its row names
|
|
82
|
-
deltas ("heading is 28px, `type.heading.lg` is 24px"; "padding 19px is not on
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
- **Check `layout:`** — the arrangement, max width, and the **order** it names
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
command printed, `Read` the `D<n>`'s `export:` image, and grade **every state you
|
|
99
|
-
have a screenshot for**, not just the happy path, on four lines:
|
|
98
|
+
- Compare each element's numbers to the tokens its row names, reporting concrete
|
|
99
|
+
deltas ("heading is 28px, `type.heading.lg` is 24px"; "padding 19px is not on the
|
|
100
|
+
`space.*` scale"). **Do not pixel-diff the screenshot** — font antialiasing makes
|
|
101
|
+
image comparison flaky enough that the check gets ignored, which is how visual
|
|
102
|
+
gates die.
|
|
103
|
+
- **Check `layout:`** — the arrangement, max width, and the **order** it names, from
|
|
104
|
+
the bounding rectangles. Every token can be correct on an element in the wrong place.
|
|
105
|
+
- **Exercise every state the entry names** — empty, loading, error, success. Drive the
|
|
106
|
+
UI into each one and observe what changes; **finding the word in the source is not
|
|
107
|
+
evidence**, and it false-passes routinely (a comment saying the loading state is
|
|
108
|
+
missing contains "loading"). A surface that renders its happy path and nothing else
|
|
109
|
+
is a fail, not a nit.
|
|
110
|
+
- Finally, **judge the render against the export.** Everything above proves the numbers
|
|
111
|
+
are right; none of it can see an element that is present, on-token, and invisible —
|
|
112
|
+
so now look. `Read` each `SCREENSHOT <state> <path>` the `visual:` command printed
|
|
113
|
+
and the `D<n>`'s `export:` image, and grade **every state you have a screenshot
|
|
114
|
+
for**, not just the happy path, on four lines:
|
|
100
115
|
|
|
101
116
|
1. **Presence** — is every `elements:` row actually *visible* in the render? At
|
|
102
117
|
`opacity: 0`, zero height, clipped out of view, hidden behind a sibling, or the
|
|
@@ -104,24 +119,21 @@ acceptance and the constitution — verified, not inferred?**
|
|
|
104
119
|
not there. **FAIL** — this is why the step exists.
|
|
105
120
|
2. **Legibility & occlusion** — text clipped, truncated mid-word, overlapping
|
|
106
121
|
another element, or on a background it can't be read against. **FAIL.**
|
|
107
|
-
3. **Correspondence** —
|
|
108
|
-
|
|
109
|
-
|
|
122
|
+
3. **Correspondence** — the same regions as the export, in the reading order
|
|
123
|
+
`layout:` names? A whole region missing is a **FAIL**; a stylistic difference is
|
|
124
|
+
a **flag**.
|
|
110
125
|
4. **Hierarchy & polish** — emphasis, balance, crowding, alignment. **Always a
|
|
111
126
|
flag**, never a fail.
|
|
112
127
|
|
|
113
|
-
Three rules bound it
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
- **Cite what you saw.** Name the state whose screenshot the finding came from and
|
|
123
|
-
what was in it ("state `error`: caption present in DOM but renders at opacity 0").
|
|
124
|
-
An uncited visual claim reads as an opinion and gets ignored.
|
|
128
|
+
Three rules bound it. **The numbers beat your eye on anything they already
|
|
129
|
+
measured** — on-scale padding that looks cramped, or an exact token that looks
|
|
130
|
+
washed out, is a `flag:`; you may fail only on what measurement *cannot* see, since
|
|
131
|
+
contradicting your own numbers sends the executor a task it cannot fix and the loop
|
|
132
|
+
is bounded at two attempts. **This is still not a pixel diff** — the export is
|
|
133
|
+
usually a wireframe, so judge structure and legibility, never visual identity. And
|
|
134
|
+
**cite what you saw**, naming the state the finding came from ("state `error`:
|
|
135
|
+
caption present in DOM but renders at opacity 0") — an uncited visual claim reads as
|
|
136
|
+
an opinion and gets ignored.
|
|
125
137
|
|
|
126
138
|
**If the command printed no `SCREENSHOT` line**, do not run this sub-step and do
|
|
127
139
|
not eyeball a substitute. Grade on steps 1–3 above, write `judge: not run — visual:
|
|
@@ -87,10 +87,10 @@ Each case runs in a throwaway sandbox seeded with `IDEA.md` and the vendored ski
|
|
|
87
87
|
### What gets measured
|
|
88
88
|
|
|
89
89
|
- **structural conformance** (deterministic) — BREAKDOWN has Problem and
|
|
90
|
-
Goal & non-goals filled, a Decisions block, ≥1 Feature,
|
|
90
|
+
Goal & non-goals filled, a Decisions block, ≥1 Feature, each with a slug and ≥1 AC; the
|
|
91
91
|
planning skills left no `.spec/` behind.
|
|
92
92
|
- **hand-off integrity** (LLM judge) — **coverage** (every PRD capability lands in ≥1
|
|
93
|
-
|
|
93
|
+
feature, nothing dropped) and **no fabrication** (every feature traces to a capability,
|
|
94
94
|
nothing invented). These are the high-value signals; a case PASSes only if structural
|
|
95
95
|
conformance holds *and* the judge confirms both.
|
|
96
96
|
- **quality** (LLM judge, reported not gated) — atomicity / user-observable phrasing,
|
|
@@ -10,6 +10,9 @@ workflow generates before you run it on your own work.
|
|
|
10
10
|
> retouched to match the current artifact formats — a doctored example would defeat
|
|
11
11
|
> the point. Later versions add a `## Design` section to `SPEC.md` (unused here: this
|
|
12
12
|
> is a CLI with no design surface) and richer status flags. The shape is the same.
|
|
13
|
+
> One format move was applied after the fact: in 2.0 the checklist stopped being its
|
|
14
|
+
> own `tasks.md` and became `PLAN.md`'s `## Tasks` section, so the original `tasks.md`
|
|
15
|
+
> was moved there verbatim — same tasks, same ticks, no content edited.
|
|
13
16
|
|
|
14
17
|
## What's here
|
|
15
18
|
|
|
@@ -20,14 +23,13 @@ TICKET.md the input — a small "todo CLI" ask
|
|
|
20
23
|
definition of done, and the verification gate
|
|
21
24
|
todo-cli/
|
|
22
25
|
SPEC.md intent + REQ-1..REQ-9 (what "done" means)
|
|
23
|
-
PLAN.md the design
|
|
24
|
-
|
|
26
|
+
PLAN.md the design, how each requirement is covered, and
|
|
27
|
+
the ordered ## Tasks checklist, all ticked [x]
|
|
25
28
|
todo.py the produced CLI (stdlib only, 117 lines)
|
|
26
29
|
test_todo.py the produced end-to-end test (129 lines)
|
|
27
30
|
```
|
|
28
31
|
|
|
29
|
-
Read them in flow order: `TICKET.md` → `SPEC.md` → `PLAN.md` → `
|
|
30
|
-
`todo.py`. The `constitution.md` is the persistent context injected into every task.
|
|
32
|
+
Read them in flow order: `TICKET.md` → `SPEC.md` → `PLAN.md` → `todo.py`. The `constitution.md` is the persistent context injected into every task.
|
|
31
33
|
|
|
32
34
|
## It passes its own gate
|
|
33
35
|
|