tiny-spec 1.0.0__tar.gz → 1.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/PKG-INFO +81 -101
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/README.md +80 -100
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/agents/tiny-spec-build-executor.md +9 -2
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/agents/tiny-spec-build-reviewer.md +26 -6
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/pyproject.toml +1 -1
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/tiny-spec-build/SKILL.md +38 -6
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/tiny-spec-plan/SKILL.md +65 -19
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/tiny_spec/__init__.py +1 -1
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/.gitignore +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/LICENSE +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/docs/eval/README.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/examples/todo-cli/README.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/tiny-spec-adopt/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/tiny-spec-create/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/tiny-spec-design/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/tiny-spec-run/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/tiny-spec-scope/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/tiny_spec/cli.py +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.1.0}/tiny_spec/manifest.json +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: tiny-spec
|
|
3
|
-
Version: 1.
|
|
3
|
+
Version: 1.1.0
|
|
4
4
|
Summary: A tiny, opinionated take on spec-driven development.
|
|
5
5
|
Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
|
|
6
6
|
Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
|
|
@@ -115,54 +115,36 @@ receipt rather than the adjective:
|
|
|
115
115
|
| [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
|
|
116
116
|
| [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
|
|
117
117
|
|
|
118
|
-
The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents.
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
the skill count went down: `tiny-spec-adopt` is genuinely new capability, and merging
|
|
122
|
-
four skills into two removed commands to learn rather than words to read.)
|
|
123
|
-
|
|
124
|
-
<sub>Counts taken from each project's own repository on 2026-08-03 by listing its
|
|
125
|
-
command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
|
|
126
|
-
`templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
|
|
127
|
-
OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
|
|
128
|
-
personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
|
|
129
|
-
`customize.toml` files (35). These projects are all substantially more popular than
|
|
130
|
-
tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
|
|
131
|
-
bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
|
|
118
|
+
The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents. No
|
|
119
|
+
config file, no build step, no orchestrator, no CLI to initialize a project — the skills
|
|
120
|
+
are the product.
|
|
132
121
|
|
|
133
|
-
|
|
122
|
+
<sub>Counted on 2026-08-03 from each project's own repository, not its marketing copy —
|
|
123
|
+
Spec Kit's `templates/commands/`, OpenSpec's `skills/`, BMAD's `SKILL.md` and
|
|
124
|
+
`customize.toml` files on `main`. All three are far more popular than tiny-spec, and
|
|
125
|
+
being smaller is a bet, not a proof of quality — re-run the counts yourself before
|
|
126
|
+
believing them.</sub>
|
|
134
127
|
|
|
135
|
-
|
|
136
|
-
many phases, many agents, many generated documents. tiny-spec makes the opposite
|
|
137
|
-
bet. Keep one safeguard, drop the rest.
|
|
128
|
+
## Why it's small
|
|
138
129
|
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
it earns its place.
|
|
130
|
+
Most spec frameworks are generous by default: many phases, many agents, many generated
|
|
131
|
+
documents. tiny-spec makes the opposite bet — keep one safeguard, drop the rest. That
|
|
132
|
+
safeguard is the independent reviewer, which exercises acceptance criteria end to end,
|
|
133
|
+
because a green unit test suite is not the same as working software.
|
|
144
134
|
|
|
145
135
|
The case for staying small:
|
|
146
136
|
|
|
147
|
-
- **Documents are context, and context isn't free.**
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
- **
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
impose ceremony on work that doesn't need it. tiny-spec's extra structure is
|
|
159
|
-
optional by design — add shape where it pays, skip it where it doesn't.
|
|
160
|
-
- **More moving parts is more to maintain.** Orchestrators, ownership contracts,
|
|
161
|
-
checkpoint matrices, and config files are themselves a system you have to learn
|
|
162
|
-
and keep in sync. A few small skills and two agents are not.
|
|
163
|
-
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks
|
|
164
|
-
like progress, but it isn't proof. The proof is the reviewer running your real
|
|
165
|
-
tests before each commit.
|
|
137
|
+
- **Documents are context, and context isn't free.** Every paragraph the agent has to
|
|
138
|
+
hold is room it no longer has for your actual code. tiny-spec keeps the spine small —
|
|
139
|
+
a constitution and a short memory — and injects only what each task needs.
|
|
140
|
+
- **Real work is a ticket inside a system, not a greenfield repo.** tiny-spec binds to a
|
|
141
|
+
ticket and references your task platform instead of re-describing the world.
|
|
142
|
+
- **Rigid pipelines fight the user.** Mandatory phases impose ceremony on work that
|
|
143
|
+
doesn't need it. The extra structure here is optional — add shape where it pays.
|
|
144
|
+
- **More moving parts is more to maintain.** Orchestrators, ownership contracts, and
|
|
145
|
+
config files are themselves a system you have to keep in sync. Two agents are not.
|
|
146
|
+
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks like
|
|
147
|
+
progress, but the proof is the reviewer running your real tests before each commit.
|
|
166
148
|
|
|
167
149
|
That's the whole trade: where larger kits add machinery, tiny-spec adds one
|
|
168
150
|
independent reviewer and stops.
|
|
@@ -172,10 +154,10 @@ independent reviewer and stops.
|
|
|
172
154
|
Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
|
|
173
155
|
runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
|
|
174
156
|
code with held-out tests the suite never sees, and checks the reviewer's own verdict
|
|
175
|
-
against that ground truth
|
|
176
|
-
|
|
157
|
+
against that ground truth. The number that matters is **false-PASS rate** — how often
|
|
158
|
+
the gate blesses code that is actually broken.
|
|
177
159
|
|
|
178
|
-
Most recent run — 5 benchmark tasks
|
|
160
|
+
Most recent run — 5 benchmark tasks on v1.0.0:
|
|
179
161
|
|
|
180
162
|
| | |
|
|
181
163
|
|---|---|
|
|
@@ -184,29 +166,24 @@ Most recent run — 5 benchmark tasks, v1.0.0 (`b18e32a`), 2026-09-07:
|
|
|
184
166
|
| **false-PASS rate** | **0%** — it never called broken code done |
|
|
185
167
|
| blocker rate | 0% |
|
|
186
168
|
|
|
187
|
-
Two earlier
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
1.0 rebuilt the planning chain around it — `tiny-spec-plan` now writes the task list
|
|
191
|
-
that `tiny-spec-build` executes — and the gate came through unchanged.
|
|
169
|
+
Two earlier versions scored identically, which is the more interesting result: the loop
|
|
170
|
+
is stable across releases rather than tuned to one commit — including through 1.0, which
|
|
171
|
+
rebuilt the planning chain underneath it.
|
|
192
172
|
|
|
193
173
|
See [`docs/eval/README.md`](docs/eval/README.md) for the method and
|
|
194
|
-
[`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
The harness is in the repo precisely so you don't have to take my word for it.
|
|
174
|
+
[`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the rubric. **Read
|
|
175
|
+
both skeptically: it is my benchmark, of my own tool, scored by me**, on small
|
|
176
|
+
self-contained tasks — so it measures the loop on well-specified work, not
|
|
177
|
+
large-codebase performance. The harness is in the repo so you don't have to take my
|
|
178
|
+
word for it.
|
|
200
179
|
|
|
201
180
|
## New to spec-driven development?
|
|
202
181
|
|
|
203
|
-
Spec-driven development (SDD) means writing down *what* you want and *why* before
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
before a single line is written, and the result is checked against the spec
|
|
209
|
-
rather than vibes. tiny-spec is one small take on that idea.
|
|
182
|
+
Spec-driven development (SDD) means writing down *what* you want and *why* before any
|
|
183
|
+
code exists, then letting that spec drive the build. Instead of prompting an agent and
|
|
184
|
+
hoping, you hand it a small, explicit contract — the intent, a design, and an ordered
|
|
185
|
+
list of tasks. The payoff: the agent stops guessing, you review the plan before a line
|
|
186
|
+
is written, and the result is checked against the spec rather than vibes.
|
|
210
187
|
|
|
211
188
|
## Quickstart
|
|
212
189
|
|
|
@@ -297,13 +274,12 @@ hardens it with concrete engineering rules, and `tiny-spec-build` injects it who
|
|
|
297
274
|
every task. It holds your style, standards, invariants, definition of done, and
|
|
298
275
|
verification commands.
|
|
299
276
|
|
|
300
|
-
**On an existing codebase, deriving beats asking.** Your test command,
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
into theatre.
|
|
277
|
+
**On an existing codebase, deriving beats asking.** Your test command, layout, and
|
|
278
|
+
conventions are already written down — in `package.json`, in CI, in your linter config,
|
|
279
|
+
in the shape of the tree. `tiny-spec-adopt` reads them, marks each section as declared
|
|
280
|
+
or inferred so you know which parts to distrust, then *runs* the verification commands
|
|
281
|
+
it derived and reports which went green. A gate that has never been run turns every
|
|
282
|
+
future review into theatre.
|
|
307
283
|
|
|
308
284
|
Because it is project-wide the constitution can also go missing — deleted, or never
|
|
309
285
|
committed — while your specs survive. Re-running `tiny-spec-adopt` (or
|
|
@@ -404,13 +380,12 @@ FINDINGS:
|
|
|
404
380
|
|
|
405
381
|
Numbers first, and never a screenshot diff — pixel comparison goes flaky on font
|
|
406
382
|
antialiasing and teams end up muting it. But no measurement catches an element that is
|
|
407
|
-
present, on-token, and still not on screen
|
|
383
|
+
present, on-token, and still not on screen: occluded, clipped, truncated, or the same
|
|
408
384
|
color as what's behind it. So the last step reads a screenshot of each state next to
|
|
409
385
|
your export and grades presence, legibility, and correspondence. **Where a number
|
|
410
|
-
already settled the question the eye may only flag
|
|
411
|
-
|
|
412
|
-
|
|
413
|
-
task with no `design:` tag is graded exactly as before.
|
|
386
|
+
already settled the question the eye may only flag**, which keeps the two halves from
|
|
387
|
+
contradicting each other — measurable or visible violations fail, taste comes back as
|
|
388
|
+
`flag:` notes. A task with no `design:` tag is graded exactly as before.
|
|
414
389
|
|
|
415
390
|
<details>
|
|
416
391
|
<summary>The <code>visual:</code> command (you write this once)</summary>
|
|
@@ -454,14 +429,12 @@ await b.close();
|
|
|
454
429
|
```
|
|
455
430
|
|
|
456
431
|
The missing-selector branch **must print something** — that is what turns a renamed
|
|
457
|
-
element into a failure instead of a silent skip. `top`
|
|
458
|
-
|
|
459
|
-
|
|
460
|
-
|
|
461
|
-
|
|
462
|
-
|
|
463
|
-
rather than quietly skipping it. Without a `visual:` command at all, a task tagged
|
|
464
|
-
`design:` raises a blocker rather than passing quietly.
|
|
432
|
+
element into a failure instead of a silent skip. `top` feeds the layout-order check, and
|
|
433
|
+
`opacity`/`visibility` turn the cheapest kind of invisible element into a numeric
|
|
434
|
+
failure. The `SCREENSHOT` line arms the judge; drop it and the numeric half still gates
|
|
435
|
+
exactly as before, with the reviewer reporting `judge: not run` rather than quietly
|
|
436
|
+
skipping it. Without a `visual:` command at all, a task tagged `design:` raises a
|
|
437
|
+
blocker rather than passing quietly.
|
|
465
438
|
|
|
466
439
|
</details>
|
|
467
440
|
|
|
@@ -476,10 +449,19 @@ unnoticed.
|
|
|
476
449
|
1. Plan the task against the constitution (inline, brief).
|
|
477
450
|
2. Implement it with a fresh `tiny-spec-build-executor` agent.
|
|
478
451
|
3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
|
|
479
|
-
|
|
452
|
+
and grades against the constitution and the task's acceptance.
|
|
480
453
|
4. On pass, commit the code plus a checklist tick. On fail, loop back to the
|
|
481
454
|
executor with the findings. After two failed attempts the run halts `exhausted`.
|
|
482
455
|
|
|
456
|
+
Every task's acceptance is exercised end to end, black-box, by the reviewer — that part
|
|
457
|
+
never varies. What does vary is how much of the toolchain gets re-proven around it: a
|
|
458
|
+
clean `install → build` runs on the first task, the last one, anything touching
|
|
459
|
+
dependency or build config, and anything carrying `design:`; in between the reviewer runs
|
|
460
|
+
lint and tests against the existing build. It escalates back to the full gate whenever the
|
|
461
|
+
narrower one can't carry the verdict, and the authoritative clean run happens once more
|
|
462
|
+
over the whole project as the final smoke. Re-installing from scratch on every task
|
|
463
|
+
proved the same thing repeatedly and was the slowest part of the loop.
|
|
464
|
+
|
|
483
465
|
```mermaid
|
|
484
466
|
flowchart TB
|
|
485
467
|
SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
|
|
@@ -487,7 +469,7 @@ flowchart TB
|
|
|
487
469
|
TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
|
|
488
470
|
TASKS --> P[Plan task]
|
|
489
471
|
P --> I[Implement<br/>executor]
|
|
490
|
-
I --> R[Review + run gate<br/>reviewer]
|
|
472
|
+
I --> R[Review + run gate<br/>full or scoped<br/>reviewer]
|
|
491
473
|
I -->|blocker| B
|
|
492
474
|
R -->|pass| C[Commit + tick]
|
|
493
475
|
C --> TASKS
|
|
@@ -543,10 +525,9 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
|
|
|
543
525
|
| `fork` | a real either/or the plan doesn't answer |
|
|
544
526
|
| `conflict` | a story's branch wouldn't merge cleanly |
|
|
545
527
|
|
|
546
|
-
**Only `done` means the work is built** — and in a story run that means *all* of them.
|
|
547
|
-
|
|
548
|
-
|
|
549
|
-
many stories are still untouched.
|
|
528
|
+
**Only `done` means the work is built** — and in a story run, that means *all* of them.
|
|
529
|
+
Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
|
|
530
|
+
often, so the state is always named alongside what merged and what's still untouched.
|
|
550
531
|
|
|
551
532
|
**A halt stops the whole run, not just that story.** Later stories in a list you wrote
|
|
552
533
|
top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
|
|
@@ -573,27 +554,26 @@ never opens a PR. Merges are local, so a bad run is one `git reset` away; publis
|
|
|
573
554
|
stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
|
|
574
555
|
the undo command rather than running it.
|
|
575
556
|
|
|
576
|
-
**Walk away and come back.**
|
|
577
|
-
|
|
578
|
-
|
|
579
|
-
|
|
580
|
-
|
|
557
|
+
**Walk away and come back.** Progress isn't written down, it's derived: a story whose
|
|
558
|
+
ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
|
|
559
|
+
progress, no directory means not started. Ask again tomorrow in a fresh session and it
|
|
560
|
+
picks up where it stopped. No run-state file, no lock, no budget to configure — the
|
|
561
|
+
story list *is* the budget.
|
|
581
562
|
|
|
582
563
|
**It never fixes a blocker for you.** A blocker means one of your documents is wrong,
|
|
583
564
|
and a run allowed to rewrite the requirement its own task just failed would be grading
|
|
584
565
|
its own homework. It reports and stops; the upstream edit is yours.
|
|
585
566
|
|
|
586
567
|
**And it never decides to build on its own.** Whether a run stops before the build or
|
|
587
|
-
goes all the way is fixed from your opening request and never revised
|
|
588
|
-
|
|
589
|
-
yourself into it is the one thing a router doesn't get to do.
|
|
568
|
+
goes all the way is fixed from your opening request and never revised. Build is your
|
|
569
|
+
review gate, so promoting itself into it is the one thing a router doesn't get to do.
|
|
590
570
|
|
|
591
571
|
## Project layout
|
|
592
572
|
|
|
593
573
|
Each skill is one self-contained `SKILL.md`, with every document skeleton inline in
|
|
594
|
-
it — no companion template files, no absolute paths,
|
|
595
|
-
|
|
596
|
-
|
|
574
|
+
it — no companion template files, no absolute paths, no shared parent required at
|
|
575
|
+
runtime. A skill works wherever you drop it, and a run never stops to ask permission to
|
|
576
|
+
read a template out of your Claude config directory.
|
|
597
577
|
|
|
598
578
|
tiny-spec creates a `.spec/` directory in your project root, never inside a skill.
|
|
599
579
|
It is namespaced per ticket, with a shared spine at the root:
|
|
@@ -78,54 +78,36 @@ receipt rather than the adjective:
|
|
|
78
78
|
| [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
|
|
79
79
|
| [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
|
|
80
80
|
|
|
81
|
-
The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents.
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
the skill count went down: `tiny-spec-adopt` is genuinely new capability, and merging
|
|
85
|
-
four skills into two removed commands to learn rather than words to read.)
|
|
86
|
-
|
|
87
|
-
<sub>Counts taken from each project's own repository on 2026-08-03 by listing its
|
|
88
|
-
command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
|
|
89
|
-
`templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
|
|
90
|
-
OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
|
|
91
|
-
personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
|
|
92
|
-
`customize.toml` files (35). These projects are all substantially more popular than
|
|
93
|
-
tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
|
|
94
|
-
bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
|
|
81
|
+
The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents. No
|
|
82
|
+
config file, no build step, no orchestrator, no CLI to initialize a project — the skills
|
|
83
|
+
are the product.
|
|
95
84
|
|
|
96
|
-
|
|
85
|
+
<sub>Counted on 2026-08-03 from each project's own repository, not its marketing copy —
|
|
86
|
+
Spec Kit's `templates/commands/`, OpenSpec's `skills/`, BMAD's `SKILL.md` and
|
|
87
|
+
`customize.toml` files on `main`. All three are far more popular than tiny-spec, and
|
|
88
|
+
being smaller is a bet, not a proof of quality — re-run the counts yourself before
|
|
89
|
+
believing them.</sub>
|
|
97
90
|
|
|
98
|
-
|
|
99
|
-
many phases, many agents, many generated documents. tiny-spec makes the opposite
|
|
100
|
-
bet. Keep one safeguard, drop the rest.
|
|
91
|
+
## Why it's small
|
|
101
92
|
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
it earns its place.
|
|
93
|
+
Most spec frameworks are generous by default: many phases, many agents, many generated
|
|
94
|
+
documents. tiny-spec makes the opposite bet — keep one safeguard, drop the rest. That
|
|
95
|
+
safeguard is the independent reviewer, which exercises acceptance criteria end to end,
|
|
96
|
+
because a green unit test suite is not the same as working software.
|
|
107
97
|
|
|
108
98
|
The case for staying small:
|
|
109
99
|
|
|
110
|
-
- **Documents are context, and context isn't free.**
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
- **
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
impose ceremony on work that doesn't need it. tiny-spec's extra structure is
|
|
122
|
-
optional by design — add shape where it pays, skip it where it doesn't.
|
|
123
|
-
- **More moving parts is more to maintain.** Orchestrators, ownership contracts,
|
|
124
|
-
checkpoint matrices, and config files are themselves a system you have to learn
|
|
125
|
-
and keep in sync. A few small skills and two agents are not.
|
|
126
|
-
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks
|
|
127
|
-
like progress, but it isn't proof. The proof is the reviewer running your real
|
|
128
|
-
tests before each commit.
|
|
100
|
+
- **Documents are context, and context isn't free.** Every paragraph the agent has to
|
|
101
|
+
hold is room it no longer has for your actual code. tiny-spec keeps the spine small —
|
|
102
|
+
a constitution and a short memory — and injects only what each task needs.
|
|
103
|
+
- **Real work is a ticket inside a system, not a greenfield repo.** tiny-spec binds to a
|
|
104
|
+
ticket and references your task platform instead of re-describing the world.
|
|
105
|
+
- **Rigid pipelines fight the user.** Mandatory phases impose ceremony on work that
|
|
106
|
+
doesn't need it. The extra structure here is optional — add shape where it pays.
|
|
107
|
+
- **More moving parts is more to maintain.** Orchestrators, ownership contracts, and
|
|
108
|
+
config files are themselves a system you have to keep in sync. Two agents are not.
|
|
109
|
+
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks like
|
|
110
|
+
progress, but the proof is the reviewer running your real tests before each commit.
|
|
129
111
|
|
|
130
112
|
That's the whole trade: where larger kits add machinery, tiny-spec adds one
|
|
131
113
|
independent reviewer and stops.
|
|
@@ -135,10 +117,10 @@ independent reviewer and stops.
|
|
|
135
117
|
Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
|
|
136
118
|
runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
|
|
137
119
|
code with held-out tests the suite never sees, and checks the reviewer's own verdict
|
|
138
|
-
against that ground truth
|
|
139
|
-
|
|
120
|
+
against that ground truth. The number that matters is **false-PASS rate** — how often
|
|
121
|
+
the gate blesses code that is actually broken.
|
|
140
122
|
|
|
141
|
-
Most recent run — 5 benchmark tasks
|
|
123
|
+
Most recent run — 5 benchmark tasks on v1.0.0:
|
|
142
124
|
|
|
143
125
|
| | |
|
|
144
126
|
|---|---|
|
|
@@ -147,29 +129,24 @@ Most recent run — 5 benchmark tasks, v1.0.0 (`b18e32a`), 2026-09-07:
|
|
|
147
129
|
| **false-PASS rate** | **0%** — it never called broken code done |
|
|
148
130
|
| blocker rate | 0% |
|
|
149
131
|
|
|
150
|
-
Two earlier
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
1.0 rebuilt the planning chain around it — `tiny-spec-plan` now writes the task list
|
|
154
|
-
that `tiny-spec-build` executes — and the gate came through unchanged.
|
|
132
|
+
Two earlier versions scored identically, which is the more interesting result: the loop
|
|
133
|
+
is stable across releases rather than tuned to one commit — including through 1.0, which
|
|
134
|
+
rebuilt the planning chain underneath it.
|
|
155
135
|
|
|
156
136
|
See [`docs/eval/README.md`](docs/eval/README.md) for the method and
|
|
157
|
-
[`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
The harness is in the repo precisely so you don't have to take my word for it.
|
|
137
|
+
[`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the rubric. **Read
|
|
138
|
+
both skeptically: it is my benchmark, of my own tool, scored by me**, on small
|
|
139
|
+
self-contained tasks — so it measures the loop on well-specified work, not
|
|
140
|
+
large-codebase performance. The harness is in the repo so you don't have to take my
|
|
141
|
+
word for it.
|
|
163
142
|
|
|
164
143
|
## New to spec-driven development?
|
|
165
144
|
|
|
166
|
-
Spec-driven development (SDD) means writing down *what* you want and *why* before
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
before a single line is written, and the result is checked against the spec
|
|
172
|
-
rather than vibes. tiny-spec is one small take on that idea.
|
|
145
|
+
Spec-driven development (SDD) means writing down *what* you want and *why* before any
|
|
146
|
+
code exists, then letting that spec drive the build. Instead of prompting an agent and
|
|
147
|
+
hoping, you hand it a small, explicit contract — the intent, a design, and an ordered
|
|
148
|
+
list of tasks. The payoff: the agent stops guessing, you review the plan before a line
|
|
149
|
+
is written, and the result is checked against the spec rather than vibes.
|
|
173
150
|
|
|
174
151
|
## Quickstart
|
|
175
152
|
|
|
@@ -260,13 +237,12 @@ hardens it with concrete engineering rules, and `tiny-spec-build` injects it who
|
|
|
260
237
|
every task. It holds your style, standards, invariants, definition of done, and
|
|
261
238
|
verification commands.
|
|
262
239
|
|
|
263
|
-
**On an existing codebase, deriving beats asking.** Your test command,
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
into theatre.
|
|
240
|
+
**On an existing codebase, deriving beats asking.** Your test command, layout, and
|
|
241
|
+
conventions are already written down — in `package.json`, in CI, in your linter config,
|
|
242
|
+
in the shape of the tree. `tiny-spec-adopt` reads them, marks each section as declared
|
|
243
|
+
or inferred so you know which parts to distrust, then *runs* the verification commands
|
|
244
|
+
it derived and reports which went green. A gate that has never been run turns every
|
|
245
|
+
future review into theatre.
|
|
270
246
|
|
|
271
247
|
Because it is project-wide the constitution can also go missing — deleted, or never
|
|
272
248
|
committed — while your specs survive. Re-running `tiny-spec-adopt` (or
|
|
@@ -367,13 +343,12 @@ FINDINGS:
|
|
|
367
343
|
|
|
368
344
|
Numbers first, and never a screenshot diff — pixel comparison goes flaky on font
|
|
369
345
|
antialiasing and teams end up muting it. But no measurement catches an element that is
|
|
370
|
-
present, on-token, and still not on screen
|
|
346
|
+
present, on-token, and still not on screen: occluded, clipped, truncated, or the same
|
|
371
347
|
color as what's behind it. So the last step reads a screenshot of each state next to
|
|
372
348
|
your export and grades presence, legibility, and correspondence. **Where a number
|
|
373
|
-
already settled the question the eye may only flag
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
task with no `design:` tag is graded exactly as before.
|
|
349
|
+
already settled the question the eye may only flag**, which keeps the two halves from
|
|
350
|
+
contradicting each other — measurable or visible violations fail, taste comes back as
|
|
351
|
+
`flag:` notes. A task with no `design:` tag is graded exactly as before.
|
|
377
352
|
|
|
378
353
|
<details>
|
|
379
354
|
<summary>The <code>visual:</code> command (you write this once)</summary>
|
|
@@ -417,14 +392,12 @@ await b.close();
|
|
|
417
392
|
```
|
|
418
393
|
|
|
419
394
|
The missing-selector branch **must print something** — that is what turns a renamed
|
|
420
|
-
element into a failure instead of a silent skip. `top`
|
|
421
|
-
|
|
422
|
-
|
|
423
|
-
|
|
424
|
-
|
|
425
|
-
|
|
426
|
-
rather than quietly skipping it. Without a `visual:` command at all, a task tagged
|
|
427
|
-
`design:` raises a blocker rather than passing quietly.
|
|
395
|
+
element into a failure instead of a silent skip. `top` feeds the layout-order check, and
|
|
396
|
+
`opacity`/`visibility` turn the cheapest kind of invisible element into a numeric
|
|
397
|
+
failure. The `SCREENSHOT` line arms the judge; drop it and the numeric half still gates
|
|
398
|
+
exactly as before, with the reviewer reporting `judge: not run` rather than quietly
|
|
399
|
+
skipping it. Without a `visual:` command at all, a task tagged `design:` raises a
|
|
400
|
+
blocker rather than passing quietly.
|
|
428
401
|
|
|
429
402
|
</details>
|
|
430
403
|
|
|
@@ -439,10 +412,19 @@ unnoticed.
|
|
|
439
412
|
1. Plan the task against the constitution (inline, brief).
|
|
440
413
|
2. Implement it with a fresh `tiny-spec-build-executor` agent.
|
|
441
414
|
3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
|
|
442
|
-
|
|
415
|
+
and grades against the constitution and the task's acceptance.
|
|
443
416
|
4. On pass, commit the code plus a checklist tick. On fail, loop back to the
|
|
444
417
|
executor with the findings. After two failed attempts the run halts `exhausted`.
|
|
445
418
|
|
|
419
|
+
Every task's acceptance is exercised end to end, black-box, by the reviewer — that part
|
|
420
|
+
never varies. What does vary is how much of the toolchain gets re-proven around it: a
|
|
421
|
+
clean `install → build` runs on the first task, the last one, anything touching
|
|
422
|
+
dependency or build config, and anything carrying `design:`; in between the reviewer runs
|
|
423
|
+
lint and tests against the existing build. It escalates back to the full gate whenever the
|
|
424
|
+
narrower one can't carry the verdict, and the authoritative clean run happens once more
|
|
425
|
+
over the whole project as the final smoke. Re-installing from scratch on every task
|
|
426
|
+
proved the same thing repeatedly and was the slowest part of the loop.
|
|
427
|
+
|
|
446
428
|
```mermaid
|
|
447
429
|
flowchart TB
|
|
448
430
|
SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
|
|
@@ -450,7 +432,7 @@ flowchart TB
|
|
|
450
432
|
TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
|
|
451
433
|
TASKS --> P[Plan task]
|
|
452
434
|
P --> I[Implement<br/>executor]
|
|
453
|
-
I --> R[Review + run gate<br/>reviewer]
|
|
435
|
+
I --> R[Review + run gate<br/>full or scoped<br/>reviewer]
|
|
454
436
|
I -->|blocker| B
|
|
455
437
|
R -->|pass| C[Commit + tick]
|
|
456
438
|
C --> TASKS
|
|
@@ -506,10 +488,9 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
|
|
|
506
488
|
| `fork` | a real either/or the plan doesn't answer |
|
|
507
489
|
| `conflict` | a story's branch wouldn't merge cleanly |
|
|
508
490
|
|
|
509
|
-
**Only `done` means the work is built** — and in a story run that means *all* of them.
|
|
510
|
-
|
|
511
|
-
|
|
512
|
-
many stories are still untouched.
|
|
491
|
+
**Only `done` means the work is built** — and in a story run, that means *all* of them.
|
|
492
|
+
Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
|
|
493
|
+
often, so the state is always named alongside what merged and what's still untouched.
|
|
513
494
|
|
|
514
495
|
**A halt stops the whole run, not just that story.** Later stories in a list you wrote
|
|
515
496
|
top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
|
|
@@ -536,27 +517,26 @@ never opens a PR. Merges are local, so a bad run is one `git reset` away; publis
|
|
|
536
517
|
stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
|
|
537
518
|
the undo command rather than running it.
|
|
538
519
|
|
|
539
|
-
**Walk away and come back.**
|
|
540
|
-
|
|
541
|
-
|
|
542
|
-
|
|
543
|
-
|
|
520
|
+
**Walk away and come back.** Progress isn't written down, it's derived: a story whose
|
|
521
|
+
ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
|
|
522
|
+
progress, no directory means not started. Ask again tomorrow in a fresh session and it
|
|
523
|
+
picks up where it stopped. No run-state file, no lock, no budget to configure — the
|
|
524
|
+
story list *is* the budget.
|
|
544
525
|
|
|
545
526
|
**It never fixes a blocker for you.** A blocker means one of your documents is wrong,
|
|
546
527
|
and a run allowed to rewrite the requirement its own task just failed would be grading
|
|
547
528
|
its own homework. It reports and stops; the upstream edit is yours.
|
|
548
529
|
|
|
549
530
|
**And it never decides to build on its own.** Whether a run stops before the build or
|
|
550
|
-
goes all the way is fixed from your opening request and never revised
|
|
551
|
-
|
|
552
|
-
yourself into it is the one thing a router doesn't get to do.
|
|
531
|
+
goes all the way is fixed from your opening request and never revised. Build is your
|
|
532
|
+
review gate, so promoting itself into it is the one thing a router doesn't get to do.
|
|
553
533
|
|
|
554
534
|
## Project layout
|
|
555
535
|
|
|
556
536
|
Each skill is one self-contained `SKILL.md`, with every document skeleton inline in
|
|
557
|
-
it — no companion template files, no absolute paths,
|
|
558
|
-
|
|
559
|
-
|
|
537
|
+
it — no companion template files, no absolute paths, no shared parent required at
|
|
538
|
+
runtime. A skill works wherever you drop it, and a run never stops to ask permission to
|
|
539
|
+
read a template out of your Claude config directory.
|
|
560
540
|
|
|
561
541
|
tiny-spec creates a `.spec/` directory in your project root, never inside a skill.
|
|
562
542
|
It is namespaced per ticket, with a shared spine at the root:
|
|
@@ -24,7 +24,10 @@ Everything you need and nothing you don't:
|
|
|
24
24
|
so you don't re-learn a pitfall a past run already paid for;
|
|
25
25
|
- **if the task carries `design:`** — that screen's `D<n>` entry from `SPEC.md` and
|
|
26
26
|
the path to its committed export;
|
|
27
|
-
- the specific existing files that are your starting point, named explicitly
|
|
27
|
+
- the specific existing files that are your starting point, named explicitly;
|
|
28
|
+
- sometimes, **the files the previous task left behind** — paths only. Start there when
|
|
29
|
+
orienting: it is the most recently moved ground in the codebase and the most likely
|
|
30
|
+
place your task connects to.
|
|
28
31
|
|
|
29
32
|
You are **blind to the workflow, not to the codebase.** You don't get the plan,
|
|
30
33
|
other tasks, or shared state. But the named files are a launch point, not the whole
|
|
@@ -64,7 +67,11 @@ violated.
|
|
|
64
67
|
6. You MAY run a **narrow self-check** of your own work (the one test file you
|
|
65
68
|
wrote, a syntax/import check). You do **not** need to run the full gate — the
|
|
66
69
|
independent **reviewer** runs the authoritative Verification commands next.
|
|
67
|
-
Leave the tree in a clean, buildable state for it
|
|
70
|
+
**Leave the tree in a clean, buildable, installed state for it** — the reviewer may
|
|
71
|
+
run against the existing build rather than rebuilding from scratch, so a tree you
|
|
72
|
+
left half-migrated or with a stale artifact costs the task a review cycle. If you
|
|
73
|
+
changed dependencies or build configuration, apply them (install, regenerate the
|
|
74
|
+
lockfile) rather than leaving it for someone else, and say so in `CHANGES`.
|
|
68
75
|
|
|
69
76
|
## Hard constraints
|
|
70
77
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: tiny-spec-build-reviewer
|
|
3
|
-
description: Independently reviews a single finished task — runs the project's real gate end-to-end and checks the code against the constitution
|
|
3
|
+
description: Independently reviews a single finished task — runs the project's real gate at the scope it was given (full or scoped, escalating when scoped can't carry the verdict), exercises the acceptance end-to-end, and checks the code against the constitution. Blind to how the code was written. Returns PASS/FAIL plus findings. Spawned (one per task) by tiny-spec-build. Does not fix code, plan, spawn agents, or invoke skills.
|
|
4
4
|
tools: Read, Write, Edit, Bash, Grep, Glob
|
|
5
5
|
---
|
|
6
6
|
|
|
@@ -17,6 +17,9 @@ back to `tiny-spec-build`; return data, not pleasantries.
|
|
|
17
17
|
- the full **constitution** (`constitution.md`) — especially **Guiding invariants**,
|
|
18
18
|
**Definition of Done**, and **Verification commands**, plus the **Design system**
|
|
19
19
|
token table if the project has one;
|
|
20
|
+
- the **gate scope** for this task — `full` or `scoped`. `tiny-spec-build` sets it; it
|
|
21
|
+
knows where this task sits in the run and you don't. If the brief names no scope,
|
|
22
|
+
treat it as `full`;
|
|
20
23
|
- the project's **memory** if any (`memory.md`) — operational lessons (e.g. the
|
|
21
24
|
gate needs the package installed first); honor them so you don't false-fail on a
|
|
22
25
|
known precondition;
|
|
@@ -32,10 +35,26 @@ acceptance and the constitution — verified, not inferred?**
|
|
|
32
35
|
1. **Read the changed code.** Check it against the constitution: does it honor the
|
|
33
36
|
**Guiding invariants**, match the **Style** and **Layout**, meet the
|
|
34
37
|
**Definition of Done**? Note any violation as a finding.
|
|
35
|
-
2. **Run the real gate.**
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
38
|
+
2. **Run the real gate, at the scope you were given.** Capture the real output either
|
|
39
|
+
way — never a test-runner shortcut, never an inferred result.
|
|
40
|
+
|
|
41
|
+
- **`full`** — execute the constitution's **Verification commands** end-to-end
|
|
42
|
+
(install → lint → test → build → run, as applicable) from a clean state, after the
|
|
43
|
+
documented setup.
|
|
44
|
+
- **`scoped`** — run `lint:` and `test:` against the working tree, and rely on the
|
|
45
|
+
existing install/build state rather than rebuilding it from clean. Step 3 below is
|
|
46
|
+
**not** scoped: you still exercise the acceptance end-to-end, black-box, with real
|
|
47
|
+
input. A scoped gate skips re-proving the toolchain; it never skips proving the task.
|
|
48
|
+
|
|
49
|
+
**Escalate rather than guess.** Move a `scoped` gate up to `full` the moment the
|
|
50
|
+
narrower one can't carry the verdict — you need a build artifact that isn't there or
|
|
51
|
+
looks stale, `install`/`build` state seems inconsistent with the changed files, a
|
|
52
|
+
command fails in a way that might be environmental, or you simply cannot exercise the
|
|
53
|
+
acceptance without building. Escalation is always allowed and always safe. **Passing a
|
|
54
|
+
task on evidence you found insufficient is not** — that is the one thing this whole
|
|
55
|
+
role exists to prevent. Never narrow a `full` gate to `scoped`.
|
|
56
|
+
|
|
57
|
+
Say which scope you actually ran in `GATE:`, including that you escalated and why.
|
|
39
58
|
3. **Exercise the acceptance.** Trigger the task's stated outcome the most
|
|
40
59
|
black-box way available (CLI > HTTP > public API) with realistic input,
|
|
41
60
|
including a negative case if the acceptance implies a boundary or rejection.
|
|
@@ -158,7 +177,8 @@ Never spawn subagents or invoke skills.
|
|
|
158
177
|
```
|
|
159
178
|
TASK: <task id>
|
|
160
179
|
VERDICT: PASS | FAIL
|
|
161
|
-
GATE: <
|
|
180
|
+
GATE: <scope: full | scoped (+ "escalated from scoped: <why>" if you moved it up);
|
|
181
|
+
the Verification commands you ran + the real result (pass/fail + key output)>
|
|
162
182
|
ACCEPTANCE: <how you exercised it + the observed effect, or why you couldn't>
|
|
163
183
|
DESIGN: <omit unless the task carried `design:`. The D<n> checked, the measurements
|
|
164
184
|
you read back vs the tokens they should match, and which states you exercised —
|
|
@@ -68,7 +68,11 @@ Spawn one **`tiny-spec-build-executor`** with a fresh, self-contained prompt:
|
|
|
68
68
|
`.spec/<active>/SPEC.md` plus its `export:` path, so the executor can look at the
|
|
69
69
|
design instead of guessing at it;
|
|
70
70
|
- only the specific existing files the task starts from, named explicitly (so it
|
|
71
|
-
edits with the real current contents, not blind)
|
|
71
|
+
edits with the real current contents, not blind);
|
|
72
|
+
- the **previous passed task's `CHANGES` list**, if there is one, labelled as what the
|
|
73
|
+
last task left behind. Every executor starts cold and re-derives the codebase; naming
|
|
74
|
+
the files that just moved is the cheapest way to cut that, and you already have the
|
|
75
|
+
list in hand. Paths only — not the previous task's description, findings, or report.
|
|
72
76
|
|
|
73
77
|
Do **not** pass the plan, sibling tasks, or other chatter. It returns a structured
|
|
74
78
|
report (`STATUS`, `CHANGES`, `DECISIONS`, `BLOCKER`). A `STATUS: blocked` →
|
|
@@ -80,11 +84,29 @@ Spawn one **`tiny-spec-build-reviewer`**, **blind to step 2**, with:
|
|
|
80
84
|
- the task id, description, and **acceptance**;
|
|
81
85
|
- the **whole** `.spec/constitution.md`;
|
|
82
86
|
- the list of changed files (from the executor's `CHANGES`) to read;
|
|
83
|
-
- the **Verification commands** from the constitution to run
|
|
87
|
+
- the **Verification commands** from the constitution to run, and the **gate scope**
|
|
88
|
+
for this task — `full` or `scoped` (decide it as below; you have the run context,
|
|
89
|
+
the reviewer does not);
|
|
84
90
|
- **if the task has a `design:` field** — the same `D<n>` entry and `export:` path
|
|
85
91
|
you gave the executor, so it grades against the contract rather than its taste.
|
|
86
92
|
|
|
87
|
-
|
|
93
|
+
**Gate scope.** A clean `install → build` proves the same thing on task 7 that it proved
|
|
94
|
+
on task 1 unless a task changed what it installs or builds, so running it every time buys
|
|
95
|
+
little and costs the most wall-clock in the loop. Name the scope in the brief:
|
|
96
|
+
|
|
97
|
+
- **`full`** — the whole gate from a clean state (install → lint → test → build → run).
|
|
98
|
+
Use it on: the **first** task of this run; any task whose `CHANGES` touch dependency or
|
|
99
|
+
build configuration (`package.json`/lockfiles, `pyproject.toml`, `go.mod`, `Cargo.toml`,
|
|
100
|
+
`Dockerfile`, CI config, build scripts); any task carrying `design:`; and the **last**
|
|
101
|
+
task in `tasks.md`.
|
|
102
|
+
- **`scoped`** — otherwise. `lint:` + `test:` plus the acceptance exercised end-to-end,
|
|
103
|
+
black-box, exactly as on a full gate. The acceptance exercise is never scoped away; it
|
|
104
|
+
is the whole reason the reviewer exists.
|
|
105
|
+
|
|
106
|
+
The reviewer may **escalate** `scoped` to `full` on its own and will say so — that
|
|
107
|
+
direction is always safe. It may never go the other way.
|
|
108
|
+
|
|
109
|
+
It runs the gate at the scope you named, checks the code against the constitution's
|
|
88
110
|
**Definition of Done** and **invariants**, confirms the **acceptance** actually
|
|
89
111
|
holds (exercised, not inferred), and returns `VERDICT: PASS | FAIL` + findings.
|
|
90
112
|
On a `design:` task it also runs the constitution's `visual:` command, measures the
|
|
@@ -94,11 +116,15 @@ numbers can't reach, like an element that measures perfectly and shows up invisi
|
|
|
94
116
|
It fails on what it measured or saw, and merely flags what is taste.
|
|
95
117
|
|
|
96
118
|
> Why independent: unit-green ≠ working, and the author is the worst judge of its
|
|
97
|
-
> own blind spots. The reviewer
|
|
98
|
-
> safeguard that keeps scope and quality honest without an
|
|
119
|
+
> own blind spots. The reviewer exercising the acceptance black-box — on every task,
|
|
120
|
+
> at either scope — is the safeguard that keeps scope and quality honest without an
|
|
121
|
+
> ownership contract.
|
|
99
122
|
|
|
100
123
|
### 4. CONVERGE (on FAIL)
|
|
101
|
-
Re-dispatch the **executor** with the reviewer's findings appended to its brief
|
|
124
|
+
Re-dispatch the **executor** with the reviewer's findings appended to its brief, then
|
|
125
|
+
review again — recomputing the **gate scope** from the fix's `CHANGES`, and using `full`
|
|
126
|
+
if the failure was in the gate itself or the reviewer escalated last time. A task that
|
|
127
|
+
already failed once has earned the wider check.
|
|
102
128
|
Bound this to **2 fix attempts**. If it still fails after that, stop and **halt
|
|
103
129
|
`exhausted`** (see **Halting** below) — don't keep grinding or hand-fix past the loop
|
|
104
130
|
silently. `exhausted` is recorded as a blocker, because a task that three attempts
|
|
@@ -274,6 +300,12 @@ When every task in `tasks.md` is `[x]`:
|
|
|
274
300
|
setup — install/build, not a test-runner shortcut). There is no separate
|
|
275
301
|
verify skill — this final smoke confirms the requirements actually work end-to-end, not
|
|
276
302
|
just that tasks are ticked.
|
|
303
|
+
|
|
304
|
+
**This is the authoritative clean-state run for the whole build**, and it is never
|
|
305
|
+
scoped or skipped. Per-task gates run at the scope step 3 named; this one always runs
|
|
306
|
+
everything from clean. A build or install regression that a `scoped` task gate did not
|
|
307
|
+
re-prove surfaces here — which is why a red final smoke is a gap to fix or route
|
|
308
|
+
upstream, never a pass and never `done`.
|
|
277
309
|
2. **Report** — the terminal state by name (`done` here, since every task is `[x]`
|
|
278
310
|
and the smoke passed), what was built, the commits made (with the branch), and any
|
|
279
311
|
open `decisions.md` items (blockers, tasks unchecked by a reconcile). If the final
|
|
@@ -138,25 +138,59 @@ past it into a task list — say what's unresolved and route back to `tiny-spec-
|
|
|
138
138
|
## Step 3 — slice the approach into tasks
|
|
139
139
|
|
|
140
140
|
Walk the `## Approach` you just wrote and break it into tasks. No waves, no
|
|
141
|
-
parallelism, no `owns:` contracts — tasks run one at a time, top to bottom.
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
141
|
+
parallelism, no `owns:` contracts — tasks run one at a time, top to bottom.
|
|
142
|
+
|
|
143
|
+
**The unit is one coherent commit** — what a competent engineer does in one focused
|
|
144
|
+
sitting and commits as a single unit of work. That typically touches several files and
|
|
145
|
+
often satisfies several `REQ-N` at once. It is the size you would open as one reviewable
|
|
146
|
+
change, not the smallest thing you could name.
|
|
147
|
+
|
|
148
|
+
**Know what a task costs, so size has something to trade against.** Every task you write
|
|
149
|
+
spends two cold-start agents that must re-orient in the codebase from scratch, a reviewer
|
|
150
|
+
that runs the gate and exercises the acceptance end-to-end, and two commits. That
|
|
151
|
+
overhead is **fixed** — it does not shrink for a small task. A task the executor
|
|
152
|
+
finishes in thirty seconds still pays all of it. Size each task so the work inside it
|
|
153
|
+
clearly outweighs the machinery around it.
|
|
154
|
+
|
|
155
|
+
So:
|
|
156
|
+
|
|
157
|
+
- **Split on independent failure, not on sentence length.** Split a task when it carries
|
|
158
|
+
two **unrelated** observable outcomes that could fail independently of each other. Do
|
|
159
|
+
**not** split because the acceptance got long — a good acceptance is usually several
|
|
160
|
+
clauses covering the happy path *and* its negatives. See `T4` in the `examples/todo-cli`
|
|
161
|
+
task list: two CLI commands, four requirements, and three negative cases, in one task,
|
|
162
|
+
with one acceptance. That is the calibration point, not the exception.
|
|
163
|
+
- **One task may satisfy several `REQ-N`** — `req:` takes a list. Coverage means every
|
|
164
|
+
requirement has a *home*, not that every requirement gets its *own* task. **A 1:1
|
|
165
|
+
REQ→task mapping is the single most common way this list comes out too granular.**
|
|
166
|
+
Group the requirements that one coherent change delivers together.
|
|
146
167
|
- **Ordered so each builds on the last.** Tasks run sequentially, so a later task may
|
|
147
168
|
freely assume an earlier task's code already exists. Put foundational work (types,
|
|
148
169
|
schema, scaffolding) first. Order by dependency, not by guesswork.
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
170
|
+
|
|
171
|
+
**Smells that mean you sliced below the commit line** — fold each of these back into the
|
|
172
|
+
task it belongs to:
|
|
173
|
+
|
|
174
|
+
- a task that only defines types, interfaces, or schema with no behavior behind them;
|
|
175
|
+
- a task that only adds tests for the task before it (the constitution's **Definition of
|
|
176
|
+
Done** already requires the tests to ship with the code);
|
|
177
|
+
- one task per file, or one task per function;
|
|
178
|
+
- a "wire it up" / "integrate the pieces" task trailing the pieces it wires.
|
|
179
|
+
|
|
180
|
+
**Count is a smell, not a cap.** A story sized the way `tiny-spec-scope` describes
|
|
181
|
+
usually lands in **3–6 tasks**. If you are past about eight, re-read the list: you have
|
|
182
|
+
either sliced below the commit line, or the story itself was too big and should have been
|
|
183
|
+
split upstream. Check the list against that; do **not** enforce a number, and never drop
|
|
184
|
+
or merge coverage just to hit one.
|
|
152
185
|
|
|
153
186
|
For each task, write:
|
|
154
187
|
|
|
155
188
|
```
|
|
156
189
|
- [ ] T<n> — <imperative description>
|
|
157
|
-
- acceptance: <one user-observable outcome that proves it's done
|
|
190
|
+
- acceptance: <one user-observable outcome that proves it's done — happy path and the
|
|
191
|
+
negatives that bound it, in one entry>
|
|
158
192
|
- type: feat # optional; Conventional Commit type (defaults to feat)
|
|
159
|
-
- req: REQ-n
|
|
193
|
+
- req: REQ-n, REQ-n # optional; the REQ-N this task delivers — a list, not one
|
|
160
194
|
- design: D-n # optional; the SPEC.md D<n> screen this task builds — arms the visual gate
|
|
161
195
|
- pause: <why> # optional; halt the build before this task so a human looks first
|
|
162
196
|
- files: <comma-separated hint of files it will touch>
|
|
@@ -164,7 +198,11 @@ For each task, write:
|
|
|
164
198
|
|
|
165
199
|
The **acceptance** is what the reviewer checks against — make it observable
|
|
166
200
|
("`spec --version` prints the version and exits 0"), not internal ("version logic
|
|
167
|
-
added"). **
|
|
201
|
+
added"). It states **one outcome**, but one outcome is not one clause: spell out the
|
|
202
|
+
happy path and the negatives that bound it in the same acceptance, separated by
|
|
203
|
+
semicolons. A long acceptance is a well-specified task, not an oversized one — it is the
|
|
204
|
+
*number of unrelated things that could fail* that decides whether to split, not the
|
|
205
|
+
length of the line. **type** picks the Conventional Commit type `tiny-spec-build` uses for this
|
|
168
206
|
task's code commit (`feat | fix | docs | refactor | test | chore | build | ci | perf | style`);
|
|
169
207
|
set it when the task is clearly not a feature, otherwise omit and it defaults to `feat`.
|
|
170
208
|
**req** ties the task to the requirement it satisfies (traceability). The **files** line
|
|
@@ -211,7 +249,9 @@ anyway. A policy that matches nothing here is not an error — say so and move o
|
|
|
211
249
|
than stretching a task to fit it.
|
|
212
250
|
|
|
213
251
|
Cover **every** part of the approach — together the tasks must deliver all `REQ-N`.
|
|
214
|
-
Don't leave a requirement with no task.
|
|
252
|
+
Don't leave a requirement with no task. Coverage is about requirements having a home,
|
|
253
|
+
not about the shape of the mapping: several `REQ-N` on one task is the normal case, and
|
|
254
|
+
a task per requirement is the anti-pattern. Likewise, if `SPEC.md` has a `## Design`
|
|
215
255
|
section, every `D<n>` in it needs at least one task carrying that `design:` reference —
|
|
216
256
|
a screen nobody is graded against is a screen that will be built wrong.
|
|
217
257
|
|
|
@@ -236,25 +276,31 @@ updated: <ISO date>
|
|
|
236
276
|
|
|
237
277
|
## Tasks
|
|
238
278
|
|
|
239
|
-
- [ ] T1 — <one
|
|
240
|
-
- acceptance: <
|
|
279
|
+
- [ ] T1 — <one coherent commit's worth of foundational work; usually several files>
|
|
280
|
+
- acceptance: <the observable outcome; happy path; and the negative case that bounds it>
|
|
241
281
|
- type: feat # optional; Conventional Commit type for this task's commit (defaults to feat)
|
|
242
|
-
- req: REQ-1
|
|
243
|
-
- files: <path, path>
|
|
282
|
+
- req: REQ-1, REQ-2 # optional; the REQ-N this task delivers — several is normal
|
|
283
|
+
- files: <path, path, path>
|
|
244
284
|
|
|
245
|
-
- [ ] T2 — <next
|
|
246
|
-
- acceptance: <observable outcome>
|
|
285
|
+
- [ ] T2 — <next coherent change; assume T1's code exists>
|
|
286
|
+
- acceptance: <observable outcome; plus what it rejects and how it fails>
|
|
247
287
|
- type: feat
|
|
248
|
-
- req: REQ-
|
|
288
|
+
- req: REQ-3, REQ-4, REQ-5
|
|
249
289
|
- design: D1 # optional; only on tasks that build the visible surface
|
|
250
290
|
- files: <path, path>
|
|
251
291
|
|
|
252
292
|
- [ ] T3 — <…>
|
|
253
293
|
- acceptance: <observable outcome>
|
|
294
|
+
- req: REQ-6
|
|
254
295
|
- pause: <optional; what to check before this runs — irreversible work only>
|
|
255
296
|
- files: <path, path>
|
|
256
297
|
```
|
|
257
298
|
|
|
299
|
+
The `req:` lists above are the shape to aim for, not filler: a handful of tasks each
|
|
300
|
+
carrying the requirements one coherent change delivers. A skeleton filled in as
|
|
301
|
+
`REQ-1`, `REQ-2`, `REQ-3` down a column of single-requirement tasks is the granularity
|
|
302
|
+
failure described in Step 3.
|
|
303
|
+
|
|
258
304
|
## Update mode (SPEC changed → PLAN and tasks are stale)
|
|
259
305
|
|
|
260
306
|
When `PLAN.md` or `tasks.md` is `status: stale`, reconcile **both in one pass** — they
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|