tiny-spec 0.4.0__tar.gz → 0.5.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/PKG-INFO +200 -37
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/README.md +198 -35
- tiny_spec-0.5.0/examples/todo-cli/README.md +69 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/pyproject.toml +3 -1
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/tiny-spec-build/SKILL.md +87 -11
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/tiny-spec-create/SKILL.md +15 -0
- tiny_spec-0.5.0/tiny-spec-loop/SKILL.md +188 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/tiny-spec-run/SKILL.md +19 -6
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/tiny-spec-tasks/SKILL.md +36 -5
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/tiny_spec/__init__.py +1 -1
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/tiny_spec/manifest.json +1 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/.gitignore +0 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/LICENSE +0 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/agents/tiny-spec-build-executor.md +0 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/agents/tiny-spec-build-reviewer.md +0 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/docs/eval/README.md +0 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/tiny-spec-breakdown/SKILL.md +0 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/tiny-spec-plan/SKILL.md +0 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/tiny-spec-prd/SKILL.md +0 -0
- {tiny_spec-0.4.0 → tiny_spec-0.5.0}/tiny_spec/cli.py +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
|
-
Metadata-Version: 2.
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
2
|
Name: tiny-spec
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.5.0
|
|
4
4
|
Summary: A tiny, opinionated take on spec-driven development.
|
|
5
5
|
Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
|
|
6
6
|
Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
|
|
@@ -44,10 +44,24 @@ Description-Content-Type: text/markdown
|
|
|
44
44
|
<p align="center">A tiny, opinionated take on spec-driven development.</p>
|
|
45
45
|
|
|
46
46
|
<p align="center">
|
|
47
|
+
<a href="https://pypi.org/project/tiny-spec/"><img src="https://img.shields.io/pypi/v/tiny-spec.svg?color=d97757" alt="PyPI version"></a>
|
|
48
|
+
<a href="https://pypi.org/project/tiny-spec/"><img src="https://img.shields.io/pypi/dm/tiny-spec.svg?color=d97757" alt="PyPI downloads"></a>
|
|
47
49
|
<a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue.svg" alt="License: MIT"></a>
|
|
48
50
|
<a href="https://docs.claude.com/en/docs/claude-code/overview"><img src="https://img.shields.io/badge/Claude%20Code-skills-d97757.svg" alt="Claude Code"></a>
|
|
51
|
+
<a href="https://github.com/GrayMa77er/tiny-spec/stargazers"><img src="https://img.shields.io/github/stars/GrayMa77er/tiny-spec?color=d97757" alt="Stars"></a>
|
|
49
52
|
</p>
|
|
50
53
|
|
|
54
|
+
<p align="center">
|
|
55
|
+
<code>uvx tiny-spec install</code>
|
|
56
|
+
</p>
|
|
57
|
+
|
|
58
|
+
<!-- DEMO GIF SLOT — images/demo.gif
|
|
59
|
+
Record the FAIL -> fix -> PASS loop, then uncomment:
|
|
60
|
+
<p align="center">
|
|
61
|
+
<img src="images/demo.gif" alt="tiny-spec build loop: the reviewer fails a task, the executor fixes it, the task commits" width="800">
|
|
62
|
+
</p>
|
|
63
|
+
-->
|
|
64
|
+
|
|
51
65
|
tiny-spec is a four-step workflow for Claude Code that turns a ticket into shipped,
|
|
52
66
|
reviewed code. You write the intent, it produces a design, a task list, and then
|
|
53
67
|
builds the work one task at a time. Every task is implemented by one agent and
|
|
@@ -57,7 +71,9 @@ committed.
|
|
|
57
71
|
That core is **four skills and two agents**. In front of it sit **two optional
|
|
58
72
|
planning on-ramps** — `tiny-spec-prd` (idea → PRD) and `tiny-spec-breakdown`
|
|
59
73
|
(PRD → stories) — for when you're starting from an idea rather than a ready ticket.
|
|
60
|
-
`tiny-spec-run` walks the three planning steps
|
|
74
|
+
Two optional routers sit over the top: `tiny-spec-run` walks the three planning steps
|
|
75
|
+
in one command, and `tiny-spec-loop` works a whole list of stories — branch, plan,
|
|
76
|
+
build, merge, next — until they're built or something stops it. No config file, no
|
|
61
77
|
build step.
|
|
62
78
|
|
|
63
79
|
```
|
|
@@ -69,6 +85,10 @@ PLANNING (optional on-ramps) EXECUTION (the core loop, one story at
|
|
|
69
85
|
idea → PRD PRD → stories intent design tasks per-task loop
|
|
70
86
|
PRD.md BREAKDOWN.md SPEC.md PLAN.md + tasks.md plan → implement → review → commit
|
|
71
87
|
constitution
|
|
88
|
+
▲ ▲
|
|
89
|
+
└──────────────── tiny-spec-loop ─────────────────┘
|
|
90
|
+
optional: per story — branch, run, build, merge,
|
|
91
|
+
then the next story. Halts on a wall.
|
|
72
92
|
```
|
|
73
93
|
|
|
74
94
|
The two on-ramps are **optional** and stack. Have nothing written down? Run
|
|
@@ -79,6 +99,97 @@ work? Skip both and start at `tiny-spec-create`. Both on-ramps write a regenerab
|
|
|
79
99
|
file at your project root (not under `.spec/`); `tiny-spec-create` then reads the
|
|
80
100
|
breakdown one story at a time.
|
|
81
101
|
|
|
102
|
+
## How small
|
|
103
|
+
|
|
104
|
+
Every other kit in this space is bigger. That is the whole pitch, so here is the
|
|
105
|
+
receipt rather than the adjective:
|
|
106
|
+
|
|
107
|
+
| | skills / commands | agents | config | artifacts per feature |
|
|
108
|
+
|---|---|---|---|---|
|
|
109
|
+
| **tiny-spec** | **8** (4 core + 4 optional) | **2** | **none** | **`SPEC` `PLAN` `tasks`** |
|
|
110
|
+
| [GitHub Spec Kit](https://github.com/github/spec-kit) | 10 | — | `specify init` | `spec` `plan` `tasks` `checklist` `constitution` `research` `data-model` `contracts/` `quickstart` |
|
|
111
|
+
| [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
|
|
112
|
+
| [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
|
|
113
|
+
|
|
114
|
+
The whole of tiny-spec is **2,079 lines** of markdown across 8 skills and 2 agents.
|
|
115
|
+
There is no config file, no build step, no orchestrator, and no CLI to initialize a
|
|
116
|
+
project — the skills are the product.
|
|
117
|
+
|
|
118
|
+
<sub>Counts taken from each project's own repository on 2026-08-03 by listing its
|
|
119
|
+
command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
|
|
120
|
+
`templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
|
|
121
|
+
OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
|
|
122
|
+
personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
|
|
123
|
+
`customize.toml` files (35). These projects are all substantially more popular than
|
|
124
|
+
tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
|
|
125
|
+
bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
|
|
126
|
+
|
|
127
|
+
## Why it's small
|
|
128
|
+
|
|
129
|
+
Most spec frameworks are generous by default:
|
|
130
|
+
many phases, many agents, many generated documents. tiny-spec makes the opposite
|
|
131
|
+
bet. Keep one safeguard, drop the rest.
|
|
132
|
+
|
|
133
|
+
A green unit test suite is not the same as working software, so the reviewer
|
|
134
|
+
exercises acceptance criteria end to end and a final smoke test confirms the whole
|
|
135
|
+
spec. That independent review is the safeguard — not the volume of planning
|
|
136
|
+
artifacts. One task, one commit, an external reviewer. Nothing gets added unless
|
|
137
|
+
it earns its place.
|
|
138
|
+
|
|
139
|
+
The case for staying small:
|
|
140
|
+
|
|
141
|
+
- **Documents are context, and context isn't free.** Generating large `spec.md`,
|
|
142
|
+
`plan.md`, `research.md`, and `data-model.md` files costs tokens to write, then
|
|
143
|
+
costs context to carry. Every paragraph the agent has to hold is room it no
|
|
144
|
+
longer has for your actual code. tiny-spec keeps the spine small — a
|
|
145
|
+
constitution and a short memory — and injects only what each task needs.
|
|
146
|
+
- **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
|
|
147
|
+
assume you're bootstrapping a project from a blank page. Day to day, you pick up
|
|
148
|
+
a ticket and change part of a system that already exists. tiny-spec binds to a
|
|
149
|
+
ticket, works one at a time, and references your task platform instead of
|
|
150
|
+
re-describing the world.
|
|
151
|
+
- **Rigid pipelines fight the user.** Mandatory phases and required sections
|
|
152
|
+
impose ceremony on work that doesn't need it. tiny-spec's extra structure is
|
|
153
|
+
optional by design — add shape where it pays, skip it where it doesn't.
|
|
154
|
+
- **More moving parts is more to maintain.** Orchestrators, ownership contracts,
|
|
155
|
+
checkpoint matrices, and config files are themselves a system you have to learn
|
|
156
|
+
and keep in sync. A few small skills and two agents are not.
|
|
157
|
+
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks
|
|
158
|
+
like progress, but it isn't proof. The proof is the reviewer running your real
|
|
159
|
+
tests before each commit.
|
|
160
|
+
|
|
161
|
+
That's the whole trade: where larger kits add machinery, tiny-spec adds one
|
|
162
|
+
independent reviewer and stops.
|
|
163
|
+
|
|
164
|
+
### Does the trade hold up?
|
|
165
|
+
|
|
166
|
+
Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
|
|
167
|
+
runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
|
|
168
|
+
code with held-out tests the suite never sees, and checks the reviewer's own verdict
|
|
169
|
+
against that ground truth — the number that matters being **false-PASS rate**, how
|
|
170
|
+
often the gate blesses code that is actually broken.
|
|
171
|
+
|
|
172
|
+
Most recent run — 5 benchmark tasks, v0.4.0 (`7810074`), 2026-08-03:
|
|
173
|
+
|
|
174
|
+
| | |
|
|
175
|
+
|---|---|
|
|
176
|
+
| held-out pass rate | **100%** (5/5) — the code works, judged by tests it never saw |
|
|
177
|
+
| suite/truth agreement | **100%** — the reviewer's verdict matched the held-out grader every time |
|
|
178
|
+
| **false-PASS rate** | **0%** — it never called broken code done |
|
|
179
|
+
| blocker rate | 0% |
|
|
180
|
+
|
|
181
|
+
The same run a month earlier on `f973c85` scored identically, which is the more
|
|
182
|
+
interesting result: the loop is at least stable across versions rather than tuned to a
|
|
183
|
+
single commit.
|
|
184
|
+
|
|
185
|
+
See [`docs/eval/README.md`](docs/eval/README.md) for the method and
|
|
186
|
+
[`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the scoring rubric
|
|
187
|
+
this repo grades itself against. **Read both skeptically: it is my benchmark, of my own
|
|
188
|
+
tool, scored by me.** The rubric says so itself — *"treat the tiny-spec column as
|
|
189
|
+
self-assessment to pressure-test, not gospel."* The tasks are small and self-contained,
|
|
190
|
+
so the harness measures the loop on well-specified work, not large-codebase performance.
|
|
191
|
+
The harness is in the repo precisely so you don't have to take my word for it.
|
|
192
|
+
|
|
82
193
|
## New to spec-driven development?
|
|
83
194
|
|
|
84
195
|
Spec-driven development (SDD) means writing down *what* you want and *why* before
|
|
@@ -121,6 +232,17 @@ Or collapse the three planning steps into one and go straight to building:
|
|
|
121
232
|
**stops before `tiny-spec-build`** — that's where you actually review the work — and
|
|
122
233
|
it writes nothing itself, it only delegates.
|
|
123
234
|
|
|
235
|
+
Or hand it a whole list of stories:
|
|
236
|
+
|
|
237
|
+
```
|
|
238
|
+
/tiny-spec-loop # per story: branch → plan → build → merge → next
|
|
239
|
+
```
|
|
240
|
+
|
|
241
|
+
`tiny-spec-loop` reads your `BREAKDOWN.md` (or a list you paste) and works the stories
|
|
242
|
+
one after another, merging each finished branch into `main` locally before starting the
|
|
243
|
+
next — until it's done or reaches a **terminal state** it names out loud. See
|
|
244
|
+
[Working a whole list](#working-a-whole-list--tiny-spec-loop).
|
|
245
|
+
|
|
124
246
|
**Building from a mockup?** There is no design flag to pass. Commit your exports to
|
|
125
247
|
`design/` before the run and `tiny-spec-create` — whether you invoke it directly or
|
|
126
248
|
reach it through `tiny-spec-run` — reads what's there and asks which screens this
|
|
@@ -141,7 +263,7 @@ git clone https://github.com/GrayMa77er/tiny-spec.git
|
|
|
141
263
|
cd tiny-spec
|
|
142
264
|
|
|
143
265
|
mkdir -p "$HOME/.claude/skills" "$HOME/.claude/agents"
|
|
144
|
-
for s in tiny-spec-prd tiny-spec-breakdown tiny-spec-run tiny-spec-create tiny-spec-plan tiny-spec-tasks tiny-spec-build; do
|
|
266
|
+
for s in tiny-spec-prd tiny-spec-breakdown tiny-spec-run tiny-spec-loop tiny-spec-create tiny-spec-plan tiny-spec-tasks tiny-spec-build; do
|
|
145
267
|
cp -R "$s" "$HOME/.claude/skills/$s"
|
|
146
268
|
done
|
|
147
269
|
cp agents/*.md "$HOME/.claude/agents/"
|
|
@@ -152,6 +274,13 @@ or install one set at a time.
|
|
|
152
274
|
|
|
153
275
|
</details>
|
|
154
276
|
|
|
277
|
+
### See a finished run first
|
|
278
|
+
|
|
279
|
+
[`examples/todo-cli/`](examples/todo-cli/) is a real run of the flow on one small
|
|
280
|
+
ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`,
|
|
281
|
+
`tasks.md` and `constitution.md` the suite wrote, and the code and tests it produced.
|
|
282
|
+
The tests pass; you can clone it and run the gate yourself.
|
|
283
|
+
|
|
155
284
|
## How it works
|
|
156
285
|
|
|
157
286
|
The constitution (`constitution.md`) is the spine. `tiny-spec-create` seeds it from a
|
|
@@ -332,19 +461,21 @@ unnoticed.
|
|
|
332
461
|
3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
|
|
333
462
|
end to end and grades against the constitution and the task's acceptance.
|
|
334
463
|
4. On pass, commit the code plus a checklist tick. On fail, loop back to the
|
|
335
|
-
executor with the findings. After two failed attempts
|
|
464
|
+
executor with the findings. After two failed attempts the run halts `exhausted`.
|
|
336
465
|
|
|
337
466
|
```mermaid
|
|
338
467
|
flowchart TB
|
|
339
468
|
SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
|
|
340
469
|
|
|
470
|
+
TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
|
|
341
471
|
TASKS --> P[Plan task]
|
|
342
472
|
P --> I[Implement<br/>executor]
|
|
343
473
|
I --> R[Review + run gate<br/>reviewer]
|
|
474
|
+
I -->|blocker| B
|
|
344
475
|
R -->|pass| C[Commit + tick]
|
|
345
476
|
C --> TASKS
|
|
346
477
|
R -->|fail| I
|
|
347
|
-
R -->|fail twice| B[
|
|
478
|
+
R -->|fail twice| B[Halt — blocked / exhausted<br/>logged to decisions.md]
|
|
348
479
|
|
|
349
480
|
CON([constitution.md]) -.-> P & I & R
|
|
350
481
|
MEM([memory.md]) -.-> I & R
|
|
@@ -365,42 +496,74 @@ and logs a blocker instead of hacking around it. You fix the gap upstream in
|
|
|
365
496
|
`tiny-spec-plan` or `tiny-spec-create`, then resume. Work runs one ticket at a time and
|
|
366
497
|
resumes from the checklist state.
|
|
367
498
|
|
|
368
|
-
|
|
499
|
+
### Working a whole list — `tiny-spec-loop`
|
|
369
500
|
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
bet. Keep one safeguard, drop the rest.
|
|
501
|
+
`/tiny-spec-loop` takes a list of stories and works them one after another. Per story
|
|
502
|
+
it does the same four moves:
|
|
373
503
|
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
|
|
377
|
-
artifacts. One task, one commit, an external reviewer. Nothing gets added unless
|
|
378
|
-
it earns its place.
|
|
504
|
+
```
|
|
505
|
+
cut a branch from main → tiny-spec-run → tiny-spec-build → merge back to main
|
|
506
|
+
```
|
|
379
507
|
|
|
380
|
-
|
|
508
|
+
Then the next story. Each branch is cut **fresh from main**, so story 3 sees stories 1
|
|
509
|
+
and 2 already merged — which is what makes an ordered list build correctly.
|
|
381
510
|
|
|
382
|
-
|
|
383
|
-
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
- **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
|
|
388
|
-
assume you're bootstrapping a project from a blank page. Day to day, you pick up
|
|
389
|
-
a ticket and change part of a system that already exists. tiny-spec binds to a
|
|
390
|
-
ticket, works one at a time, and references your task platform instead of
|
|
391
|
-
re-describing the world.
|
|
392
|
-
- **Rigid pipelines fight the user.** Mandatory phases and required sections
|
|
393
|
-
impose ceremony on work that doesn't need it. tiny-spec's extra structure is
|
|
394
|
-
optional by design — add shape where it pays, skip it where it doesn't.
|
|
395
|
-
- **More moving parts is more to maintain.** Orchestrators, ownership contracts,
|
|
396
|
-
checkpoint matrices, and config files are themselves a system you have to learn
|
|
397
|
-
and keep in sync. A few small skills and two agents are not.
|
|
398
|
-
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks
|
|
399
|
-
like progress, but it isn't proof. The proof is the reviewer running your real
|
|
400
|
-
tests before each commit.
|
|
511
|
+
**The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
|
|
512
|
+
each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
|
|
513
|
+
a list at invocation instead and that wins; but a bare feature name has no acceptance
|
|
514
|
+
criteria, so `tiny-spec-create` will interview you when it reaches it. That's the
|
|
515
|
+
honest trade: a breakdown runs unattended, a pasted list is supervised.
|
|
401
516
|
|
|
402
|
-
|
|
403
|
-
|
|
517
|
+
**Every run ends in exactly one of six states, and it says which:**
|
|
518
|
+
|
|
519
|
+
| | |
|
|
520
|
+
|---|---|
|
|
521
|
+
| `done` | every story built **and merged** |
|
|
522
|
+
| `blocked` | an upstream document is wrong — go fix the spec or the plan |
|
|
523
|
+
| `exhausted` | a task stayed red past two fix attempts |
|
|
524
|
+
| `paused` | it reached a `pause:` point |
|
|
525
|
+
| `fork` | a real either/or the plan doesn't answer |
|
|
526
|
+
| `conflict` | a story's branch wouldn't merge cleanly |
|
|
527
|
+
|
|
528
|
+
**Only `done` means the work is built** — and in a story loop that means *all* of them.
|
|
529
|
+
A run that stopped at story 2 of 7 reporting "done" is the single thing autonomous loops
|
|
530
|
+
get wrong most often, so the state is always named, along with what got merged and how
|
|
531
|
+
many stories are still untouched.
|
|
532
|
+
|
|
533
|
+
**A halt stops the whole run, not just that story.** Later stories in a list you wrote
|
|
534
|
+
top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
|
|
535
|
+
just produces a second, more confusing failure downstream.
|
|
536
|
+
|
|
537
|
+
**Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
|
|
538
|
+
the build halts *before* running it:
|
|
539
|
+
|
|
540
|
+
```
|
|
541
|
+
- [ ] T7 — run the schema migration against the restored snapshot
|
|
542
|
+
- acceptance: every row in orders has a non-null tenant_id
|
|
543
|
+
- pause: confirm the snapshot is current before this touches real rows
|
|
544
|
+
```
|
|
545
|
+
|
|
546
|
+
`tiny-spec-tasks` proposes these for genuinely irreversible work — migrations,
|
|
547
|
+
destructive file operations, a new dependency, an auth boundary, a public API contract.
|
|
548
|
+
You can also give the run a standing policy up front ("halt before anything that touches
|
|
549
|
+
auth") and it gets applied as each story's tasks are sliced.
|
|
550
|
+
|
|
551
|
+
**What it will not do to your repo.** It runs exactly five git commands — `switch`,
|
|
552
|
+
`switch -c`, `merge --no-ff`, `merge --abort`, and reads. It refuses to start on a dirty
|
|
553
|
+
tree. It **never pushes**, never rebases, never resets, never deletes a branch, and
|
|
554
|
+
never opens a PR. Merges are local, so a bad run is one `git reset` away; publishing
|
|
555
|
+
stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
|
|
556
|
+
the undo command rather than running it.
|
|
557
|
+
|
|
558
|
+
**Walk away and come back.** Nothing is written down to track progress — it's derived:
|
|
559
|
+
a story whose ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked
|
|
560
|
+
task is in progress, no directory means not started. So re-running `/tiny-spec-loop`
|
|
561
|
+
tomorrow in a fresh session picks up exactly where it stopped. No run-state file, no
|
|
562
|
+
lock, and no budget to configure: the story list *is* the budget.
|
|
563
|
+
|
|
564
|
+
**It never fixes a blocker for you.** A blocker means one of your documents is wrong,
|
|
565
|
+
and a loop allowed to rewrite the requirement its own task just failed would be grading
|
|
566
|
+
its own homework. It reports and stops; the upstream edit is yours.
|
|
404
567
|
|
|
405
568
|
## Project layout
|
|
406
569
|
|
|
@@ -7,10 +7,24 @@
|
|
|
7
7
|
<p align="center">A tiny, opinionated take on spec-driven development.</p>
|
|
8
8
|
|
|
9
9
|
<p align="center">
|
|
10
|
+
<a href="https://pypi.org/project/tiny-spec/"><img src="https://img.shields.io/pypi/v/tiny-spec.svg?color=d97757" alt="PyPI version"></a>
|
|
11
|
+
<a href="https://pypi.org/project/tiny-spec/"><img src="https://img.shields.io/pypi/dm/tiny-spec.svg?color=d97757" alt="PyPI downloads"></a>
|
|
10
12
|
<a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue.svg" alt="License: MIT"></a>
|
|
11
13
|
<a href="https://docs.claude.com/en/docs/claude-code/overview"><img src="https://img.shields.io/badge/Claude%20Code-skills-d97757.svg" alt="Claude Code"></a>
|
|
14
|
+
<a href="https://github.com/GrayMa77er/tiny-spec/stargazers"><img src="https://img.shields.io/github/stars/GrayMa77er/tiny-spec?color=d97757" alt="Stars"></a>
|
|
12
15
|
</p>
|
|
13
16
|
|
|
17
|
+
<p align="center">
|
|
18
|
+
<code>uvx tiny-spec install</code>
|
|
19
|
+
</p>
|
|
20
|
+
|
|
21
|
+
<!-- DEMO GIF SLOT — images/demo.gif
|
|
22
|
+
Record the FAIL -> fix -> PASS loop, then uncomment:
|
|
23
|
+
<p align="center">
|
|
24
|
+
<img src="images/demo.gif" alt="tiny-spec build loop: the reviewer fails a task, the executor fixes it, the task commits" width="800">
|
|
25
|
+
</p>
|
|
26
|
+
-->
|
|
27
|
+
|
|
14
28
|
tiny-spec is a four-step workflow for Claude Code that turns a ticket into shipped,
|
|
15
29
|
reviewed code. You write the intent, it produces a design, a task list, and then
|
|
16
30
|
builds the work one task at a time. Every task is implemented by one agent and
|
|
@@ -20,7 +34,9 @@ committed.
|
|
|
20
34
|
That core is **four skills and two agents**. In front of it sit **two optional
|
|
21
35
|
planning on-ramps** — `tiny-spec-prd` (idea → PRD) and `tiny-spec-breakdown`
|
|
22
36
|
(PRD → stories) — for when you're starting from an idea rather than a ready ticket.
|
|
23
|
-
`tiny-spec-run` walks the three planning steps
|
|
37
|
+
Two optional routers sit over the top: `tiny-spec-run` walks the three planning steps
|
|
38
|
+
in one command, and `tiny-spec-loop` works a whole list of stories — branch, plan,
|
|
39
|
+
build, merge, next — until they're built or something stops it. No config file, no
|
|
24
40
|
build step.
|
|
25
41
|
|
|
26
42
|
```
|
|
@@ -32,6 +48,10 @@ PLANNING (optional on-ramps) EXECUTION (the core loop, one story at
|
|
|
32
48
|
idea → PRD PRD → stories intent design tasks per-task loop
|
|
33
49
|
PRD.md BREAKDOWN.md SPEC.md PLAN.md + tasks.md plan → implement → review → commit
|
|
34
50
|
constitution
|
|
51
|
+
▲ ▲
|
|
52
|
+
└──────────────── tiny-spec-loop ─────────────────┘
|
|
53
|
+
optional: per story — branch, run, build, merge,
|
|
54
|
+
then the next story. Halts on a wall.
|
|
35
55
|
```
|
|
36
56
|
|
|
37
57
|
The two on-ramps are **optional** and stack. Have nothing written down? Run
|
|
@@ -42,6 +62,97 @@ work? Skip both and start at `tiny-spec-create`. Both on-ramps write a regenerab
|
|
|
42
62
|
file at your project root (not under `.spec/`); `tiny-spec-create` then reads the
|
|
43
63
|
breakdown one story at a time.
|
|
44
64
|
|
|
65
|
+
## How small
|
|
66
|
+
|
|
67
|
+
Every other kit in this space is bigger. That is the whole pitch, so here is the
|
|
68
|
+
receipt rather than the adjective:
|
|
69
|
+
|
|
70
|
+
| | skills / commands | agents | config | artifacts per feature |
|
|
71
|
+
|---|---|---|---|---|
|
|
72
|
+
| **tiny-spec** | **8** (4 core + 4 optional) | **2** | **none** | **`SPEC` `PLAN` `tasks`** |
|
|
73
|
+
| [GitHub Spec Kit](https://github.com/github/spec-kit) | 10 | — | `specify init` | `spec` `plan` `tasks` `checklist` `constitution` `research` `data-model` `contracts/` `quickstart` |
|
|
74
|
+
| [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
|
|
75
|
+
| [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
|
|
76
|
+
|
|
77
|
+
The whole of tiny-spec is **2,079 lines** of markdown across 8 skills and 2 agents.
|
|
78
|
+
There is no config file, no build step, no orchestrator, and no CLI to initialize a
|
|
79
|
+
project — the skills are the product.
|
|
80
|
+
|
|
81
|
+
<sub>Counts taken from each project's own repository on 2026-08-03 by listing its
|
|
82
|
+
command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
|
|
83
|
+
`templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
|
|
84
|
+
OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
|
|
85
|
+
personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
|
|
86
|
+
`customize.toml` files (35). These projects are all substantially more popular than
|
|
87
|
+
tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
|
|
88
|
+
bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
|
|
89
|
+
|
|
90
|
+
## Why it's small
|
|
91
|
+
|
|
92
|
+
Most spec frameworks are generous by default:
|
|
93
|
+
many phases, many agents, many generated documents. tiny-spec makes the opposite
|
|
94
|
+
bet. Keep one safeguard, drop the rest.
|
|
95
|
+
|
|
96
|
+
A green unit test suite is not the same as working software, so the reviewer
|
|
97
|
+
exercises acceptance criteria end to end and a final smoke test confirms the whole
|
|
98
|
+
spec. That independent review is the safeguard — not the volume of planning
|
|
99
|
+
artifacts. One task, one commit, an external reviewer. Nothing gets added unless
|
|
100
|
+
it earns its place.
|
|
101
|
+
|
|
102
|
+
The case for staying small:
|
|
103
|
+
|
|
104
|
+
- **Documents are context, and context isn't free.** Generating large `spec.md`,
|
|
105
|
+
`plan.md`, `research.md`, and `data-model.md` files costs tokens to write, then
|
|
106
|
+
costs context to carry. Every paragraph the agent has to hold is room it no
|
|
107
|
+
longer has for your actual code. tiny-spec keeps the spine small — a
|
|
108
|
+
constitution and a short memory — and injects only what each task needs.
|
|
109
|
+
- **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
|
|
110
|
+
assume you're bootstrapping a project from a blank page. Day to day, you pick up
|
|
111
|
+
a ticket and change part of a system that already exists. tiny-spec binds to a
|
|
112
|
+
ticket, works one at a time, and references your task platform instead of
|
|
113
|
+
re-describing the world.
|
|
114
|
+
- **Rigid pipelines fight the user.** Mandatory phases and required sections
|
|
115
|
+
impose ceremony on work that doesn't need it. tiny-spec's extra structure is
|
|
116
|
+
optional by design — add shape where it pays, skip it where it doesn't.
|
|
117
|
+
- **More moving parts is more to maintain.** Orchestrators, ownership contracts,
|
|
118
|
+
checkpoint matrices, and config files are themselves a system you have to learn
|
|
119
|
+
and keep in sync. A few small skills and two agents are not.
|
|
120
|
+
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks
|
|
121
|
+
like progress, but it isn't proof. The proof is the reviewer running your real
|
|
122
|
+
tests before each commit.
|
|
123
|
+
|
|
124
|
+
That's the whole trade: where larger kits add machinery, tiny-spec adds one
|
|
125
|
+
independent reviewer and stops.
|
|
126
|
+
|
|
127
|
+
### Does the trade hold up?
|
|
128
|
+
|
|
129
|
+
Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
|
|
130
|
+
runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
|
|
131
|
+
code with held-out tests the suite never sees, and checks the reviewer's own verdict
|
|
132
|
+
against that ground truth — the number that matters being **false-PASS rate**, how
|
|
133
|
+
often the gate blesses code that is actually broken.
|
|
134
|
+
|
|
135
|
+
Most recent run — 5 benchmark tasks, v0.4.0 (`7810074`), 2026-08-03:
|
|
136
|
+
|
|
137
|
+
| | |
|
|
138
|
+
|---|---|
|
|
139
|
+
| held-out pass rate | **100%** (5/5) — the code works, judged by tests it never saw |
|
|
140
|
+
| suite/truth agreement | **100%** — the reviewer's verdict matched the held-out grader every time |
|
|
141
|
+
| **false-PASS rate** | **0%** — it never called broken code done |
|
|
142
|
+
| blocker rate | 0% |
|
|
143
|
+
|
|
144
|
+
The same run a month earlier on `f973c85` scored identically, which is the more
|
|
145
|
+
interesting result: the loop is at least stable across versions rather than tuned to a
|
|
146
|
+
single commit.
|
|
147
|
+
|
|
148
|
+
See [`docs/eval/README.md`](docs/eval/README.md) for the method and
|
|
149
|
+
[`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the scoring rubric
|
|
150
|
+
this repo grades itself against. **Read both skeptically: it is my benchmark, of my own
|
|
151
|
+
tool, scored by me.** The rubric says so itself — *"treat the tiny-spec column as
|
|
152
|
+
self-assessment to pressure-test, not gospel."* The tasks are small and self-contained,
|
|
153
|
+
so the harness measures the loop on well-specified work, not large-codebase performance.
|
|
154
|
+
The harness is in the repo precisely so you don't have to take my word for it.
|
|
155
|
+
|
|
45
156
|
## New to spec-driven development?
|
|
46
157
|
|
|
47
158
|
Spec-driven development (SDD) means writing down *what* you want and *why* before
|
|
@@ -84,6 +195,17 @@ Or collapse the three planning steps into one and go straight to building:
|
|
|
84
195
|
**stops before `tiny-spec-build`** — that's where you actually review the work — and
|
|
85
196
|
it writes nothing itself, it only delegates.
|
|
86
197
|
|
|
198
|
+
Or hand it a whole list of stories:
|
|
199
|
+
|
|
200
|
+
```
|
|
201
|
+
/tiny-spec-loop # per story: branch → plan → build → merge → next
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
`tiny-spec-loop` reads your `BREAKDOWN.md` (or a list you paste) and works the stories
|
|
205
|
+
one after another, merging each finished branch into `main` locally before starting the
|
|
206
|
+
next — until it's done or reaches a **terminal state** it names out loud. See
|
|
207
|
+
[Working a whole list](#working-a-whole-list--tiny-spec-loop).
|
|
208
|
+
|
|
87
209
|
**Building from a mockup?** There is no design flag to pass. Commit your exports to
|
|
88
210
|
`design/` before the run and `tiny-spec-create` — whether you invoke it directly or
|
|
89
211
|
reach it through `tiny-spec-run` — reads what's there and asks which screens this
|
|
@@ -104,7 +226,7 @@ git clone https://github.com/GrayMa77er/tiny-spec.git
|
|
|
104
226
|
cd tiny-spec
|
|
105
227
|
|
|
106
228
|
mkdir -p "$HOME/.claude/skills" "$HOME/.claude/agents"
|
|
107
|
-
for s in tiny-spec-prd tiny-spec-breakdown tiny-spec-run tiny-spec-create tiny-spec-plan tiny-spec-tasks tiny-spec-build; do
|
|
229
|
+
for s in tiny-spec-prd tiny-spec-breakdown tiny-spec-run tiny-spec-loop tiny-spec-create tiny-spec-plan tiny-spec-tasks tiny-spec-build; do
|
|
108
230
|
cp -R "$s" "$HOME/.claude/skills/$s"
|
|
109
231
|
done
|
|
110
232
|
cp agents/*.md "$HOME/.claude/agents/"
|
|
@@ -115,6 +237,13 @@ or install one set at a time.
|
|
|
115
237
|
|
|
116
238
|
</details>
|
|
117
239
|
|
|
240
|
+
### See a finished run first
|
|
241
|
+
|
|
242
|
+
[`examples/todo-cli/`](examples/todo-cli/) is a real run of the flow on one small
|
|
243
|
+
ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`,
|
|
244
|
+
`tasks.md` and `constitution.md` the suite wrote, and the code and tests it produced.
|
|
245
|
+
The tests pass; you can clone it and run the gate yourself.
|
|
246
|
+
|
|
118
247
|
## How it works
|
|
119
248
|
|
|
120
249
|
The constitution (`constitution.md`) is the spine. `tiny-spec-create` seeds it from a
|
|
@@ -295,19 +424,21 @@ unnoticed.
|
|
|
295
424
|
3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
|
|
296
425
|
end to end and grades against the constitution and the task's acceptance.
|
|
297
426
|
4. On pass, commit the code plus a checklist tick. On fail, loop back to the
|
|
298
|
-
executor with the findings. After two failed attempts
|
|
427
|
+
executor with the findings. After two failed attempts the run halts `exhausted`.
|
|
299
428
|
|
|
300
429
|
```mermaid
|
|
301
430
|
flowchart TB
|
|
302
431
|
SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
|
|
303
432
|
|
|
433
|
+
TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
|
|
304
434
|
TASKS --> P[Plan task]
|
|
305
435
|
P --> I[Implement<br/>executor]
|
|
306
436
|
I --> R[Review + run gate<br/>reviewer]
|
|
437
|
+
I -->|blocker| B
|
|
307
438
|
R -->|pass| C[Commit + tick]
|
|
308
439
|
C --> TASKS
|
|
309
440
|
R -->|fail| I
|
|
310
|
-
R -->|fail twice| B[
|
|
441
|
+
R -->|fail twice| B[Halt — blocked / exhausted<br/>logged to decisions.md]
|
|
311
442
|
|
|
312
443
|
CON([constitution.md]) -.-> P & I & R
|
|
313
444
|
MEM([memory.md]) -.-> I & R
|
|
@@ -328,42 +459,74 @@ and logs a blocker instead of hacking around it. You fix the gap upstream in
|
|
|
328
459
|
`tiny-spec-plan` or `tiny-spec-create`, then resume. Work runs one ticket at a time and
|
|
329
460
|
resumes from the checklist state.
|
|
330
461
|
|
|
331
|
-
|
|
462
|
+
### Working a whole list — `tiny-spec-loop`
|
|
332
463
|
|
|
333
|
-
|
|
334
|
-
|
|
335
|
-
bet. Keep one safeguard, drop the rest.
|
|
464
|
+
`/tiny-spec-loop` takes a list of stories and works them one after another. Per story
|
|
465
|
+
it does the same four moves:
|
|
336
466
|
|
|
337
|
-
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
artifacts. One task, one commit, an external reviewer. Nothing gets added unless
|
|
341
|
-
it earns its place.
|
|
467
|
+
```
|
|
468
|
+
cut a branch from main → tiny-spec-run → tiny-spec-build → merge back to main
|
|
469
|
+
```
|
|
342
470
|
|
|
343
|
-
|
|
471
|
+
Then the next story. Each branch is cut **fresh from main**, so story 3 sees stories 1
|
|
472
|
+
and 2 already merged — which is what makes an ordered list build correctly.
|
|
344
473
|
|
|
345
|
-
|
|
346
|
-
|
|
347
|
-
|
|
348
|
-
|
|
349
|
-
|
|
350
|
-
- **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
|
|
351
|
-
assume you're bootstrapping a project from a blank page. Day to day, you pick up
|
|
352
|
-
a ticket and change part of a system that already exists. tiny-spec binds to a
|
|
353
|
-
ticket, works one at a time, and references your task platform instead of
|
|
354
|
-
re-describing the world.
|
|
355
|
-
- **Rigid pipelines fight the user.** Mandatory phases and required sections
|
|
356
|
-
impose ceremony on work that doesn't need it. tiny-spec's extra structure is
|
|
357
|
-
optional by design — add shape where it pays, skip it where it doesn't.
|
|
358
|
-
- **More moving parts is more to maintain.** Orchestrators, ownership contracts,
|
|
359
|
-
checkpoint matrices, and config files are themselves a system you have to learn
|
|
360
|
-
and keep in sync. A few small skills and two agents are not.
|
|
361
|
-
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks
|
|
362
|
-
like progress, but it isn't proof. The proof is the reviewer running your real
|
|
363
|
-
tests before each commit.
|
|
474
|
+
**The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
|
|
475
|
+
each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
|
|
476
|
+
a list at invocation instead and that wins; but a bare feature name has no acceptance
|
|
477
|
+
criteria, so `tiny-spec-create` will interview you when it reaches it. That's the
|
|
478
|
+
honest trade: a breakdown runs unattended, a pasted list is supervised.
|
|
364
479
|
|
|
365
|
-
|
|
366
|
-
|
|
480
|
+
**Every run ends in exactly one of six states, and it says which:**
|
|
481
|
+
|
|
482
|
+
| | |
|
|
483
|
+
|---|---|
|
|
484
|
+
| `done` | every story built **and merged** |
|
|
485
|
+
| `blocked` | an upstream document is wrong — go fix the spec or the plan |
|
|
486
|
+
| `exhausted` | a task stayed red past two fix attempts |
|
|
487
|
+
| `paused` | it reached a `pause:` point |
|
|
488
|
+
| `fork` | a real either/or the plan doesn't answer |
|
|
489
|
+
| `conflict` | a story's branch wouldn't merge cleanly |
|
|
490
|
+
|
|
491
|
+
**Only `done` means the work is built** — and in a story loop that means *all* of them.
|
|
492
|
+
A run that stopped at story 2 of 7 reporting "done" is the single thing autonomous loops
|
|
493
|
+
get wrong most often, so the state is always named, along with what got merged and how
|
|
494
|
+
many stories are still untouched.
|
|
495
|
+
|
|
496
|
+
**A halt stops the whole run, not just that story.** Later stories in a list you wrote
|
|
497
|
+
top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
|
|
498
|
+
just produces a second, more confusing failure downstream.
|
|
499
|
+
|
|
500
|
+
**Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
|
|
501
|
+
the build halts *before* running it:
|
|
502
|
+
|
|
503
|
+
```
|
|
504
|
+
- [ ] T7 — run the schema migration against the restored snapshot
|
|
505
|
+
- acceptance: every row in orders has a non-null tenant_id
|
|
506
|
+
- pause: confirm the snapshot is current before this touches real rows
|
|
507
|
+
```
|
|
508
|
+
|
|
509
|
+
`tiny-spec-tasks` proposes these for genuinely irreversible work — migrations,
|
|
510
|
+
destructive file operations, a new dependency, an auth boundary, a public API contract.
|
|
511
|
+
You can also give the run a standing policy up front ("halt before anything that touches
|
|
512
|
+
auth") and it gets applied as each story's tasks are sliced.
|
|
513
|
+
|
|
514
|
+
**What it will not do to your repo.** It runs exactly five git commands — `switch`,
|
|
515
|
+
`switch -c`, `merge --no-ff`, `merge --abort`, and reads. It refuses to start on a dirty
|
|
516
|
+
tree. It **never pushes**, never rebases, never resets, never deletes a branch, and
|
|
517
|
+
never opens a PR. Merges are local, so a bad run is one `git reset` away; publishing
|
|
518
|
+
stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
|
|
519
|
+
the undo command rather than running it.
|
|
520
|
+
|
|
521
|
+
**Walk away and come back.** Nothing is written down to track progress — it's derived:
|
|
522
|
+
a story whose ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked
|
|
523
|
+
task is in progress, no directory means not started. So re-running `/tiny-spec-loop`
|
|
524
|
+
tomorrow in a fresh session picks up exactly where it stopped. No run-state file, no
|
|
525
|
+
lock, and no budget to configure: the story list *is* the budget.
|
|
526
|
+
|
|
527
|
+
**It never fixes a blocker for you.** A blocker means one of your documents is wrong,
|
|
528
|
+
and a loop allowed to rewrite the requirement its own task just failed would be grading
|
|
529
|
+
its own homework. It reports and stops; the upstream edit is yours.
|
|
367
530
|
|
|
368
531
|
## Project layout
|
|
369
532
|
|
|
@@ -0,0 +1,69 @@
|
|
|
1
|
+
# Example: todo CLI
|
|
2
|
+
|
|
3
|
+
A **real run** of the tiny-spec suite on one small ticket — not a hand-written ideal.
|
|
4
|
+
Everything in this folder (the `.spec/` artifacts and the code) was produced by
|
|
5
|
+
actually driving `tiny-spec-create → plan → tasks → build` autonomously on
|
|
6
|
+
[`TICKET.md`](TICKET.md), then committed verbatim. Browse it to see exactly what the
|
|
7
|
+
workflow generates before you run it on your own work.
|
|
8
|
+
|
|
9
|
+
> Produced on 2026-06-29 with tiny-spec **v0.2.1**, and kept as-run rather than
|
|
10
|
+
> retouched to match the current artifact formats — a doctored example would defeat
|
|
11
|
+
> the point. Later versions add a `## Design` section to `SPEC.md` (unused here: this
|
|
12
|
+
> is a CLI with no design surface) and richer status flags. The shape is the same.
|
|
13
|
+
|
|
14
|
+
## What's here
|
|
15
|
+
|
|
16
|
+
```
|
|
17
|
+
TICKET.md the input — a small "todo CLI" ask
|
|
18
|
+
.spec/
|
|
19
|
+
constitution.md the shared spine: style, standards, invariants,
|
|
20
|
+
definition of done, and the verification gate
|
|
21
|
+
todo-cli/
|
|
22
|
+
SPEC.md intent + REQ-1..REQ-9 (what "done" means)
|
|
23
|
+
PLAN.md the design and how each requirement is covered
|
|
24
|
+
tasks.md the ordered checklist, all ticked [x]
|
|
25
|
+
todo.py the produced CLI (stdlib only, 117 lines)
|
|
26
|
+
test_todo.py the produced end-to-end test (129 lines)
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Read them in flow order: `TICKET.md` → `SPEC.md` → `PLAN.md` → `tasks.md` →
|
|
30
|
+
`todo.py`. The `constitution.md` is the persistent context injected into every task.
|
|
31
|
+
|
|
32
|
+
## It passes its own gate
|
|
33
|
+
|
|
34
|
+
The committed code clears the gate defined in its own
|
|
35
|
+
[`constitution.md`](.spec/constitution.md) — the same check `tiny-spec-build-reviewer`
|
|
36
|
+
runs. From this directory:
|
|
37
|
+
|
|
38
|
+
```sh
|
|
39
|
+
python3 -m pytest test_todo.py # 8 passed
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
And the ticket's example session reproduces exactly:
|
|
43
|
+
|
|
44
|
+
```sh
|
|
45
|
+
python3 todo.py add "buy milk" # -> 1
|
|
46
|
+
python3 todo.py add "write tests" # -> 2
|
|
47
|
+
python3 todo.py list # two open lines
|
|
48
|
+
python3 todo.py done 1
|
|
49
|
+
python3 todo.py list # 1 now shows [x]
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
> Running these creates a `todos.json` in your working directory — delete it
|
|
53
|
+
> afterward, and don't commit it.
|
|
54
|
+
|
|
55
|
+
## How it was generated
|
|
56
|
+
|
|
57
|
+
Driven headlessly the same way the eval harness drives its benchmark tasks: a
|
|
58
|
+
hermetic sandbox with the suite vendored into a local `.claude/`, seeded with
|
|
59
|
+
`TICKET.md`, and run via `claude -p` with instructions to complete the flow
|
|
60
|
+
autonomously. See [`docs/eval/harness/run.sh`](../../docs/eval/harness/run.sh) for the
|
|
61
|
+
exact recipe — the `DRIVER` prompt near the top is what was handed to the model.
|
|
62
|
+
|
|
63
|
+
This run finished in 55 turns at roughly $3.49 with no blockers; all five tasks
|
|
64
|
+
passed review on the first or second attempt.
|
|
65
|
+
|
|
66
|
+
To produce your own equivalent, point that harness recipe at this `TICKET.md`, or
|
|
67
|
+
just run the flow interactively in a scratch directory. Output won't be
|
|
68
|
+
byte-identical — the model is nondeterministic — but it will be an equivalent,
|
|
69
|
+
gate-passing run.
|
|
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
|
|
|
4
4
|
|
|
5
5
|
[project]
|
|
6
6
|
name = "tiny-spec"
|
|
7
|
-
version = "0.
|
|
7
|
+
version = "0.5.0"
|
|
8
8
|
description = "A tiny, opinionated take on spec-driven development."
|
|
9
9
|
readme = "README.md"
|
|
10
10
|
requires-python = ">=3.9"
|
|
@@ -37,6 +37,7 @@ packages = ["tiny_spec"]
|
|
|
37
37
|
"tiny-spec-prd" = "tiny_spec/_bundle/tiny-spec-prd"
|
|
38
38
|
"tiny-spec-breakdown" = "tiny_spec/_bundle/tiny-spec-breakdown"
|
|
39
39
|
"tiny-spec-run" = "tiny_spec/_bundle/tiny-spec-run"
|
|
40
|
+
"tiny-spec-loop" = "tiny_spec/_bundle/tiny-spec-loop"
|
|
40
41
|
"tiny-spec-create" = "tiny_spec/_bundle/tiny-spec-create"
|
|
41
42
|
"tiny-spec-plan" = "tiny_spec/_bundle/tiny-spec-plan"
|
|
42
43
|
"tiny-spec-tasks" = "tiny_spec/_bundle/tiny-spec-tasks"
|
|
@@ -48,6 +49,7 @@ include = [
|
|
|
48
49
|
"tiny_spec",
|
|
49
50
|
"tiny-spec-prd",
|
|
50
51
|
"tiny-spec-run",
|
|
52
|
+
"tiny-spec-loop",
|
|
51
53
|
"tiny-spec-breakdown",
|
|
52
54
|
"tiny-spec-create",
|
|
53
55
|
"tiny-spec-plan",
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: tiny-spec-build
|
|
3
|
-
description: Build the spec — run the per-task loop plan→implement→review→commit, one task at a time. Implements with a fresh executor, grades with an independent reviewer running the real gate, commits per passed task, keeps a lean memory. Resumes from the checkbox state.
|
|
3
|
+
description: Build the spec — run the per-task loop plan→implement→review→commit, one task at a time. Implements with a fresh executor, grades with an independent reviewer running the real gate, commits per passed task, keeps a lean memory. Halts on a blocker, a spent convergence budget, a task's pause: point, or a genuine fork, and names which. Resumes from the checkbox state.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# tiny-spec-build
|
|
@@ -42,6 +42,15 @@ alongside this skill (see the suite README).
|
|
|
42
42
|
For the selected task, run these steps in order. **Do not tick a task until its
|
|
43
43
|
reviewer passes.**
|
|
44
44
|
|
|
45
|
+
**Before step 1, check the task's `pause:` field.** If it has one and the user hasn't
|
|
46
|
+
waived it for this run, **halt `paused`** (see **Halting** below): leave the task
|
|
47
|
+
`[ ]`, dispatch nothing, and don't touch `tasks.md`. A pause stops *before* the work so
|
|
48
|
+
the approach gets reviewed while redirecting it is still cheap.
|
|
49
|
+
|
|
50
|
+
A pause can be waived or added for the current run only, conversationally — "skip the
|
|
51
|
+
pause on T5", "also stop before T9". Honor it for this run and **never write it back
|
|
52
|
+
into `tasks.md`**: the file records the standing intent, not one run's override.
|
|
53
|
+
|
|
45
54
|
### 1. PLAN (inline, brief)
|
|
46
55
|
Restate the task as a 2–4 step micro-plan against the constitution: which
|
|
47
56
|
**invariants** apply, which files it touches, which **Definition of Done** items
|
|
@@ -90,8 +99,12 @@ It fails on what it measured or saw, and merely flags what is taste.
|
|
|
90
99
|
|
|
91
100
|
### 4. CONVERGE (on FAIL)
|
|
92
101
|
Re-dispatch the **executor** with the reviewer's findings appended to its brief.
|
|
93
|
-
Bound this to **2 fix attempts**. If it still fails after that, stop and
|
|
94
|
-
|
|
102
|
+
Bound this to **2 fix attempts**. If it still fails after that, stop and **halt
|
|
103
|
+
`exhausted`** (see **Halting** below) — don't keep grinding or hand-fix past the loop
|
|
104
|
+
silently. `exhausted` is recorded as a blocker, because a task that three attempts
|
|
105
|
+
couldn't turn green is a signal the upstream document is wrong — but name it
|
|
106
|
+
`exhausted`, not `blocked`: the executor never said it couldn't proceed, it just
|
|
107
|
+
never got there.
|
|
95
108
|
|
|
96
109
|
### 5. COMMIT + TICK (on PASS)
|
|
97
110
|
Two commits, in order (keeps code history clean of planning churn), both in
|
|
@@ -156,12 +169,62 @@ On first use, create the file with this structure:
|
|
|
156
169
|
Report the task outcome (built, reviewed, committed). Then:
|
|
157
170
|
- **Interactive default:** continue to the next unchecked task. Pausing for the
|
|
158
171
|
user between tasks is fine and expected.
|
|
159
|
-
- If the user asked to **run it through** ("do it all", "build everything"),
|
|
160
|
-
|
|
172
|
+
- If the user asked to **run it through** ("do it all", "build everything"), or
|
|
173
|
+
`tiny-spec-loop` invoked you, keep looping until one of the five terminal states
|
|
174
|
+
below is reached — committing per passed task as you go.
|
|
175
|
+
|
|
176
|
+
**The task list is the budget.** The loop ends when the tasks end. There is no turn
|
|
177
|
+
ceiling, no token cap, and no max-tasks knob to set — adding one would put a second,
|
|
178
|
+
weaker stopping rule next to the real one.
|
|
161
179
|
|
|
162
180
|
There is **no** separate autonomous mode and **no** checkpoint config: one commit
|
|
163
181
|
per passed task, always, on the current branch. (If the user wants a feature
|
|
164
|
-
branch, create it once up front — that's their call, not a knob here.)
|
|
182
|
+
branch, create it once up front — that's their call, not a knob here.) Running
|
|
183
|
+
through is a caller's instruction, not a mode — every step above is identical either
|
|
184
|
+
way, which is why a loop and a hand-driven build can't drift apart.
|
|
185
|
+
|
|
186
|
+
## Halting
|
|
187
|
+
|
|
188
|
+
Every build run ends in exactly **one** of five terminal states. Name it out loud in
|
|
189
|
+
the closing report:
|
|
190
|
+
|
|
191
|
+
| State | Cause |
|
|
192
|
+
|---|---|
|
|
193
|
+
| `done` | every task `[x]` and the final smoke passed |
|
|
194
|
+
| `blocked` | the executor reported a `BLOCKER`, or the reviewer found a missing `visual:` command |
|
|
195
|
+
| `exhausted` | convergence (step 4) spent its 2 fix attempts and the task is still red |
|
|
196
|
+
| `paused` | a task's `pause:` point was reached |
|
|
197
|
+
| `fork` | a genuine either/or the plan doesn't pin down |
|
|
198
|
+
|
|
199
|
+
**Never round up.** Only `done` may report the work as built. A run that ended at
|
|
200
|
+
`blocked`, `exhausted`, `paused`, or `fork` has unbuilt work in it, and a closing
|
|
201
|
+
report that calls it done is worse than the halt itself — it converts a stop the user
|
|
202
|
+
could act on into a false completion they won't check. Say the state, then say what's
|
|
203
|
+
left and which command resolves it.
|
|
204
|
+
|
|
205
|
+
**`blocked` and `exhausted`** log **one** entry with `type: blocker` — an upstream doc
|
|
206
|
+
must change; see **Blockers** below. **`fork`** logs `type: halt`:
|
|
207
|
+
|
|
208
|
+
```
|
|
209
|
+
## D-NNN — <short title>
|
|
210
|
+
- type: halt
|
|
211
|
+
- state: fork
|
|
212
|
+
- date: <ISO date>
|
|
213
|
+
- affects: REQ-N, T<n>
|
|
214
|
+
- note: <the options, your recommendation, and what the user has to decide>
|
|
215
|
+
```
|
|
216
|
+
|
|
217
|
+
Write one entry per halt, not one per cause — a blocker that halts the loop is a single
|
|
218
|
+
event and reads as two records if you log it twice.
|
|
219
|
+
|
|
220
|
+
**`done` and `paused` log nothing.** A green tree with every box ticked already says
|
|
221
|
+
`done`. And a `paused` run stopped at a task whose `pause:` line is sitting right there
|
|
222
|
+
in `tasks.md` saying why — an entry would restate what the file already tells you, and
|
|
223
|
+
the log is for things that aren't otherwise visible. Report it, don't record it.
|
|
224
|
+
|
|
225
|
+
The task the loop stopped at stays `[ ]`. That is the whole resume mechanism:
|
|
226
|
+
re-running `tiny-spec-build` (or `tiny-spec-loop`) picks up at exactly that task with no
|
|
227
|
+
conversational context needed.
|
|
165
228
|
|
|
166
229
|
## Blockers (never hack around)
|
|
167
230
|
|
|
@@ -170,16 +233,22 @@ command** for a `design:` task, or when convergence (step 4) exhausts its attemp
|
|
|
170
233
|
|
|
171
234
|
1. Leave the task `[ ]`.
|
|
172
235
|
2. Log it to `.spec/<active>/decisions.md` using the fixed skeleton (`type: blocker`,
|
|
173
|
-
|
|
174
|
-
`affects:`). Create the file if absent:
|
|
236
|
+
the terminal state in `state:`, the upstream doc to fix named in `note:`, and the
|
|
237
|
+
affected `REQ-N`/`T<n>` in `affects:`). Create the file if absent:
|
|
175
238
|
|
|
176
239
|
```
|
|
177
240
|
## D-NNN — <short title>
|
|
178
241
|
- type: blocker
|
|
242
|
+
- state: blocked
|
|
179
243
|
- date: <ISO date>
|
|
180
244
|
- affects: REQ-N, T<n>
|
|
181
245
|
- note: <what stopped you; which upstream doc must change>
|
|
182
246
|
```
|
|
247
|
+
|
|
248
|
+
`state:` is `blocked` for an executor blocker or a missing `visual:`, and
|
|
249
|
+
`exhausted` when convergence ran out. Both route upstream the same way; the
|
|
250
|
+
distinction is what the user reads to know whether the task was impossible or just
|
|
251
|
+
never converged.
|
|
183
252
|
3. **Surface and route upstream:** `tiny-spec-plan` (a design gap, a missing `visual:`
|
|
184
253
|
command, or a token the Design system doesn't define) or `tiny-spec-create` (a
|
|
185
254
|
requirement is wrong/impossible), in update mode. Never resolve a missing `visual:`
|
|
@@ -190,6 +259,12 @@ command** for a `design:` task, or when convergence (step 4) exhausts its attemp
|
|
|
190
259
|
A genuine fork the plan doesn't pin down → don't guess: present the options + your
|
|
191
260
|
recommendation, get the user's call, record it in `decisions.md`, then continue.
|
|
192
261
|
|
|
262
|
+
**When running through, a fork is a halt** (`type: halt`, `state: fork`) — record the
|
|
263
|
+
options and your recommendation in the `note:` and stop, rather than blocking on an
|
|
264
|
+
answer nobody may be present to give. Ask only when the user is in the conversation
|
|
265
|
+
with you; a question asked into an empty room stalls the run without ending it, and
|
|
266
|
+
the halt is what makes the same fork readable hours later.
|
|
267
|
+
|
|
193
268
|
## Completion
|
|
194
269
|
|
|
195
270
|
When every task in `tasks.md` is `[x]`:
|
|
@@ -199,10 +274,11 @@ When every task in `tasks.md` is `[x]`:
|
|
|
199
274
|
setup — install/build, not a test-runner shortcut). There is no separate
|
|
200
275
|
verify skill — this final smoke confirms the requirements actually work end-to-end, not
|
|
201
276
|
just that tasks are ticked.
|
|
202
|
-
2. **Report** —
|
|
203
|
-
|
|
277
|
+
2. **Report** — the terminal state by name (`done` here, since every task is `[x]`
|
|
278
|
+
and the smoke passed), what was built, the commits made (with the branch), and any
|
|
279
|
+
open `decisions.md` items (blockers, tasks unchecked by a reconcile). If the final
|
|
204
280
|
smoke reveals a gap, it's a bug to fix now (new/edited task) or a blocker to
|
|
205
|
-
route upstream — not a pass
|
|
281
|
+
route upstream — **not** a pass, and not `done`.
|
|
206
282
|
3. **Close the ticket (reference-only).** If a ticket is bound, this is where the
|
|
207
283
|
**closing** keyword belongs — on the final / PR commit, not the per-task ones:
|
|
208
284
|
GitHub `Closes #123`, ADO `Fixes AB#123`. Jira/Monday have no closing keyword,
|
|
@@ -111,6 +111,21 @@ Only the **project-wide** questions collapse — still confirm this story's bind
|
|
|
111
111
|
requirements. If there is **no `BREAKDOWN.md`**, or no entry matches, run the **full
|
|
112
112
|
interview** below unchanged.
|
|
113
113
|
|
|
114
|
+
**Exception — a loop run.** If the caller says this is a **loop run** and names the
|
|
115
|
+
story's slug (`tiny-spec-loop` does, through `tiny-spec-run`), skip step 6's
|
|
116
|
+
confirmation and write `SPEC.md` straight out. The `AC:` lines *are* the approval: the
|
|
117
|
+
user reviewed them when they wrote or accepted `BREAKDOWN.md`, and re-confirming them
|
|
118
|
+
once per story is asking the same question twice — across a seven-story run it is the
|
|
119
|
+
difference between walking away and being interrupted seven times. Report the captured
|
|
120
|
+
`REQ-N` instead of asking about them.
|
|
121
|
+
|
|
122
|
+
This waives **only** that confirmation. Everything that is a genuine question still
|
|
123
|
+
stops the run: a story whose `AC:` lines contradict each other or the Decisions block,
|
|
124
|
+
an `AC:` you cannot turn into a testable `REQ-N`, a missing ticket id the platform
|
|
125
|
+
needs, a design export the story names that isn't on disk. Those are not ceremony —
|
|
126
|
+
they are the human input the loop is supposed to stop for. Never guess past one because
|
|
127
|
+
a loop is running.
|
|
128
|
+
|
|
114
129
|
## Pick the slug (resolve the active dir)
|
|
115
130
|
|
|
116
131
|
**Fresh modes only** — in reseed or update mode the active dir already exists and you
|
|
@@ -0,0 +1,188 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: tiny-spec-loop
|
|
3
|
+
description: Work a list of stories end to end, one after another — for each story cut a branch from main, plan it, build it, merge it back locally, then move to the next. Reads BREAKDOWN.md by default, or a list pasted at invocation. Halts the WHOLE run on a blocker, a spent convergence budget, a task's pause: point, a genuine fork, a merge conflict, or a red gate after a merge — and names which. Use for "build the backlog", "work through the breakdown", "build these features one after another". NOT for a single story — for that use tiny-spec-run then tiny-spec-build.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# tiny-spec-loop
|
|
7
|
+
|
|
8
|
+
A **router**, not a stage. It owns no artifact, writes nothing, and holds no state
|
|
9
|
+
file. It walks a list of stories and, for each one, drives the existing flow to a
|
|
10
|
+
finished, merged branch — then moves to the next.
|
|
11
|
+
|
|
12
|
+
Per story the cycle is always the same four moves:
|
|
13
|
+
|
|
14
|
+
```
|
|
15
|
+
cut a branch from main → tiny-spec-run → tiny-spec-build → merge back to main
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
**Everything it decides comes from files on disk and from git.** Which stories are
|
|
19
|
+
built, which are merged, which haven't started — all derived, never recorded. That is
|
|
20
|
+
what makes a run abandoned overnight resume correctly tomorrow, and it is why this
|
|
21
|
+
skill needs no run log, pointer, or lock. Do not "optimize" this into stored state.
|
|
22
|
+
|
|
23
|
+
**It supplies one thing the flow was missing: a stopping rule.** The task list is the
|
|
24
|
+
goal, the independent reviewer running the real gate is the verification, `memory.md`
|
|
25
|
+
plus a commit per passed task is the memory. This skill adds when to stop, and nothing
|
|
26
|
+
else — no budget, no turn ceiling, no config.
|
|
27
|
+
|
|
28
|
+
## Step 0 — the run brief
|
|
29
|
+
|
|
30
|
+
The one interactive moment. Everything after it runs without asking unless a stage
|
|
31
|
+
genuinely needs an answer.
|
|
32
|
+
|
|
33
|
+
1. **Check the preconditions.** All three, before touching anything:
|
|
34
|
+
- it is a git repo (`git rev-parse --is-inside-work-tree`);
|
|
35
|
+
- the working tree is **clean** (`git status --porcelain` prints nothing) — a dirty
|
|
36
|
+
tree would get swept into the first story's commits;
|
|
37
|
+
- an integration branch exists — `main`, or `master` if there is no `main`.
|
|
38
|
+
|
|
39
|
+
Any of them failing → say which and stop. Do not offer to stash, commit, or create
|
|
40
|
+
the branch for the user.
|
|
41
|
+
|
|
42
|
+
2. **Resolve the story list** (Step 1) and play it back in order — slug and title —
|
|
43
|
+
with the state you derived for each (Step 2). This is the user's chance to
|
|
44
|
+
reorder, drop, or narrow before anything is cut.
|
|
45
|
+
|
|
46
|
+
3. **Take the pause policy.** Ask for, or accept, standing technical stop points for
|
|
47
|
+
this run — *"halt before anything that touches auth"*, *"stop before any schema
|
|
48
|
+
migration"*. Carry the wording verbatim into every story's `tiny-spec-tasks` stage
|
|
49
|
+
(Step 3), where it becomes a real `pause:` line on the matching task. Pause points
|
|
50
|
+
are **technical**, not per-story: there is no way to mark a whole story for review,
|
|
51
|
+
because the thing worth looking at is a migration or an auth boundary, not a
|
|
52
|
+
feature heading.
|
|
53
|
+
|
|
54
|
+
Step 0 is a **once-per-run** check. Do not re-run it between stories.
|
|
55
|
+
|
|
56
|
+
## Step 1 — the story list
|
|
57
|
+
|
|
58
|
+
**Default: `BREAKDOWN.md` at the project root**, in file order. Each `- Story:` under
|
|
59
|
+
each `## Feature:` heading is one item; take its **`slug:`** — that names both the
|
|
60
|
+
branch and the `.spec/<slug>/` dir. A `## Feature:` heading is a grouping, not an item:
|
|
61
|
+
it carries no slug, so a Feature with three stories is three branches and three merges.
|
|
62
|
+
|
|
63
|
+
**A list pasted at invocation wins** over `BREAKDOWN.md` when the user gives one. Treat
|
|
64
|
+
each line as a story title and derive a kebab-case slug from it. Such a story has **no
|
|
65
|
+
acceptance criteria**, so `tiny-spec-create` will run its full interview when it reaches
|
|
66
|
+
that story — which is correct, not a failure: a one-line feature name is not enough to
|
|
67
|
+
build from, and that interview *is* the human input the run stops for. Say so in Step 0
|
|
68
|
+
so the user knows a pasted list is a supervised run, not a walk-away one.
|
|
69
|
+
|
|
70
|
+
If neither exists, stop and say so. Never invent the list.
|
|
71
|
+
|
|
72
|
+
## Step 2 — where each story stands (derive, don't record)
|
|
73
|
+
|
|
74
|
+
For each story slug, in order, ask git — **in this order**, first match wins:
|
|
75
|
+
|
|
76
|
+
1. **`git show <integration>:.spec/<slug>/tasks.md`** succeeds and every task is `[x]`
|
|
77
|
+
→ **built and merged.** Skip it.
|
|
78
|
+
2. Otherwise, if the branch exists (`git rev-parse --verify <slug>`), read
|
|
79
|
+
**`git show <slug>:.spec/<slug>/tasks.md`**:
|
|
80
|
+
- succeeds, every task `[x]` → **built, not merged** → resume at the **merge** (4).
|
|
81
|
+
- succeeds, at least one `[ ]` → **in progress** → resume at the **build** (3).
|
|
82
|
+
- fails (no task list yet) → **planning incomplete** → resume at the **plan** (2).
|
|
83
|
+
3. Otherwise → **not started.** Run the full cycle.
|
|
84
|
+
|
|
85
|
+
Ask **git**, not the working tree. A story that isn't merged yet has no `.spec/<slug>/`
|
|
86
|
+
on the integration branch at all, so "does the directory exist" can't tell "not started"
|
|
87
|
+
apart from "built on a branch you haven't merged" — reading each ref explicitly can.
|
|
88
|
+
Sourcing step 1 from the integration branch is also what survives a deleted branch: once
|
|
89
|
+
a story is merged, its ticked task list is part of `main` whether or not the branch that
|
|
90
|
+
built it still exists.
|
|
91
|
+
|
|
92
|
+
## Step 3 — the per-story cycle
|
|
93
|
+
|
|
94
|
+
For the first story that isn't already built and merged:
|
|
95
|
+
|
|
96
|
+
1. **Branch.** `git switch <slug>` if it already exists; otherwise
|
|
97
|
+
`git switch -c <slug> <integration>` — cut **fresh from the integration branch** so
|
|
98
|
+
this story sees every story merged before it. That is what makes an ordered list
|
|
99
|
+
build correctly: story 3 gets stories 1 and 2 already in its tree.
|
|
100
|
+
|
|
101
|
+
2. **Plan.** Invoke **`tiny-spec-run`**, once. It resolves `.spec/<slug>/` from the
|
|
102
|
+
branch name automatically, so you don't name the ticket dir. Hand it two things:
|
|
103
|
+
- a **create-stage brief**: *this is a loop run; seed from the `BREAKDOWN.md` story
|
|
104
|
+
with slug `<slug>` and do not stop to confirm the requirements — its `AC:` lines
|
|
105
|
+
are already approved.* Omit the seeding clause for a pasted-list story; there is
|
|
106
|
+
nothing to seed from and the interview is correct.
|
|
107
|
+
- a **tasks-stage brief**: the pause policy from Step 0, verbatim, so the technical
|
|
108
|
+
stop points land as `pause:` lines.
|
|
109
|
+
|
|
110
|
+
If `tiny-spec-run` stops anywhere other than **L9** (tasks ready) or **L10** (already
|
|
111
|
+
built), that is a halt — report and stop the whole run. **Never invoke it twice to
|
|
112
|
+
push past its own stop:** each of those is a human decision it deliberately declined
|
|
113
|
+
to make, and running it again declines again.
|
|
114
|
+
|
|
115
|
+
3. **Build.** Invoke **`tiny-spec-build`**, once, briefed to run it through. It owns
|
|
116
|
+
the per-task loop and writes its own halt record. Anything other than `done` halts
|
|
117
|
+
the whole run — **do not merge a story that didn't finish**, and do not invoke build
|
|
118
|
+
a second time: it resumes from the checkbox state, so it lands on the very task that
|
|
119
|
+
just halted and halts there again.
|
|
120
|
+
|
|
121
|
+
4. **Merge — only on `done`.** In order:
|
|
122
|
+
- `git switch <integration>`
|
|
123
|
+
- `git merge --no-ff <slug>` — the merge commit keeps each story legible in history.
|
|
124
|
+
- **Conflict** → `git merge --abort`, then halt `conflict` (below).
|
|
125
|
+
- **Run the constitution's Verification commands on the merged result**, exercised
|
|
126
|
+
the way a user would. A story that was green alone can still break against work
|
|
127
|
+
merged before it, and that is exactly what this catches. Red → halt `blocked`.
|
|
128
|
+
- **Never push.** Merging locally keeps a bad run one `git reset` away; sending it
|
|
129
|
+
to a remote is the user's call, and this suite makes no network calls.
|
|
130
|
+
|
|
131
|
+
5. **Next story.** Return to Step 2. Do not re-run Step 0.
|
|
132
|
+
|
|
133
|
+
## Step 4 — report
|
|
134
|
+
|
|
135
|
+
Name, in this order:
|
|
136
|
+
|
|
137
|
+
1. **The terminal state** — exactly one of `done`, `blocked`, `exhausted`, `paused`,
|
|
138
|
+
`fork`, `conflict`. Use the word.
|
|
139
|
+
2. **Stories built and merged**, in order, with their merge commits.
|
|
140
|
+
3. **The story it stopped on**, the task within it, and why in one line.
|
|
141
|
+
4. **Stories never started** — say how many are left, by name. A run that stopped at
|
|
142
|
+
story 2 of 7 must not read like a finished backlog.
|
|
143
|
+
5. **The one command that resolves it** — `tiny-spec-create`/`tiny-spec-plan` in update
|
|
144
|
+
mode for `blocked`/`exhausted`, `tiny-spec-loop` again for `paused`, the decision
|
|
145
|
+
the user owes you for a `fork`, or the conflicted paths for a `conflict`.
|
|
146
|
+
|
|
147
|
+
**Only `done` — every story merged — may report the work as built.** A run that halted
|
|
148
|
+
has unbuilt stories in it, and a report that rounds `blocked`, `exhausted`, `paused`,
|
|
149
|
+
`fork`, or `conflict` up to done converts a stop the user could act on into a false
|
|
150
|
+
completion they won't check. Say the state, then say what's left.
|
|
151
|
+
|
|
152
|
+
### Halting
|
|
153
|
+
|
|
154
|
+
A halt stops **the whole run**, not just the current story. Later stories in a list you
|
|
155
|
+
wrote top to bottom usually assume the earlier ones landed, so skipping ahead past a
|
|
156
|
+
failure produces a second, more confusing failure downstream.
|
|
157
|
+
|
|
158
|
+
`tiny-spec-build` records its own halts (`blocked`, `exhausted`, `fork`) in the story's
|
|
159
|
+
`decisions.md`. **You record nothing** — the merge-stage halts are already legible
|
|
160
|
+
without a log: a `conflict` leaves the story's branch unmerged and git itself reports
|
|
161
|
+
the conflicted paths, and a red gate after a merge is reported by the gate. A log entry
|
|
162
|
+
restating what git already shows is a second source of truth with extra steps.
|
|
163
|
+
|
|
164
|
+
**On a red gate after a merge, leave the merge in place.** Report it, name the story,
|
|
165
|
+
and tell the user that `git reset --hard HEAD~1` on the integration branch undoes it.
|
|
166
|
+
Do not undo it yourself: fixing forward and rolling back are both reasonable, the
|
|
167
|
+
choice is theirs, and discarding a real merge is not a call a router gets to make.
|
|
168
|
+
|
|
169
|
+
### Hard rules
|
|
170
|
+
|
|
171
|
+
- **Never write, edit, or flip anything.** No `status:`, no checkbox, no
|
|
172
|
+
`decisions.md` entry, no code, no `BREAKDOWN.md` edit. Delegate or stop.
|
|
173
|
+
- **Never push, force, rebase, reset, or delete a branch.** The only git commands this
|
|
174
|
+
skill issues are `switch`, `switch -c`, `merge --no-ff`, `merge --abort`, and
|
|
175
|
+
read-only queries. Everything outward-facing or destructive is the user's.
|
|
176
|
+
- **Never merge a story whose build didn't return `done`.**
|
|
177
|
+
- **At most one `tiny-spec-run` and one `tiny-spec-build` per story.**
|
|
178
|
+
- **Never invoke `tiny-spec-loop`.** Re-entering means re-reading these steps, not
|
|
179
|
+
calling yourself. Self-invocation compounds context and does not terminate.
|
|
180
|
+
- **Never invoke `tiny-spec-create`, `tiny-spec-plan`, or `tiny-spec-tasks` directly.**
|
|
181
|
+
That ladder is `tiny-spec-run`'s; walking it here would duplicate it, and two ladders
|
|
182
|
+
drift. Briefs are passed *through* run, not around it.
|
|
183
|
+
- **Never resolve a halt yourself.** A blocker means an upstream document is wrong,
|
|
184
|
+
which is the user's call. Routing to `plan`/`create` in update mode automatically
|
|
185
|
+
would let the loop rewrite the requirement its own task just failed to satisfy — the
|
|
186
|
+
agent grading its own homework, one level up.
|
|
187
|
+
- **There is no budget to set.** The story list is the budget: the run ends when the
|
|
188
|
+
stories end. No turn ceiling, no token cap, no max-stories knob.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: tiny-spec-run
|
|
3
|
-
description: Run the spec flow end to end — read each artifact's status flag to work out where the active ticket stands, then invoke tiny-spec-create → tiny-spec-plan → tiny-spec-tasks in order, taking anything marked stale before anything missing. STOPS before tiny-spec-build; it never builds, never commits, and never writes an artifact itself. Use for "run the whole spec flow", "spec this out end to end", or "pick the chain back up after a change". NOT for a single stage — for that, invoke tiny-spec-create, tiny-spec-plan, or tiny-spec-tasks directly.
|
|
3
|
+
description: Run the spec flow end to end — read each artifact's status flag to work out where the active ticket stands, then invoke tiny-spec-create → tiny-spec-plan → tiny-spec-tasks in order, taking anything marked stale before anything missing. STOPS before tiny-spec-build; it never builds, never commits, and never writes an artifact itself. Use for "run the whole spec flow", "spec this out end to end", or "pick the chain back up after a change". NOT for a single stage — for that, invoke tiny-spec-create, tiny-spec-plan, or tiny-spec-tasks directly. NOT for building — for "spec it out and build it", use tiny-spec-loop.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# tiny-spec-run
|
|
@@ -29,9 +29,11 @@ to resolve. `tiny-spec-create` picks `BREAKDOWN.md` up on its own when it exists
|
|
|
29
29
|
If the user named a **single stage** ("update the spec", "redo the tasks"), hand off
|
|
30
30
|
to that skill and stop. `run` is for walking the chain, not for wrapping one stage.
|
|
31
31
|
|
|
32
|
-
If they asked for the flow **and** the build ("spec this out and build it"),
|
|
33
|
-
|
|
34
|
-
|
|
32
|
+
If they asked for the flow **and** the build ("spec this out and build it"), that is
|
|
33
|
+
`tiny-spec-loop` — hand off to it and stop. If you are *already* inside a
|
|
34
|
+
`tiny-spec-loop` run (it invoked you), run the chain as normal and stop at the end
|
|
35
|
+
anyway; the loop takes it from there. Either way **you** never enter the build. The
|
|
36
|
+
stop is not negotiable; see the hard rules.
|
|
35
37
|
|
|
36
38
|
Step 0 is a **once-per-run** check on the user's opening request. Do not re-run it
|
|
37
39
|
when you return to Step 1 after a stage.
|
|
@@ -181,6 +183,13 @@ something new:
|
|
|
181
183
|
fall back to an existing ticket dir — the user confirmed this is new work.*
|
|
182
184
|
(`tiny-spec-create` honors that phrase by skipping its sole-dir fallback.)
|
|
183
185
|
|
|
186
|
+
**Caller briefs — pass them through verbatim.** A caller (today, `tiny-spec-loop`) may
|
|
187
|
+
hand you a brief addressed to a specific stage: *"for the create stage: …"*, *"for the
|
|
188
|
+
tasks stage: …"*. Append it, **word for word**, to the scope you send that stage, and
|
|
189
|
+
send nothing to the others. Do not summarize it, act on it yourself, or let it change
|
|
190
|
+
which rung fires — the ladder is still yours, and a brief is cargo, not an instruction
|
|
191
|
+
to you. A brief for a stage this run never reaches is simply never delivered.
|
|
192
|
+
|
|
184
193
|
**A stage's closing "point the user at X" is not a terminus.** Each stage ends by
|
|
185
194
|
naming the next skill ("point the user at `tiny-spec-plan`"). Inside a run that
|
|
186
195
|
sentence is a *report*, not a stop — when a stage finishes, return to Step 1 and keep
|
|
@@ -224,5 +233,9 @@ invent a `## Design` section and the condition would still hold on the next pass
|
|
|
224
233
|
Telling the user once and letting them decide is the version that terminates.
|
|
225
234
|
|
|
226
235
|
Then hand off explicitly: **run `tiny-spec-build` when you're ready to build** (it
|
|
227
|
-
starts at the first unchecked task
|
|
228
|
-
|
|
236
|
+
starts at the first unchecked task, reviewing as it goes), or **`tiny-spec-loop`** to
|
|
237
|
+
run it through until the work is built or something stops it. If you stopped at L4 or
|
|
238
|
+
on a bound, say exactly what stopped you and what the user needs to decide.
|
|
239
|
+
|
|
240
|
+
(If `tiny-spec-loop` invoked you, this report is what it reads to decide whether to
|
|
241
|
+
build — so say which rung you stopped on either way.)
|
|
@@ -45,6 +45,7 @@ For each task, write:
|
|
|
45
45
|
- type: feat # optional; Conventional Commit type (defaults to feat)
|
|
46
46
|
- req: REQ-n # optional; the REQ-N this task delivers
|
|
47
47
|
- design: D-n # optional; the SPEC.md D<n> screen this task builds — arms the visual gate
|
|
48
|
+
- pause: <why> # optional; halt the loop before this task so a human looks first
|
|
48
49
|
- files: <comma-separated hint of files it will touch>
|
|
49
50
|
```
|
|
50
51
|
|
|
@@ -70,6 +71,32 @@ other. Two rules:
|
|
|
70
71
|
- Only set it if the constitution has a `visual:` verification command. Without one
|
|
71
72
|
the reviewer cannot render anything and will raise a blocker instead of a verdict.
|
|
72
73
|
|
|
74
|
+
**pause** is the other field that changes what happens rather than what gets written —
|
|
75
|
+
but where `design:` changes how a task is **graded**, `pause:` changes whether the loop
|
|
76
|
+
**continues**. Set it and `tiny-spec-build` halts *before* that task, leaving it
|
|
77
|
+
unchecked, so a human reviews the approach while redirecting it is still cheap. The
|
|
78
|
+
value is one line saying what to look at.
|
|
79
|
+
|
|
80
|
+
Propose it only on work that is genuinely **irreversible or wide-blast-radius**:
|
|
81
|
+
|
|
82
|
+
- a data migration, or anything that writes to real rows;
|
|
83
|
+
- a destructive or bulk file operation;
|
|
84
|
+
- pulling in a new third-party dependency;
|
|
85
|
+
- an auth, permissions, or trust boundary;
|
|
86
|
+
- a public API or schema contract other people's code depends on.
|
|
87
|
+
|
|
88
|
+
**When in doubt, leave it out.** A pause the user didn't want is worse than no pause at
|
|
89
|
+
all — it trains them to wave past the ones they did want, which is exactly the reflex
|
|
90
|
+
that makes the mechanism useless the one time it matters. Most task lists should carry
|
|
91
|
+
zero or one.
|
|
92
|
+
|
|
93
|
+
**A caller may hand you a standing pause policy** — *"halt before anything that touches
|
|
94
|
+
auth"*, *"stop before any schema migration"* (`tiny-spec-loop` passes one through
|
|
95
|
+
`tiny-spec-run`). Apply it to **this** task list: any task matching the description gets
|
|
96
|
+
a `pause:` naming the policy that put it there, on top of whatever you'd have set
|
|
97
|
+
anyway. A policy that matches nothing here is not an error — say so and move on, rather
|
|
98
|
+
than stretching a task to fit it.
|
|
99
|
+
|
|
73
100
|
Cover **every** part of the approach — together the tasks must deliver all
|
|
74
101
|
`REQ-N`. Don't leave a requirement with no task. Likewise, if `SPEC.md` has a
|
|
75
102
|
`## Design` section, every `D<n>` in it needs at least one task carrying that
|
|
@@ -91,9 +118,10 @@ updated: <ISO date>
|
|
|
91
118
|
# Tasks — <project / feature name>
|
|
92
119
|
|
|
93
120
|
> Executed top to bottom, one at a time. A checked `[x]` task is implemented AND
|
|
94
|
-
> reviewed. `type:`, `req:`, and `
|
|
95
|
-
> ownership contract. A task with `design:` is also graded against that screen's
|
|
96
|
-
> `D<n>` entry and the constitution's Design system.
|
|
121
|
+
> reviewed. `type:`, `req:`, `design:`, and `pause:` are optional; `files:` is a hint,
|
|
122
|
+
> not an ownership contract. A task with `design:` is also graded against that screen's
|
|
123
|
+
> `D<n>` entry and the constitution's Design system. A task with `pause:` halts the
|
|
124
|
+
> build before it runs, so a human looks first.
|
|
97
125
|
|
|
98
126
|
## Tasks
|
|
99
127
|
|
|
@@ -112,6 +140,7 @@ updated: <ISO date>
|
|
|
112
140
|
|
|
113
141
|
- [ ] T3 — <…>
|
|
114
142
|
- acceptance: <observable outcome>
|
|
143
|
+
- pause: <optional; what to check before this runs — irreversible work only>
|
|
115
144
|
- files: <path, path>
|
|
116
145
|
```
|
|
117
146
|
|
|
@@ -138,5 +167,7 @@ When `tasks.md` is `status: stale`:
|
|
|
138
167
|
|
|
139
168
|
## When done
|
|
140
169
|
|
|
141
|
-
Report the task count and
|
|
142
|
-
|
|
170
|
+
Report the task count — and any `pause:` points you set, with their reason, so the user
|
|
171
|
+
can drop one before it fires. Then point them at `tiny-spec-build` (one task at a time,
|
|
172
|
+
reviewing as it goes) or `tiny-spec-loop` (run it through until it's built or something
|
|
173
|
+
stops it).
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|