tiny-spec 0.5.0__tar.gz → 1.0.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (26) hide show
  1. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/PKG-INFO +151 -157
  2. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/README.md +150 -156
  3. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/docs/eval/README.md +6 -6
  4. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/pyproject.toml +7 -9
  5. tiny_spec-1.0.1/tiny-spec-adopt/SKILL.md +213 -0
  6. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/tiny-spec-build/SKILL.md +4 -4
  7. tiny_spec-1.0.1/tiny-spec-create/SKILL.md +343 -0
  8. tiny_spec-1.0.1/tiny-spec-design/SKILL.md +234 -0
  9. tiny_spec-1.0.1/tiny-spec-plan/SKILL.md +290 -0
  10. tiny_spec-1.0.1/tiny-spec-run/SKILL.md +360 -0
  11. tiny_spec-1.0.1/tiny-spec-scope/SKILL.md +200 -0
  12. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/tiny_spec/__init__.py +1 -1
  13. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/tiny_spec/manifest.json +3 -4
  14. tiny_spec-0.5.0/tiny-spec-breakdown/SKILL.md +0 -156
  15. tiny_spec-0.5.0/tiny-spec-create/SKILL.md +0 -454
  16. tiny_spec-0.5.0/tiny-spec-loop/SKILL.md +0 -188
  17. tiny_spec-0.5.0/tiny-spec-plan/SKILL.md +0 -152
  18. tiny_spec-0.5.0/tiny-spec-prd/SKILL.md +0 -160
  19. tiny_spec-0.5.0/tiny-spec-run/SKILL.md +0 -241
  20. tiny_spec-0.5.0/tiny-spec-tasks/SKILL.md +0 -173
  21. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/.gitignore +0 -0
  22. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/LICENSE +0 -0
  23. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/agents/tiny-spec-build-executor.md +0 -0
  24. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/agents/tiny-spec-build-reviewer.md +0 -0
  25. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/examples/todo-cli/README.md +0 -0
  26. {tiny_spec-0.5.0 → tiny_spec-1.0.1}/tiny_spec/cli.py +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: tiny-spec
3
- Version: 0.5.0
3
+ Version: 1.0.1
4
4
  Summary: A tiny, opinionated take on spec-driven development.
5
5
  Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
6
6
  Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
@@ -68,36 +68,40 @@ builds the work one task at a time. Every task is implemented by one agent and
68
68
  graded by an independent reviewer that runs the real tests before anything is
69
69
  committed.
70
70
 
71
- That core is **four skills and two agents**. In front of it sit **two optional
72
- planning on-ramps** — `tiny-spec-prd` (idea → PRD) and `tiny-spec-breakdown`
73
- (PRD → stories) — for when you're starting from an idea rather than a ready ticket.
74
- Two optional routers sit over the top: `tiny-spec-run` walks the three planning steps
75
- in one command, and `tiny-spec-loop` works a whole list of stories — branch, plan,
76
- build, merge, next — until they're built or something stops it. No config file, no
77
- build step.
71
+ That core is **three skills and two agents**. In front of it sit **two front doors** —
72
+ pick the one that matches where you're starting. Over the top sits **one router**,
73
+ `tiny-spec-run`, which drives the chain and, when you ask it to, works a whole list of
74
+ stories through to merged code. No config file, no build step.
78
75
 
79
76
  ```
80
- ┌───────────────── tiny-spec-run ─────────────────┐
81
- │ optional: one command, stops before build │
82
- ▼ ▼
83
- PLANNING (optional on-ramps) EXECUTION (the core loop, one story at a time)
84
- tiny-spec-prd ⇢ tiny-spec-breakdown ⇢ tiny-spec-create → tiny-spec-plan → tiny-spec-tasks → tiny-spec-build
85
- idea → PRD PRD → stories intent design tasks per-task loop
86
- PRD.md BREAKDOWN.md SPEC.md PLAN.md + tasks.md plan → implement → review → commit
87
- constitution
88
- ▲ ▲
89
- └──────────────── tiny-spec-loop ─────────────────┘
90
- optional: per story — branch, run, build, merge,
91
- then the next story. Halts on a wall.
77
+ GREENFIELD BROWNFIELD
78
+ starting from an idea starting from a codebase
79
+ tiny-spec-scope tiny-spec-adopt
80
+ idea → Features → Stories real code → constitution
81
+ BREAKDOWN.md constitution.md
82
+ \ /
83
+ └───────────┬───────────┘
84
+ ▼
85
+ tiny-spec-create → tiny-spec-plan → tiny-spec-build
86
+ intent design + tasks per-task loop
87
+ SPEC.md PLAN.md plan → implement
88
+ tasks.md → review → commit
89
+
90
+ tiny-spec-run one router. Walks the chain and stops before build —
91
+ or, asked to, builds each story and merges it.
92
+ tiny-spec-design optional. Wireframes → tokens + gradeable screens.
92
93
  ```
93
94
 
94
- The two on-ramps are **optional** and stack. Have nothing written down? Run
95
- `tiny-spec-prd` to interview your idea into a `PRD.md`. Have a PRD already? Run
96
- `tiny-spec-breakdown` to carve it into a `BREAKDOWN.md` — a flat list of
97
- Features → Stories with draft acceptance criteria. Have a single known piece of
98
- work? Skip both and start at `tiny-spec-create`. Both on-ramps write a regenerable
99
- file at your project root (not under `.spec/`); `tiny-spec-create` then reads the
100
- breakdown one story at a time.
95
+ **Pick one front door, once per project.** Starting from an idea with no code yet? Run
96
+ `tiny-spec-scope` — it interviews the idea into a `BREAKDOWN.md`, a flat list of
97
+ Features → Stories with draft acceptance criteria. Working in a codebase that already
98
+ exists? Run `tiny-spec-adopt` — it reads your repo and derives the constitution from
99
+ what's actually there: your real lint and test commands, your real layout, your real
100
+ conventions. Have a single known ticket in a project that's already set up? Skip both
101
+ and start at `tiny-spec-create`.
102
+
103
+ Neither builds anything. They set the project up so the three core skills have
104
+ something true to work from, and both write a regenerable file you can edit freely.
101
105
 
102
106
  ## How small
103
107
 
@@ -106,57 +110,41 @@ receipt rather than the adjective:
106
110
 
107
111
  | | skills / commands | agents | config | artifacts per feature |
108
112
  |---|---|---|---|---|
109
- | **tiny-spec** | **8** (4 core + 4 optional) | **2** | **none** | **`SPEC` `PLAN` `tasks`** |
113
+ | **tiny-spec** | **7** (3 core + 1 router + 3 optional) | **2** | **none** | **`SPEC` `PLAN` `tasks`** |
110
114
  | [GitHub Spec Kit](https://github.com/github/spec-kit) | 10 | — | `specify init` | `spec` `plan` `tasks` `checklist` `constitution` `research` `data-model` `contracts/` `quickstart` |
111
115
  | [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
112
116
  | [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
113
117
 
114
- The whole of tiny-spec is **2,079 lines** of markdown across 8 skills and 2 agents.
115
- There is no config file, no build step, no orchestrator, and no CLI to initialize a
116
- project — the skills are the product.
118
+ The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents. No
119
+ config file, no build step, no orchestrator, no CLI to initialize a project — the skills
120
+ are the product.
117
121
 
118
- <sub>Counts taken from each project's own repository on 2026-08-03 by listing its
119
- command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
120
- `templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
121
- OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
122
- personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
123
- `customize.toml` files (35). These projects are all substantially more popular than
124
- tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
125
- bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
122
+ <sub>Counted on 2026-08-03 from each project's own repository, not its marketing copy —
123
+ Spec Kit's `templates/commands/`, OpenSpec's `skills/`, BMAD's `SKILL.md` and
124
+ `customize.toml` files on `main`. All three are far more popular than tiny-spec, and
125
+ being smaller is a bet, not a proof of quality — re-run the counts yourself before
126
+ believing them.</sub>
126
127
 
127
128
  ## Why it's small
128
129
 
129
- Most spec frameworks are generous by default:
130
- many phases, many agents, many generated documents. tiny-spec makes the opposite
131
- bet. Keep one safeguard, drop the rest.
132
-
133
- A green unit test suite is not the same as working software, so the reviewer
134
- exercises acceptance criteria end to end and a final smoke test confirms the whole
135
- spec. That independent review is the safeguard — not the volume of planning
136
- artifacts. One task, one commit, an external reviewer. Nothing gets added unless
137
- it earns its place.
130
+ Most spec frameworks are generous by default: many phases, many agents, many generated
131
+ documents. tiny-spec makes the opposite bet — keep one safeguard, drop the rest. That
132
+ safeguard is the independent reviewer, which exercises acceptance criteria end to end,
133
+ because a green unit test suite is not the same as working software.
138
134
 
139
135
  The case for staying small:
140
136
 
141
- - **Documents are context, and context isn't free.** Generating large `spec.md`,
142
- `plan.md`, `research.md`, and `data-model.md` files costs tokens to write, then
143
- costs context to carry. Every paragraph the agent has to hold is room it no
144
- longer has for your actual code. tiny-spec keeps the spine small — a
145
- constitution and a short memory — and injects only what each task needs.
146
- - **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
147
- assume you're bootstrapping a project from a blank page. Day to day, you pick up
148
- a ticket and change part of a system that already exists. tiny-spec binds to a
149
- ticket, works one at a time, and references your task platform instead of
150
- re-describing the world.
151
- - **Rigid pipelines fight the user.** Mandatory phases and required sections
152
- impose ceremony on work that doesn't need it. tiny-spec's extra structure is
153
- optional by design — add shape where it pays, skip it where it doesn't.
154
- - **More moving parts is more to maintain.** Orchestrators, ownership contracts,
155
- checkpoint matrices, and config files are themselves a system you have to learn
156
- and keep in sync. A few small skills and two agents are not.
157
- - **Generated docs can fake rigor.** A folder of polished planning artifacts looks
158
- like progress, but it isn't proof. The proof is the reviewer running your real
159
- tests before each commit.
137
+ - **Documents are context, and context isn't free.** Every paragraph the agent has to
138
+ hold is room it no longer has for your actual code. tiny-spec keeps the spine small —
139
+ a constitution and a short memory — and injects only what each task needs.
140
+ - **Real work is a ticket inside a system, not a greenfield repo.** tiny-spec binds to a
141
+ ticket and references your task platform instead of re-describing the world.
142
+ - **Rigid pipelines fight the user.** Mandatory phases impose ceremony on work that
143
+ doesn't need it. The extra structure here is optional — add shape where it pays.
144
+ - **More moving parts is more to maintain.** Orchestrators, ownership contracts, and
145
+ config files are themselves a system you have to keep in sync. Two agents are not.
146
+ - **Generated docs can fake rigor.** A folder of polished planning artifacts looks like
147
+ progress, but the proof is the reviewer running your real tests before each commit.
160
148
 
161
149
  That's the whole trade: where larger kits add machinery, tiny-spec adds one
162
150
  independent reviewer and stops.
@@ -166,10 +154,10 @@ independent reviewer and stops.
166
154
  Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
167
155
  runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
168
156
  code with held-out tests the suite never sees, and checks the reviewer's own verdict
169
- against that ground truth — the number that matters being **false-PASS rate**, how
170
- often the gate blesses code that is actually broken.
157
+ against that ground truth. The number that matters is **false-PASS rate** — how often
158
+ the gate blesses code that is actually broken.
171
159
 
172
- Most recent run — 5 benchmark tasks, v0.4.0 (`7810074`), 2026-08-03:
160
+ Most recent run — 5 benchmark tasks on v1.0.0:
173
161
 
174
162
  | | |
175
163
  |---|---|
@@ -178,27 +166,24 @@ Most recent run — 5 benchmark tasks, v0.4.0 (`7810074`), 2026-08-03:
178
166
  | **false-PASS rate** | **0%** — it never called broken code done |
179
167
  | blocker rate | 0% |
180
168
 
181
- The same run a month earlier on `f973c85` scored identically, which is the more
182
- interesting result: the loop is at least stable across versions rather than tuned to a
183
- single commit.
169
+ Two earlier versions scored identically, which is the more interesting result: the loop
170
+ is stable across releases rather than tuned to one commit — including through 1.0, which
171
+ rebuilt the planning chain underneath it.
184
172
 
185
173
  See [`docs/eval/README.md`](docs/eval/README.md) for the method and
186
- [`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the scoring rubric
187
- this repo grades itself against. **Read both skeptically: it is my benchmark, of my own
188
- tool, scored by me.** The rubric says so itself — *"treat the tiny-spec column as
189
- self-assessment to pressure-test, not gospel."* The tasks are small and self-contained,
190
- so the harness measures the loop on well-specified work, not large-codebase performance.
191
- The harness is in the repo precisely so you don't have to take my word for it.
174
+ [`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the rubric. **Read
175
+ both skeptically: it is my benchmark, of my own tool, scored by me**, on small
176
+ self-contained tasks — so it measures the loop on well-specified work, not
177
+ large-codebase performance. The harness is in the repo so you don't have to take my
178
+ word for it.
192
179
 
193
180
  ## New to spec-driven development?
194
181
 
195
- Spec-driven development (SDD) means writing down *what* you want and *why* before
196
- any code exists, then letting that spec drive the build. Instead of prompting an
197
- agent and hoping, you hand it a small, explicit contract — the intent, a design,
198
- and an ordered list of tasks — and it implements against that. The payoff: the
199
- agent stops guessing. It knows what "done" looks like, you can review the plan
200
- before a single line is written, and the result is checked against the spec
201
- rather than vibes. tiny-spec is one small take on that idea.
182
+ Spec-driven development (SDD) means writing down *what* you want and *why* before any
183
+ code exists, then letting that spec drive the build. Instead of prompting an agent and
184
+ hoping, you hand it a small, explicit contract — the intent, a design, and an ordered
185
+ list of tasks. The payoff: the agent stops guessing, you review the plan before a line
186
+ is written, and the result is checked against the spec rather than vibes.
202
187
 
203
188
  ## Quickstart
204
189
 
@@ -212,41 +197,41 @@ uvx tiny-spec install
212
197
  Restart Claude Code so it picks up the new skills, then run the flow in your project:
213
198
 
214
199
  ```
215
- /tiny-spec-prd # optional: interview a rough idea into a PRD (PRD.md)
216
- /tiny-spec-breakdown # optional: carve a PRD + wireframes into stories (BREAKDOWN.md)
200
+ /tiny-spec-scope # starting from an idea: interview it into stories (BREAKDOWN.md)
201
+ /tiny-spec-adopt # starting from a codebase: derive the constitution from real code
217
202
  /tiny-spec-create # capture intent and requirements (binds a ticket, optional)
218
- /tiny-spec-plan # turn the spec into a design and harden the constitution
219
- /tiny-spec-tasks # slice the plan into an ordered checklist
203
+ /tiny-spec-plan # design it, harden the constitution, slice the task list
220
204
  /tiny-spec-build # build each task: implement, review, commit
221
205
  ```
222
206
 
223
- Or collapse the three planning steps into one and go straight to building:
207
+ Run one of the two front doors once per project, then `create → plan → build` per
208
+ ticket. Or let the router drive:
224
209
 
225
210
  ```
226
- /tiny-spec-run # create → plan → tasks in one pass; stops before build
211
+ /tiny-spec-run # walks the chain, reconciling anything stale; stops before build
227
212
  /tiny-spec-build # build each task: implement, review, commit
228
213
  ```
229
214
 
230
- `tiny-spec-run` resolves where your ticket stands and invokes whichever of
231
- `create` / `plan` / `tasks` moves it forward, reconciling anything stale first. It
232
- **stops before `tiny-spec-build`** — that's where you actually review the work — and
233
- it writes nothing itself, it only delegates.
215
+ `tiny-spec-run` resolves where your work stands and invokes whichever skill moves it
216
+ forward. It writes nothing itself, it only delegates. By default it **stops before
217
+ `tiny-spec-build`** — that's where you actually review the work.
234
218
 
235
- Or hand it a whole list of stories:
219
+ Ask it to build and it goes all the way instead:
236
220
 
237
221
  ```
238
- /tiny-spec-loop # per story: branch → plan → build → merge → next
222
+ /tiny-spec-run build the backlog # per story: branch → plan → build → merge → next
239
223
  ```
240
224
 
241
- `tiny-spec-loop` reads your `BREAKDOWN.md` (or a list you paste) and works the stories
242
- one after another, merging each finished branch into `main` locally before starting the
243
- next — until it's done or reaches a **terminal state** it names out loud. See
244
- [Working a whole list](#working-a-whole-list--tiny-spec-loop).
225
+ It reads your `BREAKDOWN.md` (or a list you paste) and works the stories one after
226
+ another, merging each finished branch into `main` locally before starting the next —
227
+ until it's done or reaches a **terminal state** it names out loud. See
228
+ [Working a whole list](#working-a-whole-list). Which of the two it does is decided from
229
+ your opening request and **never changes mid-run** — a run you started as "get it
230
+ ready" will not talk itself into building.
245
231
 
246
- **Building from a mockup?** There is no design flag to pass. Commit your exports to
247
- `design/` before the run and `tiny-spec-create` — whether you invoke it directly or
248
- reach it through `tiny-spec-run` — reads what's there and asks which screens this
249
- ticket covers. [Designs, if you have them](#designs-if-you-have-them) walks through it.
232
+ **Building from a mockup?** Run `/tiny-spec-design` after `create`. Commit your exports
233
+ to `design/` first; it reads what's there and asks which screens this ticket covers.
234
+ [Designs, if you have them](#designs-if-you-have-them) walks through it.
250
235
 
251
236
  Re-run `install` any time to update; `tiny-spec uninstall` removes only what it
252
237
  installed. Each skill is copied (not symlinked) so every install is
@@ -263,7 +248,7 @@ git clone https://github.com/GrayMa77er/tiny-spec.git
263
248
  cd tiny-spec
264
249
 
265
250
  mkdir -p "$HOME/.claude/skills" "$HOME/.claude/agents"
266
- for s in tiny-spec-prd tiny-spec-breakdown tiny-spec-run tiny-spec-loop tiny-spec-create tiny-spec-plan tiny-spec-tasks tiny-spec-build; do
251
+ for s in tiny-spec-scope tiny-spec-adopt tiny-spec-run tiny-spec-create tiny-spec-design tiny-spec-plan tiny-spec-build; do
267
252
  cp -R "$s" "$HOME/.claude/skills/$s"
268
253
  done
269
254
  cp agents/*.md "$HOME/.claude/agents/"
@@ -283,14 +268,22 @@ The tests pass; you can clone it and run the gate yourself.
283
268
 
284
269
  ## How it works
285
270
 
286
- The constitution (`constitution.md`) is the spine. `tiny-spec-create` seeds it from a
287
- short interview, `tiny-spec-plan` hardens it with concrete engineering rules, and
288
- `tiny-spec-build` injects it whole into every task. It holds your style, standards,
289
- invariants, definition of done, and verification commands.
290
-
291
- Because it is project-wide it can also go missing — deleted, or never committed —
292
- while your specs survive. Re-running `tiny-spec-create` then repairs it: it rebuilds
293
- the constitution from whatever is already written down and marks completed tasks
271
+ The constitution (`constitution.md`) is the spine. `tiny-spec-adopt` derives it from
272
+ your codebase, or `tiny-spec-create` seeds it from a short interview; `tiny-spec-plan`
273
+ hardens it with concrete engineering rules, and `tiny-spec-build` injects it whole into
274
+ every task. It holds your style, standards, invariants, definition of done, and
275
+ verification commands.
276
+
277
+ **On an existing codebase, deriving beats asking.** Your test command, layout, and
278
+ conventions are already written down — in `package.json`, in CI, in your linter config,
279
+ in the shape of the tree. `tiny-spec-adopt` reads them, marks each section as declared
280
+ or inferred so you know which parts to distrust, then *runs* the verification commands
281
+ it derived and reports which went green. A gate that has never been run turns every
282
+ future review into theatre.
283
+
284
+ Because it is project-wide the constitution can also go missing — deleted, or never
285
+ committed — while your specs survive. Re-running `tiny-spec-adopt` (or
286
+ `tiny-spec-create`, on a project with no code yet) repairs it, and marks completed tasks
294
287
  stale, since they were reviewed against a document that wasn't there.
295
288
 
296
289
  ### Designs, if you have them
@@ -302,11 +295,11 @@ on. Skip all of it for a CLI or a library; the constitution simply has no design
302
295
  section.
303
296
 
304
297
  **Designs enter by convention, not by argument.** No skill takes a design flag or a
305
- path parameter. `tiny-spec-create` reads every file in `design/` at your project root —
306
- directly, or when `tiny-spec-run` reaches it — and asks which screens this ticket
307
- covers; you can also just hand it paths during the interview. Change an export later and
308
- re-running `create` (or `run`) re-hashes it, marking the spec stale exactly like editing
309
- a requirement.
298
+ path parameter. `tiny-spec-design` reads every file in `design/` at your project root
299
+ and asks which screens this ticket covers; you can also just hand it paths. Change an
300
+ export later and `tiny-spec-run` notices — it re-hashes every anchored export on its way
301
+ down the chain, and routes a mismatch back to `tiny-spec-design`, marking the spec stale
302
+ exactly like editing a requirement.
310
303
 
311
304
  **1. Commit your exports.** Any format an agent can read — a Figma export, an HTML
312
305
  mockup, an Excalidraw file, a photo of a whiteboard.
@@ -318,7 +311,7 @@ your-project/
318
311
  dashboard.png
319
312
  ```
320
313
 
321
- **2. `tiny-spec-create` looks at them and proposes a design system.** Actually looks —
314
+ **2. `tiny-spec-design` looks at them and proposes a design system.** Actually looks —
322
315
  they are read as images. It infers *one* coherent scale across all of them rather than
323
316
  measuring each screen separately, tells you what it rounded ("your wireframes had 19px
324
317
  and 21px — proposing `space.5`=20px for both"), and on your approval writes it into
@@ -387,13 +380,12 @@ FINDINGS:
387
380
 
388
381
  Numbers first, and never a screenshot diff — pixel comparison goes flaky on font
389
382
  antialiasing and teams end up muting it. But no measurement catches an element that is
390
- present, on-token, and still not on screen — occluded, clipped, truncated, or the same
383
+ present, on-token, and still not on screen: occluded, clipped, truncated, or the same
391
384
  color as what's behind it. So the last step reads a screenshot of each state next to
392
385
  your export and grades presence, legibility, and correspondence. **Where a number
393
- already settled the question the eye may only flag** — on-scale-but-cramped is never a
394
- fail — which keeps the two halves from contradicting each other. Measurable or visible
395
- violations fail; taste comes back as `flag:` notes so a bounded fix loop can't thrash. A
396
- task with no `design:` tag is graded exactly as before.
386
+ already settled the question the eye may only flag**, which keeps the two halves from
387
+ contradicting each other — measurable or visible violations fail, taste comes back as
388
+ `flag:` notes. A task with no `design:` tag is graded exactly as before.
397
389
 
398
390
  <details>
399
391
  <summary>The <code>visual:</code> command (you write this once)</summary>
@@ -437,14 +429,12 @@ await b.close();
437
429
  ```
438
430
 
439
431
  The missing-selector branch **must print something** — that is what turns a renamed
440
- element into a failure instead of a silent skip. `top` is what the layout-order check
441
- reads, and `opacity`/`visibility` are worth printing because they turn the cheapest kind
442
- of invisible element into a numeric failure. The `SCREENSHOT` line is what the reviewer
443
- reads back to look at the render and catch the rest — occlusion, clipping, truncation,
444
- same-color-on-same-color — none of which any single property reports. Drop it and the
445
- numeric half still gates exactly as before, with the reviewer reporting `judge: not run`
446
- rather than quietly skipping it. Without a `visual:` command at all, a task tagged
447
- `design:` raises a blocker rather than passing quietly.
432
+ element into a failure instead of a silent skip. `top` feeds the layout-order check, and
433
+ `opacity`/`visibility` turn the cheapest kind of invisible element into a numeric
434
+ failure. The `SCREENSHOT` line arms the judge; drop it and the numeric half still gates
435
+ exactly as before, with the reviewer reporting `judge: not run` rather than quietly
436
+ skipping it. Without a `visual:` command at all, a task tagged `design:` raises a
437
+ blocker rather than passing quietly.
448
438
 
449
439
  </details>
450
440
 
@@ -496,13 +486,14 @@ and logs a blocker instead of hacking around it. You fix the gap upstream in
496
486
  `tiny-spec-plan` or `tiny-spec-create`, then resume. Work runs one ticket at a time and
497
487
  resumes from the checklist state.
498
488
 
499
- ### Working a whole list — `tiny-spec-loop`
489
+ ### Working a whole list
500
490
 
501
- `/tiny-spec-loop` takes a list of stories and works them one after another. Per story
502
- it does the same four moves:
491
+ Ask `/tiny-spec-run` to build — "build the backlog", "work through the breakdown",
492
+ "spec it out and build it" — and it takes a list of stories and works them one after
493
+ another. Per story it does the same four moves:
503
494
 
504
495
  ```
505
- cut a branch from main → tiny-spec-run → tiny-spec-build → merge back to main
496
+ cut a branch from main → walk the chain → tiny-spec-build → merge back to main
506
497
  ```
507
498
 
508
499
  Then the next story. Each branch is cut **fresh from main**, so story 3 sees stories 1
@@ -525,10 +516,9 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
525
516
  | `fork` | a real either/or the plan doesn't answer |
526
517
  | `conflict` | a story's branch wouldn't merge cleanly |
527
518
 
528
- **Only `done` means the work is built** — and in a story loop that means *all* of them.
529
- A run that stopped at story 2 of 7 reporting "done" is the single thing autonomous loops
530
- get wrong most often, so the state is always named, along with what got merged and how
531
- many stories are still untouched.
519
+ **Only `done` means the work is built** — and in a story run, that means *all* of them.
520
+ Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
521
+ often, so the state is always named alongside what merged and what's still untouched.
532
522
 
533
523
  **A halt stops the whole run, not just that story.** Later stories in a list you wrote
534
524
  top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
@@ -543,7 +533,7 @@ the build halts *before* running it:
543
533
  - pause: confirm the snapshot is current before this touches real rows
544
534
  ```
545
535
 
546
- `tiny-spec-tasks` proposes these for genuinely irreversible work — migrations,
536
+ `tiny-spec-plan` proposes these for genuinely irreversible work — migrations,
547
537
  destructive file operations, a new dependency, an auth boundary, a public API contract.
548
538
  You can also give the run a standing policy up front ("halt before anything that touches
549
539
  auth") and it gets applied as each story's tasks are sliced.
@@ -555,22 +545,26 @@ never opens a PR. Merges are local, so a bad run is one `git reset` away; publis
555
545
  stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
556
546
  the undo command rather than running it.
557
547
 
558
- **Walk away and come back.** Nothing is written down to track progress — it's derived:
559
- a story whose ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked
560
- task is in progress, no directory means not started. So re-running `/tiny-spec-loop`
561
- tomorrow in a fresh session picks up exactly where it stopped. No run-state file, no
562
- lock, and no budget to configure: the story list *is* the budget.
548
+ **Walk away and come back.** Progress isn't written down, it's derived: a story whose
549
+ ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
550
+ progress, no directory means not started. Ask again tomorrow in a fresh session and it
551
+ picks up where it stopped. No run-state file, no lock, no budget to configure — the
552
+ story list *is* the budget.
563
553
 
564
554
  **It never fixes a blocker for you.** A blocker means one of your documents is wrong,
565
- and a loop allowed to rewrite the requirement its own task just failed would be grading
555
+ and a run allowed to rewrite the requirement its own task just failed would be grading
566
556
  its own homework. It reports and stops; the upstream edit is yours.
567
557
 
558
+ **And it never decides to build on its own.** Whether a run stops before the build or
559
+ goes all the way is fixed from your opening request and never revised. Build is your
560
+ review gate, so promoting itself into it is the one thing a router doesn't get to do.
561
+
568
562
  ## Project layout
569
563
 
570
564
  Each skill is one self-contained `SKILL.md`, with every document skeleton inline in
571
- it — no companion template files, no absolute paths, and no shared parent required
572
- at runtime, so a skill works wherever you drop it. (It also means a run never stops
573
- to ask permission to read a template out of your Claude config directory.)
565
+ it — no companion template files, no absolute paths, no shared parent required at
566
+ runtime. A skill works wherever you drop it, and a run never stops to ask permission to
567
+ read a template out of your Claude config directory.
574
568
 
575
569
  tiny-spec creates a `.spec/` directory in your project root, never inside a skill.
576
570
  It is namespaced per ticket, with a shared spine at the root:
@@ -583,12 +577,12 @@ It is namespaced per ticket, with a shared spine at the root:
583
577
  SPEC.md PLAN.md tasks.md decisions.md
584
578
  ```
585
579
 
586
- `PRD.md`, `BREAKDOWN.md`, and `design/` sit at your project root rather than inside
587
- `.spec/`, because they are yours: the first two are regenerable pre-spec planning
588
- files, and the design exports are project files no skill ever writes.
580
+ `BREAKDOWN.md` and `design/` sit at your project root rather than inside `.spec/`,
581
+ because they are yours: the breakdown is a regenerable pre-spec planning file, and the
582
+ design exports are project files no skill ever writes.
589
583
 
590
- `CONTRACTS.md` documents the formats for maintainers. The skills do not read it at
591
- runtime; each is self-sufficient.
584
+ Every format lives inline in the skill that writes it — there is no separate contract
585
+ document to keep in sync, and no skill reads anything but its own `SKILL.md` at runtime.
592
586
 
593
587
  ## Integrations
594
588