tiny-spec 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,25 @@
1
+ ### macOS
2
+ # Finder metadata
3
+ .DS_Store
4
+
5
+ # Thumbnails
6
+ ._*
7
+
8
+ # Custom folder icons
9
+ Icon
10
+
11
+ # Volume root files
12
+ .DocumentRevisions-V100
13
+ .fseventsd
14
+ .Spotlight-V100
15
+ .TemporaryItems
16
+ .Trashes
17
+ .VolumeIcon.icns
18
+ .com.apple.timemachine.donotpresent
19
+ ### Python
20
+ __pycache__/
21
+ *.py[cod]
22
+ build/
23
+ dist/
24
+ *.egg-info/
25
+ .venv/
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Snir Orlanczyk
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,630 @@
1
+ Metadata-Version: 2.5
2
+ Name: tiny-spec
3
+ Version: 1.0.0
4
+ Summary: A tiny, opinionated take on spec-driven development.
5
+ Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
6
+ Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
7
+ Project-URL: Issues, https://github.com/GrayMa77er/tiny-spec/issues
8
+ Author: Snir Orlanczyk
9
+ License: MIT License
10
+
11
+ Copyright (c) 2026 Snir Orlanczyk
12
+
13
+ Permission is hereby granted, free of charge, to any person obtaining a copy
14
+ of this software and associated documentation files (the "Software"), to deal
15
+ in the Software without restriction, including without limitation the rights
16
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
17
+ copies of the Software, and to permit persons to whom the Software is
18
+ furnished to do so, subject to the following conditions:
19
+
20
+ The above copyright notice and this permission notice shall be included in all
21
+ copies or substantial portions of the Software.
22
+
23
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
24
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
25
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
26
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
27
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
28
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
29
+ SOFTWARE.
30
+ License-File: LICENSE
31
+ Keywords: ai-agents,claude,claude-code,skills,spec-driven-development
32
+ Classifier: Environment :: Console
33
+ Classifier: License :: OSI Approved :: MIT License
34
+ Classifier: Programming Language :: Python :: 3
35
+ Requires-Python: >=3.9
36
+ Description-Content-Type: text/markdown
37
+
38
+ <p align="center">
39
+ <img src="images/logo.png" alt="tiny-spec" width="200">
40
+ </p>
41
+
42
+ <h1 align="center">tiny-spec</h1>
43
+
44
+ <p align="center">A tiny, opinionated take on spec-driven development.</p>
45
+
46
+ <p align="center">
47
+ <a href="https://pypi.org/project/tiny-spec/"><img src="https://img.shields.io/pypi/v/tiny-spec.svg?color=d97757" alt="PyPI version"></a>
48
+ <a href="https://pypi.org/project/tiny-spec/"><img src="https://img.shields.io/pypi/dm/tiny-spec.svg?color=d97757" alt="PyPI downloads"></a>
49
+ <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue.svg" alt="License: MIT"></a>
50
+ <a href="https://docs.claude.com/en/docs/claude-code/overview"><img src="https://img.shields.io/badge/Claude%20Code-skills-d97757.svg" alt="Claude Code"></a>
51
+ <a href="https://github.com/GrayMa77er/tiny-spec/stargazers"><img src="https://img.shields.io/github/stars/GrayMa77er/tiny-spec?color=d97757" alt="Stars"></a>
52
+ </p>
53
+
54
+ <p align="center">
55
+ <code>uvx tiny-spec install</code>
56
+ </p>
57
+
58
+ <!-- DEMO GIF SLOT — images/demo.gif
59
+ Record the FAIL -> fix -> PASS loop, then uncomment:
60
+ <p align="center">
61
+ <img src="images/demo.gif" alt="tiny-spec build loop: the reviewer fails a task, the executor fixes it, the task commits" width="800">
62
+ </p>
63
+ -->
64
+
65
+ tiny-spec is a four-step workflow for Claude Code that turns a ticket into shipped,
66
+ reviewed code. You write the intent, it produces a design, a task list, and then
67
+ builds the work one task at a time. Every task is implemented by one agent and
68
+ graded by an independent reviewer that runs the real tests before anything is
69
+ committed.
70
+
71
+ That core is **three skills and two agents**. In front of it sit **two front doors** —
72
+ pick the one that matches where you're starting. Over the top sits **one router**,
73
+ `tiny-spec-run`, which drives the chain and, when you ask it to, works a whole list of
74
+ stories through to merged code. No config file, no build step.
75
+
76
+ ```
77
+ GREENFIELD BROWNFIELD
78
+ starting from an idea starting from a codebase
79
+ tiny-spec-scope tiny-spec-adopt
80
+ idea → Features → Stories real code → constitution
81
+ BREAKDOWN.md constitution.md
82
+ \ /
83
+ └───────────┬───────────┘
84
+ ▼
85
+ tiny-spec-create → tiny-spec-plan → tiny-spec-build
86
+ intent design + tasks per-task loop
87
+ SPEC.md PLAN.md plan → implement
88
+ tasks.md → review → commit
89
+
90
+ tiny-spec-run one router. Walks the chain and stops before build —
91
+ or, asked to, builds each story and merges it.
92
+ tiny-spec-design optional. Wireframes → tokens + gradeable screens.
93
+ ```
94
+
95
+ **Pick one front door, once per project.** Starting from an idea with no code yet? Run
96
+ `tiny-spec-scope` — it interviews the idea into a `BREAKDOWN.md`, a flat list of
97
+ Features → Stories with draft acceptance criteria. Working in a codebase that already
98
+ exists? Run `tiny-spec-adopt` — it reads your repo and derives the constitution from
99
+ what's actually there: your real lint and test commands, your real layout, your real
100
+ conventions. Have a single known ticket in a project that's already set up? Skip both
101
+ and start at `tiny-spec-create`.
102
+
103
+ Neither builds anything. They set the project up so the three core skills have
104
+ something true to work from, and both write a regenerable file you can edit freely.
105
+
106
+ ## How small
107
+
108
+ Every other kit in this space is bigger. That is the whole pitch, so here is the
109
+ receipt rather than the adjective:
110
+
111
+ | | skills / commands | agents | config | artifacts per feature |
112
+ |---|---|---|---|---|
113
+ | **tiny-spec** | **7** (3 core + 1 router + 3 optional) | **2** | **none** | **`SPEC` `PLAN` `tasks`** |
114
+ | [GitHub Spec Kit](https://github.com/github/spec-kit) | 10 | — | `specify init` | `spec` `plan` `tasks` `checklist` `constitution` `research` `data-model` `contracts/` `quickstart` |
115
+ | [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
116
+ | [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
117
+
118
+ The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents.
119
+ There is no config file, no build step, no orchestrator, and no CLI to initialize a
120
+ project — the skills are the product. (That line count went *up* slightly in 1.0 while
121
+ the skill count went down: `tiny-spec-adopt` is genuinely new capability, and merging
122
+ four skills into two removed commands to learn rather than words to read.)
123
+
124
+ <sub>Counts taken from each project's own repository on 2026-08-03 by listing its
125
+ command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
126
+ `templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
127
+ OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
128
+ personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
129
+ `customize.toml` files (35). These projects are all substantially more popular than
130
+ tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
131
+ bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
132
+
133
+ ## Why it's small
134
+
135
+ Most spec frameworks are generous by default:
136
+ many phases, many agents, many generated documents. tiny-spec makes the opposite
137
+ bet. Keep one safeguard, drop the rest.
138
+
139
+ A green unit test suite is not the same as working software, so the reviewer
140
+ exercises acceptance criteria end to end and a final smoke test confirms the whole
141
+ spec. That independent review is the safeguard — not the volume of planning
142
+ artifacts. One task, one commit, an external reviewer. Nothing gets added unless
143
+ it earns its place.
144
+
145
+ The case for staying small:
146
+
147
+ - **Documents are context, and context isn't free.** Generating large `spec.md`,
148
+ `plan.md`, `research.md`, and `data-model.md` files costs tokens to write, then
149
+ costs context to carry. Every paragraph the agent has to hold is room it no
150
+ longer has for your actual code. tiny-spec keeps the spine small — a
151
+ constitution and a short memory — and injects only what each task needs.
152
+ - **Real work is a ticket inside a system, not a greenfield repo.** Bigger kits
153
+ assume you're bootstrapping a project from a blank page. Day to day, you pick up
154
+ a ticket and change part of a system that already exists. tiny-spec binds to a
155
+ ticket, works one at a time, and references your task platform instead of
156
+ re-describing the world.
157
+ - **Rigid pipelines fight the user.** Mandatory phases and required sections
158
+ impose ceremony on work that doesn't need it. tiny-spec's extra structure is
159
+ optional by design — add shape where it pays, skip it where it doesn't.
160
+ - **More moving parts is more to maintain.** Orchestrators, ownership contracts,
161
+ checkpoint matrices, and config files are themselves a system you have to learn
162
+ and keep in sync. A few small skills and two agents are not.
163
+ - **Generated docs can fake rigor.** A folder of polished planning artifacts looks
164
+ like progress, but it isn't proof. The proof is the reviewer running your real
165
+ tests before each commit.
166
+
167
+ That's the whole trade: where larger kits add machinery, tiny-spec adds one
168
+ independent reviewer and stops.
169
+
170
+ ### Does the trade hold up?
171
+
172
+ Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
173
+ runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
174
+ code with held-out tests the suite never sees, and checks the reviewer's own verdict
175
+ against that ground truth — the number that matters being **false-PASS rate**, how
176
+ often the gate blesses code that is actually broken.
177
+
178
+ Most recent run — 5 benchmark tasks, v1.0.0 (`b18e32a`), 2026-09-07:
179
+
180
+ | | |
181
+ |---|---|
182
+ | held-out pass rate | **100%** (5/5) — the code works, judged by tests it never saw |
183
+ | suite/truth agreement | **100%** — the reviewer's verdict matched the held-out grader every time |
184
+ | **false-PASS rate** | **0%** — it never called broken code done |
185
+ | blocker rate | 0% |
186
+
187
+ Two earlier runs — `7810074` (2026-08-03) and `f973c85` a month before it — scored
188
+ identically, which is the more interesting result: the loop is stable across versions
189
+ rather than tuned to a single commit. The 1.0 run matters most of the three, because
190
+ 1.0 rebuilt the planning chain around it — `tiny-spec-plan` now writes the task list
191
+ that `tiny-spec-build` executes — and the gate came through unchanged.
192
+
193
+ See [`docs/eval/README.md`](docs/eval/README.md) for the method and
194
+ [`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the scoring rubric
195
+ this repo grades itself against. **Read both skeptically: it is my benchmark, of my own
196
+ tool, scored by me.** The rubric says so itself — *"treat the tiny-spec column as
197
+ self-assessment to pressure-test, not gospel."* The tasks are small and self-contained,
198
+ so the harness measures the loop on well-specified work, not large-codebase performance.
199
+ The harness is in the repo precisely so you don't have to take my word for it.
200
+
201
+ ## New to spec-driven development?
202
+
203
+ Spec-driven development (SDD) means writing down *what* you want and *why* before
204
+ any code exists, then letting that spec drive the build. Instead of prompting an
205
+ agent and hoping, you hand it a small, explicit contract — the intent, a design,
206
+ and an ordered list of tasks — and it implements against that. The payoff: the
207
+ agent stops guessing. It knows what "done" looks like, you can review the plan
208
+ before a single line is written, and the result is checked against the spec
209
+ rather than vibes. tiny-spec is one small take on that idea.
210
+
211
+ ## Quickstart
212
+
213
+ Install the skills and agents into your Claude Code config with
214
+ [uv](https://docs.astral.sh/uv/):
215
+
216
+ ```sh
217
+ uvx tiny-spec install
218
+ ```
219
+
220
+ Restart Claude Code so it picks up the new skills, then run the flow in your project:
221
+
222
+ ```
223
+ /tiny-spec-scope # starting from an idea: interview it into stories (BREAKDOWN.md)
224
+ /tiny-spec-adopt # starting from a codebase: derive the constitution from real code
225
+ /tiny-spec-create # capture intent and requirements (binds a ticket, optional)
226
+ /tiny-spec-plan # design it, harden the constitution, slice the task list
227
+ /tiny-spec-build # build each task: implement, review, commit
228
+ ```
229
+
230
+ Run one of the two front doors once per project, then `create → plan → build` per
231
+ ticket. Or let the router drive:
232
+
233
+ ```
234
+ /tiny-spec-run # walks the chain, reconciling anything stale; stops before build
235
+ /tiny-spec-build # build each task: implement, review, commit
236
+ ```
237
+
238
+ `tiny-spec-run` resolves where your work stands and invokes whichever skill moves it
239
+ forward. It writes nothing itself, it only delegates. By default it **stops before
240
+ `tiny-spec-build`** — that's where you actually review the work.
241
+
242
+ Ask it to build and it goes all the way instead:
243
+
244
+ ```
245
+ /tiny-spec-run build the backlog # per story: branch → plan → build → merge → next
246
+ ```
247
+
248
+ It reads your `BREAKDOWN.md` (or a list you paste) and works the stories one after
249
+ another, merging each finished branch into `main` locally before starting the next —
250
+ until it's done or reaches a **terminal state** it names out loud. See
251
+ [Working a whole list](#working-a-whole-list). Which of the two it does is decided from
252
+ your opening request and **never changes mid-run** — a run you started as "get it
253
+ ready" will not talk itself into building.
254
+
255
+ **Building from a mockup?** Run `/tiny-spec-design` after `create`. Commit your exports
256
+ to `design/` first; it reads what's there and asks which screens this ticket covers.
257
+ [Designs, if you have them](#designs-if-you-have-them) walks through it.
258
+
259
+ Re-run `install` any time to update; `tiny-spec uninstall` removes only what it
260
+ installed. Each skill is copied (not symlinked) so every install is
261
+ self-contained.
262
+
263
+ <details>
264
+ <summary>Manual install (no uv)</summary>
265
+
266
+ The skills and agents are plain markdown — copy them in by hand. Claude Code
267
+ loads skills from `~/.claude/skills/` and agents from `~/.claude/agents/`:
268
+
269
+ ```sh
270
+ git clone https://github.com/GrayMa77er/tiny-spec.git
271
+ cd tiny-spec
272
+
273
+ mkdir -p "$HOME/.claude/skills" "$HOME/.claude/agents"
274
+ for s in tiny-spec-scope tiny-spec-adopt tiny-spec-run tiny-spec-create tiny-spec-design tiny-spec-plan tiny-spec-build; do
275
+ cp -R "$s" "$HOME/.claude/skills/$s"
276
+ done
277
+ cp agents/*.md "$HOME/.claude/agents/"
278
+ ```
279
+
280
+ If a skill name collides with one you already have, rename these before copying,
281
+ or install one set at a time.
282
+
283
+ </details>
284
+
285
+ ### See a finished run first
286
+
287
+ [`examples/todo-cli/`](examples/todo-cli/) is a real run of the flow on one small
288
+ ticket, committed verbatim — the `TICKET.md` that went in, the `SPEC.md`, `PLAN.md`,
289
+ `tasks.md` and `constitution.md` the suite wrote, and the code and tests it produced.
290
+ The tests pass; you can clone it and run the gate yourself.
291
+
292
+ ## How it works
293
+
294
+ The constitution (`constitution.md`) is the spine. `tiny-spec-adopt` derives it from
295
+ your codebase, or `tiny-spec-create` seeds it from a short interview; `tiny-spec-plan`
296
+ hardens it with concrete engineering rules, and `tiny-spec-build` injects it whole into
297
+ every task. It holds your style, standards, invariants, definition of done, and
298
+ verification commands.
299
+
300
+ **On an existing codebase, deriving beats asking.** Your test command, your layout, and
301
+ your conventions are already written down — in `package.json`, in CI, in your linter
302
+ config, in the shape of the tree. `tiny-spec-adopt` reads them and marks each section
303
+ as declared or inferred, so you know which parts to distrust. It then *runs* the
304
+ verification commands it derived and reports which went green: a gate that has never
305
+ been run is the most dangerous thing in the file, because it turns every future review
306
+ into theatre.
307
+
308
+ Because it is project-wide the constitution can also go missing — deleted, or never
309
+ committed — while your specs survive. Re-running `tiny-spec-adopt` (or
310
+ `tiny-spec-create`, on a project with no code yet) repairs it, and marks completed tasks
311
+ stale, since they were reviewed against a document that wasn't there.
312
+
313
+ ### Designs, if you have them
314
+
315
+ Wireframes usually get read once and forgotten. tiny-spec turns them into two durable
316
+ things — a project-wide token system in the constitution, and a per-screen entry in
317
+ the spec — so "does this look right?" becomes something the reviewer can fail a task
318
+ on. Skip all of it for a CLI or a library; the constitution simply has no design
319
+ section.
320
+
321
+ **Designs enter by convention, not by argument.** No skill takes a design flag or a
322
+ path parameter. `tiny-spec-design` reads every file in `design/` at your project root
323
+ and asks which screens this ticket covers; you can also just hand it paths. Change an
324
+ export later and `tiny-spec-run` notices — it re-hashes every anchored export on its way
325
+ down the chain, and routes a mismatch back to `tiny-spec-design`, marking the spec stale
326
+ exactly like editing a requirement.
327
+
328
+ **1. Commit your exports.** Any format an agent can read — a Figma export, an HTML
329
+ mockup, an Excalidraw file, a photo of a whiteboard.
330
+
331
+ ```
332
+ your-project/
333
+ design/
334
+ signup.png
335
+ dashboard.png
336
+ ```
337
+
338
+ **2. `tiny-spec-design` looks at them and proposes a design system.** Actually looks —
339
+ they are read as images. It infers *one* coherent scale across all of them rather than
340
+ measuring each screen separately, tells you what it rounded ("your wireframes had 19px
341
+ and 21px — proposing `space.5`=20px for both"), and on your approval writes it into
342
+ `.spec/constitution.md`:
343
+
344
+ ```markdown
345
+ ## Design system
346
+ - color: `color.surface.base` #FFFFFF · `color.text.default` #111111
347
+ `color.text.muted` #6B7280 · `color.text.danger` #B91C1C
348
+ - space: `space.1`=4px · `space.2`=8px · `space.4`=16px · `space.6`=24px
349
+ - type: `type.heading.lg` 24px/600/1.25 · `type.body` 16px/400/1.5
350
+ `type.caption` 13px/400/1.4
351
+ - states: every interactive surface defines default, focus, disabled, loading, empty, error
352
+ ```
353
+
354
+ Since it is project-wide, every screen from here on speaks this vocabulary — and a
355
+ redesign edits this one table instead of every file.
356
+
357
+ **3. Each screen becomes a `D<n>` in `SPEC.md`**, written in those token names:
358
+
359
+ ```markdown
360
+ ## Design
361
+
362
+ - D1 — Signup form
363
+ - source: figma.com/file/abc#node-12:34 (view-only)
364
+ - export: design/signup.png
365
+ - sha256: d21d6330648c504edeb924b5398bf7fb6485d3a4c1907e43d800970f39622a1d
366
+ - layout: single centered column, max 420px; title → field → error → submit
367
+ - elements:
368
+ - card `[data-testid="signup-card"]` → space.6 padding, color.surface.base
369
+ - title `[data-testid="signup-title"]` → type.heading.lg, space.4 below
370
+ - field `[data-testid="signup-email"]` → type.body, space.2 below
371
+ - error `[data-testid="signup-error"]` → type.caption, color.text.danger
372
+ - states: error (caption under the field), loading (spinner replaces button label)
373
+ ```
374
+
375
+ The selectors are a **contract**, not a hint — the reviewer measures exactly these, so
376
+ your markup carries them verbatim. Prefer test ids over CSS classes: classes get
377
+ renamed by refactors and mangled by CSS-in-JS, and a selector that silently stops
378
+ matching is the failure this exists to prevent.
379
+
380
+ **4. Tag the tasks that build the surface** — and only those, not the API call behind
381
+ it. This is your blast radius:
382
+
383
+ ```
384
+ - [ ] T4 — Build the signup form
385
+ - acceptance: submitting a valid email advances to the verify step
386
+ - design: D1
387
+ ```
388
+
389
+ **5. The reviewer measures it, then looks at it, and fails on both:**
390
+
391
+ ```
392
+ DESIGN: D1 — Signup form
393
+ [data-testid="signup-card"] padding 19px — not on the space.* scale (4/8/16/24)
394
+ [data-testid="signup-email"] MISSING FROM DOM — never built, or renamed
395
+ [data-testid="signup-error"] rgb(204,0,0) — color.text.danger is #B91C1C
396
+ state "loading" never renders: button label stays "Continue", no spinner
397
+ judge: read states default, error — compared against design/signup.png
398
+ state "error": signup-error sits behind the card — every token right, and the
399
+ user sees nothing where the export shows the red caption
400
+ state "default": submit label "Create account" truncates to "Create acco…"
401
+ FINDINGS:
402
+ - flag: title/field gap feels tight (on-scale — does not fail the task)
403
+ ```
404
+
405
+ Numbers first, and never a screenshot diff — pixel comparison goes flaky on font
406
+ antialiasing and teams end up muting it. But no measurement catches an element that is
407
+ present, on-token, and still not on screen — occluded, clipped, truncated, or the same
408
+ color as what's behind it. So the last step reads a screenshot of each state next to
409
+ your export and grades presence, legibility, and correspondence. **Where a number
410
+ already settled the question the eye may only flag** — on-scale-but-cramped is never a
411
+ fail — which keeps the two halves from contradicting each other. Measurable or visible
412
+ violations fail; taste comes back as `flag:` notes so a bounded fix loop can't thrash. A
413
+ task with no `design:` tag is graded exactly as before.
414
+
415
+ <details>
416
+ <summary>The <code>visual:</code> command (you write this once)</summary>
417
+
418
+ tiny-spec ships no script — it can't know your stack. Write one, put it in the
419
+ constitution's **Verification commands**, and the reviewer runs it:
420
+
421
+ ```
422
+ ## Verification commands
423
+ - test: `npm test`
424
+ - visual: `node visual.mjs`
425
+ ```
426
+
427
+ ```js
428
+ // visual.mjs — node visual.mjs '<selector>' ['<selector>'...]
429
+ import { chromium } from 'playwright';
430
+
431
+ const b = await chromium.launch({ channel: 'chrome' }); // your Chrome, no download
432
+ const p = await b.newPage({ viewport: { width: 900, height: 700 } });
433
+ await p.goto('http://localhost:3000/signup'); // your dev server + route
434
+
435
+ for (const sel of process.argv.slice(2)) {
436
+ const el = await p.$(sel);
437
+ if (!el) { console.log(`${sel}\n MISSING FROM DOM`); continue; } // required
438
+ console.log(sel, JSON.stringify(await el.evaluate(n => {
439
+ const c = getComputedStyle(n), r = n.getBoundingClientRect();
440
+ return { padding: c.padding, margin: c.margin, fontSize: c.fontSize,
441
+ fontWeight: c.fontWeight, lineHeight: c.lineHeight, color: c.color,
442
+ background: c.backgroundColor, border: c.border,
443
+ opacity: c.opacity, visibility: c.visibility,
444
+ top: Math.round(r.top), w: Math.round(r.width), h: Math.round(r.height) };
445
+ })));
446
+ }
447
+
448
+ const state = process.env.STATE ?? 'default'; // however you drive states
449
+ const shot = `/tmp/visual-${state}.png`;
450
+ await p.screenshot({ path: shot, fullPage: true });
451
+ console.log('SCREENSHOT', state, shot); // this line arms the judge
452
+
453
+ await b.close();
454
+ ```
455
+
456
+ The missing-selector branch **must print something** — that is what turns a renamed
457
+ element into a failure instead of a silent skip. `top` is what the layout-order check
458
+ reads, and `opacity`/`visibility` are worth printing because they turn the cheapest kind
459
+ of invisible element into a numeric failure. The `SCREENSHOT` line is what the reviewer
460
+ reads back to look at the render and catch the rest — occlusion, clipping, truncation,
461
+ same-color-on-same-color — none of which any single property reports. Drop it and the
462
+ numeric half still gates exactly as before, with the reviewer reporting `judge: not run`
463
+ rather than quietly skipping it. Without a `visual:` command at all, a task tagged
464
+ `design:` raises a blocker rather than passing quietly.
465
+
466
+ </details>
467
+
468
+ **No Figma token, plugin, or design SaaS.** A view-only account works fine: the
469
+ committed export is what the agents read, and `source:` keeps the trail back. Change
470
+ an export and its recorded `sha256` stops matching, which marks the spec stale exactly
471
+ like editing a requirement — so a design that moved under finished work can't pass
472
+ unnoticed.
473
+
474
+ `tiny-spec-build` walks the task list top to bottom. Each task runs through one loop:
475
+
476
+ 1. Plan the task against the constitution (inline, brief).
477
+ 2. Implement it with a fresh `tiny-spec-build-executor` agent.
478
+ 3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
479
+ end to end and grades against the constitution and the task's acceptance.
480
+ 4. On pass, commit the code plus a checklist tick. On fail, loop back to the
481
+ executor with the findings. After two failed attempts the run halts `exhausted`.
482
+
483
+ ```mermaid
484
+ flowchart TB
485
+ SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
486
+
487
+ TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
488
+ TASKS --> P[Plan task]
489
+ P --> I[Implement<br/>executor]
490
+ I --> R[Review + run gate<br/>reviewer]
491
+ I -->|blocker| B
492
+ R -->|pass| C[Commit + tick]
493
+ C --> TASKS
494
+ R -->|fail| I
495
+ R -->|fail twice| B[Halt — blocked / exhausted<br/>logged to decisions.md]
496
+
497
+ CON([constitution.md]) -.-> P & I & R
498
+ MEM([memory.md]) -.-> I & R
499
+ DES([SPEC.md D-n + design/ export]) -.->|only on a design: task| I & R
500
+ ```
501
+
502
+ Solid arrows are the flow. Dotted arrows show the persistent context injected into
503
+ a step: the `constitution.md` goes into planning, implementation, and review, while
504
+ `memory.md` is handed to the executor and reviewer. A task tagged `design:` also
505
+ carries its screen's `D<n>` entry into both agents — and the review step then runs
506
+ the `visual:` gate on top of the usual one.
507
+
508
+ A small `memory.md` carries operational lessons between runs, so the executor and
509
+ reviewer (which start fresh each time) don't relearn the same pitfalls.
510
+
511
+ When a task can't pass because of a gap in the design or spec, the executor stops
512
+ and logs a blocker instead of hacking around it. You fix the gap upstream in
513
+ `tiny-spec-plan` or `tiny-spec-create`, then resume. Work runs one ticket at a time and
514
+ resumes from the checklist state.
515
+
516
+ ### Working a whole list
517
+
518
+ Ask `/tiny-spec-run` to build — "build the backlog", "work through the breakdown",
519
+ "spec it out and build it" — and it takes a list of stories and works them one after
520
+ another. Per story it does the same four moves:
521
+
522
+ ```
523
+ cut a branch from main → walk the chain → tiny-spec-build → merge back to main
524
+ ```
525
+
526
+ Then the next story. Each branch is cut **fresh from main**, so story 3 sees stories 1
527
+ and 2 already merged — which is what makes an ordered list build correctly.
528
+
529
+ **The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
530
+ each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
531
+ a list at invocation instead and that wins; but a bare feature name has no acceptance
532
+ criteria, so `tiny-spec-create` will interview you when it reaches it. That's the
533
+ honest trade: a breakdown runs unattended, a pasted list is supervised.
534
+
535
+ **Every run ends in exactly one of six states, and it says which:**
536
+
537
+ | | |
538
+ |---|---|
539
+ | `done` | every story built **and merged** |
540
+ | `blocked` | an upstream document is wrong — go fix the spec or the plan |
541
+ | `exhausted` | a task stayed red past two fix attempts |
542
+ | `paused` | it reached a `pause:` point |
543
+ | `fork` | a real either/or the plan doesn't answer |
544
+ | `conflict` | a story's branch wouldn't merge cleanly |
545
+
546
+ **Only `done` means the work is built** — and in a story run that means *all* of them.
547
+ A run that stopped at story 2 of 7 reporting "done" is the single thing autonomous loops
548
+ get wrong most often, so the state is always named, along with what got merged and how
549
+ many stories are still untouched.
550
+
551
+ **A halt stops the whole run, not just that story.** Later stories in a list you wrote
552
+ top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
553
+ just produces a second, more confusing failure downstream.
554
+
555
+ **Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
556
+ the build halts *before* running it:
557
+
558
+ ```
559
+ - [ ] T7 — run the schema migration against the restored snapshot
560
+ - acceptance: every row in orders has a non-null tenant_id
561
+ - pause: confirm the snapshot is current before this touches real rows
562
+ ```
563
+
564
+ `tiny-spec-plan` proposes these for genuinely irreversible work — migrations,
565
+ destructive file operations, a new dependency, an auth boundary, a public API contract.
566
+ You can also give the run a standing policy up front ("halt before anything that touches
567
+ auth") and it gets applied as each story's tasks are sliced.
568
+
569
+ **What it will not do to your repo.** It runs exactly five git commands — `switch`,
570
+ `switch -c`, `merge --no-ff`, `merge --abort`, and reads. It refuses to start on a dirty
571
+ tree. It **never pushes**, never rebases, never resets, never deletes a branch, and
572
+ never opens a PR. Merges are local, so a bad run is one `git reset` away; publishing
573
+ stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
574
+ the undo command rather than running it.
575
+
576
+ **Walk away and come back.** Nothing is written down to track progress — it's derived:
577
+ a story whose ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked
578
+ task is in progress, no directory means not started. So asking `/tiny-spec-run` to build
579
+ the backlog again tomorrow in a fresh session picks up exactly where it stopped. No
580
+ run-state file, no lock, and no budget to configure: the story list *is* the budget.
581
+
582
+ **It never fixes a blocker for you.** A blocker means one of your documents is wrong,
583
+ and a run allowed to rewrite the requirement its own task just failed would be grading
584
+ its own homework. It reports and stops; the upstream edit is yours.
585
+
586
+ **And it never decides to build on its own.** Whether a run stops before the build or
587
+ goes all the way is fixed from your opening request and never revised — not by a stage's
588
+ closing line, not by a follow-up message. Build is your review gate, so promoting
589
+ yourself into it is the one thing a router doesn't get to do.
590
+
591
+ ## Project layout
592
+
593
+ Each skill is one self-contained `SKILL.md`, with every document skeleton inline in
594
+ it — no companion template files, no absolute paths, and no shared parent required
595
+ at runtime, so a skill works wherever you drop it. (It also means a run never stops
596
+ to ask permission to read a template out of your Claude config directory.)
597
+
598
+ tiny-spec creates a `.spec/` directory in your project root, never inside a skill.
599
+ It is namespaced per ticket, with a shared spine at the root:
600
+
601
+ ```
602
+ .spec/
603
+ constitution.md project-wide, shared across tickets
604
+ memory.md operational lessons, shared across tickets
605
+ <ticket-id>/ one directory per ticket (PROJ-123/, gh-42/, …)
606
+ SPEC.md PLAN.md tasks.md decisions.md
607
+ ```
608
+
609
+ `BREAKDOWN.md` and `design/` sit at your project root rather than inside `.spec/`,
610
+ because they are yours: the breakdown is a regenerable pre-spec planning file, and the
611
+ design exports are project files no skill ever writes.
612
+
613
+ Every format lives inline in the skill that writes it — there is no separate contract
614
+ document to keep in sync, and no skill reads anything but its own `SKILL.md` at runtime.
615
+
616
+ ## Integrations
617
+
618
+ tiny-spec binds to a task platform (Jira, GitHub Issues, Azure DevOps, Monday) by
619
+ reference only: a `ticket` block in the spec and a `Refs:` footer on each
620
+ [Conventional Commit](https://www.conventionalcommits.org/en/v1.0.0/), so the
621
+ platform auto-links the work. No API calls or credentials are required.
622
+
623
+ ## Contributing
624
+
625
+ Issues and pull requests are welcome. See [CONTRIBUTING.md](CONTRIBUTING.md), and
626
+ read [AGENTS.md](AGENTS.md) before changing any skill or agent.
627
+
628
+ ## License
629
+
630
+ [MIT](LICENSE)