tiny-spec 1.0.0__tar.gz → 1.0.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/PKG-INFO +70 -99
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/README.md +69 -98
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/pyproject.toml +1 -1
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/tiny_spec/__init__.py +1 -1
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/.gitignore +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/LICENSE +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/agents/tiny-spec-build-executor.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/agents/tiny-spec-build-reviewer.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/docs/eval/README.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/examples/todo-cli/README.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/tiny-spec-adopt/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/tiny-spec-build/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/tiny-spec-create/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/tiny-spec-design/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/tiny-spec-plan/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/tiny-spec-run/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/tiny-spec-scope/SKILL.md +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/tiny_spec/cli.py +0 -0
- {tiny_spec-1.0.0 → tiny_spec-1.0.1}/tiny_spec/manifest.json +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: tiny-spec
|
|
3
|
-
Version: 1.0.
|
|
3
|
+
Version: 1.0.1
|
|
4
4
|
Summary: A tiny, opinionated take on spec-driven development.
|
|
5
5
|
Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
|
|
6
6
|
Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
|
|
@@ -115,54 +115,36 @@ receipt rather than the adjective:
|
|
|
115
115
|
| [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
|
|
116
116
|
| [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
|
|
117
117
|
|
|
118
|
-
The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents.
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
the skill count went down: `tiny-spec-adopt` is genuinely new capability, and merging
|
|
122
|
-
four skills into two removed commands to learn rather than words to read.)
|
|
123
|
-
|
|
124
|
-
<sub>Counts taken from each project's own repository on 2026-08-03 by listing its
|
|
125
|
-
command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
|
|
126
|
-
`templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
|
|
127
|
-
OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
|
|
128
|
-
personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
|
|
129
|
-
`customize.toml` files (35). These projects are all substantially more popular than
|
|
130
|
-
tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
|
|
131
|
-
bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
|
|
118
|
+
The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents. No
|
|
119
|
+
config file, no build step, no orchestrator, no CLI to initialize a project — the skills
|
|
120
|
+
are the product.
|
|
132
121
|
|
|
133
|
-
|
|
122
|
+
<sub>Counted on 2026-08-03 from each project's own repository, not its marketing copy —
|
|
123
|
+
Spec Kit's `templates/commands/`, OpenSpec's `skills/`, BMAD's `SKILL.md` and
|
|
124
|
+
`customize.toml` files on `main`. All three are far more popular than tiny-spec, and
|
|
125
|
+
being smaller is a bet, not a proof of quality — re-run the counts yourself before
|
|
126
|
+
believing them.</sub>
|
|
134
127
|
|
|
135
|
-
|
|
136
|
-
many phases, many agents, many generated documents. tiny-spec makes the opposite
|
|
137
|
-
bet. Keep one safeguard, drop the rest.
|
|
128
|
+
## Why it's small
|
|
138
129
|
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
it earns its place.
|
|
130
|
+
Most spec frameworks are generous by default: many phases, many agents, many generated
|
|
131
|
+
documents. tiny-spec makes the opposite bet — keep one safeguard, drop the rest. That
|
|
132
|
+
safeguard is the independent reviewer, which exercises acceptance criteria end to end,
|
|
133
|
+
because a green unit test suite is not the same as working software.
|
|
144
134
|
|
|
145
135
|
The case for staying small:
|
|
146
136
|
|
|
147
|
-
- **Documents are context, and context isn't free.**
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
- **
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
impose ceremony on work that doesn't need it. tiny-spec's extra structure is
|
|
159
|
-
optional by design — add shape where it pays, skip it where it doesn't.
|
|
160
|
-
- **More moving parts is more to maintain.** Orchestrators, ownership contracts,
|
|
161
|
-
checkpoint matrices, and config files are themselves a system you have to learn
|
|
162
|
-
and keep in sync. A few small skills and two agents are not.
|
|
163
|
-
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks
|
|
164
|
-
like progress, but it isn't proof. The proof is the reviewer running your real
|
|
165
|
-
tests before each commit.
|
|
137
|
+
- **Documents are context, and context isn't free.** Every paragraph the agent has to
|
|
138
|
+
hold is room it no longer has for your actual code. tiny-spec keeps the spine small —
|
|
139
|
+
a constitution and a short memory — and injects only what each task needs.
|
|
140
|
+
- **Real work is a ticket inside a system, not a greenfield repo.** tiny-spec binds to a
|
|
141
|
+
ticket and references your task platform instead of re-describing the world.
|
|
142
|
+
- **Rigid pipelines fight the user.** Mandatory phases impose ceremony on work that
|
|
143
|
+
doesn't need it. The extra structure here is optional — add shape where it pays.
|
|
144
|
+
- **More moving parts is more to maintain.** Orchestrators, ownership contracts, and
|
|
145
|
+
config files are themselves a system you have to keep in sync. Two agents are not.
|
|
146
|
+
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks like
|
|
147
|
+
progress, but the proof is the reviewer running your real tests before each commit.
|
|
166
148
|
|
|
167
149
|
That's the whole trade: where larger kits add machinery, tiny-spec adds one
|
|
168
150
|
independent reviewer and stops.
|
|
@@ -172,10 +154,10 @@ independent reviewer and stops.
|
|
|
172
154
|
Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
|
|
173
155
|
runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
|
|
174
156
|
code with held-out tests the suite never sees, and checks the reviewer's own verdict
|
|
175
|
-
against that ground truth
|
|
176
|
-
|
|
157
|
+
against that ground truth. The number that matters is **false-PASS rate** — how often
|
|
158
|
+
the gate blesses code that is actually broken.
|
|
177
159
|
|
|
178
|
-
Most recent run — 5 benchmark tasks
|
|
160
|
+
Most recent run — 5 benchmark tasks on v1.0.0:
|
|
179
161
|
|
|
180
162
|
| | |
|
|
181
163
|
|---|---|
|
|
@@ -184,29 +166,24 @@ Most recent run — 5 benchmark tasks, v1.0.0 (`b18e32a`), 2026-09-07:
|
|
|
184
166
|
| **false-PASS rate** | **0%** — it never called broken code done |
|
|
185
167
|
| blocker rate | 0% |
|
|
186
168
|
|
|
187
|
-
Two earlier
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
1.0 rebuilt the planning chain around it — `tiny-spec-plan` now writes the task list
|
|
191
|
-
that `tiny-spec-build` executes — and the gate came through unchanged.
|
|
169
|
+
Two earlier versions scored identically, which is the more interesting result: the loop
|
|
170
|
+
is stable across releases rather than tuned to one commit — including through 1.0, which
|
|
171
|
+
rebuilt the planning chain underneath it.
|
|
192
172
|
|
|
193
173
|
See [`docs/eval/README.md`](docs/eval/README.md) for the method and
|
|
194
|
-
[`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
The harness is in the repo precisely so you don't have to take my word for it.
|
|
174
|
+
[`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the rubric. **Read
|
|
175
|
+
both skeptically: it is my benchmark, of my own tool, scored by me**, on small
|
|
176
|
+
self-contained tasks — so it measures the loop on well-specified work, not
|
|
177
|
+
large-codebase performance. The harness is in the repo so you don't have to take my
|
|
178
|
+
word for it.
|
|
200
179
|
|
|
201
180
|
## New to spec-driven development?
|
|
202
181
|
|
|
203
|
-
Spec-driven development (SDD) means writing down *what* you want and *why* before
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
before a single line is written, and the result is checked against the spec
|
|
209
|
-
rather than vibes. tiny-spec is one small take on that idea.
|
|
182
|
+
Spec-driven development (SDD) means writing down *what* you want and *why* before any
|
|
183
|
+
code exists, then letting that spec drive the build. Instead of prompting an agent and
|
|
184
|
+
hoping, you hand it a small, explicit contract — the intent, a design, and an ordered
|
|
185
|
+
list of tasks. The payoff: the agent stops guessing, you review the plan before a line
|
|
186
|
+
is written, and the result is checked against the spec rather than vibes.
|
|
210
187
|
|
|
211
188
|
## Quickstart
|
|
212
189
|
|
|
@@ -297,13 +274,12 @@ hardens it with concrete engineering rules, and `tiny-spec-build` injects it who
|
|
|
297
274
|
every task. It holds your style, standards, invariants, definition of done, and
|
|
298
275
|
verification commands.
|
|
299
276
|
|
|
300
|
-
**On an existing codebase, deriving beats asking.** Your test command,
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
into theatre.
|
|
277
|
+
**On an existing codebase, deriving beats asking.** Your test command, layout, and
|
|
278
|
+
conventions are already written down — in `package.json`, in CI, in your linter config,
|
|
279
|
+
in the shape of the tree. `tiny-spec-adopt` reads them, marks each section as declared
|
|
280
|
+
or inferred so you know which parts to distrust, then *runs* the verification commands
|
|
281
|
+
it derived and reports which went green. A gate that has never been run turns every
|
|
282
|
+
future review into theatre.
|
|
307
283
|
|
|
308
284
|
Because it is project-wide the constitution can also go missing — deleted, or never
|
|
309
285
|
committed — while your specs survive. Re-running `tiny-spec-adopt` (or
|
|
@@ -404,13 +380,12 @@ FINDINGS:
|
|
|
404
380
|
|
|
405
381
|
Numbers first, and never a screenshot diff — pixel comparison goes flaky on font
|
|
406
382
|
antialiasing and teams end up muting it. But no measurement catches an element that is
|
|
407
|
-
present, on-token, and still not on screen
|
|
383
|
+
present, on-token, and still not on screen: occluded, clipped, truncated, or the same
|
|
408
384
|
color as what's behind it. So the last step reads a screenshot of each state next to
|
|
409
385
|
your export and grades presence, legibility, and correspondence. **Where a number
|
|
410
|
-
already settled the question the eye may only flag
|
|
411
|
-
|
|
412
|
-
|
|
413
|
-
task with no `design:` tag is graded exactly as before.
|
|
386
|
+
already settled the question the eye may only flag**, which keeps the two halves from
|
|
387
|
+
contradicting each other — measurable or visible violations fail, taste comes back as
|
|
388
|
+
`flag:` notes. A task with no `design:` tag is graded exactly as before.
|
|
414
389
|
|
|
415
390
|
<details>
|
|
416
391
|
<summary>The <code>visual:</code> command (you write this once)</summary>
|
|
@@ -454,14 +429,12 @@ await b.close();
|
|
|
454
429
|
```
|
|
455
430
|
|
|
456
431
|
The missing-selector branch **must print something** — that is what turns a renamed
|
|
457
|
-
element into a failure instead of a silent skip. `top`
|
|
458
|
-
|
|
459
|
-
|
|
460
|
-
|
|
461
|
-
|
|
462
|
-
|
|
463
|
-
rather than quietly skipping it. Without a `visual:` command at all, a task tagged
|
|
464
|
-
`design:` raises a blocker rather than passing quietly.
|
|
432
|
+
element into a failure instead of a silent skip. `top` feeds the layout-order check, and
|
|
433
|
+
`opacity`/`visibility` turn the cheapest kind of invisible element into a numeric
|
|
434
|
+
failure. The `SCREENSHOT` line arms the judge; drop it and the numeric half still gates
|
|
435
|
+
exactly as before, with the reviewer reporting `judge: not run` rather than quietly
|
|
436
|
+
skipping it. Without a `visual:` command at all, a task tagged `design:` raises a
|
|
437
|
+
blocker rather than passing quietly.
|
|
465
438
|
|
|
466
439
|
</details>
|
|
467
440
|
|
|
@@ -543,10 +516,9 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
|
|
|
543
516
|
| `fork` | a real either/or the plan doesn't answer |
|
|
544
517
|
| `conflict` | a story's branch wouldn't merge cleanly |
|
|
545
518
|
|
|
546
|
-
**Only `done` means the work is built** — and in a story run that means *all* of them.
|
|
547
|
-
|
|
548
|
-
|
|
549
|
-
many stories are still untouched.
|
|
519
|
+
**Only `done` means the work is built** — and in a story run, that means *all* of them.
|
|
520
|
+
Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
|
|
521
|
+
often, so the state is always named alongside what merged and what's still untouched.
|
|
550
522
|
|
|
551
523
|
**A halt stops the whole run, not just that story.** Later stories in a list you wrote
|
|
552
524
|
top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
|
|
@@ -573,27 +545,26 @@ never opens a PR. Merges are local, so a bad run is one `git reset` away; publis
|
|
|
573
545
|
stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
|
|
574
546
|
the undo command rather than running it.
|
|
575
547
|
|
|
576
|
-
**Walk away and come back.**
|
|
577
|
-
|
|
578
|
-
|
|
579
|
-
|
|
580
|
-
|
|
548
|
+
**Walk away and come back.** Progress isn't written down, it's derived: a story whose
|
|
549
|
+
ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
|
|
550
|
+
progress, no directory means not started. Ask again tomorrow in a fresh session and it
|
|
551
|
+
picks up where it stopped. No run-state file, no lock, no budget to configure — the
|
|
552
|
+
story list *is* the budget.
|
|
581
553
|
|
|
582
554
|
**It never fixes a blocker for you.** A blocker means one of your documents is wrong,
|
|
583
555
|
and a run allowed to rewrite the requirement its own task just failed would be grading
|
|
584
556
|
its own homework. It reports and stops; the upstream edit is yours.
|
|
585
557
|
|
|
586
558
|
**And it never decides to build on its own.** Whether a run stops before the build or
|
|
587
|
-
goes all the way is fixed from your opening request and never revised
|
|
588
|
-
|
|
589
|
-
yourself into it is the one thing a router doesn't get to do.
|
|
559
|
+
goes all the way is fixed from your opening request and never revised. Build is your
|
|
560
|
+
review gate, so promoting itself into it is the one thing a router doesn't get to do.
|
|
590
561
|
|
|
591
562
|
## Project layout
|
|
592
563
|
|
|
593
564
|
Each skill is one self-contained `SKILL.md`, with every document skeleton inline in
|
|
594
|
-
it — no companion template files, no absolute paths,
|
|
595
|
-
|
|
596
|
-
|
|
565
|
+
it — no companion template files, no absolute paths, no shared parent required at
|
|
566
|
+
runtime. A skill works wherever you drop it, and a run never stops to ask permission to
|
|
567
|
+
read a template out of your Claude config directory.
|
|
597
568
|
|
|
598
569
|
tiny-spec creates a `.spec/` directory in your project root, never inside a skill.
|
|
599
570
|
It is namespaced per ticket, with a shared spine at the root:
|
|
@@ -78,54 +78,36 @@ receipt rather than the adjective:
|
|
|
78
78
|
| [OpenSpec](https://github.com/Fission-AI/OpenSpec) | 12 | — | `.openspec.yaml` | `proposal` `design` `tasks` `specs/` |
|
|
79
79
|
| [BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD) | 58 | 5 personas | 35 × `customize.toml` | `PRD` `architecture` `epics` `stories` `UX` `brief` `sprint-plan` |
|
|
80
80
|
|
|
81
|
-
The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents.
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
the skill count went down: `tiny-spec-adopt` is genuinely new capability, and merging
|
|
85
|
-
four skills into two removed commands to learn rather than words to read.)
|
|
86
|
-
|
|
87
|
-
<sub>Counts taken from each project's own repository on 2026-08-03 by listing its
|
|
88
|
-
command/skill/agent directories and templates, not from its marketing copy. Spec Kit:
|
|
89
|
-
`templates/commands/` (10 files) and the artifact set named in `templates/plan-template.md`.
|
|
90
|
-
OpenSpec: `skills/` (12 directories). BMAD: `SKILL.md` files on `main` (58), agent
|
|
91
|
-
personas in `src/bmm-skills/agents/` (analyst, architect, dev, pm, ux-designer), and
|
|
92
|
-
`customize.toml` files (35). These projects are all substantially more popular than
|
|
93
|
-
tiny-spec — Spec Kit ~125k stars, OpenSpec ~64k, BMAD ~51k — and being smaller is a
|
|
94
|
-
bet, not a proof of quality. Re-run the counts yourself before believing them.</sub>
|
|
81
|
+
The whole of tiny-spec is **2,195 lines** of markdown across 7 skills and 2 agents. No
|
|
82
|
+
config file, no build step, no orchestrator, no CLI to initialize a project — the skills
|
|
83
|
+
are the product.
|
|
95
84
|
|
|
96
|
-
|
|
85
|
+
<sub>Counted on 2026-08-03 from each project's own repository, not its marketing copy —
|
|
86
|
+
Spec Kit's `templates/commands/`, OpenSpec's `skills/`, BMAD's `SKILL.md` and
|
|
87
|
+
`customize.toml` files on `main`. All three are far more popular than tiny-spec, and
|
|
88
|
+
being smaller is a bet, not a proof of quality — re-run the counts yourself before
|
|
89
|
+
believing them.</sub>
|
|
97
90
|
|
|
98
|
-
|
|
99
|
-
many phases, many agents, many generated documents. tiny-spec makes the opposite
|
|
100
|
-
bet. Keep one safeguard, drop the rest.
|
|
91
|
+
## Why it's small
|
|
101
92
|
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
it earns its place.
|
|
93
|
+
Most spec frameworks are generous by default: many phases, many agents, many generated
|
|
94
|
+
documents. tiny-spec makes the opposite bet — keep one safeguard, drop the rest. That
|
|
95
|
+
safeguard is the independent reviewer, which exercises acceptance criteria end to end,
|
|
96
|
+
because a green unit test suite is not the same as working software.
|
|
107
97
|
|
|
108
98
|
The case for staying small:
|
|
109
99
|
|
|
110
|
-
- **Documents are context, and context isn't free.**
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
- **
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
impose ceremony on work that doesn't need it. tiny-spec's extra structure is
|
|
122
|
-
optional by design — add shape where it pays, skip it where it doesn't.
|
|
123
|
-
- **More moving parts is more to maintain.** Orchestrators, ownership contracts,
|
|
124
|
-
checkpoint matrices, and config files are themselves a system you have to learn
|
|
125
|
-
and keep in sync. A few small skills and two agents are not.
|
|
126
|
-
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks
|
|
127
|
-
like progress, but it isn't proof. The proof is the reviewer running your real
|
|
128
|
-
tests before each commit.
|
|
100
|
+
- **Documents are context, and context isn't free.** Every paragraph the agent has to
|
|
101
|
+
hold is room it no longer has for your actual code. tiny-spec keeps the spine small —
|
|
102
|
+
a constitution and a short memory — and injects only what each task needs.
|
|
103
|
+
- **Real work is a ticket inside a system, not a greenfield repo.** tiny-spec binds to a
|
|
104
|
+
ticket and references your task platform instead of re-describing the world.
|
|
105
|
+
- **Rigid pipelines fight the user.** Mandatory phases impose ceremony on work that
|
|
106
|
+
doesn't need it. The extra structure here is optional — add shape where it pays.
|
|
107
|
+
- **More moving parts is more to maintain.** Orchestrators, ownership contracts, and
|
|
108
|
+
config files are themselves a system you have to keep in sync. Two agents are not.
|
|
109
|
+
- **Generated docs can fake rigor.** A folder of polished planning artifacts looks like
|
|
110
|
+
progress, but the proof is the reviewer running your real tests before each commit.
|
|
129
111
|
|
|
130
112
|
That's the whole trade: where larger kits add machinery, tiny-spec adds one
|
|
131
113
|
independent reviewer and stops.
|
|
@@ -135,10 +117,10 @@ independent reviewer and stops.
|
|
|
135
117
|
Smaller is easy to claim, so the suite ships a harness that measures it. `docs/eval/`
|
|
136
118
|
runs tiny-spec headlessly on benchmark tasks in hermetic sandboxes, grades the produced
|
|
137
119
|
code with held-out tests the suite never sees, and checks the reviewer's own verdict
|
|
138
|
-
against that ground truth
|
|
139
|
-
|
|
120
|
+
against that ground truth. The number that matters is **false-PASS rate** — how often
|
|
121
|
+
the gate blesses code that is actually broken.
|
|
140
122
|
|
|
141
|
-
Most recent run — 5 benchmark tasks
|
|
123
|
+
Most recent run — 5 benchmark tasks on v1.0.0:
|
|
142
124
|
|
|
143
125
|
| | |
|
|
144
126
|
|---|---|
|
|
@@ -147,29 +129,24 @@ Most recent run — 5 benchmark tasks, v1.0.0 (`b18e32a`), 2026-09-07:
|
|
|
147
129
|
| **false-PASS rate** | **0%** — it never called broken code done |
|
|
148
130
|
| blocker rate | 0% |
|
|
149
131
|
|
|
150
|
-
Two earlier
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
1.0 rebuilt the planning chain around it — `tiny-spec-plan` now writes the task list
|
|
154
|
-
that `tiny-spec-build` executes — and the gate came through unchanged.
|
|
132
|
+
Two earlier versions scored identically, which is the more interesting result: the loop
|
|
133
|
+
is stable across releases rather than tuned to one commit — including through 1.0, which
|
|
134
|
+
rebuilt the planning chain underneath it.
|
|
155
135
|
|
|
156
136
|
See [`docs/eval/README.md`](docs/eval/README.md) for the method and
|
|
157
|
-
[`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
The harness is in the repo precisely so you don't have to take my word for it.
|
|
137
|
+
[`docs/sdd-evaluation-rubric.md`](docs/sdd-evaluation-rubric.md) for the rubric. **Read
|
|
138
|
+
both skeptically: it is my benchmark, of my own tool, scored by me**, on small
|
|
139
|
+
self-contained tasks — so it measures the loop on well-specified work, not
|
|
140
|
+
large-codebase performance. The harness is in the repo so you don't have to take my
|
|
141
|
+
word for it.
|
|
163
142
|
|
|
164
143
|
## New to spec-driven development?
|
|
165
144
|
|
|
166
|
-
Spec-driven development (SDD) means writing down *what* you want and *why* before
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
before a single line is written, and the result is checked against the spec
|
|
172
|
-
rather than vibes. tiny-spec is one small take on that idea.
|
|
145
|
+
Spec-driven development (SDD) means writing down *what* you want and *why* before any
|
|
146
|
+
code exists, then letting that spec drive the build. Instead of prompting an agent and
|
|
147
|
+
hoping, you hand it a small, explicit contract — the intent, a design, and an ordered
|
|
148
|
+
list of tasks. The payoff: the agent stops guessing, you review the plan before a line
|
|
149
|
+
is written, and the result is checked against the spec rather than vibes.
|
|
173
150
|
|
|
174
151
|
## Quickstart
|
|
175
152
|
|
|
@@ -260,13 +237,12 @@ hardens it with concrete engineering rules, and `tiny-spec-build` injects it who
|
|
|
260
237
|
every task. It holds your style, standards, invariants, definition of done, and
|
|
261
238
|
verification commands.
|
|
262
239
|
|
|
263
|
-
**On an existing codebase, deriving beats asking.** Your test command,
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
into theatre.
|
|
240
|
+
**On an existing codebase, deriving beats asking.** Your test command, layout, and
|
|
241
|
+
conventions are already written down — in `package.json`, in CI, in your linter config,
|
|
242
|
+
in the shape of the tree. `tiny-spec-adopt` reads them, marks each section as declared
|
|
243
|
+
or inferred so you know which parts to distrust, then *runs* the verification commands
|
|
244
|
+
it derived and reports which went green. A gate that has never been run turns every
|
|
245
|
+
future review into theatre.
|
|
270
246
|
|
|
271
247
|
Because it is project-wide the constitution can also go missing — deleted, or never
|
|
272
248
|
committed — while your specs survive. Re-running `tiny-spec-adopt` (or
|
|
@@ -367,13 +343,12 @@ FINDINGS:
|
|
|
367
343
|
|
|
368
344
|
Numbers first, and never a screenshot diff — pixel comparison goes flaky on font
|
|
369
345
|
antialiasing and teams end up muting it. But no measurement catches an element that is
|
|
370
|
-
present, on-token, and still not on screen
|
|
346
|
+
present, on-token, and still not on screen: occluded, clipped, truncated, or the same
|
|
371
347
|
color as what's behind it. So the last step reads a screenshot of each state next to
|
|
372
348
|
your export and grades presence, legibility, and correspondence. **Where a number
|
|
373
|
-
already settled the question the eye may only flag
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
task with no `design:` tag is graded exactly as before.
|
|
349
|
+
already settled the question the eye may only flag**, which keeps the two halves from
|
|
350
|
+
contradicting each other — measurable or visible violations fail, taste comes back as
|
|
351
|
+
`flag:` notes. A task with no `design:` tag is graded exactly as before.
|
|
377
352
|
|
|
378
353
|
<details>
|
|
379
354
|
<summary>The <code>visual:</code> command (you write this once)</summary>
|
|
@@ -417,14 +392,12 @@ await b.close();
|
|
|
417
392
|
```
|
|
418
393
|
|
|
419
394
|
The missing-selector branch **must print something** — that is what turns a renamed
|
|
420
|
-
element into a failure instead of a silent skip. `top`
|
|
421
|
-
|
|
422
|
-
|
|
423
|
-
|
|
424
|
-
|
|
425
|
-
|
|
426
|
-
rather than quietly skipping it. Without a `visual:` command at all, a task tagged
|
|
427
|
-
`design:` raises a blocker rather than passing quietly.
|
|
395
|
+
element into a failure instead of a silent skip. `top` feeds the layout-order check, and
|
|
396
|
+
`opacity`/`visibility` turn the cheapest kind of invisible element into a numeric
|
|
397
|
+
failure. The `SCREENSHOT` line arms the judge; drop it and the numeric half still gates
|
|
398
|
+
exactly as before, with the reviewer reporting `judge: not run` rather than quietly
|
|
399
|
+
skipping it. Without a `visual:` command at all, a task tagged `design:` raises a
|
|
400
|
+
blocker rather than passing quietly.
|
|
428
401
|
|
|
429
402
|
</details>
|
|
430
403
|
|
|
@@ -506,10 +479,9 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
|
|
|
506
479
|
| `fork` | a real either/or the plan doesn't answer |
|
|
507
480
|
| `conflict` | a story's branch wouldn't merge cleanly |
|
|
508
481
|
|
|
509
|
-
**Only `done` means the work is built** — and in a story run that means *all* of them.
|
|
510
|
-
|
|
511
|
-
|
|
512
|
-
many stories are still untouched.
|
|
482
|
+
**Only `done` means the work is built** — and in a story run, that means *all* of them.
|
|
483
|
+
Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
|
|
484
|
+
often, so the state is always named alongside what merged and what's still untouched.
|
|
513
485
|
|
|
514
486
|
**A halt stops the whole run, not just that story.** Later stories in a list you wrote
|
|
515
487
|
top to bottom usually assume the earlier ones landed, so skipping ahead past a failure
|
|
@@ -536,27 +508,26 @@ never opens a PR. Merges are local, so a bad run is one `git reset` away; publis
|
|
|
536
508
|
stays yours. If the gate goes red after a merge it leaves the merge alone and tells you
|
|
537
509
|
the undo command rather than running it.
|
|
538
510
|
|
|
539
|
-
**Walk away and come back.**
|
|
540
|
-
|
|
541
|
-
|
|
542
|
-
|
|
543
|
-
|
|
511
|
+
**Walk away and come back.** Progress isn't written down, it's derived: a story whose
|
|
512
|
+
ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
|
|
513
|
+
progress, no directory means not started. Ask again tomorrow in a fresh session and it
|
|
514
|
+
picks up where it stopped. No run-state file, no lock, no budget to configure — the
|
|
515
|
+
story list *is* the budget.
|
|
544
516
|
|
|
545
517
|
**It never fixes a blocker for you.** A blocker means one of your documents is wrong,
|
|
546
518
|
and a run allowed to rewrite the requirement its own task just failed would be grading
|
|
547
519
|
its own homework. It reports and stops; the upstream edit is yours.
|
|
548
520
|
|
|
549
521
|
**And it never decides to build on its own.** Whether a run stops before the build or
|
|
550
|
-
goes all the way is fixed from your opening request and never revised
|
|
551
|
-
|
|
552
|
-
yourself into it is the one thing a router doesn't get to do.
|
|
522
|
+
goes all the way is fixed from your opening request and never revised. Build is your
|
|
523
|
+
review gate, so promoting itself into it is the one thing a router doesn't get to do.
|
|
553
524
|
|
|
554
525
|
## Project layout
|
|
555
526
|
|
|
556
527
|
Each skill is one self-contained `SKILL.md`, with every document skeleton inline in
|
|
557
|
-
it — no companion template files, no absolute paths,
|
|
558
|
-
|
|
559
|
-
|
|
528
|
+
it — no companion template files, no absolute paths, no shared parent required at
|
|
529
|
+
runtime. A skill works wherever you drop it, and a run never stops to ask permission to
|
|
530
|
+
read a template out of your Claude config directory.
|
|
560
531
|
|
|
561
532
|
tiny-spec creates a `.spec/` directory in your project root, never inside a skill.
|
|
562
533
|
It is namespaced per ticket, with a shared spine at the root:
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|