tiny-spec 1.0.1__tar.gz → 1.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/PKG-INFO +55 -16
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/README.md +54 -15
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/agents/tiny-spec-build-executor.md +15 -2
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/agents/tiny-spec-build-reviewer.md +72 -40
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/pyproject.toml +1 -1
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/tiny-spec-build/SKILL.md +108 -18
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/tiny-spec-plan/SKILL.md +83 -32
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/tiny-spec-run/SKILL.md +139 -87
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/tiny-spec-scope/SKILL.md +20 -1
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/tiny_spec/__init__.py +1 -1
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/.gitignore +0 -0
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/LICENSE +0 -0
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/docs/eval/README.md +0 -0
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/examples/todo-cli/README.md +0 -0
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/tiny-spec-adopt/SKILL.md +0 -0
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/tiny-spec-create/SKILL.md +0 -0
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/tiny-spec-design/SKILL.md +0 -0
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/tiny_spec/cli.py +0 -0
- {tiny_spec-1.0.1 → tiny_spec-1.2.0}/tiny_spec/manifest.json +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: tiny-spec
|
|
3
|
-
Version: 1.0
|
|
3
|
+
Version: 1.2.0
|
|
4
4
|
Summary: A tiny, opinionated take on spec-driven development.
|
|
5
5
|
Project-URL: Homepage, https://github.com/GrayMa77er/tiny-spec
|
|
6
6
|
Project-URL: Source, https://github.com/GrayMa77er/tiny-spec
|
|
@@ -449,10 +449,19 @@ unnoticed.
|
|
|
449
449
|
1. Plan the task against the constitution (inline, brief).
|
|
450
450
|
2. Implement it with a fresh `tiny-spec-build-executor` agent.
|
|
451
451
|
3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
|
|
452
|
-
|
|
452
|
+
and grades against the constitution and the task's acceptance.
|
|
453
453
|
4. On pass, commit the code plus a checklist tick. On fail, loop back to the
|
|
454
454
|
executor with the findings. After two failed attempts the run halts `exhausted`.
|
|
455
455
|
|
|
456
|
+
Every task's acceptance is exercised end to end, black-box, by the reviewer — that part
|
|
457
|
+
never varies. What does vary is how much of the toolchain gets re-proven around it: a
|
|
458
|
+
clean `install → build` runs on the first task, the last one, anything touching
|
|
459
|
+
dependency or build config, and anything carrying `design:`; in between the reviewer runs
|
|
460
|
+
lint and tests against the existing build. It escalates back to the full gate whenever the
|
|
461
|
+
narrower one can't carry the verdict, and the authoritative clean run happens once more
|
|
462
|
+
over the whole project as the final smoke. Re-installing from scratch on every task
|
|
463
|
+
proved the same thing repeatedly and was the slowest part of the loop.
|
|
464
|
+
|
|
456
465
|
```mermaid
|
|
457
466
|
flowchart TB
|
|
458
467
|
SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
|
|
@@ -460,7 +469,7 @@ flowchart TB
|
|
|
460
469
|
TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
|
|
461
470
|
TASKS --> P[Plan task]
|
|
462
471
|
P --> I[Implement<br/>executor]
|
|
463
|
-
I --> R[Review + run gate<br/>reviewer]
|
|
472
|
+
I --> R[Review + run gate<br/>full or scoped<br/>reviewer]
|
|
464
473
|
I -->|blocker| B
|
|
465
474
|
R -->|pass| C[Commit + tick]
|
|
466
475
|
C --> TASKS
|
|
@@ -489,15 +498,38 @@ resumes from the checklist state.
|
|
|
489
498
|
### Working a whole list
|
|
490
499
|
|
|
491
500
|
Ask `/tiny-spec-run` to build — "build the backlog", "work through the breakdown",
|
|
492
|
-
"spec it out and build it" — and it takes a list of stories and works them
|
|
493
|
-
|
|
501
|
+
"spec it out and build it" — and it takes a list of stories and works them in batches.
|
|
502
|
+
Per story it does the same four moves:
|
|
494
503
|
|
|
495
504
|
```
|
|
496
505
|
cut a branch from main → walk the chain → tiny-spec-build → merge back to main
|
|
497
506
|
```
|
|
498
507
|
|
|
499
|
-
|
|
500
|
-
|
|
508
|
+
Each branch is cut **fresh from main**, so a later story sees the earlier ones already
|
|
509
|
+
merged — which is what makes an ordered list build correctly.
|
|
510
|
+
|
|
511
|
+
**Independent stories build at the same time.** A story can declare what it must follow
|
|
512
|
+
with a `needs:` line in `BREAKDOWN.md`; everything with no unmet `needs:` forms a batch
|
|
513
|
+
and builds **concurrently, one git worktree per story**, three at a time by default.
|
|
514
|
+
The batch merges, then the next one starts.
|
|
515
|
+
|
|
516
|
+
```
|
|
517
|
+
## Feature: interface
|
|
518
|
+
|
|
519
|
+
- Story: expose both helpers on a CLI slug: cli
|
|
520
|
+
- AC: `textkit slugify "Hi There"` prints "hi-there"
|
|
521
|
+
- needs: slugify, wordwrap
|
|
522
|
+
```
|
|
523
|
+
|
|
524
|
+
Omit `needs:` when a story stands alone — that's the common case, and the field is meant
|
|
525
|
+
to be rare. A `needs:` you didn't need costs you parallelism forever; one you missed
|
|
526
|
+
costs a single merge conflict, which the run already catches and halts on. You can also
|
|
527
|
+
just name the set yourself at invocation ("build these three at once"), which overrides
|
|
528
|
+
the graph. A cycle, or a `needs:` naming a story that isn't there, stops the run rather
|
|
529
|
+
than being guessed past.
|
|
530
|
+
|
|
531
|
+
**Tasks *inside* a story never run in parallel.** They share files and each one assumes
|
|
532
|
+
the last landed, so they stay strictly sequential. Parallelism is across stories only.
|
|
501
533
|
|
|
502
534
|
**The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
|
|
503
535
|
each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
|
|
@@ -520,9 +552,13 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
|
|
|
520
552
|
Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
|
|
521
553
|
often, so the state is always named alongside what merged and what's still untouched.
|
|
522
554
|
|
|
523
|
-
**A halt stops the
|
|
524
|
-
|
|
525
|
-
|
|
555
|
+
**A halt stops the lane it happened in, and ends the run after that batch.** Its siblings
|
|
556
|
+
were declared independent, so they finish and merge — killing working lanes because one
|
|
557
|
+
failed throws away good work. But the run does not start the next batch: later stories
|
|
558
|
+
usually assume the earlier ones landed, so skipping ahead past a failure just produces a
|
|
559
|
+
second, more confusing failure downstream. With more than one lane you get each story's
|
|
560
|
+
own state, and the run's state is the worst of them — four green lanes and one `blocked`
|
|
561
|
+
is a `blocked` run.
|
|
526
562
|
|
|
527
563
|
**Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
|
|
528
564
|
the build halts *before* running it:
|
|
@@ -538,12 +574,15 @@ destructive file operations, a new dependency, an auth boundary, a public API co
|
|
|
538
574
|
You can also give the run a standing policy up front ("halt before anything that touches
|
|
539
575
|
auth") and it gets applied as each story's tasks are sliced.
|
|
540
576
|
|
|
541
|
-
**What it will not do to your repo.** It runs exactly
|
|
542
|
-
`switch -c`, `merge --no-ff`, `merge --abort`,
|
|
543
|
-
|
|
544
|
-
|
|
545
|
-
|
|
546
|
-
the
|
|
577
|
+
**What it will not do to your repo.** It runs exactly seven git commands — `switch`,
|
|
578
|
+
`switch -c`, `merge --no-ff`, `merge --abort`, `worktree add`, `worktree list`, and
|
|
579
|
+
reads. It refuses to start on a dirty tree, or if a worktree path it needs already
|
|
580
|
+
exists. It **never pushes**, never rebases, never resets, never deletes a branch, never
|
|
581
|
+
removes a worktree, and never opens a PR. Merges are local, so a bad run is one
|
|
582
|
+
`git reset` away; publishing stays yours. If the gate goes red after a merge it leaves
|
|
583
|
+
the merge alone and tells you the undo command rather than running it — and it hands
|
|
584
|
+
back the `git worktree remove` commands for the lanes instead of running those either,
|
|
585
|
+
since a halted lane's worktree is the tree you need to look at.
|
|
547
586
|
|
|
548
587
|
**Walk away and come back.** Progress isn't written down, it's derived: a story whose
|
|
549
588
|
ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
|
|
@@ -412,10 +412,19 @@ unnoticed.
|
|
|
412
412
|
1. Plan the task against the constitution (inline, brief).
|
|
413
413
|
2. Implement it with a fresh `tiny-spec-build-executor` agent.
|
|
414
414
|
3. Review it with an independent `tiny-spec-build-reviewer` agent that runs the gate
|
|
415
|
-
|
|
415
|
+
and grades against the constitution and the task's acceptance.
|
|
416
416
|
4. On pass, commit the code plus a checklist tick. On fail, loop back to the
|
|
417
417
|
executor with the findings. After two failed attempts the run halts `exhausted`.
|
|
418
418
|
|
|
419
|
+
Every task's acceptance is exercised end to end, black-box, by the reviewer — that part
|
|
420
|
+
never varies. What does vary is how much of the toolchain gets re-proven around it: a
|
|
421
|
+
clean `install → build` runs on the first task, the last one, anything touching
|
|
422
|
+
dependency or build config, and anything carrying `design:`; in between the reviewer runs
|
|
423
|
+
lint and tests against the existing build. It escalates back to the full gate whenever the
|
|
424
|
+
narrower one can't carry the verdict, and the authoritative clean run happens once more
|
|
425
|
+
over the whole project as the final smoke. Re-installing from scratch on every task
|
|
426
|
+
proved the same thing repeatedly and was the slowest part of the loop.
|
|
427
|
+
|
|
419
428
|
```mermaid
|
|
420
429
|
flowchart TB
|
|
421
430
|
SPEC[SPEC.md<br/>intent] --> PLAN[PLAN.md<br/>design] --> TASKS[tasks.md<br/>checklist]
|
|
@@ -423,7 +432,7 @@ flowchart TB
|
|
|
423
432
|
TASKS -->|pause: set| H[Halt — paused<br/>task stays unchecked]
|
|
424
433
|
TASKS --> P[Plan task]
|
|
425
434
|
P --> I[Implement<br/>executor]
|
|
426
|
-
I --> R[Review + run gate<br/>reviewer]
|
|
435
|
+
I --> R[Review + run gate<br/>full or scoped<br/>reviewer]
|
|
427
436
|
I -->|blocker| B
|
|
428
437
|
R -->|pass| C[Commit + tick]
|
|
429
438
|
C --> TASKS
|
|
@@ -452,15 +461,38 @@ resumes from the checklist state.
|
|
|
452
461
|
### Working a whole list
|
|
453
462
|
|
|
454
463
|
Ask `/tiny-spec-run` to build — "build the backlog", "work through the breakdown",
|
|
455
|
-
"spec it out and build it" — and it takes a list of stories and works them
|
|
456
|
-
|
|
464
|
+
"spec it out and build it" — and it takes a list of stories and works them in batches.
|
|
465
|
+
Per story it does the same four moves:
|
|
457
466
|
|
|
458
467
|
```
|
|
459
468
|
cut a branch from main → walk the chain → tiny-spec-build → merge back to main
|
|
460
469
|
```
|
|
461
470
|
|
|
462
|
-
|
|
463
|
-
|
|
471
|
+
Each branch is cut **fresh from main**, so a later story sees the earlier ones already
|
|
472
|
+
merged — which is what makes an ordered list build correctly.
|
|
473
|
+
|
|
474
|
+
**Independent stories build at the same time.** A story can declare what it must follow
|
|
475
|
+
with a `needs:` line in `BREAKDOWN.md`; everything with no unmet `needs:` forms a batch
|
|
476
|
+
and builds **concurrently, one git worktree per story**, three at a time by default.
|
|
477
|
+
The batch merges, then the next one starts.
|
|
478
|
+
|
|
479
|
+
```
|
|
480
|
+
## Feature: interface
|
|
481
|
+
|
|
482
|
+
- Story: expose both helpers on a CLI slug: cli
|
|
483
|
+
- AC: `textkit slugify "Hi There"` prints "hi-there"
|
|
484
|
+
- needs: slugify, wordwrap
|
|
485
|
+
```
|
|
486
|
+
|
|
487
|
+
Omit `needs:` when a story stands alone — that's the common case, and the field is meant
|
|
488
|
+
to be rare. A `needs:` you didn't need costs you parallelism forever; one you missed
|
|
489
|
+
costs a single merge conflict, which the run already catches and halts on. You can also
|
|
490
|
+
just name the set yourself at invocation ("build these three at once"), which overrides
|
|
491
|
+
the graph. A cycle, or a `needs:` naming a story that isn't there, stops the run rather
|
|
492
|
+
than being guessed past.
|
|
493
|
+
|
|
494
|
+
**Tasks *inside* a story never run in parallel.** They share files and each one assumes
|
|
495
|
+
the last landed, so they stay strictly sequential. Parallelism is across stories only.
|
|
464
496
|
|
|
465
497
|
**The list is `BREAKDOWN.md` by default** — its `- Story:` entries, in file order,
|
|
466
498
|
each already carrying a `slug:` (the branch and directory name) and `AC:` lines. Paste
|
|
@@ -483,9 +515,13 @@ honest trade: a breakdown runs unattended, a pasted list is supervised.
|
|
|
483
515
|
Stopping at story 2 of 7 and reporting "done" is what autonomous loops get wrong most
|
|
484
516
|
often, so the state is always named alongside what merged and what's still untouched.
|
|
485
517
|
|
|
486
|
-
**A halt stops the
|
|
487
|
-
|
|
488
|
-
|
|
518
|
+
**A halt stops the lane it happened in, and ends the run after that batch.** Its siblings
|
|
519
|
+
were declared independent, so they finish and merge — killing working lanes because one
|
|
520
|
+
failed throws away good work. But the run does not start the next batch: later stories
|
|
521
|
+
usually assume the earlier ones landed, so skipping ahead past a failure just produces a
|
|
522
|
+
second, more confusing failure downstream. With more than one lane you get each story's
|
|
523
|
+
own state, and the run's state is the worst of them — four green lanes and one `blocked`
|
|
524
|
+
is a `blocked` run.
|
|
489
525
|
|
|
490
526
|
**Pause points are technical, not per-story.** Any task can carry a `pause:` line, and
|
|
491
527
|
the build halts *before* running it:
|
|
@@ -501,12 +537,15 @@ destructive file operations, a new dependency, an auth boundary, a public API co
|
|
|
501
537
|
You can also give the run a standing policy up front ("halt before anything that touches
|
|
502
538
|
auth") and it gets applied as each story's tasks are sliced.
|
|
503
539
|
|
|
504
|
-
**What it will not do to your repo.** It runs exactly
|
|
505
|
-
`switch -c`, `merge --no-ff`, `merge --abort`,
|
|
506
|
-
|
|
507
|
-
|
|
508
|
-
|
|
509
|
-
the
|
|
540
|
+
**What it will not do to your repo.** It runs exactly seven git commands — `switch`,
|
|
541
|
+
`switch -c`, `merge --no-ff`, `merge --abort`, `worktree add`, `worktree list`, and
|
|
542
|
+
reads. It refuses to start on a dirty tree, or if a worktree path it needs already
|
|
543
|
+
exists. It **never pushes**, never rebases, never resets, never deletes a branch, never
|
|
544
|
+
removes a worktree, and never opens a PR. Merges are local, so a bad run is one
|
|
545
|
+
`git reset` away; publishing stays yours. If the gate goes red after a merge it leaves
|
|
546
|
+
the merge alone and tells you the undo command rather than running it — and it hands
|
|
547
|
+
back the `git worktree remove` commands for the lanes instead of running those either,
|
|
548
|
+
since a halted lane's worktree is the tree you need to look at.
|
|
510
549
|
|
|
511
550
|
**Walk away and come back.** Progress isn't written down, it's derived: a story whose
|
|
512
551
|
ticked `tasks.md` is on `main` is done, a `.spec/<slug>/` with an unchecked task is in
|
|
@@ -15,6 +15,12 @@ pleasantries.
|
|
|
15
15
|
|
|
16
16
|
Everything you need and nothing you don't:
|
|
17
17
|
|
|
18
|
+
- **the working directory** to operate in. Every path you read, write, or run a command
|
|
19
|
+
against resolves against it. It may be a **git worktree** rather than the main checkout
|
|
20
|
+
— a build can run several stories at once, each in its own worktree. **Never read or
|
|
21
|
+
write outside the directory you were given**, and never `cd` to a sibling worktree to
|
|
22
|
+
"check something": another story is being built there right now, and what you find will
|
|
23
|
+
be wrong by the time you act on it;
|
|
18
24
|
- the **task id**, **description**, and **acceptance** (the outcome that proves it done);
|
|
19
25
|
- a **`files:` hint** — likely paths to touch (guidance, not a hard boundary);
|
|
20
26
|
- the full **constitution** (`constitution.md`): Style, Engineering standards,
|
|
@@ -24,7 +30,10 @@ Everything you need and nothing you don't:
|
|
|
24
30
|
so you don't re-learn a pitfall a past run already paid for;
|
|
25
31
|
- **if the task carries `design:`** — that screen's `D<n>` entry from `SPEC.md` and
|
|
26
32
|
the path to its committed export;
|
|
27
|
-
- the specific existing files that are your starting point, named explicitly
|
|
33
|
+
- the specific existing files that are your starting point, named explicitly;
|
|
34
|
+
- sometimes, **the files the previous task left behind** — paths only. Start there when
|
|
35
|
+
orienting: it is the most recently moved ground in the codebase and the most likely
|
|
36
|
+
place your task connects to.
|
|
28
37
|
|
|
29
38
|
You are **blind to the workflow, not to the codebase.** You don't get the plan,
|
|
30
39
|
other tasks, or shared state. But the named files are a launch point, not the whole
|
|
@@ -64,7 +73,11 @@ violated.
|
|
|
64
73
|
6. You MAY run a **narrow self-check** of your own work (the one test file you
|
|
65
74
|
wrote, a syntax/import check). You do **not** need to run the full gate — the
|
|
66
75
|
independent **reviewer** runs the authoritative Verification commands next.
|
|
67
|
-
Leave the tree in a clean, buildable state for it
|
|
76
|
+
**Leave the tree in a clean, buildable, installed state for it** — the reviewer may
|
|
77
|
+
run against the existing build rather than rebuilding from scratch, so a tree you
|
|
78
|
+
left half-migrated or with a stale artifact costs the task a review cycle. If you
|
|
79
|
+
changed dependencies or build configuration, apply them (install, regenerate the
|
|
80
|
+
lockfile) rather than leaving it for someone else, and say so in `CHANGES`.
|
|
68
81
|
|
|
69
82
|
## Hard constraints
|
|
70
83
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: tiny-spec-build-reviewer
|
|
3
|
-
description: Independently reviews a single finished task — runs the project's real gate end-to-end and checks the code against the constitution
|
|
3
|
+
description: Independently reviews a single finished task — runs the project's real gate at the scope it was given (full or scoped, escalating when scoped can't carry the verdict), exercises the acceptance end-to-end, and checks the code against the constitution. Blind to how the code was written. Returns PASS/FAIL plus findings. Spawned (one per task) by tiny-spec-build. Does not fix code, plan, spawn agents, or invoke skills.
|
|
4
4
|
tools: Read, Write, Edit, Bash, Grep, Glob
|
|
5
5
|
---
|
|
6
6
|
|
|
@@ -13,10 +13,18 @@ back to `tiny-spec-build`; return data, not pleasantries.
|
|
|
13
13
|
|
|
14
14
|
## What you receive (the context contract)
|
|
15
15
|
|
|
16
|
+
- **the working directory** to operate in — every path and every gate command resolves
|
|
17
|
+
against it. It may be a **git worktree** rather than the main checkout, since a build
|
|
18
|
+
can run several stories at once. **Never read, write, or run a gate outside the
|
|
19
|
+
directory you were given**: a sibling worktree holds a different story mid-build, and
|
|
20
|
+
measuring it would make your verdict meaningless;
|
|
16
21
|
- the **task id**, **description**, and **acceptance** (the outcome that must hold);
|
|
17
22
|
- the full **constitution** (`constitution.md`) — especially **Guiding invariants**,
|
|
18
23
|
**Definition of Done**, and **Verification commands**, plus the **Design system**
|
|
19
24
|
token table if the project has one;
|
|
25
|
+
- the **gate scope** for this task — `full` or `scoped`. `tiny-spec-build` sets it; it
|
|
26
|
+
knows where this task sits in the run and you don't. If the brief names no scope,
|
|
27
|
+
treat it as `full`;
|
|
20
28
|
- the project's **memory** if any (`memory.md`) — operational lessons (e.g. the
|
|
21
29
|
gate needs the package installed first); honor them so you don't false-fail on a
|
|
22
30
|
known precondition;
|
|
@@ -32,10 +40,38 @@ acceptance and the constitution — verified, not inferred?**
|
|
|
32
40
|
1. **Read the changed code.** Check it against the constitution: does it honor the
|
|
33
41
|
**Guiding invariants**, match the **Style** and **Layout**, meet the
|
|
34
42
|
**Definition of Done**? Note any violation as a finding.
|
|
35
|
-
2. **Run the real gate.**
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
43
|
+
2. **Run the real gate, at the scope you were given.** Capture the real output either
|
|
44
|
+
way — never a test-runner shortcut, never an inferred result.
|
|
45
|
+
|
|
46
|
+
- **`full`** — execute the constitution's **Verification commands** end-to-end
|
|
47
|
+
(install → lint → test → build → run, as applicable) from a clean state, after the
|
|
48
|
+
documented setup.
|
|
49
|
+
- **`scoped`** — run `lint:` and `test:` against the working tree, and rely on the
|
|
50
|
+
existing install/build state rather than rebuilding it from clean. Step 3 below is
|
|
51
|
+
**not** scoped: you still exercise the acceptance end-to-end, black-box, with real
|
|
52
|
+
input. A scoped gate skips re-proving the toolchain; it never skips proving the task.
|
|
53
|
+
|
|
54
|
+
**Escalate rather than guess.** Move a `scoped` gate up to `full` the moment the
|
|
55
|
+
narrower one can't carry the verdict — you need a build artifact that isn't there or
|
|
56
|
+
looks stale, `install`/`build` state seems inconsistent with the changed files, a
|
|
57
|
+
command fails in a way that might be environmental, or you simply cannot exercise the
|
|
58
|
+
acceptance without building. Escalation is always allowed and always safe. **Passing a
|
|
59
|
+
task on evidence you found insufficient is not** — that is the one thing this whole
|
|
60
|
+
role exists to prevent. Never narrow a `full` gate to `scoped`.
|
|
61
|
+
|
|
62
|
+
**A Verification command that cannot run in this environment is a `FAIL` and a
|
|
63
|
+
blocker — never a `PASS`.** If a documented command is red for a reason the task's
|
|
64
|
+
code cannot fix (the interpreter is too old, a tool isn't installed, the command was
|
|
65
|
+
never runnable as written), say exactly that in `GATE:`, return `FAIL`, and name the
|
|
66
|
+
**constitution** as the upstream document to fix. Proving the task green *somewhere
|
|
67
|
+
else* — a fresh venv, a container, an install you fixed by hand — is **not** a pass:
|
|
68
|
+
you verified a different environment than the one the gate names, and the next task
|
|
69
|
+
will hit the identical red. Diagnosing the cause and demonstrating the task's own code
|
|
70
|
+
is sound is genuinely useful; put it in `FINDINGS` so the fix is cheap. It does not
|
|
71
|
+
change the verdict. A constitution whose gate cannot run is the single most dangerous
|
|
72
|
+
file in the project, because every later task inherits the same false signal.
|
|
73
|
+
|
|
74
|
+
Say which scope you actually ran in `GATE:`, including that you escalated and why.
|
|
39
75
|
3. **Exercise the acceptance.** Trigger the task's stated outcome the most
|
|
40
76
|
black-box way available (CLI > HTTP > public API) with realistic input,
|
|
41
77
|
including a negative case if the acceptance implies a boundary or rejection.
|
|
@@ -59,25 +95,23 @@ acceptance and the constitution — verified, not inferred?**
|
|
|
59
95
|
- **A selector that matches nothing is a `FAIL`, never a skip.** Either the code
|
|
60
96
|
didn't build the element or it named it something else — both are real, and both
|
|
61
97
|
are invisible if you quietly move on. Report the selector and that it was absent.
|
|
62
|
-
- Compare each element's numbers to the tokens its row names
|
|
63
|
-
deltas ("heading is 28px, `type.heading.lg` is 24px"; "padding 19px is not on
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
- **Check `layout:`** — the arrangement, max width, and the **order** it names
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
command printed, `Read` the `D<n>`'s `export:` image, and grade **every state you
|
|
80
|
-
have a screenshot for**, not just the happy path, on four lines:
|
|
98
|
+
- Compare each element's numbers to the tokens its row names, reporting concrete
|
|
99
|
+
deltas ("heading is 28px, `type.heading.lg` is 24px"; "padding 19px is not on the
|
|
100
|
+
`space.*` scale"). **Do not pixel-diff the screenshot** — font antialiasing makes
|
|
101
|
+
image comparison flaky enough that the check gets ignored, which is how visual
|
|
102
|
+
gates die.
|
|
103
|
+
- **Check `layout:`** — the arrangement, max width, and the **order** it names, from
|
|
104
|
+
the bounding rectangles. Every token can be correct on an element in the wrong place.
|
|
105
|
+
- **Exercise every state the entry names** — empty, loading, error, success. Drive the
|
|
106
|
+
UI into each one and observe what changes; **finding the word in the source is not
|
|
107
|
+
evidence**, and it false-passes routinely (a comment saying the loading state is
|
|
108
|
+
missing contains "loading"). A surface that renders its happy path and nothing else
|
|
109
|
+
is a fail, not a nit.
|
|
110
|
+
- Finally, **judge the render against the export.** Everything above proves the numbers
|
|
111
|
+
are right; none of it can see an element that is present, on-token, and invisible —
|
|
112
|
+
so now look. `Read` each `SCREENSHOT <state> <path>` the `visual:` command printed
|
|
113
|
+
and the `D<n>`'s `export:` image, and grade **every state you have a screenshot
|
|
114
|
+
for**, not just the happy path, on four lines:
|
|
81
115
|
|
|
82
116
|
1. **Presence** — is every `elements:` row actually *visible* in the render? At
|
|
83
117
|
`opacity: 0`, zero height, clipped out of view, hidden behind a sibling, or the
|
|
@@ -85,24 +119,21 @@ acceptance and the constitution — verified, not inferred?**
|
|
|
85
119
|
not there. **FAIL** — this is why the step exists.
|
|
86
120
|
2. **Legibility & occlusion** — text clipped, truncated mid-word, overlapping
|
|
87
121
|
another element, or on a background it can't be read against. **FAIL.**
|
|
88
|
-
3. **Correspondence** —
|
|
89
|
-
|
|
90
|
-
|
|
122
|
+
3. **Correspondence** — the same regions as the export, in the reading order
|
|
123
|
+
`layout:` names? A whole region missing is a **FAIL**; a stylistic difference is
|
|
124
|
+
a **flag**.
|
|
91
125
|
4. **Hierarchy & polish** — emphasis, balance, crowding, alignment. **Always a
|
|
92
126
|
flag**, never a fail.
|
|
93
127
|
|
|
94
|
-
Three rules bound it
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
- **Cite what you saw.** Name the state whose screenshot the finding came from and
|
|
104
|
-
what was in it ("state `error`: caption present in DOM but renders at opacity 0").
|
|
105
|
-
An uncited visual claim reads as an opinion and gets ignored.
|
|
128
|
+
Three rules bound it. **The numbers beat your eye on anything they already
|
|
129
|
+
measured** — on-scale padding that looks cramped, or an exact token that looks
|
|
130
|
+
washed out, is a `flag:`; you may fail only on what measurement *cannot* see, since
|
|
131
|
+
contradicting your own numbers sends the executor a task it cannot fix and the loop
|
|
132
|
+
is bounded at two attempts. **This is still not a pixel diff** — the export is
|
|
133
|
+
usually a wireframe, so judge structure and legibility, never visual identity. And
|
|
134
|
+
**cite what you saw**, naming the state the finding came from ("state `error`:
|
|
135
|
+
caption present in DOM but renders at opacity 0") — an uncited visual claim reads as
|
|
136
|
+
an opinion and gets ignored.
|
|
106
137
|
|
|
107
138
|
**If the command printed no `SCREENSHOT` line**, do not run this sub-step and do
|
|
108
139
|
not eyeball a substitute. Grade on steps 1–3 above, write `judge: not run — visual:
|
|
@@ -158,7 +189,8 @@ Never spawn subagents or invoke skills.
|
|
|
158
189
|
```
|
|
159
190
|
TASK: <task id>
|
|
160
191
|
VERDICT: PASS | FAIL
|
|
161
|
-
GATE: <
|
|
192
|
+
GATE: <scope: full | scoped (+ "escalated from scoped: <why>" if you moved it up);
|
|
193
|
+
the Verification commands you ran + the real result (pass/fail + key output)>
|
|
162
194
|
ACCEPTANCE: <how you exercised it + the observed effect, or why you couldn't>
|
|
163
195
|
DESIGN: <omit unless the task carried `design:`. The D<n> checked, the measurements
|
|
164
196
|
you read back vs the tokens they should match, and which states you exercised —
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: tiny-spec-build
|
|
3
|
-
description: Build the spec — run the per-task loop plan→implement→review→commit, one task at a time. Implements with a fresh executor, grades with an independent reviewer running the real gate, commits per passed task, keeps a lean memory. Halts on a blocker, a spent convergence budget, a task's pause: point, or a genuine fork, and names which. Resumes from the checkbox state.
|
|
3
|
+
description: Build the spec — run the per-task loop plan→implement→review→commit, one task at a time. Implements with a fresh executor, grades with an independent reviewer running the real gate, commits per passed task, keeps a lean memory. Can be handed several lanes (one worktree + slug per independent story) and runs them concurrently, one task at a time within each. Halts on a blocker, a spent convergence budget, a task's pause: point, or a genuine fork, and names which. Resumes from the checkbox state.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# tiny-spec-build
|
|
@@ -19,16 +19,14 @@ alongside this skill (see the suite README).
|
|
|
19
19
|
|
|
20
20
|
## Inputs
|
|
21
21
|
|
|
22
|
-
1. **Resolve the active ticket dir**
|
|
23
|
-
whose slug matches the branch
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
**degraded** case, not an ask case — branch match is simply unavailable, so fall
|
|
31
|
-
through to sole-dir and ask as written.
|
|
22
|
+
1. **Resolve the active ticket dir** (call it `<active>`), in order: the `.spec/<slug>/`
|
|
23
|
+
whose slug matches the current git branch (one branch per ticket); else the sole
|
|
24
|
+
ticket dir if exactly one exists; else ask. **Ask instead** when more than one dir
|
|
25
|
+
matches the branch, or when ticket dirs exist while you are on `main`/`master` with no
|
|
26
|
+
name match — neither has a safe tie-break. Detached HEAD or no git repo is
|
|
27
|
+
**degraded**, not an ask: fall through to sole-dir and ask as written.
|
|
28
|
+
In **multi-lane mode** (below) this step does not run: each lane is handed its
|
|
29
|
+
worktree and slug explicitly.
|
|
32
30
|
2. Read `.spec/constitution.md` (**the shared constitution**), `.spec/memory.md` if
|
|
33
31
|
it exists (**shared**), and `.spec/<active>/tasks.md`. The constitution + memory
|
|
34
32
|
get injected **whole** into every executor and reviewer. Also note the `ticket`
|
|
@@ -37,6 +35,52 @@ alongside this skill (see the suite README).
|
|
|
37
35
|
`tiny-spec-plan` to reconcile first.
|
|
38
36
|
4. Pick the **first unchecked `[ ]`** task. If all are `[x]`, jump to **Completion**.
|
|
39
37
|
|
|
38
|
+
## Multi-lane mode (optional — several stories at once)
|
|
39
|
+
|
|
40
|
+
A caller may hand you **lanes** instead of a single active dir: one or more
|
|
41
|
+
`(worktree path, slug)` pairs, each an independent story in its own git worktree.
|
|
42
|
+
`tiny-spec-run` does this for a batch of stories whose `needs:` say they don't depend on
|
|
43
|
+
each other. **With one lane, or none, everything below is inert and this skill behaves
|
|
44
|
+
exactly as it always has** — that is the common case and it must not drift.
|
|
45
|
+
|
|
46
|
+
Given lanes, run them **in rounds**:
|
|
47
|
+
|
|
48
|
+
1. For every live lane, pick its first `[ ]` task and build the executor brief exactly as
|
|
49
|
+
step 2 describes — resolved against **that lane's worktree**, not the main checkout.
|
|
50
|
+
2. **Dispatch every lane's executor in one message**, so they run concurrently.
|
|
51
|
+
3. **Dispatch every lane's reviewer in one message**, once the executors are back.
|
|
52
|
+
4. Commit and tick per passing lane, in that lane's worktree.
|
|
53
|
+
|
|
54
|
+
Then start the next round with whichever lanes are still live. A fast lane waits at the
|
|
55
|
+
round boundary for a slow one; that barrier is deliberate, and much simpler to follow than
|
|
56
|
+
letting lanes free-run.
|
|
57
|
+
|
|
58
|
+
**Within a lane nothing changes.** Tasks still run one at a time, top to bottom, with the
|
|
59
|
+
same convergence bound, the same gate scope rules, and the same halting states. There is
|
|
60
|
+
no parallelism *inside* a story and no `owns:` contract — tasks in one story share files
|
|
61
|
+
and assume their predecessors landed, which is exactly why they stay sequential.
|
|
62
|
+
|
|
63
|
+
**A halt stops that lane only.** The other lanes run to completion; you report each lane's
|
|
64
|
+
terminal state separately. The stories were declared independent, so killing working lanes
|
|
65
|
+
because one failed throws away finished work for nothing.
|
|
66
|
+
|
|
67
|
+
**Every dispatched agent is told its working directory**, and that all paths resolve
|
|
68
|
+
against it. An executor or reviewer must never read or write another lane's worktree.
|
|
69
|
+
|
|
70
|
+
### `memory.md` in multi-lane mode
|
|
71
|
+
|
|
72
|
+
`.spec/memory.md` is **shared at the `.spec/` root**, and step 6 below writes to it. Two
|
|
73
|
+
lanes appending to their own copy in their own worktree would conflict on *every* parallel
|
|
74
|
+
batch — a guaranteed merge failure that has nothing to do with the code.
|
|
75
|
+
|
|
76
|
+
So: **in multi-lane mode, lanes do not write `memory.md` at all.** Collect each lane's
|
|
77
|
+
distilled lessons as it goes, and write them **once, in the main checkout on the
|
|
78
|
+
integration branch, after the batch has merged** — pruning superseded entries there, as
|
|
79
|
+
step 6 describes. Single-lane runs are unaffected and write it in place as before.
|
|
80
|
+
|
|
81
|
+
The other artifacts are safe and need no special handling: `tasks.md` and `decisions.md`
|
|
82
|
+
are per-story, and `constitution.md` is read-only for the whole build.
|
|
83
|
+
|
|
40
84
|
## The per-task loop
|
|
41
85
|
|
|
42
86
|
For the selected task, run these steps in order. **Do not tick a task until its
|
|
@@ -60,6 +104,9 @@ it short and concrete.
|
|
|
60
104
|
### 2. IMPLEMENT (dispatch `tiny-spec-build-executor`)
|
|
61
105
|
Spawn one **`tiny-spec-build-executor`** with a fresh, self-contained prompt:
|
|
62
106
|
|
|
107
|
+
- **the working directory** it must operate in — this lane's worktree, or the project
|
|
108
|
+
root in a single-lane run. Every path resolves against it, and it must never read or
|
|
109
|
+
write another lane's worktree;
|
|
63
110
|
- the task id, description, and **acceptance**;
|
|
64
111
|
- the `files:` hint;
|
|
65
112
|
- the **whole** `.spec/constitution.md`;
|
|
@@ -68,7 +115,16 @@ Spawn one **`tiny-spec-build-executor`** with a fresh, self-contained prompt:
|
|
|
68
115
|
`.spec/<active>/SPEC.md` plus its `export:` path, so the executor can look at the
|
|
69
116
|
design instead of guessing at it;
|
|
70
117
|
- only the specific existing files the task starts from, named explicitly (so it
|
|
71
|
-
edits with the real current contents, not blind)
|
|
118
|
+
edits with the real current contents, not blind);
|
|
119
|
+
- the **accumulated `CHANGES` paths from every passed task in this story so far**,
|
|
120
|
+
labelled as the ground this story has already moved. Every executor starts cold and
|
|
121
|
+
re-derives the codebase from scratch — that re-derivation is the **largest single cost
|
|
122
|
+
in the loop**, far larger than the prompt text around it — and naming the files this
|
|
123
|
+
story has already touched is the cheapest way to cut it. You already have the lists in
|
|
124
|
+
hand. **Paths only**, deduplicated, most-recently-touched first — never the earlier
|
|
125
|
+
tasks' descriptions, findings, or reports. (The reviewer does *not* get this list: it
|
|
126
|
+
is scoped to the task's own changed files on purpose, and widening it would dilute the
|
|
127
|
+
independence that makes its verdict worth anything.)
|
|
72
128
|
|
|
73
129
|
Do **not** pass the plan, sibling tasks, or other chatter. It returns a structured
|
|
74
130
|
report (`STATUS`, `CHANGES`, `DECISIONS`, `BLOCKER`). A `STATUS: blocked` →
|
|
@@ -77,14 +133,34 @@ report (`STATUS`, `CHANGES`, `DECISIONS`, `BLOCKER`). A `STATUS: blocked` →
|
|
|
77
133
|
### 3. REVIEW (dispatch `tiny-spec-build-reviewer` — independent)
|
|
78
134
|
Spawn one **`tiny-spec-build-reviewer`**, **blind to step 2**, with:
|
|
79
135
|
|
|
136
|
+
- **the working directory** it must operate in — this lane's worktree, or the project
|
|
137
|
+
root in a single-lane run. It runs the gate there and touches no other lane;
|
|
80
138
|
- the task id, description, and **acceptance**;
|
|
81
139
|
- the **whole** `.spec/constitution.md`;
|
|
82
140
|
- the list of changed files (from the executor's `CHANGES`) to read;
|
|
83
|
-
- the **Verification commands** from the constitution to run
|
|
141
|
+
- the **Verification commands** from the constitution to run, and the **gate scope**
|
|
142
|
+
for this task — `full` or `scoped` (decide it as below; you have the run context,
|
|
143
|
+
the reviewer does not);
|
|
84
144
|
- **if the task has a `design:` field** — the same `D<n>` entry and `export:` path
|
|
85
145
|
you gave the executor, so it grades against the contract rather than its taste.
|
|
86
146
|
|
|
87
|
-
|
|
147
|
+
**Gate scope.** A clean `install → build` proves the same thing on task 7 that it proved
|
|
148
|
+
on task 1 unless a task changed what it installs or builds, so running it every time buys
|
|
149
|
+
little and costs the most wall-clock in the loop. Name the scope in the brief:
|
|
150
|
+
|
|
151
|
+
- **`full`** — the whole gate from a clean state (install → lint → test → build → run).
|
|
152
|
+
Use it on: the **first** task of this run; any task whose `CHANGES` touch dependency or
|
|
153
|
+
build configuration (`package.json`/lockfiles, `pyproject.toml`, `go.mod`, `Cargo.toml`,
|
|
154
|
+
`Dockerfile`, CI config, build scripts); any task carrying `design:`; and the **last**
|
|
155
|
+
task in `tasks.md`.
|
|
156
|
+
- **`scoped`** — otherwise. `lint:` + `test:` plus the acceptance exercised end-to-end,
|
|
157
|
+
black-box, exactly as on a full gate. The acceptance exercise is never scoped away; it
|
|
158
|
+
is the whole reason the reviewer exists.
|
|
159
|
+
|
|
160
|
+
The reviewer may **escalate** `scoped` to `full` on its own and will say so — that
|
|
161
|
+
direction is always safe. It may never go the other way.
|
|
162
|
+
|
|
163
|
+
It runs the gate at the scope you named, checks the code against the constitution's
|
|
88
164
|
**Definition of Done** and **invariants**, confirms the **acceptance** actually
|
|
89
165
|
holds (exercised, not inferred), and returns `VERDICT: PASS | FAIL` + findings.
|
|
90
166
|
On a `design:` task it also runs the constitution's `visual:` command, measures the
|
|
@@ -94,11 +170,15 @@ numbers can't reach, like an element that measures perfectly and shows up invisi
|
|
|
94
170
|
It fails on what it measured or saw, and merely flags what is taste.
|
|
95
171
|
|
|
96
172
|
> Why independent: unit-green ≠ working, and the author is the worst judge of its
|
|
97
|
-
> own blind spots. The reviewer
|
|
98
|
-
> safeguard that keeps scope and quality honest without an
|
|
173
|
+
> own blind spots. The reviewer exercising the acceptance black-box — on every task,
|
|
174
|
+
> at either scope — is the safeguard that keeps scope and quality honest without an
|
|
175
|
+
> ownership contract.
|
|
99
176
|
|
|
100
177
|
### 4. CONVERGE (on FAIL)
|
|
101
|
-
Re-dispatch the **executor** with the reviewer's findings appended to its brief
|
|
178
|
+
Re-dispatch the **executor** with the reviewer's findings appended to its brief, then
|
|
179
|
+
review again — recomputing the **gate scope** from the fix's `CHANGES`, and using `full`
|
|
180
|
+
if the failure was in the gate itself or the reviewer escalated last time. A task that
|
|
181
|
+
already failed once has earned the wider check.
|
|
102
182
|
Bound this to **2 fix attempts**. If it still fails after that, stop and **halt
|
|
103
183
|
`exhausted`** (see **Halting** below) — don't keep grinding or hand-fix past the loop
|
|
104
184
|
silently. `exhausted` is recorded as a blocker, because a task that three attempts
|
|
@@ -143,6 +223,9 @@ entry to the **shared** `.spec/memory.md` (the root — lessons are project-wide
|
|
|
143
223
|
pruning any entry the new one supersedes. Skip code-style rules (→ shared
|
|
144
224
|
`constitution.md`) and one-off history (→ the ticket's `decisions.md`). Keep it lean.
|
|
145
225
|
|
|
226
|
+
**In multi-lane mode, do not write the file here.** Hold the entry and write it once
|
|
227
|
+
after the batch merges — see **Multi-lane mode** above for why.
|
|
228
|
+
|
|
146
229
|
On first use, create the file with this structure:
|
|
147
230
|
|
|
148
231
|
```markdown
|
|
@@ -186,7 +269,8 @@ way, which is why a loop and a hand-driven build can't drift apart.
|
|
|
186
269
|
## Halting
|
|
187
270
|
|
|
188
271
|
Every build run ends in exactly **one** of five terminal states. Name it out loud in
|
|
189
|
-
the closing report
|
|
272
|
+
the closing report — **per lane**, when there is more than one; a lane's halt never
|
|
273
|
+
rounds up into another lane's `done`, and the run's own state is the worst of them:
|
|
190
274
|
|
|
191
275
|
| State | Cause |
|
|
192
276
|
|---|---|
|
|
@@ -274,6 +358,12 @@ When every task in `tasks.md` is `[x]`:
|
|
|
274
358
|
setup — install/build, not a test-runner shortcut). There is no separate
|
|
275
359
|
verify skill — this final smoke confirms the requirements actually work end-to-end, not
|
|
276
360
|
just that tasks are ticked.
|
|
361
|
+
|
|
362
|
+
**This is the authoritative clean-state run for the whole build**, and it is never
|
|
363
|
+
scoped or skipped. Per-task gates run at the scope step 3 named; this one always runs
|
|
364
|
+
everything from clean. A build or install regression that a `scoped` task gate did not
|
|
365
|
+
re-prove surfaces here — which is why a red final smoke is a gap to fix or route
|
|
366
|
+
upstream, never a pass and never `done`.
|
|
277
367
|
2. **Report** — the terminal state by name (`done` here, since every task is `[x]`
|
|
278
368
|
and the smoke passed), what was built, the commits made (with the branch), and any
|
|
279
369
|
open `decisions.md` items (blockers, tasks unchecked by a reconcile). If the final
|
|
@@ -18,19 +18,15 @@ sets `tasks.md`'s `status:`.
|
|
|
18
18
|
|
|
19
19
|
Artifacts live under `.spec/`: the **shared** constitution at the root
|
|
20
20
|
(`.spec/constitution.md`), the per-ticket `SPEC.md`/`PLAN.md`/`tasks.md` under
|
|
21
|
-
`.spec/<slug>/`.
|
|
22
|
-
`.spec/<
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
`main`/`master` with no name match (the usual cause is a forgotten `git switch`, and
|
|
31
|
-
the sole-dir fallback would otherwise swallow it). Detached HEAD or no git repo is a
|
|
32
|
-
**degraded** case, not an ask case — branch match is simply unavailable, so fall
|
|
33
|
-
through to sole-dir and ask as written.
|
|
21
|
+
`.spec/<slug>/`. Both skeletons are inline below — write them from there, no file to
|
|
22
|
+
read. Requires `.spec/<active>/SPEC.md`.
|
|
23
|
+
|
|
24
|
+
**Resolve the active ticket dir**, in order: the `.spec/<slug>/` whose slug matches the
|
|
25
|
+
current git branch (one branch per ticket); else the sole ticket dir if exactly one
|
|
26
|
+
exists; else ask. **Ask instead** when more than one dir matches the branch, or when
|
|
27
|
+
ticket dirs exist while you are on `main`/`master` with no name match — neither has a
|
|
28
|
+
safe tie-break. Detached HEAD or no git repo is **degraded**, not an ask: branch match is
|
|
29
|
+
simply unavailable, so fall through to sole-dir and ask as written.
|
|
34
30
|
|
|
35
31
|
## Step 1 — harden the constitution (`constitution.md`)
|
|
36
32
|
|
|
@@ -138,25 +134,68 @@ past it into a task list — say what's unresolved and route back to `tiny-spec-
|
|
|
138
134
|
## Step 3 — slice the approach into tasks
|
|
139
135
|
|
|
140
136
|
Walk the `## Approach` you just wrote and break it into tasks. No waves, no
|
|
141
|
-
parallelism, no `owns:` contracts — tasks run one at a time, top to bottom.
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
137
|
+
parallelism, no `owns:` contracts — tasks run one at a time, top to bottom.
|
|
138
|
+
|
|
139
|
+
**The unit is one coherent commit** — what a competent engineer does in one focused
|
|
140
|
+
sitting and commits as a single unit of work. That typically touches several files and
|
|
141
|
+
often satisfies several `REQ-N` at once. It is the size you would open as one reviewable
|
|
142
|
+
change, not the smallest thing you could name.
|
|
143
|
+
|
|
144
|
+
**Know what a task costs, so size has something to trade against.** Every task you write
|
|
145
|
+
spends two cold-start agents that must re-orient in the codebase from scratch, a reviewer
|
|
146
|
+
that runs the gate and exercises the acceptance end-to-end, and two commits. That
|
|
147
|
+
overhead is **fixed** — it does not shrink for a small task. A task the executor
|
|
148
|
+
finishes in thirty seconds still pays all of it. Size each task so the work inside it
|
|
149
|
+
clearly outweighs the machinery around it.
|
|
150
|
+
|
|
151
|
+
So:
|
|
152
|
+
|
|
153
|
+
- **Split on independent failure, not on sentence length.** Split a task when it carries
|
|
154
|
+
two **unrelated** observable outcomes that could fail independently of each other. Do
|
|
155
|
+
**not** split because the acceptance got long — a good acceptance is usually several
|
|
156
|
+
clauses covering the happy path *and* its negatives. See `T4` in the `examples/todo-cli`
|
|
157
|
+
task list: two CLI commands, four requirements, and three negative cases, in one task,
|
|
158
|
+
with one acceptance. That is the calibration point, not the exception.
|
|
159
|
+
- **One task may satisfy several `REQ-N`** — `req:` takes a list. Coverage means every
|
|
160
|
+
requirement has a *home*, not that every requirement gets its *own* task. **A 1:1
|
|
161
|
+
REQ→task mapping is the single most common way this list comes out too granular.**
|
|
162
|
+
Group the requirements that one coherent change delivers together.
|
|
146
163
|
- **Ordered so each builds on the last.** Tasks run sequentially, so a later task may
|
|
147
164
|
freely assume an earlier task's code already exists. Put foundational work (types,
|
|
148
165
|
schema, scaffolding) first. Order by dependency, not by guesswork.
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
166
|
+
|
|
167
|
+
**Smells that mean you sliced below the commit line** — fold each of these back into the
|
|
168
|
+
task it belongs to:
|
|
169
|
+
|
|
170
|
+
- a task that only defines types, interfaces, or schema with no behavior behind them;
|
|
171
|
+
- a task that only adds tests for the task before it (the constitution's **Definition of
|
|
172
|
+
Done** already requires the tests to ship with the code);
|
|
173
|
+
- one task per file, or one task per function;
|
|
174
|
+
- a "wire it up" / "integrate the pieces" task trailing the pieces it wires;
|
|
175
|
+
- a **leading pure-scaffold task** — a skeleton, a dispatch stub, a module that imports
|
|
176
|
+
cleanly and does nothing. Fold it into the first task that gives it behavior;
|
|
177
|
+
- a **trailing end-to-end verification task**. `tiny-spec-build`'s Completion step already
|
|
178
|
+
runs the whole gate against the whole project from a clean state, exercised the way a
|
|
179
|
+
user would. A task that re-does it buys nothing and costs the full per-task overhead.
|
|
180
|
+
|
|
181
|
+
The shipped `examples/todo-cli` task list predates these two smells and shows both: its
|
|
182
|
+
`T1` is a pure scaffold and its `T5` is an end-to-end verification pass. Today `T1` folds
|
|
183
|
+
into `T2` and `T5` doesn't exist. Read that list for `T4`'s sizing, not for its edges.
|
|
184
|
+
|
|
185
|
+
**Count is a smell, not a cap.** A story sized the way `tiny-spec-scope` describes
|
|
186
|
+
usually lands in **2–4 tasks**. If you are past about six, re-read the list: you have
|
|
187
|
+
either sliced below the commit line, or the story itself was too big and should have been
|
|
188
|
+
split upstream. Check the list against that; do **not** enforce a number, and never drop
|
|
189
|
+
or merge coverage just to hit one.
|
|
152
190
|
|
|
153
191
|
For each task, write:
|
|
154
192
|
|
|
155
193
|
```
|
|
156
194
|
- [ ] T<n> — <imperative description>
|
|
157
|
-
- acceptance: <one user-observable outcome that proves it's done
|
|
195
|
+
- acceptance: <one user-observable outcome that proves it's done — happy path and the
|
|
196
|
+
negatives that bound it, in one entry>
|
|
158
197
|
- type: feat # optional; Conventional Commit type (defaults to feat)
|
|
159
|
-
- req: REQ-n
|
|
198
|
+
- req: REQ-n, REQ-n # optional; the REQ-N this task delivers — a list, not one
|
|
160
199
|
- design: D-n # optional; the SPEC.md D<n> screen this task builds — arms the visual gate
|
|
161
200
|
- pause: <why> # optional; halt the build before this task so a human looks first
|
|
162
201
|
- files: <comma-separated hint of files it will touch>
|
|
@@ -164,7 +203,11 @@ For each task, write:
|
|
|
164
203
|
|
|
165
204
|
The **acceptance** is what the reviewer checks against — make it observable
|
|
166
205
|
("`spec --version` prints the version and exits 0"), not internal ("version logic
|
|
167
|
-
added"). **
|
|
206
|
+
added"). It states **one outcome**, but one outcome is not one clause: spell out the
|
|
207
|
+
happy path and the negatives that bound it in the same acceptance, separated by
|
|
208
|
+
semicolons. A long acceptance is a well-specified task, not an oversized one — it is the
|
|
209
|
+
*number of unrelated things that could fail* that decides whether to split, not the
|
|
210
|
+
length of the line. **type** picks the Conventional Commit type `tiny-spec-build` uses for this
|
|
168
211
|
task's code commit (`feat | fix | docs | refactor | test | chore | build | ci | perf | style`);
|
|
169
212
|
set it when the task is clearly not a feature, otherwise omit and it defaults to `feat`.
|
|
170
213
|
**req** ties the task to the requirement it satisfies (traceability). The **files** line
|
|
@@ -211,7 +254,9 @@ anyway. A policy that matches nothing here is not an error — say so and move o
|
|
|
211
254
|
than stretching a task to fit it.
|
|
212
255
|
|
|
213
256
|
Cover **every** part of the approach — together the tasks must deliver all `REQ-N`.
|
|
214
|
-
Don't leave a requirement with no task.
|
|
257
|
+
Don't leave a requirement with no task. Coverage is about requirements having a home,
|
|
258
|
+
not about the shape of the mapping: several `REQ-N` on one task is the normal case, and
|
|
259
|
+
a task per requirement is the anti-pattern. Likewise, if `SPEC.md` has a `## Design`
|
|
215
260
|
section, every `D<n>` in it needs at least one task carrying that `design:` reference —
|
|
216
261
|
a screen nobody is graded against is a screen that will be built wrong.
|
|
217
262
|
|
|
@@ -236,25 +281,31 @@ updated: <ISO date>
|
|
|
236
281
|
|
|
237
282
|
## Tasks
|
|
238
283
|
|
|
239
|
-
- [ ] T1 — <one
|
|
240
|
-
- acceptance: <
|
|
284
|
+
- [ ] T1 — <one coherent commit's worth of foundational work; usually several files>
|
|
285
|
+
- acceptance: <the observable outcome; happy path; and the negative case that bounds it>
|
|
241
286
|
- type: feat # optional; Conventional Commit type for this task's commit (defaults to feat)
|
|
242
|
-
- req: REQ-1
|
|
243
|
-
- files: <path, path>
|
|
287
|
+
- req: REQ-1, REQ-2 # optional; the REQ-N this task delivers — several is normal
|
|
288
|
+
- files: <path, path, path>
|
|
244
289
|
|
|
245
|
-
- [ ] T2 — <next
|
|
246
|
-
- acceptance: <observable outcome>
|
|
290
|
+
- [ ] T2 — <next coherent change; assume T1's code exists>
|
|
291
|
+
- acceptance: <observable outcome; plus what it rejects and how it fails>
|
|
247
292
|
- type: feat
|
|
248
|
-
- req: REQ-
|
|
293
|
+
- req: REQ-3, REQ-4, REQ-5
|
|
249
294
|
- design: D1 # optional; only on tasks that build the visible surface
|
|
250
295
|
- files: <path, path>
|
|
251
296
|
|
|
252
297
|
- [ ] T3 — <…>
|
|
253
298
|
- acceptance: <observable outcome>
|
|
299
|
+
- req: REQ-6
|
|
254
300
|
- pause: <optional; what to check before this runs — irreversible work only>
|
|
255
301
|
- files: <path, path>
|
|
256
302
|
```
|
|
257
303
|
|
|
304
|
+
The `req:` lists above are the shape to aim for, not filler: a handful of tasks each
|
|
305
|
+
carrying the requirements one coherent change delivers. A skeleton filled in as
|
|
306
|
+
`REQ-1`, `REQ-2`, `REQ-3` down a column of single-requirement tasks is the granularity
|
|
307
|
+
failure described in Step 3.
|
|
308
|
+
|
|
258
309
|
## Update mode (SPEC changed → PLAN and tasks are stale)
|
|
259
310
|
|
|
260
311
|
When `PLAN.md` or `tasks.md` is `status: stale`, reconcile **both in one pass** — they
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: tiny-spec-run
|
|
3
|
-
description: The one router — read each artifact's status flag to work out where work stands, then invoke the skill that moves it forward. Two stop points, chosen once at the start: by default it walks tiny-spec-adopt/create → design → plan and STOPS before tiny-spec-build; asked to build ("spec it out and build it") or handed a story list, it drives each story branch → plan → build → merge until the list ends or something halts it. Use for "run the spec flow", "pick the chain back up", "build the backlog", "work through the breakdown". NOT for a single stage — for that invoke tiny-spec-create, tiny-spec-plan, or tiny-spec-design directly.
|
|
3
|
+
description: The one router — read each artifact's status flag to work out where work stands, then invoke the skill that moves it forward. Two stop points, chosen once at the start: by default it walks tiny-spec-adopt/create → design → plan and STOPS before tiny-spec-build; asked to build ("spec it out and build it") or handed a story list, it drives each story branch → plan → build → merge until the list ends or something halts it. Independent stories — grouped by their BREAKDOWN.md needs: field, or named by the user — build concurrently, one git worktree per lane, merging batch by batch. Use for "run the spec flow", "pick the chain back up", "build the backlog", "work through the breakdown". NOT for a single stage — for that invoke tiny-spec-create, tiny-spec-plan, or tiny-spec-design directly.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# tiny-spec-run
|
|
@@ -53,8 +53,8 @@ return to Step 2 after a stage, or between stories.
|
|
|
53
53
|
the branch for the user.
|
|
54
54
|
|
|
55
55
|
2. **Resolve the story list** (below) and play it back in order — slug and title — with
|
|
56
|
-
the state you derived for each
|
|
57
|
-
before anything is cut.
|
|
56
|
+
the state you derived for each, **grouped into the batches you will run**. This is the
|
|
57
|
+
user's chance to reorder, drop, narrow, or correct a dependency before anything is cut.
|
|
58
58
|
|
|
59
59
|
3. **Take the pause policy.** Ask for, or accept, standing technical stop points for
|
|
60
60
|
this run — *"halt before anything that touches auth"*, *"stop before any schema
|
|
@@ -63,6 +63,17 @@ return to Step 2 after a stage, or between stories.
|
|
|
63
63
|
**technical**, not per-story: the thing worth looking at is a migration or an auth
|
|
64
64
|
boundary, not a feature heading.
|
|
65
65
|
|
|
66
|
+
4. **Fix the lane count.** Default **3**; take a number from the user if they gave one.
|
|
67
|
+
This is a concurrency limit on the machine, not a budget — the story list is still the
|
|
68
|
+
budget, and a batch wider than the lane count simply runs in successive fills. **One
|
|
69
|
+
lane must stay available**: any phrasing asking for one story at a time ("one at a
|
|
70
|
+
time", "no parallelism", "serially") sets it to 1, which collapses Step 4 to the
|
|
71
|
+
sequential behavior — no worktrees, no batching.
|
|
72
|
+
|
|
73
|
+
Then check one more precondition, once you know the slugs in the first batch: **no
|
|
74
|
+
target worktree path already exists.** Name the collisions and stop; do not reuse or
|
|
75
|
+
remove a directory you found.
|
|
76
|
+
|
|
66
77
|
## Step 1 — resolve the scope
|
|
67
78
|
|
|
68
79
|
### Stop-before-build: the active ticket dir
|
|
@@ -126,7 +137,26 @@ on the integration branch at all, so "does the directory exist" can't tell "not
|
|
|
126
137
|
apart from "built on a branch you haven't merged" — reading each ref explicitly can.
|
|
127
138
|
Sourcing step 1 from the integration branch is also what survives a deleted branch: once
|
|
128
139
|
a story is merged, its ticked task list is part of `main` whether or not the branch that
|
|
129
|
-
built it still exists.
|
|
140
|
+
built it still exists. **This survives parallelism unchanged** — a worktree is just a
|
|
141
|
+
branch, so a story built in a worktree but not yet merged lands on rule 2 exactly as one
|
|
142
|
+
built in place.
|
|
143
|
+
|
|
144
|
+
**Group the remaining stories into batches.** Level 0 is every unbuilt story whose
|
|
145
|
+
`needs:` are empty or already built-and-merged; level 1 is every story whose `needs:` are
|
|
146
|
+
all in level 0; and so on. Batches run **in order**; the stories inside one batch run
|
|
147
|
+
**concurrently**, capped at the lane count.
|
|
148
|
+
|
|
149
|
+
- **A parallel set the user named at invocation overrides the graph** for this run —
|
|
150
|
+
"build these three at once", or a pasted list they called independent. Explicit beats
|
|
151
|
+
derived, the same way a pasted story list already wins over `BREAKDOWN.md`. Say that you
|
|
152
|
+
are overriding, and name what the graph would have done.
|
|
153
|
+
- **A cycle in `needs:` is a stop.** Print the cycle and stop. Never break it by picking
|
|
154
|
+
an order — `tiny-spec-scope` says two stories that need each other are really one, and
|
|
155
|
+
choosing for the user hides a carve that needs fixing.
|
|
156
|
+
- **A `needs:` naming a slug that isn't in the list is a stop**, not a shrug. It is either
|
|
157
|
+
a typo or a story someone dropped, and both want a human.
|
|
158
|
+
- **A batch of one is the sequential behavior** — Step 4 collapses to exactly what it did
|
|
159
|
+
before lanes existed. This is the common case for a single ticket and must not change.
|
|
130
160
|
|
|
131
161
|
## Step 2 — the ladder
|
|
132
162
|
|
|
@@ -146,67 +176,62 @@ re-evaluate against the new state on disk.
|
|
|
146
176
|
| — | **no rung matched** | **stop** — report the exact state you found and ask; never improvise a stage |
|
|
147
177
|
|
|
148
178
|
**Upstream beats downstream — that is what the table order encodes.** Always fix the
|
|
149
|
-
earliest artifact in the chain that needs attention,
|
|
150
|
-
|
|
151
|
-
|
|
179
|
+
earliest artifact in the chain that needs attention, stale *or* missing. Deriving a task
|
|
180
|
+
list from a design you already know is wrong wastes the run, and the reconcile then has
|
|
181
|
+
to uncheck completed work all over again.
|
|
152
182
|
|
|
153
183
|
**L0/L1 carry the greenfield/brownfield fork.** "Does this repo already contain source?"
|
|
154
|
-
means:
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
**L1 is a reseed, not a fresh start
|
|
161
|
-
missing while a perfectly good `SPEC.md` sits next to it.
|
|
162
|
-
**reseed/refresh mode** — say so explicitly, and add: *do not
|
|
163
|
-
do not create a ticket dir, and do not touch `SPEC.md`.* If the
|
|
164
|
-
doesn't exist afterwards, stop and tell the user; do not loop.
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
is
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
whether two documents agree. A design that moved under a finished spec is otherwise
|
|
182
|
-
completely invisible — no status flips, and the build reviews against a screen that no
|
|
183
|
-
longer exists.
|
|
184
|
+
means: real code beyond config and docs. If yes, the constitution must be *derived from
|
|
185
|
+
that code* (`tiny-spec-adopt`) rather than interviewed out of the user — an interviewed
|
|
186
|
+
constitution on an existing codebase is how you get a gate that names commands the
|
|
187
|
+
project doesn't have. Starting from an idea with no code yet, `tiny-spec-scope` is the
|
|
188
|
+
user's to run first; it is not in the ladder (see below).
|
|
189
|
+
|
|
190
|
+
**L1 is a reseed, not a fresh start**, and it **outranks L2**. The constitution is
|
|
191
|
+
project-wide, so it can go missing while a perfectly good `SPEC.md` sits next to it.
|
|
192
|
+
Invoke the chosen skill in **reseed/refresh mode** — say so explicitly, and add: *do not
|
|
193
|
+
re-interview from scratch, do not create a ticket dir, and do not touch `SPEC.md`.* If the
|
|
194
|
+
constitution still doesn't exist afterwards, stop and tell the user; do not loop. The next
|
|
195
|
+
pass lands on L2 and creates the spec properly — a *different rung*, so the same-skill
|
|
196
|
+
bound below does not fire.
|
|
197
|
+
|
|
198
|
+
**L3 is a stop, not a stage.** Nothing in the suite ever *sets* `SPEC.md` to stale — it is
|
|
199
|
+
the root of the chain — so a stale SPEC means someone hand-edited it. Surface it and offer
|
|
200
|
+
`tiny-spec-create` in update mode rather than assuming intent. A go-ahead **in the same
|
|
201
|
+
turn** is enough to invoke it and carry on down the ladder; the stop exists to get a human
|
|
202
|
+
decision, not to force a second command. (Update mode *clears* the flag on its way out, so
|
|
203
|
+
this rung resolves rather than repeating.)
|
|
204
|
+
|
|
205
|
+
**L4 is the one thing `run` checks that isn't a `status:` flag**, and it is deliberately
|
|
206
|
+
*not* a validator: for each `D<n>` in the active `SPEC.md`, run `shasum -a 256 <export>`
|
|
207
|
+
and compare with the `sha256:` the entry declares. You are checking whether an anchor
|
|
208
|
+
still points at what it says it points at, not inferring whether two documents agree. A
|
|
209
|
+
design that moved under a finished spec is otherwise completely invisible — no status
|
|
210
|
+
flips, and the build reviews against a screen that no longer exists.
|
|
184
211
|
|
|
185
212
|
- **Mismatch** → `tiny-spec-design` in re-anchor mode.
|
|
186
213
|
- **Missing file** → **stop and tell the user**, naming the entry and the path. Don't
|
|
187
|
-
route it: a deleted export can mean a rename, a move, or a design
|
|
188
|
-
|
|
214
|
+
route it: a deleted export can mean a rename, a move, or a withdrawn design, and each
|
|
215
|
+
wants a different answer. Never quietly drop the entry.
|
|
189
216
|
- No `## Design` section, or every hash matches → the rung doesn't fire; fall through.
|
|
190
217
|
|
|
191
218
|
**Any `status:` that isn't exactly `current`** — `stale`, missing, unreadable, or an
|
|
192
219
|
unrecognized value like `draft` — counts as **stale**, on `SPEC.md`, `PLAN.md`, and
|
|
193
|
-
`tasks.md` alike. Say so out loud.
|
|
194
|
-
|
|
195
|
-
`SPEC.md` it means L3 — stop and ask.
|
|
220
|
+
`tasks.md` alike. Say so out loud. `PLAN.md`/`tasks.md` → reconcile (update mode preserves
|
|
221
|
+
existing ids, the non-destructive way to be wrong); `SPEC.md` → L3, stop and ask.
|
|
196
222
|
|
|
197
|
-
**A `tasks.md` with no tasks at all is not "built"** —
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
time, stop and tell the user — the plan has nothing derivable in it.
|
|
223
|
+
**A `tasks.md` with no tasks at all is not "built"** — L6 and L7 both require at least one
|
|
224
|
+
task, so an empty checklist matches neither and L5 catches it. Empty a second time → stop
|
|
225
|
+
and tell the user; the plan has nothing derivable in it.
|
|
201
226
|
|
|
202
227
|
`run` trusts the `status:` flags. It does **not** second-guess hand edits, diff
|
|
203
|
-
timestamps, or validate the chain — there is no validator in this suite by design. (L4
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
228
|
+
timestamps, or validate the chain — there is no validator in this suite by design. (L4 is
|
|
229
|
+
not an exception: a hash is a value the spec itself declares about a file it names, so
|
|
230
|
+
checking it is reading state, not judging consistency. Not a precedent for adding
|
|
231
|
+
cross-document checks.) So it inherits each stage's propagation: if `tiny-spec-plan`
|
|
232
|
+
update mode doesn't flip `tasks.md`, `run` walks right past it. That trust is the price of
|
|
233
|
+
having no validator; when a run's result looks wrong, suspect the stage's propagation
|
|
234
|
+
before the ladder.
|
|
210
235
|
|
|
211
236
|
**`tiny-spec-scope` is not in the ladder.** It is a pre-spec on-ramp that writes
|
|
212
237
|
`BREAKDOWN.md` at the project root, carries no `status:` frontmatter, and is the user's
|
|
@@ -253,39 +278,60 @@ to enter `tiny-spec-build` in a stop-before-build run.**
|
|
|
253
278
|
## Step 4 — build-through only: build and merge
|
|
254
279
|
|
|
255
280
|
Reached only when Step 0 fixed the stop point at build-through and the ladder is at L6
|
|
256
|
-
(or L7 with an unmerged branch).
|
|
257
|
-
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
|
|
281
|
+
(or L7 with an unmerged branch). Work **one batch at a time**, in batch order. For every
|
|
282
|
+
story in the current batch that isn't already built and merged:
|
|
283
|
+
|
|
284
|
+
1. **Lane.** Give each story its own **git worktree**, so lanes cannot collide on the
|
|
285
|
+
filesystem: `git worktree add ../<repo>-<slug> -b <slug> <integration>`, or
|
|
286
|
+
`git worktree add ../<repo>-<slug> <slug>` when the branch already exists. Cut **fresh
|
|
287
|
+
from the integration branch** so the story sees every story merged before it — that is
|
|
288
|
+
what makes an ordered backlog build correctly, and it is why batches merge before the
|
|
289
|
+
next one starts.
|
|
290
|
+
|
|
291
|
+
With a single lane you may stay in the main checkout and `git switch` as before; a
|
|
292
|
+
worktree is only required when a batch has more than one story.
|
|
293
|
+
|
|
294
|
+
2. **Ladder.** Walk Steps 1–3 for each story, in its own worktree, until it reaches L6.
|
|
295
|
+
A story that stops anywhere else does **not** enter the build — its lane is halted;
|
|
296
|
+
see the halt rule below. **Never walk the ladder twice to push past its own stop:**
|
|
297
|
+
each of those is a human decision it deliberately declined to make, and running it
|
|
298
|
+
again declines again.
|
|
299
|
+
|
|
300
|
+
3. **Build.** Invoke **`tiny-spec-build`** **once for the whole batch**, briefed to run
|
|
301
|
+
through, and hand it every live lane as a `(worktree path, slug)` pair. It owns the
|
|
302
|
+
per-task loop, runs the lanes concurrently, and writes each story's halt record.
|
|
303
|
+
Do not invoke it a second time for a lane that halted: it resumes from the checkbox
|
|
304
|
+
state, so it lands on the very task that just halted and halts there again.
|
|
305
|
+
|
|
306
|
+
4. **Merge — only the lanes that returned `done`, one at a time, in batch order.**
|
|
307
|
+
For each, in the main checkout:
|
|
275
308
|
- `git switch <integration>`
|
|
276
309
|
- `git merge --no-ff <slug>` — the merge commit keeps each story legible in history.
|
|
277
|
-
- **Conflict** → `git merge --abort`,
|
|
278
|
-
|
|
279
|
-
|
|
280
|
-
|
|
310
|
+
- **Conflict** → `git merge --abort`, halt that story `conflict`, and carry on with
|
|
311
|
+
the rest of the batch. Report every conflicted story at the end.
|
|
312
|
+
- **Run the constitution's Verification commands on the merged result**, exercised the
|
|
313
|
+
way a user would. Red → halt `blocked`. **This check is load-bearing under
|
|
314
|
+
parallelism and is never skipped:** every lane gated against a tree that did not
|
|
315
|
+
contain its siblings, so the merged result is the first time they meet.
|
|
281
316
|
- **Never push.** Merging locally keeps a bad run one `git reset` away; sending it to
|
|
282
317
|
a remote is the user's call, and this suite makes no network calls.
|
|
283
318
|
|
|
284
|
-
5. **Next
|
|
319
|
+
5. **Next batch.** If **every** lane in this batch merged cleanly, return to Step 1 for
|
|
320
|
+
the next batch. Do not re-run Step 0. **If any lane halted, stop the run here** — a
|
|
321
|
+
later batch may well depend on the story that failed, and working out which stories
|
|
322
|
+
are still safe is exactly the machinery this suite refuses to grow. Report and let the
|
|
323
|
+
user decide.
|
|
324
|
+
|
|
325
|
+
6. **Leave the worktrees.** Do **not** run `git worktree remove`. Cleanup is the user's,
|
|
326
|
+
for the same reason branch deletion is — a halted lane's worktree holds the tree they
|
|
327
|
+
need to look at. List every path you created in the closing report, with the
|
|
328
|
+
`git worktree remove` commands, and let them run it.
|
|
285
329
|
|
|
286
|
-
**A halt stops the
|
|
287
|
-
|
|
288
|
-
|
|
330
|
+
**A halt stops the lane it happened in, and ends the run at the end of that batch.**
|
|
331
|
+
Within a batch the other lanes run to completion: the stories were declared independent,
|
|
332
|
+
so stopping the ones that are working buys nothing and throws away finished work. Across
|
|
333
|
+
batches the old rule stands unchanged — later stories usually assume earlier ones landed,
|
|
334
|
+
so never skip ahead past a failure.
|
|
289
335
|
|
|
290
336
|
`tiny-spec-build` records its own halts (`blocked`, `exhausted`, `fork`) in the story's
|
|
291
337
|
`decisions.md`. **You record nothing** — the merge-stage halts are already legible
|
|
@@ -309,11 +355,15 @@ L3 or on a bound, say exactly what stopped you and what the user needs to decide
|
|
|
309
355
|
**Build-through.** Name, in this order:
|
|
310
356
|
|
|
311
357
|
1. **The terminal state** — exactly one of `done`, `blocked`, `exhausted`, `paused`,
|
|
312
|
-
`fork`, `conflict`. Use the word.
|
|
313
|
-
|
|
314
|
-
|
|
358
|
+
`fork`, `conflict`. Use the word. With more than one lane, report **each story's own
|
|
359
|
+
state**, and give the run's state as the **worst** of them: four green lanes and one
|
|
360
|
+
`blocked` is a `blocked` run, never a `done` one with a footnote.
|
|
361
|
+
2. **Stories built and merged**, in batch order, with their merge commits.
|
|
362
|
+
3. **The story or stories it stopped on**, the task within each, and why in one line.
|
|
315
363
|
4. **Stories never started** — say how many are left, by name.
|
|
316
|
-
5. **
|
|
364
|
+
5. **Worktrees left on disk**, with the `git worktree remove` command for each. You do
|
|
365
|
+
not run them.
|
|
366
|
+
6. **The one command that resolves it** — `tiny-spec-create`/`tiny-spec-plan` in update
|
|
317
367
|
mode for `blocked`/`exhausted`, this skill again for `paused`, the decision the user
|
|
318
368
|
owes you for a `fork`, or the conflicted paths for a `conflict`.
|
|
319
369
|
|
|
@@ -338,9 +388,11 @@ ticket that legitimately has no visual surface.
|
|
|
338
388
|
- **Never write, edit, or flip anything.** No `status:`, no checkbox, no `decisions.md`
|
|
339
389
|
entry, no code, no `BREAKDOWN.md` edit. Delegate or stop. A router that regenerates
|
|
340
390
|
documents is a second source of truth.
|
|
341
|
-
- **Never push, force, rebase, reset, or
|
|
342
|
-
skill issues are `switch`, `switch -c`, `merge --no-ff`,
|
|
343
|
-
|
|
391
|
+
- **Never push, force, rebase, reset, delete a branch, or remove a worktree.** The only
|
|
392
|
+
git commands this skill issues are `switch`, `switch -c`, `merge --no-ff`,
|
|
393
|
+
`merge --abort`, `worktree add`, `worktree list`, and read-only queries. Everything
|
|
394
|
+
outward-facing or destructive stays the user's — `worktree remove` included, which is
|
|
395
|
+
why Step 4 hands those commands over instead of running them.
|
|
344
396
|
- **Never merge a story whose build didn't return `done`.**
|
|
345
397
|
- **Never invoke `tiny-spec-run`.** Re-entering means re-reading these steps, not
|
|
346
398
|
calling yourself. Self-invocation compounds context and does not terminate.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: tiny-spec-scope
|
|
3
|
-
description: The greenfield on-ramp — turn a rough idea, or an existing PRD, into BREAKDOWN.md at the project root - a flat list of Features → user Stories, each with draft acceptance criteria
|
|
3
|
+
description: The greenfield on-ramp — turn a rough idea, or an existing PRD, into BREAKDOWN.md at the project root - a flat list of Features → user Stories, each with draft acceptance criteria, a slug, and an optional needs: naming the stories it must follow, plus a shared Decisions block. The one skill that works from a blank page. Does not scaffold .spec/ or touch the constitution — tiny-spec-create reads BREAKDOWN.md in seeded mode and does that. For an existing codebase use tiny-spec-adopt instead. The suite works without either.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# tiny-spec-scope
|
|
@@ -135,6 +135,7 @@ Write `BREAKDOWN.md` at the **project root** (the user's cwd) with the structure
|
|
|
135
135
|
- Story: <one user-observable capability> slug: <ado-__ | kebab>
|
|
136
136
|
- AC: <a single user-observable, testable outcome — becomes REQ-1 in this story's SPEC.md>
|
|
137
137
|
- AC: <another — keep each atomic; no "and" hiding two capabilities>
|
|
138
|
+
- needs: <slug, slug> <!-- optional; stories this one must be built after. Omit when independent. -->
|
|
138
139
|
- design: <path(s) to the wireframe(s) covering this story — omit if it has no visual surface>
|
|
139
140
|
- Story: <one user-observable capability> slug: <ado-__ | kebab>
|
|
140
141
|
- AC: <…>
|
|
@@ -171,6 +172,24 @@ story it belongs to, so `tiny-spec-design` knows which files to open for that sp
|
|
|
171
172
|
instead of re-deriving the mapping from filenames. A wireframe covering several stories
|
|
172
173
|
is listed on each of them.
|
|
173
174
|
|
|
175
|
+
**`needs:` is what lets stories build in parallel — so keep it honest and keep it rare.**
|
|
176
|
+
It names the stories this one must be built *after*. `tiny-spec-run` uses it to group the
|
|
177
|
+
backlog into batches: stories with no unmet `needs:` build concurrently, each in its own
|
|
178
|
+
git worktree, and the next batch starts once they merge. Three rules:
|
|
179
|
+
|
|
180
|
+
- **Name a story only for a real code dependency** — this story reads or edits files that
|
|
181
|
+
one writes, or builds on a schema, API, or type it introduces. A shared theme, a shared
|
|
182
|
+
`## Feature:` heading, or "it feels like it comes second" is **not** a dependency.
|
|
183
|
+
- **When unsure, omit it.** The two errors are not symmetric. A `needs:` you didn't need
|
|
184
|
+
costs parallelism on every future run of this backlog, silently and forever. A `needs:`
|
|
185
|
+
you missed costs one merge conflict, which `tiny-spec-run` already detects and halts on.
|
|
186
|
+
- **Never point across a cycle.** If two stories each need the other, they are one story —
|
|
187
|
+
merge them, or split the shared part out into a third that both need.
|
|
188
|
+
|
|
189
|
+
A `## Feature:` heading carries no `needs:` — it is a grouping, not a work item. Most
|
|
190
|
+
backlogs need the field on a minority of stories; a `needs:` on nearly every story means
|
|
191
|
+
the carve is really a sequence, and you should say so rather than encode it here.
|
|
192
|
+
|
|
174
193
|
**Slugs** mirror `tiny-spec-create`: derive from the platform key when bound — ADO
|
|
175
194
|
`AB#77`→`ado-77`, GitHub `#42`→`gh-42`, Monday item→`monday-<id>`, Jira `PROJ-123`
|
|
176
195
|
verbatim; **ad-hoc** → a short kebab-case slug (`email-login`). The tracker id is
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|