@luizsantiago/spec-guardrails 3.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +206 -0
- package/index.js +335 -0
- package/lib/archive.js +208 -0
- package/lib/assets.js +145 -0
- package/lib/brownfield.js +446 -0
- package/lib/config.js +293 -0
- package/lib/constants.js +262 -0
- package/lib/cursorrules.js +92 -0
- package/lib/delta-merge.js +248 -0
- package/lib/doctor.js +343 -0
- package/lib/download.js +133 -0
- package/lib/feature.js +272 -0
- package/lib/fs-utils.js +114 -0
- package/lib/gates.js +138 -0
- package/lib/install.js +140 -0
- package/lib/memory.js +34 -0
- package/lib/next-steps.js +50 -0
- package/lib/presets.js +176 -0
- package/lib/project-rules.js +210 -0
- package/lib/specs-utils.js +117 -0
- package/lib/token-cost.js +124 -0
- package/package.json +46 -0
- package/rules/engineering-baseline.mdc +56 -0
- package/scripts/_common.py +356 -0
- package/scripts/analyze_artifacts.py +187 -0
- package/scripts/check_commit.py +140 -0
- package/scripts/lessons.py +447 -0
- package/scripts/loop_plan.py +217 -0
- package/scripts/validate_spec.py +345 -0
- package/scripts/validate_state.py +385 -0
- package/scripts/validate_tasks.py +379 -0
- package/skills/agent-architecture.md +221 -0
- package/skills/appsec.md +83 -0
- package/skills/code-simplify.md +49 -0
- package/skills/engineering-standards.md +98 -0
- package/skills/git-handoff.md +213 -0
- package/skills/qa-strategy.md +83 -0
- package/skills/references/analyze.md +56 -0
- package/skills/references/archive.md +60 -0
- package/skills/references/constitution.md +66 -0
- package/skills/references/context-limits.md +73 -0
- package/skills/references/converge.md +47 -0
- package/skills/references/design.md +88 -0
- package/skills/references/discuss.md +68 -0
- package/skills/references/explore.md +61 -0
- package/skills/references/implement.md +175 -0
- package/skills/references/lessons.md +71 -0
- package/skills/references/memory.md +98 -0
- package/skills/references/project-init.md +62 -0
- package/skills/references/quick-mode.md +84 -0
- package/skills/references/specify.md +144 -0
- package/skills/references/sub-agents.md +117 -0
- package/skills/references/tasks.md +178 -0
- package/skills/references/validate.md +210 -0
- package/skills/security-review.md +120 -0
- package/skills/ship-ready.md +50 -0
- package/skills/task-graph-engineering.md +180 -0
- package/templates/GETTING_STARTED.md +61 -0
- package/templates/config.yaml.example +28 -0
- package/templates/presets/default.yaml +16 -0
- package/templates/presets/node-ts.yaml +22 -0
- package/templates/presets/python.yaml +22 -0
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# Context Limits
|
|
2
|
+
|
|
3
|
+
How much to load, in what order, and when to stop. Cross-cutting: every phase reads this when the session is long or the feature is large.
|
|
4
|
+
|
|
5
|
+
The hub already says read a reference completely before acting on it. This file is the budget for everything else — specs, diffs, sister skills, and prior features.
|
|
6
|
+
|
|
7
|
+
## Budget
|
|
8
|
+
|
|
9
|
+
Treat context as a working set, not an archive. Prefer a complete read of a small set over a partial read of a large set.
|
|
10
|
+
|
|
11
|
+
| Slot | Load | Do not load |
|
|
12
|
+
| --- | --- | --- |
|
|
13
|
+
| Contract | Hub + the current phase reference + the sister skill that table names | The other phase references |
|
|
14
|
+
| Feature | This feature's `spec.md`; `context.md` / `design.md` / `tasks.md` only when that phase ran | Sibling feature specs |
|
|
15
|
+
| Memory | `STATE.md` Next Step, Blockers, and `AD-NNN` that constrain this area; `lessons.py list --status confirmed` | Candidates, quarantined entries, other features' `validation.md` |
|
|
16
|
+
| Code | Files named on the current task | The rest of the module "for orientation" |
|
|
17
|
+
| Verify | Spec + diff range + the tests the spec names; `security-review.md` | The author's session notes |
|
|
18
|
+
| Conditional sister | **At most one** of `appsec.md`, `qa-strategy.md`, `code-simplify.md`, or `ship-ready.md`, and only when that skill’s trigger says so | Two or more at once; AppSec/QA on Quick/Simple without a trigger; ship-ready during normal Verify |
|
|
19
|
+
|
|
20
|
+
If the working set no longer fits, drop code search leftovers first, then prior-phase artifacts you have already turned into the current artifact, then sister skills that are not in the phase map cell. Drop a finished conditional sister before loading the next (e.g. AppSec before QA; never hold simplify or ship with AppSec/QA).
|
|
21
|
+
|
|
22
|
+
## Load order
|
|
23
|
+
|
|
24
|
+
1. Reconcile `STATE.md` against git (`memory.md`).
|
|
25
|
+
2. Open the hub only if the phase or the router is in doubt.
|
|
26
|
+
3. Open the current phase reference completely.
|
|
27
|
+
4. Open the sister skill the phase map names, if any.
|
|
28
|
+
5. On Verify only: if an AppSec trigger fired, open `appsec.md`, act, then drop it before opening `qa-strategy.md` (never both). Skip both on Quick/Simple without triggers. Do **not** auto-load `ship-ready.md` here.
|
|
29
|
+
6. On Execute (Medium+ after A–D, or owner ask): optionally open **only** `code-simplify.md`, then drop it.
|
|
30
|
+
7. On explicit ship/deploy ask (after Verify PASS): open **only** `ship-ready.md`.
|
|
31
|
+
8. Open this feature's artifacts for the current phase — never two features at once.
|
|
32
|
+
9. Open source files the current task lists.
|
|
33
|
+
|
|
34
|
+
Skip a step when its output is already in the working set from earlier in the same session.
|
|
35
|
+
|
|
36
|
+
## Artifact size
|
|
37
|
+
|
|
38
|
+
Write artifacts so they can be loaded whole.
|
|
39
|
+
|
|
40
|
+
| Artifact | Soft limit | If it overflows |
|
|
41
|
+
| --- | --- | --- |
|
|
42
|
+
| `spec.md` | ~150 lines | Split a second feature; do not hide requirements in `context.md` |
|
|
43
|
+
| `design.md` | ~150 lines | Link to existing code instead of pasting it |
|
|
44
|
+
| `tasks.md` | ~200 lines | Group under `### Phase N`; keep fields to one line each |
|
|
45
|
+
| `validation.md` | ~150 lines | One row per criterion and per mutant; no log dumps |
|
|
46
|
+
| `STATE.md` | keep Next Step to one item | Archive resolved blockers; never let it become a diary |
|
|
47
|
+
| Phase reference | ~220 lines | Cut an example before cutting a rule |
|
|
48
|
+
|
|
49
|
+
These are authoring limits, not gate checks. The gates cannot count lines; you can.
|
|
50
|
+
|
|
51
|
+
## Session rules
|
|
52
|
+
|
|
53
|
+
- **One feature in focus.** Loading a second spec to "stay consistent" is how the wrong ID lands in `tasks.md`.
|
|
54
|
+
- **Do not reload a reference you already followed** in this session unless a gate failed and you need the checklist again.
|
|
55
|
+
- **Search is not loading.** A grep hit is a pointer; read the file only when the current task names it or the spec cites it.
|
|
56
|
+
- **Sub-agents get a subset.** A worker receives the task, the spec IDs it serves, and its `Files` list — not the whole hub and not sibling tasks.
|
|
57
|
+
- **Verify starts empty.** The verifier does not inherit the author's working set. That is the point of Author ≠ verifier.
|
|
58
|
+
|
|
59
|
+
## When the budget is already blown
|
|
60
|
+
|
|
61
|
+
1. Stop generating. Write what you know is missing into `STATE.md` Next Step.
|
|
62
|
+
2. Drop everything that is not the current task's files and the current phase reference.
|
|
63
|
+
3. Resume from the phase reference, not from memory of a previous plan.
|
|
64
|
+
|
|
65
|
+
A confused session produces a confused spec. Resetting context is cheaper than a wrong `REQ`.
|
|
66
|
+
|
|
67
|
+
## Related
|
|
68
|
+
|
|
69
|
+
- `agent-architecture.md` — phase map (which reference belongs to which phase); conditional AppSec / QA
|
|
70
|
+
- `appsec.md` / `qa-strategy.md` — conditional sisters (one at a time on Verify)
|
|
71
|
+
- `memory.md` — resume protocol (what to reconcile before loading)
|
|
72
|
+
- `task-graph-engineering.md` — what a sub-agent is allowed to receive
|
|
73
|
+
- `references/implement.md` — per-task cycle (load only the current task's files)
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# Converge
|
|
2
|
+
|
|
3
|
+
Reassess the codebase against spec/plan/tasks and append remaining work.
|
|
4
|
+
|
|
5
|
+
## When to Use
|
|
6
|
+
|
|
7
|
+
- Long Execute session — drift suspected
|
|
8
|
+
- Owner asks "what's left?"
|
|
9
|
+
- Partial implementation landed outside the task list
|
|
10
|
+
- After merging upstream changes into a feature branch
|
|
11
|
+
|
|
12
|
+
## When NOT to Use
|
|
13
|
+
|
|
14
|
+
- Before Tasks exist
|
|
15
|
+
- During Verify (use `validate.md` instead)
|
|
16
|
+
|
|
17
|
+
## Inputs
|
|
18
|
+
|
|
19
|
+
- Current feature artifacts (`spec.md`, `tasks.md`, `design.md`)
|
|
20
|
+
- Git diff and test status
|
|
21
|
+
- `.specs/STATE.md`
|
|
22
|
+
|
|
23
|
+
## Output
|
|
24
|
+
|
|
25
|
+
Updated `tasks.md` with new tasks for uncovered work (append only — do not rewrite completed tasks).
|
|
26
|
+
|
|
27
|
+
## Procedure
|
|
28
|
+
|
|
29
|
+
1. **Re-read the spec** — list every REQ and whether tests exist for its outcome.
|
|
30
|
+
2. **Audit the diff** — files changed outside task `Files` fields are scope drift or missing tasks.
|
|
31
|
+
3. **Run analyze** — `analyze_artifacts.py [feature]`.
|
|
32
|
+
4. **Append tasks** for gaps using the standard task template in `tasks.md`.
|
|
33
|
+
5. **Update STATE** Next Step to the first open task.
|
|
34
|
+
6. **Commit** `.specs/` changes (Tier 0) — no push.
|
|
35
|
+
|
|
36
|
+
## Rules
|
|
37
|
+
|
|
38
|
+
- Never delete completed task checkboxes.
|
|
39
|
+
- Never weaken tests to match partial implementation.
|
|
40
|
+
- New tasks need Requirement, Files, Depends on, Tests, Gate, Done when.
|
|
41
|
+
- Re-run `validate_tasks.py` after editing tasks.
|
|
42
|
+
|
|
43
|
+
## Next
|
|
44
|
+
|
|
45
|
+
- Resume → `implement.md`
|
|
46
|
+
- Spec itself wrong → `specify.md` (delta spec for brownfield)
|
|
47
|
+
- Back → `agent-architecture.md`
|
|
@@ -0,0 +1,88 @@
|
|
|
1
|
+
# Design
|
|
2
|
+
|
|
3
|
+
Define HOW to build it: architecture, components, reuse, and risk. Optional phase.
|
|
4
|
+
|
|
5
|
+
## When to Use
|
|
6
|
+
|
|
7
|
+
- New architecture, new API surface, or infrastructure change
|
|
8
|
+
- Unfamiliar technology or a pattern the codebase does not have yet
|
|
9
|
+
- More than one defensible approach with real trade-offs
|
|
10
|
+
|
|
11
|
+
## When NOT to Use
|
|
12
|
+
|
|
13
|
+
- Straightforward changes with no architectural decision — design inline during Execute
|
|
14
|
+
- Bug fixes, copy changes, config tweaks
|
|
15
|
+
|
|
16
|
+
Skipping Design is the default for Simple and Medium tiers.
|
|
17
|
+
|
|
18
|
+
## Inputs
|
|
19
|
+
|
|
20
|
+
- Approved `spec.md`
|
|
21
|
+
- `context.md` when Discuss ran
|
|
22
|
+
- `.specs/STATE.md` decisions (`AD-NNN`)
|
|
23
|
+
- Confirmed lessons: `python3 .specs/guardrails/scripts/lessons.py list --status confirmed`
|
|
24
|
+
- Existing codebase structure and conventions
|
|
25
|
+
- `context-limits.md` — this feature only
|
|
26
|
+
|
|
27
|
+
## Output
|
|
28
|
+
|
|
29
|
+
`.specs/features/[feature]/design.md`
|
|
30
|
+
|
|
31
|
+
## Procedure
|
|
32
|
+
|
|
33
|
+
1. **Map what already exists.** List the components you will reuse before proposing new ones. Reuse beats invention.
|
|
34
|
+
2. **Load confirmed lessons** that constrain this design (`lessons.py list --status confirmed`).
|
|
35
|
+
3. **Follow the Knowledge Verification Chain** for anything unfamiliar: codebase → project docs → MCP/Context → web search → flag uncertainty. Never fabricate an API.
|
|
36
|
+
4. **Choose an approach and state it definitively.** Record the alternatives you rejected and why — that is the part future readers need.
|
|
37
|
+
5. **Draw the shape.** A small component table and a step-by-step data-flow table beat three vague paragraphs.
|
|
38
|
+
6. **Name the risks** and the mitigation for each. A risk without a mitigation is a blocker.
|
|
39
|
+
7. **Link every component back to requirement IDs.** Anything that serves no `REQ` is scope creep.
|
|
40
|
+
8. **Promote project-wide decisions** to `STATE.md` as `AD-NNN` (see `memory.md`).
|
|
41
|
+
|
|
42
|
+
## Template
|
|
43
|
+
|
|
44
|
+
```markdown
|
|
45
|
+
# Design: [Feature]
|
|
46
|
+
|
|
47
|
+
## Approach
|
|
48
|
+
[The chosen approach in two or three sentences.]
|
|
49
|
+
|
|
50
|
+
## Components
|
|
51
|
+
|
|
52
|
+
| Component | Responsibility | New or reuse | Serves |
|
|
53
|
+
| --- | --- | --- | --- |
|
|
54
|
+
| [name] | [what it does] | reuse `src/...` | REQ-001 |
|
|
55
|
+
|
|
56
|
+
## Data Flow
|
|
57
|
+
|
|
58
|
+
| Step | From | To | Payload / notes |
|
|
59
|
+
| ---: | --- | --- | --- |
|
|
60
|
+
| 1 | Client | API | [request] |
|
|
61
|
+
| 2 | API | Service | [validated input] |
|
|
62
|
+
| 3 | Service | Store | [persistence] |
|
|
63
|
+
|
|
64
|
+
## Decisions
|
|
65
|
+
|
|
66
|
+
### AD-00X: [Decision]
|
|
67
|
+
- **Chosen**: [option]
|
|
68
|
+
- **Rejected**: [option] because [reason]
|
|
69
|
+
|
|
70
|
+
## Risks
|
|
71
|
+
|
|
72
|
+
| Risk | Impact | Mitigation |
|
|
73
|
+
| --- | --- | --- |
|
|
74
|
+
| [risk] | [impact] | [mitigation] |
|
|
75
|
+
|
|
76
|
+
## Out of Scope for This Design
|
|
77
|
+
- [explicitly excluded]
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
## Rules
|
|
81
|
+
|
|
82
|
+
- No implementation code in `design.md` — interfaces and signatures only when they clarify a contract.
|
|
83
|
+
- If the design reveals that the spec is wrong or incomplete, stop and update `spec.md` first; re-run the spec gate.
|
|
84
|
+
- Prefer the smallest design that satisfies the spec. Extensibility that no requirement asks for is speculation.
|
|
85
|
+
|
|
86
|
+
## Next
|
|
87
|
+
|
|
88
|
+
`tasks.md` — break the design into atomic tasks.
|
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
# Discuss
|
|
2
|
+
|
|
3
|
+
Resolve gray areas with the owner before they become guesses in code. Runs inside Specify, never as a standalone phase.
|
|
4
|
+
|
|
5
|
+
## When to Use
|
|
6
|
+
|
|
7
|
+
Trigger automatically when the feature has any of these dimensions and the spec does not already answer them:
|
|
8
|
+
|
|
9
|
+
| Dimension | Typical gray area |
|
|
10
|
+
| --- | --- |
|
|
11
|
+
| Persistence / state | What survives a refresh, a restart, a logout? |
|
|
12
|
+
| External calls | Timeout, retry, and failure behavior |
|
|
13
|
+
| Auth | Who can see or do this; what happens when denied |
|
|
14
|
+
| Payments | Partial failure, refunds, idempotency |
|
|
15
|
+
| Concurrency | Two users or tabs acting at once |
|
|
16
|
+
| State transitions | Which transitions are legal; what is irreversible |
|
|
17
|
+
| User-facing behavior | Empty, loading, and error states |
|
|
18
|
+
|
|
19
|
+
Also trigger when the owner's request has more than one reasonable interpretation.
|
|
20
|
+
|
|
21
|
+
## When NOT to Use
|
|
22
|
+
|
|
23
|
+
- The answer is already in `spec.md`, `STATE.md`, or established codebase convention
|
|
24
|
+
- The decision is purely technical with no owner-visible consequence — that belongs in `design.md`
|
|
25
|
+
|
|
26
|
+
## Inputs
|
|
27
|
+
|
|
28
|
+
- Draft `spec.md`
|
|
29
|
+
- `.specs/STATE.md` decisions
|
|
30
|
+
|
|
31
|
+
## Output
|
|
32
|
+
|
|
33
|
+
`.specs/features/[feature]/context.md` — created only when Discuss actually runs.
|
|
34
|
+
|
|
35
|
+
## Procedure
|
|
36
|
+
|
|
37
|
+
1. **List the gray areas** you found, grouped by dimension. Keep it short — no more than five at a time.
|
|
38
|
+
2. **Ask one gray area at a time** (unless the owner asks for a batch).
|
|
39
|
+
3. **Offer concrete options**, not open questions. "Should the session expire after 24h, 7d, or on browser close?" beats "How should sessions work?"
|
|
40
|
+
4. **State your recommendation and why.** The owner should be able to answer "yes" and move on.
|
|
41
|
+
5. **Wait for an explicit yes** or option letter before writing `context.md` — soft “sounds fine” is not enough.
|
|
42
|
+
6. **Record the decision verbatim** in `context.md`, with the rationale and the date.
|
|
43
|
+
7. **Fold consequences back into `spec.md`** as acceptance criteria — `context.md` records the decision; the spec records the testable outcome.
|
|
44
|
+
8. **Escalate irreversible choices.** Anything expensive to undo gets an explicit confirmation (see the human gate in `task-graph-engineering.md`).
|
|
45
|
+
|
|
46
|
+
## Template
|
|
47
|
+
|
|
48
|
+
```markdown
|
|
49
|
+
# Context: [Feature]
|
|
50
|
+
|
|
51
|
+
## D-001: [Question in one line]
|
|
52
|
+
- **Options considered**: A) ... B) ... C) ...
|
|
53
|
+
- **Decision**: [what the owner chose]
|
|
54
|
+
- **Rationale**: [why]
|
|
55
|
+
- **Consequences**: [which REQ IDs this affects]
|
|
56
|
+
- **Date**: [ISO date]
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
## Rules
|
|
60
|
+
|
|
61
|
+
- Write `context.md` in **English**, whatever language the conversation happens in (see `engineering-standards.md`).
|
|
62
|
+
- Never invent an answer to keep momentum. An unanswered gray area is a blocker, not a detail.
|
|
63
|
+
- If the owner defers a decision, record it as an open question in `STATE.md` under Blockers.
|
|
64
|
+
- Project-wide decisions graduate to `STATE.md` as `AD-NNN` (see `memory.md`).
|
|
65
|
+
|
|
66
|
+
## Next
|
|
67
|
+
|
|
68
|
+
Return to `specify.md`, update the spec, and re-run the spec gate.
|
|
@@ -0,0 +1,61 @@
|
|
|
1
|
+
# Explore
|
|
2
|
+
|
|
3
|
+
Think through an idea before committing to a spec. No artifacts required.
|
|
4
|
+
|
|
5
|
+
## When to Use
|
|
6
|
+
|
|
7
|
+
- The goal is fuzzy ("make auth better", "add dark mode somehow")
|
|
8
|
+
- You need to compare options against the existing codebase
|
|
9
|
+
- The owner is not sure what to build yet
|
|
10
|
+
- Before `/specify` on any non-Quick change
|
|
11
|
+
|
|
12
|
+
## When NOT to Use
|
|
13
|
+
|
|
14
|
+
- The owner already gave a clear, testable goal — go to `/specify`
|
|
15
|
+
- Quick-tier fixes (≤3 files) — use `quick-mode.md`
|
|
16
|
+
|
|
17
|
+
## Inputs
|
|
18
|
+
|
|
19
|
+
- Owner's question or idea, in their own words
|
|
20
|
+
- Existing codebase (read relevant modules only)
|
|
21
|
+
- `.specs/project/CONSTITUTION.md` when it exists
|
|
22
|
+
- `.specs/project/PROJECT.md` when it exists
|
|
23
|
+
- `context-limits.md` — do not load sibling feature specs
|
|
24
|
+
|
|
25
|
+
## Output
|
|
26
|
+
|
|
27
|
+
None required. Optionally capture decisions in chat. When the idea crystallizes, transition to `/specify`.
|
|
28
|
+
|
|
29
|
+
## Procedure
|
|
30
|
+
|
|
31
|
+
1. **Read the code that matters.** Skim the area the idea touches; name files and patterns you found.
|
|
32
|
+
2. **Restate the problem** in one sentence and confirm it with the owner.
|
|
33
|
+
3. **Offer 2–3 options** with trade-offs (complexity, dependencies, risk). Recommend one.
|
|
34
|
+
4. **Surface unknowns** with `[NEEDS CLARIFICATION: specific question]` — do not guess.
|
|
35
|
+
5. **Estimate complexity** using the hub router (Quick / Simple / Medium / Complex).
|
|
36
|
+
6. **Transition** when the owner picks a direction:
|
|
37
|
+
```bash
|
|
38
|
+
npx @luizsantiago/spec-guardrails feature-init "chat with presence"
|
|
39
|
+
```
|
|
40
|
+
Then open `references/specify.md`.
|
|
41
|
+
|
|
42
|
+
## Rules
|
|
43
|
+
|
|
44
|
+
- No production code during Explore.
|
|
45
|
+
- No `spec.md` until `/specify` runs (after `feature-init`).
|
|
46
|
+
- No scaffolding empty `design.md` or `tasks.md`.
|
|
47
|
+
- Keep the working set small — hub + this file + targeted code reads.
|
|
48
|
+
|
|
49
|
+
## Anti-Patterns
|
|
50
|
+
|
|
51
|
+
| Avoid | Prefer |
|
|
52
|
+
| --- | --- |
|
|
53
|
+
| Jumping to implementation | Options + recommendation first |
|
|
54
|
+
| Writing a full spec during Explore | Transition to Specify when scope is clear |
|
|
55
|
+
| Guessing auth/data choices | `[NEEDS CLARIFICATION: ...]` markers |
|
|
56
|
+
|
|
57
|
+
## Next
|
|
58
|
+
|
|
59
|
+
- Scope is clear → `feature-init` then `specify.md`
|
|
60
|
+
- Gray areas remain → `discuss.md` inside Specify
|
|
61
|
+
- Back → `agent-architecture.md`
|
|
@@ -0,0 +1,175 @@
|
|
|
1
|
+
# Implement (Execute / Loop)
|
|
2
|
+
|
|
3
|
+
Orchestrate tasks from `tasks.md` and `task-graph.md`: **parallel waves with sub-agents** when files are disjoint, otherwise **one task at a time** inline. Always required after Specify (and Tasks, when it ran).
|
|
4
|
+
|
|
5
|
+
## When to Use
|
|
6
|
+
|
|
7
|
+
- Every feature after Specify (and Tasks, when it ran)
|
|
8
|
+
|
|
9
|
+
## Inputs
|
|
10
|
+
|
|
11
|
+
- Approved `spec.md`, plus `design.md`, `tasks.md`, `task-graph.md` when they exist
|
|
12
|
+
- Confirmed lessons: `python3 .specs/guardrails/scripts/lessons.py list --status confirmed`
|
|
13
|
+
- `engineering-standards.md` for code quality and artifact language
|
|
14
|
+
- `context-limits.md` — load only this feature and the files the current task names
|
|
15
|
+
|
|
16
|
+
## Outputs
|
|
17
|
+
|
|
18
|
+
- Production code and tests
|
|
19
|
+
- One atomic commit per task
|
|
20
|
+
- Updated `tasks.md` checkboxes
|
|
21
|
+
|
|
22
|
+
## Before Starting
|
|
23
|
+
|
|
24
|
+
1. Read this file completely.
|
|
25
|
+
2. **Discover this repo’s test command** from `package.json`, `Makefile`, CI, or README — prefer the focused command the task `Gate` names. Do not assume `npm test` if the project uses another runner.
|
|
26
|
+
3. Run `python3 .specs/guardrails/scripts/validate_tasks.py` when a formal `tasks.md` exists.
|
|
27
|
+
4. If Tasks was skipped, list the atomic steps inline now. More than 5 steps or real dependencies means the Tasks phase was skipped in error — stop and create `tasks.md`.
|
|
28
|
+
5. **Plan the wave** — run `python3 .specs/guardrails/scripts/loop_plan.py [feature]` (or `loop-plan --json`) at the start of Execute and after every batch completes. It lists the next runnable tasks and marks **parallel groups** (disjoint `Files`) vs inline work.
|
|
29
|
+
6. When `loop-plan` shows a **parallel group** (2+ tasks), offer sub-agent dispatch per `task-graph-engineering.md` and `sub-agents.md`. Offer and wait; never auto-spawn. Large features (roughly 8+ tasks total) also warrant batching across waves.
|
|
30
|
+
7. Confirm you are the only writer for each file this task names. Two parallel tasks never share a file in the same round.
|
|
31
|
+
|
|
32
|
+
## Orchestration (each /loop round)
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
loop-plan → dispatch (parallel sub-agents | inline) → merge → loop-plan → … → /verify
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
1. **loop-plan** — Read the next wave: which tasks have satisfied dependencies and disjoint file ownership.
|
|
39
|
+
2. **Dispatch**
|
|
40
|
+
- **Parallel group (2+ tasks):** one sub-agent per task with the worker brief in `sub-agents.md`. Cap at 3 workers + 1 verifier (see `task-graph-engineering.md`).
|
|
41
|
+
- **Single task:** run the per-task cycle below inline (or as sole worker).
|
|
42
|
+
3. **Merge** — After a parallel round, confirm every task is `[x]`, commits exist, and the project harness passes once on the integrated tree.
|
|
43
|
+
4. **Repeat** — Run `loop-plan` again until all tasks are complete, then close Execute.
|
|
44
|
+
|
|
45
|
+
Do not start the next wave until the current wave is fully committed. Parallelism is **inside** a wave only when `Files` do not overlap.
|
|
46
|
+
|
|
47
|
+
## Per-Task Cycle
|
|
48
|
+
|
|
49
|
+
```
|
|
50
|
+
Plan → Test → Implement → Gate → Commit → Next
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
1. **Plan** — Restate the task's `Done when` criterion and the files you will touch. Nothing else gets touched.
|
|
54
|
+
2. **Test first** — Write the test derived from the acceptance criteria. It must fail for the right reason before you write production code.
|
|
55
|
+
3. **Implement** — The smallest change that makes the test pass, following the conventions already in the codebase.
|
|
56
|
+
4. **Gate** — Run the task's `Gate` command. The runner decides, not your judgment. On failure, follow the playbook below.
|
|
57
|
+
5. **Mark complete** — Check the task box in `tasks.md` in the same change set.
|
|
58
|
+
6. **Commit** — One atomic commit including the code, the tests, and the `tasks.md` update — only after Adequacy A–D pass.
|
|
59
|
+
|
|
60
|
+
```bash
|
|
61
|
+
python3 .specs/guardrails/scripts/check_commit.py --message "feat(auth): add token refresh"
|
|
62
|
+
git add [files] .specs/features/[feature]/tasks.md
|
|
63
|
+
git commit -m "feat(auth): add token refresh"
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
## Adequacy (before commit)
|
|
67
|
+
|
|
68
|
+
Verifier judgment, not a structural gate. If any check is No, do not commit — continue the cycle or escalate.
|
|
69
|
+
|
|
70
|
+
| Check | Ask |
|
|
71
|
+
| --- | --- |
|
|
72
|
+
| **A Outcome** | Does the new test assert this task’s `Done when` / spec criterion (not merely that code ran)? |
|
|
73
|
+
| **B Scope** | Do `git status` and the task `Files` list agree — no sibling files in the index? |
|
|
74
|
+
| **C Gate** | Did this task’s `Gate` command pass? |
|
|
75
|
+
| **D Spec** | No extra behavior outside the spec? If the spec is wrong, stop and follow Spec Deviation (`SPEC_DEVIATION`) — do not adapt in code. |
|
|
76
|
+
|
|
77
|
+
A gate that passes while A or D fails is not done.
|
|
78
|
+
|
|
79
|
+
**Anti-rationalizations (judgment):** “test after code” → no, RED first (**A**). “Gate green = done” → still need **A**/**D**. “Nearby sibling files” → stay in `Files` (**B**). “Spec is slightly wrong” → Spec Deviation, never silent adapt (**D**).
|
|
80
|
+
|
|
81
|
+
**Optional simplify.** On Medium+ after A–D (or when the owner asks), offer to load **only** `code-simplify.md`, then drop it — never with AppSec/QA/ship-ready. Judgment; not gated.
|
|
82
|
+
|
|
83
|
+
## Gate Failure Playbook
|
|
84
|
+
|
|
85
|
+
A failing gate is information. Do not skip, delete, or loosen a test to make it pass.
|
|
86
|
+
|
|
87
|
+
| Failure | First move | Then |
|
|
88
|
+
| --- | --- | --- |
|
|
89
|
+
| New test fails for the wrong reason | Fix the test until it asserts the spec outcome and fails because the production code is missing | Re-run |
|
|
90
|
+
| New test fails for the right reason | Implement the smallest production change | Re-run |
|
|
91
|
+
| Existing test regresses | Revert unrelated edits; keep the diff inside the task's `Files` list | Re-run |
|
|
92
|
+
| Lint / types | Fix only what this task introduced | Re-run |
|
|
93
|
+
| `check_commit.py` | Rewrite the message; do not commit with `--no-verify` | Re-run |
|
|
94
|
+
| Flaky or order-dependent test | Isolate the external dependency; never retry-until-green | Re-run |
|
|
95
|
+
|
|
96
|
+
Retry the same task up to **3 times**. After the third failure, stop the loop and escalate to the owner with: the command that failed, the last error, the files touched, and the option you recommend.
|
|
97
|
+
|
|
98
|
+
A gate that passes while `Done when` is still false is not done. Write the missing assertion and continue the cycle.
|
|
99
|
+
|
|
100
|
+
**Right-reason test.** The new test must fail because the production behavior is absent, not because of a typo, a wrong import, or an assertion that cannot succeed. Read the failure once. If the message is about setup, fix the test; if it is the spec outcome (401 vs 200, missing field, wrong code), then implement.
|
|
101
|
+
|
|
102
|
+
**Commit checklist.** Before `git commit`: the task box is checked, `check_commit.py` has passed on the exact message, `Files` in the task and `git status` agree, and no sibling file is in the index. An unchecked box in a "done" commit fails `validate_state.py` later and wastes a verify round.
|
|
103
|
+
|
|
104
|
+
## Spec Deviation
|
|
105
|
+
|
|
106
|
+
Implementation sometimes proves the spec wrong — an ID that cannot be satisfied, an outcome the platform cannot produce, a missing error path.
|
|
107
|
+
|
|
108
|
+
1. **Stop the loop.** Do not "adapt" the spec in code.
|
|
109
|
+
2. Update `spec.md`, keep the original requirement ID, and add a new ID only for genuinely new behavior.
|
|
110
|
+
3. Record the change in `STATE.md` under Decisions or Deferred Ideas.
|
|
111
|
+
4. Re-run `validate_spec.py`. Re-derive affected tests.
|
|
112
|
+
5. Resume Execute only after the owner approves the delta.
|
|
113
|
+
|
|
114
|
+
A silent spec change during Execute is a process failure, not a shortcut.
|
|
115
|
+
|
|
116
|
+
## Parallel Task Conflicts
|
|
117
|
+
|
|
118
|
+
When this task is part of a parallel group (see `task-graph.md`):
|
|
119
|
+
|
|
120
|
+
- Touch only the files listed on this task. A file that another in-flight task owns is out of bounds, even for an import fix — leave a note in `STATE.md` instead.
|
|
121
|
+
- If you discover the split was wrong (two tasks must edit the same file), stop both, merge them into one sequential task, and re-run `validate_tasks.py`.
|
|
122
|
+
- Do not "just finish" a sibling's work because you are already in the module. That is how one-writer-per-file breaks.
|
|
123
|
+
|
|
124
|
+
## When to Stop and Escalate
|
|
125
|
+
|
|
126
|
+
Stop immediately, with a written recommendation, when any of these are true:
|
|
127
|
+
|
|
128
|
+
- The third gate retry still fails
|
|
129
|
+
- The work needs a file this task does not own, and the owner of that file is another in-flight task
|
|
130
|
+
- A gray area appears that Specify did not settle — return to `discuss.md`, do not guess
|
|
131
|
+
- Auth, payments, or data destruction is required and was not in the approved spec
|
|
132
|
+
- The remaining work is no longer the approved task list (scope doubled, architecture flipped)
|
|
133
|
+
|
|
134
|
+
Escalation is a handoff, not a pause to keep coding. Update `STATE.md` Next Step to the decision you need, then wait.
|
|
135
|
+
|
|
136
|
+
## Mid-loop Discoveries
|
|
137
|
+
|
|
138
|
+
| Discovery | Action |
|
|
139
|
+
| --- | --- |
|
|
140
|
+
| Extra file required by this task | Amend `Files` on the current task, confirm no sibling owns it, continue |
|
|
141
|
+
| Extra behavior required by the spec | New task, or a spec delta — not a drive-by in this commit |
|
|
142
|
+
| Better design than `design.md` | Record in `STATE.md` Deferred Ideas; do not redesign mid-task |
|
|
143
|
+
| Dead code adjacent to the change | Leave it unless the task's `Done when` names it |
|
|
144
|
+
| Test harness cannot express the criterion | Escalate; do not lower the criterion to what is easy to test |
|
|
145
|
+
|
|
146
|
+
## Rules
|
|
147
|
+
|
|
148
|
+
- **Surgical changes** — Touch only what the task requires. No drive-by refactors.
|
|
149
|
+
- **No scope creep** — Good ideas that are not in the task go to `STATE.md` under Deferred Ideas, not into the diff.
|
|
150
|
+
- **One writer per file** — Two parallel tasks never mutate the same file in the same round.
|
|
151
|
+
- **Never weaken tests** — Do not skip, delete, or loosen a test to make a gate pass.
|
|
152
|
+
- **Blast radius** — Local commits are authorized by task approval. `git push`, deploy, and destructive operations need an explicit go-ahead.
|
|
153
|
+
|
|
154
|
+
## Commit Format
|
|
155
|
+
|
|
156
|
+
Conventional Commits, English, one concern per commit:
|
|
157
|
+
|
|
158
|
+
```
|
|
159
|
+
feat(auth): add token refresh
|
|
160
|
+
fix(cart): prevent negative quantity on decrement
|
|
161
|
+
test(auth): cover session expiry edge case
|
|
162
|
+
docs(spec): record validation report for auth
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
## Closing Execute
|
|
166
|
+
|
|
167
|
+
When the last task is complete:
|
|
168
|
+
|
|
169
|
+
1. Run the full project harness once more (tests, linter, build).
|
|
170
|
+
2. Trigger `/verify` with a fresh context — mandatory, never prompted. See `validate.md`.
|
|
171
|
+
3. Do not declare the feature done until `validate_state.py` passes.
|
|
172
|
+
|
|
173
|
+
## Next
|
|
174
|
+
|
|
175
|
+
`validate.md` — independent verification.
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
# Lessons
|
|
2
|
+
|
|
3
|
+
How to record, promote, and load grounded lessons. The engine owns the files; this reference owns the judgment.
|
|
4
|
+
|
|
5
|
+
## When to Use
|
|
6
|
+
|
|
7
|
+
- The last step of `/verify` when the verdict is FAIL for a grounded reason
|
|
8
|
+
- The first step of `/specify` and `/plan`, loading only confirmed lessons
|
|
9
|
+
- After a confirmed lesson was loaded and the same failure still happened (`penalize`)
|
|
10
|
+
|
|
11
|
+
## When NOT to Use
|
|
12
|
+
|
|
13
|
+
- A clean PASS — record nothing
|
|
14
|
+
- A preference, a style nit, or anything not evidenced in `validation.md`
|
|
15
|
+
- Hand-editing `.specs/LESSONS.md` or `.specs/lessons.json`
|
|
16
|
+
|
|
17
|
+
## Files
|
|
18
|
+
|
|
19
|
+
| Path | Owner | Role |
|
|
20
|
+
| --- | --- | --- |
|
|
21
|
+
| `.specs/lessons.json` | `lessons.py` | Canonical store |
|
|
22
|
+
| `.specs/LESSONS.md` | `lessons.py` | Rendered playbook — read, never write |
|
|
23
|
+
| `.specs/guardrails/scripts/lessons.py` | Harness | `add`, `list`, `penalize`, `prune`, `status` |
|
|
24
|
+
|
|
25
|
+
## How to Phrase
|
|
26
|
+
|
|
27
|
+
A lesson is a trigger plus a rule the next agent can apply without the original thread:
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
python3 .specs/guardrails/scripts/lessons.py add \
|
|
31
|
+
--title "Assert error codes, not just status" \
|
|
32
|
+
--trigger "mutant returning 403 instead of 401 survived" \
|
|
33
|
+
--rule "Acceptance criteria must name the error code, and tests must assert it" \
|
|
34
|
+
--source .specs/features/auth/validation.md:41
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
`--source` is mandatory and must point at a non-empty `validation.md`. A lesson without evidence is opinion, and the engine refuses it.
|
|
38
|
+
|
|
39
|
+
Titles and rules are deduped after normalization (casefold, accents stripped, punctuation ignored). Rephrase the same idea and it attaches to the existing ID instead of creating a twin.
|
|
40
|
+
|
|
41
|
+
## Lifecycle
|
|
42
|
+
|
|
43
|
+
```
|
|
44
|
+
FAIL with evidence → add (candidate)
|
|
45
|
+
candidate seen in 2 distinct features → confirmed (guidance)
|
|
46
|
+
confirmed loaded but the same failure recurs → penalize
|
|
47
|
+
2 penalties → quarantined (stop loading)
|
|
48
|
+
candidate idle 90 days → prune
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
Same-feature recurrence does not promote. One noisy feature must not turn a guess into a house rule.
|
|
52
|
+
|
|
53
|
+
## When to Load
|
|
54
|
+
|
|
55
|
+
At the start of Specify and Design:
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
python3 .specs/guardrails/scripts/lessons.py list --status confirmed
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
Apply every confirmed rule that matches this work. Candidates are not guidance — they live in `lessons.json` until a second feature corroborates them.
|
|
62
|
+
|
|
63
|
+
If a confirmed lesson was in the working set and the same gap still appears in `/verify`, penalize it with the new `validation.md` as `--source`. Two penalties quarantine it.
|
|
64
|
+
|
|
65
|
+
## Degraded mode
|
|
66
|
+
|
|
67
|
+
Without Python the engine does not run. Do not write `LESSONS.md` by hand to compensate — that file is generated, and a hand-written entry has no store, no dedup, and no promotion. Note the gap in `STATE.md` and record the lesson once Python is available.
|
|
68
|
+
|
|
69
|
+
## Next
|
|
70
|
+
|
|
71
|
+
Return to `validate.md` after recording, or to `specify.md` / `design.md` after loading.
|