@luizsantiago/spec-guardrails 3.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +206 -0
- package/index.js +335 -0
- package/lib/archive.js +208 -0
- package/lib/assets.js +145 -0
- package/lib/brownfield.js +446 -0
- package/lib/config.js +293 -0
- package/lib/constants.js +262 -0
- package/lib/cursorrules.js +92 -0
- package/lib/delta-merge.js +248 -0
- package/lib/doctor.js +343 -0
- package/lib/download.js +133 -0
- package/lib/feature.js +272 -0
- package/lib/fs-utils.js +114 -0
- package/lib/gates.js +138 -0
- package/lib/install.js +140 -0
- package/lib/memory.js +34 -0
- package/lib/next-steps.js +50 -0
- package/lib/presets.js +176 -0
- package/lib/project-rules.js +210 -0
- package/lib/specs-utils.js +117 -0
- package/lib/token-cost.js +124 -0
- package/package.json +46 -0
- package/rules/engineering-baseline.mdc +56 -0
- package/scripts/_common.py +356 -0
- package/scripts/analyze_artifacts.py +187 -0
- package/scripts/check_commit.py +140 -0
- package/scripts/lessons.py +447 -0
- package/scripts/loop_plan.py +217 -0
- package/scripts/validate_spec.py +345 -0
- package/scripts/validate_state.py +385 -0
- package/scripts/validate_tasks.py +379 -0
- package/skills/agent-architecture.md +221 -0
- package/skills/appsec.md +83 -0
- package/skills/code-simplify.md +49 -0
- package/skills/engineering-standards.md +98 -0
- package/skills/git-handoff.md +213 -0
- package/skills/qa-strategy.md +83 -0
- package/skills/references/analyze.md +56 -0
- package/skills/references/archive.md +60 -0
- package/skills/references/constitution.md +66 -0
- package/skills/references/context-limits.md +73 -0
- package/skills/references/converge.md +47 -0
- package/skills/references/design.md +88 -0
- package/skills/references/discuss.md +68 -0
- package/skills/references/explore.md +61 -0
- package/skills/references/implement.md +175 -0
- package/skills/references/lessons.md +71 -0
- package/skills/references/memory.md +98 -0
- package/skills/references/project-init.md +62 -0
- package/skills/references/quick-mode.md +84 -0
- package/skills/references/specify.md +144 -0
- package/skills/references/sub-agents.md +117 -0
- package/skills/references/tasks.md +178 -0
- package/skills/references/validate.md +210 -0
- package/skills/security-review.md +120 -0
- package/skills/ship-ready.md +50 -0
- package/skills/task-graph-engineering.md +180 -0
- package/templates/GETTING_STARTED.md +61 -0
- package/templates/config.yaml.example +28 -0
- package/templates/presets/default.yaml +16 -0
- package/templates/presets/node-ts.yaml +22 -0
- package/templates/presets/python.yaml +22 -0
|
@@ -0,0 +1,210 @@
|
|
|
1
|
+
# Validate
|
|
2
|
+
|
|
3
|
+
Independent verification of the delivered feature. Always required, never prompted.
|
|
4
|
+
|
|
5
|
+
## When to Use
|
|
6
|
+
|
|
7
|
+
- After the last task of a feature is committed
|
|
8
|
+
- After any fix round that follows a FAIL verdict
|
|
9
|
+
|
|
10
|
+
## Who Runs It
|
|
11
|
+
|
|
12
|
+
A **fresh verifier context that never wrote the code**. Author ≠ verifier is non-negotiable. The verifier re-derives coverage from the spec instead of inheriting the author's mental model.
|
|
13
|
+
|
|
14
|
+
When sub-agents are available, dispatch the verifier as a separate agent (see `task-graph-engineering.md`). Without sub-agents, start a clean context and run this file as a fresh-eyes pass.
|
|
15
|
+
|
|
16
|
+
## Inputs
|
|
17
|
+
|
|
18
|
+
- `spec.md` acceptance criteria, `context.md` when it exists
|
|
19
|
+
- The diff range for the feature
|
|
20
|
+
- `security-review.md` for the security checklist
|
|
21
|
+
- `appsec.md` only when Complex or an attack-surface trigger fires — then **drop** it before QA
|
|
22
|
+
- `qa-strategy.md` only when Complex, multi-step UI, or an explicit regression ask — never together with `appsec.md`
|
|
23
|
+
- Do **not** load `ship-ready.md` or `code-simplify.md` during Verify (ship is owner-triggered after PASS; simplify is Execute-side)
|
|
24
|
+
- `context-limits.md` — load the spec, the diff, and the tests the spec names; do not load the author's chat; at most one conditional sister at a time
|
|
25
|
+
|
|
26
|
+
## Output
|
|
27
|
+
|
|
28
|
+
`.specs/features/[feature]/validation.md`
|
|
29
|
+
|
|
30
|
+
## Procedure
|
|
31
|
+
|
|
32
|
+
### 1. Spec-anchored outcome check
|
|
33
|
+
|
|
34
|
+
For each acceptance criterion, confirm that a test asserts the **spec-defined outcome** — not merely that the code runs. Flag criteria where the test asserts an implementation detail, and flag spec text too imprecise to test.
|
|
35
|
+
|
|
36
|
+
### 2. Discrimination sensor
|
|
37
|
+
|
|
38
|
+
Confirm the tests can actually fail:
|
|
39
|
+
|
|
40
|
+
1. Inject behavior-level faults one at a time in an **isolated scratch copy** — a temp worktree or file copies. Never use `git stash` and never mutate the working tree.
|
|
41
|
+
2. Confirm the relevant test fails for each mutant.
|
|
42
|
+
3. Discard the scratch and verify the real tree is unchanged (`git status --porcelain` matches the pre-sensor baseline).
|
|
43
|
+
4. Any mutant that survives becomes a fix task — the tests do not discriminate.
|
|
44
|
+
|
|
45
|
+
### 3. Security review
|
|
46
|
+
|
|
47
|
+
Run the checklist in `security-review.md`. Features with no auth, API, input, payment, or infrastructure surface may take the documented lightweight path — with the justification written into the report.
|
|
48
|
+
|
|
49
|
+
### 4. Evidence-or-zero
|
|
50
|
+
|
|
51
|
+
A requirement is satisfied only with a `file:line` reference to an assertive test that passes. No reference means not done, regardless of how the code looks.
|
|
52
|
+
|
|
53
|
+
### 5. Conditional AppSec (optional)
|
|
54
|
+
|
|
55
|
+
If the hub AppSec trigger fires, load **only** `appsec.md`, write `## AppSec` (or `skipped — reason`), then **drop** that skill from context. Do not load `qa-strategy.md` yet. If the trigger does not fire, record a one-line skip. Judgment only — not gated.
|
|
56
|
+
|
|
57
|
+
### 6. Conditional QA (optional)
|
|
58
|
+
|
|
59
|
+
If the hub QA trigger fires, load **only** `qa-strategy.md` (after AppSec is done or skipped), write `## QA`, then continue. Judgment only — not gated.
|
|
60
|
+
|
|
61
|
+
### 7. Verdict
|
|
62
|
+
|
|
63
|
+
Write `validation.md`, then run the completion gate.
|
|
64
|
+
|
|
65
|
+
## Mutant catalog
|
|
66
|
+
|
|
67
|
+
Pick mutants from the kind of code that landed, not from a generic list. One killed mutant per risky behavior is the minimum; surviving mutants are fix tasks.
|
|
68
|
+
|
|
69
|
+
| Code kind | Mutant | Expected killer |
|
|
70
|
+
| --- | --- | --- |
|
|
71
|
+
| Auth / session | Skip the expiry check; accept an empty token; invert role comparison | Test that names the denied case |
|
|
72
|
+
| HTTP handler | Return 200 on the error path; swap 401/403; drop the error code body | Status + code assertion |
|
|
73
|
+
| Validation | Remove an upper/lower bound; accept the empty string; skip a required field | Boundary test |
|
|
74
|
+
| Persistence | Skip the unique constraint; write without a transaction; ignore not-found | Duplicate / rollback / 404 test |
|
|
75
|
+
| Payments / money | Off-by-one on minor units; skip idempotency key; apply a refund twice | Amount and replay tests |
|
|
76
|
+
| Concurrency | Drop a lock or compare-and-swap; process the same event twice | Idempotency or conflict test |
|
|
77
|
+
| State machine | Allow a backward transition; skip the terminal-state guard | Illegal-transition test |
|
|
78
|
+
|
|
79
|
+
A mutant that the compiler or typechecker rejects before a test runs does not count. Change behavior, not syntax.
|
|
80
|
+
|
|
81
|
+
**Weak assertion tells.** Treat the test as weak when it asserts any of: `toBeDefined()`, a 2xx status with no body, "the function was called", or a snapshot of an entire module. The spec names an outcome; the test must name it too (status + error code, exact field, exact transition).
|
|
82
|
+
|
|
83
|
+
Do not inject mutants that the spec does not constrain. A surviving mutant of unspecified behavior is a spec gap, not a test gap — send it back to Specify.
|
|
84
|
+
|
|
85
|
+
## Gap catalog
|
|
86
|
+
|
|
87
|
+
Every FAIL names a gap type so the author knows which loop to re-enter.
|
|
88
|
+
|
|
89
|
+
| Gap | Meaning | Return to |
|
|
90
|
+
| --- | --- | --- |
|
|
91
|
+
| Missing evidence | No `file:line` for a criterion | Execute — add the assertion |
|
|
92
|
+
| Weak assertion | Test checks that code ran, not the spec outcome | Execute — rewrite the test |
|
|
93
|
+
| Surviving mutant | Discriminating test is absent | Execute — add the killer test, then the production fix if needed |
|
|
94
|
+
| Imprecise criterion | Spec cannot be tested as written | Specify — rewrite with `SHALL`/`MUST` and a trigger |
|
|
95
|
+
| Spec deviation | Implementation does something the spec forbids, or omits something it requires | Specify + Execute |
|
|
96
|
+
| Security finding | Checklist item failed | Execute — fix, then re-verify |
|
|
97
|
+
| Open task | `tasks.md` still has `- [ ]` | Execute — finish or drop the task with owner approval |
|
|
98
|
+
| Sensor skipped | No mutant outcome in the report | Verify — run the sensor. On **Medium+** (`design.md` with content, 4+ tasks, or 2+ phases) do not pass — the completion gate blocks. Below Medium+ the gate warns (`--strict` promotes); still run the sensor when risk warrants it |
|
|
99
|
+
|
|
100
|
+
Rank gaps: security and spec deviation first, then surviving mutants, then missing evidence, then imprecise criteria. The author fixes in that order.
|
|
101
|
+
|
|
102
|
+
## Evidence format
|
|
103
|
+
|
|
104
|
+
The completion gate searches for `file:line` (for example `test/routes/login.test.ts:24`). A URL, a CI job name, or "covered by the suite" is not evidence.
|
|
105
|
+
|
|
106
|
+
- Cite the assertive test, not the production file
|
|
107
|
+
- One evidence line per criterion; reuse a test only when it truly asserts both outcomes
|
|
108
|
+
- After a fix round, cite the new line numbers — stale citations fail the reader even if they pass the regex
|
|
109
|
+
|
|
110
|
+
## Gate
|
|
111
|
+
|
|
112
|
+
```bash
|
|
113
|
+
python3 .specs/guardrails/scripts/validate_state.py .specs/features/[feature]
|
|
114
|
+
python3 .specs/guardrails/scripts/validate_state.py [feature]
|
|
115
|
+
python3 .specs/guardrails/scripts/validate_state.py # single-feature projects
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
Checks that the report exists, the verdict is exactly PASS in the **preamble** (before the first `##` section) or under a dedicated `## Verdict` / `## Result` / `## Status` heading, every spec requirement ID shares a line with test `file:line` evidence, and no task remains open. A `- Verdict: PASS` buried under Discrimination Sensor or Coverage does not count. Preamble and `## Verdict` must not disagree. Evidence inside fenced samples or HTML comments does not count. `PASS` with any surviving mutant on a sensor/mutant line fails. `PASS` with open `Gaps` bullets or Security Review `Result: fail` fails. On Medium+ features (`design.md` with content, 4+ tasks, or 2+ phases) a discrimination-sensor **outcome** is **blocking** — the section heading alone is not enough — and a Medium+ `PASS` requires at least one `killed` mutant in the sensor focus (`injected` alone, or `killed` only under Gaps, is not enough). Below Medium+ a missing outcome is a warning (`--strict` still promotes warnings). Non-zero exit means the feature is not done.
|
|
119
|
+
|
|
120
|
+
The gate cannot judge whether a cited test actually asserts the criterion. That judgment is the verifier's; a green gate with a weak assertion is still a FAIL in the report.
|
|
121
|
+
|
|
122
|
+
**Gate-enforced vs verifier judgment.** `validate_state.py` enforces form: verdict scope, test-path evidence, REQ↔evidence lines, sensor outcomes on Medium+, open Gaps, Security `Result: fail`, open tasks. The following stay **verifier judgment** (not structural gates): whether each coverage row's test truly asserts the outcome, whether a lightweight Security path is justified, Interactive UAT / walkthrough success, and optional `## AppSec` / `## QA` sections. See the [gate stability contract](https://github.com/luizssantiago92/spec-guardrails/blob/main/prd/gate-stability.md).
|
|
123
|
+
|
|
124
|
+
**Red flags before PASS (judgment):** evidence only in fences/comments; open Gaps with PASS; soft `PASS WITH GAPS`; verdict buried under Sensor/Coverage; config path as evidence; surviving mutant or Medium+ without a kill — these already fail the gate or the reader.
|
|
125
|
+
|
|
126
|
+
## Template
|
|
127
|
+
|
|
128
|
+
```markdown
|
|
129
|
+
# Validation: [Feature]
|
|
130
|
+
|
|
131
|
+
- Verifier: independent agent (clean context)
|
|
132
|
+
- Date: [ISO date]
|
|
133
|
+
- Diff range: [base..head]
|
|
134
|
+
- Verdict: PASS
|
|
135
|
+
|
|
136
|
+
## Coverage
|
|
137
|
+
|
|
138
|
+
| Requirement | Test evidence | Result |
|
|
139
|
+
| --- | --- | --- |
|
|
140
|
+
| REQ-001 | test/routes/login.test.ts:24 | pass |
|
|
141
|
+
| REQ-002 | test/auth/token.test.ts:41 | pass |
|
|
142
|
+
|
|
143
|
+
## Discrimination Sensor
|
|
144
|
+
|
|
145
|
+
| Mutant | Expected killer | Result |
|
|
146
|
+
| --- | --- | --- |
|
|
147
|
+
| Removed expiry check | test/auth/token.test.ts:41 | killed |
|
|
148
|
+
|
|
149
|
+
## Security Review
|
|
150
|
+
[Full checklist result, or the justified lightweight path.]
|
|
151
|
+
|
|
152
|
+
## AppSec
|
|
153
|
+
- Applied: yes | skipped — [reason]
|
|
154
|
+
- Boundaries: ...
|
|
155
|
+
- Top risks: ...
|
|
156
|
+
- Result: pass | fail | escalate | skipped
|
|
157
|
+
|
|
158
|
+
## QA
|
|
159
|
+
- Applied: yes | skipped — [reason]
|
|
160
|
+
- Smoke: ...
|
|
161
|
+
- Regression focus: ...
|
|
162
|
+
- Result: pass | fail | skipped
|
|
163
|
+
|
|
164
|
+
## Gaps
|
|
165
|
+
[Ranked list, or "none".]
|
|
166
|
+
|
|
167
|
+
## Interactive UAT
|
|
168
|
+
- Applied: yes | skipped — [reason]
|
|
169
|
+
- Steps:
|
|
170
|
+
1. [action] → expect [observable]
|
|
171
|
+
2. [action] → expect [observable]
|
|
172
|
+
- Result: pass | fail | skipped
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
Lightweight security path (only when the feature has no auth, API, input, payment, or infrastructure surface):
|
|
176
|
+
|
|
177
|
+
```markdown
|
|
178
|
+
## Security Review
|
|
179
|
+
- Path: lightweight
|
|
180
|
+
- Justification: copy change in docs/README.md, no input or trust boundary
|
|
181
|
+
- Result: not applicable
|
|
182
|
+
```
|
|
183
|
+
|
|
184
|
+
An unjustified lightweight path is a gap.
|
|
185
|
+
|
|
186
|
+
## Failure Handling
|
|
187
|
+
|
|
188
|
+
**PASS requirements.** Write `Verdict: PASS` only when every coverage row passes (verifier judgment); no surviving mutant appears on a sensor/mutant line (gate blocks that); on Medium+ a sensor **outcome** is present and at least one mutant is `killed` (below Medium+ a missing outcome is a gate warning unless `--strict`); Security is pass or a justified lightweight path (gate blocks `Result: fail`; justification quality is verifier judgment); Gaps is `none` (gate blocks open Gaps bullets); AppSec/QA/Interactive UAT are pass or correctly skipped when their triggers apply (verifier judgment; not gated); and `validate_state.py` exits 0. Anything else is FAIL — do not write PASS and list gaps underneath.
|
|
189
|
+
|
|
190
|
+
- FAIL verdict → gaps become fix tasks; return to `implement.md`.
|
|
191
|
+
- The fix → re-verify loop is bounded to **3 iterations**, then escalate to the owner with the blocking gap.
|
|
192
|
+
- Every grounded failure — surviving mutant, imprecise spec, failed criterion — is recorded with `lessons.py add --source` pointing at this `validation.md`. A clean PASS records nothing. See `lessons.md`.
|
|
193
|
+
|
|
194
|
+
## Interactive UAT
|
|
195
|
+
|
|
196
|
+
Lean walkthrough for **Complex-tier user-facing** work (UI or a human-observable flow). Backend-only, infra, or library changes: skip and record one line in the report (`Interactive UAT: skipped — no user-facing surface`). Not Complex, or not user-facing: skip the same way. Prefer running UAT **after** conditional AppSec/QA steps when those ran.
|
|
197
|
+
|
|
198
|
+
**When applied**
|
|
199
|
+
|
|
200
|
+
1. After the automated coverage / sensor / security pass, write a numbered script: each step is `action → expected observable outcome`.
|
|
201
|
+
2. If the owner is in the session, run **one step at a time** and wait for a reply. Otherwise hand them the script.
|
|
202
|
+
3. Interpret replies as: **pass** (“yes” / “works” / “next”), **skip** (“can’t test” / “n/a”), or **issue** (anything else — log the words, add a Gaps bullet, open a fix task).
|
|
203
|
+
|
|
204
|
+
**FAIL rule.** Automated structural PASS with a failed walkthrough is still a FAIL for the verifier: do not leave `Verdict: PASS` while UAT Result is fail. Put the issue under Gaps and return to Execute. **Interactive UAT is verifier judgment; `validate_state.py` does not require this section and does not run the walkthrough.**
|
|
205
|
+
|
|
206
|
+
## Next
|
|
207
|
+
|
|
208
|
+
After `validate_state.py` passes → `archive.md` to fold the feature into domain truth and reset STATE.
|
|
209
|
+
|
|
210
|
+
Otherwise → `memory.md` and `git-handoff.md` — record decisions, commit `.specs/`, hand off.
|
|
@@ -0,0 +1,120 @@
|
|
|
1
|
+
# Security Review
|
|
2
|
+
|
|
3
|
+
Independent security checklist for the **Verify** phase.
|
|
4
|
+
Use alongside `agent-architecture.md` and `engineering-standards.md`.
|
|
5
|
+
|
|
6
|
+
## When to Use
|
|
7
|
+
|
|
8
|
+
- Running `/verify` on any feature touching auth, data, APIs, or infrastructure
|
|
9
|
+
- Before merging PRs that handle user input, payments, or sensitive data
|
|
10
|
+
- After dependency updates (supply chain review)
|
|
11
|
+
|
|
12
|
+
## Lightweight Path (non-security changes)
|
|
13
|
+
|
|
14
|
+
For changes with **no** auth, API, user input, payments, or infrastructure impact (e.g. copy, docs, pure UI styling):
|
|
15
|
+
|
|
16
|
+
1. Verifier still uses clean context (Author ≠ verifier)
|
|
17
|
+
2. Run discrimination sensor on affected tests
|
|
18
|
+
3. Document in `validation.md`:
|
|
19
|
+
|
|
20
|
+
```markdown
|
|
21
|
+
## Security Review — Skipped (justified)
|
|
22
|
+
- Reason: [no auth/API/input surface touched]
|
|
23
|
+
- Mutants tested: [list]
|
|
24
|
+
- Evidence: [file:line references]
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Full OWASP checklist remains mandatory for anything touching auth, data, APIs, or infra.
|
|
28
|
+
|
|
29
|
+
## Pre-Review Setup
|
|
30
|
+
|
|
31
|
+
- Verifier must have **clean context** (not the code author).
|
|
32
|
+
- Read `.specs/features/[feature]/spec.md` acceptance criteria.
|
|
33
|
+
- Identify attack surface: inputs, outputs, auth boundaries, data stores.
|
|
34
|
+
|
|
35
|
+
## OWASP-Oriented Checklist
|
|
36
|
+
|
|
37
|
+
### Injection
|
|
38
|
+
- [ ] SQL/NoSQL queries use parameterization or ORM safely
|
|
39
|
+
- [ ] Shell commands avoid unsanitized user input
|
|
40
|
+
- [ ] Template rendering escapes user content (XSS prevention)
|
|
41
|
+
|
|
42
|
+
### Broken Authentication
|
|
43
|
+
- [ ] Sessions/tokens expire appropriately
|
|
44
|
+
- [ ] Passwords hashed with modern algorithms (bcrypt, argon2)
|
|
45
|
+
- [ ] No credentials in URLs, logs, or client-side storage
|
|
46
|
+
|
|
47
|
+
### Sensitive Data Exposure
|
|
48
|
+
- [ ] Secrets in environment variables, not source code
|
|
49
|
+
- [ ] TLS for data in transit; encryption at rest where required
|
|
50
|
+
- [ ] PII minimized and masked in logs/responses
|
|
51
|
+
|
|
52
|
+
### Access Control
|
|
53
|
+
- [ ] Authorization checked on every protected endpoint
|
|
54
|
+
- [ ] IDOR prevented — users cannot access others' resources by ID manipulation
|
|
55
|
+
- [ ] Role/permission checks server-side, not client-only
|
|
56
|
+
|
|
57
|
+
### Security Misconfiguration
|
|
58
|
+
- [ ] Default credentials removed
|
|
59
|
+
- [ ] Error messages do not leak stack traces or internals in production
|
|
60
|
+
- [ ] CORS configured restrictively (not `*` with credentials)
|
|
61
|
+
|
|
62
|
+
### Vulnerable Components
|
|
63
|
+
- [ ] `npm audit` / equivalent run; critical/high addressed or documented
|
|
64
|
+
- [ ] No known-vulnerable dependencies without mitigation
|
|
65
|
+
|
|
66
|
+
### SSRF / External Requests
|
|
67
|
+
- [ ] URL fetchers validate allowed domains/schemes
|
|
68
|
+
- [ ] Internal network not reachable from user-controlled URLs
|
|
69
|
+
|
|
70
|
+
## Discrimination Sensor (Mutants)
|
|
71
|
+
|
|
72
|
+
Confirm tests catch intentional failures:
|
|
73
|
+
|
|
74
|
+
1. Remove or bypass an auth check → test must fail
|
|
75
|
+
2. Skip input validation → test must fail
|
|
76
|
+
3. Return wrong status code on error → test must fail
|
|
77
|
+
|
|
78
|
+
Inject mutants in an **isolated scratch copy** — a temp worktree or file copies. Never use `git stash` and never mutate the working tree. Discard the scratch afterwards and confirm `git status --porcelain` matches the pre-sensor baseline.
|
|
79
|
+
|
|
80
|
+
Document mutants tested in `.specs/features/[feature]/validation.md`.
|
|
81
|
+
|
|
82
|
+
## Evidence-or-Zero
|
|
83
|
+
|
|
84
|
+
Each security requirement from spec must have:
|
|
85
|
+
- Test file and line proving the control works
|
|
86
|
+
- Or explicit documented exception with owner approval
|
|
87
|
+
|
|
88
|
+
## Output
|
|
89
|
+
|
|
90
|
+
Write findings into the verifier's report at `.specs/features/[feature]/validation.md`, under the Security Review section defined in `references/validate.md`:
|
|
91
|
+
|
|
92
|
+
```markdown
|
|
93
|
+
## Security Review
|
|
94
|
+
- Reviewer: independent agent (clean context)
|
|
95
|
+
- Date: [ISO date]
|
|
96
|
+
- Mutants tested: [list]
|
|
97
|
+
- Findings: [pass/fail per checklist item]
|
|
98
|
+
- Evidence: [file:line references]
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
The completion gate reads this file:
|
|
102
|
+
|
|
103
|
+
```bash
|
|
104
|
+
python3 .specs/guardrails/scripts/validate_state.py .specs/features/[feature]
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
## Escalation
|
|
108
|
+
|
|
109
|
+
If critical vulnerability found:
|
|
110
|
+
1. Do not merge
|
|
111
|
+
2. Record the lesson with `lessons.py add --source` pointing at `validation.md`
|
|
112
|
+
3. Notify the project owner with severity and remediation steps
|
|
113
|
+
|
|
114
|
+
## Related Skills
|
|
115
|
+
|
|
116
|
+
- `agent-architecture.md` — SDD hub, execution contract, gates
|
|
117
|
+
- `references/validate.md` — verifier procedure and report schema
|
|
118
|
+
- `engineering-standards.md` — secure coding, commit format, artifact language
|
|
119
|
+
- `git-handoff.md` — git sync and session handoff for `.specs/`
|
|
120
|
+
- `task-graph-engineering.md` — verify node in diamond pattern
|
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
# Ship Ready
|
|
2
|
+
|
|
3
|
+
Lean pre-launch checklist when the owner asks to **ship, deploy, or go live**. Complements Verify — it does **not** replace `/verify` or authorize remote actions.
|
|
4
|
+
|
|
5
|
+
**Judgment only.** Not gated by `validate_state.py`. Blast radius still applies: `git push`, deploy, and destructive ops need an **explicit** go-ahead for that action.
|
|
6
|
+
|
|
7
|
+
## When to Use
|
|
8
|
+
|
|
9
|
+
Load only when the owner explicitly asks for ship / deploy / launch / go-live checklist **after** (or alongside a completed) Verify PASS for the feature in scope.
|
|
10
|
+
|
|
11
|
+
## When NOT to Use
|
|
12
|
+
|
|
13
|
+
- Normal Execute or Verify loops
|
|
14
|
+
- “Are we done coding?” — that is `/verify`, not this skill
|
|
15
|
+
- Never load with `appsec.md`, `qa-strategy.md`, or `code-simplify.md` in the same window
|
|
16
|
+
|
|
17
|
+
## Checklist
|
|
18
|
+
|
|
19
|
+
Confirm each item or record N/A with reason:
|
|
20
|
+
|
|
21
|
+
| Item | Ask |
|
|
22
|
+
| --- | --- |
|
|
23
|
+
| Verify | Feature `validation.md` is PASS; `validate_state.py` exits 0 |
|
|
24
|
+
| Tests / CI | Project suite or CI green for this change |
|
|
25
|
+
| Secrets | No secrets in the diff; env/config documented |
|
|
26
|
+
| Migrations | Backward-compatible or rollback noted |
|
|
27
|
+
| Observability | Critical path has log/metric/trace if the project expects it |
|
|
28
|
+
| Rollback | How to undo the release if it fails |
|
|
29
|
+
| Owner go-ahead | Explicit approval for push/deploy named in chat |
|
|
30
|
+
|
|
31
|
+
## Output shape
|
|
32
|
+
|
|
33
|
+
```markdown
|
|
34
|
+
## Ship ready
|
|
35
|
+
- Applied: yes | skipped — [reason]
|
|
36
|
+
- Verify: PASS | blocked — [why]
|
|
37
|
+
- CI/tests: pass | fail | n/a
|
|
38
|
+
- Secrets/migrations/obs/rollback: [ok or gap]
|
|
39
|
+
- Push/deploy: waiting for explicit go-ahead | approved — [quote]
|
|
40
|
+
- Result: ready | not ready
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
`not ready` → Gaps or blockers in `STATE.md`; do not push.
|
|
44
|
+
|
|
45
|
+
## Related
|
|
46
|
+
|
|
47
|
+
- `references/validate.md` — Verify first; do not auto-load this on Verify
|
|
48
|
+
- `git-handoff.md` — commit `.specs/`; push still needs go-ahead
|
|
49
|
+
- `agent-architecture.md` — blast radius
|
|
50
|
+
- `references/context-limits.md` — at most one conditional sister
|
|
@@ -0,0 +1,180 @@
|
|
|
1
|
+
# Task Graph Engineering
|
|
2
|
+
|
|
3
|
+
Design the **topology** of agent work — which jobs run, in what order, and what can run in parallel.
|
|
4
|
+
Sister skill to `agent-architecture.md` — SDD defines *when* (phases); this skill defines *how jobs connect* within Execute and Verify.
|
|
5
|
+
|
|
6
|
+
Adapted from task-graph orchestration patterns (MIT). Knowledge-graph content intentionally omitted — this skill covers agent orchestration only.
|
|
7
|
+
|
|
8
|
+
## When to Use
|
|
9
|
+
|
|
10
|
+
- Breaking down `/tasks` for multi-step or multi-file features
|
|
11
|
+
- Planning parallel subagent work before `/loop`
|
|
12
|
+
- Deciding whether to spawn multiple agents or keep one sequential context
|
|
13
|
+
- Structuring `/verify` with separate verifier contexts (diamond pattern)
|
|
14
|
+
- Any feature where tasks have real dependencies — or look independent but aren't
|
|
15
|
+
|
|
16
|
+
## Sister Skills (use together)
|
|
17
|
+
|
|
18
|
+
| Skill | Role |
|
|
19
|
+
| --- | --- |
|
|
20
|
+
| `agent-architecture.md` | Process — hub, contract, complexity router |
|
|
21
|
+
| `engineering-standards.md` | Quality — one writer per file, commit format |
|
|
22
|
+
| `security-review.md` | Verification — OWASP checklist for `/verify` |
|
|
23
|
+
| `git-handoff.md` | Persistence — git sync for `.specs/` |
|
|
24
|
+
| **`task-graph-engineering.md`** | **Topology — DAG of jobs, parallelism, verify separation** |
|
|
25
|
+
|
|
26
|
+
## Core Model
|
|
27
|
+
|
|
28
|
+
A **task graph** is a DAG (directed acyclic graph):
|
|
29
|
+
|
|
30
|
+
- **Nodes** = jobs (one unit of work you'd hand to a single agent)
|
|
31
|
+
- **Edges** = real dependencies (job B needs job A's *output*, not just "comes after")
|
|
32
|
+
|
|
33
|
+
Draw the graph in `.specs/features/[feature]/task-graph.md` before `/loop` when the feature has 3+ tasks or any parallel work.
|
|
34
|
+
|
|
35
|
+
```markdown
|
|
36
|
+
## Task Graph: [feature]
|
|
37
|
+
|
|
38
|
+
| Node | Depends on | Parallel group | Owner |
|
|
39
|
+
| --- | --- | --- | --- |
|
|
40
|
+
| Task 1: schema | — | A | agent-1 |
|
|
41
|
+
| Task 2: types | — | A | agent-2 |
|
|
42
|
+
| Task 3: API | 1, 2 | — | agent-1 |
|
|
43
|
+
| Verify | 3 | — | verifier (clean context) |
|
|
44
|
+
| Merge + handoff | Verify | — | merge owner |
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## The Stop Rule
|
|
48
|
+
|
|
49
|
+
Before parallelizing, ask: *where does work split into pieces that never read each other's results?*
|
|
50
|
+
|
|
51
|
+
| Work shape | Strategy |
|
|
52
|
+
| --- | --- |
|
|
53
|
+
| **Splittable** — independent files, modules, or research angles | Parallel workers → separate verify → one merge owner |
|
|
54
|
+
| **Sequential** — each step needs the full picture from the previous step | Single agent, no fan-out |
|
|
55
|
+
|
|
56
|
+
Multi-agent setups win on splittable work and **lose** on sequential work. More agents is not a strategy — the shape of the work decides.
|
|
57
|
+
|
|
58
|
+
## Fake Edges
|
|
59
|
+
|
|
60
|
+
For every "and then" in a task list, ask: does the next job actually read the previous job's output?
|
|
61
|
+
|
|
62
|
+
- **Real edge**: Task B imports or depends on artifacts from Task A → keep the dependency
|
|
63
|
+
- **Fake edge**: "Write tests and then update README" when README doesn't use test output → delete the edge; run in parallel or reorder
|
|
64
|
+
|
|
65
|
+
Most hand-built task lists contain 2–3 fake edges. Removing them unlocks safe parallelism.
|
|
66
|
+
|
|
67
|
+
## The Diamond Pattern
|
|
68
|
+
|
|
69
|
+
The shape serious systems converge to for splittable work:
|
|
70
|
+
|
|
71
|
+
| Stage | What runs | Notes |
|
|
72
|
+
| --- | --- | --- |
|
|
73
|
+
| Plan | Single planner | Produces `tasks.md` + `task-graph.md` |
|
|
74
|
+
| Workers | 1–3 parallel agents | Disjoint file ownership only |
|
|
75
|
+
| Verify | Fresh verifier context | Never the code author |
|
|
76
|
+
| Merge | One merge owner | Resolves conflicts, runs harness |
|
|
77
|
+
| Result | Commits + evidence | Handoff to `/archive` when done |
|
|
78
|
+
|
|
79
|
+
Rules:
|
|
80
|
+
|
|
81
|
+
1. **Split** only at real boundaries (see Stop Rule)
|
|
82
|
+
2. **Workers** run in parallel with disjoint file ownership (see engineering-standards)
|
|
83
|
+
3. **Verify** in a **separate context** — never the code author; ask diverse questions (correct? current? tested?)
|
|
84
|
+
4. **Merge** has one owner who resolves conflicts and runs the project harness
|
|
85
|
+
|
|
86
|
+
Maps to SDD: `/tasks` → `/loop` (workers) → `/verify` (diamond verify node) → `/handoff` (merge + persist).
|
|
87
|
+
|
|
88
|
+
## Sub-Agent Delegation
|
|
89
|
+
|
|
90
|
+
**Trigger** — Count the tasks. Roughly **8 or fewer** fits one batch: execute inline. More than that: offer sub-agents.
|
|
91
|
+
|
|
92
|
+
**Offer-then-confirm** — Never auto-spawn. Present the proposed split and wait for the owner to accept.
|
|
93
|
+
|
|
94
|
+
**Batching**
|
|
95
|
+
|
|
96
|
+
- A **batch** is the execution unit: one or more consecutive whole phases packed to about **7 tasks**.
|
|
97
|
+
- Phases stay the semantic unit — never split a phase across workers.
|
|
98
|
+
- Batches run **sequentially**; a batch starts only after the previous one reports every task complete.
|
|
99
|
+
- Each worker implements → gates → commits each of its tasks in order, then reports a compact summary: tasks done, commit hashes, test counts, deviations.
|
|
100
|
+
- Workers never spawn further sub-agents.
|
|
101
|
+
|
|
102
|
+
Full operational contract (worker payload, compact summary template, failure table, merge owner steps): `references/sub-agents.md`.
|
|
103
|
+
|
|
104
|
+
**Verifier** — After the final task, dispatch a fresh verifier regardless of batch count. It is the closing step of Execute, never prompted. See `references/validate.md`.
|
|
105
|
+
|
|
106
|
+
**Model tier per role** — When Spec Guardrails allows choosing a model per sub-agent:
|
|
107
|
+
|
|
108
|
+
| Role | Tier |
|
|
109
|
+
| --- | --- |
|
|
110
|
+
| Mechanical batch (config, wiring, CRUD) | Fast / cost-effective |
|
|
111
|
+
| Core-domain or ambiguous batch | High reasoning |
|
|
112
|
+
| Design phase | High reasoning |
|
|
113
|
+
| Verifier | Mid-to-high — adversarial reasoning, mutant design |
|
|
114
|
+
|
|
115
|
+
If Spec Guardrails cannot set a per-agent model, ignore this and invest more care on the heavy steps.
|
|
116
|
+
|
|
117
|
+
## Human Gate
|
|
118
|
+
|
|
119
|
+
Route **irreversible** actions through explicit human approval:
|
|
120
|
+
|
|
121
|
+
- Deploy, publish, push to shared branch
|
|
122
|
+
- Delete data, refund, send external communications
|
|
123
|
+
- Merge to main/production
|
|
124
|
+
|
|
125
|
+
**Placement rule**: gate where a mistake is expensive to undo — not on every step.
|
|
126
|
+
|
|
127
|
+
| Action | Gate? |
|
|
128
|
+
| --- | --- |
|
|
129
|
+
| Local commit | No (handoff handles this) |
|
|
130
|
+
| `git push` | Yes — human or explicit instruction |
|
|
131
|
+
| Deploy / release | Yes — always |
|
|
132
|
+
| Continue after 3 failed correction loops | Yes — escalate to human |
|
|
133
|
+
|
|
134
|
+
## Guardrails
|
|
135
|
+
|
|
136
|
+
1. **Max correction loops** — 3 per task, then escalate (see `agent-architecture.md`)
|
|
137
|
+
2. **One writer per file** — no two jobs mutate the same file in the same round
|
|
138
|
+
3. **Plan in writing** — routing lives in `task-graph.md` and `tasks.md`; agents fill jobs, not the plan
|
|
139
|
+
4. **Cap subagents** — hard limit on parallel agents (default: 3 workers + 1 verifier)
|
|
140
|
+
5. **Evidence over self-report** — judge on harness output (tests ran, linter passed), not agent claims
|
|
141
|
+
|
|
142
|
+
## Integration with SDD Phases
|
|
143
|
+
|
|
144
|
+
| SDD Phase | Task graph action |
|
|
145
|
+
| --- | --- |
|
|
146
|
+
| `/tasks` | Identify fake edges; mark parallel groups; link to REQ IDs |
|
|
147
|
+
| `/task-graph` | Draw or revise the DAG in `task-graph.md` |
|
|
148
|
+
| `/loop` | Execute graph — `loop-plan` each round; parallel sub-agents when files are disjoint |
|
|
149
|
+
| `/verify` | Diamond verify node — clean context, diverse checks |
|
|
150
|
+
| `/handoff` | Commit `task-graph.md` with other `.specs/` artifacts |
|
|
151
|
+
|
|
152
|
+
## When NOT to Use
|
|
153
|
+
|
|
154
|
+
Skip the task graph for:
|
|
155
|
+
|
|
156
|
+
- Single-file bug fixes
|
|
157
|
+
- Copy or config-only changes
|
|
158
|
+
- Trivial changes covered by the complexity router in `agent-architecture.md`
|
|
159
|
+
|
|
160
|
+
## Available Commands
|
|
161
|
+
|
|
162
|
+
| Command | Action |
|
|
163
|
+
| --- | --- |
|
|
164
|
+
| `/task-graph` | Draw or revise DAG in `.specs/features/[feature]/task-graph.md` |
|
|
165
|
+
| `/tasks` | Atomic breakdown — apply fake-edge and stop-rule checks |
|
|
166
|
+
| `/loop` | Execute graph respecting parallelism and one-writer-per-file |
|
|
167
|
+
| `/verify` | Independent verification (diamond verify node) |
|
|
168
|
+
|
|
169
|
+
## Related Skills
|
|
170
|
+
|
|
171
|
+
- `agent-architecture.md` — SDD hub, contract, complexity router
|
|
172
|
+
- `references/tasks.md` — task schema and the `validate_tasks.py` gate
|
|
173
|
+
- `references/validate.md` — verifier procedure for the diamond verify node
|
|
174
|
+
- `engineering-standards.md` — one writer per file, git hygiene
|
|
175
|
+
- `security-review.md` — security checklist for verify node
|
|
176
|
+
- `git-handoff.md` — commit `task-graph.md` at phase boundaries
|
|
177
|
+
|
|
178
|
+
## Credits
|
|
179
|
+
|
|
180
|
+
Task-graph patterns adapted under MIT license from [graph-engineering](https://github.com/codejunkie99/graph-engineering) (task-graph half only).
|
|
@@ -0,0 +1,61 @@
|
|
|
1
|
+
# Getting started
|
|
2
|
+
|
|
3
|
+
You installed the **Spec Guardrails**. You do **not** need to memorize CLI commands.
|
|
4
|
+
|
|
5
|
+
## What to do now
|
|
6
|
+
|
|
7
|
+
1. Open **Cursor** or **Claude Code** in this project.
|
|
8
|
+
2. Start with **Specify** (an **agent command** — chat, not terminal):
|
|
9
|
+
|
|
10
|
+
```
|
|
11
|
+
/specify
|
|
12
|
+
|
|
13
|
+
Add CSV export to the reports page. Users pick a date range.
|
|
14
|
+
Out of scope: PDF export.
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
3. Review `.specs/features/…/spec.md` and **approve** before implementation.
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## Agent commands (chat — not terminal)
|
|
22
|
+
|
|
23
|
+
Type these in **Cursor or Claude Code**. They load phase procedures from `.cursor/skills/references/`. The agent runs gates for you.
|
|
24
|
+
|
|
25
|
+
| Command | When |
|
|
26
|
+
| --- | --- |
|
|
27
|
+
| `/specify` | **Start here** — written requirements |
|
|
28
|
+
| `/tasks` | Break into jobs after spec approval |
|
|
29
|
+
| `/loop` | Implement — agent runs `loop-plan` each wave |
|
|
30
|
+
| `/verify` | Fresh-context proof after last task |
|
|
31
|
+
| `/quick` | Tiny fix only (≤3 files) |
|
|
32
|
+
|
|
33
|
+
**Typical order:** `/specify` → `/tasks` → `/analyze` → `/loop` → `/verify` → `/archive`
|
|
34
|
+
|
|
35
|
+
**Full reference** (every command, examples, CLI): [Agent commands](https://github.com/luizssantiago92/spec-guardrails/blob/main/docs/guide/agent-commands.md)
|
|
36
|
+
|
|
37
|
+
---
|
|
38
|
+
|
|
39
|
+
## CLI you might run yourself (terminal)
|
|
40
|
+
|
|
41
|
+
| Command | When |
|
|
42
|
+
| --- | --- |
|
|
43
|
+
| `install` | First time or upgrade |
|
|
44
|
+
| `project-init` | Brownfield repo (optional) |
|
|
45
|
+
| `doctor` | Install looks broken |
|
|
46
|
+
| `validate-spec` / `validate-state` | Double-check gates manually |
|
|
47
|
+
|
|
48
|
+
Everything else (`loop-plan`, `validate-tasks`, `check-commit`, …) is normally run **by the agent**.
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## Where things live
|
|
53
|
+
|
|
54
|
+
| Path | Purpose |
|
|
55
|
+
| --- | --- |
|
|
56
|
+
| `.cursor/skills/agent-architecture.md` | Hub — phase map |
|
|
57
|
+
| `.specs/STATE.md` | Where you left off |
|
|
58
|
+
| `.specs/features/` | One folder per feature |
|
|
59
|
+
| `.specs/guardrails/scripts/` | Automatic gates |
|
|
60
|
+
|
|
61
|
+
More guides: [docs/guide/Home.md](https://github.com/luizssantiago92/spec-guardrails/blob/main/docs/guide/Home.md)
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Project guardrails config (optional). Copy to .specs/config.yaml and edit.
|
|
2
|
+
# Or run: npx @luizsantiago/spec-guardrails init-config --preset node-ts
|
|
3
|
+
# Injected into phase procedures via phase-context.
|
|
4
|
+
|
|
5
|
+
schema: spec-driven
|
|
6
|
+
|
|
7
|
+
# Extend a built-in preset (default | node-ts | python):
|
|
8
|
+
# extends: node-ts
|
|
9
|
+
|
|
10
|
+
context: |
|
|
11
|
+
Tech stack: (fill in)
|
|
12
|
+
Test command: npm test
|
|
13
|
+
Branch prefix: feat
|
|
14
|
+
|
|
15
|
+
rules:
|
|
16
|
+
specify:
|
|
17
|
+
- Prefer EARS acceptance criteria
|
|
18
|
+
- Mark unknowns with [NEEDS CLARIFICATION: question]
|
|
19
|
+
tasks:
|
|
20
|
+
- Every REQ must appear in the Test Coverage Matrix
|
|
21
|
+
verify:
|
|
22
|
+
- Evidence must cite test file:line paths
|
|
23
|
+
|
|
24
|
+
# Project-specific overrides (appended on top of preset + rules above):
|
|
25
|
+
# overrides:
|
|
26
|
+
# rules:
|
|
27
|
+
# verify:
|
|
28
|
+
# - Also run playwright e2e before PASS
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
# Default preset — general-purpose projects
|
|
2
|
+
schema: spec-driven
|
|
3
|
+
|
|
4
|
+
context: |
|
|
5
|
+
Tech stack: (fill in)
|
|
6
|
+
Test command: npm test
|
|
7
|
+
Branch prefix: feat
|
|
8
|
+
|
|
9
|
+
rules:
|
|
10
|
+
specify:
|
|
11
|
+
- Prefer EARS acceptance criteria
|
|
12
|
+
- Mark unknowns with [NEEDS CLARIFICATION: question]
|
|
13
|
+
tasks:
|
|
14
|
+
- Every REQ must appear in the Test Coverage Matrix
|
|
15
|
+
verify:
|
|
16
|
+
- Evidence must cite test file:line paths
|