@luizsantiago/spec-guardrails 3.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (63) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +206 -0
  3. package/index.js +335 -0
  4. package/lib/archive.js +208 -0
  5. package/lib/assets.js +145 -0
  6. package/lib/brownfield.js +446 -0
  7. package/lib/config.js +293 -0
  8. package/lib/constants.js +262 -0
  9. package/lib/cursorrules.js +92 -0
  10. package/lib/delta-merge.js +248 -0
  11. package/lib/doctor.js +343 -0
  12. package/lib/download.js +133 -0
  13. package/lib/feature.js +272 -0
  14. package/lib/fs-utils.js +114 -0
  15. package/lib/gates.js +138 -0
  16. package/lib/install.js +140 -0
  17. package/lib/memory.js +34 -0
  18. package/lib/next-steps.js +50 -0
  19. package/lib/presets.js +176 -0
  20. package/lib/project-rules.js +210 -0
  21. package/lib/specs-utils.js +117 -0
  22. package/lib/token-cost.js +124 -0
  23. package/package.json +46 -0
  24. package/rules/engineering-baseline.mdc +56 -0
  25. package/scripts/_common.py +356 -0
  26. package/scripts/analyze_artifacts.py +187 -0
  27. package/scripts/check_commit.py +140 -0
  28. package/scripts/lessons.py +447 -0
  29. package/scripts/loop_plan.py +217 -0
  30. package/scripts/validate_spec.py +345 -0
  31. package/scripts/validate_state.py +385 -0
  32. package/scripts/validate_tasks.py +379 -0
  33. package/skills/agent-architecture.md +221 -0
  34. package/skills/appsec.md +83 -0
  35. package/skills/code-simplify.md +49 -0
  36. package/skills/engineering-standards.md +98 -0
  37. package/skills/git-handoff.md +213 -0
  38. package/skills/qa-strategy.md +83 -0
  39. package/skills/references/analyze.md +56 -0
  40. package/skills/references/archive.md +60 -0
  41. package/skills/references/constitution.md +66 -0
  42. package/skills/references/context-limits.md +73 -0
  43. package/skills/references/converge.md +47 -0
  44. package/skills/references/design.md +88 -0
  45. package/skills/references/discuss.md +68 -0
  46. package/skills/references/explore.md +61 -0
  47. package/skills/references/implement.md +175 -0
  48. package/skills/references/lessons.md +71 -0
  49. package/skills/references/memory.md +98 -0
  50. package/skills/references/project-init.md +62 -0
  51. package/skills/references/quick-mode.md +84 -0
  52. package/skills/references/specify.md +144 -0
  53. package/skills/references/sub-agents.md +117 -0
  54. package/skills/references/tasks.md +178 -0
  55. package/skills/references/validate.md +210 -0
  56. package/skills/security-review.md +120 -0
  57. package/skills/ship-ready.md +50 -0
  58. package/skills/task-graph-engineering.md +180 -0
  59. package/templates/GETTING_STARTED.md +61 -0
  60. package/templates/config.yaml.example +28 -0
  61. package/templates/presets/default.yaml +16 -0
  62. package/templates/presets/node-ts.yaml +22 -0
  63. package/templates/presets/python.yaml +22 -0
@@ -0,0 +1,210 @@
1
+ # Validate
2
+
3
+ Independent verification of the delivered feature. Always required, never prompted.
4
+
5
+ ## When to Use
6
+
7
+ - After the last task of a feature is committed
8
+ - After any fix round that follows a FAIL verdict
9
+
10
+ ## Who Runs It
11
+
12
+ A **fresh verifier context that never wrote the code**. Author ≠ verifier is non-negotiable. The verifier re-derives coverage from the spec instead of inheriting the author's mental model.
13
+
14
+ When sub-agents are available, dispatch the verifier as a separate agent (see `task-graph-engineering.md`). Without sub-agents, start a clean context and run this file as a fresh-eyes pass.
15
+
16
+ ## Inputs
17
+
18
+ - `spec.md` acceptance criteria, `context.md` when it exists
19
+ - The diff range for the feature
20
+ - `security-review.md` for the security checklist
21
+ - `appsec.md` only when Complex or an attack-surface trigger fires — then **drop** it before QA
22
+ - `qa-strategy.md` only when Complex, multi-step UI, or an explicit regression ask — never together with `appsec.md`
23
+ - Do **not** load `ship-ready.md` or `code-simplify.md` during Verify (ship is owner-triggered after PASS; simplify is Execute-side)
24
+ - `context-limits.md` — load the spec, the diff, and the tests the spec names; do not load the author's chat; at most one conditional sister at a time
25
+
26
+ ## Output
27
+
28
+ `.specs/features/[feature]/validation.md`
29
+
30
+ ## Procedure
31
+
32
+ ### 1. Spec-anchored outcome check
33
+
34
+ For each acceptance criterion, confirm that a test asserts the **spec-defined outcome** — not merely that the code runs. Flag criteria where the test asserts an implementation detail, and flag spec text too imprecise to test.
35
+
36
+ ### 2. Discrimination sensor
37
+
38
+ Confirm the tests can actually fail:
39
+
40
+ 1. Inject behavior-level faults one at a time in an **isolated scratch copy** — a temp worktree or file copies. Never use `git stash` and never mutate the working tree.
41
+ 2. Confirm the relevant test fails for each mutant.
42
+ 3. Discard the scratch and verify the real tree is unchanged (`git status --porcelain` matches the pre-sensor baseline).
43
+ 4. Any mutant that survives becomes a fix task — the tests do not discriminate.
44
+
45
+ ### 3. Security review
46
+
47
+ Run the checklist in `security-review.md`. Features with no auth, API, input, payment, or infrastructure surface may take the documented lightweight path — with the justification written into the report.
48
+
49
+ ### 4. Evidence-or-zero
50
+
51
+ A requirement is satisfied only with a `file:line` reference to an assertive test that passes. No reference means not done, regardless of how the code looks.
52
+
53
+ ### 5. Conditional AppSec (optional)
54
+
55
+ If the hub AppSec trigger fires, load **only** `appsec.md`, write `## AppSec` (or `skipped — reason`), then **drop** that skill from context. Do not load `qa-strategy.md` yet. If the trigger does not fire, record a one-line skip. Judgment only — not gated.
56
+
57
+ ### 6. Conditional QA (optional)
58
+
59
+ If the hub QA trigger fires, load **only** `qa-strategy.md` (after AppSec is done or skipped), write `## QA`, then continue. Judgment only — not gated.
60
+
61
+ ### 7. Verdict
62
+
63
+ Write `validation.md`, then run the completion gate.
64
+
65
+ ## Mutant catalog
66
+
67
+ Pick mutants from the kind of code that landed, not from a generic list. One killed mutant per risky behavior is the minimum; surviving mutants are fix tasks.
68
+
69
+ | Code kind | Mutant | Expected killer |
70
+ | --- | --- | --- |
71
+ | Auth / session | Skip the expiry check; accept an empty token; invert role comparison | Test that names the denied case |
72
+ | HTTP handler | Return 200 on the error path; swap 401/403; drop the error code body | Status + code assertion |
73
+ | Validation | Remove an upper/lower bound; accept the empty string; skip a required field | Boundary test |
74
+ | Persistence | Skip the unique constraint; write without a transaction; ignore not-found | Duplicate / rollback / 404 test |
75
+ | Payments / money | Off-by-one on minor units; skip idempotency key; apply a refund twice | Amount and replay tests |
76
+ | Concurrency | Drop a lock or compare-and-swap; process the same event twice | Idempotency or conflict test |
77
+ | State machine | Allow a backward transition; skip the terminal-state guard | Illegal-transition test |
78
+
79
+ A mutant that the compiler or typechecker rejects before a test runs does not count. Change behavior, not syntax.
80
+
81
+ **Weak assertion tells.** Treat the test as weak when it asserts any of: `toBeDefined()`, a 2xx status with no body, "the function was called", or a snapshot of an entire module. The spec names an outcome; the test must name it too (status + error code, exact field, exact transition).
82
+
83
+ Do not inject mutants that the spec does not constrain. A surviving mutant of unspecified behavior is a spec gap, not a test gap — send it back to Specify.
84
+
85
+ ## Gap catalog
86
+
87
+ Every FAIL names a gap type so the author knows which loop to re-enter.
88
+
89
+ | Gap | Meaning | Return to |
90
+ | --- | --- | --- |
91
+ | Missing evidence | No `file:line` for a criterion | Execute — add the assertion |
92
+ | Weak assertion | Test checks that code ran, not the spec outcome | Execute — rewrite the test |
93
+ | Surviving mutant | Discriminating test is absent | Execute — add the killer test, then the production fix if needed |
94
+ | Imprecise criterion | Spec cannot be tested as written | Specify — rewrite with `SHALL`/`MUST` and a trigger |
95
+ | Spec deviation | Implementation does something the spec forbids, or omits something it requires | Specify + Execute |
96
+ | Security finding | Checklist item failed | Execute — fix, then re-verify |
97
+ | Open task | `tasks.md` still has `- [ ]` | Execute — finish or drop the task with owner approval |
98
+ | Sensor skipped | No mutant outcome in the report | Verify — run the sensor. On **Medium+** (`design.md` with content, 4+ tasks, or 2+ phases) do not pass — the completion gate blocks. Below Medium+ the gate warns (`--strict` promotes); still run the sensor when risk warrants it |
99
+
100
+ Rank gaps: security and spec deviation first, then surviving mutants, then missing evidence, then imprecise criteria. The author fixes in that order.
101
+
102
+ ## Evidence format
103
+
104
+ The completion gate searches for `file:line` (for example `test/routes/login.test.ts:24`). A URL, a CI job name, or "covered by the suite" is not evidence.
105
+
106
+ - Cite the assertive test, not the production file
107
+ - One evidence line per criterion; reuse a test only when it truly asserts both outcomes
108
+ - After a fix round, cite the new line numbers — stale citations fail the reader even if they pass the regex
109
+
110
+ ## Gate
111
+
112
+ ```bash
113
+ python3 .specs/guardrails/scripts/validate_state.py .specs/features/[feature]
114
+ python3 .specs/guardrails/scripts/validate_state.py [feature]
115
+ python3 .specs/guardrails/scripts/validate_state.py # single-feature projects
116
+ ```
117
+
118
+ Checks that the report exists, the verdict is exactly PASS in the **preamble** (before the first `##` section) or under a dedicated `## Verdict` / `## Result` / `## Status` heading, every spec requirement ID shares a line with test `file:line` evidence, and no task remains open. A `- Verdict: PASS` buried under Discrimination Sensor or Coverage does not count. Preamble and `## Verdict` must not disagree. Evidence inside fenced samples or HTML comments does not count. `PASS` with any surviving mutant on a sensor/mutant line fails. `PASS` with open `Gaps` bullets or Security Review `Result: fail` fails. On Medium+ features (`design.md` with content, 4+ tasks, or 2+ phases) a discrimination-sensor **outcome** is **blocking** — the section heading alone is not enough — and a Medium+ `PASS` requires at least one `killed` mutant in the sensor focus (`injected` alone, or `killed` only under Gaps, is not enough). Below Medium+ a missing outcome is a warning (`--strict` still promotes warnings). Non-zero exit means the feature is not done.
119
+
120
+ The gate cannot judge whether a cited test actually asserts the criterion. That judgment is the verifier's; a green gate with a weak assertion is still a FAIL in the report.
121
+
122
+ **Gate-enforced vs verifier judgment.** `validate_state.py` enforces form: verdict scope, test-path evidence, REQ↔evidence lines, sensor outcomes on Medium+, open Gaps, Security `Result: fail`, open tasks. The following stay **verifier judgment** (not structural gates): whether each coverage row's test truly asserts the outcome, whether a lightweight Security path is justified, Interactive UAT / walkthrough success, and optional `## AppSec` / `## QA` sections. See the [gate stability contract](https://github.com/luizssantiago92/spec-guardrails/blob/main/prd/gate-stability.md).
123
+
124
+ **Red flags before PASS (judgment):** evidence only in fences/comments; open Gaps with PASS; soft `PASS WITH GAPS`; verdict buried under Sensor/Coverage; config path as evidence; surviving mutant or Medium+ without a kill — these already fail the gate or the reader.
125
+
126
+ ## Template
127
+
128
+ ```markdown
129
+ # Validation: [Feature]
130
+
131
+ - Verifier: independent agent (clean context)
132
+ - Date: [ISO date]
133
+ - Diff range: [base..head]
134
+ - Verdict: PASS
135
+
136
+ ## Coverage
137
+
138
+ | Requirement | Test evidence | Result |
139
+ | --- | --- | --- |
140
+ | REQ-001 | test/routes/login.test.ts:24 | pass |
141
+ | REQ-002 | test/auth/token.test.ts:41 | pass |
142
+
143
+ ## Discrimination Sensor
144
+
145
+ | Mutant | Expected killer | Result |
146
+ | --- | --- | --- |
147
+ | Removed expiry check | test/auth/token.test.ts:41 | killed |
148
+
149
+ ## Security Review
150
+ [Full checklist result, or the justified lightweight path.]
151
+
152
+ ## AppSec
153
+ - Applied: yes | skipped — [reason]
154
+ - Boundaries: ...
155
+ - Top risks: ...
156
+ - Result: pass | fail | escalate | skipped
157
+
158
+ ## QA
159
+ - Applied: yes | skipped — [reason]
160
+ - Smoke: ...
161
+ - Regression focus: ...
162
+ - Result: pass | fail | skipped
163
+
164
+ ## Gaps
165
+ [Ranked list, or "none".]
166
+
167
+ ## Interactive UAT
168
+ - Applied: yes | skipped — [reason]
169
+ - Steps:
170
+ 1. [action] → expect [observable]
171
+ 2. [action] → expect [observable]
172
+ - Result: pass | fail | skipped
173
+ ```
174
+
175
+ Lightweight security path (only when the feature has no auth, API, input, payment, or infrastructure surface):
176
+
177
+ ```markdown
178
+ ## Security Review
179
+ - Path: lightweight
180
+ - Justification: copy change in docs/README.md, no input or trust boundary
181
+ - Result: not applicable
182
+ ```
183
+
184
+ An unjustified lightweight path is a gap.
185
+
186
+ ## Failure Handling
187
+
188
+ **PASS requirements.** Write `Verdict: PASS` only when every coverage row passes (verifier judgment); no surviving mutant appears on a sensor/mutant line (gate blocks that); on Medium+ a sensor **outcome** is present and at least one mutant is `killed` (below Medium+ a missing outcome is a gate warning unless `--strict`); Security is pass or a justified lightweight path (gate blocks `Result: fail`; justification quality is verifier judgment); Gaps is `none` (gate blocks open Gaps bullets); AppSec/QA/Interactive UAT are pass or correctly skipped when their triggers apply (verifier judgment; not gated); and `validate_state.py` exits 0. Anything else is FAIL — do not write PASS and list gaps underneath.
189
+
190
+ - FAIL verdict → gaps become fix tasks; return to `implement.md`.
191
+ - The fix → re-verify loop is bounded to **3 iterations**, then escalate to the owner with the blocking gap.
192
+ - Every grounded failure — surviving mutant, imprecise spec, failed criterion — is recorded with `lessons.py add --source` pointing at this `validation.md`. A clean PASS records nothing. See `lessons.md`.
193
+
194
+ ## Interactive UAT
195
+
196
+ Lean walkthrough for **Complex-tier user-facing** work (UI or a human-observable flow). Backend-only, infra, or library changes: skip and record one line in the report (`Interactive UAT: skipped — no user-facing surface`). Not Complex, or not user-facing: skip the same way. Prefer running UAT **after** conditional AppSec/QA steps when those ran.
197
+
198
+ **When applied**
199
+
200
+ 1. After the automated coverage / sensor / security pass, write a numbered script: each step is `action → expected observable outcome`.
201
+ 2. If the owner is in the session, run **one step at a time** and wait for a reply. Otherwise hand them the script.
202
+ 3. Interpret replies as: **pass** (“yes” / “works” / “next”), **skip** (“can’t test” / “n/a”), or **issue** (anything else — log the words, add a Gaps bullet, open a fix task).
203
+
204
+ **FAIL rule.** Automated structural PASS with a failed walkthrough is still a FAIL for the verifier: do not leave `Verdict: PASS` while UAT Result is fail. Put the issue under Gaps and return to Execute. **Interactive UAT is verifier judgment; `validate_state.py` does not require this section and does not run the walkthrough.**
205
+
206
+ ## Next
207
+
208
+ After `validate_state.py` passes → `archive.md` to fold the feature into domain truth and reset STATE.
209
+
210
+ Otherwise → `memory.md` and `git-handoff.md` — record decisions, commit `.specs/`, hand off.
@@ -0,0 +1,120 @@
1
+ # Security Review
2
+
3
+ Independent security checklist for the **Verify** phase.
4
+ Use alongside `agent-architecture.md` and `engineering-standards.md`.
5
+
6
+ ## When to Use
7
+
8
+ - Running `/verify` on any feature touching auth, data, APIs, or infrastructure
9
+ - Before merging PRs that handle user input, payments, or sensitive data
10
+ - After dependency updates (supply chain review)
11
+
12
+ ## Lightweight Path (non-security changes)
13
+
14
+ For changes with **no** auth, API, user input, payments, or infrastructure impact (e.g. copy, docs, pure UI styling):
15
+
16
+ 1. Verifier still uses clean context (Author ≠ verifier)
17
+ 2. Run discrimination sensor on affected tests
18
+ 3. Document in `validation.md`:
19
+
20
+ ```markdown
21
+ ## Security Review — Skipped (justified)
22
+ - Reason: [no auth/API/input surface touched]
23
+ - Mutants tested: [list]
24
+ - Evidence: [file:line references]
25
+ ```
26
+
27
+ Full OWASP checklist remains mandatory for anything touching auth, data, APIs, or infra.
28
+
29
+ ## Pre-Review Setup
30
+
31
+ - Verifier must have **clean context** (not the code author).
32
+ - Read `.specs/features/[feature]/spec.md` acceptance criteria.
33
+ - Identify attack surface: inputs, outputs, auth boundaries, data stores.
34
+
35
+ ## OWASP-Oriented Checklist
36
+
37
+ ### Injection
38
+ - [ ] SQL/NoSQL queries use parameterization or ORM safely
39
+ - [ ] Shell commands avoid unsanitized user input
40
+ - [ ] Template rendering escapes user content (XSS prevention)
41
+
42
+ ### Broken Authentication
43
+ - [ ] Sessions/tokens expire appropriately
44
+ - [ ] Passwords hashed with modern algorithms (bcrypt, argon2)
45
+ - [ ] No credentials in URLs, logs, or client-side storage
46
+
47
+ ### Sensitive Data Exposure
48
+ - [ ] Secrets in environment variables, not source code
49
+ - [ ] TLS for data in transit; encryption at rest where required
50
+ - [ ] PII minimized and masked in logs/responses
51
+
52
+ ### Access Control
53
+ - [ ] Authorization checked on every protected endpoint
54
+ - [ ] IDOR prevented — users cannot access others' resources by ID manipulation
55
+ - [ ] Role/permission checks server-side, not client-only
56
+
57
+ ### Security Misconfiguration
58
+ - [ ] Default credentials removed
59
+ - [ ] Error messages do not leak stack traces or internals in production
60
+ - [ ] CORS configured restrictively (not `*` with credentials)
61
+
62
+ ### Vulnerable Components
63
+ - [ ] `npm audit` / equivalent run; critical/high addressed or documented
64
+ - [ ] No known-vulnerable dependencies without mitigation
65
+
66
+ ### SSRF / External Requests
67
+ - [ ] URL fetchers validate allowed domains/schemes
68
+ - [ ] Internal network not reachable from user-controlled URLs
69
+
70
+ ## Discrimination Sensor (Mutants)
71
+
72
+ Confirm tests catch intentional failures:
73
+
74
+ 1. Remove or bypass an auth check → test must fail
75
+ 2. Skip input validation → test must fail
76
+ 3. Return wrong status code on error → test must fail
77
+
78
+ Inject mutants in an **isolated scratch copy** — a temp worktree or file copies. Never use `git stash` and never mutate the working tree. Discard the scratch afterwards and confirm `git status --porcelain` matches the pre-sensor baseline.
79
+
80
+ Document mutants tested in `.specs/features/[feature]/validation.md`.
81
+
82
+ ## Evidence-or-Zero
83
+
84
+ Each security requirement from spec must have:
85
+ - Test file and line proving the control works
86
+ - Or explicit documented exception with owner approval
87
+
88
+ ## Output
89
+
90
+ Write findings into the verifier's report at `.specs/features/[feature]/validation.md`, under the Security Review section defined in `references/validate.md`:
91
+
92
+ ```markdown
93
+ ## Security Review
94
+ - Reviewer: independent agent (clean context)
95
+ - Date: [ISO date]
96
+ - Mutants tested: [list]
97
+ - Findings: [pass/fail per checklist item]
98
+ - Evidence: [file:line references]
99
+ ```
100
+
101
+ The completion gate reads this file:
102
+
103
+ ```bash
104
+ python3 .specs/guardrails/scripts/validate_state.py .specs/features/[feature]
105
+ ```
106
+
107
+ ## Escalation
108
+
109
+ If critical vulnerability found:
110
+ 1. Do not merge
111
+ 2. Record the lesson with `lessons.py add --source` pointing at `validation.md`
112
+ 3. Notify the project owner with severity and remediation steps
113
+
114
+ ## Related Skills
115
+
116
+ - `agent-architecture.md` — SDD hub, execution contract, gates
117
+ - `references/validate.md` — verifier procedure and report schema
118
+ - `engineering-standards.md` — secure coding, commit format, artifact language
119
+ - `git-handoff.md` — git sync and session handoff for `.specs/`
120
+ - `task-graph-engineering.md` — verify node in diamond pattern
@@ -0,0 +1,50 @@
1
+ # Ship Ready
2
+
3
+ Lean pre-launch checklist when the owner asks to **ship, deploy, or go live**. Complements Verify — it does **not** replace `/verify` or authorize remote actions.
4
+
5
+ **Judgment only.** Not gated by `validate_state.py`. Blast radius still applies: `git push`, deploy, and destructive ops need an **explicit** go-ahead for that action.
6
+
7
+ ## When to Use
8
+
9
+ Load only when the owner explicitly asks for ship / deploy / launch / go-live checklist **after** (or alongside a completed) Verify PASS for the feature in scope.
10
+
11
+ ## When NOT to Use
12
+
13
+ - Normal Execute or Verify loops
14
+ - “Are we done coding?” — that is `/verify`, not this skill
15
+ - Never load with `appsec.md`, `qa-strategy.md`, or `code-simplify.md` in the same window
16
+
17
+ ## Checklist
18
+
19
+ Confirm each item or record N/A with reason:
20
+
21
+ | Item | Ask |
22
+ | --- | --- |
23
+ | Verify | Feature `validation.md` is PASS; `validate_state.py` exits 0 |
24
+ | Tests / CI | Project suite or CI green for this change |
25
+ | Secrets | No secrets in the diff; env/config documented |
26
+ | Migrations | Backward-compatible or rollback noted |
27
+ | Observability | Critical path has log/metric/trace if the project expects it |
28
+ | Rollback | How to undo the release if it fails |
29
+ | Owner go-ahead | Explicit approval for push/deploy named in chat |
30
+
31
+ ## Output shape
32
+
33
+ ```markdown
34
+ ## Ship ready
35
+ - Applied: yes | skipped — [reason]
36
+ - Verify: PASS | blocked — [why]
37
+ - CI/tests: pass | fail | n/a
38
+ - Secrets/migrations/obs/rollback: [ok or gap]
39
+ - Push/deploy: waiting for explicit go-ahead | approved — [quote]
40
+ - Result: ready | not ready
41
+ ```
42
+
43
+ `not ready` → Gaps or blockers in `STATE.md`; do not push.
44
+
45
+ ## Related
46
+
47
+ - `references/validate.md` — Verify first; do not auto-load this on Verify
48
+ - `git-handoff.md` — commit `.specs/`; push still needs go-ahead
49
+ - `agent-architecture.md` — blast radius
50
+ - `references/context-limits.md` — at most one conditional sister
@@ -0,0 +1,180 @@
1
+ # Task Graph Engineering
2
+
3
+ Design the **topology** of agent work — which jobs run, in what order, and what can run in parallel.
4
+ Sister skill to `agent-architecture.md` — SDD defines *when* (phases); this skill defines *how jobs connect* within Execute and Verify.
5
+
6
+ Adapted from task-graph orchestration patterns (MIT). Knowledge-graph content intentionally omitted — this skill covers agent orchestration only.
7
+
8
+ ## When to Use
9
+
10
+ - Breaking down `/tasks` for multi-step or multi-file features
11
+ - Planning parallel subagent work before `/loop`
12
+ - Deciding whether to spawn multiple agents or keep one sequential context
13
+ - Structuring `/verify` with separate verifier contexts (diamond pattern)
14
+ - Any feature where tasks have real dependencies — or look independent but aren't
15
+
16
+ ## Sister Skills (use together)
17
+
18
+ | Skill | Role |
19
+ | --- | --- |
20
+ | `agent-architecture.md` | Process — hub, contract, complexity router |
21
+ | `engineering-standards.md` | Quality — one writer per file, commit format |
22
+ | `security-review.md` | Verification — OWASP checklist for `/verify` |
23
+ | `git-handoff.md` | Persistence — git sync for `.specs/` |
24
+ | **`task-graph-engineering.md`** | **Topology — DAG of jobs, parallelism, verify separation** |
25
+
26
+ ## Core Model
27
+
28
+ A **task graph** is a DAG (directed acyclic graph):
29
+
30
+ - **Nodes** = jobs (one unit of work you'd hand to a single agent)
31
+ - **Edges** = real dependencies (job B needs job A's *output*, not just "comes after")
32
+
33
+ Draw the graph in `.specs/features/[feature]/task-graph.md` before `/loop` when the feature has 3+ tasks or any parallel work.
34
+
35
+ ```markdown
36
+ ## Task Graph: [feature]
37
+
38
+ | Node | Depends on | Parallel group | Owner |
39
+ | --- | --- | --- | --- |
40
+ | Task 1: schema | — | A | agent-1 |
41
+ | Task 2: types | — | A | agent-2 |
42
+ | Task 3: API | 1, 2 | — | agent-1 |
43
+ | Verify | 3 | — | verifier (clean context) |
44
+ | Merge + handoff | Verify | — | merge owner |
45
+ ```
46
+
47
+ ## The Stop Rule
48
+
49
+ Before parallelizing, ask: *where does work split into pieces that never read each other's results?*
50
+
51
+ | Work shape | Strategy |
52
+ | --- | --- |
53
+ | **Splittable** — independent files, modules, or research angles | Parallel workers → separate verify → one merge owner |
54
+ | **Sequential** — each step needs the full picture from the previous step | Single agent, no fan-out |
55
+
56
+ Multi-agent setups win on splittable work and **lose** on sequential work. More agents is not a strategy — the shape of the work decides.
57
+
58
+ ## Fake Edges
59
+
60
+ For every "and then" in a task list, ask: does the next job actually read the previous job's output?
61
+
62
+ - **Real edge**: Task B imports or depends on artifacts from Task A → keep the dependency
63
+ - **Fake edge**: "Write tests and then update README" when README doesn't use test output → delete the edge; run in parallel or reorder
64
+
65
+ Most hand-built task lists contain 2–3 fake edges. Removing them unlocks safe parallelism.
66
+
67
+ ## The Diamond Pattern
68
+
69
+ The shape serious systems converge to for splittable work:
70
+
71
+ | Stage | What runs | Notes |
72
+ | --- | --- | --- |
73
+ | Plan | Single planner | Produces `tasks.md` + `task-graph.md` |
74
+ | Workers | 1–3 parallel agents | Disjoint file ownership only |
75
+ | Verify | Fresh verifier context | Never the code author |
76
+ | Merge | One merge owner | Resolves conflicts, runs harness |
77
+ | Result | Commits + evidence | Handoff to `/archive` when done |
78
+
79
+ Rules:
80
+
81
+ 1. **Split** only at real boundaries (see Stop Rule)
82
+ 2. **Workers** run in parallel with disjoint file ownership (see engineering-standards)
83
+ 3. **Verify** in a **separate context** — never the code author; ask diverse questions (correct? current? tested?)
84
+ 4. **Merge** has one owner who resolves conflicts and runs the project harness
85
+
86
+ Maps to SDD: `/tasks` → `/loop` (workers) → `/verify` (diamond verify node) → `/handoff` (merge + persist).
87
+
88
+ ## Sub-Agent Delegation
89
+
90
+ **Trigger** — Count the tasks. Roughly **8 or fewer** fits one batch: execute inline. More than that: offer sub-agents.
91
+
92
+ **Offer-then-confirm** — Never auto-spawn. Present the proposed split and wait for the owner to accept.
93
+
94
+ **Batching**
95
+
96
+ - A **batch** is the execution unit: one or more consecutive whole phases packed to about **7 tasks**.
97
+ - Phases stay the semantic unit — never split a phase across workers.
98
+ - Batches run **sequentially**; a batch starts only after the previous one reports every task complete.
99
+ - Each worker implements → gates → commits each of its tasks in order, then reports a compact summary: tasks done, commit hashes, test counts, deviations.
100
+ - Workers never spawn further sub-agents.
101
+
102
+ Full operational contract (worker payload, compact summary template, failure table, merge owner steps): `references/sub-agents.md`.
103
+
104
+ **Verifier** — After the final task, dispatch a fresh verifier regardless of batch count. It is the closing step of Execute, never prompted. See `references/validate.md`.
105
+
106
+ **Model tier per role** — When Spec Guardrails allows choosing a model per sub-agent:
107
+
108
+ | Role | Tier |
109
+ | --- | --- |
110
+ | Mechanical batch (config, wiring, CRUD) | Fast / cost-effective |
111
+ | Core-domain or ambiguous batch | High reasoning |
112
+ | Design phase | High reasoning |
113
+ | Verifier | Mid-to-high — adversarial reasoning, mutant design |
114
+
115
+ If Spec Guardrails cannot set a per-agent model, ignore this and invest more care on the heavy steps.
116
+
117
+ ## Human Gate
118
+
119
+ Route **irreversible** actions through explicit human approval:
120
+
121
+ - Deploy, publish, push to shared branch
122
+ - Delete data, refund, send external communications
123
+ - Merge to main/production
124
+
125
+ **Placement rule**: gate where a mistake is expensive to undo — not on every step.
126
+
127
+ | Action | Gate? |
128
+ | --- | --- |
129
+ | Local commit | No (handoff handles this) |
130
+ | `git push` | Yes — human or explicit instruction |
131
+ | Deploy / release | Yes — always |
132
+ | Continue after 3 failed correction loops | Yes — escalate to human |
133
+
134
+ ## Guardrails
135
+
136
+ 1. **Max correction loops** — 3 per task, then escalate (see `agent-architecture.md`)
137
+ 2. **One writer per file** — no two jobs mutate the same file in the same round
138
+ 3. **Plan in writing** — routing lives in `task-graph.md` and `tasks.md`; agents fill jobs, not the plan
139
+ 4. **Cap subagents** — hard limit on parallel agents (default: 3 workers + 1 verifier)
140
+ 5. **Evidence over self-report** — judge on harness output (tests ran, linter passed), not agent claims
141
+
142
+ ## Integration with SDD Phases
143
+
144
+ | SDD Phase | Task graph action |
145
+ | --- | --- |
146
+ | `/tasks` | Identify fake edges; mark parallel groups; link to REQ IDs |
147
+ | `/task-graph` | Draw or revise the DAG in `task-graph.md` |
148
+ | `/loop` | Execute graph — `loop-plan` each round; parallel sub-agents when files are disjoint |
149
+ | `/verify` | Diamond verify node — clean context, diverse checks |
150
+ | `/handoff` | Commit `task-graph.md` with other `.specs/` artifacts |
151
+
152
+ ## When NOT to Use
153
+
154
+ Skip the task graph for:
155
+
156
+ - Single-file bug fixes
157
+ - Copy or config-only changes
158
+ - Trivial changes covered by the complexity router in `agent-architecture.md`
159
+
160
+ ## Available Commands
161
+
162
+ | Command | Action |
163
+ | --- | --- |
164
+ | `/task-graph` | Draw or revise DAG in `.specs/features/[feature]/task-graph.md` |
165
+ | `/tasks` | Atomic breakdown — apply fake-edge and stop-rule checks |
166
+ | `/loop` | Execute graph respecting parallelism and one-writer-per-file |
167
+ | `/verify` | Independent verification (diamond verify node) |
168
+
169
+ ## Related Skills
170
+
171
+ - `agent-architecture.md` — SDD hub, contract, complexity router
172
+ - `references/tasks.md` — task schema and the `validate_tasks.py` gate
173
+ - `references/validate.md` — verifier procedure for the diamond verify node
174
+ - `engineering-standards.md` — one writer per file, git hygiene
175
+ - `security-review.md` — security checklist for verify node
176
+ - `git-handoff.md` — commit `task-graph.md` at phase boundaries
177
+
178
+ ## Credits
179
+
180
+ Task-graph patterns adapted under MIT license from [graph-engineering](https://github.com/codejunkie99/graph-engineering) (task-graph half only).
@@ -0,0 +1,61 @@
1
+ # Getting started
2
+
3
+ You installed the **Spec Guardrails**. You do **not** need to memorize CLI commands.
4
+
5
+ ## What to do now
6
+
7
+ 1. Open **Cursor** or **Claude Code** in this project.
8
+ 2. Start with **Specify** (an **agent command** — chat, not terminal):
9
+
10
+ ```
11
+ /specify
12
+
13
+ Add CSV export to the reports page. Users pick a date range.
14
+ Out of scope: PDF export.
15
+ ```
16
+
17
+ 3. Review `.specs/features/…/spec.md` and **approve** before implementation.
18
+
19
+ ---
20
+
21
+ ## Agent commands (chat — not terminal)
22
+
23
+ Type these in **Cursor or Claude Code**. They load phase procedures from `.cursor/skills/references/`. The agent runs gates for you.
24
+
25
+ | Command | When |
26
+ | --- | --- |
27
+ | `/specify` | **Start here** — written requirements |
28
+ | `/tasks` | Break into jobs after spec approval |
29
+ | `/loop` | Implement — agent runs `loop-plan` each wave |
30
+ | `/verify` | Fresh-context proof after last task |
31
+ | `/quick` | Tiny fix only (≤3 files) |
32
+
33
+ **Typical order:** `/specify` → `/tasks` → `/analyze` → `/loop` → `/verify` → `/archive`
34
+
35
+ **Full reference** (every command, examples, CLI): [Agent commands](https://github.com/luizssantiago92/spec-guardrails/blob/main/docs/guide/agent-commands.md)
36
+
37
+ ---
38
+
39
+ ## CLI you might run yourself (terminal)
40
+
41
+ | Command | When |
42
+ | --- | --- |
43
+ | `install` | First time or upgrade |
44
+ | `project-init` | Brownfield repo (optional) |
45
+ | `doctor` | Install looks broken |
46
+ | `validate-spec` / `validate-state` | Double-check gates manually |
47
+
48
+ Everything else (`loop-plan`, `validate-tasks`, `check-commit`, …) is normally run **by the agent**.
49
+
50
+ ---
51
+
52
+ ## Where things live
53
+
54
+ | Path | Purpose |
55
+ | --- | --- |
56
+ | `.cursor/skills/agent-architecture.md` | Hub — phase map |
57
+ | `.specs/STATE.md` | Where you left off |
58
+ | `.specs/features/` | One folder per feature |
59
+ | `.specs/guardrails/scripts/` | Automatic gates |
60
+
61
+ More guides: [docs/guide/Home.md](https://github.com/luizssantiago92/spec-guardrails/blob/main/docs/guide/Home.md)
@@ -0,0 +1,28 @@
1
+ # Project guardrails config (optional). Copy to .specs/config.yaml and edit.
2
+ # Or run: npx @luizsantiago/spec-guardrails init-config --preset node-ts
3
+ # Injected into phase procedures via phase-context.
4
+
5
+ schema: spec-driven
6
+
7
+ # Extend a built-in preset (default | node-ts | python):
8
+ # extends: node-ts
9
+
10
+ context: |
11
+ Tech stack: (fill in)
12
+ Test command: npm test
13
+ Branch prefix: feat
14
+
15
+ rules:
16
+ specify:
17
+ - Prefer EARS acceptance criteria
18
+ - Mark unknowns with [NEEDS CLARIFICATION: question]
19
+ tasks:
20
+ - Every REQ must appear in the Test Coverage Matrix
21
+ verify:
22
+ - Evidence must cite test file:line paths
23
+
24
+ # Project-specific overrides (appended on top of preset + rules above):
25
+ # overrides:
26
+ # rules:
27
+ # verify:
28
+ # - Also run playwright e2e before PASS
@@ -0,0 +1,16 @@
1
+ # Default preset — general-purpose projects
2
+ schema: spec-driven
3
+
4
+ context: |
5
+ Tech stack: (fill in)
6
+ Test command: npm test
7
+ Branch prefix: feat
8
+
9
+ rules:
10
+ specify:
11
+ - Prefer EARS acceptance criteria
12
+ - Mark unknowns with [NEEDS CLARIFICATION: question]
13
+ tasks:
14
+ - Every REQ must appear in the Test Coverage Matrix
15
+ verify:
16
+ - Evidence must cite test file:line paths