@mohammadhprp/system-prompt 0.11.1 → 0.11.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (95) hide show
  1. package/framework/agents/backend-architect.md +1 -1
  2. package/framework/mcps/figma-mcp-go/README.md +0 -1
  3. package/framework/mcps/gitlab-mcp/README.md +0 -1
  4. package/framework/mcps/jira-mcp/README.md +0 -1
  5. package/framework/mcps/laravel-boost/README.md +0 -1
  6. package/framework/mcps/notion-mcp/README.md +0 -1
  7. package/framework/mcps/supabase-mcp/README.md +0 -1
  8. package/framework/plugins/opencode-goal-plugin/README.md +0 -1
  9. package/framework/references/standards/api.md +0 -1
  10. package/framework/references/standards/architecture.md +0 -1
  11. package/framework/references/standards/database.md +0 -1
  12. package/framework/references/standards/debugging.md +0 -1
  13. package/framework/references/standards/documentation.md +0 -2
  14. package/framework/references/standards/logging.md +0 -1
  15. package/framework/references/standards/naming.md +0 -1
  16. package/framework/references/standards/observability.md +0 -1
  17. package/framework/references/standards/performance.md +0 -1
  18. package/framework/references/standards/pull-requests.md +0 -1
  19. package/framework/references/standards/security.md +0 -1
  20. package/framework/references/standards/testing.md +0 -1
  21. package/framework/skills/README.md +15 -3
  22. package/framework/skills/codenavi/SKILL.md +306 -0
  23. package/framework/skills/codenavi/examples.md +33 -0
  24. package/framework/skills/codenavi/references/coding-principles.md +143 -0
  25. package/framework/skills/codenavi/references/notebook-spec.md +171 -0
  26. package/framework/skills/create-adr/SKILL.md +429 -0
  27. package/framework/skills/create-adr/examples.md +35 -0
  28. package/framework/skills/docs-writer/SKILL.md +39 -0
  29. package/framework/skills/docs-writer/examples.md +34 -0
  30. package/framework/skills/docs-writer/references/style-guide.md +72 -0
  31. package/framework/skills/frontend-design/SKILL.md +55 -0
  32. package/framework/skills/frontend-design/examples.md +45 -0
  33. package/framework/skills/humanizer/SKILL.md +412 -0
  34. package/framework/skills/humanizer/examples.md +46 -0
  35. package/framework/skills/learning-opportunities/SKILL.md +140 -0
  36. package/framework/skills/learning-opportunities/examples.md +34 -0
  37. package/framework/skills/learning-opportunities/references/PRINCIPLES.md +42 -0
  38. package/framework/skills/perf-web-optimization/SKILL.md +163 -0
  39. package/framework/skills/perf-web-optimization/examples.md +35 -0
  40. package/framework/skills/perf-web-optimization/references/bundle-optimization.md +180 -0
  41. package/framework/skills/perf-web-optimization/references/core-web-vitals.md +154 -0
  42. package/framework/skills/perf-web-optimization/references/image-optimization.md +170 -0
  43. package/framework/skills/security-best-practices/LICENSE.txt +201 -0
  44. package/framework/skills/security-best-practices/SKILL.md +89 -0
  45. package/framework/skills/security-best-practices/examples.md +35 -0
  46. package/framework/skills/security-best-practices/references/golang-general-backend-security.md +988 -0
  47. package/framework/skills/security-best-practices/references/javascript-express-web-server-security.md +1151 -0
  48. package/framework/skills/security-best-practices/references/javascript-general-web-frontend-security.md +725 -0
  49. package/framework/skills/security-best-practices/references/javascript-jquery-web-frontend-security.md +672 -0
  50. package/framework/skills/security-best-practices/references/javascript-typescript-nextjs-web-server-security.md +1138 -0
  51. package/framework/skills/security-best-practices/references/javascript-typescript-react-web-frontend-security.md +975 -0
  52. package/framework/skills/security-best-practices/references/javascript-typescript-vue-web-frontend-security.md +789 -0
  53. package/framework/skills/security-best-practices/references/python-django-web-server-security.md +880 -0
  54. package/framework/skills/security-best-practices/references/python-fastapi-web-server-security.md +1030 -0
  55. package/framework/skills/security-best-practices/references/python-flask-web-server-security.md +835 -0
  56. package/framework/skills/sentry/SKILL.md +127 -0
  57. package/framework/skills/sentry/examples.md +34 -0
  58. package/framework/skills/sentry/scripts/sentry_api.py +238 -0
  59. package/framework/skills/show-me/SKILL.md +127 -0
  60. package/framework/skills/show-me/examples.md +78 -0
  61. package/framework/skills/spec-driven-eval/SKILL.md +341 -0
  62. package/framework/skills/spec-driven-eval/examples.md +35 -0
  63. package/framework/skills/spec-driven-eval/references/quickstart.md +118 -0
  64. package/framework/skills/spec-driven-eval/references/reference.md +295 -0
  65. package/framework/skills/technical-design-doc-creator/README.md +411 -0
  66. package/framework/skills/technical-design-doc-creator/SKILL.md +1484 -0
  67. package/framework/skills/technical-design-doc-creator/examples.md +35 -0
  68. package/framework/skills/tlc-spec-driven/SKILL.md +184 -0
  69. package/framework/skills/tlc-spec-driven/examples.md +34 -0
  70. package/framework/skills/tlc-spec-driven/references/code-analysis.md +98 -0
  71. package/framework/skills/tlc-spec-driven/references/coding-principles.md +72 -0
  72. package/framework/skills/tlc-spec-driven/references/context-limits.md +31 -0
  73. package/framework/skills/tlc-spec-driven/references/design.md +199 -0
  74. package/framework/skills/tlc-spec-driven/references/discuss.md +159 -0
  75. package/framework/skills/tlc-spec-driven/references/implement.md +436 -0
  76. package/framework/skills/tlc-spec-driven/references/lessons.md +115 -0
  77. package/framework/skills/tlc-spec-driven/references/memory.md +144 -0
  78. package/framework/skills/tlc-spec-driven/references/specify.md +228 -0
  79. package/framework/skills/tlc-spec-driven/references/sub-agents.md +147 -0
  80. package/framework/skills/tlc-spec-driven/references/tasks.md +451 -0
  81. package/framework/skills/tlc-spec-driven/references/validate.md +355 -0
  82. package/framework/skills/tlc-spec-driven/scripts/check_commit.py +115 -0
  83. package/framework/skills/tlc-spec-driven/scripts/lessons.py +412 -0
  84. package/framework/skills/tlc-spec-driven/scripts/validate_spec.py +260 -0
  85. package/framework/skills/tlc-spec-driven/scripts/validate_state.py +162 -0
  86. package/framework/skills/tlc-spec-driven/scripts/validate_tasks.py +251 -0
  87. package/framework/skills/web-design-guidelines/SKILL.md +65 -0
  88. package/framework/skills/web-design-guidelines/examples.md +32 -0
  89. package/framework/skills/web-design-guidelines/references/guideline.md +174 -0
  90. package/package.json +1 -1
  91. package/src/catalog.js +15 -3
  92. package/framework/skills/backend-engineer/SKILL.md +0 -76
  93. package/framework/skills/backend-engineer/examples.md +0 -31
  94. package/framework/skills/documentation/SKILL.md +0 -74
  95. package/framework/skills/documentation/examples.md +0 -31
@@ -0,0 +1,115 @@
1
+ # Lessons - Self-Improving Layer
2
+
3
+ **Purpose**: Turn verification failures into reusable, project-local guidance that actually changes future behavior - without the lessons file rotting into a dead log.
4
+
5
+ **The split that keeps it alive**: the agent (you) supplies *judgment* - read the failure, phrase the lesson, cite its grounding. The script `scripts/lessons.py` owns everything *mechanical* - IDs, recurrence counting across distinct features, candidate→confirmed promotion, pruning, demotion, and rendering. Hand-kept bookkeeping is exactly what rots, so it is not your job; the script's job.
6
+
7
+ **What feeds it**: only the execution signals already produced by the Verifier in [validate.md](validate.md) and written to `.specs/features/[feature]/validation.md`. No signal → no lesson. This is the hard gate: a lesson with no grounding in a real verification outcome is an opinion, and the script refuses it.
8
+
9
+ **Scope discipline (critical)**: this layer captures *execution* lessons that are project-local and grounded in a signal. It does **NOT** capture methodology opinions about the SDD process itself ("we should always discuss earlier"). Those are maintainer decisions that ship in a version bump - never auto-written. If a candidate lesson is really about how to run the skill rather than about this codebase, do not record it.
10
+
11
+ ---
12
+
13
+ ## Files
14
+
15
+ | File | Owner | Purpose |
16
+ | ---- | ----- | ------- |
17
+ | `.specs/lessons.json` | script | Canonical machine state. Never hand-edit. |
18
+ | `.specs/LESSONS.md` | script (rendered) | Human/agent-readable playbook. Read it; never write it by hand. |
19
+ | `<skill-dir>/scripts/lessons.py` | package | The only way to mutate lessons. Invoke via the skill directory - never `python3 scripts/lessons.py` from the project root. |
20
+
21
+ `confirmed` lessons are the playbook the agent loads. `candidate` lessons are tracked but NOT trusted until corroborated across `promote_threshold` distinct features (default 2). `quarantined` lessons failed when applied and are ignored.
22
+
23
+ **Invocation:** resolve `<skill-dir>` as the directory that contains this skill's `SKILL.md`, then run `python3 <skill-dir>/scripts/lessons.py ...`. The store under `.specs/` is still relative to the project root (use `--root` when cwd differs).
24
+
25
+ ---
26
+
27
+ ## WRITE - distill lessons (runs inside Execute, after validation)
28
+
29
+ This is **not a new phase**. It is the final action of the Verifier step in [validate.md](validate.md), grafted onto a step that already always runs. Do it immediately after `validation.md` is written, before reporting completion.
30
+
31
+ ### When to write
32
+
33
+ Walk the just-written `validation.md`. For each **grounded** signal, record one lesson:
34
+
35
+ | validation.md signal | `--signal` value |
36
+ | -------------------- | ---------------- |
37
+ | An acceptance criterion failed or had no evidence | `ac_gap` |
38
+ | A discrimination-sensor mutant survived (weak test) | `surviving_mutant` |
39
+ | A criterion flagged ⚠️ Spec-precision gap | `spec_precision_gap` |
40
+ | A `// SPEC_DEVIATION` marker was added during implement | `spec_deviation` |
41
+ | The build-level gate check failed | `gate_fail` |
42
+
43
+ If `validation.md` is a clean PASS with no surviving mutants, no spec-precision gaps, and no deviations → **write nothing**. A clean run produces no lessons. This is correct, not a miss.
44
+
45
+ ### How to write
46
+
47
+ For each signal, phrase the lesson as **one terse, actionable, codebase-general sentence** - a rule a future feature could apply, not a restatement of this bug. Then call the script:
48
+
49
+ ```bash
50
+ python3 <skill-dir>/scripts/lessons.py add \
51
+ --feature "[feature folder name]" \
52
+ --signal "[signal value from table above]" \
53
+ --source "[file:line | AC id | mutant id | SPEC_DEVIATION ref from validation.md]" \
54
+ --text "[the one-sentence lesson]" \
55
+ --scope "[optional: path/layer/tag, e.g. billing, routes, repo-layer]"
56
+ ```
57
+
58
+ **Phrasing rules** (they make recurrences actually merge - dedup is exact-after-normalization, not semantic):
59
+
60
+ - Write the general rule, not the incident. ✅ `"Assert the exact persisted status value, not just that a status field exists"` ❌ `"The subscription test on line 88 was too weak"`.
61
+ - Be canonical and terse. Two lessons that mean the same thing must read the same way, or the script counts them as different and neither gets promoted.
62
+ - One lesson per signal. Don't bundle.
63
+
64
+ `--source` is **mandatory**. The script exits non-zero if it is empty - that is the grounding gate working, not an error to route around.
65
+
66
+ ### Self-check (do not skip)
67
+
68
+ After distilling, if `validation.md` contained any FAIL, surviving mutant, spec-precision gap, or SPEC_DEVIATION but you recorded zero lessons, state plainly in chat: *"Validation had signal X but no lesson was recorded - recording now / here's why it's out of scope."* Silent skipping is how the file dies.
69
+
70
+ ### Demotion
71
+
72
+ If a `confirmed` lesson was loaded for this feature (see READ below) and the *same* failure recurred anyway, the guidance is not working:
73
+
74
+ ```bash
75
+ python3 <skill-dir>/scripts/lessons.py penalize --id L-NNN
76
+ ```
77
+
78
+ Two penalties quarantine it. Use sparingly and only on real repeats.
79
+
80
+ ---
81
+
82
+ ## READ - load lessons (runs at Specify and Design)
83
+
84
+ A lessons file nobody reads is dead by definition. Loading is **mandatory**, not optional.
85
+
86
+ At the start of **Specify** (and again at **Design** for Large/Complex), load the confirmed lessons relevant to this feature:
87
+
88
+ ```bash
89
+ # All confirmed lessons:
90
+ python3 <skill-dir>/scripts/lessons.py list --status confirmed
91
+
92
+ # Or filter by the area this feature touches:
93
+ python3 <skill-dir>/scripts/lessons.py list --status confirmed --scope billing
94
+ python3 <skill-dir>/scripts/lessons.py list --status confirmed --query "idempotency"
95
+ ```
96
+
97
+ Apply the returned lessons as guidance while writing the spec / design. Do **not** load `candidate` or `quarantined` lessons as guidance - they are not trusted. Keep the loaded set small; this runs inside the <40k token budget.
98
+
99
+ ---
100
+
101
+ ## Fallback when code execution is unavailable
102
+
103
+ Some harnesses cannot run Python. Only then: maintain `.specs/LESSONS.md` by hand, following the exact same rules - grounded entries only, candidate→confirmed after 2 distinct features, prune stale candidates. **This path is degraded**: hand bookkeeping is the failure mode this layer exists to avoid, so prefer the script wherever a code tool exists. State once in chat that you are in the no-script fallback so the user knows accounting is best-effort.
104
+
105
+ ---
106
+
107
+ ## Disable
108
+
109
+ This layer is additive and self-gating (no signal → no write). To turn it off for a project, delete `.specs/lessons.json` and `.specs/LESSONS.md` and skip the WRITE/READ steps. The core Specify→Design→Tasks→Execute flow is unaffected.
110
+
111
+ ---
112
+
113
+ ## Known limitation
114
+
115
+ Deduplication is exact-after-normalization (Unicode casefold, diacritic-stripped, punctuation-stripped, any-script alnum preserved) - there are no embeddings (stdlib-only, zero-dependency by design). Near-duplicate lessons phrased differently will not merge and will each sit as separate candidates that never promote. Mitigation: follow the phrasing rules above. A future version may add embedding-based dedup.
@@ -0,0 +1,144 @@
1
+ # Memory Layer
2
+
3
+ **File:** `.specs/STATE.md`
4
+
5
+ A single file with two section-scoped parts. Each section has its own lifecycle; writes are always targeted - never whole-file overwrites.
6
+
7
+ ---
8
+
9
+ ## Sections
10
+
11
+ ### `## Decisions` - append-only log
12
+
13
+ Records **project-level** decisions only: conventions, patterns, constraints, or cross-cutting technology choices that future features must follow or supersede.
14
+
15
+ **Not project-level → stays in the feature's `design.md` Tech Decisions table.**
16
+ Heuristic: would a different feature need to know about this? If yes → project-level. If no → feature-local.
17
+
18
+ **Record sparingly - the log stays useful only by staying small.** Even a project-level decision earns an `AD-NNN` entry only when all three hold:
19
+
20
+ 1. **Hard to reverse** - changing course later carries real cost.
21
+ 2. **Surprising without context** - a future reader will look at the result and wonder "why did they do it this way?"
22
+ 3. **The product of a real trade-off** - there were genuine alternatives and you chose one for specific reasons.
23
+
24
+ If any one is missing, skip it: an easily-reversed choice you will just reverse; an unsurprising one nobody questions; a no-alternative choice records nothing beyond "we did the obvious thing." What typically qualifies: architectural shape, integration patterns between areas, technology choices that carry lock-in, boundary and ownership decisions, and deliberate deviations from the obvious path. A choice that clears all three but is only feature-local still stays in `design.md`.
25
+
26
+ **Format** (one entry per decision):
27
+
28
+ ```markdown
29
+ ## Decisions
30
+
31
+ ### AD-001
32
+ - **Decision**: [what was decided - one sentence]
33
+ - **Reason**: [why this option was chosen]
34
+ - **Trade-off**: [what was given up]
35
+ - **Scope**: [which features / packages / layers this governs]
36
+ - **Date**: YYYY-MM-DD
37
+ - **Status**: active | superseded by AD-NNN
38
+ ```
39
+
40
+ **Supersession rule:** When a new decision replaces an old one, append a new `AD-NNN` entry and update the old entry's `status` field to `superseded by AD-NNN`. Never delete old entries - the history is the audit trail.
41
+
42
+ ---
43
+
44
+ ### `## Handoff` - pause snapshot (~500 tokens, overwritten each pause)
45
+
46
+ Captures mid-task / in-flight state so work can resume without re-reading the full task history. It complements `tasks.md` and git evidence: on resume, the Handoff is a starting hypothesis that must be reconciled against the real branch, commits, and working tree (see Resume below).
47
+
48
+ **Format:**
49
+
50
+ ```markdown
51
+ ## Handoff
52
+
53
+ - **Feature**: [feature name / .specs path]
54
+ - **Phase / Task**: [e.g., Phase 2 / T4 - implement repository layer]
55
+ - **Completed**: [comma-separated task IDs or "none"]
56
+ - **In-progress** (file:line): [e.g., `src/billing/subscription.service.ts:88` - mid-write]
57
+ - **Next step**: [one sentence - exactly what to do next]
58
+ - **Blockers**: [none | description]
59
+ - **Uncommitted files**: [list or "none"]
60
+ - **Branch**: [git branch name]
61
+ ```
62
+
63
+ ---
64
+
65
+ ## File shape
66
+
67
+ ```markdown
68
+ # STATE
69
+
70
+ ## Decisions
71
+
72
+ [AD-NNN entries…]
73
+
74
+ ## Handoff
75
+
76
+ [latest snapshot…]
77
+ ```
78
+
79
+ If the file does not yet exist, create it with both section headers and empty bodies.
80
+
81
+ ---
82
+
83
+ ## Read / Write Triggers
84
+
85
+ | Trigger | Section | Operation |
86
+ | ------- | ------- | --------- |
87
+ | Design phase, Step 1 (Load Context) | `## Decisions` | **Read** - conform to active decisions or supersede |
88
+ | Design phase, Tech Decisions step | `## Decisions` | **Append** - only for project-level decisions |
89
+ | Pause work / end of session | `## Handoff` | **Replace** - overwrite Handoff section only |
90
+ | Resume work / start of session | `## Handoff` | **Read** - load snapshot, then reconcile with git before acting |
91
+ | Resume work / start of session | `## Decisions` | **Read** - re-confirm active constraints before designing |
92
+
93
+ ---
94
+
95
+ ## Section-scoped write rule (critical)
96
+
97
+ One file holds two lifecycles. Writes MUST target their section only:
98
+
99
+ - **Design appends** to `## Decisions`. It MUST NOT touch `## Handoff`.
100
+ - **Pause replaces** `## Handoff`. It MUST NOT rewrite, reorder, or drop any entry in `## Decisions`.
101
+
102
+ The correct technique: locate the target section header, replace only the content between it and the next `##` header (or end of file). Never overwrite the full file.
103
+
104
+ Violating this rule causes one of two failures:
105
+ 1. A pause write clobbers the decisions log → decisions are silently lost.
106
+ 2. A design append touches the handoff snapshot → mid-task state is corrupted.
107
+
108
+ Both are silent data loss. The section-scoped write rule is the single correctness invariant of this memory layer.
109
+
110
+ ---
111
+
112
+ ## Pause / Resume Procedure
113
+
114
+ ### Pause
115
+
116
+ 1. Locate the `## Handoff` section in `.specs/STATE.md`.
117
+ 2. Replace its body (everything between `## Handoff` and the next `##` or EOF) with the current snapshot.
118
+ 3. Do NOT modify anything above or before `## Handoff`.
119
+ 4. Commit or stash outstanding changes as appropriate.
120
+
121
+ ### Resume
122
+
123
+ 1. Read `.specs/STATE.md` - both sections.
124
+ 2. Re-confirm active decisions from `## Decisions` - nothing superseded since last session?
125
+ 3. Read `## Handoff` - treat it as a **hypothesis** for feature, phase/task, next step, blockers, uncommitted files, branch - not as ground truth by itself.
126
+ 4. **Reconcile with git before editing anything:**
127
+ - Current branch vs Handoff `Branch`
128
+ - `git status --porcelain` (uncommitted / unexpected paths)
129
+ - Recent commits on the branch (messages and touched files)
130
+ - `tasks.md` completion marks and, when present, gate evidence / commit references
131
+ 5. **Resolve conflicts with evidence, not narrative:**
132
+ - A task with a green gate and an atomic commit already on the branch → do **not** redo it; mark it complete in `tasks.md` if the file still shows it open, then continue from the next incomplete task
133
+ - Partial unverified work in the working tree → preserve it, re-run the relevant gate, then finish the status+commit cycle
134
+ - Stale or missing Handoff → rebuild next-step from git + `tasks.md`, then propose that to the user
135
+ - Unexplained local changes you cannot map to the current task → STOP and ask; do not discard them
136
+ 6. Propose the reconciled next step to the user before writing any code.
137
+
138
+ ---
139
+
140
+ ## AD-NNN numbering
141
+
142
+ - Numbers are sequential, project-scoped, and permanent - never reused.
143
+ - The counter starts at `AD-001`. Check existing entries before assigning the next number.
144
+ - If `.specs/STATE.md` does not exist, the first decision is `AD-001`.
@@ -0,0 +1,228 @@
1
+ # Specify
2
+
3
+ **Goal**: Capture WHAT to build with testable, traceable requirements.
4
+
5
+ If the feature has ambiguous gray areas (multiple valid approaches for user-facing behavior), the agent will automatically trigger the [discuss gray areas](discuss.md) process within this phase. For clear, well-defined features, it goes straight to the next phase.
6
+
7
+ ## Implicit-Requirement Dimensions
8
+
9
+ The canonical rubric for requirements that are easy to miss. Referenced by [discuss.md](discuss.md) - defined here, not duplicated.
10
+
11
+ | Dimension | What to cover |
12
+ | --------- | ------------- |
13
+ | Input validation & bounds | Limits, formats, sanitization |
14
+ | Failure / partial-failure states | Timeouts, partial saves, rollbacks |
15
+ | Idempotency / retry / duplicate handling | Safe retries, dedup keys |
16
+ | Auth boundaries & rate limits | Who can call what, throttle rules |
17
+ | Concurrency / ordering | Race conditions, ordering guarantees |
18
+ | Data lifecycle / expiry | TTL, archival, deletion |
19
+ | Observability | Logging, metrics, tracing hooks |
20
+ | External-dependency failure | Circuit breakers, fallbacks |
21
+ | State-transition integrity | Valid transitions, guards |
22
+
23
+ ---
24
+
25
+ ## Process
26
+
27
+ ### 1. Clarify Requirements
28
+
29
+ **Load confirmed lessons first:** Before clarifying, load the project's confirmed lessons so past verification failures shape this spec instead of repeating. Run `python3 <skill-dir>/scripts/lessons.py list --status confirmed` (optionally `--scope [area]` or `--query [term]` for the area this feature touches) and apply what comes back as guidance. Load only `confirmed` - never `candidate` or `quarantined`. If no store exists yet or no code tool is available, skip silently. See [lessons.md](lessons.md).
30
+
31
+ **Lightweight context scan first (Knowledge Verification Chain Step 1):** Before asking questions, briefly scan existing code, patterns, and neighboring features relevant to this feature. Use what you find to ground your clarifying questions in reality - not to constrain the spec to current implementation. Keep it lightweight (stay within the <40k token budget; reuse the chain, no new machinery). The spec captures WHAT is needed, not only what exists.
32
+
33
+ You are a thinking partner, not an interviewer. Start open - let the user dump their mental model. Follow the energy: whatever they emphasize, dig into that.
34
+
35
+ Ask conversationally (not as a checklist):
36
+
37
+ - "What problem are you solving?"
38
+ - "Who is the user and what's their pain?"
39
+ - "What does success look like?"
40
+
41
+ If needed:
42
+
43
+ - "What are the constraints (time, tech, resources)?"
44
+ - "What is explicitly out of scope?"
45
+
46
+ **Facts you look up; decisions you ask.** Anything discoverable by reading the environment (the codebase, config, docs, existing conventions) you resolve yourself through the Knowledge Verification Chain - do not spend the user's attention asking for it. Reserve questions for genuine decisions that are the user's to make: scope, priorities, product behavior, trade-offs. A question you could have answered by reading the code erodes trust and wastes a turn.
47
+
48
+ **Challenge vagueness.** Never accept fuzzy answers. "Good" means what? "Users" means who? "Simple" means how? Make the abstract concrete: "Walk me through using this." "What does that actually look like?"
49
+
50
+ **Know when to stop - then run the dimensions sweep.** When you understand what they're building, why, who it's for, and what done looks like, run a closing **implicit-requirement dimensions sweep** before offering to proceed:
51
+
52
+ - **Large / Complex:** Cover every dimension above - each must resolve to a requirement OR an explicit `N/A because [reason]`. No blank entries allowed.
53
+ - **Medium:** Cover only dimensions obviously present for this feature's domain; collapse the rest to a single `remaining dimensions N/A for this scope`.
54
+ - **Small:** Skip the sweep entirely.
55
+
56
+ The `N/A because...` escape is mandatory - it prevents inventing requirements to fill the checklist. Bound the sweep to THIS feature's scope; never add requirements outside the feature boundary.
57
+
58
+ ### 2. Capture User Stories with Priorities
59
+
60
+ **P1 = MVP** (must ship), **P2** (should have), **P3** (nice to have)
61
+
62
+ Each story MUST be **independently testable** - you can implement and demo just that story.
63
+
64
+ ### 3. Write Acceptance Criteria (EARS notation)
65
+
66
+ Write every acceptance criterion in **EARS** (Easy Approach to Requirements Syntax). Each criterion resolves to exactly one pattern, which keeps it unambiguous and directly testable. Choose the pattern that fits the requirement instead of forcing everything into a single shape:
67
+
68
+ | Pattern | Keyword | Template | Use for |
69
+ | ------- | ------- | -------- | ------- |
70
+ | Ubiquitous | (none) | The [system] SHALL [response] | Always-on invariants and constraints |
71
+ | Event-driven | WHEN | WHEN [trigger] THEN the [system] SHALL [response] | A response to a discrete trigger |
72
+ | State-driven | WHILE | WHILE [state] the [system] SHALL [response] | Behavior that holds during a state |
73
+ | Optional-feature | WHERE | WHERE [feature is present] the [system] SHALL [response] | Behavior gated behind an optional capability or flag |
74
+ | Unwanted-behavior | IF / THEN | IF [undesired condition] THEN the [system] SHALL [response] | Errors, failures, invalid input, timeouts |
75
+ | Complex | combination | WHILE [state], WHEN [trigger] the [system] SHALL [response] | Richer behavior combining the above |
76
+
77
+ **Why patterns beat one shape:** failure states, state transitions, and optional behavior become first-class criteria instead of footnotes squeezed into WHEN/THEN. The patterns map onto the implicit-requirement dimensions above: state-transition integrity to State-driven; failure and external-dependency failure to Unwanted-behavior; feature flags to Optional-feature.
78
+
79
+ **Rules:** one requirement per criterion (never bundle two behaviors); use concrete values (a specific status code, a specific message, a bound) rather than "quickly" or "gracefully"; every criterion contains a SHALL and is measurable. `python3 <skill-dir>/scripts/validate_spec.py` flags any criterion without a SHALL and any that matches no recognized pattern.
80
+
81
+ ### 4. Requirement Closure Gate (before confirm)
82
+
83
+ Before presenting the spec for confirmation, run the three checks below. The spec is not presentable for confirmation until every item is resolved or assumption-logged - this is the guarantee that no requirement leaves the spec silently unclear.
84
+
85
+ **Scope-tiered:** Large/Complex = full gate; Medium = resolve obvious ambiguities, log the rest as assumptions; Small = skip entirely (consistent with skipping the sweep).
86
+
87
+ 1. **Unambiguity + precision (hard).** Every AC must (a) have a single interpretation and (b) define a precise, spec-defined expected outcome. Any AC that fails either check: resolve with the user, split it, or log it as an explicit assumption with the chosen interpretation and rationale. No AC proceeds readable two ways or with an undefined outcome.
88
+
89
+ 2. **Open-questions / assumptions closure.** Enumerate every unresolved decision that surfaced during clarification. Each must be either (a) resolved with the user OR (b) recorded as an **assumption** (chosen default + rationale) in the spec's Assumptions & Open Questions section. Nothing proceeds unmarked.
90
+
91
+ 3. **Declined gray areas become assumptions.** Any gray area the user declined to discuss or that went undiscussed is written to the spec's Assumptions & Open Questions section (agent's chosen default + rationale) - never silently dropped. See [discuss.md](discuss.md).
92
+
93
+ Fix inline. This gate is bounded to THIS feature's stated dimensions and actual behavior - never to "anything imaginable." The Out of Scope table and anti-scope-creep rules remain the counterweights: the gate clarifies existing requirements, it never invents new ones.
94
+
95
+ **Deterministic backing (run before you present the spec).** The structural half of this gate is enforced by a script so it cannot drift when a step is forgotten: `python3 <skill-dir>/scripts/validate_spec.py <spec-path-or-feature>` checks that required sections exist, every AC is EARS-shaped (has a SHALL), no Assumptions row has an empty default or rationale, and requirement IDs are well-formed. A non-zero exit means fix before confirming. The script checks structure; you still own the judgment calls (is the interpretation right, is the outcome precise). If no code-execution tool is available, run the same checks by reading the spec.
96
+
97
+ ---
98
+
99
+ ## Template: `.specs/features/[feature]/spec.md`
100
+
101
+ ```markdown
102
+ # [Feature Name] Specification
103
+
104
+ ## Problem Statement
105
+
106
+ [Describe the problem in 2-3 sentences. What pain point are we solving? Why now?]
107
+
108
+ ## Goals
109
+
110
+ - [ ] [Primary goal with measurable outcome]
111
+ - [ ] [Secondary goal with measurable outcome]
112
+
113
+ ## Out of Scope
114
+
115
+ Explicitly excluded. Documented to prevent scope creep.
116
+
117
+ | Feature | Reason |
118
+ | ----------- | -------------- |
119
+ | [Feature X] | [Why excluded] |
120
+ | [Feature Y] | [Why excluded] |
121
+
122
+ ---
123
+
124
+ ## Assumptions & Open Questions
125
+
126
+ Every ambiguity is resolved or recorded here - nothing is left silently unclear.
127
+
128
+ | Assumption / decision | Chosen default | Rationale | Confirmed? |
129
+ | --------------------- | --------------- | --------- | ---------- |
130
+ | [ambiguity] | [what we'll do] | [why] | [y/n] |
131
+
132
+ **Open questions:** none - all resolved or logged above (required before the spec is confirmed).
133
+
134
+ ---
135
+
136
+ ## User Stories
137
+
138
+ ### P1: [Story Title] ⭐ MVP
139
+
140
+ **User Story**: As a [role], I want [capability] so that [benefit].
141
+
142
+ **Why P1**: [Why this is critical for MVP]
143
+
144
+ **Acceptance Criteria** (each line is one EARS pattern):
145
+
146
+ 1. WHEN [user action/event] THEN system SHALL [expected behavior] <!-- event-driven -->
147
+ 2. IF [invalid input / failure] THEN system SHALL [graceful handling] <!-- unwanted-behavior -->
148
+ 3. WHILE [state holds] system SHALL [behavior during that state] <!-- state-driven -->
149
+ 4. The system SHALL [always-on invariant] <!-- ubiquitous -->
150
+
151
+ **Independent Test**: [How to verify this story works alone - e.g., "Can demo by doing X and seeing Y"]
152
+
153
+ ---
154
+
155
+ ### P2: [Story Title]
156
+
157
+ **User Story**: As a [role], I want [capability] so that [benefit].
158
+
159
+ **Why P2**: [Why this isn't MVP but important]
160
+
161
+ **Acceptance Criteria**:
162
+
163
+ 1. WHEN [event] THEN system SHALL [behavior]
164
+ 2. WHEN [event] THEN system SHALL [behavior]
165
+
166
+ **Independent Test**: [How to verify]
167
+
168
+ ---
169
+
170
+ ### P3: [Story Title]
171
+
172
+ **User Story**: As a [role], I want [capability] so that [benefit].
173
+
174
+ **Why P3**: [Why this is nice-to-have]
175
+
176
+ **Acceptance Criteria**:
177
+
178
+ 1. WHEN [event] THEN system SHALL [behavior]
179
+
180
+ ---
181
+
182
+ ## Edge Cases
183
+
184
+ Edge cases are usually unwanted-behavior (IF/THEN) or boundary (WHEN) criteria:
185
+
186
+ - IF [error scenario] THEN system SHALL [graceful handling]
187
+ - IF [unexpected input] THEN system SHALL [validation response]
188
+ - WHEN [boundary condition] THEN system SHALL [behavior]
189
+
190
+ ---
191
+
192
+ ## Requirement Traceability
193
+
194
+ Each requirement gets a unique ID for tracking across design, tasks, and validation.
195
+
196
+ | Requirement ID | Story | Phase | Status |
197
+ | -------------- | ----------- | ------ | ------- |
198
+ | [FEAT]-01 | P1: [Story] | Design | Pending |
199
+ | [FEAT]-02 | P1: [Story] | Design | Pending |
200
+ | [FEAT]-03 | P2: [Story] | - | Pending |
201
+
202
+ **ID format:** `[CATEGORY]-[NUMBER]` (e.g., `AUTH-01`, `CART-03`, `NOTIF-02`)
203
+
204
+ **Status values:** Pending → In Design → In Tasks → Implementing → Verified
205
+
206
+ **Coverage:** X total, Y mapped to tasks, Z unmapped ⚠️
207
+
208
+ ---
209
+
210
+ ## Success Criteria
211
+
212
+ How we know the feature is successful:
213
+
214
+ - [ ] [Measurable outcome - e.g., "User can complete X in < 2 minutes"]
215
+ - [ ] [Measurable outcome - e.g., "Zero errors in Y scenario"]
216
+ ```
217
+
218
+ ---
219
+
220
+ ## Tips
221
+
222
+ - **P1 = Vertical Slice** - A complete, demo-able feature, not just backend or frontend
223
+ - **EARS is code** - If you can't write a criterion as a test, rewrite it; pick the pattern (WHEN / WHILE / WHERE / IF / ubiquitous) that fits
224
+ - **Requirement IDs are mandatory** - Every story maps to trackable IDs
225
+ - **Edge cases matter** - What breaks? What's empty? What's huge?
226
+ - **Out of Scope prevents creep** - If it's not here, it doesn't get built
227
+ - **Closure gate before confirm** - Three checks: unambiguity + precision, open-questions/assumptions closure, declined gray areas logged; scope-tiered; bounded to stated dimensions; never invents requirements
228
+ - **Confirm after the gate passes** - Present the spec for user confirmation only after the closure gate passes (no unresolved-and-unmarked items remain) and `validate_spec.py` exits clean; user approves spec before moving to discuss phase
@@ -0,0 +1,147 @@
1
+ # Sub-Agent Delegation
2
+
3
+ Full mechanics for phase-batch workers and the Verifier sub-agent used during Execute.
4
+
5
+ ## Phase-Batch Workers
6
+
7
+ **Two layers - keep them distinct:**
8
+
9
+ - **Phase** = the semantic / dependency unit (Foundation → Core → Integration), authored during Tasks. Indivisible.
10
+ - **Batch** = the execution / logistics unit - one or more *consecutive whole phases* assigned to a single worker.
11
+
12
+ Conflating the two (one worker per phase) is what fragments execution: a feature's dependency-layer count has nothing to do with the ideal per-worker workload. Batching by task budget separates the two concerns without breaking phases.
13
+
14
+ **Trigger:** Count total tasks across all phases. If the feature packs into **more than one batch** (> ~8 tasks), offer the user phase-batch sub-agents before starting Execute. If it fits a single batch (≤ ~8 tasks), execute inline in the main window - no sub-agents spawned.
15
+
16
+ **Batching algorithm (task budget ≈ 7 tasks/worker, phase-aligned):**
17
+
18
+ The benchmarked sweet spot is ~7 tasks of context per worker (~20 tasks → 3 workers). Pack whole phases into that budget:
19
+
20
+ 1. Count total tasks `T`.
21
+ 2. If `T ≤ ~8` → inline, no sub-agents.
22
+ 3. Otherwise walk phases **in order**, accumulating whole phases into the current batch. When the batch's running task count reaches ~7 **and** phases remain, close the batch and start the next.
23
+ 4. **Never split a phase** across workers - the cut only ever lands on a phase boundary. This preserves dependency ordering and keeps a phase's tasks + shared context in one worker.
24
+ 5. If the final batch is a lone tail (1-2 tasks), fold it into the previous batch.
25
+
26
+ Result ≈ `ceil(T / 7)` workers, scaling linearly. Unevenness is absorbed by greedy packing - phases never need to divide evenly. Worked examples (20 tasks):
27
+
28
+ - Phases `[3,3,3,3,4,4]` → `{P1+P2=6, P3+P4=6, P5+P6=8}` = **3 workers**
29
+ - Phases `[8,2,2,8]` → `{P1=8, P2+P3=4, P4=8}` = **3 workers** (no even split needed)
30
+ - Phases `[5,5,5,5]` → `{P1+P2=10, P3+P4=10}` = **2 workers** (phases too coarse to hit 3 - see below)
31
+
32
+ **Coarse-phase caveat:** Because the cut lands only on phase boundaries, very coarse phases limit how finely you can pack. If a single phase alone exceeds ~1.5× the budget (~10+ tasks), that is a Tasks-authoring smell - split it into real sub-phases during Tasks (at a genuine dependency/cohesion boundary), never at dispatch time.
33
+
34
+ **Offer-then-confirm (never auto-spawn):**
35
+
36
+ > "This feature has [T] tasks across [N] phases. I can pack them into [K] sub-agents (~7 tasks each, whole phases per worker) - every worker runs its phases in order, reports a compact summary, and the orchestrator advances to the next batch. This keeps the main window lean without over-fragmenting. Want to proceed that way?"
37
+
38
+ The user must explicitly accept. If they decline (or if the feature fits one batch), execute inline.
39
+
40
+ **Execution model - one worker per task-budgeted batch, sequential:**
41
+
42
+ ```
43
+ Phases 1+2 (7 tasks) ------→ Batch Worker 1 ------→ compact summary ------→ orchestrator updates tasks.md
44
+ Phases 3+4 (6 tasks) ------→ Batch Worker 2 ------→ compact summary ------→ orchestrator updates tasks.md
45
+ Phase 5 (7 tasks) ------→ Batch Worker 3 ------→ compact summary ------→ orchestrator updates tasks.md
46
+ ...
47
+ ```
48
+
49
+ Batches run strictly sequentially: a batch never starts until the previous batch's summary shows all its tasks complete.
50
+
51
+ **What a batch worker receives:**
52
+
53
+ - The task definitions for **every** phase in its batch (from `tasks.md`)
54
+ - The Test Coverage Matrix and Gate Check Commands (from `tasks.md`)
55
+ - `references/coding-principles.md`
56
+ - Relevant `spec.md` and `design.md` context for the feature (not all specs)
57
+
58
+ **What a batch worker does:**
59
+
60
+ Executes ALL tasks in its assigned batch **in order** - finishing every task in one phase before starting the next phase in the batch - following the `implement.md` cycle for each task (implement → gate → atomic commit). It does NOT spawn further sub-agents. After completing all tasks in the batch, the worker reports a **compact summary** to the orchestrator:
61
+
62
+ ```
63
+ Batch (phases [N]-[M]) complete:
64
+ - Tasks done: [list with commit hashes]
65
+ - Tests: [N passed, 0 failed]
66
+ - Deviations/blockers: [none | description]
67
+ ```
68
+
69
+ No raw logs, no full test output - only the above fields keep the main context clean.
70
+
71
+ **No nesting:** Batch workers execute their tasks themselves. They never spawn sub-sub-agents. Execution is strictly sequential within and across batches - there is no intra-phase or intra-batch parallelism.
72
+
73
+ **The orchestrating agent's role during Execute:**
74
+
75
+ 1. Count total tasks and pack phases into task-budgeted batches (~7 tasks each) - if that yields more than one batch, offer batch sub-agents and wait for the user to accept
76
+ 2. Dispatch the next batch to a worker (or execute inline if not using sub-agents)
77
+ 3. Receive the compact summary
78
+ 4. Update `tasks.md` with results
79
+ 5. If all tasks in the summary show complete: dispatch the next batch
80
+ 6. If a task failed: the worker has already stopped; decide fix/escalate before dispatching the next batch
81
+
82
+ **Failure handling:** If a task in a batch fails (gate does not pass, blocker hit), the worker stops and includes the failure in its summary. The next batch does not start until the current batch's summary shows all tasks complete. The orchestrator decides: fix and re-run, or escalate to the user.
83
+
84
+ **Context sizing signal:** If a batch's task list would likely push the worker's context beyond ~40k tokens, close the batch at an earlier phase boundary (fewer phases per worker). If a *single* phase alone would blow the budget, that phase is too coarse - split it during Tasks per the granularity guidance in `references/tasks.md`.
85
+
86
+ ---
87
+
88
+ ## Verifier Sub-Agent
89
+
90
+ **Always-on, never prompted - one per feature completion.** The Verifier is a separate role from the batch worker. It runs once - after the last task of the feature is committed - as an independent quality gate, dispatched automatically by the orchestrator. It is **not** gated behind the batching offer; it always runs. Do NOT ask the user whether to run validation; it is mandatory.
91
+
92
+ **Author ≠ verifier:** The agent (or batch worker) that wrote the code and tests is the author. The Verifier is a fresh sub-agent dispatched by the orchestrator after the final commit. It does not inherit the author's context, mental model, or assumptions. This separation is what makes the gate trustworthy.
93
+
94
+ **What the Verifier receives:**
95
+ - `spec.md` for the feature (ACs = source of truth)
96
+ - The git diff surface for the feature (scoped to the feature branch or commit range)
97
+ - The test files in scope
98
+ - `references/validate.md` as its operating checklist
99
+
100
+ **What the Verifier does (full process in `validate.md`):**
101
+ 1. **Spec-anchored coverage check** - re-derives coverage evidence-or-zero: every AC traced to `file:line` + assertion expression. For each covered criterion, confirms the test's asserted value matches the **spec-defined expected outcome** (not just that an assertion exists). Where the spec does not define a precise outcome, flags a **spec-precision gap** rather than passing silently.
102
+ 2. **Discrimination sensor** - injects a small behavior-level fault (flip a condition, change a return value, off-by-one, remove a required side effect) in an **isolated scratch** (temporary `git worktree` or temp file copies - never `git stash`), runs the relevant tests there, confirms they FAIL (kill the mutant), discards the scratch, and verifies the real worktree's `git status --porcelain` matches the pre-sensor baseline. Tiered by risk: lightweight (1-3 mutations) for standard features; expanded (≥5 mutations or full mutation tooling) for P0/critical paths. Surviving mutants become fix tasks.
103
+ 3. Applies the **payload/conjunction rule**: checks payload fields are asserted on value/state, not just that the call occurred.
104
+ 4. **Writes the persisted report** to `.specs/features/[feature]/validation.md` - PASS/FAIL, per-AC evidence (`file:line` + assertion + spec outcome), sensor result (killed/survived per mutation), gate exit results, diff/commit range.
105
+ 5. **Returns a compact verdict in chat** to the orchestrator.
106
+ 6. Does **NOT** write, modify, or fix any code or tests - the real working tree is never mutated (sensor mutations run in scratch state only).
107
+
108
+ **What the Verifier reports back (compact chat format):**
109
+ ```
110
+ ## Validation: [feature name] - [PASS ✅ | FAIL ❌]
111
+
112
+ **Spec-anchored check**: [N/N ACs matched spec outcome | M spec-precision gaps flagged]
113
+ **Gate**: [X passed, 0 failed]
114
+ **Sensor**: [N mutations injected, N killed, N survived]
115
+ **Report**: `.specs/features/[feature]/validation.md`
116
+
117
+ **Ranked gaps** (if FAIL):
118
+ 1. [Gap description] - [AC or criterion] - [file:line or "no evidence"]
119
+ 2. ...
120
+ ```
121
+
122
+ **Failure handling:** The orchestrator routes the ranked gaps to an implementer as fix tasks, then re-dispatches the Verifier. This fix→re-verify loop is bounded to a maximum of **3 iterations**. If gaps remain after 3 iterations, escalate to the user.
123
+
124
+ **Standalone fallback:** When running without sub-agents (a single agent executing the full feature), run `validate.md` as an independent fresh-eyes pass - re-read `spec.md` and the diff from scratch, apply evidence-or-zero, run the spec-anchored check and discrimination sensor, write the report file, then run `python3 <skill-dir>/scripts/validate_state.py <feature>` to confirm the report is a real PASS, and report PASS/FAIL before marking the feature done.
125
+
126
+ ---
127
+
128
+ ## Model Tier per Role
129
+
130
+ **Applies only if the harness can assign a model per sub-agent.** If it cannot, ignore this section and run everything on the default model - the workflow is correct either way. The point is to spend high-reasoning capacity where ambiguity and consequence are high, and a faster tier where the work is mechanical, instead of paying top-tier cost uniformly.
131
+
132
+ Judge the tier by the work in front of the role, not by the role's title:
133
+
134
+ | Role / work | Characteristic | Suggested tier |
135
+ | ----------- | -------------- | -------------- |
136
+ | Design phase | High ambiguity, hard-to-reverse structural decisions | High-reasoning |
137
+ | Batch worker - core-domain or high-ambiguity phase | Non-obvious logic, tricky edge cases, novel integration | High-reasoning |
138
+ | Batch worker - mechanical phase | Entities, DTOs, config, wiring, straightforward CRUD against a settled pattern | Faster / cheaper |
139
+ | Verifier | Adversarial reasoning: designs mutations, re-derives coverage, judges outcome precision | Mid-to-high |
140
+ | Specify / Tasks authoring | Structured but judgment-heavy | Mid-to-high |
141
+
142
+ **Rules of thumb:**
143
+
144
+ - When unsure, size up, not down. An under-powered worker on ambiguous logic produces gaps the Verifier then has to catch - more expensive than paying for reasoning once.
145
+ - The Verifier is never the cheapest tier; a weak Verifier defeats the author ≠ verifier gate.
146
+ - Set the tier per batch, from that batch's phases. A feature can mix tiers across batches.
147
+ - This is advisory metadata only. No gate, commit, or verification step depends on it.