@mohammadhprp/system-prompt 0.11.0 → 0.11.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/framework/agents/backend-architect.md +1 -1
  2. package/framework/commands/commit.md +0 -3
  3. package/framework/mcps/figma-mcp-go/README.md +0 -1
  4. package/framework/mcps/gitlab-mcp/README.md +0 -1
  5. package/framework/mcps/jira-mcp/README.md +0 -1
  6. package/framework/mcps/laravel-boost/README.md +0 -1
  7. package/framework/mcps/notion-mcp/README.md +0 -1
  8. package/framework/mcps/supabase-mcp/README.md +0 -1
  9. package/framework/plugins/opencode-goal-plugin/README.md +0 -1
  10. package/framework/references/standards/api.md +0 -1
  11. package/framework/references/standards/architecture.md +0 -1
  12. package/framework/references/standards/database.md +0 -1
  13. package/framework/references/standards/debugging.md +0 -1
  14. package/framework/references/standards/documentation.md +0 -2
  15. package/framework/references/standards/logging.md +0 -1
  16. package/framework/references/standards/naming.md +0 -1
  17. package/framework/references/standards/observability.md +0 -1
  18. package/framework/references/standards/performance.md +0 -1
  19. package/framework/references/standards/pull-requests.md +0 -1
  20. package/framework/references/standards/security.md +0 -1
  21. package/framework/references/standards/testing.md +0 -1
  22. package/framework/skills/README.md +15 -3
  23. package/framework/skills/codenavi/SKILL.md +306 -0
  24. package/framework/skills/codenavi/examples.md +33 -0
  25. package/framework/skills/codenavi/references/coding-principles.md +143 -0
  26. package/framework/skills/codenavi/references/notebook-spec.md +171 -0
  27. package/framework/skills/create-adr/SKILL.md +429 -0
  28. package/framework/skills/create-adr/examples.md +35 -0
  29. package/framework/skills/docs-writer/SKILL.md +39 -0
  30. package/framework/skills/docs-writer/examples.md +34 -0
  31. package/framework/skills/docs-writer/references/style-guide.md +72 -0
  32. package/framework/skills/frontend-design/SKILL.md +55 -0
  33. package/framework/skills/frontend-design/examples.md +45 -0
  34. package/framework/skills/humanizer/SKILL.md +412 -0
  35. package/framework/skills/humanizer/examples.md +46 -0
  36. package/framework/skills/learning-opportunities/SKILL.md +140 -0
  37. package/framework/skills/learning-opportunities/examples.md +34 -0
  38. package/framework/skills/learning-opportunities/references/PRINCIPLES.md +42 -0
  39. package/framework/skills/perf-web-optimization/SKILL.md +163 -0
  40. package/framework/skills/perf-web-optimization/examples.md +35 -0
  41. package/framework/skills/perf-web-optimization/references/bundle-optimization.md +180 -0
  42. package/framework/skills/perf-web-optimization/references/core-web-vitals.md +154 -0
  43. package/framework/skills/perf-web-optimization/references/image-optimization.md +170 -0
  44. package/framework/skills/security-best-practices/LICENSE.txt +201 -0
  45. package/framework/skills/security-best-practices/SKILL.md +89 -0
  46. package/framework/skills/security-best-practices/examples.md +35 -0
  47. package/framework/skills/security-best-practices/references/golang-general-backend-security.md +988 -0
  48. package/framework/skills/security-best-practices/references/javascript-express-web-server-security.md +1151 -0
  49. package/framework/skills/security-best-practices/references/javascript-general-web-frontend-security.md +725 -0
  50. package/framework/skills/security-best-practices/references/javascript-jquery-web-frontend-security.md +672 -0
  51. package/framework/skills/security-best-practices/references/javascript-typescript-nextjs-web-server-security.md +1138 -0
  52. package/framework/skills/security-best-practices/references/javascript-typescript-react-web-frontend-security.md +975 -0
  53. package/framework/skills/security-best-practices/references/javascript-typescript-vue-web-frontend-security.md +789 -0
  54. package/framework/skills/security-best-practices/references/python-django-web-server-security.md +880 -0
  55. package/framework/skills/security-best-practices/references/python-fastapi-web-server-security.md +1030 -0
  56. package/framework/skills/security-best-practices/references/python-flask-web-server-security.md +835 -0
  57. package/framework/skills/sentry/SKILL.md +127 -0
  58. package/framework/skills/sentry/examples.md +34 -0
  59. package/framework/skills/sentry/scripts/sentry_api.py +238 -0
  60. package/framework/skills/show-me/SKILL.md +127 -0
  61. package/framework/skills/show-me/examples.md +78 -0
  62. package/framework/skills/spec-driven-eval/SKILL.md +341 -0
  63. package/framework/skills/spec-driven-eval/examples.md +35 -0
  64. package/framework/skills/spec-driven-eval/references/quickstart.md +118 -0
  65. package/framework/skills/spec-driven-eval/references/reference.md +295 -0
  66. package/framework/skills/technical-design-doc-creator/README.md +411 -0
  67. package/framework/skills/technical-design-doc-creator/SKILL.md +1484 -0
  68. package/framework/skills/technical-design-doc-creator/examples.md +35 -0
  69. package/framework/skills/tlc-spec-driven/SKILL.md +184 -0
  70. package/framework/skills/tlc-spec-driven/examples.md +34 -0
  71. package/framework/skills/tlc-spec-driven/references/code-analysis.md +98 -0
  72. package/framework/skills/tlc-spec-driven/references/coding-principles.md +72 -0
  73. package/framework/skills/tlc-spec-driven/references/context-limits.md +31 -0
  74. package/framework/skills/tlc-spec-driven/references/design.md +199 -0
  75. package/framework/skills/tlc-spec-driven/references/discuss.md +159 -0
  76. package/framework/skills/tlc-spec-driven/references/implement.md +436 -0
  77. package/framework/skills/tlc-spec-driven/references/lessons.md +115 -0
  78. package/framework/skills/tlc-spec-driven/references/memory.md +144 -0
  79. package/framework/skills/tlc-spec-driven/references/specify.md +228 -0
  80. package/framework/skills/tlc-spec-driven/references/sub-agents.md +147 -0
  81. package/framework/skills/tlc-spec-driven/references/tasks.md +451 -0
  82. package/framework/skills/tlc-spec-driven/references/validate.md +355 -0
  83. package/framework/skills/tlc-spec-driven/scripts/check_commit.py +115 -0
  84. package/framework/skills/tlc-spec-driven/scripts/lessons.py +412 -0
  85. package/framework/skills/tlc-spec-driven/scripts/validate_spec.py +260 -0
  86. package/framework/skills/tlc-spec-driven/scripts/validate_state.py +162 -0
  87. package/framework/skills/tlc-spec-driven/scripts/validate_tasks.py +251 -0
  88. package/framework/skills/web-design-guidelines/SKILL.md +65 -0
  89. package/framework/skills/web-design-guidelines/examples.md +32 -0
  90. package/framework/skills/web-design-guidelines/references/guideline.md +174 -0
  91. package/package.json +1 -1
  92. package/src/catalog.js +15 -3
  93. package/src/installer.js +66 -1
  94. package/framework/skills/backend-engineer/SKILL.md +0 -76
  95. package/framework/skills/backend-engineer/examples.md +0 -31
  96. package/framework/skills/documentation/SKILL.md +0 -74
  97. package/framework/skills/documentation/examples.md +0 -31
@@ -0,0 +1,159 @@
1
+ # Specify: Discuss Gray Areas
2
+
3
+ **Goal:** Capture HOW the user envisions the feature when the spec has ambiguous areas. This is NOT a separate phase - it's triggered within Specify when the agent detects gray areas that need user input.
4
+
5
+ **Trigger:** Automatically when gray areas are detected during spec creation, or explicitly via "discuss feature", "how should this work?", "capture context"
6
+
7
+ **When to trigger (auto-detect):** The spec contains user-facing behavior that could go multiple ways AND the user hasn't expressed a preference. If the spec is clear and unambiguous, skip this entirely.
8
+
9
+ **When NOT to trigger:** Genuinely trivial features - a pure read endpoint, a config tweak, features with no [implicit-requirement dimensions](specify.md#implicit-requirement-dimensions) present (no persistence/state, external calls, auth, payments, concurrency, or state transitions). When any dimension is present, trigger discuss.
10
+
11
+ ## Why This Phase Exists
12
+
13
+ Specifications capture WHAT to build. Design captures the architecture. But neither captures the user's vision for ambiguous areas - layout preferences, interaction patterns, error handling style, content tone. Without this, the agent guesses. With this, the agent builds what the user actually imagined.
14
+
15
+ The output - `context.md` - feeds directly into Design and Tasks:
16
+
17
+ - **Design reads it** to know what decisions are locked vs. flexible
18
+ - **Tasks reads it** to include specific behaviors in task definitions
19
+
20
+ ## Process
21
+
22
+ ### 1. Analyze the Feature
23
+
24
+ Read `.specs/features/[feature]/spec.md` and identify the domain:
25
+
26
+ | Domain | Gray areas to explore |
27
+ | ------------------------------ | ------------------------------------------------------------- |
28
+ | Something users **SEE** | Layout, density, interactions, empty states, visual hierarchy |
29
+ | Something users **CALL** (API) | Response format, errors, auth, versioning, rate limiting |
30
+ | Something users **RUN** (CLI) | Output format, flags, modes, error handling, verbosity |
31
+ | Something users **READ** | Structure, tone, depth, flow, navigation |
32
+ | Something being **ORGANIZED** | Grouping criteria, naming, duplicates, exceptions |
33
+ | Something with **backend / state / contract** | Failure & partial-failure states, idempotency/retry/dedup, auth boundaries & rate limits, data lifecycle/expiry, concurrency/ordering - see [implicit-requirement dimensions](specify.md#implicit-requirement-dimensions) |
34
+
35
+ Generate 3-4 **feature-specific** gray areas. Not generic categories, but concrete decisions for THIS feature.
36
+
37
+ ### 2. Present Gray Areas
38
+
39
+ Present the feature boundary (from spec.md) and the gray areas to the user. Let them choose which to discuss. Do NOT include a "skip all" option - the user invoked this phase to discuss.
40
+
41
+ Any gray area the user **declines** to discuss, or that goes undiscussed, is written to the spec's **Assumptions & Open Questions** section (agent's chosen default + rationale) - never silently dropped. This ensures the spec's closure gate can pass: every gray area is either resolved through discussion or recorded as a signed-off assumption.
42
+
43
+ ### 3. Choose discussion pace (once)
44
+
45
+ Before deep-diving, ask **one** pace question. Recommend **Guided** as the default. If the user skips, says "whatever", or "you choose", use Guided.
46
+
47
+ | Pace | When it fits | Cadence |
48
+ | ------------ | ------------------------------------------------- | ----------------------------------------------------------------------- |
49
+ | **Quick** | User wants speed; trusts defaults | Propose defaults per area (rationale included); user accepts / overrides |
50
+ | **Guided** | Default - balances depth and turn count | Adaptive elicitation (see below) |
51
+ | **Detailed** | High ambiguity; user wants Socratic control | Exactly one decision per turn, dependency order |
52
+
53
+ Honor mid-discussion switches immediately ("go faster", "slow down", "just decide") - change pace without restarting or re-asking settled decisions.
54
+
55
+ ### 4. Deep-Dive Each Area
56
+
57
+ Shared rules for every pace:
58
+
59
+ 1. Options must be concrete ("Card layout" or "Table layout" - not "Option A" or "how should it look?").
60
+ 2. Lead with your recommended answer and one line of reasoning. You have read the codebase; the user should be able to accept or override in a word.
61
+ 3. Offer "You decide" when reasonable - it records agent discretion explicitly.
62
+ 4. Resolve anything discoverable from the code yourself (Knowledge Verification Chain); only put genuine product decisions to the user.
63
+ 5. When an area is settled: "More on [area], or move on?" After all areas: "Ready to create context?"
64
+
65
+ **Quick:** For each selected gray area, present the recommended decisions for that area in one turn (defaults + short rationale). Wait for accept / override. Do not drip-feed single questions unless the user challenges a default and opens a real fork.
66
+
67
+ **Guided:** Adaptive elicitation - questions are a decision tree to prune, not a checklist to finish.
68
+
69
+ 1. Classify upcoming decisions as **independent** vs **dependent**.
70
+ 2. Low-stakes / safe-to-default → state the assumption and invite correction (no blocking question).
71
+ 3. Independent product decisions → ask **at most 2** in the same turn, each with options + recommended default.
72
+ 4. Dependent decisions → ask **exactly one**, wait, then continue (the earlier answer should prune later questions).
73
+ 5. Never dump 3+ questions in one turn. Never ask what the code already answers.
74
+ 6. Stop the area as soon as enough is decided.
75
+
76
+ **Detailed:** Walk selected gray areas as a strict decision tree - one concrete question per turn, dependency order, wait for each answer before the next. Use when the user wants maximum control or the feature is highly ambiguous.
77
+
78
+ ### 5. Scope Guardrail (CRITICAL)
79
+
80
+ The feature boundary from spec.md is **fixed**. Discussion clarifies HOW to implement, never WHETHER to add new capabilities.
81
+
82
+ **Allowed:** "How should posts be displayed?" (clarifying ambiguity)
83
+ **Not allowed:** "Should we also add comments?" (new capability)
84
+
85
+ When user suggests scope creep: "That sounds like a separate feature. I'll note it in Deferred Ideas. Back to [current area]."
86
+
87
+ ### 6. Write context.md
88
+
89
+ ---
90
+
91
+ ## Template: `.specs/features/[feature]/context.md`
92
+
93
+ ```markdown
94
+ # [Feature] Context
95
+
96
+ **Gathered:** [date]
97
+ **Spec:** `.specs/features/[feature]/spec.md`
98
+ **Status:** Ready for design
99
+
100
+ ---
101
+
102
+ ## Feature Boundary
103
+
104
+ [Clear statement of what this feature delivers - the scope anchor from spec.md]
105
+
106
+ ---
107
+
108
+ ## Implementation Decisions
109
+
110
+ ### [Area 1 that was discussed]
111
+
112
+ - [Specific decision made]
113
+ - [Another decision if applicable]
114
+
115
+ ### [Area 2 that was discussed]
116
+
117
+ - [Specific decision made]
118
+
119
+ ### [Area 3 that was discussed]
120
+
121
+ - [Specific decision made]
122
+
123
+ ### Agent's Discretion
124
+
125
+ [Areas where user explicitly said "you decide" - agent has flexibility here during design/implementation]
126
+
127
+ ### Declined / Undiscussed Gray Areas → Assumptions
128
+
129
+ [Gray areas the user declined to discuss or that were not covered. Each entry is written to the spec's Assumptions & Open Questions section with the agent's chosen default and rationale - not left silently unresolved.]
130
+
131
+ ---
132
+
133
+ ## Specific References
134
+
135
+ [Any "I want it like X" moments, product references, specific behaviors, interaction patterns mentioned during discussion]
136
+
137
+ [If none: "No specific requirements - open to standard approaches"]
138
+
139
+ ---
140
+
141
+ ## Deferred Ideas
142
+
143
+ [Ideas that came up during discussion but belong in other features/phases. Captured here so they're not lost, but explicitly out of scope]
144
+
145
+ [If none: "None - discussion stayed within feature scope"]
146
+ ```
147
+
148
+ ---
149
+
150
+ ## Tips
151
+
152
+ - **Pace is a user choice; Guided is the default** - Quick for speed, Guided for balance, Detailed for Socratic depth; honor mid-discussion switches
153
+ - **Guided ≠ interrogation and ≠ form dump** - Assume-first when safe, ≤2 independent questions per turn, one-at-a-time only when answers depend on each other
154
+ - **Look it up, don't ask** - Resolve anything discoverable from the code yourself; ask only genuine product decisions
155
+ - **Decisions, not vision** - "Card-based layout with subtle shadows" is a decision. "Should feel modern" is not.
156
+ - **Scope is sacred** - Deferred Ideas captures scope creep without losing ideas
157
+ - **User = visionary, Agent = builder** - Ask about how they imagine it, not about technical implementation
158
+ - **Don't ask about:** Technical architecture, performance, implementation details - that's Design's job
159
+ - **Confirm before Design** - User approves context.md before moving to design phase
@@ -0,0 +1,436 @@
1
+ # Execute
2
+
3
+ **Goal**: Implement ONE task at a time. Surgical changes. Verify. Commit. Repeat.
4
+
5
+ This is where code gets written. Every task follows the same cycle: plan → implement → verify → commit. Verification is built into every task, not a separate phase.
6
+
7
+ ---
8
+
9
+ ## MANDATORY: Before Starting Any Implementation
10
+
11
+ **Read [coding-principles.md](coding-principles.md) and state:**
12
+
13
+ 1. **Assumptions** - What am I assuming? Any uncertainty?
14
+ 2. **Files to touch** - List ONLY files this task requires
15
+ 3. **Success criteria** - How will I verify this works?
16
+
17
+ ⚠️ **Do not proceed without stating these explicitly.**
18
+
19
+ ---
20
+
21
+ ## Process
22
+
23
+ **Batch worker context:** When this task is executed as part of a phase-batch sub-agent, the worker
24
+ receives the task definitions for every phase in its batch, coding principles, the generated Test
25
+ Coverage Matrix and Gate Check Commands from tasks.md, and relevant spec/design context. A batch is
26
+ one or more consecutive whole phases packed to ~7 tasks. The worker executes ALL tasks in its
27
+ assigned batch in order - finishing every task in one phase before starting the next phase in the
28
+ batch - and each task follows every step below (implement → gate → atomic commit) before moving to
29
+ the next. After all tasks in the batch are complete, the worker reports a compact summary (tasks
30
+ done, commit hashes, test counts, deviations/blockers) to the orchestrator. See
31
+ [sub-agents.md](sub-agents.md) for the full model.
32
+
33
+ ### Before implementing: assess sub-agent delegation (MANDATORY - before the first task)
34
+
35
+ Before implementing anything, if a formal `tasks.md` with an Execution Plan exists, **count its total tasks** and pack the phases into task-budgeted batches (~7 tasks per worker, whole phases - see [sub-agents.md](sub-agents.md)). If that yields **more than one batch** (> ~8 tasks), you MUST present the sub-agent offer to the user and wait for their choice before starting Execute - do not silently proceed inline. If the feature fits a single batch (≤ ~8 tasks, or the user declines), execute inline. Skip this check only when you are already a batch worker executing a delegated batch (the orchestrator already made the delegation decision).
36
+
37
+ ### 0. List Atomic Steps (MANDATORY when Tasks phase was skipped)
38
+
39
+ If there is no `tasks.md` for this feature, you MUST list atomic steps before writing any code. This is non-negotiable - it prevents the agent from losing focus and doing too many things at once.
40
+
41
+ ```
42
+ ## Execution Plan
43
+
44
+ 1. [Step] → files: [list] → verify: [how] → commit: [message]
45
+ 2. [Step] → files: [list] → verify: [how] → commit: [message]
46
+ 3. [Step] → files: [list] → verify: [how] → commit: [message]
47
+ ```
48
+
49
+ **Each step must be:**
50
+
51
+ - ONE deliverable (one component, one function, one endpoint, one file change)
52
+ - Independently verifiable (can prove it works before moving on)
53
+ - Independently committable (gets its own atomic git commit)
54
+
55
+ If listing steps reveals >5 steps or complex dependencies, STOP and create a formal `tasks.md` instead. The Tasks phase was wrongly skipped.
56
+
57
+ ### 1. Pick Task
58
+
59
+ From tasks.md (if exists) or from the execution plan above. User specifies ("implement T3") or suggest next available.
60
+
61
+ ### 2. Verify Dependencies
62
+
63
+ If tasks.md exists, check dependencies. If using inline plan, follow the order listed.
64
+
65
+ ❌ If blocked: "T3 depends on T2 which isn't done. Should I do T2 first?"
66
+
67
+ ### 3. State Implementation Plan
68
+
69
+ Before writing code:
70
+
71
+ ```
72
+ Files: [list]
73
+ Approach: [brief description]
74
+ Success: [how to verify]
75
+ ```
76
+
77
+ ### 4. Write Tests (derived from spec, not from implementation)
78
+
79
+ If the task includes tests (per the Tests field and **Test Coverage Matrix** in tasks.md):
80
+
81
+ 1. Write the test file(s) covering the task's acceptance criteria.
82
+ 2. Tests MUST be derived from the task's "Done when" criteria and `spec.md` ACs - **not** from the implementation. Each test encodes what the spec requires; never write tests by reading the code and asserting what it currently does.
83
+ 3. Each acceptance criterion from "Done when" maps to at least one test assertion whose asserted value matches the **spec-defined expected outcome**. Where the spec does not define a precise outcome, note it as a **spec-precision gap** rather than writing a vague assertion and passing silently.
84
+ 4. Edge cases from spec.md that apply to this task get test cases too.
85
+
86
+ **HARD CONSTRAINTS (test integrity - never violate):**
87
+
88
+ - Do NOT weaken assertions (making them less specific to pass more easily)
89
+ - Do NOT delete or skip test cases
90
+ - Do NOT use the test framework's skip/disable/pending mechanism to bypass failing tests
91
+
92
+ If a test is genuinely wrong (tests the wrong behavior per spec), STOP and ask the user
93
+ before modifying it. Never silently change a test.
94
+
95
+ If the task does NOT include tests (e.g., entity-only, config-only), skip to Step 4b.
96
+
97
+ ### 4b. Implement
98
+
99
+ Write the minimum implementation needed to satisfy the task's success criteria: pass all relevant tests (when present) and meet the defined verification/gate checks when there are no direct tests.
100
+
101
+ **HARD CONSTRAINTS:**
102
+
103
+ - The test-integrity rules from step 4 still hold: do NOT weaken, delete, or skip/disable tests. The tests are the spec - implementation conforms to them, not the reverse.
104
+ - Modify a test only to fix a genuinely wrong assertion, and ask the user first.
105
+ - Minimum code to pass - save structural improvements for a refactor task
106
+
107
+ Follow [coding-principles.md](coding-principles.md):
108
+
109
+ - Simplest code that works
110
+ - Touch ONLY listed files
111
+ - No scope creep
112
+
113
+ ### 5. Gate Check (VERIFY)
114
+
115
+ Run the gate check command from the task definition. This is MANDATORY - not "if applicable."
116
+
117
+ 1. Look up the command for the task's Gate level (quick/full/build) in the **Gate Check Commands** section of tasks.md, then run it
118
+ 2. Non-zero exit code = STOP. Fix the failure. Re-run. Do not proceed until it passes.
119
+ 3. Confirm the test count matches expectations (no tests were silently deleted or skipped)
120
+
121
+ **Tiered gates (from the Gate Check Commands section of tasks.md):**
122
+
123
+ | Task includes | Gate level | What runs |
124
+ | -------------------------------- | ---------- | ------------------------ |
125
+ | Unit tests only | Quick | Unit test command |
126
+ | E2E or integration tests | Full | Unit + E2E commands |
127
+ | Last task in a phase | Build | Build + lint + all tests |
128
+ | No tests (config, entities, etc) | Build | Build + lint only |
129
+
130
+ The gate check is deterministic. The test runner decides if the code is correct,
131
+ not the agent's self-assessment.
132
+
133
+ ### 6. Post-Gate Review
134
+
135
+ After the gate check passes:
136
+
137
+ 1. Verify test count: Are there at least as many test cases as before? (prevents silent deletion)
138
+ 2. Verify no SPEC_DEVIATION: If implementation diverged from spec/design, add a marker:
139
+
140
+ ```
141
+ // SPEC_DEVIATION: [what diverged]
142
+ // Reason: [why the deviation was necessary]
143
+ ```
144
+
145
+ 3. Quick complexity check: "Would senior engineer flag this as overcomplicated?"
146
+ - Yes → Simplify, re-run gate
147
+ - No → Proceed
148
+
149
+ 4. **Test Adequacy Review (MANDATORY - hard gate).**
150
+
151
+ A task cannot be committed or marked done until all four checks below pass. Tests must be both **necessary** (every test traces to a requirement) and **sufficient** (every requirement is covered). The scope boundary is the feature spec - do not test beyond it.
152
+
153
+ **Check A - Sufficient coverage (per-layer depth).** Build and output this table:
154
+
155
+ | Done-when criterion / spec AC / listed edge case | `file:line` + assertion expression | Spec-defined outcome | Covered? |
156
+ | ------------------------------------------------- | ---------------------------------- | -------------------- | -------- |
157
+ | [criterion from task or spec] | `path/to/test.ts:42` - `expect(result.field).toBe(expected)` | [expected value from spec] | ✅ Yes / ❌ No / ⚠️ Spec-precision gap |
158
+
159
+ **Evidence-or-zero rule:** Each covered cell MUST cite the exact `file:line` where the assertion lives AND reproduce the assertion expression (not just the `describe`/`it` name). A criterion with no located `file:line` evidence counts as **NOT covered**; the task cannot be marked done. Do not declare a criterion absent without first searching the test files - show the search before concluding it is missing (mirror: evidence or zero, never a guess).
160
+
161
+ **Spec-anchored outcome check:** For each covered criterion, derive the expected outcome from `spec.md` (or the task's "Done when" field) and confirm the test's asserted value matches it - not just that an assertion exists. Where the spec defines a precise outcome (e.g., a specific status code, a specific field value, a specific error message), the test assertion MUST target that exact outcome. Where the spec does not define a precise outcome, mark the cell as **⚠️ Spec-precision gap** and add a note; do NOT silently pass a vague assertion as if it were covered.
162
+
163
+ Every "Done when" criterion, every spec.md acceptance criterion, and every listed edge case that applies to this task must map to at least one concrete test assertion. Enforce the layer's Coverage Expectation from the Test Coverage Matrix:
164
+
165
+ - Domain / service layer: assertions map 1:1 to spec ACs; every listed edge case has a dedicated test.
166
+ - Route / controller / e2e layer: every route the task adds or modifies must have a happy-path test, a test for each listed edge case, and a test for each documented error/failure path.
167
+
168
+ No criterion left unverified.
169
+
170
+ **Check B - Non-shallow litmus.** Reject each of the following shallow patterns:
171
+ - Assertion-free tests or `expect(true)` / `expect(1).toBe(1)` style tautologies
172
+ - "No error thrown" as the only assertion - unless not-throwing IS the specified behavior
173
+ - Asserting only on mock call counts when the actual output/state is what the criterion demands
174
+ - Happy-path only when the task's "Done when" or spec.md lists edge cases
175
+
176
+ **Payload/conjunction rule.** For each named field in an emitted event, returned object, or persisted record, apply a separate check:
177
+ 1. Open the constructed object at its `file:line` and confirm the field is present in the assertion.
178
+ 2. Confirm the assertion targets the field's **value or state**, not just the call that produced it.
179
+ 3. A present `emit(...)` / `return ...` / `save(...)` call does NOT prove the field - only an assertion on the result does.
180
+ 4. Asserting a method was called (spy/mock) != asserting the resulting state. Both may be needed; neither substitutes for the other.
181
+
182
+ Apply this check to every payload-bearing criterion before marking it covered.
183
+
184
+ **Stack-agnostic litmus:** An assertion is shallow if it would still pass under a plausible *wrong* implementation. If so, strengthen it before committing.
185
+
186
+ **Check C - Necessary (no tests beyond the spec).** Reverse-map every test back to a spec AC, a listed edge case, or a "Done when" criterion. Build this table:
187
+
188
+ | `file:line` + assertion expression | Maps to (AC / edge case / Done-when criterion) | Keep? |
189
+ | ---------------------------------- | ---------------------------------------------- | ----- |
190
+ | `path/to/test.ts:42` - `expect(result.field).toBe(expected)` | [requirement ID or criterion text] | ✅ Keep / ❌ Remove |
191
+
192
+ Any test that maps to nothing → remove it. A test with no requirement is scope creep - it proves nothing about the feature and expands scope beyond the spec. Do not write speculative "what if" tests, do not test framework or library behavior, and do not duplicate an assertion that is already covered at another layer for the same scenario.
193
+
194
+ **Check D - Guideline conformance.** If project quality/testing guidelines were found in step 0 of tasks.md step 1.5, verify this task's tests conform to them (naming conventions, file locations, coverage thresholds, etc.). Note the guideline file followed.
195
+
196
+ **Bound:** Tests prove the work; they do not expand it. Thoroughness is scoped to the feature + spec. Repo depth is a floor (never less thorough than existing tests for the same layer); the spec is the ceiling. Do not invent requirements or tests that have no spec anchor.
197
+
198
+ **Anti-patterns - known verification cheats (treat any of these as an automatic Check failure):**
199
+
200
+ | Anti-pattern | Why it fails |
201
+ | ------------ | ------------ |
202
+ | Committing before the gate check passes | Skips the deterministic verifier - the gate is not optional |
203
+ | Asserting call count / spy invocation instead of the resulting state | Proves the method ran, not that it did the right thing |
204
+ | Marking a criterion covered without a `file:line` citation | Violates evidence-or-zero; suspicion of coverage is not coverage |
205
+ | Weakening an assertion (making it less specific) to force a pass | Moves the goalposts instead of fixing the code |
206
+ | Deleting or skipping a test to make the suite pass | Destroys coverage permanently; a failing test is a signal, not noise |
207
+ | "Tested elsewhere" deferral without citing where | Coverage gaps hide behind vague claims; cite the file:line or it doesn't count |
208
+ | Speculative "what if" tests with no spec anchor | Expands scope beyond the ceiling; remove them in Check C |
209
+ | Testing framework or library behavior | Tests a dependency, not the feature; remove them in Check C |
210
+
211
+ **On any failure** → rewrite or remove the affected test(s), re-run the gate, then re-run this review.
212
+
213
+ *Honest caveat:* This is an inspection-based review (model judgment), complementary to - not a replacement for - the deterministic gate. The gate confirms the test suite runs; the feature-level discrimination sensor (step 9) confirms the tests can detect regressions. This review confirms the suite is meaningful and bounded.
214
+
215
+ Add the two mapping tables and a one-line adequacy verdict to the Execution Template's Post-Gate section.
216
+
217
+ ### 7. Status + Atomic Commit (same commit)
218
+
219
+ After the gate is green, close the task record **before** creating the commit, then commit code and status together. Never leave `tasks.md` still open after a successful task commit - a crash between those steps is how resume redoes finished work.
220
+
221
+ 1. Mark the task complete in `tasks.md`. Update requirement traceability in `spec.md` if requirement IDs are used.
222
+ 2. Create **one** atomic commit that includes the implementation, its tests, and those status/traceability updates.
223
+
224
+ Each task gets its own commit immediately after verification. Never batch multiple tasks into one commit.
225
+
226
+ **Format ([Conventional Commits 1.0.0](https://www.conventionalcommits.org/en/v1.0.0/)):**
227
+
228
+ ```
229
+ <type>(<scope>): <description>
230
+
231
+ [optional body]
232
+
233
+ [optional footer(s)]
234
+ ```
235
+
236
+ **Types:**
237
+
238
+ | Type | When to use |
239
+ | ---------- | ------------------------------------------------------- |
240
+ | `feat` | New feature or capability |
241
+ | `fix` | Bug fix |
242
+ | `refactor` | Code change that neither fixes a bug nor adds a feature |
243
+ | `docs` | Documentation only |
244
+ | `test` | Adding or correcting tests |
245
+ | `style` | Formatting, missing semicolons, etc. (no code change) |
246
+ | `perf` | Performance improvement |
247
+ | `build` | Build system or external dependencies |
248
+ | `ci` | CI configuration files and scripts |
249
+ | `chore` | Maintenance tasks that don't modify src or test files |
250
+
251
+ **Scope:** Feature name or module area, lowercase, e.g., `auth`, `cart`, `api`
252
+
253
+ **Description rules:**
254
+
255
+ - Imperative mood ("add", not "added" or "adds")
256
+ - Lowercase first letter
257
+ - No period at the end
258
+ - Complete the sentence: "If applied, this commit will _[your description]_"
259
+
260
+ **Breaking changes:** Append `!` after type/scope AND add `BREAKING CHANGE:` footer:
261
+
262
+ ```
263
+ feat(api)!: change authentication endpoint response format
264
+
265
+ BREAKING CHANGE: login endpoint now returns JWT in body instead of cookie
266
+ ```
267
+
268
+ **Examples:**
269
+
270
+ ```
271
+ feat(auth): add email validation to login form
272
+ ```
273
+
274
+ ```
275
+ fix(cart): prevent negative quantity on item decrement
276
+ ```
277
+
278
+ ```
279
+ refactor(api): extract token refresh logic into service
280
+
281
+ Move token refresh from inline handler to dedicated AuthTokenService
282
+ for reuse across multiple endpoints.
283
+ ```
284
+
285
+ **Rules:**
286
+
287
+ - One task = one commit
288
+ - Description references what was DONE, not what was planned
289
+ - Include only files listed in the task - plus the `tasks.md` / `spec.md` status updates for this task
290
+ - Never sneak in "while I'm here" changes
291
+ - If tests are part of the task, include them in the same commit
292
+
293
+ **Deterministic check.** Validate the message before committing: `python3 <skill-dir>/scripts/check_commit.py --message "<your message>"`. A non-zero exit means fix the format first. This makes the format rule enforceable instead of memory-dependent.
294
+
295
+ **Optional git-level guard (git only, no agent dependency).** In a git repo the same check can run on every commit by wiring it as a `commit-msg` hook, so a malformed message is rejected regardless of who or what drives the commit:
296
+
297
+ ```bash
298
+ # from the repo root, one time (resolve <skill-dir> to the directory that contains this skill's SKILL.md):
299
+ ln -sf <skill-dir>/scripts/check_commit.py .git/hooks/commit-msg && chmod +x .git/hooks/commit-msg
300
+ ```
301
+
302
+ This is a plain git hook, not tied to any editor or assistant. Skip it if the project manages hooks its own way (for example a pre-commit framework); the manual check above still applies.
303
+
304
+ ### 8. Scope Guardrail
305
+
306
+ During implementation, you will notice things that could be improved, refactored, or added. **Do not act on them.** Instead:
307
+
308
+ - If it's a bug: surface it to the user (or capture it as a separate task)
309
+ - If it's an improvement: add it to the feature's `context.md` under "Deferred Ideas" (or surface it to the user if there is no `context.md`)
310
+ - If it's related to the current task: only include it if it's in the "Done when" criteria
311
+
312
+ **The heuristic:** "Is this in my task definition?" If no, don't touch it.
313
+
314
+ **Blast radius (approval ≠ remote authority):** Approving a spec or tasks authorizes local implementation and local commits only. Before `git push`, force-push, deploy, production DB migration, or any other remote / externally visible / destructive operation, STOP and get an explicit go-ahead for that action - even if Execute was already approved.
315
+
316
+ ### 9. Feature-Level Validation (after the LAST task - MANDATORY, always runs)
317
+
318
+ When the task you just completed is the **last task of the feature** (or of a priority group being delivered on its own, e.g. all P1 tasks), you MUST run feature-level validation before reporting the work as done. **This is not optional and is never prompted - it runs automatically.** Do not stop at the final task's commit.
319
+
320
+ **Author ≠ verifier.** An author checking their own work reapplies the mental model that may have produced the gaps. The Verifier is a fresh sub-agent that re-derives coverage from the spec independently - this separation is the quality gate, not a style preference.
321
+
322
+ **Layering:**
323
+ - Per-task adequacy self-check (steps 5-6): cheap, always runs, author does it, confirms each task in isolation.
324
+ - Feature-level validation (step 9): one trustworthy independent gate at completion, always-on, Verifier sub-agent does it.
325
+
326
+ **How to delegate to the Verifier:**
327
+ Dispatch a fresh sub-agent following the **Verifier** role described in [sub-agents.md](sub-agents.md). Provide it with:
328
+ - `spec.md` (ACs = source of truth)
329
+ - The git diff surface for this feature (commit range)
330
+ - The test files in scope
331
+ - `validate.md` as its operating checklist
332
+
333
+ **What the Verifier does** (full procedure in [sub-agents.md](sub-agents.md); operating checklist in [validate.md](validate.md)): a spec-anchored coverage check (evidence-or-zero, each asserted value matched to the spec outcome) plus a discrimination sensor (behavior-level mutations run in a scratch state and then discarded), after which it writes `.specs/features/[feature]/validation.md` (PASS/FAIL, per-AC evidence, sensor result, diff range) and returns a compact verdict + ranked gaps in chat. It runs read-only over the real tree and does NOT fix.
334
+
335
+ If the Verifier returns FAIL, the orchestrator routes the ranked gaps back to an implementer as fix tasks, then re-dispatches the Verifier - bounded to **3 fix→re-verify iterations** before escalating to the user.
336
+
337
+ If you are unsure whether more tasks remain, check `tasks.md`: if every task is marked complete, dispatch the Verifier now.
338
+
339
+ ---
340
+
341
+ ## Execution Template
342
+
343
+ ```markdown
344
+ ## Implementing T[X]: [Task Title]
345
+
346
+ **Reading**: task definition from tasks.md
347
+ **Dependencies**: [All done? ✅ | Blocked by: TY]
348
+ **Tests**: [unit/e2e/integration/none]
349
+ **Gate**: [quick/full/build]
350
+
351
+ ### Pre-Implementation (MANDATORY)
352
+
353
+ - **Assumptions**: [state explicitly]
354
+ - **Files to touch**: [list ONLY these]
355
+ - **Success criteria**: [how to verify]
356
+
357
+ ### Tests: Write tests derived from spec ACs
358
+
359
+ - Test file(s): [paths]
360
+ - Test count: [N test cases]
361
+ - Spec-derived: each test's asserted value maps to spec-defined outcome (or gap flagged)
362
+
363
+ ### Implement
364
+
365
+ [Write minimum code to pass tests]
366
+
367
+ - Tests modified: None
368
+ - Tests skipped/deleted: None
369
+
370
+ ### VERIFY: Gate Check
371
+
372
+ - Command: [gate check command]
373
+ - Result: [X passed, 0 failed]
374
+ - Test count: [N - matches planned test count]
375
+
376
+ ### Post-Gate
377
+
378
+ - [x] No SPEC_DEVIATION (or markers added)
379
+ - [x] No unnecessary changes made
380
+ - [x] Matches existing patterns
381
+
382
+ **Test Adequacy Review:**
383
+
384
+ *Check A - Sufficient (coverage mapping):*
385
+
386
+ | Done-when criterion / spec AC / listed edge case | `file:line` + assertion expression | Spec-defined outcome | Covered? |
387
+ | ------------------------------------------------- | ---------------------------------- | -------------------- | -------- |
388
+ | [criterion] | `path/to/test.ts:42` - `expect(result.field).toBe(expected)` | [spec value] | ✅ Yes / ⚠️ Gap |
389
+
390
+ *Check C - Necessary (reverse mapping):*
391
+
392
+ | `file:line` + assertion expression | Maps to (AC / edge case / Done-when criterion) | Keep? |
393
+ | ---------------------------------- | ---------------------------------------------- | ----- |
394
+ | `path/to/test.ts:42` - `expect(result.field).toBe(expected)` | [requirement or criterion text] | ✅ Keep |
395
+
396
+ - [ ] Check A: every criterion covered with `file:line` evidence; spec-defined outcomes matched or gap flagged; per-layer depth met
397
+ - [ ] Check B: no shallow assertions; payload/conjunction rule applied to every payload-bearing criterion
398
+ - [ ] Check C: every test maps to a requirement - no speculative or unclaimed tests
399
+ - [ ] Check D: guideline conformance - [guideline file followed, or "none - strong defaults applied"]
400
+
401
+ **Verdict**: [All criteria covered, spec outcomes matched, no shallow assertions, all tests necessary] / [Rewritten: describe what was fixed]
402
+
403
+ **Status**: ✅ Complete | ❌ Blocked | ⚠️ Partial
404
+ ```
405
+
406
+ **After the LAST task:** dispatch the Verifier sub-agent (see step 9 and [sub-agents.md](sub-agents.md)) for independent feature-level validation, including the spec-anchored check and discrimination sensor. Validation always runs automatically - never prompted. Execute is not done until the Verifier reports PASS and the validation report is written, confirmed deterministically by `python3 <skill-dir>/scripts/validate_state.py <feature>` (exit non-zero = not done); see [validate.md](validate.md).
407
+
408
+ ---
409
+
410
+ ## Tips
411
+
412
+ - **One task at a time** - Focus prevents errors
413
+ - **Tools matter** - Wrong MCP = wrong approach
414
+ - **Reuses save tokens** - Copy patterns, don't reinvent
415
+ - **Status then commit, same commit** - Mark `tasks.md` complete before the atomic commit and include that update in it
416
+ - **Stay surgical** - Touch only what's necessary
417
+ - **Commit per task** - Clean git history enables bisect and rollback
418
+ - **Never "while I'm here"** - Scope creep during implementation is the #1 quality killer
419
+ - **Approval is local** - Push, deploy, and other remote/destructive ops need an explicit go-ahead
420
+ - **Learn from mistakes** - If something goes wrong, surface it to the user so it informs the next task
421
+ - **Don't stop at the last commit** - Feature-level validation (step 9) is the final step of Execute, not optional
422
+ - **Plain voice in prose** - Commit bodies and the validation summary follow the writing rules in [coding-principles.md](coding-principles.md): lead with what changed, no filler
423
+ - **Validate the commit message** - `python3 <skill-dir>/scripts/check_commit.py --message "..."` before committing
424
+
425
+ ---
426
+
427
+ ## Pause / End of Session
428
+
429
+ When work is interrupted, paused, or a session ends before the feature is complete:
430
+
431
+ 1. Open `.specs/STATE.md`.
432
+ 2. Locate the `## Handoff` section.
433
+ 3. **Replace only that section's body** with the current snapshot (feature, phase/task, completed, in-progress `file:line`, next step, blockers, uncommitted files, branch). See [memory.md](memory.md) for the exact format.
434
+ 4. Do NOT touch the `## Decisions` section above it - decisions are written only during Design.
435
+
436
+ **Section-scoped write (critical):** Replace the content between the `## Handoff` header and the next `##` header (or end of file). Never overwrite the full file - doing so silently destroys the Decisions log.