@gobing-ai/spur 0.3.86 → 0.3.87

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/config/plugin-scripts.json +4 -0
  3. package/config/templates/feature/default.md +2 -0
  4. package/config/templates/task/brainstorm.md +2 -2
  5. package/config/templates/task/feature-impl.md +2 -2
  6. package/config/templates/task/issue.md +2 -2
  7. package/config/templates/task/meta.md +2 -2
  8. package/config/templates/task/review.md +2 -2
  9. package/config/templates/task/standard.md +2 -2
  10. package/config/workflow-candidates.json +13 -35
  11. package/config/workflows/feature-lifecycle.yaml +6 -0
  12. package/config/workflows/feature-verification.yaml +6 -5
  13. package/config/workflows/idea-pipeline.yaml +72 -36
  14. package/package.json +1 -1
  15. package/plugins/sp/README.md +7 -4
  16. package/plugins/sp/commands/dev-idea.md +9 -2
  17. package/plugins/sp/commands/dev-refactor.md +33 -0
  18. package/plugins/sp/plugin.json +1 -1
  19. package/plugins/sp/references/roles.md +1 -1
  20. package/plugins/sp/scripts/idea-coverage-check.ts +168 -0
  21. package/plugins/sp/scripts/inline-pipeline-parity-check.ts +114 -3
  22. package/plugins/sp/scripts/inline-run-setup.ts +28 -18
  23. package/plugins/sp/skills/brainstorm/SKILL.md +4 -0
  24. package/plugins/sp/skills/code-refactoring/SKILL.md +155 -0
  25. package/plugins/sp/skills/code-refactoring/references/finding-schema.md +74 -0
  26. package/plugins/sp/skills/code-refactoring/references/fix-ladder.md +52 -0
  27. package/plugins/sp/skills/code-refactoring/references/focus-detection.md +44 -0
  28. package/plugins/sp/skills/code-refactoring/references/refactor-finding.schema.json +95 -0
  29. package/plugins/sp/skills/spec-decomposition/references/decomposition.md +23 -16
  30. package/plugins/sp/skills/spur-cli/references/features/acceptance-criteria.md +4 -2
  31. package/plugins/sp/skills/spur-cli/references/features.md +6 -1
  32. package/plugins/sp/skills/spur-cli/references/workflows/authoring-workflows.md +2 -2
  33. package/plugins/sp/skills/spur-cli/references/workflows/operations.md +3 -3
  34. package/plugins/sp/skills/spur-cli/references/workflows/workflow-fit-and-tuning.md +1 -1
  35. package/plugins/sp/skills/spur-dev/references/ac-style-guide.md +24 -0
  36. package/plugins/sp/skills/spur-dev/references/dev-operations.md +18 -2
  37. package/plugins/sp/skills/spur-dev/references/flag-glossary.md +22 -3
  38. package/plugins/sp/skills/spur-dev/references/idea-evaluation.md +11 -1
  39. package/plugins/sp/skills/spur-dev/references/inline-pipeline-driver.md +33 -1
  40. package/plugins/sp/skills/taste-refactoring-api/SKILL.md +44 -1
  41. package/plugins/sp/skills/taste-refactoring-api/references/protocol-modes.md +31 -0
  42. package/plugins/sp/skills/taste-refactoring-architect/SKILL.md +43 -0
  43. package/plugins/sp/skills/taste-refactoring-tests/SKILL.md +42 -0
  44. package/plugins/sp/skills/taste-refactoring-ui/SKILL.md +43 -0
  45. package/schemas/task-batch.schema.json +2 -2
  46. package/spur.js +124 -23
@@ -0,0 +1,52 @@
1
+ # Fix ladder and eligibility
2
+
3
+ Authority: `docs/design/dev-refactor-command.md` §6. The ladder ranks what may be applied
4
+ mechanically versus what needs an operator answer. It reuses the repository severity meaning
5
+ (P1 blocker / P2 major — see [finding-schema.md](./finding-schema.md)) so `--fix blockers-first`
6
+ keeps the glossary sense.
7
+
8
+ ## Rungs
9
+
10
+ | Rung | Example | `preservation` | `fix_eligibility` |
11
+ | --- | --- | --- | --- |
12
+ | Rename / move / dedupe / inline with identical behavior | A3 consolidate, T3 fixture cleanup, ui token normalization | preserving | `auto` |
13
+ | Add missing test, contract field, a11y attribute | T4, api additive, ui P3 | preserving | `auto` |
14
+ | Remove dead code with zero references | A1 direct removal proven dead | preserving | `auto` only when a reference search finds no caller; else `confirm` |
15
+ | Remove a test, endpoint, control, or code path with callers | T1, A1/A2 live, api removal, ui control removal | cutting | `confirm` |
16
+ | Change a contract or observable behavior | api breaking, A5–A7 seam moves | breaking | `confirm` (or `suggest` when multi-task) |
17
+ | Architectural migration plan | A6–A7, ADR candidates | — | `suggest` |
18
+
19
+ Hard rules:
20
+
21
+ - An `auto` fix **never deletes or weakens a test** (no skipped assertions, no loosened
22
+ expectations, no deleted cases).
23
+ - `cutting` and `breaking` findings are **never `auto`** and **never below P2**.
24
+
25
+ ## Apply policies (`--fix`)
26
+
27
+ | Policy | Meaning |
28
+ | --- | --- |
29
+ | `none` (default) | Write both artifacts, perform no edit. |
30
+ | `blockers-first` | Apply P1/P2 findings with `fix_eligibility: auto`. |
31
+ | `all` | Apply every `auto` finding, then queue every `confirm` finding for the taste gate. |
32
+
33
+ ## Apply loop (mirrors `sp:code-simplification` / `dev-simplify`: test-after-each, revert on regression)
34
+
35
+ 1. **Green baseline first.** Run the `--check` command before any edit. A red baseline is a hard
36
+ stop: write the report, apply nothing.
37
+ 2. **One finding at a time.** Apply a single finding (highest severity first: P1 → P4), limited to
38
+ the finding's evidence files.
39
+ 3. **Re-run `--check`.**
40
+ - Passes → mark the finding `status: applied`, continue with the next.
41
+ - Fails → revert **only that finding's own edits**: reverse-apply the exact hunks the finding
42
+ introduced (keep the finding's diff; `git apply -R`, or re-edit the specific lines). Never
43
+ `git checkout -- <file>` a whole evidence file — a later finding may share that file with an
44
+ earlier **applied** finding, and a file-level checkout would revert that applied work too.
45
+ Mark the finding `status: reverted`, continue with the next. Never revert another finding's
46
+ work.
47
+ 4. After the last finding: re-run the structural check on the findings artifact, then write the
48
+ report (see `SKILL.md` phase 5).
49
+
50
+ `confirm` findings in the `all` policy follow the taste gate: applied only after an explicit
51
+ operator `yes` (see the gate matrix in `SKILL.md`). A declined finding is marked `status: rejected`
52
+ or `status: deferred` (headless) per the operator answer — never silently dropped.
@@ -0,0 +1,44 @@
1
+ # Focus auto-detection
2
+
3
+ Authority: `docs/design/dev-refactor-command.md` §8. Lens selection is **deterministic globs, not
4
+ model judgment**, so a misroute is auditable: re-running detection on the same tree yields the same
5
+ set.
6
+
7
+ ## Classification
8
+
9
+ First matching table row per file; the lens set is the **union** across all files in `--scope`.
10
+
11
+ | Order | Glob | Lens |
12
+ | --- | --- | --- |
13
+ | 1 | `**/tests/**`, `**/*.test.*`, `**/*.spec.*`, `**/__tests__/**` | tests |
14
+ | 2 | `apps/web/**`, `**/*.astro`, `**/*.tsx`, `**/*.jsx`, `**/*.css`, `**/*.vue` | ui |
15
+ | 3 | `packages/contracts/**`, `**/routes/**`, `**/openapi*`, `**/*.proto`, `**/*.graphql`, `apps/cli/src/commands/**`, `apps/server/src/**` | api |
16
+ | 4 | anything else (fallback) | architect |
17
+
18
+ Rules:
19
+
20
+ - Evaluate rows in order; a file matching an earlier row is never classified by a later one
21
+ (a `*.test.tsx` file is a **tests** file, not ui).
22
+ - Every file in scope resolves to exactly one lens; the scope set is the union of resolved lenses.
23
+ - An empty scope set (no files matched anything) collapses to the `architect` fallback for the
24
+ whole scope only when scope itself is non-empty; an empty scope is a resolve-phase error.
25
+
26
+ ## `--focus` flag
27
+
28
+ | Value | Behavior |
29
+ | --- | --- |
30
+ | `auto` (default) | Run detection above and use its lens set. |
31
+ | single lens (`api` \| `architect` \| `tests` \| `ui`) | Run only that lens. |
32
+ | comma list (`api,tests`) | Run exactly the listed lenses, in the operator's order. |
33
+
34
+ ## Report before run
35
+
36
+ The resolved lens set **must be reported before any lens runs** — one line naming the chosen
37
+ lenses and, for `auto`, the per-file counts that selected them, e.g.:
38
+
39
+ ```text
40
+ focus=auto → lenses: tests (14 files), api (3 files)
41
+ ```
42
+
43
+ Under `--auto` this line is still written (to the report and the session output); what `--auto`
44
+ skips is only the *interactive* confirmation of scope and lens set, never the report.
@@ -0,0 +1,95 @@
1
+ {
2
+ "$schema": "http://json-schema.org/draft-07/schema#",
3
+ "$id": "https://spur.gobing.ai/schemas/refactor-finding.schema.json",
4
+ "title": "Refactor findings artifact (bare array of findings, design H13 §4)",
5
+ "description": "One JSON object per refactoring finding; the artifact written to .spur/run/<run-id>-refactor-findings.json is a bare array of these objects.",
6
+ "type": "array",
7
+ "items": {
8
+ "type": "object",
9
+ "additionalProperties": false,
10
+ "required": [
11
+ "id",
12
+ "focus",
13
+ "severity",
14
+ "rung",
15
+ "title",
16
+ "evidence",
17
+ "preservation",
18
+ "fix_eligibility",
19
+ "proposal",
20
+ "verify",
21
+ "status"
22
+ ],
23
+ "properties": {
24
+ "id": {
25
+ "type": "string",
26
+ "pattern": "^RF-(api|architect|tests|ui)-[0-9]{3}$",
27
+ "description": "RF-<focus>-<nnn>, e.g. RF-api-001."
28
+ },
29
+ "focus": {
30
+ "type": "string",
31
+ "enum": ["api", "architect", "tests", "ui"],
32
+ "description": "The lens that produced the finding."
33
+ },
34
+ "severity": {
35
+ "type": "string",
36
+ "enum": ["P1", "P2", "P3", "P4"],
37
+ "description": "Repository-authority severity after the lens-native → P1–P4 map (finding-schema.md §5). P0 is outside the map."
38
+ },
39
+ "rung": {
40
+ "type": "string",
41
+ "description": "Lens-native rung kept verbatim: architect A0–A7, tests T0–T7, api compatibility class, ui pass name."
42
+ },
43
+ "title": {
44
+ "type": "string",
45
+ "minLength": 1,
46
+ "description": "One line."
47
+ },
48
+ "evidence": {
49
+ "type": "array",
50
+ "minItems": 1,
51
+ "items": {
52
+ "type": "object",
53
+ "additionalProperties": false,
54
+ "required": ["file", "line"],
55
+ "properties": {
56
+ "file": {
57
+ "type": "string",
58
+ "minLength": 1
59
+ },
60
+ "line": {
61
+ "type": "integer",
62
+ "minimum": 1
63
+ }
64
+ }
65
+ },
66
+ "description": "At least one file:line inside --scope."
67
+ },
68
+ "preservation": {
69
+ "type": "string",
70
+ "enum": ["preserving", "cutting", "breaking"],
71
+ "description": "preserving: behavior identical; cutting: a user-visible feature/test/endpoint/control is removed; breaking: contract or behavior changes for a consumer."
72
+ },
73
+ "fix_eligibility": {
74
+ "type": "string",
75
+ "enum": ["auto", "confirm", "suggest"],
76
+ "description": "auto: mechanical, behavior-preserving, checkable; confirm: needs an operator answer; suggest: report only."
77
+ },
78
+ "proposal": {
79
+ "type": "string",
80
+ "minLength": 1,
81
+ "description": "What to change, imperative."
82
+ },
83
+ "verify": {
84
+ "type": "string",
85
+ "minLength": 1,
86
+ "description": "Command or check that proves the fix; defaults to the --check command."
87
+ },
88
+ "status": {
89
+ "type": "string",
90
+ "enum": ["open", "applied", "reverted", "deferred", "rejected"],
91
+ "description": "Lifecycle of the finding through the gate/apply loop."
92
+ }
93
+ }
94
+ }
95
+ }
@@ -134,14 +134,13 @@ single run-on paragraph on render, even though they look like separate items in
134
134
  "requirements": "- [ ] R1. <text>\n- [ ] R2. <text>\n- [ ] R3. <text>"
135
135
  ```
136
136
 
137
- `R1. <text>\nR2. <text>` (no marker) is the trap. `spur task check` **accepts** it — the L3
138
- R-numbering rule matches the bare `Rn.` token — so nothing fails, and the defect only surfaces later
139
- as an unreadable paragraph in Board preview. Do not rely on `check` to catch this. Keep the `Rn.`
137
+ `R1. <text>\nR2. <text>` and `- R1 — <text>` are the traps: `spur task check` reports
138
+ `L3.requirements-checkbox` and a bare run renders as one paragraph in Board preview. Keep the `Rn.`
140
139
  (period) token inside the marker so R-numbering still resolves; see the canonical rule in
141
140
  `sp:spur-dev` → `references/planning-workflow.md`.
142
141
 
143
142
  Applies to the other body fields too: `plan` as an ordered list (`1. …\n2. …`), `acceptance_criteria`
144
- as a fenced ```` ```gherkin ```` block, and any enumeration inside `background` or `design` as a
143
+ as `- [ ] AC1 — <feature scenario title without its R-number>` bullets (see § Idea-pipeline emission), and any enumeration inside `background` or `design` as a
145
144
  `- ` list.
146
145
 
147
146
  ### Design at create (default) vs `--skip-design`
@@ -207,20 +206,22 @@ scenarios are numbered in the **feature's** namespace. Both appear in a task's
207
206
 
208
207
  The rule:
209
208
 
210
- - **Scenarios covering the task's own requirements carry the task-local R-prefix** —
211
- `Scenario: R3 — <observable outcome>`. Tasks declaring `ac_numbering: task-local` in frontmatter
212
- get these cross-checked by `spur task check` (`L3.ac-requirement-coverage`): a requirement with no
213
- scenario, or a scenario citing a requirement that does not exist, is reported.
214
- - **Scenarios carried verbatim from the feature (for DD-09 traceability) carry NO R-prefix** — copy
215
- the title text only. `normalizeTitle` (`packages/domain/src/bdd/coverage.ts:58`) strips `R\d+`
216
- before matching, so the prefix is invisible to feature coverage anyway; dropping it keeps the
217
- feature's number from being read as a task requirement id. Verified empirically: removing the
218
- prefix from a carried scenario left the feature's orphan count unchanged.
209
+ - **Task AC items are numbered `AC1, AC2, …` (task-local), never `R<n>`** — `R<n>` is the
210
+ Requirements namespace, and two R-numberings in one file is how they get confused.
211
+ - **Bind a scenario to the task requirement it covers with `(req: R<n>)`** —
212
+ `Scenario: AC1 — <observable outcome> (req: R3)` (semicolons for several: `R1; R2`). Tasks
213
+ declaring `ac_numbering: task-local` get these cross-checked by `spur task check`
214
+ (`L3.ac-requirement-coverage`): a requirement with no scenario, or a scenario citing a requirement
215
+ that does not exist, is reported.
216
+ - **Scenarios carried from the feature (for DD-09 traceability) copy the title text after the
217
+ feature's `R<n> —`** — `- [ ] AC2 — <title>`. `normalizeTitle`
218
+ (`packages/domain/src/bdd/coverage.ts`) strips both `AC\d+` and `R\d+` before matching, so the
219
+ prefix is invisible to feature coverage; the feature's number is never read as a task id.
219
220
 
220
221
  **Legacy tasks are exempt.** Most existing tasks predate this and copied feature AC wholesale,
221
222
  carrying the feature's numbers. The coverage check is opt-in precisely so they emit nothing —
222
- absent `ac_numbering`, only DD-09 applies. Opting an old task in is a pure prefix renumber; it cannot
223
- break traceability. New tasks get `ac_numbering: task-local` from the templates automatically;
223
+ absent `ac_numbering`, only DD-09 applies. Legacy `Scenario: R<n> —` titles still bind by prefix, so
224
+ opting an old task in cannot break traceability. New tasks get `ac_numbering: task-local` from the templates automatically;
224
225
  `spur task update <wbs> --ac-numbering task-local` opts in an existing one.
225
226
 
226
227
  Edge-case scenarios may map to tasks, merge into a sibling, or be deferred. Record deferrals
@@ -523,7 +524,7 @@ The payload is a top-level JSON **array** (no `tasks` wrapper):
523
524
  "requirements": "- [ ] R1. Accept a title and an optional description on POST /tasks.\n- [ ] R2. Reject an empty title with a 400 and a reason.\n- [ ] R3. Allocate the task file through the CLI-gated write path.",
524
525
  "design": "Approach: POST /tasks via existing TaskService.create.\nRejected: ad-hoc SQL in handler.\nInvariants: CLI-gated corpus writes only.",
525
526
  "plan": "1. Contract\n2. Handler\n3. Tests",
526
- "acceptance_criteria": "Scenario: create succeeds\n Given a valid title\n When POST /tasks\n Then a task file is allocated"
527
+ "acceptance_criteria": "- [ ] AC1 — User can create a task with required fields"
527
528
  },
528
529
  {
529
530
  "name": "Implement task listing endpoint",
@@ -559,6 +560,12 @@ fields and normal default planning fills them from your analysis; the per-task r
559
560
  batch-create still deepens them when a task needs more detail. Validate locally against the
560
561
  schema before emitting.
561
562
 
563
+ **Pass the deterministic task check.** Each `acceptance_criteria` bullet must be an exact feature
564
+ scenario title (L4.uncovered-task-scenario); add task-local checks as prose after the bullets, not
565
+ as extra bullets. No section body may use `HITL`, `approval`/`approved`, `merged`/`merge event`,
566
+ `content-gate`, `GATED`, or `capstone` as standalone words (L4.gate-language) — say "operator
567
+ answer" / "accepted" instead, and keep enum values out of that list too.
568
+
562
569
  **The order sidecar.** Also emit the private task-order sidecar at
563
570
  `.spur/run/<runId>-idea-task-order.json`: a JSON array (one entry per batch item) of
564
571
  `{ name: <exact batch item name>, depends_on_names: [<batch item names>] }` declaring
@@ -47,8 +47,10 @@ Every scenario/item carries an `R1, R2, …` prefix:
47
47
 
48
48
  - **Sequential within a feature**, starting at R1.
49
49
  - **Stable forever.** Never renumber once tasks exist — tasks match AC by **normalized scenario
50
- title** (the `R<n> —` prefix is stripped on comparison), so renumbering around a title is safe
51
- but *rewording* a title breaks the coverage edge.
50
+ title** (the `R<n> —` prefix is stripped on comparison, as is the task side's `AC<n> —`), so
51
+ renumbering around a title is safe but *rewording* a title breaks the coverage edge.
52
+ - **`R<n>` is the feature's namespace.** Task AC items are numbered `AC<n>` (task-local) and copy
53
+ the feature title after their own prefix; task Requirements own `R<n>.` inside the task.
52
54
  - **One R-number = one scenario.** Don't split one requirement across scenarios under a single
53
55
  R-number; don't merge two requirements into one scenario.
54
56
 
@@ -58,6 +58,9 @@ spur feature create "Planning layer" --parent H # → H<n>
58
58
  spur feature create "Task CLI" --parent H1 # → H1<n>
59
59
  ```
60
60
 
61
+ `create --json` returns `{ ref: { kind, id, filePath, folder }, content }` — read the id from
62
+ `.ref.id`, not a top-level `.id`.
63
+
61
64
  To restructure, use `move` — never hand-edit an ID. `move <id> --parent <new>` re-parents the
62
65
  subtree and **cascade-renames** every descendant; omit `--parent` to lift it to a top-level group.
63
66
 
@@ -184,7 +187,9 @@ spur feature check --strict --json # warnings → failures
184
187
  ```
185
188
 
186
189
  The 4-layer validator (frontmatter, AC syntax, children-limit/structure, L4 traceability) emits its
187
- verdict and findings as JSON. **Query this, don't re-derive it** — the rules live in the CLI, never
190
+ verdict and findings as a JSON **array**, one entry per feature (`jq '.[0].pass'`, `.[0].findings[].code`),
191
+ like `spur task check --json`. Gherkin AC must keep its `Feature:` line (`L3.ac-bdd-error`) and
192
+ every `Scenario:` title is the identity key tasks reference verbatim. **Query this, don't re-derive it** — the rules live in the CLI, never
188
193
  restated as prose here. This is what `sp:spur-dev`'s feature-check gate loop runs.
189
194
 
190
195
  ## Status sync - `sync`
@@ -137,8 +137,8 @@ runner (see [validation-and-extension.md](validation-and-extension.md)).
137
137
  Order matters for both guards and conditions: **the first that passes wins.** Put the discriminating
138
138
  guard before the unconditional fallback (`always` / no-guard edge). For multi-condition gates (doctor
139
139
  + task check, quality gate + attempt cap), prefer a **soft probe** shell that writes PASS|FAIL and
140
- always exits 0, then branch with ordered status-file guards — see shipped `basic.yaml` /
141
- `task-pipeline.yaml` (more reliable than `action-ok` alone when more than one condition decides the edge).
140
+ always exits 0, then branch with ordered status-file guards — see shipped
141
+ `task-pipeline.yaml` / `wrapup-pipeline.yaml` (more reliable than `action-ok` alone when more than one condition decides the edge).
142
142
 
143
143
  ## Template variables
144
144
 
@@ -28,9 +28,9 @@ Authored workflows default to a project-local directory, grouped by purpose:
28
28
  ```
29
29
 
30
30
  A `--file <path>` argument overrides the default. Keep one workflow per file, named for what it does
31
- (`approval.yaml`, `import-file.yaml`), not for its mode. The canonical example
32
- (`basic.yaml`) lives here; copy real schema shapes from it rather than from a
33
- half-remembered snippet.
31
+ (`approval.yaml`, `import-file.yaml`), not for its mode. Copy real schema shapes from a
32
+ retained definition such as `task-pipeline.yaml` rather than from a half-remembered
33
+ snippet.
34
34
 
35
35
  ## Sub-procedure: mode-selection gate
36
36
 
@@ -95,7 +95,7 @@ and orders capabilities, it does not contain them** (ADR-069).
95
95
  between them are one judgment step.
96
96
  - [ ] **Soft status-file probe over repeated probing.** Run the expensive check once in an action
97
97
  that always exits 0 and writes its verdict to a run-scoped file; branch with ordered cheap
98
- guards that read that file. One subprocess instead of one per branch — the `basic.yaml` and
98
+ guards that read that file. One subprocess instead of one per branch — the
99
99
  `task-pipeline.yaml` quality-gate idiom.
100
100
  - [ ] **Order guards cheapest-discriminating-first.** The first passing guard wins, so a `test -f`
101
101
  ahead of a `spur … --json` parse skips the expensive call on the common path.
@@ -27,6 +27,14 @@ Rules:
27
27
  - **One R-number = one scenario.** Never split a requirement across multiple scenarios
28
28
  under the same R-number; never merge two requirements into one scenario.
29
29
 
30
+ ## Task-side numbering (`AC<n>`)
31
+
32
+ `R<n>` is the **feature** scenario key and the **task Requirements** key. Task `### Acceptance
33
+ Criteria` items therefore use their own namespace: `- [ ] AC1 — <title>` or `Scenario: AC1 — <title>`,
34
+ numbered task-locally. Carry a feature scenario by copying its title after the prefix (`normalizeTitle`
35
+ strips both `AC<n>` and `R<n>`, so DD-09 matching is unaffected); bind a task requirement with
36
+ `(req: R<n>)`. Legacy tasks that wrote `- [ ] R<n> —` / `Scenario: R<n> —` keep working unchanged.
37
+
30
38
  ## Two AC tiers (authoring convention)
31
39
 
32
40
  A planning convention (DD-06 "permissive start"), not a `spur feature check` feature today —
@@ -73,6 +81,10 @@ scenario, it matches by title. Rules:
73
81
  Registered user can log in with email and password" is traceable.
74
82
  - **No synonyms in cross-references.** The title in the feature file and the title in the
75
83
  task's AC reference must be byte-identical.
84
+ - **Avoid gate vocabulary in titles.** `spur task check` (L4.gate-language) rejects task sections
85
+ containing `HITL`, `approval`/`approved`, `merged`/`merge event`, `content-gate`, `GATED`, or
86
+ `capstone` as standalone words; task AC bullets copy scenario titles verbatim, so a title using
87
+ them fails every child task. Write "pause for an operator answer" instead of "HITL approval".
76
88
 
77
89
  ## Verdict AC ↔ feature scenario linkage
78
90
 
@@ -199,6 +211,18 @@ Use the canonical BDD template at `templates/bdd/gherkin.md`. Key rules:
199
211
  - **When** describes the single action under test.
200
212
  - **Then** asserts the observable outcome.
201
213
  - **And** chains additional preconditions, actions, or assertions.
214
+ - **Trace the scenario to its requirements (0887 R4).** Directly under each `Scenario:`
215
+ heading add a comment line listing the requirement-inventory ids the scenario covers;
216
+ the BDD parser skips `#` comment lines, so the form is checker-inert:
217
+
218
+ ```gherkin
219
+ Scenario: Registered user can log in with email and password
220
+ # covers: I1, I3
221
+ Given ...
222
+ ```
223
+
224
+ Every inventory item without a `[deferred: ...]` marker must be covered by at least one
225
+ scenario — `idea-coverage-check` measures this at the idea-pipeline's ac-generate boundary.
202
226
 
203
227
  Avoid:
204
228
 
@@ -85,7 +85,8 @@ each would be scope creep for one-liner procedures.
85
85
  | 13a | parallel | `dev-parallel` | `Skill()` | `sp:parallel-execution` | `--tasks <selector> [--feature <id>] [--mode <fan-out\|review-panel\|investigation>] [--agent <inline\|auto\|name>] [--json]` |
86
86
  | 14 | wrap | `dev-wrap` | `Skill()` | `spur workflow run` (wrapup-pipeline) | `<wbs> [--agent <inline\|auto\|name>] [--auto] [--merge] [--dry-run]` |
87
87
  | 15 | wrapall | `dev-wrapall` | `Skill()` | `spur workflow run` (wrapup-pipeline) | `[--since <iso>] [--feature <id>] [--status <s>] [--agent <inline\|auto\|name>] [--auto] [--merge] [--dry-run]` |
88
- | 16 | idea | `dev-idea` | `Skill()` | inline driver (`idea-pipeline`) or async workflow | `"<idea>" [--auto] [--skip-design] [--approve-taste] [--agent <inline\|auto\|name>]` |
88
+ | 16 | idea | `dev-idea` | `Skill()` | inline driver (`idea-pipeline`) or async workflow | `"<idea>" [--from-file <path>] [--auto] [--skip-design] [--approve-taste] [--agent <inline\|auto\|name>]` |
89
+ | 17 | refactor | `dev-refactor` | `Skill()` | `sp:code-refactoring` skill (thin wrapper, ADR-032) | `[<description>] [--scope <path>] [--focus <api\|architect\|tests\|ui\|auto>] [--fix <none\|blockers-first\|all>] [--check <cmd>] [--agent <inline\|auto\|name>] [--auto]` |
89
90
 
90
91
  ## Skill-backed operations
91
92
 
@@ -335,11 +336,18 @@ must not be changed without updating the backing skill.
335
336
  ### 16. idea
336
337
 
337
338
  - **Purpose:** Turn a vague idea into a feature with AC and a decomposed task batch — the unified entry point for the planning half.
338
- - **Inputs:** `"<idea>"` (required, positional, quoted). Three everyday axes:
339
+ - **Inputs:** `"<idea>"` (quoted) or `--from-file <path>` — exactly one; the two are mutually
340
+ exclusive (0887 R7). Three everyday axes:
339
341
  - `--auto` — skip **objective** HITL (feature-check, batch-create); taste gates still pause.
340
342
  - `--skip-design` — design package off (system-design + task Design).
341
343
  - `--approve-taste` — with `--auto`, skip **all** remaining taste pauses this run (idea-eval + design-approval). Sets `idea_approved=true` and `design_approved=true`.
342
344
  Aliases (prefer `--approve-taste`): `--idea-approved` → `idea_approved`; `--design-approved` → `design_approved`. There is **no** `--design` force flag.
345
+ - `--from-file <path>` — read the idea text from a file instead of the positional argument
346
+ (verbatim; long/multiline asks). Mutually exclusive with `"<idea>"`.
347
+ - **Verbatim idea artifact:** before `start` executes, the driver persists the idea argument (or
348
+ `--from-file` contents) unmodified to `.spur/run/<run-id>-idea-input.md` (0887 R1); the precheck
349
+ fails the run when that file is empty or missing, and every model-bearing stage prompt treats
350
+ it as the authoritative ask (R2).
343
351
  - **Backing:** `idea-pipeline.yaml` through the inline driver for omitted/`inline`, or `spur workflow run idea-pipeline.yaml --async` for `auto`/name.
344
352
  - **Behavior:** Builds vars from the table above and drives the idea pipeline. Flow: discovery → **idea-eval** (taste; reject → cancelled) → feature-create → ac-generate → feature-check → system-design (conditional) → design-approval (taste) → decompose → batch-create (`--skip-ready`) → ready-prepare (ready checklist per created task + ready-evidence sidecar, 0788) → handoff. STOPS at handoff — no task execution, no pipeline nesting. Headless runs use one `trace --follow`; cancellation is reported stopped only when `workflow cancel --json` returns `killed: true`.
345
353
  - **Delegation:** Host-session inline driver by default; explicit executor selection uses the async workflow worker.
@@ -356,6 +364,14 @@ must not be changed without updating the backing skill.
356
364
 
357
365
  - **Taste pre-clear (`--approve-taste`):** owned with design-approval var semantics in [cross-cutting.md](cross-cutting.md) § "Design Approval Gate"; idea-eval uses the parallel `idea_approved` var. One CLI flag sets both.
358
366
 
367
+ ### 17. refactor
368
+
369
+ - **Purpose:** Lens-routed refactoring with a preservation contract — route a scope through taste lenses (api / architect / tests / ui, auto-detected by path), classify findings on the shared `refactor-finding` schema, and apply via the fix ladder without breaking preserved behavior.
370
+ - **Inputs:** `[<description>]` free-text steering, `--scope <path>` (default: working tree), `--focus <lens>` (default: auto), `--fix <policy>` (default: none), `--check <cmd>` (default: project gate), `--agent <selector>` (default: inline), `--auto` (default: off).
371
+ - **Backing:** `sp:code-refactoring` skill — the command carries zero orchestration logic (ADR-032); the skill owns focus auto-detection, lens dispatch, the P1–P4 severity map, objective/taste gates, and the fix ladder with revert-on-regression.
372
+ - **Behavior:** Green `--check` baseline → focus detection → lens dispatch → findings to `.spur/run/<run-id>-refactor-findings.json` + human report `.spur/run/<run-id>-refactor-report.md` (lens set, P1–P4 table, preservation summary, applied/reverted/deferred). `--fix blockers-first` auto-applies P1/P2 `auto`-eligible findings; `--fix all` extends to `operator`-eligible P3/P4; every fix re-runs `--check` and reverts on regression. Tests are never removed or weakened by an `auto` fix.
373
+ - **Operator taste gates:** every `cutting` or `breaking` finding pauses for an explicit operator answer in every mode — even `--auto` skips only objective gates. `--fix none` (the default) writes both artifacts with no edits.
374
+
359
375
  ---
360
376
 
361
377
  ## Inline operations
@@ -160,13 +160,21 @@ Scope the operation to all tasks under a feature id (`^[A-Z][1-9]*$`). On featur
160
160
  commands (`dev-wrapall`) it also advances the feature through legal lifecycle edges with guards
161
161
  honored.
162
162
 
163
+ ### `--check <cmd>` — validation command for iterate-and-check loops
164
+
165
+ **Anchor:** `#flag-check`.
166
+
167
+ Verification command a command iterates against (`dev-simplify`, `dev-refactor`). The command
168
+ establishes a baseline with it before the first change and re-runs it after each change.
169
+
163
170
  ### `--focus <dims>` — constrain the operation to specific dimensions
164
171
 
165
172
  **Anchor:** `#flag-focus`.
166
173
 
167
174
  Constrain the operation to a named subset of dimensions — review dimensions on `dev-review`/
168
175
  `dev-verify`/`dev-verifyall` (`all|stack|dependencies|data|flows|api|security|quality|performance`),
169
- a refine focus mode on `dev-refine`/`dev-refineall`, or a reconstruction lens on `dev-reverse`.
176
+ a refactor lens set on `dev-refactor` (`api|architect|tests|ui|auto`), a refine focus mode on
177
+ `dev-refine`/`dev-refineall`, or a reconstruction lens on `dev-reverse`.
170
178
  Narrowing reduces token cost; omitting runs
171
179
  all dimensions.
172
180
 
@@ -175,7 +183,7 @@ all dimensions.
175
183
  **Anchor:** `#flag-scope`.
176
184
 
177
185
  Limit the operation to a file or directory path (`dev-arch`, `dev-debug`, `dev-fixall`,
178
- `dev-gitmsg`, `dev-gtd`, `dev-simplify`) to bound the working set.
186
+ `dev-gitmsg`, `dev-gtd`, `dev-refactor`, `dev-simplify`) to bound the working set.
179
187
 
180
188
  ### `--all` — widen the operation to everything in its domain
181
189
 
@@ -253,7 +261,8 @@ tasks by `updated_at >= date`).
253
261
 
254
262
  **Anchor:** `#flag-fix`.
255
263
 
256
- Remediation policy on verify-family commands (`dev-verify`, `dev-verifyall`):
264
+ Remediation policy on verify-family commands (`dev-verify`, `dev-verifyall`) and the refactor
265
+ coordinator (`dev-refactor`):
257
266
  `none|blockers-first|all`. `none` reports findings without fixing; `blockers-first` fixes only P1/P2;
258
267
  `all` fixes everything found. Deprecated on `dev-review` (routes to `dev-verify --fix`).
259
268
 
@@ -294,6 +303,16 @@ verifying a task whose artifact is intentionally not yet shippable (e.g. a doc-o
294
303
  Omit the design package (system-design satellite + task `### Design`) on planning commands
295
304
  (`dev-plan`, `dev-idea`). The task is created without the design section; refine supplies it later.
296
305
 
306
+ ### `--from-file <path>` — read the idea from a file (dev-idea)
307
+
308
+ **Anchor:** `#flag-from-file`.
309
+
310
+ `dev-idea` reads the idea text from `<path>` instead of the positional argument. Mutually
311
+ exclusive with `"<idea>"` — exactly one must be present. The file's contents become the verbatim
312
+ idea text, persisted unmodified to `.spur/run/<run-id>-idea-input.md` (0887 R1) and treated as
313
+ the authoritative ask by every model-bearing stage prompt. Useful for long or multiline asks
314
+ that are awkward to quote (0887 R7).
315
+
297
316
  ### `--output <path>` — write the result to a path
298
317
 
299
318
  **Anchor:** `#flag-output`.
@@ -20,7 +20,9 @@ recommendation is mandatory, stakes in plain English, and approve/reject is the
20
20
  re-author the report.
21
21
 
22
22
  **Sidecar rule:** The enhanced idea does **not** overwrite `vars.idea`. Feature-create reads both
23
- the original idea and this report.
23
+ the original idea and this report. The operator's verbatim ask of record is the run's
24
+ `.spur/run/<run-id>-idea-input.md` (persisted by the pipeline `start` state / the inline driver
25
+ before any processing); the `## Requirement inventory` items trace back to it.
24
26
 
25
27
  ## Template
26
28
 
@@ -30,6 +32,13 @@ the original idea and this report.
30
32
  ## Enhanced Idea
31
33
  <one-paragraph refined statement of what the idea actually requires — the "real requirement" after discovery sharpens the vague input>
32
34
 
35
+ ## Requirement inventory
36
+ <mandatory — the coverage gate (idea-coverage-check) parses this section, so keep the exact `- I<n> — ` item form>
37
+ - I1 — <requirement stated as an ask, quoting or paraphrasing the source line from the run's idea-input artifact> (source: "<quoted fragment from the operator's idea>")
38
+ - I2 — <next requirement>
39
+ - I<n> — <optional: a requirement explicitly out of scope> [deferred: <reason>]
40
+ - I<n> — <optional: an ambiguous ask> [unclear: <why it is ambiguous>] — an unclear marker does not exempt the item; it still needs coverage or an explicit deferral
41
+
33
42
  ## Scores
34
43
 
35
44
  | Dimension | Score (0–5) | Rationale |
@@ -74,6 +83,7 @@ Stakes: <plain-English cost of proceeding vs not; reversibility; blast radius>
74
83
  |------|--------|
75
84
  | Filled instance path | `.spur/run/idea-eval-report.md` |
76
85
  | Template home | this file |
86
+ | Requirement inventory | mandatory `## Requirement inventory` section (0887 R3); consumed by `idea-coverage-check` (R4) |
77
87
  | HITL state | `idea-eval` in `idea-pipeline.yaml` |
78
88
  | Approve | continue → `feature-create` |
79
89
  | Reject / cancel | → `cancelled` (no feature) |
@@ -160,6 +160,35 @@ the human/native presentation layer — labels are display addresses only, never
160
160
  This is required for the normal `testing → done` provenance guard. Planning pipelines have no
161
161
  task lifecycle link and skip this task-specific action.
162
162
 
163
+ ### Idea-pipeline quick start (0887 R1/R2)
164
+
165
+ Minimum files to read for `/sp:dev-idea` inline runs — then drive `idea-pipeline.yaml`:
166
+
167
+ - `.spur/workflows/idea-pipeline.yaml` (the machine) and this driver.
168
+ - `plugins/sp/skills/spur-dev/references/idea-evaluation.md` (report template incl. the mandatory
169
+ `## Requirement inventory`) and `references/ac-style-guide.md` (scenario `# covers:` form).
170
+ - `references/dev-operations.md` § idea for the stage-by-stage surface.
171
+
172
+ **Persist the verbatim idea FIRST.** Before executing the `start` state, write the operator's
173
+ idea argument (or `--from-file` contents) **unmodified** to
174
+ `.spur/run/<run-id>-idea-input.md`; the run precheck fails when that file is empty or missing,
175
+ and every model-bearing stage prompt treats it as the authoritative ask. `--from-file` and the
176
+ positional idea are mutually exclusive — exactly one must be present.
177
+
178
+ Expected artifacts per stage (all run-scoped under `.spur/run/<run-id>-*`):
179
+
180
+ | Stage | Artifacts |
181
+ | ----- | --------- |
182
+ | start | `-idea-input.md` (verbatim idea), `-idea-precheck-doctor.status` |
183
+ | discovery | `-idea-eval-report.md` (with `## Requirement inventory`), `-idea-needs-design.json` |
184
+ | feature-create | `-idea-feature-id.txt`, `-idea-goal.md`, `-idea-scope.md` |
185
+ | ac-generate | `-idea-ac-content.md`, `-idea-ac-check.status`, `-idea-coverage.status` |
186
+ | system-design | `-idea-design-review.md`, `-idea-design-check.status` |
187
+ | decompose | `-idea-task-batch.json`, `-idea-task-order.json` |
188
+ | batch-create-run | `-idea-batch-create-result.json`, `-idea-batch-create.done`/`.failed` |
189
+ | ready-prepare | `-idea-ready.json` |
190
+ | handoff-finalize | `-idea-handoff.md` |
191
+
163
192
  ## Comprehensive-check retention and evidence (R7/R8)
164
193
 
165
194
  **R7 — comprehensive checks stay at their owning boundaries.** Quick readiness and plan projection are
@@ -428,7 +457,10 @@ The driver reaches it through the existing run delegate (`$SETUP_SCRIPT`,
428
457
  under its declared error policy and `failed` otherwise; `--duration-ms` is the wall clock the
429
458
  driver measured around the action. This writes the `action_runs` row (node, kind, status, `ok`,
430
459
  `duration_ms`, `run_id`) the engine would have written, so the run's rows are queryable by run id
431
- (`spur workflow progress <run-id>`, `ActionRunDao`) without reading the text log.
460
+ (`spur workflow progress <run-id>`, `ActionRunDao`) without reading the text log. The writer
461
+ back-dates the row's `started_at` from its own `completed_at` minus the measured duration
462
+ (0887 R8), so `completed_at − started_at == duration_ms` exactly; a back-date failure is
463
+ recorded (`action.backdate`) and never affects the run.
432
464
 
433
465
  - **At the run's declared terminal state** — before the driver reports the run complete, close the
434
466
  row so a successful inline run is never left non-terminal for `spur workflow clean` to reap as