@opengsd/gsd-core 1.7.0 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +1 -1
- package/.claude-plugin/plugin.json +1 -1
- package/.opencode/plugins/gsd-core.js +14 -0
- package/README.md +2 -0
- package/agents/gsd-debug-session-manager.md +42 -4
- package/agents/gsd-debugger.md +87 -29
- package/agents/gsd-executor.md +29 -2
- package/agents/gsd-planner.md +29 -36
- package/agents/gsd-verifier.md +2 -2
- package/bin/install.js +1152 -80
- package/commands/gsd/ai-integration-phase.md +1 -1
- package/commands/gsd/mempalace-capture.md +9 -5
- package/commands/gsd/new-milestone.md +1 -1
- package/commands/gsd/plan-phase.md +5 -3
- package/commands/gsd/plan-review-convergence.md +3 -2
- package/gsd-core/bin/gsd-tools.cjs +1878 -2507
- package/gsd-core/bin/lib/adapter-imperative.cjs +8 -1
- package/gsd-core/bin/lib/agent-command-router.cjs +20 -5
- package/gsd-core/bin/lib/api-coverage.cjs +338 -45
- package/gsd-core/bin/lib/broken-windows.cjs +716 -0
- package/gsd-core/bin/lib/capability-command-router.cjs +733 -0
- package/gsd-core/bin/lib/capability-registry.cjs +155 -86
- package/gsd-core/bin/lib/capability-writer.cjs +6 -1
- package/gsd-core/bin/lib/check-command-router.cjs +128 -25
- package/gsd-core/bin/lib/claude-orchestration-command-router.cjs +115 -27
- package/gsd-core/bin/lib/claude-orchestration.cjs +84 -9
- package/gsd-core/bin/lib/command-aliases.cjs +14 -0
- package/gsd-core/bin/lib/commands.cjs +81 -4
- package/gsd-core/bin/lib/config-loader.cjs +14 -2
- package/gsd-core/bin/lib/config.cjs +69 -18
- package/gsd-core/bin/lib/core-utils.cjs +6 -1
- package/gsd-core/bin/lib/decisions.cjs +32 -8
- package/gsd-core/bin/lib/docs.cjs +6 -0
- package/gsd-core/bin/lib/external-descriptor-trust.cjs +14 -2
- package/gsd-core/bin/lib/gap-checker.cjs +17 -2
- package/gsd-core/bin/lib/init.cjs +111 -47
- package/gsd-core/bin/lib/install-engine.cjs +298 -23
- package/gsd-core/bin/lib/install-profiles.cjs +239 -1
- package/gsd-core/bin/lib/installer-migrations/005-opencode-baseline-commands-dir.cjs +146 -0
- package/gsd-core/bin/lib/installer-migrations/006-pi-extension-cjs-to-js.cjs +91 -0
- package/gsd-core/bin/lib/installer-migrations.cjs +44 -5
- package/gsd-core/bin/lib/markdown-sectionizer.cjs +107 -0
- package/gsd-core/bin/lib/milestone.cjs +246 -12
- package/gsd-core/bin/lib/model-catalog.cjs +19 -4
- package/gsd-core/bin/lib/model-resolver.cjs +189 -7
- package/gsd-core/bin/lib/onboard-projection.cjs +11 -8
- package/gsd-core/bin/lib/phase-id.cjs +26 -4
- package/gsd-core/bin/lib/phase.cjs +201 -12
- package/gsd-core/bin/lib/plan-scan.cjs +70 -2
- package/gsd-core/bin/lib/roadmap-parser.cjs +7 -4
- package/gsd-core/bin/lib/roadmap.cjs +13 -3
- package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +7 -1
- package/gsd-core/bin/lib/runtime-artifact-layout.cjs +22 -8
- package/gsd-core/bin/lib/runtime-hooks-surface.cjs +16 -0
- package/gsd-core/bin/lib/smart-entry.cjs +69 -4
- package/gsd-core/bin/lib/state-document.cjs +7 -4
- package/gsd-core/bin/lib/state-transition.cjs +22 -1
- package/gsd-core/bin/lib/state.cjs +65 -11
- package/gsd-core/bin/lib/surface.cjs +51 -9
- package/gsd-core/bin/lib/uat.cjs +420 -5
- package/gsd-core/bin/lib/validate.cjs +12 -8
- package/gsd-core/bin/lib/verification.cjs +112 -17
- package/gsd-core/bin/lib/verify.cjs +220 -22
- package/gsd-core/bin/shared/config-schema.manifest.json +3 -2
- package/gsd-core/references/api-coverage.md +37 -7
- package/gsd-core/references/checkpoints.md +1 -1
- package/gsd-core/references/common-bug-patterns.md +13 -0
- package/gsd-core/references/debugger-bug-taxonomy.md +111 -0
- package/gsd-core/references/debugger-fix-acceptance.md +157 -0
- package/gsd-core/references/debugger-philosophy.md +1 -0
- package/gsd-core/references/debugger-prevention.md +98 -0
- package/gsd-core/references/debugger-rca-branching.md +98 -0
- package/gsd-core/references/debugger-repro-hardening.md +130 -0
- package/gsd-core/references/debugger-sbfl.md +110 -0
- package/gsd-core/references/debugger-semantic-recall.md +81 -0
- package/gsd-core/references/execute-phase-quota-recovery.md +55 -0
- package/gsd-core/references/execute-phase-requirement-revert.md +8 -0
- package/gsd-core/references/execute-phase-response-language.md +7 -0
- package/gsd-core/references/planner-antipatterns.md +6 -0
- package/gsd-core/references/planner-mvp-mode.md +12 -13
- package/gsd-core/references/planner-preconditions.md +156 -0
- package/gsd-core/references/planner-reversibility.md +132 -0
- package/gsd-core/references/reviewer-instances.md +9 -7
- package/gsd-core/references/skeleton-template.md +1 -1
- package/gsd-core/references/thinking-models-planning.md +3 -1
- package/gsd-core/templates/DEBUG.md +5 -3
- package/gsd-core/workflows/add-phase.md +2 -0
- package/gsd-core/workflows/add-tests.md +3 -1
- package/gsd-core/workflows/add-todo.md +32 -1
- package/gsd-core/workflows/ai-integration-phase.md +4 -2
- package/gsd-core/workflows/audit-fix.md +2 -2
- package/gsd-core/workflows/check-todos.md +3 -1
- package/gsd-core/workflows/cleanup.md +7 -1
- package/gsd-core/workflows/code-review.md +17 -5
- package/gsd-core/workflows/complete-milestone.md +3 -0
- package/gsd-core/workflows/debug.md +25 -5
- package/gsd-core/workflows/diagnose-issues.md +1 -1
- package/gsd-core/workflows/discovery-phase.md +7 -0
- package/gsd-core/workflows/discuss-phase/templates/context.md +16 -2
- package/gsd-core/workflows/discuss-phase-assumptions.md +3 -0
- package/gsd-core/workflows/do.md +7 -1
- package/gsd-core/workflows/docs-update.md +1 -0
- package/gsd-core/workflows/eval-review.md +3 -0
- package/gsd-core/workflows/execute-phase/steps/post-merge-gate.md +4 -4
- package/gsd-core/workflows/execute-phase/steps/regression-gate.md +2 -2
- package/gsd-core/workflows/execute-phase.md +25 -34
- package/gsd-core/workflows/execute-plan.md +15 -4
- package/gsd-core/workflows/graduation.md +3 -0
- package/gsd-core/workflows/health.md +7 -1
- package/gsd-core/workflows/help/modes/full.md +6 -2
- package/gsd-core/workflows/import.md +8 -2
- package/gsd-core/workflows/inbox.md +7 -0
- package/gsd-core/workflows/ingest-docs.md +15 -10
- package/gsd-core/workflows/manager.md +3 -1
- package/gsd-core/workflows/map-codebase.md +4 -4
- package/gsd-core/workflows/mvp-phase.md +3 -0
- package/gsd-core/workflows/new-milestone.md +69 -21
- package/gsd-core/workflows/new-project.md +17 -15
- package/gsd-core/workflows/new-workspace.md +3 -1
- package/gsd-core/workflows/onboard.md +3 -0
- package/gsd-core/workflows/plan-phase.md +14 -5
- package/gsd-core/workflows/plan-review-convergence.md +48 -3
- package/gsd-core/workflows/plant-seed.md +3 -0
- package/gsd-core/workflows/profile-user.md +7 -1
- package/gsd-core/workflows/progress.md +31 -3
- package/gsd-core/workflows/quick.md +19 -7
- package/gsd-core/workflows/remove-workspace.md +3 -0
- package/gsd-core/workflows/review.md +89 -73
- package/gsd-core/workflows/scan.md +1 -1
- package/gsd-core/workflows/secure-phase.md +3 -0
- package/gsd-core/workflows/settings-integrations.md +3 -0
- package/gsd-core/workflows/settings.md +3 -0
- package/gsd-core/workflows/ship.md +50 -3
- package/gsd-core/workflows/sketch.md +3 -0
- package/gsd-core/workflows/smart-entry.md +3 -0
- package/gsd-core/workflows/spike.md +7 -1
- package/gsd-core/workflows/ui-phase.md +3 -1
- package/gsd-core/workflows/ui-review.md +3 -0
- package/gsd-core/workflows/undo.md +7 -0
- package/gsd-core/workflows/update.md +2 -0
- package/gsd-core/workflows/validate-phase.md +3 -0
- package/gsd-core/workflows/verify-phase.md +2 -2
- package/gsd-core/workflows/verify-work.md +7 -3
- package/hooks/dist/gsd-context-monitor.js +27 -9
- package/hooks/dist/gsd-statusline.js +88 -3
- package/hooks/gsd-context-monitor.js +27 -9
- package/hooks/gsd-statusline.js +88 -3
- package/package.json +6 -4
- package/pi/gsd.cjs +8 -2
- package/scripts/changeset/lint.cjs +1 -0
- package/scripts/changeset/parse.cjs +26 -0
- package/scripts/check-glossary-refs.cjs +220 -0
- package/scripts/ci-rebase-check.cjs +48 -4
- package/scripts/gen-adr-index.cjs +526 -0
- package/scripts/gen-test-timings.cjs +201 -0
- package/scripts/lint-portable-timeout.cjs +140 -0
- package/scripts/lint-test-file-count.allowlist.json +1 -0
- package/scripts/release-tarball-smoke.cjs +18 -11
- package/scripts/run-tests.cjs +420 -58
- package/skills/gsd-ai-integration-phase/SKILL.md +1 -1
- package/skills/gsd-mempalace-capture/SKILL.md +9 -5
- package/skills/gsd-new-milestone/SKILL.md +1 -1
- package/skills/gsd-plan-phase/SKILL.md +5 -3
- package/skills/gsd-plan-review-convergence/SKILL.md +3 -2
- package/vscode/package.json +1 -1
|
@@ -200,9 +200,23 @@ function mapToolInput(args) {
|
|
|
200
200
|
* @param {string} [opts.cwd] working directory for the child
|
|
201
201
|
* @returns {{ stdout: string, exitCode: number, timedOut: boolean }}
|
|
202
202
|
*/
|
|
203
|
+
const warnedMissingHooks = new Set();
|
|
204
|
+
|
|
203
205
|
function runHook(hookFile, payload, opts = {}) {
|
|
204
206
|
const hookPath = path.join(HOOKS_DIR, hookFile);
|
|
205
207
|
if (!fs.existsSync(hookPath)) {
|
|
208
|
+
// A missing guard script means the guard is silently NOT enforced — the
|
|
209
|
+
// exact failure mode of #2305 (plugin staged, hooks bundle not). Never
|
|
210
|
+
// break the tool call (the adapter's design contract), but never be
|
|
211
|
+
// silent about it either: warn loudly, once per hook file.
|
|
212
|
+
if (!warnedMissingHooks.has(hookFile)) {
|
|
213
|
+
warnedMissingHooks.add(hookFile);
|
|
214
|
+
console.error(
|
|
215
|
+
`[gsd-core] hook script missing: ${hookPath} — ${hookFile} is NOT ` +
|
|
216
|
+
"enforced. The GSD install may be incomplete; reinstall (or run " +
|
|
217
|
+
"/gsd-update) to restage the hooks/ bundle.",
|
|
218
|
+
);
|
|
219
|
+
}
|
|
206
220
|
return { stdout: "", exitCode: 0, timedOut: false };
|
|
207
221
|
}
|
|
208
222
|
const timeout = opts.timeout ?? 8000;
|
package/README.md
CHANGED
|
@@ -60,6 +60,8 @@ New here? Follow [Your first project](docs/tutorials/your-first-project.md) for
|
|
|
60
60
|
|
|
61
61
|
## Documentation
|
|
62
62
|
|
|
63
|
+
**What's new in 1.7.0** → [docs/whats-new-1.7.0.md](docs/whats-new-1.7.0.md)
|
|
64
|
+
|
|
63
65
|
**Tutorials** — learning by doing:
|
|
64
66
|
- [Your first project](docs/tutorials/your-first-project.md)
|
|
65
67
|
- [Onboarding an existing codebase](docs/tutorials/onboarding-an-existing-codebase.md)
|
|
@@ -270,30 +270,67 @@ If user selects 1 or 2: spawn continuation agent (with any additional context pr
|
|
|
270
270
|
|
|
271
271
|
If user selects 3: proceed to Step 4 with fix = "not applied".
|
|
272
272
|
|
|
273
|
+
### 3f. FIX REJECTED BY GUARDRAIL
|
|
274
|
+
|
|
275
|
+
When agent returns `## FIX REJECTED BY GUARDRAIL`:
|
|
276
|
+
|
|
277
|
+
Present the failing signal and evidence to the user via AskUserQuestion:
|
|
278
|
+
```
|
|
279
|
+
Fix rejected by the acceptance guardrail.
|
|
280
|
+
|
|
281
|
+
Failing signal: {failing signal}
|
|
282
|
+
Evidence: {why it failed}
|
|
283
|
+
|
|
284
|
+
Options:
|
|
285
|
+
1. Revise fix — spawn continuation agent to revise the fix so the signal passes
|
|
286
|
+
2. Accept as technical debt — record the unmet signal + justification (the fix lands without the gate passing; this is never silent)
|
|
287
|
+
3. Abandon — stop; session stays unresolved
|
|
288
|
+
```
|
|
289
|
+
|
|
290
|
+
If user selects 1: spawn continuation agent with `goal: find_and_fix` naming the failing signal to revise. Loop back to Step 3.
|
|
291
|
+
|
|
292
|
+
If user selects 2: spawn continuation agent instructed to record `guardrail_verdict: accepted_debt` + the justification in the debug file, then proceed to request_human_verification. Loop back to Step 3.
|
|
293
|
+
|
|
294
|
+
If user selects 3: proceed to Step 4 with fix = "not applied (guardrail rejected)".
|
|
295
|
+
|
|
273
296
|
## Step 4: Return Compact Summary
|
|
274
297
|
|
|
298
|
+
**Non-terminal early stop — check this FIRST.** Before returning any summary below, ask: is your own turn/context budget exhausted while the debugger (`gsd-debugger`) is still investigating — i.e. you have NOT reached `DEBUG COMPLETE`, a user-chosen `ABANDONED`, or exhausted the `INVESTIGATION INCONCLUSIVE` options? If so, do NOT fabricate a `DEBUG SESSION COMPLETE` or `ABANDONED` summary to fit this shape. Return the non-terminal marker instead:
|
|
299
|
+
|
|
300
|
+
```markdown
|
|
301
|
+
## CONTINUE_REQUIRED
|
|
302
|
+
|
|
303
|
+
**Session:** {debug_file_path}
|
|
304
|
+
**Status:** {status from frontmatter, e.g. investigating}
|
|
305
|
+
**Next action:** {next_action from Current Focus}
|
|
306
|
+
**Reason:** session-manager turn/context budget exhausted — investigation still in progress
|
|
307
|
+
```
|
|
308
|
+
|
|
309
|
+
`CONTINUE_REQUIRED` is distinct from both terminal shapes below AND from `## CHECKPOINT REACHED` (Step 3d): a `CHECKPOINT REACHED` is a genuine user-input/approval checkpoint that already correctly pauses via `AskUserQuestion` before looping back to Step 3 — it is not returned to the orchestrator. `CONTINUE_REQUIRED` is emitted only when no checkpoint is pending and the loop simply cannot proceed further in this turn. The orchestrator resumes by re-spawning this agent with the SAME `slug`/`debug_file_path` — the on-disk checkpoint at `.planning/debug/{slug}.md` (its `status` and `next_action`) is the source of truth for where to pick up. Never return control to the user as if the session were complete when it is not.
|
|
310
|
+
|
|
275
311
|
Read the resolved (or current) debug file to extract final Resolution values.
|
|
276
312
|
|
|
277
|
-
Return compact summary:
|
|
313
|
+
Return compact summary (terminal — investigation resolved):
|
|
278
314
|
|
|
279
315
|
```markdown
|
|
280
316
|
## DEBUG SESSION COMPLETE
|
|
281
317
|
|
|
282
318
|
**Session:** {final path — resolved/ if archived, otherwise debug_file_path}
|
|
283
|
-
**Root Cause:** {one sentence from Resolution.root_cause
|
|
319
|
+
**Root Cause:** {one sentence, or a '; '-joined list when the AND-gate identified multiple contributing causes, from Resolution.root_cause; or "not determined"}
|
|
284
320
|
**Fix:** {one sentence from Resolution.fix, or "not applied"}
|
|
285
321
|
**Cycles:** {N} (investigation) + {M} (fix)
|
|
286
322
|
**TDD:** {yes/no}
|
|
287
323
|
**Specialist review:** {specialist_hint used, or "none"}
|
|
324
|
+
**Prevention:** {one-line from the blameless postmortem — "why not caught: <gate, or 'none (no gate existed for this class)'>; guard: <artifact>"}
|
|
288
325
|
```
|
|
289
326
|
|
|
290
|
-
If the session was abandoned by user choice, return:
|
|
327
|
+
If the session was abandoned by user choice, return (terminal — user stopped):
|
|
291
328
|
|
|
292
329
|
```markdown
|
|
293
330
|
## DEBUG SESSION COMPLETE
|
|
294
331
|
|
|
295
332
|
**Session:** {debug_file_path}
|
|
296
|
-
**Root Cause:** {one sentence if found, or "not determined"}
|
|
333
|
+
**Root Cause:** {one sentence if found (or a '; '-joined list if the AND-gate identified multiple contributing causes), or "not determined"}
|
|
297
334
|
**Fix:** not applied
|
|
298
335
|
**Cycles:** {N}
|
|
299
336
|
**TDD:** {yes/no}
|
|
@@ -311,5 +348,6 @@ If the session was abandoned by user choice, return:
|
|
|
311
348
|
- [ ] Specialist dispatch executed when specialist_dispatch_enabled and hint maps to a skill
|
|
312
349
|
- [ ] TDD gate applied when tdd_mode=true and ROOT CAUSE FOUND
|
|
313
350
|
- [ ] Loop continues until DEBUG COMPLETE, ABANDONED, or user stops
|
|
351
|
+
- [ ] Non-terminal `CONTINUE_REQUIRED` (not a fabricated terminal summary) returned when the manager's own turn/context budget is exhausted mid-investigation
|
|
314
352
|
- [ ] Compact summary returned (at most 2K tokens)
|
|
315
353
|
</success_criteria>
|
package/agents/gsd-debugger.md
CHANGED
|
@@ -253,6 +253,10 @@ reasoning_checkpoint:
|
|
|
253
253
|
falsification_test: "[what specific observation would prove this hypothesis wrong]"
|
|
254
254
|
fix_rationale: "[why the proposed fix addresses the root cause — not just the symptom]"
|
|
255
255
|
blind_spots: "[what you haven't tested that could invalidate this hypothesis]"
|
|
256
|
+
candidate_causes:
|
|
257
|
+
- "[cause in category: code|config|environment|data]"
|
|
258
|
+
- "[cause in a DIFFERENT category — single-category is not a branch]"
|
|
259
|
+
and_gate: "[could this failure require >1 contributing condition simultaneously? yes/no + why — see RCA branching]"
|
|
256
260
|
```
|
|
257
261
|
|
|
258
262
|
**Check before proceeding:**
|
|
@@ -260,8 +264,9 @@ reasoning_checkpoint:
|
|
|
260
264
|
- Is the confirming evidence direct observation, not inference?
|
|
261
265
|
- Does the fix address the root cause or a symptom?
|
|
262
266
|
- Have you documented your blind spots honestly?
|
|
267
|
+
- **Did you branch across ≥2 categories and answer the AND-gate?** (Single-cause is fine when the AND-gate is no — but you must have checked.)
|
|
263
268
|
|
|
264
|
-
If you cannot fill all
|
|
269
|
+
If you cannot fill all seven fields with specific, concrete answers — you do not have a confirmed root cause yet. Return to investigation_loop.
|
|
265
270
|
|
|
266
271
|
## Minimal Reproduction
|
|
267
272
|
|
|
@@ -274,6 +279,7 @@ If you cannot fill all five fields with specific, concrete answers — you do no
|
|
|
274
279
|
3. Test: Does it still reproduce? YES = keep removed. NO = put back.
|
|
275
280
|
4. Repeat until bare minimum
|
|
276
281
|
5. Bug is now obvious in stripped-down code
|
|
282
|
+
6. **Shrinking (input-space bugs)** — when the bug triggers on a class of inputs, wrap it in a property (fast-check for JS/TS, Hypothesis for Python) and let the shrinker auto-minimize the counterexample; store the **minimized** input as the regression seed. See `gsd-core/references/debugger-repro-hardening.md`.
|
|
277
283
|
|
|
278
284
|
**Example:**
|
|
279
285
|
```jsx
|
|
@@ -443,18 +449,21 @@ MISMATCH: Checker looks in wrong directory → hooks "not found" → reported as
|
|
|
443
449
|
|
|
444
450
|
**The discipline:** Never assume a constructed path is correct. Resolve it to its actual value and verify the other side agrees. When two systems share a resource (file, directory, key), trace the full path in both.
|
|
445
451
|
|
|
446
|
-
## Technique Selection
|
|
452
|
+
## Technique Selection (routed by bug class)
|
|
447
453
|
|
|
448
|
-
|
|
449
|
-
|
|
450
|
-
|
|
451
|
-
|
|
452
|
-
|
|
453
|
-
|
|
|
454
|
-
|
|
455
|
-
|
|
|
456
|
-
|
|
|
457
|
-
|
|
|
454
|
+
Classify the failure first (Phase 1.75), then route by class — not by ad-hoc
|
|
455
|
+
situation:
|
|
456
|
+
|
|
457
|
+
@~/.claude/gsd-core/references/debugger-bug-taxonomy.md
|
|
458
|
+
|
|
459
|
+
| bug_class | Route to | Revoke if already run |
|
|
460
|
+
|---|---|---|
|
|
461
|
+
| Bohrbug | deterministic reproduction → SBFL (Phase 1.25) → git bisect → binary search | — |
|
|
462
|
+
| Heisenbug / Mandelbug | record-replay (`rr`) → stability-stress → statistical sampling | SBFL — Phase 1.25 runs before classification; if it ran, mark its Evidence entry revoked (flaky spectrum poisons the ranking) |
|
|
463
|
+
| Concurrency | atomicity / order / deadlock checklist (see reference) FIRST | — |
|
|
464
|
+
| General (any class) | Binary search, Working backwards, Differential, Delta debugging, Comment-out-everything, Follow-the-indirection, Rubber duck, Observability first (always, before changes) | — |
|
|
465
|
+
|
|
466
|
+
The class rows pick the first move; the General lane holds situation-cued techniques that apply to any class. When the situation table and the class route disagree, the class route wins.
|
|
458
467
|
|
|
459
468
|
## Combining Techniques
|
|
460
469
|
|
|
@@ -590,6 +599,13 @@ function processUserData(user) {
|
|
|
590
599
|
// 5. Test is now regression protection forever
|
|
591
600
|
```
|
|
592
601
|
|
|
602
|
+
**Harden the regression test (so the Phase 1A mutation guardrail bites):**
|
|
603
|
+
|
|
604
|
+
@~/.claude/gsd-core/references/debugger-repro-hardening.md
|
|
605
|
+
|
|
606
|
+
- **Classify the oracle** before writing the assertion — `specified` / `derived` (contract/model) / `metamorphic` / `implicit` (crash, weakest). Record it under `Resolution.oracle_type`. Never default to implicit silently.
|
|
607
|
+
- **Add boundary neighbors** around the fixed defect's equivalence class — off-by-one (N±1), min/max (0/length), empty/singleton — the single reported value misses the adjacent off-by-one.
|
|
608
|
+
|
|
593
609
|
## Verification Checklist
|
|
594
610
|
|
|
595
611
|
```markdown
|
|
@@ -788,9 +804,11 @@ Each resolved session appends one entry:
|
|
|
788
804
|
## {slug} — {one-line description}
|
|
789
805
|
- **Date:** {ISO date}
|
|
790
806
|
- **Error patterns:** {comma-separated keywords extracted from symptoms.errors and symptoms.actual}
|
|
791
|
-
- **Root cause:** {from Resolution.root_cause}
|
|
807
|
+
- **Root cause(s):** {from Resolution.root_cause — one cause, or a '; '-joined list when the AND-gate fired}
|
|
792
808
|
- **Fix:** {from Resolution.fix}
|
|
793
809
|
- **Files changed:** {from Resolution.files_changed}
|
|
810
|
+
- **Why not caught:** {which existing gate (test/typecheck/lint/review/verify/build) should have caught it — or "no gate existed for this class"}
|
|
811
|
+
- **Recurrence guard:** {the concrete artifact preventing this class from returning — regression test (path:name) / assertion / lint rule / type refinement / config-default change / KB pattern}
|
|
794
812
|
---
|
|
795
813
|
```
|
|
796
814
|
|
|
@@ -804,9 +822,11 @@ At the **end of `archive_session`**, after the session file is moved to `resolve
|
|
|
804
822
|
|
|
805
823
|
## Matching Logic
|
|
806
824
|
|
|
807
|
-
|
|
825
|
+
**Semantic-first, keyword-fallback.** Query MemPalace with the current symptoms and surface the top-k meaning-similar prior resolutions — this catches same-root-cause/different-wording cases keyword overlap misses. Fall back to keyword overlap on `knowledge-base.md` when MemPalace is absent. See:
|
|
826
|
+
|
|
827
|
+
@~/.claude/gsd-core/references/debugger-semantic-recall.md
|
|
808
828
|
|
|
809
|
-
**Important:** A match is a **hypothesis candidate**, not a confirmed diagnosis
|
|
829
|
+
**Important:** A match is a **hypothesis candidate**, not a confirmed diagnosis — surface it in Current Focus and test it first; do not skip other hypotheses or assume correctness.
|
|
810
830
|
|
|
811
831
|
</knowledge_base_protocol>
|
|
812
832
|
|
|
@@ -966,12 +986,10 @@ At investigation decision points, apply structured reasoning:
|
|
|
966
986
|
**Autonomous investigation. Update file continuously.**
|
|
967
987
|
|
|
968
988
|
**Phase 0: Check knowledge base**
|
|
969
|
-
-
|
|
970
|
-
- Extract keywords from `Symptoms.errors` and `Symptoms.actual` (nouns, error substrings, identifiers)
|
|
971
|
-
- Scan knowledge base entries for 2+ keyword overlap (case-insensitive)
|
|
989
|
+
- Query MemPalace semantically with the current symptoms (top-k meaning-similar prior resolutions); fall back to reading `.planning/debug/knowledge-base.md` and keyword overlap when MemPalace is absent
|
|
972
990
|
- If match found:
|
|
973
991
|
- Note in Current Focus: `known_pattern_candidate: "{matched slug} — {description}"`
|
|
974
|
-
- Add to Evidence: `found: Knowledge base match on [{keywords}] → Root cause was: {root_cause}. Fix was: {fix}.`
|
|
992
|
+
- Add to Evidence: `found: Knowledge base match on [{keywords}] → Root cause was: {root_cause}. Fix was: {fix}. Why not caught: {why_not_caught}. Recurrence guard: {recurrence_guard}.` (the last two are absent on old entries — that's fine; consume them when present)
|
|
975
993
|
- Test this hypothesis FIRST in Phase 2 — but treat it as one hypothesis, not a certainty
|
|
976
994
|
- If no match: proceed normally
|
|
977
995
|
|
|
@@ -983,14 +1001,32 @@ At investigation decision points, apply structured reasoning:
|
|
|
983
1001
|
- Run app/tests to observe behavior
|
|
984
1002
|
- APPEND to Evidence after each finding
|
|
985
1003
|
|
|
1004
|
+
**Phase 1.25: Spectrum-based fault localization (optional, coverage-gated)**
|
|
1005
|
+
- When a runnable test suite with per-test coverage exists (≥1 failing AND ≥1 passing test), compute an Ochiai suspiciousness ranking and seed the top-N into Evidence before forming hypotheses — narrows the search space deterministically before LLM reasoning:
|
|
1006
|
+
|
|
1007
|
+
@~/.claude/gsd-core/references/debugger-sbfl.md
|
|
1008
|
+
|
|
1009
|
+
- Skip with a logged note when there is no test suite, no failing tests, or no per-test coverage; investigation proceeds unchanged
|
|
1010
|
+
|
|
986
1011
|
**Phase 1.5: Check common bug patterns**
|
|
987
1012
|
- Read @~/.claude/gsd-core/references/common-bug-patterns.md
|
|
988
1013
|
- Match symptoms to pattern categories using the Symptom-to-Category Quick Map
|
|
989
1014
|
- Any matching patterns become hypothesis candidates for Phase 2
|
|
990
1015
|
- If no patterns match, proceed to open-ended hypothesis formation
|
|
991
1016
|
|
|
1017
|
+
**Phase 1.75: Classify the failure**
|
|
1018
|
+
- Assign a `bug_class` — Bohrbug (deterministic) / Heisenbug-Mandelbug (transient, non-deterministic) / Concurrency — and record it in Current Focus. The class routes which investigation technique to use:
|
|
1019
|
+
|
|
1020
|
+
@~/.claude/gsd-core/references/debugger-bug-taxonomy.md
|
|
1021
|
+
|
|
1022
|
+
- Bohrbug → reproduction + SBFL + bisect; Heisenbug/Mandelbug → record-replay/stability (skip SBFL — flaky spectra poison it); Concurrency → the atomicity/order/deadlock checklist first
|
|
1023
|
+
|
|
992
1024
|
**Phase 2: Form hypothesis**
|
|
993
1025
|
- Based on evidence AND common pattern matches, form SPECIFIC, FALSIFIABLE hypothesis
|
|
1026
|
+
- **Branch, don't chain** — at hypothesis formation (so it's done before the Phase 4 commit), enumerate candidate causes across ≥2 Ishikawa categories (code / config / environment / data) and answer the AND-gate check; `root_cause` may hold a set when the AND-gate fires:
|
|
1027
|
+
|
|
1028
|
+
@~/.claude/gsd-core/references/debugger-rca-branching.md
|
|
1029
|
+
|
|
994
1030
|
- Update Current Focus with hypothesis, test, expecting, next_action
|
|
995
1031
|
|
|
996
1032
|
**Phase 3: Test hypothesis**
|
|
@@ -1043,7 +1079,7 @@ Return structured diagnosis:
|
|
|
1043
1079
|
|
|
1044
1080
|
**Debug Session:** .planning/debug/{slug}.md
|
|
1045
1081
|
|
|
1046
|
-
**Root Cause:** {from Resolution.root_cause}
|
|
1082
|
+
**Root Cause:** {from Resolution.root_cause — one cause, or a '; '-joined list when the AND-gate identified multiple contributing causes}
|
|
1047
1083
|
|
|
1048
1084
|
**Evidence Summary:**
|
|
1049
1085
|
- {key finding 1}
|
|
@@ -1083,7 +1119,7 @@ Update status to "fixing".
|
|
|
1083
1119
|
|
|
1084
1120
|
**0. Structured Reasoning Checkpoint (MANDATORY)**
|
|
1085
1121
|
- Write the `reasoning_checkpoint` block to Current Focus (see Structured Reasoning Checkpoint in investigation_techniques)
|
|
1086
|
-
- Verify
|
|
1122
|
+
- Verify every field can be filled with specific, concrete answers — including the RCA `candidate_causes` (≥2 categories) and `and_gate` fields
|
|
1087
1123
|
- If any field is vague or empty: return to investigation_loop — root cause is not confirmed
|
|
1088
1124
|
|
|
1089
1125
|
**1. Implement minimal fix**
|
|
@@ -1091,11 +1127,15 @@ Update status to "fixing".
|
|
|
1091
1127
|
- Make SMALLEST change that addresses root cause
|
|
1092
1128
|
- Update Resolution.fix and Resolution.files_changed
|
|
1093
1129
|
|
|
1094
|
-
**2. Verify**
|
|
1130
|
+
**2. Verify (Fix-Acceptance Guardrail)**
|
|
1095
1131
|
- Update status to "verifying"
|
|
1096
|
-
-
|
|
1097
|
-
|
|
1098
|
-
-
|
|
1132
|
+
- Run the multi-signal guardrail before accepting the fix:
|
|
1133
|
+
|
|
1134
|
+
@~/.claude/gsd-core/references/debugger-fix-acceptance.md
|
|
1135
|
+
|
|
1136
|
+
- Record every signal's result under `Resolution.verification` (per-signal schema in the reference)
|
|
1137
|
+
- If ANY applicable signal fails (and no documented technical-debt escape applies): return `## FIX REJECTED BY GUARDRAIL` (see structured_returns) — do NOT request human verification
|
|
1138
|
+
- If all applicable signals pass: set `guardrail_verdict: accepted`, proceed to request_human_verification
|
|
1099
1139
|
</step>
|
|
1100
1140
|
|
|
1101
1141
|
<step name="request_human_verification">
|
|
@@ -1174,9 +1214,13 @@ Then commit planning docs via CLI (respects `commit_docs` config automatically):
|
|
|
1174
1214
|
gsd_run query commit "docs: resolve debug {slug}" --files .planning/debug/resolved/{slug}.md
|
|
1175
1215
|
```
|
|
1176
1216
|
|
|
1177
|
-
**Append to knowledge base:**
|
|
1217
|
+
**Append to knowledge base (with the Prevention block):**
|
|
1218
|
+
|
|
1219
|
+
Read `.planning/debug/resolved/{slug}.md` to extract final `Resolution` values. Then produce the **Prevention block** — a blameless postmortem (branching 5-Whys per RCA, "why wasn't this caught?", and a concrete recurrence guard):
|
|
1220
|
+
|
|
1221
|
+
@~/.claude/gsd-core/references/debugger-prevention.md
|
|
1178
1222
|
|
|
1179
|
-
|
|
1223
|
+
Then append to `.planning/debug/knowledge-base.md` (create file with header if it doesn't exist):
|
|
1180
1224
|
|
|
1181
1225
|
If creating for the first time, write this header first:
|
|
1182
1226
|
```markdown
|
|
@@ -1193,9 +1237,11 @@ Then append the entry:
|
|
|
1193
1237
|
## {slug} — {one-line description of the bug}
|
|
1194
1238
|
- **Date:** {ISO date}
|
|
1195
1239
|
- **Error patterns:** {comma-separated keywords from Symptoms.errors + Symptoms.actual}
|
|
1196
|
-
- **Root cause:** {Resolution.root_cause}
|
|
1240
|
+
- **Root cause(s):** {Resolution.root_cause — joined as '; ' when multiple contributing causes were confirmed}
|
|
1197
1241
|
- **Fix:** {Resolution.fix}
|
|
1198
1242
|
- **Files changed:** {Resolution.files_changed joined as comma list}
|
|
1243
|
+
- **Why not caught:** {which existing gate (test/typecheck/lint/review/verify/build) should have caught it — or "no gate existed for this class"}
|
|
1244
|
+
- **Recurrence guard:** {concrete artifact preventing this class from returning — regression test (path:name) / assertion / lint rule / KB pattern / type refinement / config-default change}
|
|
1199
1245
|
---
|
|
1200
1246
|
|
|
1201
1247
|
```
|
|
@@ -1205,6 +1251,8 @@ Commit the knowledge base update alongside the resolved session:
|
|
|
1205
1251
|
gsd_run query commit "docs: update debug knowledge base with {slug}" --files .planning/debug/knowledge-base.md
|
|
1206
1252
|
```
|
|
1207
1253
|
|
|
1254
|
+
**Index into MemPalace (when available)** per the semantic-recall reference — the Resolution summary (not raw symptoms), redacted — so a future Phase-0 query surfaces it by meaning. Skip with a logged note when MemPalace is absent or the KB write failed; `knowledge-base.md` is the durable fallback.
|
|
1255
|
+
|
|
1208
1256
|
Report completion and offer next steps.
|
|
1209
1257
|
</step>
|
|
1210
1258
|
|
|
@@ -1298,7 +1346,7 @@ Orchestrator presents checkpoint to user, gets response, spawns fresh continuati
|
|
|
1298
1346
|
|
|
1299
1347
|
**Debug Session:** .planning/debug/{slug}.md
|
|
1300
1348
|
|
|
1301
|
-
**Root Cause:** {specific cause with evidence}
|
|
1349
|
+
**Root Cause:** {specific cause with evidence — one cause, or a '; '-joined list when the AND-gate identified multiple contributing causes}
|
|
1302
1350
|
|
|
1303
1351
|
**Evidence Summary:**
|
|
1304
1352
|
- {key finding 1}
|
|
@@ -1334,6 +1382,16 @@ Orchestrator presents checkpoint to user, gets response, spawns fresh continuati
|
|
|
1334
1382
|
|
|
1335
1383
|
Only return this after human verification confirms the fix.
|
|
1336
1384
|
|
|
1385
|
+
## FIX REJECTED BY GUARDRAIL
|
|
1386
|
+
|
|
1387
|
+
Returned when a fix-acceptance guardrail signal fails (see `@~/.claude/gsd-core/references/debugger-fix-acceptance.md`). Do **not** mark the session resolved.
|
|
1388
|
+
|
|
1389
|
+
**Debug Session:** .planning/debug/{slug}.md
|
|
1390
|
+
**Failing signal:** {signal 1–5 name}
|
|
1391
|
+
**Evidence:** {why the signal failed — e.g. "mutant at fix site survived", "deletion-only diff with no RCA justification", "bug did not return on revert"}
|
|
1392
|
+
|
|
1393
|
+
The session-manager continuation surfaces this and offers revise / accept-as-debt / abandon.
|
|
1394
|
+
|
|
1337
1395
|
## INVESTIGATION INCONCLUSIVE
|
|
1338
1396
|
|
|
1339
1397
|
```markdown
|
package/agents/gsd-executor.md
CHANGED
|
@@ -144,6 +144,10 @@ At execution decision points, apply structured reasoning:
|
|
|
144
144
|
|
|
145
145
|
For each task:
|
|
146
146
|
|
|
147
|
+
0. **Precondition check (before any other task work):** If the task carries a `<precondition>` element, evaluate that single prose line first — it names a runnable/checkable fact the task assumes (env var set, prior-phase artifact present, server responding to `/health`, `user_setup` step done). Verify with **read-only checks only** — file existence, env var presence (no value output), idempotent `GET /health`-style pings. Do NOT run commands with side effects (writes, network POSTs, secret emission) as the check; if a side-effecting check seems required, halt and surface via checkpoint instead.
|
|
148
|
+
- **Met OR absent:** continue with no visible change to execution flow. The precondition is a no-op for the rest of the task loop.
|
|
149
|
+
- **Unmet:** STOP — return a `checkpoint:human-verify` (use `checkpoint_return_format`) with `**Blocked by:** Precondition not met: <precondition text>`. Do NOT partial-commit the task. Unmet preconditions are NEVER auto-approved, even under `AUTO_CFG=true` — a missing prerequisite is not a verification step a human can rubber-stamp; it is a fact the executor cannot establish on its own. The human either satisfies the precondition (sets the env var, completes the `user_setup` step, regenerates the artifact) or reruns `/gsd:plan-phase` to restructure.
|
|
150
|
+
|
|
147
151
|
1. **If `type="auto"`:**
|
|
148
152
|
- Check for `tdd="true"` → follow TDD execution flow
|
|
149
153
|
- Execute task, apply deviation rules as needed
|
|
@@ -152,11 +156,17 @@ For each task:
|
|
|
152
156
|
- Commit (see task_commit_protocol)
|
|
153
157
|
- Track completion + commit hash for Summary
|
|
154
158
|
|
|
155
|
-
2. **If `type="
|
|
159
|
+
2. **If `type="tracer"`:** (the leading thin end-to-end slice — production-quality, never a throwaway)
|
|
160
|
+
- Execute and commit exactly like `type="auto"` (real implementation, real `<verify>`, atomic commit).
|
|
161
|
+
- **Then run the tracer feedback gate BEFORE any expansion task** — an early integration checkpoint on the proven slice:
|
|
162
|
+
- **Autonomous run (auto mode active — `AUTO_CHAIN` or `AUTO_CFG` is `"true"`, per `<auto_mode_detection>`):** re-run the tracer's `<verify>` end-to-end. If it **fails**, HALT and surface it (deviation Rule 1) — do NOT proceed to expansion tasks. Pouring more layers onto a broken foundation is exactly the failure this gate prevents. If it passes, log `⚡ Tracer verified end-to-end — expanding` and continue.
|
|
163
|
+
- **Interactive run (auto mode not active):** immediately after committing the tracer, STOP and return a `checkpoint:human-verify` for the tracer's `<verify>` (the working slice) using checkpoint_return_format, before any expansion task.
|
|
164
|
+
|
|
165
|
+
3. **If `type="checkpoint:*"`:**
|
|
156
166
|
- STOP immediately — return structured checkpoint message
|
|
157
167
|
- A fresh agent will be spawned to continue
|
|
158
168
|
|
|
159
|
-
|
|
169
|
+
4. After all tasks: run overall verification, confirm success criteria, document deviations
|
|
160
170
|
</step>
|
|
161
171
|
|
|
162
172
|
</execution_flow>
|
|
@@ -310,6 +320,8 @@ For full automation-first patterns, server lifecycle, CLI handling:
|
|
|
310
320
|
|
|
311
321
|
**Quick reference:** Users NEVER run CLI commands. Users ONLY visit URLs, click UI, evaluate visuals, provide secrets. Claude does all automation.
|
|
312
322
|
|
|
323
|
+
**Tracer feedback gate:** a `type="tracer"` task is followed by an early integration checkpoint on the proven slice (see `<execution_flow>` → `execute_tasks`) — in autonomous runs a failing tracer `<verify>` HALTS before any expansion task; in interactive runs the executor emits a `checkpoint:human-verify` for the tracer immediately after committing it.
|
|
324
|
+
|
|
313
325
|
---
|
|
314
326
|
|
|
315
327
|
**Auto-mode checkpoint behavior** (when `AUTO_CFG` is `"true"`):
|
|
@@ -660,6 +672,21 @@ Or: "None - plan executed exactly as written."
|
|
|
660
672
|
|
|
661
673
|
If any stubs exist, add a `## Known Stubs` section to the SUMMARY listing each stub with its file, line, and reason. These are tracked for the verifier to catch. Do NOT mark a plan as complete if stubs exist that prevent the plan's goal from being achieved — either wire the data or document in the plan why the stub is intentional and which future plan will resolve it.
|
|
662
674
|
|
|
675
|
+
**Broken-windows ledger (issue #1950).** For each stub, skipped test, or unrun `<verify>` recorded above, ALSO append it to the cross-phase defect register at `.planning/WINDOWS.md`. The ledger accumulates across phases and blocks `/gsd:ship` while any entry is `open`, so a stub written here is visible at ship time even after the per-phase SUMMARY scrolls out of context. Append one entry per defect:
|
|
676
|
+
|
|
677
|
+
```bash
|
|
678
|
+
gsd_run windows append \
|
|
679
|
+
--kind stub \
|
|
680
|
+
--phase "${PHASE_NUMBER}" \
|
|
681
|
+
--file "<path-relative-to-repo-root>" \
|
|
682
|
+
--line "<line-number-or-omit>" \
|
|
683
|
+
--description "<one-line description, same wording as the Known Stubs row>"
|
|
684
|
+
```
|
|
685
|
+
|
|
686
|
+
Use `--kind skipped-test` for a `t.skip(...)` / `test.todo(...)` you left behind, `--kind unrun-verify` for a `<verify>` you could not run, or `--kind deviation` for a documented plan deviation. The full kind vocabulary: `stub | todo | fixme | skipped-test | lint-warning | unmet-truth | unrun-verify | deviation`.
|
|
687
|
+
|
|
688
|
+
The ledger is **optional**: if `gsd_run windows append` returns `windows_ledger_missing` or `windows_ok` without writing, continue without error — population is best-effort and never blocks execution. Recording here is what makes the defect visible to the ship gate later; forgetting to record is the failure mode this ledger exists to prevent.
|
|
689
|
+
|
|
663
690
|
**Threat surface scan:** Before writing the SUMMARY, check if any files created/modified introduce security-relevant surface NOT in the plan's `<threat_model>` — new network endpoints, auth paths, file access patterns, or schema changes at trust boundaries. If found, add:
|
|
664
691
|
|
|
665
692
|
```markdown
|
package/agents/gsd-planner.md
CHANGED
|
@@ -66,7 +66,7 @@ The orchestrator provides user decisions in `<user_decisions>` tags from `/gsd:d
|
|
|
66
66
|
**Self-check before returning:** For each plan, verify:
|
|
67
67
|
- [ ] Every locked decision (D-01, D-02, etc.) has a task implementing it
|
|
68
68
|
- [ ] Task actions reference the decision ID they implement (e.g., "per D-03")
|
|
69
|
-
(The decision-coverage gate `check.decision-coverage-plan` reads D-NN citations from `<objective>`, `<tasks>`, `<task>`, and `<
|
|
69
|
+
(The decision-coverage gate `check.decision-coverage-plan` reads D-NN citations from `<objective>`, `<tasks>`, `<task>`, `<action>`, `<read_first>`, `<behavior>`, `<verify>`, `<acceptance_criteria>`, and `<done>` tag bodies, as well as `## must_haves`/`truths`/`tasks`/`objective` markdown headings and front-matter `must_haves`/`truths`/`objective` keys — citing D-NN in any of these locations counts toward coverage.)
|
|
70
70
|
- [ ] No task implements a deferred idea
|
|
71
71
|
- [ ] Discretion areas are handled reasonably
|
|
72
72
|
|
|
@@ -193,23 +193,21 @@ Every task has four required fields:
|
|
|
193
193
|
**Grep gate hygiene:** `grep -c` counts comments, so header prose can be self-invalidating. Use `grep -v '^#' | grep -c token`. Bare `== 0` gates on unfiltered files are forbidden.
|
|
194
194
|
|
|
195
195
|
<comment_text_discipline>
|
|
196
|
-
**Comment-text discipline (HARD GATE, #429):** A literal an acceptance criterion negative-greps for
|
|
197
|
-
|
|
198
|
-
`<!-- planner-discipline-allow: LIT -->`
|
|
199
|
-
|
|
200
|
-
Full rules + worked examples: @gsd-core/references/planner-antipatterns.md ("Comment-Text Discipline").
|
|
196
|
+
**Comment-text discipline (HARD GATE, #429):** A literal an acceptance criterion negative-greps for must NOT appear verbatim in any `<action>` body. Full rules + `<!-- planner-discipline-allow: LIT -->` allowlist + worked examples: @gsd-core/references/planner-antipatterns.md ("Comment-Text Discipline").
|
|
201
197
|
</comment_text_discipline>
|
|
202
198
|
|
|
203
199
|
<region_scoped_negative_gate>
|
|
204
|
-
**Region-scoped negative gates (WARN, #968)
|
|
205
|
-
|
|
206
|
-
**Verify-gate hygiene (#1478/#1479):** See @gsd-core/references/planner-antipatterns.md.
|
|
200
|
+
**Region-scoped negative gates (WARN, #968)** and **Verify-gate hygiene (#1478/#1479):** @gsd-core/references/planner-antipatterns.md.
|
|
207
201
|
</region_scoped_negative_gate>
|
|
208
202
|
|
|
209
203
|
**<done>:** Acceptance criteria - measurable state of completion.
|
|
210
204
|
- Good: "Valid credentials return 200 + JWT cookie, invalid credentials return 401"
|
|
211
205
|
- Bad: "Authentication is complete"
|
|
212
206
|
|
|
207
|
+
**<precondition>** (optional, one prose line): a runnable/checkable fact the task assumes that plan ordering does not guarantee — external setup (`user_setup`), a prior-phase artifact, or an env var. The executor asserts it before running the task and halts on unmet. Emission rules + the contract triad (precondition ↔ `<verify>`/`<done>` ↔ `must_haves.truths`): @~/.claude/gsd-core/references/planner-preconditions.md.
|
|
208
|
+
|
|
209
|
+
**<reversibility>** (optional): `rating="reversible|costly|one-way"` + one-line rationale for a decision this task implements. `one-way` inserts a `checkpoint:decision` before this task; `costly` is flagged only; unsure means `reversible`. Rules: @~/.claude/gsd-core/references/planner-reversibility.md
|
|
210
|
+
|
|
213
211
|
See @~/.claude/gsd-core/references/planner-guidance.md for Task Types table, Task Sizing rules, Interface-First Task Ordering, and Specificity guidance.
|
|
214
212
|
|
|
215
213
|
## TDD Detection
|
|
@@ -250,34 +248,33 @@ Exceptions where `tdd="true"` is not needed: `type="checkpoint:*"` tasks, config
|
|
|
250
248
|
|
|
251
249
|
`workflow.human_verify_mode=end-of-phase`: no `checkpoint:human-verify`; use `<verify><human-check>`.
|
|
252
250
|
|
|
253
|
-
##
|
|
254
|
-
|
|
255
|
-
**When `MVP_MODE` is enabled (passed by the plan-phase orchestrator):** Decompose tasks as **vertical feature slices**, not horizontal layers. Required reading: Read `~/.claude/gsd-core/references/planner-mvp-mode.md` for the vertical-slice rules (lazy — only on MVP runs).
|
|
251
|
+
## Tracer-First Decomposition (default)
|
|
256
252
|
|
|
257
|
-
**
|
|
253
|
+
**Every phase plan LEADS with one `type="tracer"` task** — the thinnest path that touches every layer the phase will modify, wired end-to-end, carrying a real runnable `<verify>`. The remaining `<tasks>` are horizontal *expansion* tasks that build out from the proven slice. This is the default for **every** phase; it is not gated behind a flag. Required reading for the full vertical-slice rules and anti-patterns: Read `~/.claude/gsd-core/references/planner-mvp-mode.md`.
|
|
258
254
|
|
|
259
|
-
**
|
|
255
|
+
**Why tracer-first:** proving the architecture end-to-end on the agent's best early-context tokens catches an architectural dead-end after one commit instead of after ten already-committed layers.
|
|
260
256
|
|
|
261
|
-
|
|
257
|
+
**A tracer is production-quality, not a prototype.** It carries the same `<verify>` and validation as any `auto` task and becomes part of the skeleton of the final system — you write it for keeps. Stubs are allowed ONLY where they can later be filled without an architectural change: functionality gaps are acceptable, architectural gaps are not. (Glossary: `tracer bullet` vs `prototype` in `CONTEXT.md` — GSD ships tracers, never prototypes.)
|
|
262
258
|
|
|
263
|
-
|
|
264
|
-
## Phase Goal
|
|
259
|
+
**Tracer task shape:**
|
|
265
260
|
|
|
266
|
-
|
|
267
|
-
|
|
261
|
+
```xml
|
|
262
|
+
<task type="tracer">
|
|
263
|
+
<name>End-to-end "[capability]" — one path only</name>
|
|
264
|
+
<files>[one file per layer the phase touches]</files>
|
|
265
|
+
<action>Wire ONE entry point through every layer to the far end of the stack. No other call sites, no batching. Real error handling on the single path.</action>
|
|
266
|
+
<verify>[a real, runnable END-TO-END check of the one path — not a per-layer unit test]</verify>
|
|
267
|
+
<done>The single happy path works end-to-end and is committed.</done>
|
|
268
|
+
</task>
|
|
269
|
+
```
|
|
268
270
|
|
|
269
|
-
|
|
270
|
-
- All three slots required. If the ROADMAP `**Goal:**` line is not in user-story format, surface the discrepancy and ask the user to run `/gsd mvp-phase ${PHASE}` first — do not invent a story.
|
|
271
|
-
- Bold the three keywords (`**As a**`, `**I want to**`, `**so that**`) when emitting to PLAN.md. The ROADMAP form does not use bolded keywords; the PLAN form does.
|
|
272
|
-
2. First task: failing end-to-end test for the happy path.
|
|
273
|
-
3. Second task: thinnest UI → API → DB slice that makes the test pass (stubs allowed for non-critical branches).
|
|
274
|
-
4. Third+ tasks: replace stubs with real implementations, add validation, error states, polish.
|
|
271
|
+
**Core rule (expansion tasks):** after each task a real user can do something they could not before. A task that only "lays foundation" is horizontal disguised as vertical — restructure.
|
|
275
272
|
|
|
276
|
-
|
|
273
|
+
**`--no-tracer` (`TRACER_MODE=false`):** opt out of tracer-first and decompose into horizontal layers (the legacy default). Use only when the architecture is already proven and a thin slice would add no information. Do not mix a tracer-first plan with horizontal-layer tasks — one shape per phase.
|
|
277
274
|
|
|
278
|
-
**
|
|
275
|
+
**MVP enrichment (`MVP_MODE=true`):** layered on top of the tracer-first ordering above (MVP no longer *turns on* vertical slices — that is now the default). It adds: (1) frame the phase goal as a user story at the top of `PLAN.md`, sourced from the ROADMAP `**Goal:**` line, bolding `**As a**` / `**I want to**` / `**so that**` (Read `~/.claude/gsd-core/references/user-story-template.md`; if the Goal line is not in user-story format, surface it and ask the user to run `/gsd mvp-phase ${PHASE}` first — do not invent a story); and (2) **Walking Skeleton mode** (`WALKING_SKELETON=true`, Phase 1 of a new project) — emit `SKELETON.md` from `~/.claude/gsd-core/references/skeleton-template.md` alongside `PLAN.md`. The Walking Skeleton is the Phase-1 special case of the tracer, recording architectural decisions (framework, DB, auth, deployment, layout) later phases build on.
|
|
279
276
|
|
|
280
|
-
**
|
|
277
|
+
**TDD composition (`workflow.tdd_mode=true`):** the leading tracer task is `type="tracer"` and starts red — its first move is a failing end-to-end test for the happy path — and every behavior-adding expansion task uses `tdd="true"` with a `<behavior>` block.
|
|
281
278
|
|
|
282
279
|
See @~/.claude/gsd-core/references/planner-guidance.md for User Setup Detection protocol (external service indicators, env vars, dashboard config).
|
|
283
280
|
|
|
@@ -542,15 +539,9 @@ Do NOT use for: Deploying (use CLI), creating webhooks (use API), creating datab
|
|
|
542
539
|
|
|
543
540
|
When Claude tries CLI/API and gets auth error → creates checkpoint → user authenticates → Claude retries. Auth gates are created dynamically, NOT pre-planned.
|
|
544
541
|
|
|
545
|
-
## Writing Guidelines
|
|
546
|
-
|
|
547
|
-
**DO:** Automate everything before checkpoint, be specific ("Visit https://myapp.vercel.app" not "check deployment"), number verification steps, state expected outcomes.
|
|
542
|
+
## Writing Guidelines, Anti-Patterns, and Extended Examples
|
|
548
543
|
|
|
549
|
-
|
|
550
|
-
|
|
551
|
-
## Anti-Patterns and Extended Examples
|
|
552
|
-
|
|
553
|
-
For checkpoint anti-patterns, specificity comparison tables, context section anti-patterns, and scope reduction patterns:
|
|
544
|
+
For checkpoint writing guidelines (DO/DON'T), anti-patterns, specificity comparison tables, context section anti-patterns, and scope reduction patterns:
|
|
554
545
|
@~/.claude/gsd-core/references/planner-antipatterns.md
|
|
555
546
|
|
|
556
547
|
</checkpoints>
|
|
@@ -768,6 +759,8 @@ At decision points during plan creation, apply structured reasoning:
|
|
|
768
759
|
|
|
769
760
|
Decompose phase into tasks. **Think dependencies first, not sequence.**
|
|
770
761
|
|
|
762
|
+
**Lead with the tracer.** Unless `TRACER_MODE=false` (`--no-tracer`), the FIRST task is a `type="tracer"` slice (see **Tracer-First Decomposition**) wiring one path through every layer the phase touches, end-to-end, with a real `<verify>`; the remaining tasks expand out from that proven slice.
|
|
763
|
+
|
|
771
764
|
For each task:
|
|
772
765
|
1. What does it NEED? (files, types, APIs that must exist)
|
|
773
766
|
2. What does it CREATE? (files, types, APIs others might need)
|
package/agents/gsd-verifier.md
CHANGED
|
@@ -537,11 +537,11 @@ grep -R -n -E 'probe-[^[:space:]]+\.sh|scripts/.*/tests/probe-.*\.sh' "$PHASE_DI
|
|
|
537
537
|
|
|
538
538
|
1. Build the `PROBES` list from explicit PLAN declarations first; include conventional `scripts/*/tests/probe-*.sh` when the phase is a migration/tooling phase or the success criteria mention probes.
|
|
539
539
|
2. For every documented probe path, if the file is missing or unreadable, mark `MISSING_PROBE` and set `status: gaps_found`. Do not require the executable bit because probes run through `bash "$probe"`.
|
|
540
|
-
3. Run each probe from the built `PROBES` list
|
|
540
|
+
3. Run each probe from the built `PROBES` list from the repository root:
|
|
541
541
|
|
|
542
542
|
```bash
|
|
543
543
|
for probe in "${PROBES[@]}"; do
|
|
544
|
-
timeout
|
|
544
|
+
gsd_run run-with-timeout 30 -- bash "$probe"
|
|
545
545
|
done
|
|
546
546
|
```
|
|
547
547
|
|