@opengsd/gsd-core 1.5.0-rc.3 → 1.5.0-rc.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (127) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/agents/gsd-advisor-researcher.md +1 -1
  3. package/agents/gsd-assumptions-analyzer.md +1 -1
  4. package/agents/gsd-code-fixer.md +1 -1
  5. package/agents/gsd-code-reviewer.md +1 -1
  6. package/agents/gsd-codebase-mapper.md +1 -1
  7. package/agents/gsd-debugger.md +1 -1
  8. package/agents/gsd-doc-writer.md +1 -1
  9. package/agents/gsd-eval-auditor.md +1 -1
  10. package/agents/gsd-executor.md +1 -1
  11. package/agents/gsd-integration-checker.md +1 -1
  12. package/agents/gsd-nyquist-auditor.md +1 -0
  13. package/agents/gsd-phase-researcher.md +1 -1
  14. package/agents/gsd-plan-checker.md +1 -1
  15. package/agents/gsd-planner.md +5 -1
  16. package/agents/gsd-project-researcher.md +1 -1
  17. package/agents/gsd-research-synthesizer.md +1 -1
  18. package/agents/gsd-roadmapper.md +55 -2
  19. package/agents/gsd-security-auditor.md +1 -0
  20. package/agents/gsd-ui-auditor.md +1 -1
  21. package/agents/gsd-ui-checker.md +1 -1
  22. package/agents/gsd-ui-researcher.md +1 -1
  23. package/agents/gsd-verifier.md +58 -15
  24. package/bin/install.js +39 -58
  25. package/commands/gsd/progress.md +2 -1
  26. package/gemini-extension.json +1 -1
  27. package/gsd-core/bin/gsd-tools.cjs +205 -40
  28. package/gsd-core/bin/lib/active-workstream-store.cjs +6 -0
  29. package/gsd-core/bin/lib/agent-command-router.cjs +2 -2
  30. package/gsd-core/bin/lib/agent-install-check.cjs +143 -0
  31. package/gsd-core/bin/lib/audit-command-router.cjs +4 -4
  32. package/gsd-core/bin/lib/capability-activation.cjs +37 -10
  33. package/gsd-core/bin/lib/capability-registry.cjs +2 -0
  34. package/gsd-core/bin/lib/capability-state.cjs +177 -33
  35. package/gsd-core/bin/lib/capability-writer.cjs +355 -0
  36. package/gsd-core/bin/lib/check-command-router.cjs +50 -51
  37. package/gsd-core/bin/lib/commands.cjs +20 -2
  38. package/gsd-core/bin/lib/config-loader.cjs +3 -4
  39. package/gsd-core/bin/lib/config-schema.cjs +1 -1
  40. package/gsd-core/bin/lib/config-types.cjs +2 -1
  41. package/gsd-core/bin/lib/config.cjs +85 -26
  42. package/gsd-core/bin/lib/docs.cjs +14 -2
  43. package/gsd-core/bin/lib/edge-probe.cjs +25 -2
  44. package/gsd-core/bin/lib/frontmatter.cjs +55 -3
  45. package/gsd-core/bin/lib/gap-checker.cjs +5 -2
  46. package/gsd-core/bin/lib/git-base-branch.cjs +220 -0
  47. package/gsd-core/bin/lib/graphify-command-router.cjs +6 -8
  48. package/gsd-core/bin/lib/graphify.cjs +7 -31
  49. package/gsd-core/bin/lib/gsd2-import.cjs +2 -2
  50. package/gsd-core/bin/lib/init.cjs +58 -10
  51. package/gsd-core/bin/lib/install-profiles.cjs +55 -0
  52. package/gsd-core/bin/lib/installer-migration-report.cjs +1 -0
  53. package/gsd-core/bin/lib/intel-command-router.cjs +6 -3
  54. package/gsd-core/bin/lib/intel.cjs +28 -34
  55. package/gsd-core/bin/lib/io.cjs +2 -4
  56. package/gsd-core/bin/lib/learnings.cjs +2 -2
  57. package/gsd-core/bin/lib/loop-resolver.cjs +45 -167
  58. package/gsd-core/bin/lib/milestone.cjs +13 -5
  59. package/gsd-core/bin/lib/model-resolver.cjs +3 -4
  60. package/gsd-core/bin/lib/phase-id.cjs +2 -4
  61. package/gsd-core/bin/lib/phase-locator.cjs +3 -6
  62. package/gsd-core/bin/lib/phase.cjs +46 -19
  63. package/gsd-core/bin/lib/plan-drift-guard.cjs +117 -0
  64. package/gsd-core/bin/lib/probe-core.cjs +139 -1
  65. package/gsd-core/bin/lib/profile-output.cjs +5 -2
  66. package/gsd-core/bin/lib/prohibition-enforcement.cjs +485 -0
  67. package/gsd-core/bin/lib/roadmap-command-router.cjs +2 -2
  68. package/gsd-core/bin/lib/roadmap-parser.cjs +16 -8
  69. package/gsd-core/bin/lib/roadmap.cjs +9 -4
  70. package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +246 -0
  71. package/gsd-core/bin/lib/runtime-artifact-layout.cjs +34 -2
  72. package/gsd-core/bin/lib/state.cjs +251 -61
  73. package/gsd-core/bin/lib/task-command-router.cjs +2 -2
  74. package/gsd-core/bin/lib/template.cjs +11 -2
  75. package/gsd-core/bin/lib/uat.cjs +8 -2
  76. package/gsd-core/bin/lib/verification.cjs +8 -5
  77. package/gsd-core/bin/lib/verify.cjs +384 -8
  78. package/gsd-core/bin/lib/workstream-inventory.cjs +2 -2
  79. package/gsd-core/bin/lib/workstream.cjs +8 -2
  80. package/gsd-core/bin/lib/worktree-safety.cjs +39 -3
  81. package/gsd-core/bin/shared/config-schema.manifest.json +2 -0
  82. package/gsd-core/references/edge-probe.md +11 -0
  83. package/gsd-core/references/planner-antipatterns.md +46 -0
  84. package/gsd-core/references/planning-config.md +5 -1
  85. package/gsd-core/references/prohibition-probe-fixtures/01-streak-reminder/expected.json +14 -0
  86. package/gsd-core/references/prohibition-probe-fixtures/02-clean-utility/expected.json +4 -0
  87. package/gsd-core/references/prohibition-probe-fixtures/03-multi-prohibition/expected.json +32 -0
  88. package/gsd-core/references/prohibition-probe.md +297 -0
  89. package/gsd-core/templates/spec.md +14 -0
  90. package/gsd-core/templates/verification-report.md +16 -3
  91. package/gsd-core/workflows/complete-milestone.md +1 -5
  92. package/gsd-core/workflows/execute-phase/steps/worktree-recovery-policy.md +9 -0
  93. package/gsd-core/workflows/execute-phase.md +9 -4
  94. package/gsd-core/workflows/execute-plan.md +21 -6
  95. package/gsd-core/workflows/help/modes/full.md +4 -0
  96. package/gsd-core/workflows/next.md +50 -2
  97. package/gsd-core/workflows/pause-work.md +7 -1
  98. package/gsd-core/workflows/plan-phase.md +2 -0
  99. package/gsd-core/workflows/plan-review-convergence.md +14 -4
  100. package/gsd-core/workflows/pr-branch.md +4 -2
  101. package/gsd-core/workflows/quick.md +6 -2
  102. package/gsd-core/workflows/resume-project.md +17 -1
  103. package/gsd-core/workflows/settings-advanced.md +5 -5
  104. package/gsd-core/workflows/settings-integrations.md +5 -5
  105. package/gsd-core/workflows/settings.md +27 -1
  106. package/gsd-core/workflows/ship.md +1 -5
  107. package/gsd-core/workflows/spec-phase.md +98 -0
  108. package/gsd-core/workflows/verify-phase.md +33 -8
  109. package/hooks/dist/gsd-ensure-canonical-path.js +305 -0
  110. package/hooks/dist/gsd-statusline.js +1 -1
  111. package/hooks/dist/managed-hooks-registry.cjs +1 -0
  112. package/hooks/gsd-ensure-canonical-path.js +305 -0
  113. package/hooks/gsd-statusline.js +1 -1
  114. package/hooks/hooks.json +1 -0
  115. package/hooks/managed-hooks-registry.cjs +1 -0
  116. package/package.json +4 -4
  117. package/scripts/build-hooks.js +7 -0
  118. package/scripts/changeset/new.cjs +17 -3
  119. package/scripts/fix-slash-commands.cjs +15 -3
  120. package/scripts/gen-capability-registry.cjs +41 -2
  121. package/scripts/lint-allow-test-rule-refs.allowlist.json +2 -3
  122. package/scripts/lint-test-file-count.allowlist.json +6 -0
  123. package/scripts/mutation-matrix.cjs +108 -7
  124. package/scripts/pr-target-policy.cjs +63 -0
  125. package/scripts/research-profiles.cjs +5 -5
  126. package/scripts/run-tests.cjs +107 -6
  127. package/gsd-core/bin/lib/core.cjs +0 -345
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "gsd-core",
3
3
  "displayName": "GSD Core",
4
- "version": "1.5.0-rc.3",
4
+ "version": "1.5.0-rc.5",
5
5
  "description": "GSD Core is a meta-prompting, context engineering, and spec-driven development system for AI coding agents.",
6
6
  "author": {
7
7
  "name": "open-gsd",
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-advisor-researcher
3
3
  description: Researches a single gray area decision and returns a structured comparison table with rationale. Spawned by discuss-phase advisor mode.
4
- tools: Read, Bash, Grep, Glob, WebSearch, WebFetch, mcp__context7__*
4
+ tools: Read, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__*
5
5
  color: cyan
6
6
  ---
7
7
 
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-assumptions-analyzer
3
3
  description: Deeply analyzes codebase for a phase and returns structured assumptions with evidence. Spawned by discuss-phase assumptions mode.
4
- tools: Read, Bash, Grep, Glob
4
+ tools: Read, Bash, Grep, Glob, Skill
5
5
  color: cyan
6
6
  ---
7
7
 
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-code-fixer
3
3
  description: Applies fixes to code review findings from REVIEW.md. Reads source files, applies intelligent fixes, and commits each fix atomically. Spawned by /gsd:code-review --fix.
4
- tools: Read, Edit, Write, Bash, Grep, Glob
4
+ tools: Read, Edit, Write, Bash, Grep, Glob, Skill
5
5
  color: green
6
6
  # hooks:
7
7
  # - before_write
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-code-reviewer
3
3
  description: Reviews source files for bugs, security issues, and code quality problems. Produces structured REVIEW.md with severity-classified findings. Spawned by /gsd:code-review.
4
- tools: Read, Write, Bash, Grep, Glob
4
+ tools: Read, Write, Bash, Grep, Glob, Skill
5
5
  color: orange
6
6
  # hooks:
7
7
  # - before_write
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-codebase-mapper
3
3
  description: Explores codebase and writes structured analysis documents. Spawned by map-codebase with a focus area (tech, arch, quality, concerns). Writes documents directly to reduce orchestrator context load.
4
- tools: Read, Bash, Grep, Glob, Write
4
+ tools: Read, Bash, Grep, Glob, Write, Skill
5
5
  color: cyan
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-debugger
3
3
  description: Investigates bugs using scientific method, manages debug sessions, handles checkpoints. Spawned by /gsd:debug orchestrator.
4
- tools: Read, Write, Edit, Bash, Grep, Glob, WebSearch
4
+ tools: Read, Write, Edit, Bash, Grep, Glob, Skill, WebSearch
5
5
  color: orange
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-doc-writer
3
3
  description: Writes and updates project documentation. Spawned with a doc_assignment block specifying doc type, mode (create/update/supplement), and project context.
4
- tools: Read, Bash, Grep, Glob, Write, Edit
4
+ tools: Read, Bash, Grep, Glob, Write, Edit, Skill
5
5
  color: purple
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-eval-auditor
3
3
  description: Retroactive audit of an implemented AI phase's evaluation coverage. Checks implementation against the AI-SPEC.md evaluation plan. Scores each eval dimension as COVERED/PARTIAL/MISSING. Produces a scored EVAL-REVIEW.md with findings, gaps, and remediation guidance. Spawned by /gsd:eval-review orchestrator.
4
- tools: Read, Write, Bash, Grep, Glob
4
+ tools: Read, Write, Bash, Grep, Glob, Skill
5
5
  color: red
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-executor
3
3
  description: Executes GSD plans with atomic commits, deviation handling, checkpoint protocols, and state management. Spawned by execute-phase orchestrator or execute-plan command.
4
- tools: Read, Write, Edit, Bash, Grep, Glob, mcp__context7__*
4
+ tools: Read, Write, Edit, Bash, Grep, Glob, Skill, mcp__context7__*
5
5
  color: yellow
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-integration-checker
3
3
  description: Verifies cross-phase integration and E2E flows. Checks that phases connect properly and user workflows complete end-to-end.
4
- tools: Read, Bash, Grep, Glob
4
+ tools: Read, Bash, Grep, Glob, Skill
5
5
  color: blue
6
6
  ---
7
7
 
@@ -8,6 +8,7 @@ tools:
8
8
  - Bash
9
9
  - Glob
10
10
  - Grep
11
+ - Skill
11
12
  color: purple
12
13
  ---
13
14
 
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-phase-researcher
3
3
  description: Researches how to implement a phase before planning. Produces RESEARCH.md consumed by gsd-planner. Spawned by /gsd:plan-phase orchestrator.
4
- tools: Read, Write, Edit, Bash, Grep, Glob, WebSearch, WebFetch, mcp__context7__*, mcp__firecrawl__*, mcp__exa__*, mcp__tavily__*, mcp__ref__*, mcp__jina__*
4
+ tools: Read, Write, Edit, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__*, mcp__firecrawl__*, mcp__exa__*, mcp__tavily__*, mcp__ref__*, mcp__jina__*, mcp__perplexity__*
5
5
  color: cyan
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-plan-checker
3
3
  description: Verifies plans will achieve phase goal before execution. Goal-backward analysis of plan quality. Spawned by /gsd:plan-phase orchestrator.
4
- tools: Read, Bash, Glob, Grep
4
+ tools: Read, Bash, Glob, Grep, Skill
5
5
  color: green
6
6
  ---
7
7
 
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-planner
3
3
  description: Creates executable phase plans with task breakdown, dependency analysis, and goal-backward verification. Spawned by /gsd:plan-phase orchestrator.
4
- tools: Read, Write, Edit, Bash, Glob, Grep, WebFetch, mcp__context7__*
4
+ tools: Read, Write, Edit, Bash, Glob, Grep, Skill, WebFetch, mcp__context7__*
5
5
  color: green
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -198,6 +198,10 @@ Every task has four required fields:
198
198
  Full rules + worked examples: @gsd-core/references/planner-antipatterns.md ("Comment-Text Discipline").
199
199
  </comment_text_discipline>
200
200
 
201
+ <region_scoped_negative_gate>
202
+ **Region-scoped negative gates (WARN, #968):** Region-scope a file-wide negative grep when a sibling task needs that construct elsewhere in the same file; `validate_plan` WARNS. See: @gsd-core/references/planner-antipatterns.md ("Region-Scoped Negative Gates").
203
+ </region_scoped_negative_gate>
204
+
201
205
  **<done>:** Acceptance criteria - measurable state of completion.
202
206
  - Good: "Valid credentials return 200 + JWT cookie, invalid credentials return 401"
203
207
  - Bad: "Authentication is complete"
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-project-researcher
3
3
  description: Researches domain ecosystem before roadmap creation. Produces files in .planning/research/ consumed during roadmap creation. Spawned by /gsd:new-project or /gsd:new-milestone orchestrators.
4
- tools: Read, Write, Bash, Grep, Glob, WebSearch, WebFetch, mcp__context7__*, mcp__firecrawl__*, mcp__exa__*, mcp__tavily__*, mcp__ref__*, mcp__jina__*
4
+ tools: Read, Write, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__*, mcp__firecrawl__*, mcp__exa__*, mcp__tavily__*, mcp__ref__*, mcp__jina__*, mcp__perplexity__*
5
5
  color: cyan
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-research-synthesizer
3
3
  description: Synthesizes research outputs from parallel researcher agents into SUMMARY.md. Spawned by /gsd:new-project after 4 researcher agents complete.
4
- tools: Read, Write, Bash
4
+ tools: Read, Write, Bash, Skill
5
5
  color: purple
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-roadmapper
3
3
  description: Creates project roadmaps with phase breakdown, requirement mapping, success criteria derivation, and coverage validation. Spawned by /gsd:new-project orchestrator.
4
- tools: Read, Write, Bash, Glob, Grep
4
+ tools: Read, Write, Bash, Glob, Grep, Skill
5
5
  color: purple
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -209,6 +209,23 @@ Track coverage as you go.
209
209
  - New milestone: Start at 1
210
210
  - Continuing milestone: Check existing phases, start at last + 1
211
211
 
212
+ ## Phase ID Convention
213
+
214
+ Read `phase_id_convention` from config.json. This setting controls how phase headers and
215
+ checklist entries are formatted throughout the generated ROADMAP.md.
216
+
217
+ | Convention | Summary checklist form | Detail header form |
218
+ |---|---|---|
219
+ | `sequential` (default) | `- [ ] **Phase 1: Name**` | `### Phase 1: Name` |
220
+ | `milestone-prefixed` | `- [ ] **Phase 1-01: Name**` | `### Phase 1-01: Name` |
221
+
222
+ When `phase_id_convention` is absent or set to `"sequential"`, use plain sequential phase IDs
223
+ (e.g. `Phase 1`, `Phase 2`). When set to `"milestone-prefixed"`, prefix each phase ID with the
224
+ current milestone number and a two-digit phase index within that milestone
225
+ (e.g. `Phase 1-01`, `Phase 1-02`, `Phase 2-01`). The milestone number comes from the project's
226
+ active milestone context (default: `1` for new projects). This ensures downstream tools that
227
+ parse `### Phase N-NN:` headers for milestone-scoped workflows receive correctly prefixed IDs.
228
+
212
229
  ## Granularity Calibration
213
230
 
214
231
  Read granularity from config.json. Granularity controls compression tolerance.
@@ -310,14 +327,30 @@ After roadmap creation, REQUIREMENTS.md gets updated with phase mappings:
310
327
 
311
328
  ### 1. Summary Checklist (under `## Phases`)
312
329
 
330
+ Use the form matching `phase_id_convention` from config.
331
+
332
+ **Sequential (default — when absent or `"sequential"`):**
333
+
313
334
  ```markdown
314
335
  - [ ] **Phase 1: Name** - One-line description
315
336
  - [ ] **Phase 2: Name** - One-line description
316
337
  - [ ] **Phase 3: Name** - One-line description
317
338
  ```
318
339
 
340
+ **Milestone-prefixed (when `phase_id_convention: "milestone-prefixed"`):**
341
+
342
+ ```markdown
343
+ - [ ] **Phase 1-01: Name** - One-line description
344
+ - [ ] **Phase 1-02: Name** - One-line description
345
+ - [ ] **Phase 1-03: Name** - One-line description
346
+ ```
347
+
319
348
  ### 2. Detail Sections (under `## Phase Details`)
320
349
 
350
+ Use the header form matching `phase_id_convention` from config.
351
+
352
+ **Sequential (default):**
353
+
321
354
  ```markdown
322
355
  ### Phase 1: Name
323
356
  **Goal**: What this phase delivers
@@ -334,7 +367,25 @@ After roadmap creation, REQUIREMENTS.md gets updated with phase mappings:
334
367
  ...
335
368
  ```
336
369
 
337
- **The `### Phase X:` headers are parsed by downstream tools.** If you only write the summary checklist, phase lookups will fail.
370
+ **Milestone-prefixed (when `phase_id_convention: "milestone-prefixed"`):**
371
+
372
+ ```markdown
373
+ ### Phase 1-01: Name
374
+ **Goal**: What this phase delivers
375
+ **Depends on**: Nothing (first phase)
376
+ **Requirements**: REQ-01, REQ-02
377
+ **Success Criteria** (what must be TRUE):
378
+ 1. Observable behavior from user perspective
379
+ 2. Observable behavior from user perspective
380
+ **Plans**: TBD
381
+
382
+ ### Phase 1-02: Name
383
+ **Goal**: What this phase delivers
384
+ **Depends on**: Phase 1-01
385
+ ...
386
+ ```
387
+
388
+ **The `### Phase X:` headers are parsed by downstream tools.** If you only write the summary checklist, phase lookups will fail. Use the correct form for the configured convention so downstream parsing succeeds.
338
389
 
339
390
  ### UI Phase Detection
340
391
 
@@ -476,6 +527,8 @@ Apply phase identification methodology:
476
527
  2. Identify dependencies between groups
477
528
  3. Create phases that complete coherent capabilities
478
529
  4. Check granularity setting for compression guidance
530
+ 5. Read `phase_id_convention` from config (`sequential` or `milestone-prefixed`); apply the
531
+ matching header and checklist form throughout all output sections
479
532
 
480
533
  ## Step 5: Derive Success Criteria
481
534
 
@@ -8,6 +8,7 @@ tools:
8
8
  - Bash
9
9
  - Glob
10
10
  - Grep
11
+ - Skill
11
12
  color: red
12
13
  ---
13
14
 
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-ui-auditor
3
3
  description: Retroactive 6-pillar visual audit of implemented frontend code. Produces scored UI-REVIEW.md. Spawned by /gsd:ui-review orchestrator.
4
- tools: Read, Write, Bash, Grep, Glob
4
+ tools: Read, Write, Bash, Grep, Glob, Skill
5
5
  color: pink
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-ui-checker
3
3
  description: Validates UI-SPEC.md design contracts against 6 quality dimensions. Produces BLOCK/FLAG/PASS verdicts. Spawned by /gsd:ui-phase orchestrator.
4
- tools: Read, Bash, Glob, Grep
4
+ tools: Read, Bash, Glob, Grep, Skill
5
5
  color: cyan
6
6
  ---
7
7
 
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-ui-researcher
3
3
  description: Produces UI-SPEC.md design contract for frontend phases. Reads upstream artifacts, detects design system state, asks only unanswered questions. Spawned by /gsd:ui-phase orchestrator.
4
- tools: Read, Write, Edit, Bash, Grep, Glob, WebSearch, WebFetch, mcp__context7__*, mcp__firecrawl__*, mcp__exa__*, mcp__tavily__*, mcp__ref__*, mcp__jina__*
4
+ tools: Read, Write, Edit, Bash, Grep, Glob, Skill, WebSearch, WebFetch, mcp__context7__*, mcp__firecrawl__*, mcp__exa__*, mcp__tavily__*, mcp__ref__*, mcp__jina__*
5
5
  color: purple
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-verifier
3
3
  description: Verifies phase goal achievement through goal-backward analysis. Checks codebase delivers what phase promised, not just that tasks completed. Creates VERIFICATION.md report.
4
- tools: Read, Write, Bash, Grep, Glob
4
+ tools: Read, Write, Bash, Grep, Glob, Skill
5
5
  color: green
6
6
  # hooks:
7
7
  # PostToolUse:
@@ -85,7 +85,7 @@ cat "$PHASE_DIR"/*-VERIFICATION.md 2>/dev/null
85
85
  **If previous verification exists with `gaps:` section → RE-VERIFICATION MODE:**
86
86
 
87
87
  1. Parse previous VERIFICATION.md frontmatter
88
- 2. Extract `must_haves` (truths, artifacts, key_links)
88
+ 2. Extract `must_haves` (truths, artifacts, key_links, prohibitions)
89
89
  3. Extract `gaps` (items that failed)
90
90
  4. Set `is_re_verification = true`
91
91
  5. **Skip to Step 3** with optimization:
@@ -140,8 +140,19 @@ must_haves:
140
140
  - from: "src/components/Chat.tsx"
141
141
  to: "src/app/api/chat/route.ts"
142
142
  via: "fetch in useEffect — calls /api/chat endpoint"
143
+ prohibitions:
144
+ - statement: "MUST NOT store raw SSN in plaintext"
145
+ status: "resolved"
146
+ verification: "judgment"
143
147
  ```
144
148
 
149
+ **Also extract `must_haves.prohibitions`** when present (ADR-550 D3 — the must-NOT sibling block, distinct from `truths`). Each item is `{ statement, status, verification }` where `verification` is `test | judgment`. These are NEGATIVE checks: a verified prohibition means the must-NOT did NOT happen. Route them by verification tier in the verdict assembly (ADR-550 D4, the "B-with-guard" 2026-06-12 maintainer decision):
150
+
151
+ - **judgment-tier prohibitions → mode-dependent soft-gate.** Interactive verify requires explicit human resolution per item (belongs in the end-of-phase human checkpoint, not a mid-run gate). Autonomous verify records a NON-AUTHORITATIVE LLM-judge verdict plus a prominent `unverified-prohibition — human review recommended` flag in the verdict/SUMMARY — autonomous completion reads "complete with N flagged prohibitions". NEVER a silent pass; NEVER a hard halt of an AFK run.
152
+ - **test-tier prohibitions → FAIL CLOSED (accept-and-flag, not reject-at-parse).** Accept the `verification: test` value (the SPEC↔must_haves.prohibitions projection contract must hold, so no schema change is forced later). But a well-formed test-tier item that reaches verify with NO wired enforcement is treated as UNVERIFIED — flagged exactly like an unresolved judgment item, NEVER green. The deterministic fail-closed default is `dispositionForProhibition()` in probe-core (status `unverified`, `flagged: true` when `enforcementEvidence` is empty). Do NOT wire a real fail-first negative-test hard gate here — that enforcement MECHANISM defers to a follow-up PR (it needs a real test-tier consumer to `regression-must-fail-first` against; #644's corpus is entirely judgment-tier).
153
+
154
+ A flagged prohibition counts as a human-verification item (status `human_needed`) or a gap (status `gaps_found`) per the existing decision tree — it must never be silently absorbed into a `passed` verdict.
155
+
145
156
  **Step 2c: Merge must-haves**
146
157
 
147
158
  Combine all sources into a single must-haves list:
@@ -169,21 +180,28 @@ For each truth, determine if codebase enables it.
169
180
 
170
181
  **Verification status:**
171
182
 
172
- - ✓ VERIFIED: All supporting artifacts pass all checks
183
+ - ✓ VERIFIED: All supporting artifacts pass all checks — and, for a behavior-dependent truth, a behavioral test exercises the asserted behavior (see below)
184
+ - ⚠️ PRESENT_BEHAVIOR_UNVERIFIED: Supporting artifacts are present and wired, but the truth asserts runtime behavior that no test exercises — present, not behaviorally proven. Routes to human verification (Step 8) and does NOT count toward the verified score (Step 9).
173
185
  - ✗ FAILED: One or more artifacts missing, stub, or unwired
174
186
  - ? UNCERTAIN: Can't verify programmatically (needs human)
175
187
 
188
+ **Behavior-dependent truths.** A truth is *behavior-dependent* when its correctness hinges on runtime behavior grep/presence checks cannot see — a **state transition** or a **cancellation / cleanup / ordering invariant** (e.g. "cancels the in-flight task and bumps the generation counter", "resets the busy flag on abort", "rolls back on failure"). For these, symbol presence + wiring is *necessary but not sufficient*: the code can be present and wired yet still leak state on the very path the invariant covers.
189
+
176
190
  For each truth:
177
191
 
178
192
  1. Identify supporting artifacts
179
193
  2. Check artifact status (Step 4)
180
194
  3. Check wiring status (Step 5)
181
- 4. **Before marking FAIL:** Check for override (Step 3b)
182
- 5. Determine truth status
195
+ 4. **Before marking FAIL or PRESENT_BEHAVIOR_UNVERIFIED:** Check for override (Step 3b)
196
+ 5. **Classify behavior-dependence.** If the truth asserts a state transition or a cancellation/cleanup/ordering invariant, its status cannot be VERIFIED on presence alone:
197
+ - A pre-existing test exercises the transition/invariant and passes (confirm via Step 7b's single-named-test path) → ✓ VERIFIED.
198
+ - No such test exists, or it can't run without a server/state mutation → ⚠️ PRESENT_BEHAVIOR_UNVERIFIED. Emit a human-verification item (Step 8) and do not count it toward the verified score (Step 9).
199
+ - An accepted override (Step 3b) carries the truth as PASSED (override), exactly as it does for a FAILED truth.
200
+ 6. Determine truth status
183
201
 
184
202
  ## Step 3b: Check Verification Overrides
185
203
 
186
- Before marking any must-have as FAILED, check the VERIFICATION.md frontmatter for an `overrides:` entry that matches this must-have.
204
+ Before marking any must-have as FAILED or ⚠️ PRESENT_BEHAVIOR_UNVERIFIED, check the VERIFICATION.md frontmatter for an `overrides:` entry that matches this must-have.
187
205
 
188
206
  **Override check procedure:**
189
207
 
@@ -193,12 +211,12 @@ Before marking any must-have as FAILED, check the VERIFICATION.md frontmatter fo
193
211
  4. Key technical terms (file paths, component names, API endpoints) have higher weight
194
212
 
195
213
  **If override found:**
196
- - Mark as `PASSED (override)` instead of FAIL
214
+ - Mark as `PASSED (override)` instead of FAIL/PRESENT_BEHAVIOR_UNVERIFIED
197
215
  - Evidence: `Override: {reason} — accepted by {accepted_by} on {accepted_at}`
198
- - Count toward passing score, not failing score
216
+ - Count toward passing score (`verified_truths`), not failing score
199
217
 
200
218
  **If no override found:**
201
- - Mark as FAILED as normal
219
+ - Mark as FAILED (or ⚠️ PRESENT_BEHAVIOR_UNVERIFIED, per Step 3 step 5) as normal
202
220
  - Consider suggesting an override if the failure looks intentional (alternative implementation exists)
203
221
 
204
222
  **Suggesting overrides:** When a must-have FAILs but evidence shows an alternative implementation that achieves the same intent, include an override suggestion in the report:
@@ -450,6 +468,8 @@ Anti-pattern scanning (Step 7) checks for code smells. Behavioral spot-checks go
450
468
 
451
469
  **When to run:** For phases that produce runnable code (APIs, CLI tools, build scripts, data pipelines). Skip for documentation-only or config-only phases.
452
470
 
471
+ **Behavioral evidence for behavior-dependent truths (Step 3).** When a truth asserts a state transition or a cancellation/cleanup/ordering invariant, the single named test below is what upgrades it from ⚠️ PRESENT_BEHAVIOR_UNVERIFIED to ✓ VERIFIED. Run only the one named test that exercises the transition/invariant — never the full suite (per #25/#753). If no such test exists, leave the truth ⚠️ PRESENT_BEHAVIOR_UNVERIFIED and route it to human verification (Step 8); do not mark it VERIFIED on presence.
472
+
453
473
  **How:**
454
474
 
455
475
  1. **Identify checkable behaviors** from must-haves truths. Select 2-4 that can be tested with a single command:
@@ -537,6 +557,8 @@ done
537
557
 
538
558
  **Needs human if uncertain:** Complex wiring grep can't trace, dynamic state behavior, edge cases.
539
559
 
560
+ **Behavior-unverified truths (Step 3):** Every truth left ⚠️ PRESENT_BEHAVIOR_UNVERIFIED is recorded in the `behavior_unverified_items` frontmatter list (emitted whenever the count > 0, regardless of overall status, so it survives a gaps_found phase) and surfaces for human verification; when the overall status is human_needed it also appears in the human_verification section. Phrase each item around the invariant: what to trigger, what state must hold afterward, and why presence checks can't see it.
561
+
540
562
  **Harvest deferred items from PLAN.md (#3309 / `workflow.human_verify_mode = end-of-phase`):** Scan every PLAN file in the phase for `<verify><human-check>` blocks on `auto` tasks. These are verification items the planner deliberately deferred from `checkpoint:human-verify` to end-of-phase to avoid the executor cold-start cost. Each block has the same shape used by the planner:
541
563
 
542
564
  ```xml
@@ -568,18 +590,30 @@ Classify status using this decision tree IN ORDER (most restrictive first):
568
590
  1. IF any truth FAILED, artifact MISSING/STUB, key link NOT_WIRED, or blocker anti-pattern found:
569
591
  → **status: gaps_found**
570
592
 
571
- 2. IF Step 8 produced ANY human verification items (section is non-empty):
593
+ 2. IF Step 8 produced ANY human verification items (section is non-empty) — this includes every ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truth from Step 3:
572
594
  → **status: human_needed**
573
- (Even if all truths are VERIFIED and score is N/N — human items take priority)
595
+ (Even if all other truths are VERIFIED — human items take priority)
574
596
 
575
597
  3. IF all truths VERIFIED, all artifacts pass, all links WIRED, no blockers, AND no human verification items:
576
598
  → **status: passed**
577
599
 
578
- **passed is ONLY valid when the human verification section is empty.** If you identified items requiring human testing in Step 8, status MUST be human_needed.
600
+ **passed is ONLY valid when the human verification section is empty.** If Step 8 produced any items — including any truth left ⚠️ PRESENT_BEHAVIOR_UNVERIFIED — the status is not `passed`: it is `human_needed`, or `gaps_found` when rule 1 also fires (the ordered tree keeps gaps_found's precedence).
601
+
602
+ **A ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truth is never FAILED and never VERIFIED.** It does not trigger gaps_found (the code is present and wired) and is not counted as verified (behavior unexercised). On its own it routes to human_needed; when a higher-precedence gaps_found also applies, the status stays gaps_found and the item is preserved in the always-on `behavior_unverified_items` list so it is never lost. Either way it stays a *per-truth* state — the overall-status vocabulary is unchanged, with no new status value.
579
603
 
580
604
  > **Shared status seam**: the status vocabulary (`passed`, `gaps_found`, `human_needed`) and the per-status routing (next action and next command for each value) are owned by `src/verification.cts` via `gsd_run query verification.status`. This agent is the single emitter of the frontmatter status field; consumers (ship.md, execute-phase.md) read routing from that query instead of re-deriving it.
581
605
 
582
- **Score:** `verified_truths / total_truths`
606
+ **Score (presence- vs behavior-verified split):**
607
+
608
+ - `verified_truths` counts ✓ VERIFIED truths plus PASSED (override) truths (Step 3b). For a behavior-dependent truth, VERIFIED means a behavioral test passed, not just that symbols are present.
609
+ - ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truths are the *only* ones excluded from `verified_truths`; they are reported separately as `behavior_unverified`.
610
+
611
+ ```text
612
+ score: verified_truths / total_truths # e.g. 6/7
613
+ behavior_unverified: P # truths present + wired but behavior not exercised
614
+ ```
615
+
616
+ A headline N/N therefore certifies that every behavior-dependent truth had behavioral evidence — a clean score can no longer be reached on symbol presence alone.
583
617
 
584
618
  ## Step 9b: Filter Deferred Items
585
619
 
@@ -688,6 +722,7 @@ phase: XX-name
688
722
  verified: YYYY-MM-DDTHH:MM:SSZ
689
723
  status: passed | gaps_found | human_needed
690
724
  score: N/M must-haves verified
725
+ behavior_unverified: 0 # Count of ⚠️ PRESENT_BEHAVIOR_UNVERIFIED truths (present + wired, behavior not exercised); each is detailed in behavior_unverified_items below (and in human_verification when status is human_needed)
691
726
  overrides_applied: 0 # Count of PASSED (override) items included in score
692
727
  overrides: # Only if overrides exist — carried forward or newly added
693
728
  - must_have: "Must-have text that was overridden"
@@ -714,6 +749,11 @@ deferred: # Only if deferred items exist (Step 9b)
714
749
  - truth: "Observable truth addressed in a later phase"
715
750
  addressed_in: "Phase N"
716
751
  evidence: "Matching goal or success criteria text"
752
+ behavior_unverified_items: # Only if behavior_unverified > 0 — emitted regardless of overall status, so these survive a gaps_found phase
753
+ - truth: "Observable truth whose state transition or cancellation/cleanup/ordering invariant no test exercises"
754
+ test: "What to trigger"
755
+ expected: "What state must hold afterward"
756
+ why_human: "Why presence checks can't see it"
717
757
  human_verification: # Only if status: human_needed
718
758
  - test: "What to do"
719
759
  expected: "What should happen"
@@ -735,8 +775,9 @@ human_verification: # Only if status: human_needed
735
775
  | --- | ------- | ---------- | -------------- |
736
776
  | 1 | {truth} | ✓ VERIFIED | {evidence} |
737
777
  | 2 | {truth} | ✗ FAILED | {what's wrong} |
778
+ | 3 | {truth} | ⚠️ PRESENT_BEHAVIOR_UNVERIFIED | {present + wired; no test exercises the transition/invariant — see Human Verification} |
738
779
 
739
- **Score:** {N}/{M} truths verified
780
+ **Score:** {N}/{M} truths verified ({P} present, behavior-unverified)
740
781
 
741
782
  ### Deferred Items
742
783
 
@@ -823,7 +864,7 @@ Structured gaps in VERIFICATION.md frontmatter for `/gsd:plan-phase --gaps`.
823
864
 
824
865
  {If human_needed:}
825
866
  ### Human Verification Required
826
- {N} items need human testing:
867
+ {N} items need human testing (including {P} present-but-behavior-unverified truths — code wired, transition/invariant not exercised by a test):
827
868
  1. **{Test name}** — {what to do}
828
869
  - Expected: {what should happen}
829
870
 
@@ -846,6 +887,8 @@ Automated checks passed. Awaiting human verification.
846
887
 
847
888
  **Keep verification fast.** Use grep/file checks, not running the app.
848
889
 
890
+ **Presence is not behavior.** Grep/file checks prove a symbol is present and wired — they do not prove a state transition or a cancellation/cleanup/ordering invariant holds at runtime. For a behavior-dependent truth, require a passing behavioral test (Step 7b's single named test) or mark it ⚠️ PRESENT_BEHAVIOR_UNVERIFIED and route to human verification. Never let symbol presence alone produce a VERIFIED on a behavior-dependent truth.
891
+
849
892
  **DO NOT commit.** Leave committing to the orchestrator.
850
893
 
851
894
  </critical_rules>