@opengsd/gsd-core 1.6.1 → 1.7.0-rc.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (119) hide show
  1. package/.claude-plugin/marketplace.json +20 -0
  2. package/.claude-plugin/plugin.json +1 -1
  3. package/.opencode/plugins/gsd-core.js +711 -0
  4. package/agents/gsd-advisor-researcher.md +2 -0
  5. package/agents/gsd-ai-researcher.md +1 -1
  6. package/agents/gsd-assumptions-analyzer.md +2 -0
  7. package/agents/gsd-code-fixer.md +2 -0
  8. package/agents/gsd-code-reviewer.md +2 -0
  9. package/agents/gsd-codebase-mapper.md +2 -0
  10. package/agents/gsd-debugger.md +2 -0
  11. package/agents/gsd-doc-writer.md +2 -0
  12. package/agents/gsd-eval-auditor.md +2 -0
  13. package/agents/gsd-executor.md +9 -6
  14. package/agents/gsd-integration-checker.md +2 -0
  15. package/agents/gsd-nyquist-auditor.md +2 -0
  16. package/agents/gsd-phase-researcher.md +2 -0
  17. package/agents/gsd-plan-checker.md +2 -0
  18. package/agents/gsd-planner.md +2 -0
  19. package/agents/gsd-project-researcher.md +2 -0
  20. package/agents/gsd-research-synthesizer.md +2 -0
  21. package/agents/gsd-roadmapper.md +2 -0
  22. package/agents/gsd-security-auditor.md +2 -0
  23. package/agents/gsd-ui-auditor.md +2 -0
  24. package/agents/gsd-ui-checker.md +2 -0
  25. package/agents/gsd-ui-researcher.md +2 -0
  26. package/agents/gsd-verifier.md +5 -2
  27. package/bin/gsd-mcp-server.js +31 -0
  28. package/bin/install.js +411 -1146
  29. package/commands/gsd/review.md +6 -0
  30. package/gemini-extension.json +1 -1
  31. package/gsd-core/bin/gsd-tools.cjs +134 -8
  32. package/gsd-core/bin/lib/adapter-declarative.cjs +35 -0
  33. package/gsd-core/bin/lib/adapter-imperative.cjs +52 -0
  34. package/gsd-core/bin/lib/assumption-delta.cjs +231 -0
  35. package/gsd-core/bin/lib/capability-lifecycle.cjs +7 -7
  36. package/gsd-core/bin/lib/capability-loader.cjs +45 -9
  37. package/gsd-core/bin/lib/capability-lock.cjs +2 -2
  38. package/gsd-core/bin/lib/capability-registry.cjs +891 -82
  39. package/gsd-core/bin/lib/capability-source.cjs +26 -11
  40. package/gsd-core/bin/lib/capability-validator.cjs +222 -2
  41. package/gsd-core/bin/lib/cli-skew-check.cjs +44 -0
  42. package/gsd-core/bin/lib/command-aliases.cjs +8 -0
  43. package/gsd-core/bin/lib/commands.cjs +2 -1
  44. package/gsd-core/bin/lib/config.cjs +27 -0
  45. package/gsd-core/bin/lib/embedding-adapter.cjs +27 -0
  46. package/gsd-core/bin/lib/external-descriptor-trust.cjs +70 -0
  47. package/gsd-core/bin/lib/frontmatter.cjs +53 -6
  48. package/gsd-core/bin/lib/handshake-serialized.cjs +70 -0
  49. package/gsd-core/bin/lib/hook-bus.cjs +81 -0
  50. package/gsd-core/bin/lib/host-integration-sdk.cjs +53 -0
  51. package/gsd-core/bin/lib/host-integration.cjs +469 -0
  52. package/gsd-core/bin/lib/init.cjs +35 -7
  53. package/gsd-core/bin/lib/install-engine.cjs +755 -0
  54. package/gsd-core/bin/lib/install-profiles.cjs +35 -4
  55. package/gsd-core/bin/lib/installer-migrations.cjs +1 -1
  56. package/gsd-core/bin/lib/mcp-server.cjs +194 -0
  57. package/gsd-core/bin/lib/milestone.cjs +68 -40
  58. package/gsd-core/bin/lib/model-adapter.cjs +50 -0
  59. package/gsd-core/bin/lib/phase-id.cjs +18 -0
  60. package/gsd-core/bin/lib/phase.cjs +57 -90
  61. package/gsd-core/bin/lib/phases-command-router.cjs +4 -3
  62. package/gsd-core/bin/lib/planning-workspace.cjs +1 -1
  63. package/gsd-core/bin/lib/probe-core.cjs +132 -2
  64. package/gsd-core/bin/lib/review-reviewer-selection.cjs +129 -13
  65. package/gsd-core/bin/lib/roadmap-command-router.cjs +3 -2
  66. package/gsd-core/bin/lib/roadmap-parser.cjs +21 -11
  67. package/gsd-core/bin/lib/roadmap-upgrade.cjs +3 -2
  68. package/gsd-core/bin/lib/roadmap.cjs +33 -22
  69. package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +65 -9
  70. package/gsd-core/bin/lib/runtime-artifact-install-plan.cjs +54 -4
  71. package/gsd-core/bin/lib/runtime-artifact-layout.cjs +5 -2
  72. package/gsd-core/bin/lib/runtime-hooks-surface.cjs +1 -1
  73. package/gsd-core/bin/lib/runtime-name-policy.cjs +160 -30
  74. package/gsd-core/bin/lib/shell-command-projection.cjs +16 -0
  75. package/gsd-core/bin/lib/stale-bake-guard.cjs +254 -0
  76. package/gsd-core/bin/lib/state-command-router.cjs +4 -0
  77. package/gsd-core/bin/lib/state-io.cjs +55 -0
  78. package/gsd-core/bin/lib/state-transition.cjs +1603 -0
  79. package/gsd-core/bin/lib/state.cjs +327 -683
  80. package/gsd-core/bin/lib/surface.cjs +4 -1
  81. package/gsd-core/bin/lib/validate.cjs +2 -1
  82. package/gsd-core/bin/lib/verify.cjs +6 -4
  83. package/gsd-core/bin/lib/workstream-inventory-builder.cjs +12 -2
  84. package/gsd-core/bin/lib/workstream-inventory.cjs +28 -0
  85. package/gsd-core/bin/lib/workstream.cjs +4 -4
  86. package/gsd-core/bin/shared/config-schema.manifest.json +9 -0
  87. package/gsd-core/references/agent-skills-bootstrap.md +60 -0
  88. package/gsd-core/references/honest-verifier.md +105 -0
  89. package/gsd-core/references/model-profiles.md +27 -0
  90. package/gsd-core/references/reviewer-instances.md +99 -0
  91. package/gsd-core/workflows/autonomous.md +30 -32
  92. package/gsd-core/workflows/complete-milestone.md +6 -10
  93. package/gsd-core/workflows/execute-phase.md +1 -1
  94. package/gsd-core/workflows/forensics.md +3 -3
  95. package/gsd-core/workflows/help/modes/full.md +1 -1
  96. package/gsd-core/workflows/manager.md +15 -15
  97. package/gsd-core/workflows/milestone-summary.md +3 -3
  98. package/gsd-core/workflows/new-milestone.md +6 -0
  99. package/gsd-core/workflows/plan-phase/steps/closed-phase-gate.md +42 -0
  100. package/gsd-core/workflows/plan-phase/steps/prd-express-path.md +102 -0
  101. package/gsd-core/workflows/plan-phase/steps/windows-troubleshooting.md +23 -0
  102. package/gsd-core/workflows/plan-phase.md +4 -159
  103. package/gsd-core/workflows/review.md +33 -2
  104. package/gsd-core/workflows/thread.md +4 -4
  105. package/gsd-core/workflows/verify-phase.md +11 -4
  106. package/gsd-core/workflows/verify-work.md +1 -2
  107. package/hooks/dist/gsd-graphify-update.sh +7 -1
  108. package/hooks/gsd-graphify-update.sh +7 -1
  109. package/package.json +6 -4
  110. package/scripts/ci-test-scope.cjs +38 -9
  111. package/scripts/lint-allow-test-rule-refs.allowlist.json +0 -1
  112. package/scripts/lint-regression-test-names.allowlist.json +3 -0
  113. package/scripts/lint-test-file-count.allowlist.json +19 -5
  114. package/scripts/mutation-matrix.cjs +45 -3
  115. package/scripts/prompt-injection-scan.sh +8 -0
  116. package/scripts/run-tests.cjs +51 -1
  117. package/scripts/sync-manifest-versions.cjs +66 -14
  118. package/skills/gsd-review/SKILL.md +6 -0
  119. package/scripts/lint-windows-test-portability.cjs +0 -178
@@ -43,6 +43,9 @@ const runtimeArtifactLayout = require("./runtime-artifact-layout.cjs");
43
43
  const { findInstallSourceRoot } = runtimeArtifactLayout;
44
44
  // eslint-disable-next-line @typescript-eslint/no-require-imports
45
45
  const runtimeArtifactConversion = require("./runtime-artifact-conversion.cjs");
46
+ // eslint-disable-next-line @typescript-eslint/no-require-imports
47
+ const runtimeArtifactInstallPlan = require("./runtime-artifact-install-plan.cjs");
48
+ const { assertDestWithinConfigHome } = runtimeArtifactInstallPlan;
46
49
  const SURFACE_FILE_NAME = '.gsd-surface.json';
47
50
  /**
48
51
  * Read the surface state from a runtime config directory.
@@ -302,7 +305,7 @@ function applySurface(runtimeConfigDir, layout, manifest, clusterMap, registry)
302
305
  tempDirsToClean.push(rewritten);
303
306
  }
304
307
  }
305
- const dest = node_path_1.default.join(layout.configDir, kind.destSubpath);
308
+ const dest = assertDestWithinConfigHome(layout.configDir, kind.destSubpath);
306
309
  _syncGsdDir(staged, dest, kind, skillManifest);
307
310
  }
308
311
  }
@@ -93,7 +93,8 @@ function buildRoadmapPhaseVariants(roadmapContent) {
93
93
  const roadmapPhaseVariants = new Set();
94
94
  // Matches both legacy numeric (Phase 1:), decimal (Phase 2.1:), milestone-prefixed (Phase 2-01:),
95
95
  // and bracket-prefixed (### [GSD] Phase 2-01:) headings.
96
- const phasePattern = /#{2,4}\s*(?:\[[^\]]+\]\s*)?Phase\s+([\w][\w.-]*)\s*:/gi;
96
+ // #1729: `(?:\s*\([^)\n]*\))?` tolerates a pre-colon ( ) tag (literal mirror of OPTIONAL_PHASE_TAG_SOURCE).
97
+ const phasePattern = /#{2,4}\s*(?:\[[^\]]+\]\s*)?Phase\s+([\w][\w.-]*)(?:\s*\([^)\n]*\))?\s*:/gi;
97
98
  let m;
98
99
  while ((m = phasePattern.exec(roadmapContent)) !== null) {
99
100
  roadmapPhases.add(m[1]);
@@ -40,7 +40,7 @@ const configLoaderMod = require("./config-loader.cjs");
40
40
  const { loadConfig, CONFIG_DEFAULTS } = configLoaderMod;
41
41
  // eslint-disable-next-line @typescript-eslint/no-require-imports
42
42
  const phaseIdMod = require("./phase-id.cjs");
43
- const { normalizePhaseName, phaseTokenMatches, escapeRegex, getMilestoneFromPhaseId } = phaseIdMod;
43
+ const { normalizePhaseName, phaseTokenMatches, escapeRegex, getMilestoneFromPhaseId, OPTIONAL_PHASE_TAG_SOURCE } = phaseIdMod;
44
44
  // eslint-disable-next-line @typescript-eslint/no-require-imports
45
45
  const phaseLocatorMod = require("./phase-locator.cjs");
46
46
  const { findPhaseInternal } = phaseLocatorMod;
@@ -992,7 +992,8 @@ function checkMilestonePrefixMismatches(roadmapContent, { getMilestoneFromPhaseI
992
992
  }
993
993
  for (const section of sections) {
994
994
  const content = roadmapContent.slice(section.start, section.end);
995
- const phaseRx = /#{2,4}\s*(?:\[[^\]]+\]\s*)?Phase\s+([\w][\w.-]*)\s*:/gi;
995
+ // #1729: `(?:\s*\([^)\n]*\))?` tolerates a pre-colon ( ) tag (literal mirror of OPTIONAL_PHASE_TAG_SOURCE).
996
+ const phaseRx = /#{2,4}\s*(?:\[[^\]]+\]\s*)?Phase\s+([\w][\w.-]*)(?:\s*\([^)\n]*\))?\s*:/gi;
996
997
  let pm;
997
998
  while ((pm = phaseRx.exec(content)) !== null) {
998
999
  const phaseId = pm[1];
@@ -1351,7 +1352,7 @@ function cmdValidateHealth(cwd, options, raw) {
1351
1352
  stateContent.match(/Current Phase:\s*(\S+)/i);
1352
1353
  if (currentPhaseMatch) {
1353
1354
  const statePhase = currentPhaseMatch[1].replace(/^0+/, '');
1354
- const phaseCheckboxRe = new RegExp(`-\\s*\\[x\\].*Phase\\s+0*${escapeRegex(statePhase)}[:\\s]`, 'i');
1355
+ const phaseCheckboxRe = new RegExp(`-\\s*\\[x\\].*Phase\\s+0*${escapeRegex(statePhase)}${OPTIONAL_PHASE_TAG_SOURCE}[:\\s]`, 'i');
1355
1356
  if (phaseCheckboxRe.test(roadmapContentFull)) {
1356
1357
  const stateStatus = stateContent.match(/\*\*Status:\*\*\s*(.+)/i);
1357
1358
  const statusVal = stateStatus ? stateStatus[1].trim().toLowerCase() : '';
@@ -1509,7 +1510,8 @@ function cmdValidateHealth(cwd, options, raw) {
1509
1510
  if (isMarkedComplete) {
1510
1511
  const roadmapRaw = node_fs_1.default.readFileSync(roadmapPath, 'utf-8');
1511
1512
  const scopedContent = extractCurrentMilestone(roadmapRaw, cwd);
1512
- const phasePattern = /#{2,4}\s*Phase\s+(\d+[A-Z]?(?:\.\d+)*)\s*:\s*([^\n]+)/gi;
1513
+ // #1729: `(?:\s*\([^)\n]*\))?` tolerates a pre-colon ( ) tag (literal mirror of OPTIONAL_PHASE_TAG_SOURCE).
1514
+ const phasePattern = /#{2,4}\s*Phase\s+(\d+[A-Z]?(?:\.\d+)*)(?:\s*\([^)\n]*\))?\s*:\s*([^\n]+)/gi;
1513
1515
  const unstarted = [];
1514
1516
  let pm;
1515
1517
  // Non-hoisted: load-order matters (circular dep guard)
@@ -28,7 +28,7 @@ function isCompletedInventory(status) {
28
28
  return /\bmilestone\s+complete\b/.test(s) || /\barchived\b/.test(s);
29
29
  }
30
30
  function buildWorkstreamInventory(inputs) {
31
- const { name, projectDir, workstreamDir, phaseDirNames, activeWorkstreamName, phaseFilesCounts, roadmapPhaseCount, stateProjection, filesExist, } = inputs;
31
+ const { name, projectDir, workstreamDir, phaseDirNames, activeWorkstreamName, phaseFilesCounts, roadmapPhaseCount, stateProjection, filesExist, milestoneShipped, } = inputs;
32
32
  // Index counts by directory for O(1) lookup during sort/iteration
33
33
  const countsMap = new Map();
34
34
  for (const entry of phaseFilesCounts) {
@@ -56,6 +56,14 @@ function buildWorkstreamInventory(inputs) {
56
56
  summary_count: counts.summaryCount,
57
57
  });
58
58
  }
59
+ // #1913: derive status from authoritative shipped signals rather than trusting
60
+ // the mutable STATE.md `Status` field. When a shipped signal is present, the
61
+ // workstream is "milestone complete" regardless of a stale field value.
62
+ const fieldStatus = stateProjection.status;
63
+ const useDerived = milestoneShipped;
64
+ const status = useDerived ? 'milestone complete' : fieldStatus;
65
+ const status_source = useDerived ? 'derived' : 'field';
66
+ const status_conflict = useDerived && !isCompletedInventory(fieldStatus);
59
67
  return {
60
68
  name,
61
69
  path: toPosixPath(node_path_1.default.relative(projectDir, workstreamDir)),
@@ -65,7 +73,9 @@ function buildWorkstreamInventory(inputs) {
65
73
  state: filesExist.state,
66
74
  requirements: filesExist.requirements,
67
75
  },
68
- status: stateProjection.status,
76
+ status,
77
+ status_source,
78
+ status_conflict,
69
79
  current_phase: stateProjection.current_phase,
70
80
  last_activity: stateProjection.last_activity,
71
81
  phases,
@@ -63,6 +63,33 @@ function readStateProjection(statePath) {
63
63
  };
64
64
  }
65
65
  }
66
+ /**
67
+ * #1913: detect an authoritative shipped signal for a workstream so the
68
+ * inventory status is never trusted from the mutable STATE.md `Status` field
69
+ * alone. Returns true when EITHER an archived milestone snapshot is present
70
+ * under `<planningBase>/milestones/` OR the workstream ROADMAP carries a
71
+ * SHIPPED marker — both are hard to desync, unlike the hand-maintained field.
72
+ */
73
+ function workstreamMilestoneShipped(roadmapPath, planningBase) {
74
+ try {
75
+ const milestonesDir = node_path_1.default.join(planningBase, 'milestones');
76
+ for (const entry of node_fs_1.default.readdirSync(milestonesDir, { withFileTypes: true })) {
77
+ if (entry.isFile() && /-ROADMAP\.md$/i.test(entry.name))
78
+ return true;
79
+ }
80
+ }
81
+ catch {
82
+ /* no milestones archive dir */
83
+ }
84
+ try {
85
+ if (/SHIPPED/i.test(node_fs_1.default.readFileSync(roadmapPath, 'utf-8')))
86
+ return true;
87
+ }
88
+ catch {
89
+ /* no roadmap */
90
+ }
91
+ return false;
92
+ }
66
93
  function sortWorkstreamInventories(inventories, activeWorkstreamName) {
67
94
  return [...inventories].sort((a, b) => {
68
95
  const aActive = a.name === activeWorkstreamName ? 1 : 0;
@@ -99,6 +126,7 @@ function inspectWorkstream(cwd, name, options = {}) {
99
126
  state: node_fs_1.default.existsSync(p.state),
100
127
  requirements: node_fs_1.default.existsSync(p.requirements),
101
128
  },
129
+ milestoneShipped: workstreamMilestoneShipped(p.roadmap, p.planning),
102
130
  });
103
131
  }
104
132
  function listWorkstreamInventories(cwd) {
@@ -67,7 +67,7 @@ function migrateToWorkstreams(cwd, workstreamName) {
67
67
  const src = node_path_1.default.join(baseDir, item.name);
68
68
  if (node_fs_1.default.existsSync(src)) {
69
69
  const dest = node_path_1.default.join(wsDir, item.name);
70
- node_fs_1.default.renameSync(src, dest);
70
+ (0, shell_command_projection_cjs_1.retryRenameSync)(src, dest);
71
71
  filesMoved.push(item.name);
72
72
  }
73
73
  }
@@ -75,7 +75,7 @@ function migrateToWorkstreams(cwd, workstreamName) {
75
75
  catch (err) {
76
76
  for (const name of filesMoved) {
77
77
  try {
78
- node_fs_1.default.renameSync(node_path_1.default.join(wsDir, name), node_path_1.default.join(baseDir, name));
78
+ (0, shell_command_projection_cjs_1.retryRenameSync)(node_path_1.default.join(wsDir, name), node_path_1.default.join(baseDir, name));
79
79
  }
80
80
  catch { /* ignore */ }
81
81
  }
@@ -275,14 +275,14 @@ function cmdWorkstreamComplete(cwd, name, options, raw) {
275
275
  try {
276
276
  const entries = node_fs_1.default.readdirSync(wsDir, { withFileTypes: true });
277
277
  for (const entry of entries) {
278
- node_fs_1.default.renameSync(node_path_1.default.join(wsDir, entry.name), node_path_1.default.join(archivePath, entry.name));
278
+ (0, shell_command_projection_cjs_1.retryRenameSync)(node_path_1.default.join(wsDir, entry.name), node_path_1.default.join(archivePath, entry.name));
279
279
  filesMoved.push(entry.name);
280
280
  }
281
281
  }
282
282
  catch (err) {
283
283
  for (const fname of filesMoved) {
284
284
  try {
285
- node_fs_1.default.renameSync(node_path_1.default.join(archivePath, fname), node_path_1.default.join(wsDir, fname));
285
+ (0, shell_command_projection_cjs_1.retryRenameSync)(node_path_1.default.join(archivePath, fname), node_path_1.default.join(wsDir, fname));
286
286
  }
287
287
  catch { /* ignore */ }
288
288
  }
@@ -10,6 +10,10 @@
10
10
  "brave_search",
11
11
  "firecrawl",
12
12
  "exa_search",
13
+ "tavily_search",
14
+ "ref_search",
15
+ "perplexity",
16
+ "jina",
13
17
  "workflow.plan_check",
14
18
  "workflow.verifier",
15
19
  "workflow.auto_advance",
@@ -172,6 +176,11 @@
172
176
  "source": "^review\\.max_prompt_tokens_per_reviewer\\.[a-zA-Z0-9_-]+$",
173
177
  "description": "review.max_prompt_tokens_per_reviewer.<reviewer-slug>"
174
178
  },
179
+ {
180
+ "topLevel": "review",
181
+ "source": "^review\\.reviewer_instances\\.[a-zA-Z0-9_-]+\\.(cli|model|agent)$",
182
+ "description": "review.reviewer_instances.<instance-name>.<cli|model|agent> (#1517)"
183
+ },
175
184
  {
176
185
  "topLevel": "model_policy",
177
186
  "source": "^model_policy\\.runtime_tiers\\.[a-zA-Z0-9_-]+\\.(opus|sonnet|haiku)$",
@@ -0,0 +1,60 @@
1
+ # Agent Skills Self-Load (Bootstrap)
2
+
3
+ > **Shared contract.** Every `agent_skills` consumer agent self-loads its configured
4
+ > skills in its mandatory init step, so a project's `.planning/config.json`
5
+ > `agent_skills.<agent-type>` mapping reaches the agent that actually does the work —
6
+ > even when the orchestrator did not run bash init (e.g. a runtime whose `Skill()`
7
+ > delegation does not reliably execute the delegated workflow's bash, such as Cursor;
8
+ > see open-gsd/gsd-core#1600 / #1601). This is the durable counterpart to the
9
+ > orchestrator-side injection documented under
10
+ > [Agent Skills Injection](../../docs/CONFIGURATION.md#agent-skills-injection).
11
+
12
+ ## When to run
13
+
14
+ In your mandatory init step — right after `mandatory-initial-read.md` / the
15
+ `Project skills` discovery, before any other work.
16
+
17
+ ## Steps
18
+
19
+ 1. **Dedup guard (MANDATORY).** Look at your own prompt. If it already contains an
20
+ `<agent_skills>` block, the orchestrator already injected one — **skip self-load
21
+ entirely.** Loading a second copy wastes context on runtimes where orchestrator-side
22
+ injection also runs (e.g. Claude Code). The guard is what keeps the two seams from
23
+ doubling the block.
24
+
25
+ 2. **Query your configured skills.** Use **your own agent type** — the `name:` value in
26
+ your frontmatter (e.g. an agent whose frontmatter says `name: gsd-executor` queries
27
+ `gsd-executor`). The query is read-only and idempotent — it exits 0 with an empty
28
+ block when nothing is configured for your type:
29
+
30
+ ```bash
31
+ _AGENT_SKILLS=$(gsd_run query agent-skills <YOUR-FRONTMATTER-NAME> 2>/dev/null || true)
32
+ ```
33
+
34
+ The runtime `gsd_run` resolver is the standard one from
35
+ `_runtime-launcher.snippet.sh`; your own init already defines it.
36
+
37
+ 3. **Read every listed skill.** The block emits entries as `@<path>/SKILL.md`
38
+ includes — `Read` each one before starting work. If the block is empty, there is
39
+ nothing to do (zero overhead).
40
+
41
+ ## What self-load does and does not cover
42
+
43
+ | Skill form | Self-loads? | Notes |
44
+ |---|---|---|
45
+ | Project-relative path (`skills/my-skill`) | ✅ everywhere | `Read` the `@`-include |
46
+ | Global personal (`global:<name>`) | ✅ everywhere | resolves to the runtime global skills dir, then `Read` |
47
+ | Plugin-provided (`global:<plugin>:<skill>`) | Claude only | emitted as a Skill-tool directive on Claude; **skipped with a warning on all other runtimes** — the plugin/Skill-tool model has no equivalent elsewhere (#1601, #1258). Not closeable on Cursor. |
48
+
49
+ ## Notes
50
+
51
+ - **Idempotent and read-only.** `query agent-skills` never mutates state; calling it
52
+ twice (once by the orchestrator, once by the agent) is harmless because the dedup
53
+ guard suppresses the second load.
54
+ - **No new config keys.** This reuses the existing `agent_skills` map and the existing
55
+ `buildAgentSkillsBlock` / `cmdAgentSkills` machinery (`src/init.cts`). The 22 consumer
56
+ agent types are mirrored in `tests/agent-skills.test.cjs` (`CONSUMER_AGENTS`) and
57
+ guarded against drift by `tests/agent-skills-bootstrap.test.cjs`.
58
+ - **Checkers / read-only agents.** Bash is universal across consumer agents, so
59
+ self-load works for plan-checkers, verifiers, and auditors too; the bootstrap assumes
60
+ no tool an agent lacks.
@@ -0,0 +1,105 @@
1
+ # Honest Verifier — Abstention on Non-Inferable Checks
2
+
3
+ Shared reference for the **verify** phase. The verify-time companion to the spec-time
4
+ `@~/.claude/gsd-core/references/edge-probe.md` (which *classifies* non-inferable checks) and
5
+ `@~/.claude/gsd-core/references/prohibition-probe.md` (whose judgment-tier disposition this mirrors).
6
+ This doc is written in generic `spec → predicate → verifier` terms with no tool-specific vocabulary,
7
+ so it is portable: copy it into any verification process.
8
+
9
+ ## The problem it solves
10
+
11
+ A verifier is trustworthy on **inferable** checks — defects determined by the stated spec. On a
12
+ **non-inferable** check the correct answer is *not derivable from the spec alone* (e.g. "does `[1,2]`
13
+ touching `[2,3]` merge?", "is a 'character' a grapheme or a code unit?"). On these the verifier *does
14
+ not know that it does not know*: measured behavior is a **confident PASS on the blind-spot check ~100%
15
+ of the time** (mean confidence ~0.93), because a model cannot self-detect a gap it does not perceive.
16
+
17
+ The edge-probe already detects these at spec time and tags them `verification: backstop` (ADR-550
18
+ D7a). The honest verifier consumes that tag so the verifier **abstains** instead of confidently
19
+ false-passing — converting a silent false-pass (the worst failure: you don't know to look) into an
20
+ explicit, actionable "write a held-out test." Measured: the confident-false-pass rate on the blind
21
+ spot drops **100% → 17%** (N17).
22
+
23
+ ## The two properties that define the design
24
+
25
+ 1. **Exogenous, not endogenous.** The trigger is the *external tag* (`backstop`), never the verifier's
26
+ self-judgment. Asking the verifier to "abstain if unsure" barely moves the number (100% → 67%) and
27
+ only on ambiguity it already notices; on a true blind spot it stays confidently wrong. A confidence
28
+ gate cannot reach a blind spot the model does not feel — so there is **no "are you sure?" prompt**;
29
+ routing is on the pre-existing tag only.
30
+ 2. **Routing, not diagnosis.** The verifier need not name the omitted rule (if it could, it wouldn't
31
+ be a blind spot). In testing, verifiers abstained correctly while citing the *wrong* edge. The
32
+ honest verdict requires only "I was told this is under-specified and I cannot rule it out." The
33
+ omitted rule is carried by a human-authored held-out test, not by the verifier.
34
+
35
+ ## The disposition (the protocol)
36
+
37
+ For each `must_haves.truths` item:
38
+
39
+ | Item | Confirmable with explicit evidence? | Disposition |
40
+ |---|---|---|
41
+ | Inferable (plain string, or `verification: explicit`) | n/a — graded normally | ✓ VERIFIED / ✗ FAILED as usual; **never abstained** (over-abstention guard) |
42
+ | Non-inferable (`verification: backstop`) | **yes** (a wired held-out/property-based test that passes, or a directly-observed behavior) | ✓ VERIFIED |
43
+ | Non-inferable (`verification: backstop`) | **no** | **abstain** → ⚠️ `insufficient_spec`, flagged, → `human_needed` — **never `passed`** |
44
+
45
+ - **Explicit evidence** = a wired held-out/property-based test that passes, or a behavior the verifier
46
+ directly observed. Symbol presence + wiring is **not** explicit evidence for a non-inferable truth.
47
+ - **Never silent, never a hard halt.** *Interactive:* the abstained item routes to the end-of-phase
48
+ human checkpoint. *Autonomous (AFK):* it produces a prominent `unverified — held-out test
49
+ recommended` flag and the completion line reads "complete with N unverified non-inferable checks";
50
+ the run neither silently passes the blind spot nor hard-halts.
51
+ - **Distinguishable reason.** The abstain disposition carries `reason: insufficient_spec` so the
52
+ `human_needed` outcome is never conflated with an ordinary manual-UAT `human_needed`.
53
+
54
+ This is the verify-time half of ADR-550 Decision 4 (the never-silent-pass disposition), applied to the
55
+ edge `backstop` truth tier instead of the prohibition judgment tier — the same machinery, opposite
56
+ polarity (must-HAVE under-specified vs must-NOT irreducible).
57
+
58
+ ## Deterministic engine surface
59
+
60
+ The CI-testable surface is the **deterministic disposition + projection**, never the LLM's judgment
61
+ (ADR-550 D5 — a test asserting the model's verdict is vacuous and rejected). In `probe-core`:
62
+
63
+ - `truthStatement(t)` / `truthVerification(t)` — normalizers; read a truth's statement and tier from
64
+ either the plain-string or object form (a truth-reader MUST normalize, never assume a string).
65
+ - `projectTruths(items)` — conservative serializer: a `backstop` truth → flat-scalar object
66
+ `{ statement, verification: backstop }`; every inferable truth → a bare string.
67
+ - `dispositionForUnverifiableTruth(truth, { evidence })` → `{ status, flagged, tier, reason }`:
68
+ `backstop` + no evidence → `unverified`/`flagged`/`insufficient_spec`; `backstop` + evidence →
69
+ `green`; non-`backstop` → `green` (over-abstention guard).
70
+
71
+ ## Capable-tier requirement (a documented cost)
72
+
73
+ Abstention is **model-tier dependent** and this is a standing cost, not an assumption:
74
+
75
+ - The default `gsd-verifier` tier (`sonnet`, golden/balanced) heeds the exogenous tag reliably
76
+ (2/2 under testing).
77
+ - The **budget tier (`haiku`)** is the least flag-responsive (1/2, inconsistent) and **degrades toward
78
+ current behavior** (confident false-pass). Run honest-verifier on a capable tier; treat the budget
79
+ tier as best-effort. Re-validate when the `gsd-verifier` model tier changes or a new budget model is
80
+ adopted (captured as a test so a tier regression is caught, not discovered in production).
81
+
82
+ ## Evidence and scope (stated honestly)
83
+
84
+ - **Evidence strength.** N17 is n=27 verdicts (3 models × 3 conditions × 3 tasks), 1 rep —
85
+ **direction-finding, not powered.** The blind-spot effect is large and monotone
86
+ (100% → 67% → 17%); the two costs are clean single events (a *false* tag made the strongest model
87
+ over-abstain on a real spec-determined bug; the weakest tier was flag-deaf) and they name exactly
88
+ the failure modes the over-abstention guard and the capable-tier requirement defend against.
89
+ - **Tag-precision coupling.** Quality is bounded by the edge-probe's `backstop` recall/precision — a
90
+ false non-inferable flag causes over-abstention. Positive coupling: improving the probe (#1110)
91
+ improves this for free. It adds no independent burden.
92
+ - **Explicit non-goals.** Does NOT identify the omitted rule; does NOT recalibrate decisive verdicts;
93
+ does NOT defend against *malicious compliance* (a self-graded review rationalizing away its own
94
+ findings). It raises the floor on *honest* uncertainty about non-inferable checks — that is the
95
+ whole claim.
96
+
97
+ ## Distinct from neighbours
98
+
99
+ - **vs `PRESENT_BEHAVIOR_UNVERIFIED` (#966 axis):** that is the *inferable-but-unobserved* case — the
100
+ truth **can** be verified from the spec but was shortcut-passed on symbol presence; the fix is to
101
+ demand behavioral evidence. Honest-verifier is the *non-inferable* case — the truth **cannot** be
102
+ verified from the spec at all; the fix is to abstain and route to a held-out test. Orthogonal axes
103
+ (insufficient *evidence* vs insufficient *spec*); both feed the same `human_needed` sink.
104
+ - **vs prohibition judgment-tier (#644):** that disposes **must-NOT** constraints; honest-verifier
105
+ disposes **non-inferable positive truths**. Opposite polarity, same never-silent disposition.
@@ -133,6 +133,33 @@ If you're using Claude Code with OpenRouter, a local model, or any non-Anthropic
133
133
 
134
134
  Without `inherit`, GSD's default `balanced` profile spawns specific Anthropic models (`opus`, `sonnet`, `haiku`) for each agent type, which can result in additional API costs through your non-Anthropic provider.
135
135
 
136
+ ## Advisor Tool (Claude Code)
137
+
138
+ Claude Code (v2.1.98+) can pair the session's executor model with a stronger **advisor** model that it consults mid-generation for strategy and course-correction (Anthropic's [advisor tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool)). This is a host-runtime feature, not a GSD setting — GSD selects each agent's *executor* model through the profile/tier system above; Claude Code supplies the advisor.
139
+
140
+ Set it once at the session level with `/advisor <model>` (or the `advisorModel` setting / `--advisor` flag). **Subagents inherit the session advisor automatically**, so every GSD subagent an orchestrator spawns gets the same advisor with no per-agent configuration. It composes cleanly with GSD's tiering: the profile keeps executors cheap where the work is mechanical, and the advisor adds a stronger reviewer inline on the turns that benefit.
141
+
142
+ ### Candidate pairings
143
+
144
+ Per Anthropic's advisor-tool docs the advisor must be at least as capable as the executor. Candidate pairings by profile — evaluate on your own workload; the quality/cost characterizations below are Anthropic-reported, not GSD guarantees:
145
+
146
+ | Profile | Typical executors | Candidate advisor | Rationale (per Anthropic docs) |
147
+ |---|---|---|---|
148
+ | `budget` | Haiku / Sonnet | Fable 5 or Opus | A step up in intelligence over Haiku alone, at lower cost than switching the executor to a larger model |
149
+ | `balanced` | Sonnet | Fable 5 or Opus | A quality lift at similar or lower total cost than Sonnet-solo on complex tasks |
150
+ | `quality` / `adaptive` | Opus (planning), Sonnet | Fable 5 or Opus | Marginal on turns already at top capability; most valuable on the Sonnet-executor agents |
151
+
152
+ Fable 5 is a valid advisor for Haiku 4.5, Sonnet 4.6/5, and Opus 4.8 executors, so it pairs with any tier a profile assigns.
153
+
154
+ ### When it's worth enabling
155
+
156
+ - **Worth it:** long, multi-step agent loops where the plan matters but most turns are mechanical — e.g. `execute-phase` and `debug`. Anthropic's docs note advisor prompt-caching pays off at roughly three or more advisor calls, which these long loops make.
157
+ - **Skip it:** short, one-shot agents (mappers, quick audits, single-file checks) — there is little to plan, and the advisor adds cost without a commensurate quality gain.
158
+
159
+ ### Constraint: session-level only (today)
160
+
161
+ The advisor is a single session-wide setting inherited by all subagents; there is **no per-agent advisor selection**, so GSD cannot vary the advisor by role the way it varies the executor model (e.g. "no advisor on the Haiku mapper, a Fable 5 advisor on the Sonnet executor"). Per-agent advisor control is tracked upstream at [anthropics/claude-code#73072](https://github.com/anthropics/claude-code/issues/73072); until it lands, pick one session advisor that fits the most valuable agents in your run.
162
+
136
163
  ## Dynamic Routing with Failure-Tier Escalation (#3024)
137
164
 
138
165
  When `dynamic_routing.enabled = true` in `.planning/config.json`, the resolver picks a model from a tier-mapped table based on the agent's *default tier* (light / standard / heavy) and escalates to the next tier up on orchestrator-detected soft failure.
@@ -0,0 +1,99 @@
1
+ # Reviewer Instances (#1517)
2
+
3
+ Custom reviewer instances for `/gsd:review`: run one model-capable adapter (e.g. OpenCode)
4
+ as several independent reviewer identities in a single review pass. Loaded lazily by
5
+ `gsd-core/workflows/review.md` when `review.reviewer_instances` is configured. See
6
+ [ADR-1517](../docs/adr/1517-reviewer-instances-config-surface.md) for the contract.
7
+
8
+ ---
9
+
10
+ ## Config shape
11
+
12
+ A `review.reviewer_instances` object under the `review` namespace. Each entry maps an
13
+ instance name to `{ cli, model?, agent? }`:
14
+
15
+ ```json
16
+ {
17
+ "review": {
18
+ "reviewer_instances": {
19
+ "opencode-deepseek": { "cli": "opencode", "model": "deepseek/deepseek-v4-pro", "agent": "review" },
20
+ "opencode-mimo": { "cli": "opencode", "model": "xiaomi/mimo-v2.5-pro" }
21
+ },
22
+ "default_reviewers": ["opencode-deepseek", "opencode-mimo", "codex"]
23
+ }
24
+ }
25
+ ```
26
+
27
+ - Instance name: `^[a-z0-9][a-z0-9-]*$`, must not equal a built-in slug. Validated at
28
+ `config-set` time.
29
+ - `cli`: MUST be a known adapter (`KNOWN_REVIEWER_SLUGS`) — never an arbitrary shell command.
30
+ - `model`: opaque `provider/model` string, passed through verbatim. GSD does not parse it.
31
+ - `agent`: opaque string; honoured only by adapters with a native agent concept (OpenCode
32
+ `--agent` in v1). Ignored by other adapters.
33
+
34
+ ---
35
+
36
+ ## Resolution rules (single source)
37
+
38
+ The canonical logic lives in `resolveReviewerSelection` / `normalizeReviewerInstances` in
39
+ `review-reviewer-selection.cjs`. Apply the SAME rules in the workflow so the two surfaces
40
+ cannot diverge (`DEFECT.GENERATIVE-FIX`; parity-locked in
41
+ `tests/review-reviewer-instances.test.cjs`).
42
+
43
+ 1. Instances participate ONLY via `review.default_reviewers`. They never appear under `--all`
44
+ or explicit `--<cli>` flags, and there are no per-instance CLI flags.
45
+ 2. Expand instance references BEFORE the built-in-slug check: an entry that is a key in
46
+ `review.reviewer_instances` is an **instance**; an entry that is a built-in slug is a
47
+ **builtin**.
48
+ 3. An instance is **available** iff its base `cli` is detected (e.g. `opencode-deepseek` is
49
+ available iff `opencode` is available).
50
+ 4. An entry that is NEITHER a defined instance NOR a built-in slug is a **hard error** (likely
51
+ a typo'd instance name) — stop and report it. Do NOT silently drop it. (When
52
+ `review.reviewer_instances` is absent entirely, fall back to the legacy unknown-slug
53
+ warn-and-drop behaviour for backward compatibility.)
54
+ 5. `model`/`agent`/instance-name are opaque: pass them as separate argv elements. They are
55
+ NEVER interpolated into shell strings.
56
+
57
+ ---
58
+
59
+ ## Invocation
60
+
61
+ For each selected INSTANCE, invoke its base `cli` using the instance's own `model`/`agent` —
62
+ NOT the global `review.models.<cli>`. Each instance writes to its OWN per-instance output file
63
+ and runs as a distinct reviewer identity.
64
+
65
+ For an OpenCode-backed instance (the motivating adapter):
66
+
67
+ ```bash
68
+ # $INSTANCE_MODEL / $INSTANCE_AGENT come from the instance spec; $INSTANCE_NAME is the
69
+ # reviewer identity (e.g. opencode-deepseek). --agent is OpenCode's native subagent flag;
70
+ # omit it when the instance has no agent.
71
+ if [ -n "$INSTANCE_AGENT" ] && [ "$INSTANCE_AGENT" != "null" ]; then
72
+ cat /tmp/gsd-review-prompt-{phase}.md | opencode run --model "$INSTANCE_MODEL" --agent "$INSTANCE_AGENT" - 2>/dev/null > /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md
73
+ else
74
+ cat /tmp/gsd-review-prompt-{phase}.md | opencode run --model "$INSTANCE_MODEL" - 2>/dev/null > /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md
75
+ fi
76
+ if [ ! -s /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md ]; then
77
+ echo "OpenCode review ($INSTANCE_NAME) failed or returned empty output." > /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md
78
+ fi
79
+ ```
80
+
81
+ For an instance backed by a DIFFERENT cli, reuse that cli's invocation block with two
82
+ substitutions: use the instance's `model` in place of the global `review.models.<cli>` value,
83
+ and write to `/tmp/gsd-review-${INSTANCE_NAME}-{phase}.md`. Only `opencode` honours an
84
+ `agent` field in v1; ignore `agent` for other adapters.
85
+
86
+ ---
87
+
88
+ ## REVIEWS.md contract
89
+
90
+ - **Frontmatter `reviewers:`** records the actual identities invoked. For a built-in slug use
91
+ the slug (`opencode`); for an instance use the instance name (`opencode-deepseek`), so
92
+ frontmatter distinguishes the independent voices. Example:
93
+ `reviewers: [opencode-deepseek, opencode-mimo, codex]`.
94
+ - **Section headers:** each instance gets its OWN top-level section, headed with the base
95
+ adapter's display name plus the instance name in parentheses:
96
+ `## OpenCode Review (opencode-deepseek)`. Same-cli instances are never collapsed.
97
+ - **Shared-adapter caveat:** when ≥2 invoked instances share the same base `cli`, print a
98
+ one-line caveat immediately after the frontmatter (before the first section), e.g.:
99
+ `> Note: opencode-deepseek and opencode-mimo share the opencode adapter; their consensus is cross-model, not cross-tool.`