@opengsd/gsd-core 1.6.1 → 1.7.0-rc.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +20 -0
- package/.claude-plugin/plugin.json +1 -1
- package/.opencode/plugins/gsd-core.js +711 -0
- package/agents/gsd-advisor-researcher.md +2 -0
- package/agents/gsd-ai-researcher.md +1 -1
- package/agents/gsd-assumptions-analyzer.md +2 -0
- package/agents/gsd-code-fixer.md +2 -0
- package/agents/gsd-code-reviewer.md +2 -0
- package/agents/gsd-codebase-mapper.md +2 -0
- package/agents/gsd-debugger.md +2 -0
- package/agents/gsd-doc-writer.md +2 -0
- package/agents/gsd-eval-auditor.md +2 -0
- package/agents/gsd-executor.md +9 -6
- package/agents/gsd-integration-checker.md +2 -0
- package/agents/gsd-nyquist-auditor.md +2 -0
- package/agents/gsd-phase-researcher.md +2 -0
- package/agents/gsd-plan-checker.md +2 -0
- package/agents/gsd-planner.md +2 -0
- package/agents/gsd-project-researcher.md +2 -0
- package/agents/gsd-research-synthesizer.md +2 -0
- package/agents/gsd-roadmapper.md +2 -0
- package/agents/gsd-security-auditor.md +2 -0
- package/agents/gsd-ui-auditor.md +2 -0
- package/agents/gsd-ui-checker.md +2 -0
- package/agents/gsd-ui-researcher.md +2 -0
- package/agents/gsd-verifier.md +5 -2
- package/bin/gsd-mcp-server.js +31 -0
- package/bin/install.js +411 -1146
- package/commands/gsd/review.md +6 -0
- package/gemini-extension.json +1 -1
- package/gsd-core/bin/gsd-tools.cjs +134 -8
- package/gsd-core/bin/lib/adapter-declarative.cjs +35 -0
- package/gsd-core/bin/lib/adapter-imperative.cjs +52 -0
- package/gsd-core/bin/lib/assumption-delta.cjs +231 -0
- package/gsd-core/bin/lib/capability-lifecycle.cjs +7 -7
- package/gsd-core/bin/lib/capability-loader.cjs +45 -9
- package/gsd-core/bin/lib/capability-lock.cjs +2 -2
- package/gsd-core/bin/lib/capability-registry.cjs +891 -82
- package/gsd-core/bin/lib/capability-source.cjs +26 -11
- package/gsd-core/bin/lib/capability-validator.cjs +222 -2
- package/gsd-core/bin/lib/cli-skew-check.cjs +44 -0
- package/gsd-core/bin/lib/command-aliases.cjs +8 -0
- package/gsd-core/bin/lib/commands.cjs +2 -1
- package/gsd-core/bin/lib/config.cjs +27 -0
- package/gsd-core/bin/lib/embedding-adapter.cjs +27 -0
- package/gsd-core/bin/lib/external-descriptor-trust.cjs +70 -0
- package/gsd-core/bin/lib/frontmatter.cjs +53 -6
- package/gsd-core/bin/lib/handshake-serialized.cjs +70 -0
- package/gsd-core/bin/lib/hook-bus.cjs +81 -0
- package/gsd-core/bin/lib/host-integration-sdk.cjs +53 -0
- package/gsd-core/bin/lib/host-integration.cjs +469 -0
- package/gsd-core/bin/lib/init.cjs +35 -7
- package/gsd-core/bin/lib/install-engine.cjs +755 -0
- package/gsd-core/bin/lib/install-profiles.cjs +35 -4
- package/gsd-core/bin/lib/installer-migrations.cjs +1 -1
- package/gsd-core/bin/lib/mcp-server.cjs +194 -0
- package/gsd-core/bin/lib/milestone.cjs +68 -40
- package/gsd-core/bin/lib/model-adapter.cjs +50 -0
- package/gsd-core/bin/lib/phase-id.cjs +18 -0
- package/gsd-core/bin/lib/phase.cjs +57 -90
- package/gsd-core/bin/lib/phases-command-router.cjs +4 -3
- package/gsd-core/bin/lib/planning-workspace.cjs +1 -1
- package/gsd-core/bin/lib/probe-core.cjs +132 -2
- package/gsd-core/bin/lib/review-reviewer-selection.cjs +129 -13
- package/gsd-core/bin/lib/roadmap-command-router.cjs +3 -2
- package/gsd-core/bin/lib/roadmap-parser.cjs +21 -11
- package/gsd-core/bin/lib/roadmap-upgrade.cjs +3 -2
- package/gsd-core/bin/lib/roadmap.cjs +33 -22
- package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +65 -9
- package/gsd-core/bin/lib/runtime-artifact-install-plan.cjs +54 -4
- package/gsd-core/bin/lib/runtime-artifact-layout.cjs +5 -2
- package/gsd-core/bin/lib/runtime-hooks-surface.cjs +1 -1
- package/gsd-core/bin/lib/runtime-name-policy.cjs +160 -30
- package/gsd-core/bin/lib/shell-command-projection.cjs +16 -0
- package/gsd-core/bin/lib/stale-bake-guard.cjs +254 -0
- package/gsd-core/bin/lib/state-command-router.cjs +4 -0
- package/gsd-core/bin/lib/state-io.cjs +55 -0
- package/gsd-core/bin/lib/state-transition.cjs +1603 -0
- package/gsd-core/bin/lib/state.cjs +327 -683
- package/gsd-core/bin/lib/surface.cjs +4 -1
- package/gsd-core/bin/lib/validate.cjs +2 -1
- package/gsd-core/bin/lib/verify.cjs +6 -4
- package/gsd-core/bin/lib/workstream-inventory-builder.cjs +12 -2
- package/gsd-core/bin/lib/workstream-inventory.cjs +28 -0
- package/gsd-core/bin/lib/workstream.cjs +4 -4
- package/gsd-core/bin/shared/config-schema.manifest.json +9 -0
- package/gsd-core/references/agent-skills-bootstrap.md +60 -0
- package/gsd-core/references/honest-verifier.md +105 -0
- package/gsd-core/references/model-profiles.md +27 -0
- package/gsd-core/references/reviewer-instances.md +99 -0
- package/gsd-core/workflows/autonomous.md +30 -32
- package/gsd-core/workflows/complete-milestone.md +6 -10
- package/gsd-core/workflows/execute-phase.md +1 -1
- package/gsd-core/workflows/forensics.md +3 -3
- package/gsd-core/workflows/help/modes/full.md +1 -1
- package/gsd-core/workflows/manager.md +15 -15
- package/gsd-core/workflows/milestone-summary.md +3 -3
- package/gsd-core/workflows/new-milestone.md +6 -0
- package/gsd-core/workflows/plan-phase/steps/closed-phase-gate.md +42 -0
- package/gsd-core/workflows/plan-phase/steps/prd-express-path.md +102 -0
- package/gsd-core/workflows/plan-phase/steps/windows-troubleshooting.md +23 -0
- package/gsd-core/workflows/plan-phase.md +4 -159
- package/gsd-core/workflows/review.md +33 -2
- package/gsd-core/workflows/thread.md +4 -4
- package/gsd-core/workflows/verify-phase.md +11 -4
- package/gsd-core/workflows/verify-work.md +1 -2
- package/hooks/dist/gsd-graphify-update.sh +7 -1
- package/hooks/gsd-graphify-update.sh +7 -1
- package/package.json +6 -4
- package/scripts/ci-test-scope.cjs +38 -9
- package/scripts/lint-allow-test-rule-refs.allowlist.json +0 -1
- package/scripts/lint-regression-test-names.allowlist.json +3 -0
- package/scripts/lint-test-file-count.allowlist.json +19 -5
- package/scripts/mutation-matrix.cjs +45 -3
- package/scripts/prompt-injection-scan.sh +8 -0
- package/scripts/run-tests.cjs +51 -1
- package/scripts/sync-manifest-versions.cjs +66 -14
- package/skills/gsd-review/SKILL.md +6 -0
- package/scripts/lint-windows-test-portability.cjs +0 -178
|
@@ -43,6 +43,9 @@ const runtimeArtifactLayout = require("./runtime-artifact-layout.cjs");
|
|
|
43
43
|
const { findInstallSourceRoot } = runtimeArtifactLayout;
|
|
44
44
|
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
45
45
|
const runtimeArtifactConversion = require("./runtime-artifact-conversion.cjs");
|
|
46
|
+
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
47
|
+
const runtimeArtifactInstallPlan = require("./runtime-artifact-install-plan.cjs");
|
|
48
|
+
const { assertDestWithinConfigHome } = runtimeArtifactInstallPlan;
|
|
46
49
|
const SURFACE_FILE_NAME = '.gsd-surface.json';
|
|
47
50
|
/**
|
|
48
51
|
* Read the surface state from a runtime config directory.
|
|
@@ -302,7 +305,7 @@ function applySurface(runtimeConfigDir, layout, manifest, clusterMap, registry)
|
|
|
302
305
|
tempDirsToClean.push(rewritten);
|
|
303
306
|
}
|
|
304
307
|
}
|
|
305
|
-
const dest =
|
|
308
|
+
const dest = assertDestWithinConfigHome(layout.configDir, kind.destSubpath);
|
|
306
309
|
_syncGsdDir(staged, dest, kind, skillManifest);
|
|
307
310
|
}
|
|
308
311
|
}
|
|
@@ -93,7 +93,8 @@ function buildRoadmapPhaseVariants(roadmapContent) {
|
|
|
93
93
|
const roadmapPhaseVariants = new Set();
|
|
94
94
|
// Matches both legacy numeric (Phase 1:), decimal (Phase 2.1:), milestone-prefixed (Phase 2-01:),
|
|
95
95
|
// and bracket-prefixed (### [GSD] Phase 2-01:) headings.
|
|
96
|
-
|
|
96
|
+
// #1729: `(?:\s*\([^)\n]*\))?` tolerates a pre-colon ( ) tag (literal mirror of OPTIONAL_PHASE_TAG_SOURCE).
|
|
97
|
+
const phasePattern = /#{2,4}\s*(?:\[[^\]]+\]\s*)?Phase\s+([\w][\w.-]*)(?:\s*\([^)\n]*\))?\s*:/gi;
|
|
97
98
|
let m;
|
|
98
99
|
while ((m = phasePattern.exec(roadmapContent)) !== null) {
|
|
99
100
|
roadmapPhases.add(m[1]);
|
|
@@ -40,7 +40,7 @@ const configLoaderMod = require("./config-loader.cjs");
|
|
|
40
40
|
const { loadConfig, CONFIG_DEFAULTS } = configLoaderMod;
|
|
41
41
|
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
42
42
|
const phaseIdMod = require("./phase-id.cjs");
|
|
43
|
-
const { normalizePhaseName, phaseTokenMatches, escapeRegex, getMilestoneFromPhaseId } = phaseIdMod;
|
|
43
|
+
const { normalizePhaseName, phaseTokenMatches, escapeRegex, getMilestoneFromPhaseId, OPTIONAL_PHASE_TAG_SOURCE } = phaseIdMod;
|
|
44
44
|
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
45
45
|
const phaseLocatorMod = require("./phase-locator.cjs");
|
|
46
46
|
const { findPhaseInternal } = phaseLocatorMod;
|
|
@@ -992,7 +992,8 @@ function checkMilestonePrefixMismatches(roadmapContent, { getMilestoneFromPhaseI
|
|
|
992
992
|
}
|
|
993
993
|
for (const section of sections) {
|
|
994
994
|
const content = roadmapContent.slice(section.start, section.end);
|
|
995
|
-
|
|
995
|
+
// #1729: `(?:\s*\([^)\n]*\))?` tolerates a pre-colon ( ) tag (literal mirror of OPTIONAL_PHASE_TAG_SOURCE).
|
|
996
|
+
const phaseRx = /#{2,4}\s*(?:\[[^\]]+\]\s*)?Phase\s+([\w][\w.-]*)(?:\s*\([^)\n]*\))?\s*:/gi;
|
|
996
997
|
let pm;
|
|
997
998
|
while ((pm = phaseRx.exec(content)) !== null) {
|
|
998
999
|
const phaseId = pm[1];
|
|
@@ -1351,7 +1352,7 @@ function cmdValidateHealth(cwd, options, raw) {
|
|
|
1351
1352
|
stateContent.match(/Current Phase:\s*(\S+)/i);
|
|
1352
1353
|
if (currentPhaseMatch) {
|
|
1353
1354
|
const statePhase = currentPhaseMatch[1].replace(/^0+/, '');
|
|
1354
|
-
const phaseCheckboxRe = new RegExp(`-\\s*\\[x\\].*Phase\\s+0*${escapeRegex(statePhase)}[:\\s]`, 'i');
|
|
1355
|
+
const phaseCheckboxRe = new RegExp(`-\\s*\\[x\\].*Phase\\s+0*${escapeRegex(statePhase)}${OPTIONAL_PHASE_TAG_SOURCE}[:\\s]`, 'i');
|
|
1355
1356
|
if (phaseCheckboxRe.test(roadmapContentFull)) {
|
|
1356
1357
|
const stateStatus = stateContent.match(/\*\*Status:\*\*\s*(.+)/i);
|
|
1357
1358
|
const statusVal = stateStatus ? stateStatus[1].trim().toLowerCase() : '';
|
|
@@ -1509,7 +1510,8 @@ function cmdValidateHealth(cwd, options, raw) {
|
|
|
1509
1510
|
if (isMarkedComplete) {
|
|
1510
1511
|
const roadmapRaw = node_fs_1.default.readFileSync(roadmapPath, 'utf-8');
|
|
1511
1512
|
const scopedContent = extractCurrentMilestone(roadmapRaw, cwd);
|
|
1512
|
-
|
|
1513
|
+
// #1729: `(?:\s*\([^)\n]*\))?` tolerates a pre-colon ( ) tag (literal mirror of OPTIONAL_PHASE_TAG_SOURCE).
|
|
1514
|
+
const phasePattern = /#{2,4}\s*Phase\s+(\d+[A-Z]?(?:\.\d+)*)(?:\s*\([^)\n]*\))?\s*:\s*([^\n]+)/gi;
|
|
1513
1515
|
const unstarted = [];
|
|
1514
1516
|
let pm;
|
|
1515
1517
|
// Non-hoisted: load-order matters (circular dep guard)
|
|
@@ -28,7 +28,7 @@ function isCompletedInventory(status) {
|
|
|
28
28
|
return /\bmilestone\s+complete\b/.test(s) || /\barchived\b/.test(s);
|
|
29
29
|
}
|
|
30
30
|
function buildWorkstreamInventory(inputs) {
|
|
31
|
-
const { name, projectDir, workstreamDir, phaseDirNames, activeWorkstreamName, phaseFilesCounts, roadmapPhaseCount, stateProjection, filesExist, } = inputs;
|
|
31
|
+
const { name, projectDir, workstreamDir, phaseDirNames, activeWorkstreamName, phaseFilesCounts, roadmapPhaseCount, stateProjection, filesExist, milestoneShipped, } = inputs;
|
|
32
32
|
// Index counts by directory for O(1) lookup during sort/iteration
|
|
33
33
|
const countsMap = new Map();
|
|
34
34
|
for (const entry of phaseFilesCounts) {
|
|
@@ -56,6 +56,14 @@ function buildWorkstreamInventory(inputs) {
|
|
|
56
56
|
summary_count: counts.summaryCount,
|
|
57
57
|
});
|
|
58
58
|
}
|
|
59
|
+
// #1913: derive status from authoritative shipped signals rather than trusting
|
|
60
|
+
// the mutable STATE.md `Status` field. When a shipped signal is present, the
|
|
61
|
+
// workstream is "milestone complete" regardless of a stale field value.
|
|
62
|
+
const fieldStatus = stateProjection.status;
|
|
63
|
+
const useDerived = milestoneShipped;
|
|
64
|
+
const status = useDerived ? 'milestone complete' : fieldStatus;
|
|
65
|
+
const status_source = useDerived ? 'derived' : 'field';
|
|
66
|
+
const status_conflict = useDerived && !isCompletedInventory(fieldStatus);
|
|
59
67
|
return {
|
|
60
68
|
name,
|
|
61
69
|
path: toPosixPath(node_path_1.default.relative(projectDir, workstreamDir)),
|
|
@@ -65,7 +73,9 @@ function buildWorkstreamInventory(inputs) {
|
|
|
65
73
|
state: filesExist.state,
|
|
66
74
|
requirements: filesExist.requirements,
|
|
67
75
|
},
|
|
68
|
-
status
|
|
76
|
+
status,
|
|
77
|
+
status_source,
|
|
78
|
+
status_conflict,
|
|
69
79
|
current_phase: stateProjection.current_phase,
|
|
70
80
|
last_activity: stateProjection.last_activity,
|
|
71
81
|
phases,
|
|
@@ -63,6 +63,33 @@ function readStateProjection(statePath) {
|
|
|
63
63
|
};
|
|
64
64
|
}
|
|
65
65
|
}
|
|
66
|
+
/**
|
|
67
|
+
* #1913: detect an authoritative shipped signal for a workstream so the
|
|
68
|
+
* inventory status is never trusted from the mutable STATE.md `Status` field
|
|
69
|
+
* alone. Returns true when EITHER an archived milestone snapshot is present
|
|
70
|
+
* under `<planningBase>/milestones/` OR the workstream ROADMAP carries a
|
|
71
|
+
* SHIPPED marker — both are hard to desync, unlike the hand-maintained field.
|
|
72
|
+
*/
|
|
73
|
+
function workstreamMilestoneShipped(roadmapPath, planningBase) {
|
|
74
|
+
try {
|
|
75
|
+
const milestonesDir = node_path_1.default.join(planningBase, 'milestones');
|
|
76
|
+
for (const entry of node_fs_1.default.readdirSync(milestonesDir, { withFileTypes: true })) {
|
|
77
|
+
if (entry.isFile() && /-ROADMAP\.md$/i.test(entry.name))
|
|
78
|
+
return true;
|
|
79
|
+
}
|
|
80
|
+
}
|
|
81
|
+
catch {
|
|
82
|
+
/* no milestones archive dir */
|
|
83
|
+
}
|
|
84
|
+
try {
|
|
85
|
+
if (/SHIPPED/i.test(node_fs_1.default.readFileSync(roadmapPath, 'utf-8')))
|
|
86
|
+
return true;
|
|
87
|
+
}
|
|
88
|
+
catch {
|
|
89
|
+
/* no roadmap */
|
|
90
|
+
}
|
|
91
|
+
return false;
|
|
92
|
+
}
|
|
66
93
|
function sortWorkstreamInventories(inventories, activeWorkstreamName) {
|
|
67
94
|
return [...inventories].sort((a, b) => {
|
|
68
95
|
const aActive = a.name === activeWorkstreamName ? 1 : 0;
|
|
@@ -99,6 +126,7 @@ function inspectWorkstream(cwd, name, options = {}) {
|
|
|
99
126
|
state: node_fs_1.default.existsSync(p.state),
|
|
100
127
|
requirements: node_fs_1.default.existsSync(p.requirements),
|
|
101
128
|
},
|
|
129
|
+
milestoneShipped: workstreamMilestoneShipped(p.roadmap, p.planning),
|
|
102
130
|
});
|
|
103
131
|
}
|
|
104
132
|
function listWorkstreamInventories(cwd) {
|
|
@@ -67,7 +67,7 @@ function migrateToWorkstreams(cwd, workstreamName) {
|
|
|
67
67
|
const src = node_path_1.default.join(baseDir, item.name);
|
|
68
68
|
if (node_fs_1.default.existsSync(src)) {
|
|
69
69
|
const dest = node_path_1.default.join(wsDir, item.name);
|
|
70
|
-
|
|
70
|
+
(0, shell_command_projection_cjs_1.retryRenameSync)(src, dest);
|
|
71
71
|
filesMoved.push(item.name);
|
|
72
72
|
}
|
|
73
73
|
}
|
|
@@ -75,7 +75,7 @@ function migrateToWorkstreams(cwd, workstreamName) {
|
|
|
75
75
|
catch (err) {
|
|
76
76
|
for (const name of filesMoved) {
|
|
77
77
|
try {
|
|
78
|
-
|
|
78
|
+
(0, shell_command_projection_cjs_1.retryRenameSync)(node_path_1.default.join(wsDir, name), node_path_1.default.join(baseDir, name));
|
|
79
79
|
}
|
|
80
80
|
catch { /* ignore */ }
|
|
81
81
|
}
|
|
@@ -275,14 +275,14 @@ function cmdWorkstreamComplete(cwd, name, options, raw) {
|
|
|
275
275
|
try {
|
|
276
276
|
const entries = node_fs_1.default.readdirSync(wsDir, { withFileTypes: true });
|
|
277
277
|
for (const entry of entries) {
|
|
278
|
-
|
|
278
|
+
(0, shell_command_projection_cjs_1.retryRenameSync)(node_path_1.default.join(wsDir, entry.name), node_path_1.default.join(archivePath, entry.name));
|
|
279
279
|
filesMoved.push(entry.name);
|
|
280
280
|
}
|
|
281
281
|
}
|
|
282
282
|
catch (err) {
|
|
283
283
|
for (const fname of filesMoved) {
|
|
284
284
|
try {
|
|
285
|
-
|
|
285
|
+
(0, shell_command_projection_cjs_1.retryRenameSync)(node_path_1.default.join(archivePath, fname), node_path_1.default.join(wsDir, fname));
|
|
286
286
|
}
|
|
287
287
|
catch { /* ignore */ }
|
|
288
288
|
}
|
|
@@ -10,6 +10,10 @@
|
|
|
10
10
|
"brave_search",
|
|
11
11
|
"firecrawl",
|
|
12
12
|
"exa_search",
|
|
13
|
+
"tavily_search",
|
|
14
|
+
"ref_search",
|
|
15
|
+
"perplexity",
|
|
16
|
+
"jina",
|
|
13
17
|
"workflow.plan_check",
|
|
14
18
|
"workflow.verifier",
|
|
15
19
|
"workflow.auto_advance",
|
|
@@ -172,6 +176,11 @@
|
|
|
172
176
|
"source": "^review\\.max_prompt_tokens_per_reviewer\\.[a-zA-Z0-9_-]+$",
|
|
173
177
|
"description": "review.max_prompt_tokens_per_reviewer.<reviewer-slug>"
|
|
174
178
|
},
|
|
179
|
+
{
|
|
180
|
+
"topLevel": "review",
|
|
181
|
+
"source": "^review\\.reviewer_instances\\.[a-zA-Z0-9_-]+\\.(cli|model|agent)$",
|
|
182
|
+
"description": "review.reviewer_instances.<instance-name>.<cli|model|agent> (#1517)"
|
|
183
|
+
},
|
|
175
184
|
{
|
|
176
185
|
"topLevel": "model_policy",
|
|
177
186
|
"source": "^model_policy\\.runtime_tiers\\.[a-zA-Z0-9_-]+\\.(opus|sonnet|haiku)$",
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# Agent Skills Self-Load (Bootstrap)
|
|
2
|
+
|
|
3
|
+
> **Shared contract.** Every `agent_skills` consumer agent self-loads its configured
|
|
4
|
+
> skills in its mandatory init step, so a project's `.planning/config.json`
|
|
5
|
+
> `agent_skills.<agent-type>` mapping reaches the agent that actually does the work —
|
|
6
|
+
> even when the orchestrator did not run bash init (e.g. a runtime whose `Skill()`
|
|
7
|
+
> delegation does not reliably execute the delegated workflow's bash, such as Cursor;
|
|
8
|
+
> see open-gsd/gsd-core#1600 / #1601). This is the durable counterpart to the
|
|
9
|
+
> orchestrator-side injection documented under
|
|
10
|
+
> [Agent Skills Injection](../../docs/CONFIGURATION.md#agent-skills-injection).
|
|
11
|
+
|
|
12
|
+
## When to run
|
|
13
|
+
|
|
14
|
+
In your mandatory init step — right after `mandatory-initial-read.md` / the
|
|
15
|
+
`Project skills` discovery, before any other work.
|
|
16
|
+
|
|
17
|
+
## Steps
|
|
18
|
+
|
|
19
|
+
1. **Dedup guard (MANDATORY).** Look at your own prompt. If it already contains an
|
|
20
|
+
`<agent_skills>` block, the orchestrator already injected one — **skip self-load
|
|
21
|
+
entirely.** Loading a second copy wastes context on runtimes where orchestrator-side
|
|
22
|
+
injection also runs (e.g. Claude Code). The guard is what keeps the two seams from
|
|
23
|
+
doubling the block.
|
|
24
|
+
|
|
25
|
+
2. **Query your configured skills.** Use **your own agent type** — the `name:` value in
|
|
26
|
+
your frontmatter (e.g. an agent whose frontmatter says `name: gsd-executor` queries
|
|
27
|
+
`gsd-executor`). The query is read-only and idempotent — it exits 0 with an empty
|
|
28
|
+
block when nothing is configured for your type:
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
_AGENT_SKILLS=$(gsd_run query agent-skills <YOUR-FRONTMATTER-NAME> 2>/dev/null || true)
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
The runtime `gsd_run` resolver is the standard one from
|
|
35
|
+
`_runtime-launcher.snippet.sh`; your own init already defines it.
|
|
36
|
+
|
|
37
|
+
3. **Read every listed skill.** The block emits entries as `@<path>/SKILL.md`
|
|
38
|
+
includes — `Read` each one before starting work. If the block is empty, there is
|
|
39
|
+
nothing to do (zero overhead).
|
|
40
|
+
|
|
41
|
+
## What self-load does and does not cover
|
|
42
|
+
|
|
43
|
+
| Skill form | Self-loads? | Notes |
|
|
44
|
+
|---|---|---|
|
|
45
|
+
| Project-relative path (`skills/my-skill`) | ✅ everywhere | `Read` the `@`-include |
|
|
46
|
+
| Global personal (`global:<name>`) | ✅ everywhere | resolves to the runtime global skills dir, then `Read` |
|
|
47
|
+
| Plugin-provided (`global:<plugin>:<skill>`) | Claude only | emitted as a Skill-tool directive on Claude; **skipped with a warning on all other runtimes** — the plugin/Skill-tool model has no equivalent elsewhere (#1601, #1258). Not closeable on Cursor. |
|
|
48
|
+
|
|
49
|
+
## Notes
|
|
50
|
+
|
|
51
|
+
- **Idempotent and read-only.** `query agent-skills` never mutates state; calling it
|
|
52
|
+
twice (once by the orchestrator, once by the agent) is harmless because the dedup
|
|
53
|
+
guard suppresses the second load.
|
|
54
|
+
- **No new config keys.** This reuses the existing `agent_skills` map and the existing
|
|
55
|
+
`buildAgentSkillsBlock` / `cmdAgentSkills` machinery (`src/init.cts`). The 22 consumer
|
|
56
|
+
agent types are mirrored in `tests/agent-skills.test.cjs` (`CONSUMER_AGENTS`) and
|
|
57
|
+
guarded against drift by `tests/agent-skills-bootstrap.test.cjs`.
|
|
58
|
+
- **Checkers / read-only agents.** Bash is universal across consumer agents, so
|
|
59
|
+
self-load works for plan-checkers, verifiers, and auditors too; the bootstrap assumes
|
|
60
|
+
no tool an agent lacks.
|
|
@@ -0,0 +1,105 @@
|
|
|
1
|
+
# Honest Verifier — Abstention on Non-Inferable Checks
|
|
2
|
+
|
|
3
|
+
Shared reference for the **verify** phase. The verify-time companion to the spec-time
|
|
4
|
+
`@~/.claude/gsd-core/references/edge-probe.md` (which *classifies* non-inferable checks) and
|
|
5
|
+
`@~/.claude/gsd-core/references/prohibition-probe.md` (whose judgment-tier disposition this mirrors).
|
|
6
|
+
This doc is written in generic `spec → predicate → verifier` terms with no tool-specific vocabulary,
|
|
7
|
+
so it is portable: copy it into any verification process.
|
|
8
|
+
|
|
9
|
+
## The problem it solves
|
|
10
|
+
|
|
11
|
+
A verifier is trustworthy on **inferable** checks — defects determined by the stated spec. On a
|
|
12
|
+
**non-inferable** check the correct answer is *not derivable from the spec alone* (e.g. "does `[1,2]`
|
|
13
|
+
touching `[2,3]` merge?", "is a 'character' a grapheme or a code unit?"). On these the verifier *does
|
|
14
|
+
not know that it does not know*: measured behavior is a **confident PASS on the blind-spot check ~100%
|
|
15
|
+
of the time** (mean confidence ~0.93), because a model cannot self-detect a gap it does not perceive.
|
|
16
|
+
|
|
17
|
+
The edge-probe already detects these at spec time and tags them `verification: backstop` (ADR-550
|
|
18
|
+
D7a). The honest verifier consumes that tag so the verifier **abstains** instead of confidently
|
|
19
|
+
false-passing — converting a silent false-pass (the worst failure: you don't know to look) into an
|
|
20
|
+
explicit, actionable "write a held-out test." Measured: the confident-false-pass rate on the blind
|
|
21
|
+
spot drops **100% → 17%** (N17).
|
|
22
|
+
|
|
23
|
+
## The two properties that define the design
|
|
24
|
+
|
|
25
|
+
1. **Exogenous, not endogenous.** The trigger is the *external tag* (`backstop`), never the verifier's
|
|
26
|
+
self-judgment. Asking the verifier to "abstain if unsure" barely moves the number (100% → 67%) and
|
|
27
|
+
only on ambiguity it already notices; on a true blind spot it stays confidently wrong. A confidence
|
|
28
|
+
gate cannot reach a blind spot the model does not feel — so there is **no "are you sure?" prompt**;
|
|
29
|
+
routing is on the pre-existing tag only.
|
|
30
|
+
2. **Routing, not diagnosis.** The verifier need not name the omitted rule (if it could, it wouldn't
|
|
31
|
+
be a blind spot). In testing, verifiers abstained correctly while citing the *wrong* edge. The
|
|
32
|
+
honest verdict requires only "I was told this is under-specified and I cannot rule it out." The
|
|
33
|
+
omitted rule is carried by a human-authored held-out test, not by the verifier.
|
|
34
|
+
|
|
35
|
+
## The disposition (the protocol)
|
|
36
|
+
|
|
37
|
+
For each `must_haves.truths` item:
|
|
38
|
+
|
|
39
|
+
| Item | Confirmable with explicit evidence? | Disposition |
|
|
40
|
+
|---|---|---|
|
|
41
|
+
| Inferable (plain string, or `verification: explicit`) | n/a — graded normally | ✓ VERIFIED / ✗ FAILED as usual; **never abstained** (over-abstention guard) |
|
|
42
|
+
| Non-inferable (`verification: backstop`) | **yes** (a wired held-out/property-based test that passes, or a directly-observed behavior) | ✓ VERIFIED |
|
|
43
|
+
| Non-inferable (`verification: backstop`) | **no** | **abstain** → ⚠️ `insufficient_spec`, flagged, → `human_needed` — **never `passed`** |
|
|
44
|
+
|
|
45
|
+
- **Explicit evidence** = a wired held-out/property-based test that passes, or a behavior the verifier
|
|
46
|
+
directly observed. Symbol presence + wiring is **not** explicit evidence for a non-inferable truth.
|
|
47
|
+
- **Never silent, never a hard halt.** *Interactive:* the abstained item routes to the end-of-phase
|
|
48
|
+
human checkpoint. *Autonomous (AFK):* it produces a prominent `unverified — held-out test
|
|
49
|
+
recommended` flag and the completion line reads "complete with N unverified non-inferable checks";
|
|
50
|
+
the run neither silently passes the blind spot nor hard-halts.
|
|
51
|
+
- **Distinguishable reason.** The abstain disposition carries `reason: insufficient_spec` so the
|
|
52
|
+
`human_needed` outcome is never conflated with an ordinary manual-UAT `human_needed`.
|
|
53
|
+
|
|
54
|
+
This is the verify-time half of ADR-550 Decision 4 (the never-silent-pass disposition), applied to the
|
|
55
|
+
edge `backstop` truth tier instead of the prohibition judgment tier — the same machinery, opposite
|
|
56
|
+
polarity (must-HAVE under-specified vs must-NOT irreducible).
|
|
57
|
+
|
|
58
|
+
## Deterministic engine surface
|
|
59
|
+
|
|
60
|
+
The CI-testable surface is the **deterministic disposition + projection**, never the LLM's judgment
|
|
61
|
+
(ADR-550 D5 — a test asserting the model's verdict is vacuous and rejected). In `probe-core`:
|
|
62
|
+
|
|
63
|
+
- `truthStatement(t)` / `truthVerification(t)` — normalizers; read a truth's statement and tier from
|
|
64
|
+
either the plain-string or object form (a truth-reader MUST normalize, never assume a string).
|
|
65
|
+
- `projectTruths(items)` — conservative serializer: a `backstop` truth → flat-scalar object
|
|
66
|
+
`{ statement, verification: backstop }`; every inferable truth → a bare string.
|
|
67
|
+
- `dispositionForUnverifiableTruth(truth, { evidence })` → `{ status, flagged, tier, reason }`:
|
|
68
|
+
`backstop` + no evidence → `unverified`/`flagged`/`insufficient_spec`; `backstop` + evidence →
|
|
69
|
+
`green`; non-`backstop` → `green` (over-abstention guard).
|
|
70
|
+
|
|
71
|
+
## Capable-tier requirement (a documented cost)
|
|
72
|
+
|
|
73
|
+
Abstention is **model-tier dependent** and this is a standing cost, not an assumption:
|
|
74
|
+
|
|
75
|
+
- The default `gsd-verifier` tier (`sonnet`, golden/balanced) heeds the exogenous tag reliably
|
|
76
|
+
(2/2 under testing).
|
|
77
|
+
- The **budget tier (`haiku`)** is the least flag-responsive (1/2, inconsistent) and **degrades toward
|
|
78
|
+
current behavior** (confident false-pass). Run honest-verifier on a capable tier; treat the budget
|
|
79
|
+
tier as best-effort. Re-validate when the `gsd-verifier` model tier changes or a new budget model is
|
|
80
|
+
adopted (captured as a test so a tier regression is caught, not discovered in production).
|
|
81
|
+
|
|
82
|
+
## Evidence and scope (stated honestly)
|
|
83
|
+
|
|
84
|
+
- **Evidence strength.** N17 is n=27 verdicts (3 models × 3 conditions × 3 tasks), 1 rep —
|
|
85
|
+
**direction-finding, not powered.** The blind-spot effect is large and monotone
|
|
86
|
+
(100% → 67% → 17%); the two costs are clean single events (a *false* tag made the strongest model
|
|
87
|
+
over-abstain on a real spec-determined bug; the weakest tier was flag-deaf) and they name exactly
|
|
88
|
+
the failure modes the over-abstention guard and the capable-tier requirement defend against.
|
|
89
|
+
- **Tag-precision coupling.** Quality is bounded by the edge-probe's `backstop` recall/precision — a
|
|
90
|
+
false non-inferable flag causes over-abstention. Positive coupling: improving the probe (#1110)
|
|
91
|
+
improves this for free. It adds no independent burden.
|
|
92
|
+
- **Explicit non-goals.** Does NOT identify the omitted rule; does NOT recalibrate decisive verdicts;
|
|
93
|
+
does NOT defend against *malicious compliance* (a self-graded review rationalizing away its own
|
|
94
|
+
findings). It raises the floor on *honest* uncertainty about non-inferable checks — that is the
|
|
95
|
+
whole claim.
|
|
96
|
+
|
|
97
|
+
## Distinct from neighbours
|
|
98
|
+
|
|
99
|
+
- **vs `PRESENT_BEHAVIOR_UNVERIFIED` (#966 axis):** that is the *inferable-but-unobserved* case — the
|
|
100
|
+
truth **can** be verified from the spec but was shortcut-passed on symbol presence; the fix is to
|
|
101
|
+
demand behavioral evidence. Honest-verifier is the *non-inferable* case — the truth **cannot** be
|
|
102
|
+
verified from the spec at all; the fix is to abstain and route to a held-out test. Orthogonal axes
|
|
103
|
+
(insufficient *evidence* vs insufficient *spec*); both feed the same `human_needed` sink.
|
|
104
|
+
- **vs prohibition judgment-tier (#644):** that disposes **must-NOT** constraints; honest-verifier
|
|
105
|
+
disposes **non-inferable positive truths**. Opposite polarity, same never-silent disposition.
|
|
@@ -133,6 +133,33 @@ If you're using Claude Code with OpenRouter, a local model, or any non-Anthropic
|
|
|
133
133
|
|
|
134
134
|
Without `inherit`, GSD's default `balanced` profile spawns specific Anthropic models (`opus`, `sonnet`, `haiku`) for each agent type, which can result in additional API costs through your non-Anthropic provider.
|
|
135
135
|
|
|
136
|
+
## Advisor Tool (Claude Code)
|
|
137
|
+
|
|
138
|
+
Claude Code (v2.1.98+) can pair the session's executor model with a stronger **advisor** model that it consults mid-generation for strategy and course-correction (Anthropic's [advisor tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool)). This is a host-runtime feature, not a GSD setting — GSD selects each agent's *executor* model through the profile/tier system above; Claude Code supplies the advisor.
|
|
139
|
+
|
|
140
|
+
Set it once at the session level with `/advisor <model>` (or the `advisorModel` setting / `--advisor` flag). **Subagents inherit the session advisor automatically**, so every GSD subagent an orchestrator spawns gets the same advisor with no per-agent configuration. It composes cleanly with GSD's tiering: the profile keeps executors cheap where the work is mechanical, and the advisor adds a stronger reviewer inline on the turns that benefit.
|
|
141
|
+
|
|
142
|
+
### Candidate pairings
|
|
143
|
+
|
|
144
|
+
Per Anthropic's advisor-tool docs the advisor must be at least as capable as the executor. Candidate pairings by profile — evaluate on your own workload; the quality/cost characterizations below are Anthropic-reported, not GSD guarantees:
|
|
145
|
+
|
|
146
|
+
| Profile | Typical executors | Candidate advisor | Rationale (per Anthropic docs) |
|
|
147
|
+
|---|---|---|---|
|
|
148
|
+
| `budget` | Haiku / Sonnet | Fable 5 or Opus | A step up in intelligence over Haiku alone, at lower cost than switching the executor to a larger model |
|
|
149
|
+
| `balanced` | Sonnet | Fable 5 or Opus | A quality lift at similar or lower total cost than Sonnet-solo on complex tasks |
|
|
150
|
+
| `quality` / `adaptive` | Opus (planning), Sonnet | Fable 5 or Opus | Marginal on turns already at top capability; most valuable on the Sonnet-executor agents |
|
|
151
|
+
|
|
152
|
+
Fable 5 is a valid advisor for Haiku 4.5, Sonnet 4.6/5, and Opus 4.8 executors, so it pairs with any tier a profile assigns.
|
|
153
|
+
|
|
154
|
+
### When it's worth enabling
|
|
155
|
+
|
|
156
|
+
- **Worth it:** long, multi-step agent loops where the plan matters but most turns are mechanical — e.g. `execute-phase` and `debug`. Anthropic's docs note advisor prompt-caching pays off at roughly three or more advisor calls, which these long loops make.
|
|
157
|
+
- **Skip it:** short, one-shot agents (mappers, quick audits, single-file checks) — there is little to plan, and the advisor adds cost without a commensurate quality gain.
|
|
158
|
+
|
|
159
|
+
### Constraint: session-level only (today)
|
|
160
|
+
|
|
161
|
+
The advisor is a single session-wide setting inherited by all subagents; there is **no per-agent advisor selection**, so GSD cannot vary the advisor by role the way it varies the executor model (e.g. "no advisor on the Haiku mapper, a Fable 5 advisor on the Sonnet executor"). Per-agent advisor control is tracked upstream at [anthropics/claude-code#73072](https://github.com/anthropics/claude-code/issues/73072); until it lands, pick one session advisor that fits the most valuable agents in your run.
|
|
162
|
+
|
|
136
163
|
## Dynamic Routing with Failure-Tier Escalation (#3024)
|
|
137
164
|
|
|
138
165
|
When `dynamic_routing.enabled = true` in `.planning/config.json`, the resolver picks a model from a tier-mapped table based on the agent's *default tier* (light / standard / heavy) and escalates to the next tier up on orchestrator-detected soft failure.
|
|
@@ -0,0 +1,99 @@
|
|
|
1
|
+
# Reviewer Instances (#1517)
|
|
2
|
+
|
|
3
|
+
Custom reviewer instances for `/gsd:review`: run one model-capable adapter (e.g. OpenCode)
|
|
4
|
+
as several independent reviewer identities in a single review pass. Loaded lazily by
|
|
5
|
+
`gsd-core/workflows/review.md` when `review.reviewer_instances` is configured. See
|
|
6
|
+
[ADR-1517](../docs/adr/1517-reviewer-instances-config-surface.md) for the contract.
|
|
7
|
+
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
## Config shape
|
|
11
|
+
|
|
12
|
+
A `review.reviewer_instances` object under the `review` namespace. Each entry maps an
|
|
13
|
+
instance name to `{ cli, model?, agent? }`:
|
|
14
|
+
|
|
15
|
+
```json
|
|
16
|
+
{
|
|
17
|
+
"review": {
|
|
18
|
+
"reviewer_instances": {
|
|
19
|
+
"opencode-deepseek": { "cli": "opencode", "model": "deepseek/deepseek-v4-pro", "agent": "review" },
|
|
20
|
+
"opencode-mimo": { "cli": "opencode", "model": "xiaomi/mimo-v2.5-pro" }
|
|
21
|
+
},
|
|
22
|
+
"default_reviewers": ["opencode-deepseek", "opencode-mimo", "codex"]
|
|
23
|
+
}
|
|
24
|
+
}
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
- Instance name: `^[a-z0-9][a-z0-9-]*$`, must not equal a built-in slug. Validated at
|
|
28
|
+
`config-set` time.
|
|
29
|
+
- `cli`: MUST be a known adapter (`KNOWN_REVIEWER_SLUGS`) — never an arbitrary shell command.
|
|
30
|
+
- `model`: opaque `provider/model` string, passed through verbatim. GSD does not parse it.
|
|
31
|
+
- `agent`: opaque string; honoured only by adapters with a native agent concept (OpenCode
|
|
32
|
+
`--agent` in v1). Ignored by other adapters.
|
|
33
|
+
|
|
34
|
+
---
|
|
35
|
+
|
|
36
|
+
## Resolution rules (single source)
|
|
37
|
+
|
|
38
|
+
The canonical logic lives in `resolveReviewerSelection` / `normalizeReviewerInstances` in
|
|
39
|
+
`review-reviewer-selection.cjs`. Apply the SAME rules in the workflow so the two surfaces
|
|
40
|
+
cannot diverge (`DEFECT.GENERATIVE-FIX`; parity-locked in
|
|
41
|
+
`tests/review-reviewer-instances.test.cjs`).
|
|
42
|
+
|
|
43
|
+
1. Instances participate ONLY via `review.default_reviewers`. They never appear under `--all`
|
|
44
|
+
or explicit `--<cli>` flags, and there are no per-instance CLI flags.
|
|
45
|
+
2. Expand instance references BEFORE the built-in-slug check: an entry that is a key in
|
|
46
|
+
`review.reviewer_instances` is an **instance**; an entry that is a built-in slug is a
|
|
47
|
+
**builtin**.
|
|
48
|
+
3. An instance is **available** iff its base `cli` is detected (e.g. `opencode-deepseek` is
|
|
49
|
+
available iff `opencode` is available).
|
|
50
|
+
4. An entry that is NEITHER a defined instance NOR a built-in slug is a **hard error** (likely
|
|
51
|
+
a typo'd instance name) — stop and report it. Do NOT silently drop it. (When
|
|
52
|
+
`review.reviewer_instances` is absent entirely, fall back to the legacy unknown-slug
|
|
53
|
+
warn-and-drop behaviour for backward compatibility.)
|
|
54
|
+
5. `model`/`agent`/instance-name are opaque: pass them as separate argv elements. They are
|
|
55
|
+
NEVER interpolated into shell strings.
|
|
56
|
+
|
|
57
|
+
---
|
|
58
|
+
|
|
59
|
+
## Invocation
|
|
60
|
+
|
|
61
|
+
For each selected INSTANCE, invoke its base `cli` using the instance's own `model`/`agent` —
|
|
62
|
+
NOT the global `review.models.<cli>`. Each instance writes to its OWN per-instance output file
|
|
63
|
+
and runs as a distinct reviewer identity.
|
|
64
|
+
|
|
65
|
+
For an OpenCode-backed instance (the motivating adapter):
|
|
66
|
+
|
|
67
|
+
```bash
|
|
68
|
+
# $INSTANCE_MODEL / $INSTANCE_AGENT come from the instance spec; $INSTANCE_NAME is the
|
|
69
|
+
# reviewer identity (e.g. opencode-deepseek). --agent is OpenCode's native subagent flag;
|
|
70
|
+
# omit it when the instance has no agent.
|
|
71
|
+
if [ -n "$INSTANCE_AGENT" ] && [ "$INSTANCE_AGENT" != "null" ]; then
|
|
72
|
+
cat /tmp/gsd-review-prompt-{phase}.md | opencode run --model "$INSTANCE_MODEL" --agent "$INSTANCE_AGENT" - 2>/dev/null > /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md
|
|
73
|
+
else
|
|
74
|
+
cat /tmp/gsd-review-prompt-{phase}.md | opencode run --model "$INSTANCE_MODEL" - 2>/dev/null > /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md
|
|
75
|
+
fi
|
|
76
|
+
if [ ! -s /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md ]; then
|
|
77
|
+
echo "OpenCode review ($INSTANCE_NAME) failed or returned empty output." > /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md
|
|
78
|
+
fi
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
For an instance backed by a DIFFERENT cli, reuse that cli's invocation block with two
|
|
82
|
+
substitutions: use the instance's `model` in place of the global `review.models.<cli>` value,
|
|
83
|
+
and write to `/tmp/gsd-review-${INSTANCE_NAME}-{phase}.md`. Only `opencode` honours an
|
|
84
|
+
`agent` field in v1; ignore `agent` for other adapters.
|
|
85
|
+
|
|
86
|
+
---
|
|
87
|
+
|
|
88
|
+
## REVIEWS.md contract
|
|
89
|
+
|
|
90
|
+
- **Frontmatter `reviewers:`** records the actual identities invoked. For a built-in slug use
|
|
91
|
+
the slug (`opencode`); for an instance use the instance name (`opencode-deepseek`), so
|
|
92
|
+
frontmatter distinguishes the independent voices. Example:
|
|
93
|
+
`reviewers: [opencode-deepseek, opencode-mimo, codex]`.
|
|
94
|
+
- **Section headers:** each instance gets its OWN top-level section, headed with the base
|
|
95
|
+
adapter's display name plus the instance name in parentheses:
|
|
96
|
+
`## OpenCode Review (opencode-deepseek)`. Same-cli instances are never collapsed.
|
|
97
|
+
- **Shared-adapter caveat:** when ≥2 invoked instances share the same base `cli`, print a
|
|
98
|
+
one-line caveat immediately after the frontmatter (before the first section), e.g.:
|
|
99
|
+
`> Note: opencode-deepseek and opencode-mimo share the opencode adapter; their consensus is cross-model, not cross-tool.`
|