@opengsd/gsd-core 1.7.0 → 1.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (165) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.claude-plugin/plugin.json +1 -1
  3. package/.opencode/plugins/gsd-core.js +14 -0
  4. package/README.md +2 -0
  5. package/agents/gsd-debug-session-manager.md +42 -4
  6. package/agents/gsd-debugger.md +87 -29
  7. package/agents/gsd-executor.md +29 -2
  8. package/agents/gsd-planner.md +29 -36
  9. package/agents/gsd-verifier.md +2 -2
  10. package/bin/install.js +1152 -80
  11. package/commands/gsd/ai-integration-phase.md +1 -1
  12. package/commands/gsd/mempalace-capture.md +9 -5
  13. package/commands/gsd/new-milestone.md +1 -1
  14. package/commands/gsd/plan-phase.md +5 -3
  15. package/commands/gsd/plan-review-convergence.md +3 -2
  16. package/gsd-core/bin/gsd-tools.cjs +1878 -2507
  17. package/gsd-core/bin/lib/adapter-imperative.cjs +8 -1
  18. package/gsd-core/bin/lib/agent-command-router.cjs +20 -5
  19. package/gsd-core/bin/lib/api-coverage.cjs +338 -45
  20. package/gsd-core/bin/lib/broken-windows.cjs +716 -0
  21. package/gsd-core/bin/lib/capability-command-router.cjs +733 -0
  22. package/gsd-core/bin/lib/capability-registry.cjs +155 -86
  23. package/gsd-core/bin/lib/capability-writer.cjs +6 -1
  24. package/gsd-core/bin/lib/check-command-router.cjs +128 -25
  25. package/gsd-core/bin/lib/claude-orchestration-command-router.cjs +115 -27
  26. package/gsd-core/bin/lib/claude-orchestration.cjs +84 -9
  27. package/gsd-core/bin/lib/command-aliases.cjs +14 -0
  28. package/gsd-core/bin/lib/commands.cjs +81 -4
  29. package/gsd-core/bin/lib/config-loader.cjs +14 -2
  30. package/gsd-core/bin/lib/config.cjs +69 -18
  31. package/gsd-core/bin/lib/core-utils.cjs +6 -1
  32. package/gsd-core/bin/lib/decisions.cjs +32 -8
  33. package/gsd-core/bin/lib/docs.cjs +6 -0
  34. package/gsd-core/bin/lib/external-descriptor-trust.cjs +14 -2
  35. package/gsd-core/bin/lib/gap-checker.cjs +17 -2
  36. package/gsd-core/bin/lib/init.cjs +111 -47
  37. package/gsd-core/bin/lib/install-engine.cjs +298 -23
  38. package/gsd-core/bin/lib/install-profiles.cjs +239 -1
  39. package/gsd-core/bin/lib/installer-migrations/005-opencode-baseline-commands-dir.cjs +146 -0
  40. package/gsd-core/bin/lib/installer-migrations/006-pi-extension-cjs-to-js.cjs +91 -0
  41. package/gsd-core/bin/lib/installer-migrations.cjs +44 -5
  42. package/gsd-core/bin/lib/markdown-sectionizer.cjs +107 -0
  43. package/gsd-core/bin/lib/milestone.cjs +246 -12
  44. package/gsd-core/bin/lib/model-catalog.cjs +19 -4
  45. package/gsd-core/bin/lib/model-resolver.cjs +189 -7
  46. package/gsd-core/bin/lib/onboard-projection.cjs +11 -8
  47. package/gsd-core/bin/lib/phase-id.cjs +26 -4
  48. package/gsd-core/bin/lib/phase.cjs +201 -12
  49. package/gsd-core/bin/lib/plan-scan.cjs +70 -2
  50. package/gsd-core/bin/lib/roadmap-parser.cjs +7 -4
  51. package/gsd-core/bin/lib/roadmap.cjs +13 -3
  52. package/gsd-core/bin/lib/runtime-artifact-conversion.cjs +7 -1
  53. package/gsd-core/bin/lib/runtime-artifact-layout.cjs +22 -8
  54. package/gsd-core/bin/lib/runtime-hooks-surface.cjs +16 -0
  55. package/gsd-core/bin/lib/smart-entry.cjs +69 -4
  56. package/gsd-core/bin/lib/state-document.cjs +7 -4
  57. package/gsd-core/bin/lib/state-transition.cjs +22 -1
  58. package/gsd-core/bin/lib/state.cjs +65 -11
  59. package/gsd-core/bin/lib/surface.cjs +51 -9
  60. package/gsd-core/bin/lib/uat.cjs +420 -5
  61. package/gsd-core/bin/lib/validate.cjs +12 -8
  62. package/gsd-core/bin/lib/verification.cjs +112 -17
  63. package/gsd-core/bin/lib/verify.cjs +220 -22
  64. package/gsd-core/bin/shared/config-schema.manifest.json +3 -2
  65. package/gsd-core/references/api-coverage.md +37 -7
  66. package/gsd-core/references/checkpoints.md +1 -1
  67. package/gsd-core/references/common-bug-patterns.md +13 -0
  68. package/gsd-core/references/debugger-bug-taxonomy.md +111 -0
  69. package/gsd-core/references/debugger-fix-acceptance.md +157 -0
  70. package/gsd-core/references/debugger-philosophy.md +1 -0
  71. package/gsd-core/references/debugger-prevention.md +98 -0
  72. package/gsd-core/references/debugger-rca-branching.md +98 -0
  73. package/gsd-core/references/debugger-repro-hardening.md +130 -0
  74. package/gsd-core/references/debugger-sbfl.md +110 -0
  75. package/gsd-core/references/debugger-semantic-recall.md +81 -0
  76. package/gsd-core/references/execute-phase-quota-recovery.md +55 -0
  77. package/gsd-core/references/execute-phase-requirement-revert.md +8 -0
  78. package/gsd-core/references/execute-phase-response-language.md +7 -0
  79. package/gsd-core/references/planner-antipatterns.md +6 -0
  80. package/gsd-core/references/planner-mvp-mode.md +12 -13
  81. package/gsd-core/references/planner-preconditions.md +156 -0
  82. package/gsd-core/references/planner-reversibility.md +132 -0
  83. package/gsd-core/references/reviewer-instances.md +9 -7
  84. package/gsd-core/references/skeleton-template.md +1 -1
  85. package/gsd-core/references/thinking-models-planning.md +3 -1
  86. package/gsd-core/templates/DEBUG.md +5 -3
  87. package/gsd-core/workflows/add-phase.md +2 -0
  88. package/gsd-core/workflows/add-tests.md +3 -1
  89. package/gsd-core/workflows/add-todo.md +32 -1
  90. package/gsd-core/workflows/ai-integration-phase.md +4 -2
  91. package/gsd-core/workflows/audit-fix.md +2 -2
  92. package/gsd-core/workflows/check-todos.md +3 -1
  93. package/gsd-core/workflows/cleanup.md +7 -1
  94. package/gsd-core/workflows/code-review.md +17 -5
  95. package/gsd-core/workflows/complete-milestone.md +3 -0
  96. package/gsd-core/workflows/debug.md +25 -5
  97. package/gsd-core/workflows/diagnose-issues.md +1 -1
  98. package/gsd-core/workflows/discovery-phase.md +7 -0
  99. package/gsd-core/workflows/discuss-phase/templates/context.md +16 -2
  100. package/gsd-core/workflows/discuss-phase-assumptions.md +3 -0
  101. package/gsd-core/workflows/do.md +7 -1
  102. package/gsd-core/workflows/docs-update.md +1 -0
  103. package/gsd-core/workflows/eval-review.md +3 -0
  104. package/gsd-core/workflows/execute-phase/steps/post-merge-gate.md +4 -4
  105. package/gsd-core/workflows/execute-phase/steps/regression-gate.md +2 -2
  106. package/gsd-core/workflows/execute-phase.md +25 -34
  107. package/gsd-core/workflows/execute-plan.md +15 -4
  108. package/gsd-core/workflows/graduation.md +3 -0
  109. package/gsd-core/workflows/health.md +7 -1
  110. package/gsd-core/workflows/help/modes/full.md +6 -2
  111. package/gsd-core/workflows/import.md +8 -2
  112. package/gsd-core/workflows/inbox.md +7 -0
  113. package/gsd-core/workflows/ingest-docs.md +15 -10
  114. package/gsd-core/workflows/manager.md +3 -1
  115. package/gsd-core/workflows/map-codebase.md +4 -4
  116. package/gsd-core/workflows/mvp-phase.md +3 -0
  117. package/gsd-core/workflows/new-milestone.md +69 -21
  118. package/gsd-core/workflows/new-project.md +17 -15
  119. package/gsd-core/workflows/new-workspace.md +3 -1
  120. package/gsd-core/workflows/onboard.md +3 -0
  121. package/gsd-core/workflows/plan-phase.md +14 -5
  122. package/gsd-core/workflows/plan-review-convergence.md +48 -3
  123. package/gsd-core/workflows/plant-seed.md +3 -0
  124. package/gsd-core/workflows/profile-user.md +7 -1
  125. package/gsd-core/workflows/progress.md +31 -3
  126. package/gsd-core/workflows/quick.md +19 -7
  127. package/gsd-core/workflows/remove-workspace.md +3 -0
  128. package/gsd-core/workflows/review.md +89 -73
  129. package/gsd-core/workflows/scan.md +1 -1
  130. package/gsd-core/workflows/secure-phase.md +3 -0
  131. package/gsd-core/workflows/settings-integrations.md +3 -0
  132. package/gsd-core/workflows/settings.md +3 -0
  133. package/gsd-core/workflows/ship.md +50 -3
  134. package/gsd-core/workflows/sketch.md +3 -0
  135. package/gsd-core/workflows/smart-entry.md +3 -0
  136. package/gsd-core/workflows/spike.md +7 -1
  137. package/gsd-core/workflows/ui-phase.md +3 -1
  138. package/gsd-core/workflows/ui-review.md +3 -0
  139. package/gsd-core/workflows/undo.md +7 -0
  140. package/gsd-core/workflows/update.md +2 -0
  141. package/gsd-core/workflows/validate-phase.md +3 -0
  142. package/gsd-core/workflows/verify-phase.md +2 -2
  143. package/gsd-core/workflows/verify-work.md +7 -3
  144. package/hooks/dist/gsd-context-monitor.js +27 -9
  145. package/hooks/dist/gsd-statusline.js +88 -3
  146. package/hooks/gsd-context-monitor.js +27 -9
  147. package/hooks/gsd-statusline.js +88 -3
  148. package/package.json +6 -4
  149. package/pi/gsd.cjs +8 -2
  150. package/scripts/changeset/lint.cjs +1 -0
  151. package/scripts/changeset/parse.cjs +26 -0
  152. package/scripts/check-glossary-refs.cjs +220 -0
  153. package/scripts/ci-rebase-check.cjs +48 -4
  154. package/scripts/gen-adr-index.cjs +526 -0
  155. package/scripts/gen-test-timings.cjs +201 -0
  156. package/scripts/lint-portable-timeout.cjs +140 -0
  157. package/scripts/lint-test-file-count.allowlist.json +1 -0
  158. package/scripts/release-tarball-smoke.cjs +18 -11
  159. package/scripts/run-tests.cjs +420 -58
  160. package/skills/gsd-ai-integration-phase/SKILL.md +1 -1
  161. package/skills/gsd-mempalace-capture/SKILL.md +9 -5
  162. package/skills/gsd-new-milestone/SKILL.md +1 -1
  163. package/skills/gsd-plan-phase/SKILL.md +5 -3
  164. package/skills/gsd-plan-review-convergence/SKILL.md +3 -2
  165. package/vscode/package.json +1 -1
@@ -15,9 +15,14 @@
15
15
  // node scripts/run-tests.cjs --files-from /tmp/selected-tests.txt
16
16
  // node scripts/run-tests.cjs --suite unit --shard 1/3 # shard 1 of 3 (#1212)
17
17
  //
18
- // Sharding (issue #1212): --shard <i>/<n> runs a deterministic, balanced
19
- // round-robin slice of the SORTED selected file list (file index k → shard
20
- // k % n). i is 1-based (1..n); n >= 1; n=1 is a pure no-op (all files). The
18
+ // Sharding (issue #1212, reweighted #2472): --shard <i>/<n> runs a
19
+ // deterministic, COST-balanced slice of the SORTED selected file list. Files
20
+ // are partitioned by measured duration (tests/test-timings.json) using LPT —
21
+ // the same packing the chunker uses one level down — because equal file COUNTS
22
+ // are not equal file COST: the index-based split this replaced ran 12.4m /
23
+ // 19.2m / 15.2m against a 20-minute job cap. With no timing data every file
24
+ // weighs the same and the partition degenerates to the original k % n
25
+ // round-robin. i is 1-based (1..n); n >= 1; n=1 is a pure no-op (all files). The
21
26
  // CI windows full-test lane shards across N parallel runners so per-job
22
27
  // wall-clock scales as O(total/N) and stops hitting the job time cap. Sharding
23
28
  // composes with --suite (it slices the post-filter selection) and preserves
@@ -29,7 +34,7 @@
29
34
  // See docs/TESTING-SUITES.md for full grouping policy.
30
35
  'use strict';
31
36
 
32
- const { readdirSync } = require('fs');
37
+ const { readdirSync, readFileSync } = require('fs');
33
38
  const { join, basename } = require('path');
34
39
  const { execFileSync } = require('child_process');
35
40
  const { ExitError, runMain } = require('./lib/cli-exit.cjs');
@@ -206,7 +211,8 @@ function parseShardArg(value) {
206
211
  return { index, total };
207
212
  }
208
213
 
209
- // Deterministic, balanced round-robin partition of an ALREADY-SORTED file list.
214
+ // Deterministic partition of an ALREADY-SORTED file list. Without a weigher
215
+ // this is the original round-robin (#1212):
210
216
  // Shard `index` (1-based) receives every file whose position k in the sorted
211
217
  // list satisfies k % total === index - 1. Round-robin (not contiguous blocks)
212
218
  // spreads duration variance across shards and guarantees shard sizes differ by
@@ -215,9 +221,269 @@ function parseShardArg(value) {
215
221
  // sorts the list with the same (locale-independent) comparator. `total=1`
216
222
  // returns the input unchanged (pure no-op). A shard with no files (total >
217
223
  // file count) returns [] and is a legitimate result, not an error.
218
- function selectShard(sortedFiles, { index, total }) {
224
+ // `weightOf` (optional, #2472) switches the partition from equal COUNTS to
225
+ // equal COST. Equal counts were only ever a proxy for equal duration, and on a
226
+ // right-skewed suite the proxy fails: the real unit suite partitioned 12.4m /
227
+ // 19.2m / 15.2m by index against a 20-minute job cap, and because assignment
228
+ // keyed off array POSITION, inserting one test file re-indexed every file after
229
+ // it and could tip the heaviest shard over. Weighting by measured cost fixes
230
+ // both: LPT bounds the heaviest shard at 4/3 of optimal, and placement follows
231
+ // a file's cost rather than its neighbours' names.
232
+ //
233
+ // This is the same algorithm packChunks uses one level down (#2456/#2463), so
234
+ // both layers now share one cost model. Omitting `weightOf` keeps the legacy
235
+ // round-robin byte-identical — callers with no timing data lose nothing.
236
+ function selectShard(sortedFiles, { index, total }, weightOf) {
219
237
  if (total === 1) return sortedFiles;
220
- return sortedFiles.filter((_, k) => k % total === index - 1);
238
+ if (typeof weightOf !== 'function') {
239
+ return sortedFiles.filter((_, k) => k % total === index - 1);
240
+ }
241
+ // A non-finite or negative weight must not poison bin arithmetic — one NaN
242
+ // would make every subsequent comparison false and pile the rest of the suite
243
+ // into bin 0. Mirrors packChunks' safeWeight for the same reason.
244
+ const safeWeight = (file) => {
245
+ const w = weightOf(file);
246
+ return Number.isFinite(w) && w >= 0 ? w : 0;
247
+ };
248
+ const bins = Array.from({ length: total }, () => ({ weight: 0, picks: [] }));
249
+ // LPT: heaviest first, each into the currently-lightest bin. Ties break on
250
+ // the caller's sort position, and the lightest-bin scan takes the FIRST
251
+ // minimum, so the partition is byte-identical across Windows/macOS/Linux —
252
+ // the same determinism guarantee the round-robin path carries.
253
+ const order = sortedFiles
254
+ .map((file, k) => ({ k, weight: safeWeight(file) }))
255
+ .sort((a, b) => b.weight - a.weight || a.k - b.k);
256
+ for (const entry of order) {
257
+ let lightest = 0;
258
+ for (let i = 1; i < total; i += 1) {
259
+ const bin = bins[i];
260
+ const best = bins[lightest];
261
+ // Weight first, then FILE COUNT. The count tiebreak is load-bearing, not
262
+ // cosmetic: adding a zero-weight file leaves its bin's weight unchanged,
263
+ // so on weight alone bin 0 stays tied-minimum forever and every
264
+ // zero-weight file lands on it — all-zero weights put the whole suite on
265
+ // shard 1 and leave the other runners idle. Zero weights are reachable
266
+ // via safeWeight's clamp (a NaN/negative/Infinity entry in a hand-edited
267
+ // or corrupted timings table) and via any genuinely 0ms measurement, so
268
+ // the clamp above would otherwise reproduce the exact pile-onto-bin-0
269
+ // failure it exists to prevent. Counting picks makes ties rotate.
270
+ if (bin.weight < best.weight
271
+ || (bin.weight === best.weight && bin.picks.length < best.picks.length)) {
272
+ lightest = i;
273
+ }
274
+ }
275
+ bins[lightest].weight += entry.weight;
276
+ bins[lightest].picks.push(entry.k);
277
+ }
278
+ // Restore the caller's order within the shard: downstream chunking and argv
279
+ // batching assume the list arrives sorted as the caller sorted it.
280
+ return bins[index - 1].picks.sort((a, b) => a - b).map((k) => sortedFiles[k]);
281
+ }
282
+
283
+ // Read an operator-supplied numeric env knob, falling back to the default for
284
+ // anything that is not a positive finite number.
285
+ //
286
+ // This is a strict-input boundary (Postel's Law: a typo must fail SAFE, not
287
+ // silently poison arithmetic downstream). `Number('abc')` is NaN and
288
+ // `Number('')` is 0, and both are load-bearing here: a NaN chunk budget makes
289
+ // the chunk-count computation NaN, which spins packChunks' retry loop forever
290
+ // (a hung CI job with no output); a zero budget makes it Infinity, which throws
291
+ // `RangeError: Invalid array length`. Neither is an acceptable response to a
292
+ // mistyped environment variable.
293
+ function positiveNumberEnv(raw, fallback) {
294
+ if (raw === undefined || raw === null || String(raw).trim() === '') return fallback;
295
+ const n = Number(raw);
296
+ return Number.isFinite(n) && n > 0 ? n : fallback;
297
+ }
298
+
299
+ // Per-file measured durations, regenerated by scripts/gen-test-timings.cjs from
300
+ // gsd-test reporter event streams. Overridable so tests can inject a synthetic
301
+ // table instead of depending on the real suite's cost profile.
302
+ const DEFAULT_TIMINGS_PATH = join(__dirname, '..', 'tests', 'test-timings.json');
303
+ // Must track SCHEMA_VERSION in scripts/gen-test-timings.cjs.
304
+ const SUPPORTED_TIMINGS_SCHEMA = 1;
305
+
306
+ // Load the timing table and reduce it to what the packer needs.
307
+ //
308
+ // Weights are normalized by the table's MEAN duration, so an average-cost file
309
+ // weighs exactly 1 and `MAX_FILES_PER_CHUNK` keeps its original meaning ("about
310
+ // N average files per chunk"). When every file costs the same, total weight
311
+ // equals file count, so the chunk COUNT matches count-based packing exactly.
312
+ // The chunk COMPOSITION still differs — LPT balances where first-fit filled
313
+ // greedily, so 7 uniform files at budget 3 pack {3,2,2} rather than {3,3,1}.
314
+ //
315
+ // `medianWeight` is the fallback for a file absent from the table (a new test,
316
+ // or a table that has drifted). The median — not the mean — because the cost
317
+ // distribution is heavily right-skewed (median 0.28s vs mean 4.6s across the
318
+ // suite), so the median is the honest estimate for an unknown file.
319
+ //
320
+ // Returns null when the table is missing or unusable; the caller then treats
321
+ // every file as weight 1, which reproduces the pre-#2456 count-based balance.
322
+ function loadTestTimings(timingsPath) {
323
+ let parsed;
324
+ try {
325
+ parsed = JSON.parse(readFileSync(timingsPath, 'utf8'));
326
+ } catch {
327
+ return null;
328
+ }
329
+ if (!parsed || typeof parsed !== 'object') return null;
330
+ // Refuse a table written by a future generator: a v2 schema could change the
331
+ // unit or the key format, and consuming it under v1 semantics would silently
332
+ // mis-weight every file. Returning null falls back to uniform weight, which
333
+ // is the same graceful degradation as a missing table.
334
+ if (parsed.schema_version !== undefined && parsed.schema_version !== SUPPORTED_TIMINGS_SCHEMA) {
335
+ return null;
336
+ }
337
+ const timings = parsed.timings;
338
+ // Array.isArray guard: `typeof [] === 'object'`, so a hand-edit that turned
339
+ // the map into a list would pass a bare typeof check and be accepted as a
340
+ // valid table. It degrades harmlessly (no basename ever matches an array
341
+ // index, so every file takes medianWeight), but silently accepting a
342
+ // malformed table is worse than rejecting it — reject, and fall back to
343
+ // uniform weight the same way a missing file does.
344
+ if (!timings || typeof timings !== 'object' || Array.isArray(timings)) return null;
345
+ const values = Object.values(timings).filter(
346
+ (v) => typeof v === 'number' && Number.isFinite(v) && v >= 0,
347
+ );
348
+ if (values.length === 0) return null;
349
+ const mean = values.reduce((sum, v) => sum + v, 0) / values.length;
350
+ if (!(mean > 0)) return null;
351
+ const sorted = [...values].sort((a, b) => a - b);
352
+ const mid = sorted.length >> 1;
353
+ const median = sorted.length % 2 === 1 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2;
354
+ return { timings, mean, medianWeight: median / mean };
355
+ }
356
+
357
+ // Build the packer's weight function from a loaded timing table.
358
+ //
359
+ // A file present in the table weighs its measured duration relative to the
360
+ // table mean. A file ABSENT from it weighs the table's median — this is the
361
+ // "advisory, not gated" contract: a new test or a drifted table costs chunk
362
+ // balance, never a red build. A null table (missing or unparseable file) makes
363
+ // every file weigh 1, reproducing the pre-#2456 count-based balance exactly.
364
+ function makeFileWeigher(timings) {
365
+ if (!timings) return () => 1;
366
+ return (f) => {
367
+ const key = basename(f);
368
+ // Own-property check before the lookup. This is defense-in-depth, NOT a
369
+ // behavior change: the table is JSON-parsed, so a bare `timings[key]` would
370
+ // walk the prototype chain, but the only keys that resolve there are
371
+ // Object.prototype members (`constructor`, `toString`, …) and every real
372
+ // selection is a `*.test.cjs` basename, which can never equal one. Even if
373
+ // it could, the `typeof ms === 'number'` guard below already rejects the
374
+ // function it would return. `Object.hasOwn` makes the intent explicit and
375
+ // keeps the lookup correct for arbitrary input, since this function is
376
+ // exported and does not control its caller's strings.
377
+ const ms = Object.hasOwn(timings.timings, key) ? timings.timings[key] : undefined;
378
+ return typeof ms === 'number' && Number.isFinite(ms) && ms >= 0
379
+ ? ms / timings.mean
380
+ : timings.medianWeight;
381
+ };
382
+ }
383
+
384
+ // Pack `files` into chunks using LPT (longest-processing-time-first): sort by
385
+ // weight descending, then place each file into the currently-LIGHTEST chunk.
386
+ //
387
+ // #2456: the previous packer was a sequential first-fit that appended files in
388
+ // selection order and closed a chunk once its weight budget was hit. Because
389
+ // sorted-adjacent files land together, the two heaviest files in a shard packed
390
+ // into the SAME chunk, leaving the slowest chunk ~3.9x the lightest and sitting
391
+ // near the 600s per-chunk timeout while other chunks idled. LPT is the standard
392
+ // greedy approximation for exactly this makespan problem and balanced the same
393
+ // real shard to ~1.0x.
394
+ //
395
+ // Chunk COUNT is fixed before placement so LPT has bins to balance across:
396
+ // ceil(totalWeight / maxWeight) — the weighted budget, and
397
+ // ceil(fileCount / maxWeight) — a floor that pins the count at what the
398
+ // old count-based packing would produce.
399
+ // The floor is what makes a stale or missing timings table safe: unknown files
400
+ // fall back to a small median weight, which on its own would collapse many files
401
+ // into few fat chunks. With the floor, a degraded table can only ever reproduce
402
+ // today's chunking, never something coarser.
403
+ //
404
+ // `maxChars` still bounds each chunk's argv (Windows CreateProcess caps
405
+ // lpCommandLine at 32,767). A chunk that cannot fit the next file is skipped for
406
+ // that file; when no chunk has room, the chunk count grows and packing restarts.
407
+ // A single file longer than the budget lands alone rather than looping forever.
408
+ //
409
+ // Ordering is fully deterministic — ties break on the separator-normalized file
410
+ // path, and each chunk's files are emitted in their original selection order —
411
+ // so the packing is byte-identical across Windows/macOS/Linux.
412
+ function packChunks(files, { weightOf, maxWeight, maxChars, fixedOverhead }) {
413
+ if (files.length === 0) return [];
414
+ // packChunks is exported, so it cannot assume its caller normalized these.
415
+ // A non-finite or non-positive budget makes the chunk-count arithmetic
416
+ // non-finite, which spins the retry loop below forever or throws from
417
+ // Array.from; a non-finite weight propagates into the same computation.
418
+ // Degrade to a safe bound instead.
419
+ const weightBudget = Number.isFinite(maxWeight) && maxWeight > 0 ? maxWeight : files.length;
420
+ const charBudget = Number.isFinite(maxChars) && maxChars > 0 ? maxChars : Number.MAX_SAFE_INTEGER;
421
+ const overhead = Number.isFinite(fixedOverhead) && fixedOverhead >= 0 ? fixedOverhead : 0;
422
+ const safeWeight = (file) => {
423
+ const w = weightOf(file);
424
+ return Number.isFinite(w) && w >= 0 ? w : 0;
425
+ };
426
+ const entries = files.map((file, index) => ({
427
+ file,
428
+ index,
429
+ weight: safeWeight(file),
430
+ chars: file.length + 1, // +1 for the inter-arg separator
431
+ }));
432
+ const totalWeight = entries.reduce((sum, e) => sum + e.weight, 0);
433
+ // Ties break on a SEPARATOR-NORMALIZED path so a subdir file orders the same
434
+ // on Windows as on POSIX: '/' is 0x2F and '\' is 0x5C, which straddle the
435
+ // uppercase range, so comparing raw paths can order `sub/x.test.cjs` against
436
+ // `subZ.test.cjs` differently per platform and silently produce a different
437
+ // (still valid, but non-reproducible) packing.
438
+ const sortKey = (f) => f.replace(/\\/g, '/');
439
+ const heaviestFirst = [...entries].sort((a, b) => {
440
+ if (b.weight !== a.weight) return b.weight - a.weight;
441
+ const ka = sortKey(a.file);
442
+ const kb = sortKey(b.file);
443
+ return ka < kb ? -1 : ka > kb ? 1 : 0;
444
+ });
445
+
446
+ // Termination: the empty-bin rule below guarantees every file is placeable
447
+ // once chunkCount reaches files.length, so the retry loop cannot run forever.
448
+ // The upper clamp matters as much as the lower bound: a legitimate but tiny
449
+ // budget (RUN_TESTS_MAX_FILES_PER_CHUNK=1e-9) would otherwise ask for
450
+ // 637,000,000,000 bins and throw `RangeError: Invalid array length`. More
451
+ // chunks than files is never useful — one file per chunk is the finest
452
+ // possible packing.
453
+ let chunkCount = Math.min(
454
+ files.length,
455
+ Math.max(1, Math.ceil(totalWeight / weightBudget), Math.ceil(files.length / weightBudget)),
456
+ );
457
+ for (;;) {
458
+ const bins = Array.from({ length: chunkCount }, () => ({
459
+ entries: [],
460
+ weight: 0,
461
+ chars: overhead,
462
+ }));
463
+ let overflowed = false;
464
+ for (const entry of heaviestFirst) {
465
+ let target = null;
466
+ for (const bin of bins) {
467
+ // An empty bin always accepts, so an over-long single file lands alone
468
+ // instead of growing the chunk count forever.
469
+ if (bin.entries.length > 0 && bin.chars + entry.chars > charBudget) continue;
470
+ if (target === null || bin.weight < target.weight) target = bin;
471
+ }
472
+ if (target === null) {
473
+ overflowed = true;
474
+ break;
475
+ }
476
+ target.entries.push(entry);
477
+ target.weight += entry.weight;
478
+ target.chars += entry.chars;
479
+ }
480
+ if (!overflowed) {
481
+ return bins
482
+ .filter((bin) => bin.entries.length > 0)
483
+ .map((bin) => bin.entries.sort((a, b) => a.index - b.index).map((e) => e.file));
484
+ }
485
+ chunkCount++;
486
+ }
221
487
  }
222
488
 
223
489
  function parseArgs(argv) {
@@ -448,7 +714,7 @@ function main() {
448
714
  }
449
715
 
450
716
  // Shard partitioning (#1212): when --shard i/n is given, keep only this
451
- // shard's deterministic round-robin slice of the selected list. Applied
717
+ // shard's deterministic cost-balanced slice of the selected list. Applied
452
718
  // AFTER suite/explicit selection so it composes with --suite (each shard
453
719
  // runs i/n of the post-filter selection).
454
720
  //
@@ -463,11 +729,35 @@ function main() {
463
729
  // from a non-empty list" (total > file count — a valid no-op) from "the
464
730
  // selection was already empty before sharding" (a genuinely empty suite,
465
731
  // which must still hit the discovery hard-error below — Codex #1212 review).
732
+ // Loaded before sharding because BOTH layers weigh by it now (#2472): the
733
+ // shard partition below and the chunk packer further down share this one cost
734
+ // model. Advisory in both places — a missing table yields uniform weight 1,
735
+ // which makes the shard partition degenerate to the legacy equal-count split.
736
+ // Lazily memoized: BOTH layers weigh by it now (#2472) — the shard partition
737
+ // just below and the chunk packer further down share this one cost model —
738
+ // but neither should charge a readFileSync + JSON.parse to an invocation that
739
+ // exits before it needs one (an empty selection, or `--files` with nothing
740
+ // matched). Memoized so the two consumers still read the table at most once.
741
+ // Advisory in both places: a missing table yields uniform weight 1, under
742
+ // which the shard partition degenerates to the legacy equal-count split.
743
+ let weigherMemo = null;
744
+ const fileWeightOf = () => {
745
+ if (weigherMemo === null) {
746
+ const timingsPath = process.env.RUN_TESTS_TIMINGS_FILE || DEFAULT_TIMINGS_PATH;
747
+ weigherMemo = makeFileWeigher(loadTestTimings(timingsPath));
748
+ }
749
+ return weigherMemo;
750
+ };
751
+
466
752
  const usingShard = parsed.shard !== null;
467
753
  let emptyBeforeShard = false;
754
+ // The full pre-partition input, kept for the cross-job fingerprint below.
755
+ // It must be the list every shard job sees, not this job's slice.
756
+ let shardInput = null;
468
757
  if (usingShard) {
469
758
  emptyBeforeShard = selectedNames.length === 0;
470
- selectedNames = selectShard([...selectedNames].sort(), parsed.shard);
759
+ shardInput = [...selectedNames].sort();
760
+ selectedNames = selectShard(shardInput, parsed.shard, fileWeightOf());
471
761
  }
472
762
 
473
763
  const selected = selectedNames.map(f => join(testDir, f));
@@ -536,6 +826,51 @@ function main() {
536
826
  .join(' ')}`,
537
827
  );
538
828
 
829
+ // Shard diagnostics (#2472). File COUNT stopped being a balance signal the
830
+ // moment the partition started weighing by cost — two shards can now hold
831
+ // very different counts by design — so the count line above can no longer be
832
+ // eyeballed to spot a bad split. Worse, each shard job computes its partition
833
+ // independently on its own runner: if the inputs differ between jobs (the
834
+ // file list, or this table), two jobs can place the same file in different
835
+ // shards, or in none, and every job still looks internally consistent. That
836
+ // failure is silent — a test simply never runs and CI stays green.
837
+ //
838
+ // `sig` is the defense: a cheap fingerprint of the exact inputs the partition
839
+ // consumed. Every shard job of a given run must print the SAME sig; a
840
+ // mismatch across jobs is proof the runners disagreed about the input and
841
+ // therefore about the partition. `weighed` reports how many of this shard's
842
+ // files matched a real measurement — a table that silently failed to parse
843
+ // shows weighed=0 instead of being indistinguishable from a healthy load.
844
+ if (usingShard) {
845
+ const weigher = fileWeightOf();
846
+ const table = loadTestTimings(process.env.RUN_TESTS_TIMINGS_FILE || DEFAULT_TIMINGS_PATH);
847
+ const mine = selectedNames.map(f => f.split(/[\\/]/).pop());
848
+ const weighed = table
849
+ ? mine.filter(n => Object.hasOwn(table.timings, n)).length
850
+ : 0;
851
+ const myWeight = mine.reduce((sum, n) => sum + weigher(n), 0);
852
+ // Fingerprint the FULL pre-partition input — the file list and the weight
853
+ // each file was assigned — NOT this shard's slice. Every shard job of one
854
+ // run must print an identical sig; a mismatch is proof the runners
855
+ // disagreed about the input, which is the only way the union of shards can
856
+ // silently drop or duplicate a file. Order-independent sum of per-file
857
+ // (name, weight) hashes: stable across platforms, cheap for ~600 files.
858
+ let sig = 0;
859
+ for (const n of shardInput.map(f => f.split(/[\\/]/).pop())) {
860
+ let h = 2166136261;
861
+ for (let i = 0; i < n.length; i += 1) {
862
+ h = Math.imul(h ^ n.charCodeAt(i), 16777619);
863
+ }
864
+ sig = (sig + (h >>> 0) + Math.round(weigher(n) * 1000)) % 0xffffffff;
865
+ }
866
+ console.error(
867
+ `run-tests: shard=${parsed.shard.index}/${parsed.shard.total} `
868
+ + `files=${mine.length}/${shardInput.length} weighed=${weighed} `
869
+ + `weight=${myWeight.toFixed(2)} table=${table ? 'loaded' : 'absent'} `
870
+ + `sig=${sig.toString(16)}`,
871
+ );
872
+ }
873
+
539
874
  // Default concurrency: 4 on Linux/macOS, 2 on Windows.
540
875
  //
541
876
  // Windows has significantly higher per-subprocess overhead than Linux/macOS:
@@ -559,9 +894,10 @@ function main() {
559
894
  // fine there. Split into chunks sized for the tightest target so behavior
560
895
  // is identical across platforms. (#3597)
561
896
  // Operator override (also used by tests to force chunking with short paths).
562
- const MAX_CMDLINE_CHARS = process.env.RUN_TESTS_MAX_CMDLINE_CHARS
563
- ? Number(process.env.RUN_TESTS_MAX_CMDLINE_CHARS)
564
- : 28000; // headroom below the 32,767 Windows ceiling
897
+ const MAX_CMDLINE_CHARS = positiveNumberEnv(
898
+ process.env.RUN_TESTS_MAX_CMDLINE_CHARS,
899
+ 28000, // headroom below the 32,767 Windows ceiling
900
+ );
565
901
  // A full-lane shard (~171 files) fit in ONE chunk at the old cap of 180, so the
566
902
  // entire shard's wall-clock ran against a single per-chunk timeout. On the slow
567
903
  // Windows runner the install-heavy files in a shard (e.g. install-minimal-hooks
@@ -574,26 +910,31 @@ function main() {
574
910
  // node process (also relieving per-process memory pressure from 170+ files at once).
575
911
  // Lowered from 90 to 60 after #1575 — macOS Node 22 shard 2/3 chunk 2 (~80 files
576
912
  // including state.test.cjs, perf-*, worktree-cleanup) exceeded 600s with 90.
577
- const MAX_FILES_PER_CHUNK = process.env.RUN_TESTS_MAX_FILES_PER_CHUNK
578
- ? Number(process.env.RUN_TESTS_MAX_FILES_PER_CHUNK)
579
- : 60;
580
- // #2088: file COUNT is a poor proxy for a chunk's wall-clock — install-heavy
581
- // files (real installs; install-minimal-hooks.test.cjs alone runs ~250 cases
582
- // doing dozens of installs) are ~10× a unit file. When several land in the SAME
583
- // chunk — e.g. a PR touching the whole install surface, whose *targeted* lane is
584
- // unsharded (13 install-heavy files → one chunk) — that chunk blows the 600s
585
- // backstop while unit-only chunks finish in seconds. WEIGHT install-heavy files
586
- // so they fill a chunk's budget faster and therefore SPREAD across chunks
587
- // instead of clustering. Light files keep weight 1, so pure-unit chunking (and
588
- // its harness tests) is byte-for-byte unchanged. `MAX_FILES_PER_CHUNK` is now a
589
- // per-chunk WEIGHT budget (backwards-compatible: it equals the file count when
590
- // every file is light). Tune via RUN_TESTS_HEAVY_FILE_WEIGHT; classify via the
591
- // basename prefix (install*/installer*/codex-* are the real install-heavy suites).
592
- const HEAVY_TEST_RE = /^(?:install|codex-)/;
593
- const HEAVY_FILE_WEIGHT = process.env.RUN_TESTS_HEAVY_FILE_WEIGHT
594
- ? Number(process.env.RUN_TESTS_HEAVY_FILE_WEIGHT)
595
- : 12;
596
- const fileWeight = (f) => (HEAVY_TEST_RE.test(basename(f)) ? HEAVY_FILE_WEIGHT : 1);
913
+ const MAX_FILES_PER_CHUNK = positiveNumberEnv(process.env.RUN_TESTS_MAX_FILES_PER_CHUNK, 60);
914
+ // #2088 established that file COUNT is a poor proxy for a chunk's wall-clock:
915
+ // install-heavy files (real installs) cost ~10x a unit file, and when several
916
+ // land in the SAME chunk it blows the 600s backstop while unit-only chunks
917
+ // finish in seconds. #2088 approximated cost from the filename — basename
918
+ // matching /^(?:install|codex-)/ scored 12, everything else 1.
919
+ //
920
+ // #2456: that approximation is miscalibrated in BOTH directions, so chunks were
921
+ // still balanced by file count rather than by cost. Measured durations show
922
+ // installer-migration-authoring.test.cjs scoring 12 while running ~0.1s, and the
923
+ // two heaviest files in the whole suite scoring 1 — run-tests-harness.test.cjs
924
+ // (never matched the prefix) and release-tarball-smoke.install.test.cjs (the
925
+ // regex is anchored to the START of the basename, so a mid-name "install" never
926
+ // matches). Both landed in the same chunk, leaving the slowest chunk ~3.9x the
927
+ // lightest and sitting near the timeout.
928
+ //
929
+ // Weight each file by its MEASURED duration instead. `MAX_FILES_PER_CHUNK`
930
+ // remains the per-chunk weight budget and keeps its scale — weights are
931
+ // normalized so an average-cost file weighs 1 — so an all-uniform suite chunks
932
+ // exactly as it did before. Timings are ADVISORY, never gated: an unknown file
933
+ // falls back to the table's median weight and a missing table falls back to
934
+ // uniform weight 1, so staleness degrades chunk BALANCE gracefully instead of
935
+ // failing CI. Regenerate via `node scripts/gen-test-timings.cjs <events.jsonl>`.
936
+ // The cost table is loaded lazily above and memoized; both the shard
937
+ // partition and this packer consume the same weigher (#2472).
597
938
 
598
939
  // node:test does not exit until the event loop drains. A unit test that leaks
599
940
  // an open handle (un-terminated Worker, un-killed child_process, ref'd timer)
@@ -609,35 +950,19 @@ function main() {
609
950
  const forceExit = nodeMajor >= 22 && !process.env.RUN_TESTS_NO_FORCE_EXIT;
610
951
 
611
952
  const FIXED_OVERHEAD = process.execPath.length + '--test'.length + concurrency.length + (forceExit ? '--test-force-exit'.length + 1 : 0) + 8;
612
- const chunks = [];
613
- let current = [];
614
- let currentLen = FIXED_OVERHEAD;
615
- let currentWeight = 0;
616
- for (const file of selected) {
617
- const add = file.length + 1; // +1 for the inter-arg separator
618
- if (
619
- current.length > 0 &&
620
- (currentLen + add > MAX_CMDLINE_CHARS || currentWeight >= MAX_FILES_PER_CHUNK)
621
- ) {
622
- chunks.push(current);
623
- current = [];
624
- currentLen = FIXED_OVERHEAD;
625
- currentWeight = 0;
626
- }
627
- current.push(file);
628
- currentLen += add;
629
- currentWeight += fileWeight(file); // heavy install files count for more
630
- }
631
- if (current.length > 0) chunks.push(current);
953
+ const chunks = packChunks(selected, {
954
+ weightOf: fileWeightOf(),
955
+ maxWeight: MAX_FILES_PER_CHUNK,
956
+ maxChars: MAX_CMDLINE_CHARS,
957
+ fixedOverhead: FIXED_OVERHEAD,
958
+ });
632
959
 
633
960
  // A chunk that still hangs (a leak the backstop somehow misses, or a wedged
634
961
  // subprocess) must fail loudly rather than silently burn the job's wall-clock
635
962
  // budget until the CI runner cancels the whole job. Default 10 min per chunk:
636
963
  // well above a healthy chunk (~4-5 min on the windows lane) but below the 20m
637
964
  // job cap. Operator/test override via RUN_TESTS_CHUNK_TIMEOUT_MS.
638
- const chunkTimeoutMs = process.env.RUN_TESTS_CHUNK_TIMEOUT_MS
639
- ? Number(process.env.RUN_TESTS_CHUNK_TIMEOUT_MS)
640
- : 600000;
965
+ const chunkTimeoutMs = positiveNumberEnv(process.env.RUN_TESTS_CHUNK_TIMEOUT_MS, 600000);
641
966
 
642
967
  let firstFailureExit = 0;
643
968
  for (let i = 0; i < chunks.length; i++) {
@@ -673,9 +998,35 @@ function main() {
673
998
  );
674
999
  }
675
1000
  const code = err.status || 1;
676
- // Run every chunk so the operator sees all failures in one pass; report
677
- // the first non-zero exit at the end.
678
1001
  if (firstFailureExit === 0) firstFailureExit = code;
1002
+ if (timedOut) {
1003
+ // A timeout has already burned a large share of the job's budget
1004
+ // (chunkTimeoutMs defaults to 600000ms, i.e. half the 20m CI job
1005
+ // cap), so — unlike an ordinary test failure — letting the loop
1006
+ // fall through to the remaining chunks risks the CI runner
1007
+ // cancelling the whole job before they finish. That cancellation
1008
+ // replaces the loud, specific diagnostic printed above with an
1009
+ // opaque "The operation was canceled." buried at the very end of
1010
+ // the log, thousands of lines past the real cause (observed live on
1011
+ // CI run 29749380190: chunk 1/5 timed out, the loop pressed on
1012
+ // through chunks 2-4, and the job was cancelled mid-chunk-5 — the
1013
+ // timeout message was ~38,000 log lines from the end and
1014
+ // `gh run view --log-failed` returned nothing). Abort the remaining
1015
+ // chunks instead so the operator actually sees this message.
1016
+ const skipped = chunks.length - (i + 1);
1017
+ if (skipped > 0) {
1018
+ console.error(
1019
+ `run-tests: aborting — skipping the remaining ${skipped} chunk${skipped === 1 ? '' : 's'} ` +
1020
+ `after the chunk ${i + 1}/${chunks.length} timeout rather than risk the CI runner ` +
1021
+ `cancelling the job (and burying this diagnostic) before they finish.`,
1022
+ );
1023
+ }
1024
+ break;
1025
+ }
1026
+ // A non-timeout failure is cheap in wall-clock terms (the child exits
1027
+ // promptly on its own), so — unlike the timeout case above — run every
1028
+ // remaining chunk anyway: the operator sees all failures in one pass,
1029
+ // and the first non-zero exit is reported at the end.
679
1030
  }
680
1031
  }
681
1032
  if (firstFailureExit !== 0) return firstFailureExit;
@@ -685,4 +1036,15 @@ if (require.main === module) {
685
1036
  runMain(main);
686
1037
  }
687
1038
 
688
- module.exports = { suiteOf, ensureBuiltArtifacts, ensureBuiltHooks, parseShardArg, selectShard };
1039
+ module.exports = {
1040
+ suiteOf,
1041
+ ensureBuiltArtifacts,
1042
+ ensureBuiltHooks,
1043
+ parseShardArg,
1044
+ selectShard,
1045
+ positiveNumberEnv,
1046
+ loadTestTimings,
1047
+ makeFileWeigher,
1048
+ packChunks,
1049
+ DEFAULT_TIMINGS_PATH,
1050
+ };
@@ -28,7 +28,7 @@ Flow: Select Framework → Research Docs → Research Domain → Design Eval Str
28
28
  </execution_context>
29
29
 
30
30
  <context>
31
- Phase number: $ARGUMENTS — optional, auto-detects next unplanned phase if omitted.
31
+ Phase number: $ARGUMENTS — optional; when omitted, the orchestrating workflow reads ROADMAP.md and selects the next unplanned phase. This is not a `gsd-tools.cjs` CLI feature — the CLI's phase-lookup primitives require an explicit phase number.
32
32
  </context>
33
33
 
34
34
  <process>
@@ -64,12 +64,16 @@ On any error or timeout, stop and let the phase continue -- capture is best-effo
64
64
  # One-time: declare the GSD room taxonomy so detect_room() recognizes these folders
65
65
  mkdir -p "$STAGE"
66
66
  [ -f "$STAGE/mempalace.yaml" ] || cat > "$STAGE/mempalace.yaml" <<'YAML'
67
+ # Each entry MUST be a dict with a `name` key (the miner's detect_room()
68
+ # indexes room["name"] — a bare-string list crashes _mine_impl with
69
+ # TypeError: string indices must be integers, not 'str'). Optional fields:
70
+ # `description`, `keywords` (matched against folder-path segments).
67
71
  rooms:
68
- - decisions
69
- - planning
70
- - milestones
71
- - problems
72
- - general
72
+ - name: decisions
73
+ - name: planning
74
+ - name: milestones
75
+ - name: problems
76
+ - name: general
73
77
  YAML
74
78
  # Suppress MemPalace cache artifacts written into the scanned tree
75
79
  [ -f "$STAGE/.gitignore" ] || echo "mempalace_embedder.json" > "$STAGE/.gitignore"
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-new-milestone
3
3
  description: "Start a new milestone cycle — update PROJECT.md and route to requirements"
4
- argument-hint: "[milestone name, e.g., 'v1.1 Notifications']"
4
+ argument-hint: "[milestone name, e.g., 'v1.1 Notifications'] [--ws <name>]"
5
5
  allowed-tools:
6
6
  - Read
7
7
  - Write
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: gsd-plan-phase
3
3
  description: "Create detailed phase plan (PLAN.md) with verification loop"
4
- argument-hint: "[phase] [--auto] [--research] [--skip-research] [--research-phase <N>] [--view] [--gaps] [--skip-verify] [--prd <file>] [--ingest <path-or-glob>] [--ingest-format <auto|nygard|madr|narrative>] [--reviews] [--text] [--tdd] [--mvp]"
4
+ argument-hint: "[phase] [--auto] [--research] [--skip-research] [--research-phase <N>] [--view] [--gaps] [--skip-verify] [--prd <file>] [--ingest <path-or-glob>] [--ingest-format <auto|nygard|madr|narrative>] [--reviews] [--text] [--tdd] [--mvp] [--no-tracer] [--no-reversibility-gates]"
5
5
  effort: max
6
6
  allowed-tools:
7
7
  - Read
@@ -40,7 +40,7 @@ Create executable phase prompts (PLAN.md files) for a roadmap phase with integra
40
40
  </runtime_note>
41
41
 
42
42
  <context>
43
- Phase number: $ARGUMENTS (optional — auto-detects next unplanned phase if omitted)
43
+ Phase number: $ARGUMENTS (optional — when omitted, the orchestrating workflow reads ROADMAP.md and selects the next unplanned phase; `gsd-tools.cjs` itself has no auto-detect feature and requires an explicit phase number)
44
44
 
45
45
  **Flags:**
46
46
  - `--research` — Force re-research even if RESEARCH.md exists
@@ -52,7 +52,9 @@ Phase number: $ARGUMENTS (optional — auto-detects next unplanned phase if omit
52
52
  - `--ingest-format <auto|nygard|madr|narrative>` — Optional ADR parser format override (`auto` default).
53
53
  - `--reviews` — Replan incorporating cross-AI review feedback from REVIEWS.md (produced by `/gsd-review`)
54
54
  - `--text` — Use plain-text numbered lists instead of TUI menus (required for `/rc` remote sessions)
55
- - `--mvp` — Vertical MVP mode. Planner organizes tasks as feature slices (UI→API→DB) instead of horizontal layers. On Phase 1 of a new project, also emits `SKELETON.md` (Walking Skeleton). Can be persisted on a phase via `**Mode:** mvp` in ROADMAP.md.
55
+ - `--mvp` — MVP enrichment on top of the default tracer-first ordering: frames the phase goal as a user story and, on Phase 1 of a new project, also emits `SKELETON.md` (Walking Skeleton). Vertical slicing itself is now the default (see `--no-tracer`); `--mvp` no longer *turns it on*. Can be persisted on a phase via `**Mode:** mvp` in ROADMAP.md.
56
+ - `--no-tracer` — Opt out of the default **tracer-first** decomposition and plan horizontal layers (the legacy default). By default every plan LEADS with one production-quality end-to-end `tracer` slice that is verified before any expansion task.
57
+ - `--no-reversibility-gates` — Suppress the human checkpoint that a **one-way-door** decision normally earns, for runs you intend to leave unattended. By default a decision rated `one-way` (undo needs a migration, breaks a published contract, or is impossible) gets a `checkpoint:decision` before the task implementing it. Ratings are still recorded on tasks and `costly` items still flagged — the flag changes what stops the run, not what the plan remembers.
56
58
 
57
59
  Normalize phase input in step 2 before any directory lookups.
58
60
  </context>