@tekyzinc/gsd-t 5.20.15 → 5.22.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,62 @@
2
2
 
3
3
  All notable changes to GSD-T are documented here. Updated with each release.
4
4
 
5
+ ## [5.22.10] - 2026-09-27
6
+
7
+ ### Added — AI-assisted T-shirt sizing for `/gsd-t-estimate`
8
+
9
+ Estimates sized every task as if a person typed the code. Measured against the Hilo Delivery
10
+ Runway build (45 tasks, ~26 solo hours: 13 David + 13 Claude), the old scale priced the same
11
+ work at 1,262–1,577 hours. A size is now the team hours a task takes with AI-assisted
12
+ development: solo AI minutes × project multiplier (greenfield solo ×1, team ×5; yellow-field
13
+ solo ×2, team ×8 isolated / ×12 big blast radius) + task switching added after the multiplier
14
+ (7.5 / 15 / 30 min), mapped to the nearest size on a smaller scale — XS 0.1 · S 0.25 · M 0.5 ·
15
+ L 1 · XL 2 · XXL 4 person-days. The multipliers live in the estimator; the sheet layout is
16
+ unchanged.
17
+
18
+ - `bin/gsd-t-estimate-sheet.cjs`: new `size` verb (`--solo-min`, `--project`, `--switch-min`);
19
+ `write` puts the AI scale in the legend (matched by label); sizes mode halts if a sized row
20
+ outside the plan would be silently re-priced; `plan-check` previews on the new scale
21
+ - `templates/estimate-sheet-spec.md`: §1.4 sizing model; legend + verb list updated
22
+ - `commands/gsd-t-estimate.md`, `templates/playbooks/tekyz-estimation-and-prd-playbook.md`,
23
+ `templates/estimate-config.json`: size in solo AI minutes; new-team familiarization off by default
24
+ - `test/estimate-sheet-writer.test.js`: +6 tests
25
+
26
+ Existing estimate sheets keep the old scale until they are re-estimated.
27
+
28
+ ### Fixed — Red Team runs on Opus 5.5 under every profile
29
+
30
+ The `standard` profile put Red Team on Sonnet, contradicting both profile contracts and the
31
+ READMEs. It now runs on `opus` (Opus 5.5) in standard, pro and premium.
32
+
33
+ - `bin/gsd-t-model-tier-policy.cjs`, `commands/gsd-t-status.md`,
34
+ `.gsd-t/contracts/model-profile-config-contract.md`, `test/m86-policy-profiles.test.js`
35
+
36
+ ## [5.21.10] - 2026-09-22
37
+
38
+ ### Changed — the top tier now points at Opus 5.5
39
+
40
+ The `opus` tier alias resolved to `claude-opus-5`. It now resolves to
41
+ `claude-opus-5-5`, so every high-stakes stage (solution-space probe, partition
42
+ probe, competition producers and judge, pre-mortem, Red Team, both debug cycles)
43
+ runs Opus 5.5. Only the concrete model id changed: the three-tier shape, the
44
+ stage-to-tier map, and the relaxed fresh-context judge-blindness invariant are
45
+ untouched.
46
+
47
+ - `bin/gsd-t-model-tier-policy.cjs`: `MODEL_IDS.opus` → `claude-opus-5-5`; the
48
+ thinking-omission predicate still matches no current tier model
49
+ - `.gsd-t/contracts/model-tier-policy-contract.md`: → v2.1.0, alias table and
50
+ Updated line; the v2.0.0 note kept under `## Previously`
51
+ - `templates/workflows/gsd-t-{phase,verify,debug}.workflow.js`,
52
+ `templates/prompts/blind-adversary-subagent.md`: model id in the stage comments
53
+ - `templates/CLAUDE-global.md`, `README.md`, `commands/gsd-t-{help,status}.md`:
54
+ the documented `opus` = model id
55
+ - `test/m85-model-tier-policy.test.js`, `test/m86-policy-profiles.test.js`,
56
+ `test/m90-tier-policy-lint.test.js`: expected id; the drifted-literal negative
57
+ test still fails on a mismatch
58
+
59
+ No migration. A project picks the new id on its next `gsd-t update-all`.
60
+
5
61
  ## [5.20.15] - 2026-09-21
6
62
 
7
63
  ### Changed — phases must be contiguous; the Team Mix title row is the phase name only
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # GSD-T: Contract-Driven Development for Claude Code
2
2
 
3
- **v5.20.15** - A methodology for reliable, parallelizable development using Claude Code with optional Agent Teams support.
3
+ **v5.22.10** - A methodology for reliable, parallelizable development using Claude Code with optional Agent Teams support.
4
4
 
5
5
  **Eliminates context rot** — task-level fresh dispatch (one subagent per task, ~10-20% context each) means compaction never triggers.
6
6
  **Compaction-proof debug loops** — `gsd-t headless --debug-loop` runs test-fix-retest cycles as separate `claude -p` sessions. A JSONL debug ledger persists all hypothesis/fix/learning history across fresh sessions. Anti-repetition preamble injection prevents retrying failed hypotheses. Escalation tiers (sonnet → opus → human) and a hard iteration ceiling enforced externally.
@@ -17,7 +17,7 @@
17
17
  **Real-Time Agent Dashboard** — `gsd-t-stream-feed-server.js` serves a streaming UI at `127.0.0.1:7842` that renders all workers' stream-json output as a continuous feed with task/wave banners, duration + usage chips, token corner bar, localStorage filters, and replay via `WS /feed?from=N`. Dashboard auto-starts idempotently on each spawn (`scripts/gsd-t-dashboard-autostart.cjs`). Port is project-scoped via `projectScopedDefaultPort(projectDir)` so multi-project workflows do not clobber each other.
18
18
  **Rigorous User-Journey Coverage + Anti-Drift Test Quality** — `bin/journey-coverage.cjs` regex listener detector + `gsd-t check-coverage` CLI + `scripts/hooks/pre-commit-journey-coverage` commit gate blocks viewer-source commits when uncovered listeners exist. Journey specs in `e2e/journeys/` use functional assertions (zero `toBeVisible`-only tests) per the E2E Test Quality Standard in CLAUDE.md.
19
19
  **Universal Playwright Bootstrap + Deterministic UI Enforcement (M50)** — three executable enforcement layers: (1) `bin/playwright-bootstrap.cjs` + `bin/ui-detection.cjs` - idempotent installer detects package manager, installs `@playwright/test` + chromium, scaffolds `e2e/`; (2) Workflow runtime runs `playwright-bootstrap.cjs::installPlaywright()` before any E2E stage when `hasUI && !hasPlaywright`; install failure halts with `blocked-needs-human`; (3) `scripts/hooks/pre-commit-playwright-gate` (opt-in via `gsd-t doctor --install-hooks`) blocks viewer-source commits when staged files are newer than `.gsd-t/.last-playwright-pass`. The `gsd-t setup-playwright [path]` subcommand handles manual install.
20
- **Surgical model selection** — models are assigned haiku/sonnet/opus per phase (**Fable removed 2026-07-24**; `opus` = **claude-opus-5**). **Single-source tier policy:** `bin/gsd-t-model-tier-policy.cjs` is the SINGLE source of truth; every high-stakes stage (solution-space probe, partition probe, competition judge, pre-mortem, Red Team, competition producers, debug both cycles) runs Opus 5. Opus 5 shipped at the same price as Opus 4.8 but >2× its coding score, so Fable's cost premium is no longer justified. The M82 judge-blindness invariant is relaxed to "fresh independent context" — producers and judge both run opus. Drift is mechanically enforced by the M71-family lint (`test/m85-workflow-tier-policy-lint.test.js`). **M86 model profiles:** `bin/gsd-t-model-profile.cjs` adds a per-project SECOND dimension — three named profiles (`standard` / `pro` / `premium`) that dial which stages run on Opus vs. Sonnet.
20
+ **Surgical model selection** — models are assigned haiku/sonnet/opus per phase (**Fable removed 2026-07-24**; `opus` = **claude-opus-5-5**). **Single-source tier policy:** `bin/gsd-t-model-tier-policy.cjs` is the SINGLE source of truth; every high-stakes stage (solution-space probe, partition probe, competition judge, pre-mortem, Red Team, competition producers, debug both cycles) runs Opus 5. Opus 5 shipped at the same price as Opus 4.8 but >2× its coding score, so Fable's cost premium is no longer justified. The M82 judge-blindness invariant is relaxed to "fresh independent context" — producers and judge both run opus. Drift is mechanically enforced by the M71-family lint (`test/m85-workflow-tier-policy-lint.test.js`). **M86 model profiles:** `bin/gsd-t-model-profile.cjs` adds a per-project SECOND dimension — three named profiles (`standard` / `pro` / `premium`) that dial which stages run on Opus vs. Sonnet.
21
21
  **Token Telemetry** — `gsd-t-calibration-hook.js` records token usage per spawn to `.gsd-t/token-metrics.jsonl` (18-field rows). `gsd-t-token-aggregator.js` aggregates across tasks for the `/gsd-t-metrics` view. Use the native Claude Code `/context` command for live in-session context percentage.
22
22
  **Quality North Star** — projects define a `## Quality North Star` section in CLAUDE.md (1–3 sentences, e.g., "This is a published npm library. Every public API must be intuitive and backward-compatible."). `gsd-t-init` auto-detects preset (library/web-app/cli) from package.json signals; `gsd-t-setup` configures it for existing projects. Subagents read it as a quality lens; absent = silent skip (backward compatible).
23
23
  **Design Brief Artifact** — during partition, UI/frontend projects (React, Vue, Svelte, Flutter, Tailwind) automatically get `.gsd-t/contracts/design-brief.md` with color palette, typography, spacing system, component patterns, and tone/voice. Non-UI projects skip silently. User-customized briefs are preserved. Referenced in plan phase for visual consistency.
@@ -184,7 +184,7 @@ This will replace changed command files, back up your CLAUDE.md if customized, a
184
184
  | `/gsd-t-scan` | Deep codebase analysis → techdebt.md | Manual |
185
185
  | `/gsd-t-gap-analysis` | Requirements gap analysis — spec vs. existing code | Manual |
186
186
  | `/gsd-t-promote-debt` | Convert techdebt items to milestones | Manual |
187
- | `/gsd-t-estimate` | Turn any work doc (scan, gap-analysis sheet, requirements, feature/app spec) into a Tekyz client estimate Google Sheet — T-Shirt Size + Team Mix + Technology Stack, written by the deterministic `gsd-t estimate-sheet` tool from a JSON plan and read-back-audited against `templates/estimate-sheet-spec.md` — supervised, with an operator-arbitrated Estimate Red Team. No PRD (that is `/gsd-t-prd`) | Manual |
187
+ | `/gsd-t-estimate` | Turn any work doc (scan, gap-analysis sheet, requirements, feature/app spec) into a Tekyz client estimate Google Sheet — T-Shirt Size + Team Mix + Technology Stack, written by the deterministic `gsd-t estimate-sheet` tool from a JSON plan and read-back-audited against `templates/estimate-sheet-spec.md`; sizes assume AI-assisted development (solo AI minutes × project multiplier) — supervised, with an operator-arbitrated Estimate Red Team. No PRD (that is `/gsd-t-prd`) | Manual |
188
188
  | `/gsd-t-stories` | Generate a dev-team handoff doc in the Tekyz user-stories format (stories + workflows + acceptance criteria + Mermaid flow diagrams + mapped test cases) from any source | Manual |
189
189
  | `/gsd-t-demo-videos` | Narrated screen-recording walkthroughs of a running app — coverage plan, UI-driven seeding, batched TTS with a measured one-voice gate, continuous recording, mux, silence trim | Manual |
190
190
  | `/gsd-t-populate` | Auto-populate docs from existing codebase | Manual |
@@ -47,6 +47,21 @@ const TAB_TECH = "Technology Stack";
47
47
  const TAB_OVERVIEW = "Overview";
48
48
 
49
49
  const SIZE_CODES = ["XS", "S", "M", "L", "XL", "XXL"];
50
+
51
+ // AI-assisted sizing model (David, 2026-09-27; calibrated on the Hilo Delivery Runway build:
52
+ // 45 tasks, ~26 solo hours). A task is sized in SOLO AI-assisted minutes, multiplied by the
53
+ // project type, plus a per-task switching allowance added AFTER the multiplier (switching is
54
+ // one person's pickup time — it does not grow with team size). The result maps to the nearest
55
+ // size on AI_SIZE_DAYS, which `write` puts in the sheet's legend. The multipliers live here,
56
+ // in the estimator — never on the sheet.
57
+ const AI_SIZE_DAYS = { XS: 0.1, S: 0.25, M: 0.5, L: 1, XL: 2, XXL: 4 };
58
+ const PROJECT_MULTIPLIER = {
59
+ "greenfield-solo": 1,
60
+ "greenfield-team": 5,
61
+ "yellowfield-solo": 2,
62
+ "yellowfield-team-isolated": 8,
63
+ "yellowfield-team-wide": 12,
64
+ };
50
65
  const PHASES = ["MVP", "Phase 1", "Phase 2", "Phase 3"];
51
66
 
52
67
  const COLOR = {
@@ -347,10 +362,11 @@ function locateTshirt(grid) {
347
362
  cols.sizeLabels = cols.sizes.map((c) => headerText[c]);
348
363
  // legend: the XS..XXL rows in column A
349
364
  const legend = {};
365
+ const legendRow1 = {};
350
366
  let legendFirst = -1, legendLast = -1;
351
367
  for (let r = 0; r < headerRow; r++) {
352
368
  const m = textAt(grid, r, 0).trim().match(/^(XS|S|M|L|XL|XXL)\s*-/);
353
- if (m) { legend[m[1]] = numAt(grid, r, 1); if (legendFirst < 0) legendFirst = r; legendLast = r; }
369
+ if (m) { legend[m[1]] = numAt(grid, r, 1); legendRow1[m[1]] = r + 1; if (legendFirst < 0) legendFirst = r; legendLast = r; }
354
370
  }
355
371
  for (const code of SIZE_CODES) {
356
372
  if (!(code in legend) || Number.isNaN(legend[code])) throw new Halt(`${TAB_TSHIRT}: size legend missing '${code}' (looked for 'XS - …' labels in column A above the header row)`);
@@ -391,7 +407,7 @@ function locateTshirt(grid) {
391
407
  const lowDol = findCell(grid, "Low ($)", headerRow, 20), highDol = findCell(grid, "High ($)", headerRow, 20);
392
408
  if (!lowHrs || !highHrs || !lowDol || !highDol) throw new Halt(`${TAB_TSHIRT}: rollup headers 'Low ($)' / 'Low Hrs' / 'High ($)' / 'High Hrs' not all found`);
393
409
  cols.rollupPhase = mvp.c; cols.lowHrs = lowHrs.c; cols.highHrs = highHrs.c; cols.lowDol = lowDol.c; cols.highDol = highDol.c;
394
- return { headerRow, firstItemRow: headerRow + 1, legend, legendFirst1: legendFirst + 1, legendLast1: legendLast + 1, mf, mfTotal, mfTotalCell, highFactor, rate, rollupRows, cols };
410
+ return { headerRow, firstItemRow: headerRow + 1, legend, legendRow1, legendFirst1: legendFirst + 1, legendLast1: legendLast + 1, mf, mfTotal, mfTotalCell, highFactor, rate, rollupRows, cols };
395
411
  }
396
412
 
397
413
  /** The standard template's column map (what `write` produces). */
@@ -604,6 +620,33 @@ function buildRoster(plan, totalDays, mfList) {
604
620
 
605
621
  // ───────────────────────── T-Shirt totals (pure) ─────────────────────────
606
622
 
623
+ /** Switching allowance (minutes) for a task of `teamMin` minutes: small 7.5, medium 15, large 30. */
624
+ function defaultSwitchMin(teamMin) { return teamMin < 120 ? 7.5 : teamMin < 480 ? 15 : 30; }
625
+
626
+ /**
627
+ * The estimator's per-task math: solo AI minutes × project multiplier + switching (after the
628
+ * multiplier), then the nearest AI_SIZE_DAYS size (geometric midpoints between neighbours).
629
+ */
630
+ function aiTaskSize({ soloMin, project, switchMin }) {
631
+ const mult = PROJECT_MULTIPLIER[project];
632
+ if (mult == null) throw new Halt(`--project must be one of ${Object.keys(PROJECT_MULTIPLIER).join(" | ")}`, 64);
633
+ if (typeof soloMin !== "number" || !(soloMin > 0)) throw new Halt("--solo-min must be a positive number of minutes", 64);
634
+ const teamMin = soloMin * mult;
635
+ const sw = switchMin == null ? defaultSwitchMin(teamMin) : switchMin;
636
+ if (typeof sw !== "number" || sw < 0) throw new Halt("--switch-min must be a number of minutes ≥ 0", 64);
637
+ const hours = (teamMin + sw) / 60;
638
+ const days = hours / 8;
639
+ let size = SIZE_CODES[SIZE_CODES.length - 1];
640
+ for (let i = 0; i < SIZE_CODES.length - 1; i++) {
641
+ const cut = Math.sqrt(AI_SIZE_DAYS[SIZE_CODES[i]] * AI_SIZE_DAYS[SIZE_CODES[i + 1]]);
642
+ if (days < cut) { size = SIZE_CODES[i]; break; }
643
+ }
644
+ return { soloMin, project, multiplier: mult, switchMin: sw, hours: round2(hours), days: round2(days), size, sizeDays: AI_SIZE_DAYS[size] };
645
+ }
646
+
647
+ /** True when the sheet's legend already carries the AI scale. */
648
+ function legendIsAi(legend) { return SIZE_CODES.every((c) => legend[c] === AI_SIZE_DAYS[c]); }
649
+
607
650
  function sizeDays(code, legend) {
608
651
  const v = sizeOf(code);
609
652
  if (v === "") return 0;
@@ -721,11 +764,15 @@ async function writeTshirt(api, plan, opts) {
721
764
  const layout = locateTshirt(grid);
722
765
  const std = ["phase", "days", "mfactor", "total", "low", "high"].every((k) => layout.cols[k] === STANDARD_COLS[k]) && layout.cols.sizes.length === 2 && layout.legendFirst1 === 4 && layout.headerRow === 12;
723
766
  if (!std) throw new Halt(`${TAB_TSHIRT}: this sheet is not the current template layout (header row ${layout.headerRow + 1}, size columns ${layout.cols.sizeLabels.join("/")}) — 'write' supports the current template only; 'teammix' and 'audit' work on both`);
767
+ // New estimates use the AI-assisted scale: the legend is rewritten to AI_SIZE_DAYS (after the
768
+ // checks below pass, so a halted run changes nothing), and totals are computed on it.
769
+ const legendChange = !legendIsAi(layout.legend);
770
+ layout.legend = { ...AI_SIZE_DAYS };
724
771
  const totals = tshirtTotals(plan, layout);
725
772
  const phaseSrc = findPhaseSource(grid, 0);
726
773
  if (phaseSrc < 0) throw new Halt(`${TAB_TSHIRT}: no Phase dropdown source cell found in column E — add a ONE_OF_LIST validation (MVP / Phase 1 / Phase 2 / Phase 3) to E${layout.firstItemRow + 1} and re-run`);
727
774
 
728
- if (plan.tshirt.mode === "sizes") return writeTshirtSizes(api, grid, sheetId, layout, plan, totals, phaseSrc);
775
+ if (plan.tshirt.mode === "sizes") return writeTshirtSizes(api, grid, sheetId, layout, plan, totals, phaseSrc, legendChange);
729
776
 
730
777
  const first0 = layout.firstItemRow;
731
778
  let occupied = 0;
@@ -735,6 +782,7 @@ async function writeTshirt(api, plan, opts) {
735
782
  }
736
783
  const built = tshirtRows(plan, first0);
737
784
  const lastWritten1 = built.rows[built.rows.length - 1].row1;
785
+ if (legendChange) await writeAiLegend(api, layout);
738
786
 
739
787
  // 1. clear values + formats + merges in the item area (clear-then-paint)
740
788
  const clearTo = Math.max(lastWritten1 + 5, rowCount(grid) + 1);
@@ -792,7 +840,12 @@ async function writeTshirt(api, plan, opts) {
792
840
  }
793
841
 
794
842
  /** sizes mode — rows exist (gap-analysis sheet); fill E:L on rows matched by "(id)" in column C. */
795
- async function writeTshirtSizes(api, grid, sheetId, layout, plan, totals, phaseSrc) {
843
+ /** Put AI_SIZE_DAYS into the legend's value column, row by row by label (never by position). */
844
+ async function writeAiLegend(api, layout) {
845
+ for (const c of SIZE_CODES) await api.putValues(TAB_TSHIRT, `B${layout.legendRow1[c]}`, [[AI_SIZE_DAYS[c]]]);
846
+ }
847
+
848
+ async function writeTshirtSizes(api, grid, sheetId, layout, plan, totals, phaseSrc, legendChange) {
796
849
  const rowsById = new Map();
797
850
  for (let r = layout.firstItemRow; r < rowCount(grid); r++) {
798
851
  const m = textAt(grid, r, 2).match(/\(([A-Za-z]+-\d+(?:\.\d+)*)\)\s*$/);
@@ -809,6 +862,18 @@ async function writeTshirtSizes(api, grid, sheetId, layout, plan, totals, phaseS
809
862
  writes.push({ r0, phase: it.phase, values: [it.phase, sizeOf(it.fe), sizeOf(it.be), f.H, f.I, f.J, f.K, f.L] });
810
863
  }
811
864
  if (missing.length) throw new Halt(`${TAB_TSHIRT}: no row carries these ids in column C — ${missing.join(", ")}. Rows are matched BY NAME (the "(id)" suffix), never by position.`, 4, missing);
865
+ if (legendChange) {
866
+ // Changing the legend re-prices EVERY sized row on the tab. A sized row the plan does not
867
+ // cover would silently move to the new scale — halt instead, naming the rows.
868
+ const planned = new Set(writes.map((w) => w.r0));
869
+ const stray = [];
870
+ for (const [id, r0] of rowsById) {
871
+ if (planned.has(r0)) continue;
872
+ if (layout.cols.sizes.some((c) => textAt(grid, r0, c).trim())) stray.push(id);
873
+ }
874
+ if (stray.length) throw new Halt(`${TAB_TSHIRT}: the sheet is on the old day scale and these sized rows are not in the plan — ${stray.join(", ")}. Moving the legend to the AI-assisted scale would re-price them silently. Size every row in the plan, or clear those rows first.`, 4, stray);
875
+ await writeAiLegend(api, layout);
876
+ }
812
877
  let totalRow0 = -1;
813
878
  for (let r = lastItem0 + 1; r < rowCount(grid); r++) if (/^Total \(Days\)/i.test(textAt(grid, r, 0))) { totalRow0 = r; break; }
814
879
  if (totalRow0 < 0) throw new Halt(`${TAB_TSHIRT}: no 'Total (Days)' row found below the items — add it (spec §1.3) and re-run`);
@@ -1314,6 +1379,7 @@ function rostersTable(rosters) {
1314
1379
 
1315
1380
  async function verbPlanCheck(api, plan) {
1316
1381
  const layout = locateTshirt(await api.grid(TAB_TSHIRT));
1382
+ layout.legend = { ...AI_SIZE_DAYS }; // preview what `write` produces — it puts the AI scale in the legend
1317
1383
  const totals = tshirtTotals(plan, layout);
1318
1384
  const rosters = buildRosters(plan, layout);
1319
1385
  return { totals, mf: layout.mf, mfTotal: round2(layout.mfTotal), highFactor: layout.highFactor, rate: layout.rate, rosters, table: rostersTable(rosters) };
@@ -1609,7 +1675,7 @@ function printChecks(audit) {
1609
1675
  console.log(`${audit.ok ? "AUDIT PASS" : `AUDIT FAIL (${audit.failed})`} — ${audit.title}`);
1610
1676
  }
1611
1677
 
1612
- const USAGE = "usage: gsd-t estimate-sheet <read|plan-check|write|teammix|format|phases|titles|audit|plan-schema> --sheet <id|url> [--tab <name>] [--plan <plan.json>] [--replace] [--fte '{\"backend\":1.5}'] [--title <t>] [--dry-run] [--no-audit] [--key <path>] [--json]";
1678
+ const USAGE = "usage: gsd-t estimate-sheet <read|plan-check|write|teammix|format|phases|titles|audit|plan-schema|size> --sheet <id|url> [--tab <name>] [--plan <plan.json>] [--replace] [--fte '{\"backend\":1.5}'] [--title <t>] [--dry-run] [--no-audit] [--key <path>] [--json]\n gsd-t estimate-sheet size --solo-min <n> --project <greenfield-solo|greenfield-team|yellowfield-solo|yellowfield-team-isolated|yellowfield-team-wide> [--switch-min <n>] [--json]";
1613
1679
 
1614
1680
  /** Runs a verb; returns the exit code. Throws Halt (or any error) — the runner below turns that into exit 4/64. */
1615
1681
  async function main(args) {
@@ -1617,6 +1683,12 @@ async function main(args) {
1617
1683
  const json = !!args.json;
1618
1684
  if (!verb || verb === "help") { console.log(USAGE); return 0; }
1619
1685
  if (verb === "plan-schema") { console.log(JSON.stringify(PLAN_SCHEMA, null, 2)); return 0; }
1686
+ if (verb === "size") {
1687
+ const r = aiTaskSize({ soloMin: Number(args["solo-min"]), project: args.project, switchMin: args["switch-min"] == null ? undefined : Number(args["switch-min"]) });
1688
+ if (json) console.log(JSON.stringify({ ok: true, exitCode: 0, ...r }, null, 2));
1689
+ else console.log(`${r.soloMin} solo min × ${r.multiplier} (${r.project}) + ${r.switchMin} min switching = ${r.hours} h (${r.days} d) → ${r.size} (${r.sizeDays} d)`);
1690
+ return 0;
1691
+ }
1620
1692
  const sheetId = sheetIdFromArg(args.sheet);
1621
1693
  const api = new SheetsApi(await getToken(args.key), sheetId);
1622
1694
  if (verb === "read") {
@@ -1701,10 +1773,10 @@ function haltAndExit(e, json) {
1701
1773
  }
1702
1774
 
1703
1775
  module.exports = {
1704
- validatePlan, splitRoster, rosterViolations, mfCoverageViolations, monthPlan, resampleWeights, rampHours, buildRoster,
1776
+ validatePlan, aiTaskSize, legendIsAi, defaultSwitchMin, splitRoster, rosterViolations, mfCoverageViolations, monthPlan, resampleWeights, rampHours, buildRoster,
1705
1777
  tshirtTotals, phaseTotals, buildRosters, midDays, deriveFteFromTeamMix, phaseTotalsFromSheet, verbTeamMix, phaseGapMap, verbPhases, verbTitles, itemFormulasFor, rollupFormulasFor, findCell, itemFormulas, rollupFormulas, tshirtRows, teamMixValues, teamMixFormatReqs, remainderFormula, locateTshirt, findPhaseSource,
1706
1778
  auditTshirt, auditTeamMix, auditTechStack, auditOverview, colLetter, hexToColor, colorToHex, sheetIdFromArg,
1707
- constants: { SIZE_CODES, PHASES, COLOR, RAMP, ROLE_LABEL, SOFT_CEILING, FOLD_THRESHOLD, TAB_TSHIRT, TAB_TEAM, TAB_TECH, PLAN_SCHEMA },
1779
+ constants: { SIZE_CODES, AI_SIZE_DAYS, PROJECT_MULTIPLIER, PHASES, COLOR, RAMP, ROLE_LABEL, SOFT_CEILING, FOLD_THRESHOLD, TAB_TSHIRT, TAB_TEAM, TAB_TECH, PLAN_SCHEMA },
1708
1780
  Halt, SheetsApi, getToken, runAudit, main,
1709
1781
  };
1710
1782
 
@@ -18,16 +18,16 @@
18
18
  * Frozen map: tier alias → concrete model id.
19
19
  * Consumers MUST import from here — never re-hardcode these strings.
20
20
  *
21
- * THREE tiers (Fable removed 2026-07-24): `opus` is now `claude-opus-5` — the
21
+ * THREE tiers (Fable removed 2026-07-24): `opus` is now `claude-opus-5-5` — the
22
22
  * default top tier. Opus 5 shipped at the SAME price as Opus 4.8 ($5/$25 per M
23
23
  * tokens) but >2× its coding score and within 0.5% of Fable 5 at max effort, so
24
24
  * the Fable cost premium ($10/$50 — double Opus 5) is no longer justified. Every
25
- * stage formerly on `fable` OR `opus` (4.8) now runs `opus` = claude-opus-5.
25
+ * stage formerly on `fable` OR `opus` (4.8) now runs `opus` = claude-opus-5-5.
26
26
  *
27
27
  * @type {Readonly<{opus: string, sonnet: string, haiku: string}>}
28
28
  */
29
29
  const MODEL_IDS = Object.freeze({
30
- opus: 'claude-opus-5',
30
+ opus: 'claude-opus-5-5',
31
31
  sonnet: 'claude-sonnet-4-6',
32
32
  haiku: 'claude-haiku-4-5-20251001',
33
33
  });
@@ -38,7 +38,7 @@ const MODEL_IDS = Object.freeze({
38
38
 
39
39
  /**
40
40
  * Frozen map: stage key → tier alias.
41
- * Fable removed 2026-07-24: all 7 stages resolve to `opus` (= claude-opus-5).
41
+ * Fable removed 2026-07-24: all 7 stages resolve to `opus` (= claude-opus-5-5).
42
42
  * The M82 competition judge-blindness invariant is RELAXED from "different model"
43
43
  * to "fresh independent context" — producers AND judge both run Opus 5 (fresh
44
44
  * contexts remove memory-bias; the modest residual taste/blind-spot bias is
@@ -67,7 +67,7 @@ const STAGE_TIERS = Object.freeze({
67
67
  *
68
68
  * This predicate existed for `claude-fable-5`, which returned HTTP 400 when the
69
69
  * explicit thinking-disabled parameter was sent. Fable was removed 2026-07-24;
70
- * NO current tier model (opus=claude-opus-5, sonnet, haiku) is known to require
70
+ * NO current tier model (opus=claude-opus-5-5, sonnet, haiku) is known to require
71
71
  * omission — Opus 5 and Sonnet 5 default `effort:high` on the API and accept the
72
72
  * thinking params normally. Kept as a single-home predicate (callers still import
73
73
  * it) so a future model that needs omission is added HERE, never re-hardcoded.
@@ -113,10 +113,10 @@ function resolve(stageKey) {
113
113
  * Frozen profile → stage-key → tier map.
114
114
  *
115
115
  * Fable removed 2026-07-24 — profiles now dial OPUS-vs-SONNET spend (not Fable):
116
- * standard — cost-leanest: the high-stakes reasoning stages run sonnet,
117
- * only the probes stay opus.
116
+ * standard — cost-leanest: probes + red-team run opus; judge,
117
+ * pre-mortem and debug-cycle-2 run sonnet.
118
118
  * pro — mid: red-team + pre-mortem + debug-cycle-2 → opus; the rest sonnet.
119
- * premium — full opus posture: all 6 designated stages → opus (= claude-opus-5).
119
+ * premium — full opus posture: all 6 designated stages → opus (= claude-opus-5-5).
120
120
  *
121
121
  * competition-producers is held at opus in ALL profiles (always opus-5). The
122
122
  * former judge≠producers blindness clamp is REMOVED — the invariant is now
@@ -130,7 +130,7 @@ const PROFILE_STAGE_TIERS = Object.freeze({
130
130
  'partition-probe': 'opus',
131
131
  'competition-judge': 'sonnet',
132
132
  'pre-mortem': 'sonnet',
133
- 'red-team': 'sonnet',
133
+ 'red-team': 'opus',
134
134
  'debug-cycle-2': 'sonnet',
135
135
  }),
136
136
  pro: Object.freeze({
@@ -161,8 +161,8 @@ const INJECTABLE_STAGES = Object.freeze([
161
161
  'debug-cycle-2',
162
162
  ]);
163
163
 
164
- /** The HELD producers model id (always opus = claude-opus-5). */
165
- const PRODUCERS_MODEL_ID = MODEL_IDS.opus; // claude-opus-5
164
+ /** The HELD producers model id (always opus = claude-opus-5-5). */
165
+ const PRODUCERS_MODEL_ID = MODEL_IDS.opus; // claude-opus-5-5
166
166
 
167
167
  /**
168
168
  * Resolves the concrete model id for a given stage key under a profile,
@@ -25,14 +25,15 @@ Read from `$ARGUMENTS` or `.gsd-t/estimate-config.json` if present; otherwise us
25
25
  |-------|-----------------|---------|
26
26
  | `rate` | `$50/hr` | Blended hourly rate for the LOW figure. |
27
27
  | `hoursPerDay` | `8` | Hours per person-day. |
28
- | `sizeScale` | `XS 0.25 · S 0.5 · M 1 · L 3 · XL 5 · XXL 7` | T-shirt → person-days. |
28
+ | `sizeScale` | `XS 0.1 · S 0.25 · M 0.5 · L 1 · XL 2 · XXL 4` | AI-assisted T-shirt → person-days (spec §1.4). `write` puts it in the sheet legend. |
29
+ | `projectMultiplier` | greenfield solo ×1 · team ×5 · yellow-field solo ×2 · team ×8 isolated / ×12 wide | Solo AI minutes × this. Internal to the estimator — never on the sheet. |
29
30
  | `totalMF` | `0.7` | Overhead multiplier. **The sheet's own MF list (`E4:F9`) wins when a sheet exists** — read it, never overwrite it. Hilo sheets run `0.9` (QA .3 · PM .1 · Analysis .1 · Deployment .05 · StdUps/Mtgs .15 · Buffer .2). |
30
31
  | `highFactor` | `1.25` | HIGH = LOW × this (the sheet's `G4` wins when a sheet exists). |
31
32
  | `sheetTemplateId` | (blank) | Optional template to clone; normally blank — the operator supplies the target sheet. |
32
33
  | `gcpProject` | `ai-estimator-415612` | GCP project hosting the permanent Sheets-writer SA. |
33
34
  | `serviceAccountEmail` | `gsd-t-sheets-writer@ai-estimator-415612.iam.gserviceaccount.com` | **Permanent** SA — share each sheet with this as Editor. |
34
35
  | `serviceAccountKeyPath` | `~/.claude/gsd-t-secrets/gsd-t-sheets-writer-key.json` | SA key (chmod 600, outside any repo). |
35
- | `newTeamDefault` | `true` | Apply the new-team familiarization adjustment (Step 2.5a) by default. |
36
+ | `newTeamDefault` | `false` | Add a new-team familiarization adjustment (Step 2.5a) only when the operator says the team is new to the code. |
36
37
 
37
38
  ## Step 0: Inputs + Scope + Sheet
38
39
 
@@ -42,7 +43,7 @@ Read from `$ARGUMENTS` or `.gsd-t/estimate-config.json` if present; otherwise us
42
43
  4. **Resolve the input document** (`--input`, else `.gsd-t/techdebt.md`). None → "No input document found. Pass `--input <path>` or run `/gsd-t-scan` / `/gsd-t-gap-analysis` first." and stop.
43
44
  5. **Classify the input** so the line-item vocabulary matches: scan register → *findings* (`TD-n`) scoped by severity; gap-analysis sheet → *gaps* (`GA-n`, rows already on the T-Shirt tab with columns A–D filled — you fill E–G only); requirements / feature / app spec → *requirements* (`FR-n` or the doc's own numbering).
44
45
  6. **Scope**: scan default = all CRITICAL findings; `--severity high|medium|low|all` widens. Requirements default = all. Confirm scope + item count with the user before sizing.
45
- 7. **Confirm the active config values** — rate, MF list (from the sheet), high factor, and whether this is a **new-team project** (Step 2.5a; usually YES for a fresh client).
46
+ 7. **Confirm the active config values** — rate, MF list (from the sheet), high factor, and the **project type**: greenfield or yellow-field (an existing app), solo or team (spec §1.4). Nobody hand-writes code — every estimate assumes AI-assisted development by a code-familiar team unless the operator says otherwise.
46
47
 
47
48
  ## Step 1: Numbering hygiene (MECHANICAL — show result)
48
49
 
@@ -57,22 +58,22 @@ Client-facing line-items carry **sequential, rational numbering starting at 1**.
57
58
 
58
59
  For each in-scope item build a row per spec §1.2 — `A` Module · `B` User Type · `C` Functionality (**with the item id**) · `D` Low-Level Requirement · `E` Phase · `F` Web Portal size · `G` Backend/API size. `H:L` are formulas, never values.
59
60
 
60
- - **Size each column INDEPENDENTLY** (FE and BE each get their own letter; blank = 0). Scale: **XS 0.25 · S 0.5 · M 1 · L 3 · XL 5 · XXL 7** person-days.
61
- - **Bare codes in `F:G`** — `XS` `S` `M` `L` `XL` `XXL`. Never the legend text (`"XS - Extra Small"`). The lookup happens to compute either way, which is why the long form shipped unnoticed.
62
- - The sheet computes: `Days = F+G` → `MFactor Days = Days × Total MF` → `Total Days` → `LOW $ = Total × 8 × rate` → `HIGH $ = LOW × high factor`.
61
+ - **Size in solo AI minutes, then let the tool pick the size** (spec §1.4). For each column (FE, BE) estimate the SOLO AI-assisted minutes — one person directing Claude — then run `gsd-t estimate-sheet size --solo-min <n> --project <type>`: it multiplies by the project type (greenfield solo ×1 · team ×5 · yellow-field solo ×2 · team ×8 isolated / ×12 big blast radius), adds task switching after the multiplier, and prints the size. Blast radius is measured with `gsd-t graph blast-radius`, not guessed. Count switching once per item (pass `--switch-min 0` for the smaller column).
62
+ - **Bare codes in `F:G`** — `XS` `S` `M` `L` `XL` `XXL`. Never the legend text (`"XS - Extra Small"`). Scale: **XS 0.1 · S 0.25 · M 0.5 · L 1 · XL 2 · XXL 4** person-days.
63
+ - The sheet computes: `Days = F+G` → `MFactor Days = Days × Total MF` → `Total Days` → `LOW $ = Total × 8 × rate` → `HIGH $ = LOW × high factor`. The overhead factors and high factor are per-project settings the operator adjusts by hand.
63
64
  - **Cluster by fix-shape to size fast**: "add existing guard to N routes" (XS–S, repeated) vs "new backend surface" (M, +FE) vs "config / single route" (XS). Size the cluster once, apply to members.
64
65
  - **Tune the MF per project** (raise Buffer/QA when confidence is low; raise the high factor above 1.25 for more unknowns) — but change the sheet's MF list only with the operator's say-so.
65
66
  - **PAUSE:** present the sized rows (or clusters + representative sizes) and the running total (recomputed from raw sizes: `Σ(FE,BE days) × (1 + MF)`). Wait for `continue` or corrections.
66
67
 
67
68
  ## Step 2.5: Estimate Adjustments (familiarization + risk/unknowns) — JUDGMENT · PAUSE FOR REVIEW
68
69
 
69
- Base sizes assume *familiar* devs on *well-understood* work. Adjust for the two things that make real work heavier. Document each adjustment per-item so the client sees **why**. **This step is ON by default (`newTeamDefault: true`) — it was skipped on shipped estimates and had to be asked for.**
70
+ Base sizes assume *familiar* devs on *well-understood* work. Adjust for the two things that make real work heavier. Document each adjustment per-item so the client sees **why**. Part (b) always runs; part (a) only for a team new to the code.
70
71
 
71
- **(a) New-team familiarization** — bump each item's SIZE in proportion to its complexity — **NOT the MF** (the Analysis MF is for a Business Analyst, not dev ramp). Trivial config / single route → no bump. Repeated-pattern guards, few routes → +0–1 tier. High-volume sweeps + new-surface builds → +1 tier. Optionally add a one-time **"Codebase Onboarding & Downstream Analysis"** Common line (L–XL), documented as optional.
72
+ **(a) New-team familiarization — OFF by default** (`newTeamDefault: false`; estimates assume a code-familiar team). Only when the operator says the team is new to the code: bump each item's SIZE in proportion to its complexity — **NOT the MF** (the Analysis MF is for a Business Analyst, not dev ramp). Trivial config / single route → no bump. Repeated-pattern guards, few routes → +0–1 tier. High-volume sweeps + new-surface builds → +1 tier. Optionally add a one-time **"Codebase Onboarding & Downstream Analysis"** Common line (L–XL), documented as optional.
72
73
 
73
74
  **(b) R&D / unknown-approach / spike risk** — an item needing research, an unproven approach, or an unknown integration gets an uplift for the uncertainty: bump its SIZE or raise the high factor if unknowns dominate. Name the unknown explicitly ("requires spike: undocumented 3rd-party API").
74
75
 
75
- **⚠️ The scale is NON-LINEAR. M→L is a 3× cliff (1 day → 3 days).** Never push an item across M→L unless it is genuinely multi-day. Cap routine-work bumps at M. Calibration: HILO 21 criticals = $8,700 familiar → $11,730 new-team (+35%, bumps capped at M).
76
+ **⚠️ The scale doubles at every step (0.1 → 0.25 → 0.5 → 1 → 2 → 4 days).** An uplift is one step at most unless the unknown is genuinely multi-day. Calibration: the Hilo Delivery Runway build — 45 tasks, ~26 solo hours (13 David + 13 Claude) — which the old day scale priced at 1,262–1,577 hours.
76
77
 
77
78
  - **PAUSE:** present every adjustment (item, reason, before→after size, total delta). Wait for `continue` or corrections.
78
79
 
@@ -381,7 +381,7 @@ Use these when user asks for help on a specific command:
381
381
  - **Summary**: Turn any structured work document — a scan register, a gap-analysis sheet, a new-feature or new-app requirements doc — into a Tekyz client estimate Google Sheet: **T-Shirt Size + Team Mix + Technology Stack**, written and audited against `~/.claude/templates/estimate-sheet-spec.md`. No PRD (`/gsd-t-prd` owns that)
382
382
  - **Auto-invoked**: No
383
383
  - **Updates**: the Tekyz estimate Google Sheet (three tabs, written and read-back-audited by `gsd-t estimate-sheet` from `.gsd-t/estimate-plan.json`) + optional `share/<Repo>-estimate-redteam-notes.md` (and, if renumbered, the source doc/docs/scan files)
384
- - **Use when**: You need a client-facing paid estimate (T-shirt sizing, dollar range, staffed team by month) from a scan, a gap analysis, or a requirements/feature/app spec. **SUPERVISED** — judgment phases (sizing, adjustments, Team Mix, Red Team) pause for your review; **you are the final arbiter** of an Estimate Red Team that challenges the numbers. Accepts `--sheet <url>`. Rate + factors are parameterized (default Tekyz; the sheet's own MF list wins). Playbook: `~/.claude/playbooks/tekyz-estimation-and-prd-playbook.md`
384
+ - **Use when**: You need a client-facing paid estimate (T-shirt sizing, dollar range, staffed team by month) from a scan, a gap analysis, or a requirements/feature/app spec. **SUPERVISED** — judgment phases (sizing, adjustments, Team Mix, Red Team) pause for your review; **you are the final arbiter** of an Estimate Red Team that challenges the numbers. Accepts `--sheet <url>`. Sizes are AI-assisted (solo AI minutes × project multiplier + task switching → `gsd-t estimate-sheet size`; spec §1.4). Rate + factors are parameterized (default Tekyz; the sheet's own MF list wins). Playbook: `~/.claude/playbooks/tekyz-estimation-and-prd-playbook.md`
385
385
 
386
386
  ### stories
387
387
  - **Summary**: Generate a dev-team handoff document in the Tekyz user-stories format — discrete user stories with workflows, grouped acceptance criteria, per-story flow diagrams (Mermaid rendered to embedded images), and mapped test-case tables — from any source (scan register, requirements doc, design contract, or a reverse-engineered codebase)
@@ -536,7 +536,7 @@ Use these when user asks for help on a specific command:
536
536
  - **Contract**: `.gsd-t/contracts/plan-hardening-contract.md` v1.0.0 STABLE.
537
537
 
538
538
  ### model-tier-policy (M85)
539
- - **Summary**: SINGLE source of truth for GSD-T model-tier assignments. Publishes the authoritative tier set (haiku/sonnet/opus — **Fable removed 2026-07-24**; `opus` = claude-opus-5) and the 7 designated stage→tier mappings (all → opus). Opus 5 shipped at the same price as Opus 4.8 but >2× its coding score, so every stage formerly on Fable OR Opus 4.8 now runs Opus 5. The M82 competition judge-blindness invariant is RELAXED to "fresh independent context" — producers AND judge both run opus. A M71-family lint (`test/m85-workflow-tier-policy-lint.test.js`) proves every workflow `model:` literal matches the policy — a drifted literal FAILS the lint (mandatory negative test).
539
+ - **Summary**: SINGLE source of truth for GSD-T model-tier assignments. Publishes the authoritative tier set (haiku/sonnet/opus — **Fable removed 2026-07-24**; `opus` = claude-opus-5-5) and the 7 designated stage→tier mappings (all → opus). Opus 5 shipped at the same price as Opus 4.8 but >2× its coding score, so every stage formerly on Fable OR Opus 4.8 now runs Opus 5. The M82 competition judge-blindness invariant is RELAXED to "fresh independent context" — producers AND judge both run opus. A M71-family lint (`test/m85-workflow-tier-policy-lint.test.js`) proves every workflow `model:` literal matches the policy — a drifted literal FAILS the lint (mandatory negative test).
540
540
  - **Files**: `bin/gsd-t-model-tier-policy.cjs` (zero external deps — installer invariant).
541
541
  - **Use when**: Any phase that needs to resolve a concrete model id from a stage key at invoke time (M69 pattern). Workflows NEVER `require` this module (sandbox ban) — they use hard-coded tier alias literals the lint proves match the policy.
542
542
  - **CLI**: `gsd-t model-tier-policy resolve <stageKey> [--json]`. Emits `{ok, stageKey, tier, model, requiresThinkingOmitted}`. Exit 0 resolved · 1 unknown stage key.
@@ -88,8 +88,8 @@ The **Model Profile** line MUST always name the active profile — never blank,
88
88
  - If the file is absent, display the global default by name with the `(default)` marker: `Model Profile: premium (default)`.
89
89
  - If the file is present but malformed or contains an unknown profile, display: `Model Profile: premium (default, config-error)` — never silently promote to the most expensive posture.
90
90
 
91
- Profiles control which workflow stages run on Opus vs. Sonnet (Fable removed 2026-07-24 — `opus` = claude-opus-5):
92
- - `standard` — probes opus; high-stakes stages (judge/pre-mortem/red-team/debug-cycle-2) sonnet (cost-leanest)
91
+ Profiles control which workflow stages run on Opus vs. Sonnet (Fable removed 2026-07-24 — `opus` = claude-opus-5-5):
92
+ - `standard` — probes + red-team opus; judge/pre-mortem/debug-cycle-2 sonnet (cost-leanest)
93
93
  - `pro` — probes + pre-mortem + red-team + debug-cycle-2 opus; judge sonnet
94
94
  - `premium` — all 6 designated stages opus (full posture, global default)
95
95
 
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@tekyzinc/gsd-t",
3
- "version": "5.20.15",
4
- "description": "GSD-T: Contract-Driven Development for Claude Code — 54 slash commands with headless-by-default workflow spawning, unattended supervisor relay with event stream, graph-powered code analysis, real-time agent dashboard, task telemetry, doc-ripple enforcement, backlog management, impact analysis, test sync, milestone archival, and PRD generation",
3
+ "version": "5.22.10",
4
+ "description": "GSD-T: Contract-Driven Development for Claude Code \u2014 54 slash commands with headless-by-default workflow spawning, unattended supervisor relay with event stream, graph-powered code analysis, real-time agent dashboard, task telemetry, doc-ripple enforcement, backlog management, impact analysis, test sync, milestone archival, and PRD generation",
5
5
  "author": "Tekyz, Inc.",
6
6
  "license": "MIT",
7
7
  "repository": {
@@ -263,7 +263,7 @@ Every GSD-T project gets TWO logging streams scaffolded by default at `gsd-t-ini
263
263
  Every code-producing phase ends with `gsd-t-verify.workflow.js`, which runs three orthogonal validators as `parallel()` `agent()` stages with schema-validated output. Per `.gsd-t/contracts/orthogonal-validation-contract.md` v1.0.0 STABLE, they are declared orthogonal objective functions — no collapse, no substitution, no transitive trust.
264
264
 
265
265
  - **`/code-review ultra`** — cooperative correctness + cleanup. Severity: `important` / `nit` / `pre-existing`. Skippable via `args.skipUltra=true` + `args.skipUltraReason`. `skipUltra=true` is INELIGIBLE for `VERIFIED`.
266
- - **Red Team** — adversarial / security / boundaries. Non-skippable. Protocol: `templates/prompts/red-team-subagent.md`. Verdict: `FAIL` (any CRITICAL or HIGH bug — blocks completion) or `GRUDGING-PASS` (exhaustive search, nothing found). CRITICAL/HIGH bugs get up to 2 fix cycles before deferral. Runs on `model: "opus"` (= claude-opus-5; Fable removed 2026-07-24).
266
+ - **Red Team** — adversarial / security / boundaries. Non-skippable. Protocol: `templates/prompts/red-team-subagent.md`. Verdict: `FAIL` (any CRITICAL or HIGH bug — blocks completion) or `GRUDGING-PASS` (exhaustive search, nothing found). CRITICAL/HIGH bugs get up to 2 fix cycles before deferral. Runs on `model: "opus"` (= claude-opus-5-5; Fable removed 2026-07-24).
267
267
  - **QA** — test execution + shallow-test detection + contract compliance. Non-skippable. Protocol: `templates/prompts/qa-subagent.md`. Writes ZERO feature code. Any shallow E2E test blocks phase completion. Runs on `model: "sonnet"`.
268
268
 
269
269
  When `.gsd-t/contracts/design-contract.md` or `.gsd-t/contracts/design/` exists, a fourth stage runs Design Verification (protocol: `templates/prompts/design-verify-subagent.md`) — opens a browser, compares the build against the design, returns a structured element-by-element MATCH/DEVIATION schema. Deviations block completion.
@@ -272,13 +272,13 @@ Synthesis stage merges results without category collapse. Verdict: `VERIFIED` /
272
272
 
273
273
  ## Model Display (MANDATORY)
274
274
 
275
- **Each Workflow `agent()` call declares its model explicitly** via the `model:` option (`"haiku"` / `"sonnet"` / `"opus"` — **Fable removed 2026-07-24**; `opus` = claude-opus-5). The Workflow runtime emits a `⚙ [{model}] {label}` line per stage in `/workflows`, giving the user real-time visibility into which model handles each operation.
275
+ **Each Workflow `agent()` call declares its model explicitly** via the `model:` option (`"haiku"` / `"sonnet"` / `"opus"` — **Fable removed 2026-07-24**; `opus` = claude-opus-5-5). The Workflow runtime emits a `⚙ [{model}] {label}` line per stage in `/workflows`, giving the user real-time visibility into which model handles each operation.
276
276
 
277
277
  **Model assignments:**
278
278
  - `model: "haiku"` — strictly mechanical tasks: run test suites and report counts, check file existence, validate JSON structure, branch guard checks
279
279
  - `model: "sonnet"` — mid-tier reasoning: routine code changes, standard refactors, test writing, QA evaluation, straightforward synthesis
280
280
  - `model: "opus"` — high-stakes reasoning: architecture decisions, security analysis, complex debugging, cross-module refactors, quality judgment on critical paths
281
- - **Fable removed 2026-07-24** — `opus` (= claude-opus-5) is the top tier and the default for every high-stakes stage: solution-space probe, partition probe, competition judge, pre-mortem, Red Team, competition producers, debug both cycles. Opus 5 shipped at the same price as Opus 4.8 but >2× its coding score, so every stage formerly on Fable OR Opus 4.8 now runs Opus 5. The M82 competition judge-blindness invariant is RELAXED to "fresh independent context" — producers AND judge both run `opus` (fresh contexts remove memory-bias; the modest residual taste/blind-spot bias is accepted for a stronger judge). **Single source of truth for tier assignments:** `bin/gsd-t-model-tier-policy.cjs` + `.gsd-t/contracts/model-tier-policy-contract.md`. The M71-family lint (`test/m85-workflow-tier-policy-lint.test.js`) proves every workflow `model:` literal matches the policy and a drifted literal FAILS the lint (mandatory negative test).
281
+ - **Fable removed 2026-07-24** — `opus` (= claude-opus-5-5) is the top tier and the default for every high-stakes stage: solution-space probe, partition probe, competition judge, pre-mortem, Red Team, competition producers, debug both cycles. Opus 5 shipped at the same price as Opus 4.8 but >2× its coding score, so every stage formerly on Fable OR Opus 4.8 now runs Opus 5. The M82 competition judge-blindness invariant is RELAXED to "fresh independent context" — producers AND judge both run `opus` (fresh contexts remove memory-bias; the modest residual taste/blind-spot bias is accepted for a stronger judge). **Single source of truth for tier assignments:** `bin/gsd-t-model-tier-policy.cjs` + `.gsd-t/contracts/model-tier-policy-contract.md`. The M71-family lint (`test/m85-workflow-tier-policy-lint.test.js`) proves every workflow `model:` literal matches the policy and a drifted literal FAILS the lint (mandatory negative test).
282
282
 
283
283
  **Context budget:** Workflow scripts receive a `budget` global (`budget.total`, `budget.spent()`, `budget.remaining()`) tied to the user's per-turn token target. Use it for dynamic loops (`while (budget.total && budget.remaining() > 50_000) { ... }`) or to scale fleet size. Opus 4.7/4.8 ship 1M context windows; the legacy meter at `bin/token-budget.cjs` was retired in M61 — use native `/context` for live in-session usage.
284
284
 
@@ -7,8 +7,11 @@
7
7
  "hoursPerDay": 8,
8
8
  "_hoursPerDay": "Hours per person-day. Default: 8.",
9
9
 
10
- "sizeScale": { "XS": 0.25, "S": 0.5, "M": 1, "L": 3, "XL": 5, "XXL": 7 },
11
- "_sizeScale": "T-shirt size -> person-days. NON-LINEAR: M->L is a 3x cliff (1 -> 3 days). Do not push routine work across M->L.",
10
+ "sizeScale": { "XS": 0.1, "S": 0.25, "M": 0.5, "L": 1, "XL": 2, "XXL": 4 },
11
+ "_sizeScale": "AI-assisted T-shirt scale, person-days (written into the sheet legend by `gsd-t estimate-sheet write`). A size = solo AI minutes x project multiplier + task switching; see estimate-sheet-spec.md section 1.4.",
12
+
13
+ "projectMultiplier": { "greenfield-solo": 1, "greenfield-team": 5, "yellowfield-solo": 2, "yellowfield-team-isolated": 8, "yellowfield-team-wide": 12 },
14
+ "_projectMultiplier": "Solo AI minutes are multiplied by this, by project type (isolated vs wide = code-graph blast radius). Internal to the estimator; never shown on the sheet.",
12
15
 
13
16
  "totalMF": 0.7,
14
17
  "_totalMF": "Overhead multiplier applied to raw Days. Tekyz default 0.7 = QA 0.3 + PM 0.1 + Analysis 0.05 + Deployment 0.05 + Buffer 0.2. Raise Buffer/QA when confidence is low.",
@@ -34,8 +37,8 @@
34
37
  "serviceAccountKeyPath": "~/.claude/gsd-t-secrets/gsd-t-sheets-writer-key.json",
35
38
  "_serviceAccountKeyPath": "Local path to the SA JSON key (chmod 600, outside any repo — NEVER commit). Step 5 self-provisions the SA+key here if missing.",
36
39
 
37
- "newTeamDefault": true,
38
- "_newTeamDefault": "Whether to assume a new-team familiarization adjustment (Step 2.5a) by default. Usually true for a fresh client.",
40
+ "newTeamDefault": false,
41
+ "_newTeamDefault": "Whether to add a new-team familiarization adjustment (Step 2.5a). Default false: estimates assume a code-familiar, AI-assisted team. Turn on only when the operator says the team is new to the code.",
39
42
 
40
43
  "sheetSpec": "~/.claude/templates/estimate-sheet-spec.md",
41
44
  "_sheetSpec": "The layout/formula/styling/roster spec every write is audited against. Bundled at templates/estimate-sheet-spec.md in the GSD-T package."
@@ -42,7 +42,7 @@ Reference implementation (read it, don't guess): "ATP SOW Gap Analysis and Estim
42
42
  | `A1` | `Date Submitted` | bg `#D9D9D9`, bold, Calibri |
43
43
  | `B1` | date `mm/dd/yyyy` | Calibri 11, left |
44
44
  | `A3:B3` | `Legends` / `Person Days` | bg `#3D85C6`, white bold Calibri |
45
- | `A4:B9` | **Size legend** — `XS - Extra Small` 0.25 · `S - Small` 0.5 · `M - Medium` 1 · `L - Large` 3 · `XL - Extra Large` 5 · `XXL - Extra Extra Large` 7 | Calibri, right |
45
+ | `A4:B9` | **Size legend** — labels `XS - Extra Small` … `XXL - Extra Extra Large`. **Values: the AI-assisted scale** `XS 0.1 · S 0.25 · M 0.5 · L 1 · XL 2 · XXL 4` person-days (§1.4). `write` puts these values in `B4:B9` (matched by label); the labels are never touched. Sheets written before v5.22.10 keep the old `0.25 · 0.5 · 1 · 3 · 5 · 7` scale until they are re-estimated. | Calibri, right |
46
46
  | `E3` | `Multiplication Factor` | bg `#3D85C6`, white bold |
47
47
  | `E4:F9` | **MF list** — one row per factor (e.g. `QA` 0.3 · `PM` 0.1 · `Analysis` 0.1 · `Deployment` 0.05 · `StdUps/Mtgs` 0.15 · `Buffer` 0.2). **Read the live list — it varies per sheet and is the roster contract for Team Mix (§2.4).** | |
48
48
  | `E10:F10` | `Total MF` = `=sum(F4:F9)` | bg `#C9DAF8`, bold |
@@ -78,7 +78,28 @@ Column widths: `[150, 120, 300, 430, 122, 90, 90, 61, 53, 76, 81, 81]`.
78
78
  | `K` | `=J{r}*8*$H$4` | nf `$#,##0.00` |
79
79
  | `L` | `=K{r}*$G$4` | nf `$#,##0.00` |
80
80
 
81
- ### 1.3 Totals and summary (directly under the last item — NO blank rows)
81
+ #### 1.4 How a size is chosen — the AI-assisted model (David, 2026-09-27)
82
+
83
+ Nobody hand-writes code. A size is the **team hours** a task takes with AI-assisted development, and the multipliers that produce it live in the estimator — never on the sheet. Calibration: the Hilo Delivery Runway build (45 tasks, ~26 solo hours: 13 David + 13 Claude).
84
+
85
+ 1. **Estimate the task in SOLO AI-assisted minutes** — one person directing Claude, greenfield. Runway rate: roughly 5 min for a trivial change, 20 min for a typical screen element or endpoint, 1–2¼ hrs for the heaviest pieces.
86
+ 2. **Multiply by the project type:**
87
+
88
+ | Project | Multiplier |
89
+ |---|---|
90
+ | Greenfield, solo | × 1 |
91
+ | Greenfield, team | × 5 |
92
+ | Yellow-field (existing app), solo | × 2 |
93
+ | Yellow-field, team — isolated change | × 8 |
94
+ | Yellow-field, team — big blast radius | × 12 |
95
+
96
+ Blast radius comes from the code graph (`gsd-t graph blast-radius`), not a guess.
97
+ 3. **Add task switching AFTER the multiplier** — it is one person's pickup time and does not grow with team size: 5–10 min for small tasks (less when related tasks run back-to-back), ~15 min medium, up to 30 min large.
98
+ 4. **Pick the nearest size** on the §1.1 legend. `gsd-t estimate-sheet size --solo-min <n> --project <type> [--switch-min <n>]` does steps 2–4 and prints the size.
99
+
100
+ The overhead factors (`E4:F9`) and the high factor stay per-project settings the operator adjusts by hand; they are not part of the task math.
101
+
102
+ ## 1.3 Totals and summary (directly under the last item — NO blank rows)
82
103
 
83
104
  ```
84
105
  <last item row>
@@ -252,6 +273,7 @@ Everything in §1–§5 is executed by `bin/gsd-t-estimate-sheet.cjs`, not re-de
252
273
 
253
274
  ```
254
275
  gsd-t estimate-sheet plan-schema # the plan shape
276
+ gsd-t estimate-sheet size --solo-min <n> --project <type> [--switch-min <n>] # §1.4: solo minutes → team hours → size (no sheet needed)
255
277
  gsd-t estimate-sheet read --sheet <id|url> [--tab <name>] # read-before-write dump
256
278
  gsd-t estimate-sheet plan-check --sheet <id|url> --plan plan.json # validate + the roster it WOULD write (the Step 4 pause)
257
279
  gsd-t estimate-sheet write --sheet <id|url> --plan plan.json [--replace] # T-Shirt + Team Mix + Tech Stack, then audit
@@ -286,4 +308,4 @@ The plan (judgment only):
286
308
 
287
309
  - `tshirt.mode` `items` writes whole rows below the header (halts if rows exist unless `--replace`); `sizes` fills `E:L` on rows that already exist (a gap-analysis sheet), matched by the `(id)` suffix in column C — never by position.
288
310
  - `teamMix.fte` is per-discipline FTE (`backend` `frontend` `qa` `pm` `ba` `devops` `techlead` `design` `mobile`). The tool splits it into people (saturate then spill), computes months and the column count, ramps by discipline, writes the remainder formula, and refuses a roster that leaves a weighted MF factor unstaffed. It writes **one grid per phase with hours** (§2.6); `teamMix.phases: { "Phase 1": { "fte": {…} } }` overrides the mix for one phase.
289
- - The MF list, legend, rate and high factor are READ from the sheet; the plan never carries them.
311
+ - The MF list, rate and high factor are READ from the sheet; the plan never carries them. The legend VALUES are written by the tool (the AI-assisted scale, §1.4) — in `sizes` mode it HALTS if a sized row on the tab is missing from the plan, because moving the legend would silently re-price that row.
@@ -50,17 +50,21 @@ For each finding, write a row (cols A–G; leave H–L formulas alone):
50
50
  `A` Module · `B` User Type · `C` Functionality (include the `(TD-n)`) ·
51
51
  `D` Low-Level Requirement · `E` Phase (MVP) · `F` Web Portal size · `G` Backend/API size.
52
52
 
53
- - Size **each column independently** (FE and BE). Sizes: XS .25, S .5, M 1, L 3, XL 5, XXL 7.
53
+ - Size **each column independently** (FE and BE), **AI-assisted** — nobody hand-writes code.
54
+ Estimate SOLO AI minutes, then `gsd-t estimate-sheet size --solo-min <n> --project <type>`
55
+ (greenfield solo ×1 · team ×5 · yellow-field solo ×2 · team ×8 isolated / ×12 wide, + task
56
+ switching after the multiplier) picks the size. Sizes: XS .1, S .25, M .5, L 1, XL 2, XXL 4
57
+ (estimate-sheet-spec.md §1.4).
54
58
  - Sheet computes: `Days = F+G`, `MFactor = Days×MF`, `Total = Days+MFactor`,
55
59
  `LOW$ = Total×8×$50`, `HIGH$ = LOW$×1.25`.
56
60
  - **Cluster the work** to size fast: "add existing auth guard to routes" (XS–S,
57
61
  repeated pattern) vs "new backend surface" (M, +FE) vs "config/1-route" (XS).
58
62
 
59
- ### Familiarization bump (new-team projects)
60
- Base sizes assume *familiar* devs. For a new team, bump SIZE in proportion to
63
+ ### Familiarization bump (new-team projects only — off by default)
64
+ Base sizes assume a *code-familiar*, AI-assisted team. Only when the operator says the team is new, bump SIZE in proportion to
61
65
  complexity (NOT the MF — Analysis MF is for a Business Analyst):
62
66
  - Trivial → no bump. Repeated-pattern guards → +0-1 tier. High-volume sweeps +
63
- new-surface → +1 tier. **Never cross the M→L cliff (3×) unless genuinely multi-day.**
67
+ new-surface → +1 tier (the scale doubles per step — one step at most).
64
68
  - Optionally add a one-time "Codebase Onboarding & Downstream Analysis" Common line (L–XL).
65
69
 
66
70
  ### Tune the MF (per project)
@@ -4,7 +4,7 @@
4
4
  **Report concisely:** verdict/answer first, no preamble. Gloss every code/jargon term (e.g. `M93-D2` = milestone 93, domain 2) in plain words on first use. Bullets over paragraphs. Expand only if asked.
5
5
  <!-- /reader-contract -->
6
6
 
7
- **Model:** `opus` (= claude-opus-5; Fable removed 2026-07-24 — highest-leverage judgment; separate context from the proposing agent)
7
+ **Model:** `opus` (= claude-opus-5-5; Fable removed 2026-07-24 — highest-leverage judgment; separate context from the proposing agent)
8
8
 
9
9
  **Framing:** You are reviewing someone ELSE's architectural design — you did NOT propose it and have no attachment to it. Your goal is to find the **fatal flaw** in the premise being challenged, before a single line of code is committed to that premise. This framing (independent reviewer, not the author) is essential for escaping self-preference bias: the proposing agent's prior context makes it systematically less able to see its own premise's failures (source: https://arxiv.org/abs/2310.08118 — LLM self-evaluation is biased toward confirming prior outputs; https://arxiv.org/abs/2404.13076 — blind adversarial framing surfaces failures that self-critique misses).
10
10
 
@@ -57,7 +57,7 @@ const _args = (typeof args === "string") ? (() => { try { return JSON.parse(args
57
57
  // Default to {} so the premium fallback literals apply when no invoker injects overrides.
58
58
  // overrides values are CONCRETE model ids (resolver envelope); the bare literals below
59
59
  // are tier ALIASES. The sandbox runtime accepts BOTH forms in model: — proven live for
60
- // the tier alias resolves to claude-opus-5 (Fable removed 2026-07-24).
60
+ // the tier alias resolves to claude-opus-5-5 (Fable removed 2026-07-24).
61
61
  const overrides = (_args.overrides && typeof _args.overrides === "object") ? _args.overrides : {};
62
62
  const _CLI_ENVELOPE_SCHEMA = {
63
63
  type: "object", required: ["ok", "exitCode"], additionalProperties: true,
@@ -83,7 +83,7 @@ const _args = (typeof args === "string") ? (() => { try { return JSON.parse(args
83
83
  // (preserves byte-identical M85 behavior for callers that have not been updated yet).
84
84
  // overrides values are CONCRETE model ids (resolver envelope); the bare literals below
85
85
  // are tier ALIASES. The sandbox runtime accepts BOTH forms in model: — proven live for
86
- // the tier alias resolves to claude-opus-5 (Fable removed 2026-07-24).
86
+ // the tier alias resolves to claude-opus-5-5 (Fable removed 2026-07-24).
87
87
  const overrides = (_args.overrides && typeof _args.overrides === "object") ? _args.overrides : {};
88
88
  // `envelope` is typed as an OBJECT (or null), not "any".
89
89
  //
@@ -57,7 +57,7 @@ const _args = (typeof args === "string") ? (() => { try { return JSON.parse(args
57
57
  // Default to {} so the premium fallback literals apply when no invoker injects overrides.
58
58
  // overrides values are CONCRETE model ids (resolver envelope); the bare literals below
59
59
  // are tier ALIASES. The sandbox runtime accepts BOTH forms in model: — proven live for
60
- // the tier alias resolves to claude-opus-5 (Fable removed 2026-07-24).
60
+ // the tier alias resolves to claude-opus-5-5 (Fable removed 2026-07-24).
61
61
  const overrides = (_args.overrides && typeof _args.overrides === "object") ? _args.overrides : {};
62
62
  const _CLI_ENVELOPE_SCHEMA = {
63
63
  type: "object", required: ["ok", "exitCode"], additionalProperties: true,