@hanzlaa/rcode 4.11.0 β†’ 4.12.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -24,7 +24,7 @@ If a user says "just keep going" or "don't stop until done", that authorization
24
24
 
25
25
  - Follow [Conventional Commits](https://www.conventionalcommits.org/) format: `type(scope): subject`
26
26
  - Types allowed: `feat`, `fix`, `docs`, `style`, `refactor`, `test`, `chore`, `perf`, `revert`
27
- - Scopes allowed: `agents`, `skills`, `workflows`, `templates`, `dashboard`, `docs`, `config`, `github`, `github-sync`, `commands`, `memory`, `brand`, `cli`, `ci`, `release`, `meta`, `tasks`, `migrations`, `refs`, `state`, `hooks`, `init`, `install`, `parity`, `triggers`, `dogfood`, `namespace`, `planning`, `insights`, `help`, `roadmap`, `session`, `audits`, `execute`, `executor`, `plan`, `planner`, `readme`, `rcode`, `review`, `sync`, `sprint`, `agent-exp`, `extensibility`, `lens-audit`, `tiers`, `build`, `council`, `doctor`, `postinstall`, `progress`, `security`, `tools`, `uninstall`, `update`, `test`, `changelog`, `scopes`, `phases`, `references`, `kanban`, `orchestrator`, `orchpanel`, `status`, `bin`, `brain`, `dogfeed`, `new-project`, `package`, `rcode-tools`, `rihal-tools`, `team`, `usp`, `v4`, `observability`, `audit`, `agent-rules`, `cursor`, `i18n`, `phase`, `scaffold`, `campaign`, `ship`, `getting-started`, `do-router`, `milestone-health`, `modules`, `project-types`, `roadmapper`, `token`, `benchmarks`, `eval`, `scan`, `verifier`, `discuss-phase`, `ui-phase`, plus numeric phase/sprint scopes (e.g. `docs(15)`, `feat(8.3)`)
27
+ - Scopes allowed: `agents`, `skills`, `workflows`, `templates`, `dashboard`, `docs`, `config`, `github`, `github-sync`, `commands`, `memory`, `brand`, `cli`, `ci`, `release`, `meta`, `tasks`, `migrations`, `refs`, `state`, `hooks`, `init`, `install`, `parity`, `triggers`, `dogfood`, `namespace`, `planning`, `insights`, `help`, `roadmap`, `session`, `audits`, `execute`, `executor`, `plan`, `planner`, `readme`, `rcode`, `review`, `sync`, `sprint`, `agent-exp`, `extensibility`, `lens-audit`, `tiers`, `build`, `council`, `doctor`, `postinstall`, `progress`, `security`, `tools`, `uninstall`, `update`, `test`, `changelog`, `scopes`, `phases`, `references`, `kanban`, `orchestrator`, `orchpanel`, `status`, `bin`, `brain`, `dogfeed`, `new-project`, `package`, `rcode-tools`, `rihal-tools`, `team`, `usp`, `v4`, `observability`, `audit`, `agent-rules`, `cursor`, `i18n`, `phase`, `scaffold`, `campaign`, `ship`, `getting-started`, `do-router`, `milestone-health`, `modules`, `project-types`, `roadmapper`, `token`, `benchmarks`, `eval`, `scan`, `verifier`, `discuss-phase`, `ui-phase`, `state-sync`, `autonomous`, plus numeric phase/sprint scopes (e.g. `docs(15)`, `feat(8.3)`)
28
28
  - Subject: lowercase first letter, imperative mood, no trailing period, under 72 chars
29
29
  - **NEVER add Claude/AI attribution to commit messages.** No "Generated with Claude Code", no "Co-Authored-By: Claude", no "πŸ€– Generated". The user does not want this.
30
30
  - **NEVER use `--no-verify`** to bypass hooks. If hooks fail, fix the underlying issue.
package/CLAUDE.md CHANGED
@@ -24,7 +24,7 @@ If a user says "just keep going" or "don't stop until done", that authorization
24
24
 
25
25
  - Follow [Conventional Commits](https://www.conventionalcommits.org/) format: `type(scope): subject`
26
26
  - Types allowed: `feat`, `fix`, `docs`, `style`, `refactor`, `test`, `chore`, `perf`, `revert`
27
- - Scopes allowed: `agents`, `skills`, `workflows`, `templates`, `dashboard`, `docs`, `config`, `github`, `github-sync`, `commands`, `memory`, `brand`, `cli`, `ci`, `release`, `meta`, `tasks`, `migrations`, `refs`, `state`, `hooks`, `install`, `parity`, `triggers`, `dogfood`, `namespace`, `planning`, `insights`, `help`, `roadmap`, `session`, `audits`, `execute`, `executor`, `plan`, `planner`, `readme`, `rcode`, `review`, `sync`, `sprint`, `agent-exp`, `extensibility`, `lens-audit`, `tiers`, `build`, `council`, `doctor`, `postinstall`, `progress`, `security`, `tools`, `uninstall`, `update`, `test`, `changelog`, `scopes`, `phases`, `references`, `kanban`, `orchestrator`, `orchpanel`, `status`, `bin`, `brain`, `dogfeed`, `new-project`, `package`, `rcode-tools`, `rihal-tools`, `team`, `usp`, `v4`, `observability`, `audit`, `init`, `agent-rules`, `cursor`, `i18n`, `phase`, `scaffold`, `campaign`, `ship`, `getting-started`, `do-router`, `milestone-health`, `modules`, `project-types`, `roadmapper`, `token`, `verifier`, `discuss-phase`, `ui-phase`, plus numeric phase/sprint scopes (e.g. `docs(15)`, `feat(8.3)`)
27
+ - Scopes allowed: `agents`, `skills`, `workflows`, `templates`, `dashboard`, `docs`, `config`, `github`, `github-sync`, `commands`, `memory`, `brand`, `cli`, `ci`, `release`, `meta`, `tasks`, `migrations`, `refs`, `state`, `hooks`, `install`, `parity`, `triggers`, `dogfood`, `namespace`, `planning`, `insights`, `help`, `roadmap`, `session`, `audits`, `execute`, `executor`, `plan`, `planner`, `readme`, `rcode`, `review`, `sync`, `sprint`, `agent-exp`, `extensibility`, `lens-audit`, `tiers`, `build`, `council`, `doctor`, `postinstall`, `progress`, `security`, `tools`, `uninstall`, `update`, `test`, `changelog`, `scopes`, `phases`, `references`, `kanban`, `orchestrator`, `orchpanel`, `status`, `bin`, `brain`, `dogfeed`, `new-project`, `package`, `rcode-tools`, `rihal-tools`, `team`, `usp`, `v4`, `observability`, `audit`, `init`, `agent-rules`, `cursor`, `i18n`, `phase`, `scaffold`, `campaign`, `ship`, `getting-started`, `do-router`, `milestone-health`, `modules`, `project-types`, `roadmapper`, `token`, `verifier`, `discuss-phase`, `ui-phase`, `state-sync`, `autonomous`, plus numeric phase/sprint scopes (e.g. `docs(15)`, `feat(8.3)`)
28
28
  - Subject: lowercase first letter, imperative mood, no trailing period, under 72 chars
29
29
  - **NEVER add Claude/AI attribution to commit messages.** No "Generated with Claude Code", no "Co-Authored-By: Claude", no "πŸ€– Generated". The user does not want this.
30
30
  - **NEVER use `--no-verify`** to bypass hooks. If hooks fail, fix the underlying issue.
package/CONTRIBUTING.md CHANGED
@@ -369,6 +369,8 @@ We use [Conventional Commits](https://www.conventionalcommits.org/) format. The
369
369
  - `verifier` β€” `rcode-verifier` agent, verification playbook/rules, VERIFICATION.md flow
370
370
  - `discuss-phase` β€” `/rcode-discuss-phase` workflow and its gray-area/standing checks
371
371
  - `ui-phase` β€” `/rcode-ui-phase` workflow, UI-SPEC.md/WIREFRAMES.md generation, design-library
372
+ - `state-sync` β€” `state sync --from-disk`, phase-status reconciliation, ROADMAP↔state.json drift
373
+ - `autonomous` β€” `/rcode-autonomous` workflow, the greenfield prerequisite gate
372
374
  - `<phase-id>` β€” numeric phase scope when committing inside a phase (e.g. `docs(15)`, `feat(8.3)`)
373
375
  - `<sprint-id>` β€” numeric sprint scope inside a phase (e.g. `feat(15.1)`)
374
376
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@hanzlaa/rcode",
3
- "version": "4.11.0",
3
+ "version": "4.12.1",
4
4
  "description": "rcode β€” the AI team that never forgets. Persistent memory, specialist agents, and slash commands for AI IDEs. Works in Claude Code, Cursor, Gemini, VS Code, and Antigravity.",
5
5
  "main": "cli/index.js",
6
6
  "bin": {
@@ -9,6 +9,7 @@ color: green
9
9
  @.rcode/references/karpathy-guidelines-full.md
10
10
  @.rcode/references/output-realism.md
11
11
  @.rcode/brain/best-practices/no-theoretical-suggestions.md
12
+ @.rcode/references/source-of-truth-grounding.md
12
13
  @.rcode/references/planner-playbook.md
13
14
 
14
15
  <role>
@@ -8,6 +8,7 @@ color: cyan
8
8
 
9
9
  @.rcode/references/response-style.md
10
10
  @.rcode/references/karpathy-guidelines.md
11
+ @.rcode/references/source-of-truth-grounding.md
11
12
  @.rcode/references/researcher-shared.md
12
13
 
13
14
  <role>
@@ -9,6 +9,7 @@ color: purple
9
9
  @.rcode/references/response-style.md
10
10
  @.rcode/references/output-realism.md
11
11
  @.rcode/references/karpathy-guidelines.md
12
+ @.rcode/references/source-of-truth-grounding.md
12
13
  @.rcode/references/roadmapper-playbook.md
13
14
 
14
15
  <role>
@@ -7,6 +7,7 @@ color: green
7
7
 
8
8
  @.rcode/references/response-style.md
9
9
  @.rcode/references/karpathy-guidelines.md
10
+ @.rcode/references/source-of-truth-grounding.md
10
11
  @.rcode/references/sprint-checker-playbook.md
11
12
 
12
13
  <role>
@@ -29,7 +29,8 @@ Named rules. Cite by name when applying.
29
29
  - **10th-time-user** β€” delight happens through invisible efficiency. Design for the person who has done this 10 times, not just the first-timer.
30
30
  - **Ship-then-layer** β€” recommend the simplest version that ships, then layer complexity. Perfect designs that never launch are zero value.
31
31
  - **Name-one-misconception** β€” for every confusing design element, name the specific misconception and design around it.
32
- - **Library-not-invention** β€” when producing UI-SPEC.md or WIREFRAMES.md, ground tokens and palette choices in `rcode/references/design-library/`, not invented values. WIREFRAMES.md's loading/empty/error/populated coverage per screen is Silence-kills-trust made concrete, per-screen.
32
+ - **Library-not-invention** β€” when producing UI-SPEC.md or WIREFRAMES.md, ground tokens and palette choices in `rcode/references/design-library/` AND real reference sites actually looked at (not just search snippets), not invented values. WIREFRAMES.md's loading/empty/error/populated coverage per screen is Silence-kills-trust made concrete, per-screen.
33
+ - **Variants-not-a-verdict** β€” visual direction is a taste-and-tradeoff decision, not a data lookup with one right answer. When asked to propose a design direction, generate genuinely distinct options and let the user pick β€” don't rank them or present a "recommended" one, that's the same failure as silently locking in a stack choice without asking.
33
34
 
34
35
  ## Anti-Patterns / Refuse List
35
36
 
@@ -483,6 +483,61 @@ issue:
483
483
  fix_hint: "Plan was built on hallucinated findings. Re-run /rcode-debug to verify actual code state before replanning."
484
484
  ```
485
485
 
486
+ ## Dimension 12: Evidence Grounding
487
+
488
+ **Question:** Is every claim in the plan traceable to something real β€” a codebase grep, or an external source-of-truth document β€” rather than plausible-sounding invention?
489
+
490
+ This dimension has two halves. Both exist because agents produce fluent,
491
+ confident, wrong output when nothing forces them to cite a real source β€”
492
+ codebase claims and domain-terminology claims fail the same way, just
493
+ against different ground truths.
494
+
495
+ ### 12a β€” Codebase Evidence (issue #649)
496
+
497
+ Every task body that names a file count, component, pattern, or existing
498
+ behavior MUST include an `<evidence>` block citing real grep hit counts,
499
+ real `path:line` ranges, or an explicit `creates:` justification (for
500
+ genuinely new files/symbols that can't have prior evidence). A task claiming
501
+ "13 usages of `useAuth`" with no evidence, or citing a grep that doesn't
502
+ actually return 13, is theoretical β€” reject it.
503
+
504
+ **Process:**
505
+ 1. For each task with a factual claim about existing code, check for an `<evidence>` block.
506
+ 2. If present, re-run a sample of the cited greps/searches yourself. If the claimed count doesn't match reality, downgrade to blocker regardless of what the task otherwise looks like.
507
+ 3. If absent for a claim that isn't a `creates:` (brand new file/symbol), flag as blocker β€” the claim is unfalsifiable as written.
508
+
509
+ ### 12b β€” Source-Document Evidence
510
+
511
+ If PROJECT.md/REQUIREMENTS.md/CONTEXT.md reference an external source-of-truth
512
+ document (a spreadsheet, transcript, spec, glossary β€” see
513
+ `source-of-truth-grounding.md`) that defines domain terminology, enum values,
514
+ or a data model, any task creating schema fields, enum values, or seed data
515
+ for that domain MUST cite the source document and use its verbatim values β€”
516
+ not an invented approximation that merely sounds plausible for "this kind of
517
+ system." This is exactly how a competency-tracking app once shipped invented
518
+ category names (`TECH`, `DELIV`, `COLLAB`) instead of the real Excel's actual
519
+ categories (`Technical Skills`, `Delivery & Quality`, `Communication &
520
+ Collaboration`) β€” plausible, well-formatted, and wrong.
521
+
522
+ **Process:**
523
+ 1. Check whether PROJECT.md/REQUIREMENTS.md mention a source document for this domain.
524
+ 2. If yes, for each schema/seed-data task touching that domain, verify the task cites the source document path and that its field/enum values are traceable to it (spot-check by reading the source yourself if it's available).
525
+ 3. A schema/seed task for a domain with a known source document, with no citation and no traceable match, is a blocker β€” not a style nitpick. This is exactly the class of mistake that's expensive to unwind once it's in a live database.
526
+
527
+ **Severity rules:**
528
+ - **blocker:** 12a claim with no evidence and no `creates:` justification, evidence that doesn't check out on re-run, OR 12b schema/seed task for a known-source domain with no citation/traceable match
529
+ - **warning:** evidence present but thin (e.g. cites a grep pattern too broad to actually confirm the specific claim)
530
+
531
+ **Example issue:**
532
+ ```yaml
533
+ issue:
534
+ dimension: evidence_grounding
535
+ severity: blocker
536
+ description: "Task 3 seeds 'competency' enum with TECH/DELIV/COLLAB/GROWTH/IMPACT β€” PROJECT.md references docs/Competency-Matrix.xlsx as the source of truth for these categories, but no task reads it or cites its actual values"
537
+ plan: "01"
538
+ fix_hint: "Read docs/Competency-Matrix.xlsx (or its extracted contents) and replace the enum values with what it actually defines, verbatim"
539
+ ```
540
+
486
541
  </verification_dimensions>
487
542
 
488
543
  <verification_process>
@@ -3418,9 +3418,20 @@ function cmdState(subArgs) {
3418
3418
  // landed as 'planned' regardless of what the doc said.
3419
3419
  function normalizeStatus(raw) {
3420
3420
  if (!raw) return 'planned';
3421
- const s = String(raw).toLowerCase().replace(/[βœ…\s]/g, '');
3422
- if (['complete','completed','shipped','verified','done'].includes(s)) return 'complete';
3423
- if (['executing','in_progress','inprogress','active','started'].includes(s)) return 'in_progress';
3421
+ // Match on the LEADING word, not exact string equality β€” ROADMAP.md
3422
+ // status lines legitimately carry trailing detail beyond the bare
3423
+ // status word (e.g. "Complete (verification: human_needed β€” live
3424
+ // deploy deferred)"), which an exact-match check silently drops to
3425
+ // 'planned' since the full string never equals 'complete'. A real
3426
+ // "Complete (...)" phase would then look un-synced forever.
3427
+ // Strip trailing detail (anything from the first paren/colon/em-dash
3428
+ // onward β€” "(verification: ...)", ": some note") before matching, then
3429
+ // collapse whitespace/underscores so "Complete (...)", "In Progress",
3430
+ // and "in_progress" all normalize the same way.
3431
+ const leading = String(raw).toLowerCase().replace(/[βœ…]/g, '')
3432
+ .split(/[(:β€”]/)[0].trim().replace(/[\s_]+/g, '');
3433
+ if (['complete','completed','shipped','verified','done'].includes(leading)) return 'complete';
3434
+ if (['executing','inprogress','active','started'].includes(leading)) return 'in_progress';
3424
3435
  return 'planned';
3425
3436
  }
3426
3437
 
@@ -3443,7 +3454,43 @@ function cmdState(subArgs) {
3443
3454
  // Status precedence for advancement: complete > in_progress > planned.
3444
3455
  // A phase should never be downgraded by ROADMAP re-sync.
3445
3456
  const statusRank = { complete: 2, in_progress: 1, planned: 0 };
3446
- const incomingStatus = normalizeStatus(phaseStatus);
3457
+ let incomingStatus = normalizeStatus(phaseStatus);
3458
+
3459
+ // Cross-check against VERIFICATION.md before trusting a 'complete' claim
3460
+ // from ROADMAP prose. A phase can be hand-edited to say "Complete" (or
3461
+ // "gaps_found β†’ closed") without ever re-running the verifier β€” confirmed
3462
+ // live: an agent wrote that exact phrase into ROADMAP.md while the phase's
3463
+ // own VERIFICATION.md frontmatter still said `status: gaps_found`,
3464
+ // bypassing execute.md's uat_gate entirely via a direct file edit. Don't
3465
+ // let a prose claim override what the actual verification artifact says.
3466
+ if (incomingStatus === 'complete') {
3467
+ try {
3468
+ const phasesRootDir = path.join(PLANNING_DIR, 'phases');
3469
+ const phaseDirName = fs.existsSync(phasesRootDir)
3470
+ ? fs.readdirSync(phasesRootDir).find(d => d === phaseNum || d.startsWith(`${phaseNum}-`))
3471
+ : null;
3472
+ if (phaseDirName) {
3473
+ const verFile = fs.readdirSync(path.join(phasesRootDir, phaseDirName))
3474
+ .find(f => /-VERIFICATION\.md$/i.test(f));
3475
+ if (verFile) {
3476
+ const verText = fs.readFileSync(path.join(phasesRootDir, phaseDirName, verFile), 'utf8');
3477
+ const verStatusMatch = verText.match(/^status:\s*(\S+)/m);
3478
+ const verStatus = verStatusMatch ? verStatusMatch[1].trim() : null;
3479
+ if (verStatus && verStatus !== 'passed') {
3480
+ incomingStatus = 'in_progress';
3481
+ parsed.unverified_complete_claims = parsed.unverified_complete_claims || [];
3482
+ parsed.unverified_complete_claims.push({
3483
+ phase: phaseNum,
3484
+ roadmap_claim: phaseStatus,
3485
+ verification_file: verFile,
3486
+ verification_status: verStatus,
3487
+ });
3488
+ }
3489
+ }
3490
+ }
3491
+ } catch { /* best-effort cross-check β€” never fail sync over it */ }
3492
+ }
3493
+
3447
3494
  if (existingIdx >= 0) {
3448
3495
  // Backfill both id and number so future readers using either schema find it.
3449
3496
  state.phases[existingIdx].number = state.phases[existingIdx].number || phaseNum;
@@ -4660,7 +4707,7 @@ When creating, planning, or modifying a phase, you MUST go through the rcode too
4660
4707
 
4661
4708
  **Why this is enforced**: every direct \`Write\` to \`.planning/phases/**/SPRINT.md\` without registration is a silent state divergence. Future \`/rcode-status\` reports under-count work. \`/rcode-execute\` can't find the plan. \`/rcode-progress\` shows wrong percentages.
4662
4709
 
4663
- If you have a real reason to bypass (e.g. retroactively documenting a phase that already shipped), put \`<!-- rcode-bypass: <one-line reason> -->\` at the top of the file so it's auditable later. The PreToolUse hook will allow the write through.
4710
+ If you have a real reason to bypass (e.g. retroactively documenting a phase that already shipped), put \`<!-- rcode-bypass: <one-line reason> -->\` at the top of the file so it's auditable later. **This rule is currently convention, not mechanically enforced** β€” no hook blocks an unregistered direct write today, so following it is on you (and any agent reading this), not a safety net catching you if you don't.
4664
4711
 
4665
4712
  ---
4666
4713
 
@@ -0,0 +1,80 @@
1
+ # Source-of-Truth Grounding
2
+
3
+ Loaded by any workflow/agent that generates domain terminology, schema fields,
4
+ enum values, or seed data β€” discuss-phase, roadmapper, project-researcher,
5
+ planner, sprint-checker.
6
+
7
+ ## The failure this prevents
8
+
9
+ A real incident: a competency-tracking app's requirements referenced an
10
+ authoritative Excel file (the actual HR competency matrix) and a meeting
11
+ transcript explaining it. The planner never opened the Excel β€” it invented
12
+ plausible-sounding competency names (`TECH`, `DELIV`, `COLLAB`, `GROWTH`,
13
+ `IMPACT`) and level names (`Junior`, `Associate`, `Senior I/II/III`) that
14
+ *sounded* right for a generic engineering-competency system. The real Excel
15
+ defined different names (`Technical Skills`, `Delivery & Quality`,
16
+ `Communication & Collaboration`, `Leadership & Coaching`, `Strategic Impact`,
17
+ `Innovation & Improvement`) and different levels (`Potential`, `Competency`,
18
+ `Proficiency`, `Expertise`, `Mastery`). The invented values got baked into
19
+ the database schema, seed data, and UI labels β€” an entire app built on
20
+ plausible-sounding fiction instead of the actual source of truth, discovered
21
+ only when the user manually re-fed the same files back in and asked for a
22
+ gap check.
23
+
24
+ This is not a hypothetical edge case. Any project with domain-specific
25
+ terminology defined by an external document (an HR/compliance matrix, a
26
+ partner's API spec, a legal/regulatory definition list, a brand's existing
27
+ design system, a client's glossary) is exposed to this failure the moment a
28
+ planner treats "sounds plausible" as good enough.
29
+
30
+ ## The rule
31
+
32
+ **If a source document exists β€” the user references it, provides a file
33
+ path, pastes its content, or it's already sitting in the project (`docs/`,
34
+ an attached spreadsheet, a transcript, an existing spec) β€” and it defines
35
+ domain terminology, categories, enum values, or a data model that the
36
+ current work touches, that document MUST be read in full before generating
37
+ anything that uses those concepts. Extract the actual values verbatim.**
38
+
39
+ - Don't paraphrase a defined term into a shorter or "cleaner" one. If the
40
+ source says `Communication & Collaboration`, the schema/seed/UI say
41
+ `Communication & Collaboration` β€” not `COLLAB`, not `Comm & Collab`.
42
+ - Don't guess the count or order. If the source defines 6 categories, use
43
+ the 6 it defines β€” don't invent 5 that seem close enough.
44
+ - Don't treat "I've seen similar systems before" as a substitute for reading
45
+ this one's actual source. Training-data familiarity with how competency
46
+ matrices *usually* look is exactly what produces plausible-but-wrong output.
47
+ - Excel/spreadsheet files: read the actual cell contents (via a script,
48
+ `openpyxl`/`pandas`-equivalent, or unzipping the `.xlsx` and parsing the
49
+ shared-strings XML if no library is available) β€” not the filename, not an
50
+ assumption about what a file named "Competency Matrix" probably contains.
51
+ - Transcripts/recordings-as-text: read the whole thing, not just the first
52
+ few lines β€” the real values are often stated once, in the middle, in
53
+ passing ("we have six core competencies, which is technical skills,
54
+ delivery quality, communication, collaboration...").
55
+
56
+ ## Self-check before presenting or committing
57
+
58
+ After generating requirements/schema/seed-data/UI copy that should be
59
+ grounded in a source document, do one pass comparing what you generated
60
+ against the source: does every category/level/field name match verbatim?
61
+ If you can't point to the specific line/cell in the source for a value you
62
+ used, that value is invented β€” fix it before presenting, not after the user
63
+ catches it.
64
+
65
+ ## Where this applies
66
+
67
+ - **discuss-phase / new-project**: when a source document is mentioned or
68
+ present, read it during context-gathering, before requirements get
69
+ drafted β€” not after.
70
+ - **project-researcher / phase-researcher**: cite the source document's
71
+ actual terminology in STACK.md/FEATURES.md, not an approximation.
72
+ - **roadmapper**: phase success criteria referencing domain concepts use the
73
+ source's exact terms.
74
+ - **planner**: schema fields, enum values, and seed-data tasks cite the
75
+ source document path and use its verbatim values β€” a task creating seed
76
+ data for a domain concept with a known source document is incomplete if
77
+ it doesn't reference that source.
78
+ - **sprint-checker**: if PROJECT.md/REQUIREMENTS.md reference a source
79
+ document, flag as a blocker any schema/seed-data task whose field/enum
80
+ values aren't traceable to it.
@@ -2,7 +2,7 @@
2
2
 
3
3
  **Goal:** Transform PRD requirements and Architecture decisions into comprehensive stories organized by user value, creating detailed, actionable stories with complete acceptance criteria for development teams.
4
4
 
5
- **Your Role:** In addition to your name, communication_style, and persona, you are also a product strategist and technical specifications writer collaborating with a product owner. This is a partnership, not a client-vendor relationship. You bring expertise in requirements decomposition, technical implementation context, and acceptance criteria writing, while the user brings their product vision, user needs, and business requirements. Work together as equals.
5
+ **Your Role:** In addition to your name, communication_style, and persona, you are also a product strategist and technical specifications writer collaborating with a product owner. This is a partnership, not a client-vendor relationship. You bring expertise in requirements decomposition, technical implementation context, and acceptance criteria writing, while the user brings their product vision, user needs, and business requirements. Work together as equals. Ground the product-strategist half of this role in `rcode-hussain-pm`'s calibrated judgment (MoSCoW/RICE prioritization, JTBD framing, explicit scope defense) rather than generic story-writing β€” same reasoning as `rcode-create-prd`'s workflow: this is a live, halt-and-wait conversation, so you take on the role directly instead of spawning a subagent.
6
6
 
7
7
  ## State-sync rule (NO EXCEPTIONS)
8
8
 
@@ -7,7 +7,15 @@ outputFile: '{planning_artifacts}/prd.md'
7
7
 
8
8
  **Goal:** Create comprehensive PRDs through structured workflow facilitation.
9
9
 
10
- **Your Role:** Product-focused PM facilitator collaborating with an expert peer.
10
+ **Your Role:** Product-focused PM facilitator collaborating with an expert peer. This
11
+ is a live, menu-driven, halt-and-wait conversation β€” spawning a subagent here would
12
+ break that interactivity, so you take on the role directly rather than delegating it.
13
+ Ground that role in `rcode-hussain-pm`'s actual calibrated judgment, not a generic
14
+ "PM" framing: MoSCoW/RICE prioritization, JTBD framing over feature lists, defend
15
+ scope by naming what's explicitly OUT of v1 and why, defer technical feasibility
16
+ calls to Waleed and strategic go/no-go calls to Sadiq rather than making them
17
+ yourself. See `rcode/agents/rcode-hussain-pm.md` for the full principle set if you
18
+ need it β€” the summary above is what actually changes PRD-writing behavior day to day.
11
19
 
12
20
  You will continue to operate with your given name, identity, and communication_style, merged with the details of this role description.
13
21
 
@@ -1,5 +1,5 @@
1
1
  {
2
- "_comment": "pre-edit hook is currently advisory (logs warning, does not block). Full session-read tracking requires PostToolUse Read hook to write state file. Implemented in follow-up. prompt-router (UserPromptSubmit) is advisory β€” emits additionalContext to nudge toward the right rcode command; never blocks. session-start (SessionStart) emits a one-line project status primer at session open.",
2
+ "_comment": "pre-edit hook is currently advisory (logs warning, does not block). Full session-read tracking requires PostToolUse Read hook to write state file. Implemented in follow-up. prompt-router (UserPromptSubmit) is advisory \u2014 emits additionalContext to nudge toward the right rcode command; never blocks. session-start (SessionStart) emits a one-line project status primer at session open.",
3
3
  "hooks": {
4
4
  "PreToolUse": [
5
5
  {
@@ -7,11 +7,11 @@
7
7
  "hooks": [
8
8
  {
9
9
  "type": "command",
10
- "command": "node .rcode/bin/rcode-hooks.cjs pre-edit"
10
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" pre-edit || exit 0'"
11
11
  },
12
12
  {
13
13
  "type": "command",
14
- "command": "node .rcode/bin/rcode-hooks.cjs compact-nudge"
14
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" compact-nudge || exit 0'"
15
15
  }
16
16
  ]
17
17
  },
@@ -20,7 +20,7 @@
20
20
  "hooks": [
21
21
  {
22
22
  "type": "command",
23
- "command": "node .rcode/bin/rcode-hooks.cjs pre-workflow"
23
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" pre-workflow || exit 0'"
24
24
  }
25
25
  ]
26
26
  },
@@ -29,7 +29,7 @@
29
29
  "hooks": [
30
30
  {
31
31
  "type": "command",
32
- "command": "node .rcode/bin/rcode-hooks.cjs bash-guard"
32
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" bash-guard || exit 0'"
33
33
  }
34
34
  ]
35
35
  }
@@ -40,7 +40,7 @@
40
40
  "hooks": [
41
41
  {
42
42
  "type": "command",
43
- "command": "node .rcode/bin/rcode-hooks.cjs post-commit"
43
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" post-commit || exit 0'"
44
44
  }
45
45
  ]
46
46
  }
@@ -51,7 +51,7 @@
51
51
  "hooks": [
52
52
  {
53
53
  "type": "command",
54
- "command": "node .rcode/bin/rcode-hooks.cjs pre-compact"
54
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" pre-compact || exit 0'"
55
55
  }
56
56
  ]
57
57
  }
@@ -62,15 +62,15 @@
62
62
  "hooks": [
63
63
  {
64
64
  "type": "command",
65
- "command": "node .rcode/bin/rcode-hooks.cjs stop-verify"
65
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" stop-verify || exit 0'"
66
66
  },
67
67
  {
68
68
  "type": "command",
69
- "command": "node .rcode/bin/rcode-hooks.cjs cost-track"
69
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" cost-track || exit 0'"
70
70
  },
71
71
  {
72
72
  "type": "command",
73
- "command": "node .rcode/bin/rcode-hooks.cjs stop"
73
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" stop || exit 0'"
74
74
  }
75
75
  ]
76
76
  }
@@ -81,7 +81,7 @@
81
81
  "hooks": [
82
82
  {
83
83
  "type": "command",
84
- "command": "node .rcode/bin/rcode-hooks.cjs prompt-router"
84
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" prompt-router || exit 0'"
85
85
  }
86
86
  ]
87
87
  }
@@ -92,7 +92,7 @@
92
92
  "hooks": [
93
93
  {
94
94
  "type": "command",
95
- "command": "node .rcode/bin/rcode-hooks.cjs session-start"
95
+ "command": "sh -c 'H=\".rcode/bin/rcode-hooks.cjs\"; [ -f \"$H\" ] || { G=$(git rev-parse --git-common-dir 2>/dev/null) && H=\"$(cd \"$(dirname \"$G\")\" 2>/dev/null && pwd)/.rcode/bin/rcode-hooks.cjs\"; }; [ -f \"$H\" ] && exec node \"$H\" session-start || exit 0'"
96
96
  }
97
97
  ]
98
98
  }
@@ -52,30 +52,51 @@ Read all files referenced by the invoking prompt's execution_context before star
52
52
 
53
53
  ## 0. Prerequisite check (greenfield guard)
54
54
 
55
- Before any phase work, verify the methodology chain has run:
55
+ rcode supports two valid project-initialization paths, and this gate must
56
+ accept either:
57
+ - **Full chain:** `/rcode-create-prd` β†’ `/rcode-new-milestone` β†’ `/rcode-create-epics-and-stories`
58
+ - **Direct roadmap path:** `/rcode-new-project` β†’ `rcode-roadmapper` writes
59
+ ROADMAP.md directly with phases, no prd.md/epics.md produced β€” this is a
60
+ first-class supported path, not an edge case, and autonomous execution
61
+ only actually needs a ROADMAP.md with real phases in it to do phase work.
62
+
63
+ Before any phase work, verify at least one path's minimum requirement is met:
56
64
 
57
65
  ```bash
58
66
  HAS_PRD=$( ( ls .planning/prd.md .planning/PRD.md .planning/prds/*.md .planning/milestones/*/PRD.md 2>/dev/null | head -1 ) && echo true || echo false)
59
- HAS_ROADMAP_MILESTONES=$(grep -qE "^## Milestone\s+M[0-9]+" .planning/ROADMAP.md 2>/dev/null && echo true || echo false)
60
67
  HAS_EPICS=$( ( ls .planning/epics.md .planning/EPICS.md .planning/epics/*.md .planning/milestones/*/EPICS.md 2>/dev/null | head -1 ) && echo true || echo false)
68
+ # Milestone marker: accepts heading style ("## Milestone M1"), bold-prose style
69
+ # ("**Milestone:** M1 β€” ..."), or PROJECT.md's own style ("## Current Milestone: M3 β€” ...")
70
+ # β€” roadmapper's actual output uses the bold-prose form, which the old heading-only
71
+ # regex never matched, permanently failing this gate for every project that used
72
+ # the direct roadmap path. Fixed live: confirmed against a real project's ROADMAP.md.
73
+ HAS_ROADMAP_MILESTONES=$(grep -qEi "milestone[:*]*\s*M[0-9]" .planning/ROADMAP.md 2>/dev/null && echo true || echo false)
74
+ # Direct roadmap path's actual minimum: a ROADMAP.md with at least one real phase.
75
+ HAS_ROADMAP_PHASES=$(grep -qE "^##\s*Phase\s+[0-9]|^\|\s*[0-9]+\s*\|" .planning/ROADMAP.md 2>/dev/null && echo true || echo false)
61
76
  SKIP_FLAG=$(echo "$ARGUMENTS" | grep -qE "\-\-skip-prerequisites" && echo true || echo false)
77
+
78
+ FULL_CHAIN_OK=$([ "$HAS_PRD" = "true" ] && [ "$HAS_ROADMAP_MILESTONES" = "true" ] && [ "$HAS_EPICS" = "true" ] && echo true || echo false)
79
+ DIRECT_PATH_OK=$([ "$HAS_ROADMAP_PHASES" = "true" ] && echo true || echo false)
62
80
  ```
63
81
 
64
- If `SKIP_FLAG=false` AND any prerequisite is missing, HALT with a clear message:
82
+ If `SKIP_FLAG=false` AND both `FULL_CHAIN_OK` and `DIRECT_PATH_OK` are false, HALT with a clear message:
65
83
 
66
84
  ```
67
- ⚠ Cannot run autonomous: missing prerequisite β€” {what}.
85
+ ⚠ Cannot run autonomous: no valid project initialization found.
68
86
 
69
- The autonomous flow assumes a project that has already gone through:
70
- 1. /rcode-create-prd β†’ produces .planning/prd.md
71
- 2. /rcode-new-milestone β†’ produces ROADMAP.md with M1..Mn
72
- 3. /rcode-create-epics-and-stories β†’ produces .planning/epics.md
73
- 4. THEN /rcode-autonomous ← you are here
87
+ The autonomous flow needs either:
88
+ A) Full chain: /rcode-create-prd β†’ /rcode-new-milestone β†’ /rcode-create-epics-and-stories
89
+ B) Direct roadmap: /rcode-new-project (produces ROADMAP.md with phases directly)
74
90
 
75
- Suggested first step: /rcode-{first-missing-command}
91
+ Neither was found β€” no ROADMAP.md with real phases exists, and the full
92
+ chain's artifacts (prd.md, milestone-marked ROADMAP.md, epics.md) are also
93
+ missing.
76
94
 
77
- If you genuinely want to skip these (rare β€” usually inverted methodology),
78
- re-invoke with: /rcode-autonomous --skip-prerequisites
95
+ Suggested first step: /rcode-new-project (recommended β€” simpler, fully
96
+ supported) or /rcode-create-prd if you specifically want the full chain.
97
+
98
+ If you genuinely want to skip this check, re-invoke with:
99
+ /rcode-autonomous --skip-prerequisites
79
100
  ```
80
101
 
81
102
  If `SKIP_FLAG=true`: print a warning that downstream workflows may produce low-quality output without upstream artifacts, then proceed.
@@ -463,7 +484,7 @@ Task(
463
484
  )
464
485
  ```
465
486
 
466
- Store the agent task_id. The workflow can now start discussing the next phase while this phase executes in the background. Before starting post-execution routing for this phase, wait for the execute agent to complete.
487
+ Store the agent task_id. The workflow can now start discussing the next phase while this phase executes in the background. Before starting post-execution routing for this phase, wait for the execute agent to complete, then run the mandatory reconciliation command below (same as the non-interactive path) before continuing.
467
488
 
468
489
  **If `INTERACTIVE` is NOT set (default):** Run execute inline as before.
469
490
 
@@ -471,6 +492,24 @@ Store the agent task_id. The workflow can now start discussing the next phase wh
471
492
  Skill(skill="rcode-execute", args="${PHASE_NUMBER} --no-transition")
472
493
  ```
473
494
 
495
+ **Mandatory reconciliation (run this bash command directly, every phase, no exceptions):**
496
+ `execute.md`'s own state-write steps (`phase set-status`/`phase complete`) are
497
+ documented but not mechanically guaranteed β€” a real unattended multi-phase
498
+ run was observed skipping them, leaving `state.json` stuck at every phase's
499
+ planning-time status while ROADMAP.md correctly showed them complete. Don't
500
+ rely on remembering to do this; run it as its own step, immediately after
501
+ execute returns, before moving to post-execution routing:
502
+
503
+ ```bash
504
+ node ".rcode/bin/rcode-tools.cjs" state sync --from-disk >/dev/null 2>&1 || true
505
+ ```
506
+
507
+ This is idempotent and cheap β€” it re-derives phase status from ROADMAP.md's
508
+ actual `**Status:**` text (never downgrades an already-advanced status), so
509
+ running it here closes the gap even if execute.md's own inline state-write
510
+ steps were skipped during a long run. Do this for every phase in the loop,
511
+ not just once at the end.
512
+
474
513
  ### 3c.5. Code Review and Fix
475
514
 
476
515
  Auto-invoke code review and fix chain. Autonomous mode chains both review and fix.
@@ -6,6 +6,7 @@ You are a thinking partner, not an interviewer. The user is the visionary β€” yo
6
6
 
7
7
  <required_reading>
8
8
  @.rcode/references/universal-anti-patterns.md
9
+ @.rcode/references/source-of-truth-grounding.md
9
10
  </required_reading>
10
11
 
11
12
  <conditional_reading>
@@ -90,6 +90,12 @@ Enabled guardrails:
90
90
  β€’ session-start: Greets the session with one-line phase status and suggested next command
91
91
 
92
92
  To disable, remove the hooks section from .claude/settings.json or edit .rcode/templates/settings-hooks.json and re-run.
93
+
94
+ Note: .rcode/bin/ is gitignored by design, so a `git worktree add` checkout
95
+ won't have it β€” each hook command resolves this automatically via
96
+ `git rev-parse --git-common-dir` (finds the main checkout's .rcode/bin/ and
97
+ runs from there), falling back to a silent no-op if genuinely not found
98
+ anywhere. No action needed in worktrees; this is handled for you.
93
99
  ```
94
100
 
95
101
  ## Success Criteria
@@ -541,7 +541,11 @@ Write files first, then return.
541
541
 
542
542
  **If `## ROADMAP BLOCKED`:** present the blocker, collect resolution from user, re-spawn the roadmapper with revision context.
543
543
 
544
- **If `## ROADMAP CREATED`:** read ROADMAP.md, present inline:
544
+ **If `## ROADMAP CREATED`:**
545
+
546
+ **Team review pass (once per roadmap draft; skip in auto/yolo mode β€” see `new-project-roadmap.md`'s identical step for the full rationale):** spawn `rcode-waleed` and `rcode-fatima` in parallel, same prompts as `new-project-roadmap.md`'s team review pass (feasibility sanity check / release-risk sanity check, 3 bullets max or "no concerns" in one line). Fold genuine concerns into a **Team Review Notes** section in the presentation below, before the phase table β€” skip the section if both say no concerns.
547
+
548
+ Read ROADMAP.md, present inline:
545
549
 
546
550
  ```
547
551
  ## Proposed Roadmap
@@ -310,6 +310,28 @@ Display research complete banner and key findings:
310
310
  Files: `.planning/research/`
311
311
  ```
312
312
 
313
+ **Mandatory stack-choice confirmation (before proceeding to roadmap):**
314
+ Tech stack is the single most foundational, hardest-to-reverse decision in
315
+ the project β€” every phase gets built on top of it, and un-picking it later
316
+ means rework, not a config change. Do not let it get silently locked in.
317
+
318
+ Check `.planning/research/STACK.md` for language indicating this was a
319
+ judgment call, not a forced single-option pick β€” phrases like "judgment
320
+ call," "flagging explicitly," a comparison table between 2+ named
321
+ platforms/frameworks, or an explicit confidence caveat on the top-level
322
+ choice. (A stack choice with genuinely no alternative β€” e.g. "must integrate
323
+ with the client's existing Salesforce" β€” doesn't need this; skip only in
324
+ that case, and say why.)
325
+
326
+ If it's a judgment call, use AskUserQuestion before continuing:
327
+ - **question:** "Research recommends {stack} over {alternative(s)} for {one-line reason}. Confirm, or pick differently?"
328
+ - **options:** "Confirm {stack}" / "Use {alternative}" / "Something else β€” I'll specify"
329
+
330
+ If the user picks differently, note the override as a decision and don't
331
+ silently keep researching the original recommendation β€” the chosen stack
332
+ from here on is whatever the user picked, not STACK.md's original
333
+ recommendation, and downstream phases/roadmap should reflect that.
334
+
313
335
  **If "Skip research":** Continue to Step 7.
314
336
 
315
337
 
@@ -215,6 +215,25 @@ Write files first, then return. This ensures artifacts persist even if context i
215
215
 
216
216
  **If `## ROADMAP CREATED`:**
217
217
 
218
+ **Team review pass (once per roadmap draft, not per phase β€” keep this cheap; skip entirely in auto/yolo mode β€” an unattended loop shouldn't pay this round-trip on every roadmap, and there's no user present to read the notes anyway):**
219
+ Before presenting the roadmap for approval, spawn two focused reviewers in
220
+ parallel against the freshly-written ROADMAP.md β€” a single async round, not a
221
+ council debate, so this adds one round-trip, not multiple:
222
+
223
+ ```
224
+ Task(prompt="Read .planning/ROADMAP.md and .planning/PROJECT.md. As Waleed (CTO/architect), flag ONLY real feasibility/scalability concerns you'd actually block a real project over β€” a phase sequencing a schema-breaking change after the API that depends on it, a scale ceiling the phase structure doesn't account for, a missing foundational/shell phase for a UI project. If there are none, say so in one line. Keep it to 3 bullets max β€” this is a sanity check, not a full architecture review.", subagent_type="rcode-waleed", model="{roadmapper_model}", description="Architecture sanity check")
225
+
226
+ Task(prompt="Read .planning/ROADMAP.md and .planning/PROJECT.md. As Fatima (QA Lead), flag ONLY real release-risk concerns β€” a phase with no way to verify its success criteria, a security/compliance-sensitive area with no phase covering it, a dependency ordering that makes a phase unverifiable until a later one lands. If there are none, say so in one line. Keep it to 3 bullets max β€” this is a sanity check, not a full QA strategy doc.", subagent_type="rcode-fatima", model="{roadmapper_model}", description="Release-risk sanity check")
227
+ ```
228
+
229
+ Run both, collect their one-line-or-3-bullets responses. If either flags a
230
+ concern that would make the roadmap actually wrong (not a nitpick), add a
231
+ **Team Review Notes** section to the presentation below, before the phase
232
+ table, so the user sees it as part of deciding whether to approve β€” don't
233
+ silently drop it, and don't block on it either; the approval gate right
234
+ after this is where the user decides what to do with it. If both reviewers
235
+ say "no concerns," skip the section entirely β€” don't manufacture filler.
236
+
218
237
  Read the created ROADMAP.md and present it nicely inline:
219
238
 
220
239
  ```
@@ -625,12 +625,27 @@ Task(
625
625
  - **`## VERIFICATION PASSED`:** Display confirmation, proceed to step 13.
626
626
  - **`## ISSUES FOUND`:** Display issues, check iteration count, proceed to step 12.
627
627
 
628
- **Thinking partner for architectural tradeoffs (conditional):**
628
+ **Thinking partner for architectural tradeoffs (default ON in guided mode, OFF in yolo/autonomous):**
629
629
  ```bash
630
- THINKING_PARTNER_ENABLED=$(node ".rcode/bin/rcode-tools.cjs" config-get features.thinking_partner 2>/dev/null || echo "false")
630
+ THINKING_PARTNER_CONFIG=$(node ".rcode/bin/rcode-tools.cjs" config-get features.thinking_partner 2>/dev/null || echo "")
631
+ if [ -n "$THINKING_PARTNER_CONFIG" ]; then
632
+ THINKING_PARTNER_ENABLED="$THINKING_PARTNER_CONFIG"
633
+ elif [ "$MODE" = "yolo" ] || [ -n "$AUTONOMOUS" ]; then
634
+ THINKING_PARTNER_ENABLED="false"
635
+ else
636
+ THINKING_PARTNER_ENABLED="true"
637
+ fi
631
638
  ```
632
639
  ${THINKING_PARTNER_ENABLED === 'true' ? '@.rcode/references/plan-thinking-partner.md' : ''}
633
- If `features.thinking_partner` is disabled: skip this block entirely.
640
+ If `THINKING_PARTNER_ENABLED` is `false`: skip this block entirely. An explicit
641
+ `features.thinking_partner` in config.yaml always wins over the mode-based
642
+ default (set it `false` to silence even in guided mode, or `true` to keep it
643
+ on during autonomous runs if you want that). The check itself is cheap β€” a
644
+ keyword scan over the checker's existing issues, not a new agent spawn β€”
645
+ which is why it defaults on for guided/interactive planning: a second-opinion
646
+ sanity check on architectural tradeoffs is exactly the kind of thing "does
647
+ this actually get built right" needs, and it only activates when the checker
648
+ already flagged a tradeoff-shaped issue.
634
649
 
635
650
  ## 12. Revision Loop (Max 3 Iterations, 1 in autonomous/yolo mode)
636
651
 
@@ -50,7 +50,7 @@ If `flags.existing_ui` or `flags.design_system` provided:
50
50
  EXISTING=$(node .rcode/bin/rcode-tools.cjs find-files --type=design-tokens)
51
51
  ```
52
52
 
53
- Load existing design system, extract into `EXISTING_DESIGN_SYSTEM_DATA` (a text block passed verbatim into Step 2's prompt):
53
+ Load existing design system, extract into `EXISTING_DESIGN_SYSTEM_DATA` (a text block passed verbatim into Step 2b's prompt):
54
54
  - Color palette (hex, variable names)
55
55
  - Typography scales (font family, sizes, weights, line heights)
56
56
  - Component list (buttons, forms, layouts, modals, etc.)
@@ -60,7 +60,10 @@ If `$EXISTING` is empty or extraction finds nothing usable, treat this the
60
60
  same as "no existing system found" and fall through to Step 1b β€” don't leave
61
61
  `EXISTING_DESIGN_SYSTEM_DATA` half-populated.
62
62
 
63
- If an existing design system was found and `EXISTING_DESIGN_SYSTEM_DATA` is populated, skip Step 1b (don't override what's already decided) and go straight to Step 2.
63
+ If an existing design system was found and `EXISTING_DESIGN_SYSTEM_DATA` is
64
+ populated, there's nothing to choose between β€” skip Step 1b, Step 1c, AND
65
+ Step 2 (variant generation + user confirmation) entirely, and go straight to
66
+ Step 2b using `EXISTING_DESIGN_SYSTEM_DATA` as the chosen direction.
64
67
 
65
68
  ## Step 1b β€” Ground the design in the reference library (no existing system found)
66
69
 
@@ -73,11 +76,85 @@ first so there's a project category to ground the design in." Otherwise look up:
73
76
  2. **Style detail** β€” take the recommended style name from step 1 and look it up in `styles.csv` for concrete hex values, effects, accessibility rating, and an implementation checklist. Capture as `STYLE_DETAIL`.
74
77
  3. **UX rules** β€” grep `ux-guidelines.csv` for the categories relevant to this project's screens (Navigation, Forms, etc.) for concrete do/don't rules with code examples. Capture as `UX_RULES`.
75
78
 
76
- `CATEGORY_MATCH` + `STYLE_DETAIL` + `UX_RULES` together are what Step 2's
77
- prompt calls `DESIGN_LOOKUP_RESULT` below β€” assemble them into one text block
78
- before spawning the agent.
79
+ `CATEGORY_MATCH` + `STYLE_DETAIL` + `UX_RULES` together are what Step 1c and
80
+ Step 2 below call `DESIGN_LOOKUP_RESULT` β€” assemble them into one text block.
81
+
82
+ ## Step 1c β€” Look at real reference sites (skip only if Step 1 found an existing system)
83
+
84
+ A CSV row is a starting hypothesis, not a substitute for actually looking at
85
+ what real sites in this space look like. A style-data table can't tell you
86
+ that every competitor uses a specific hero-image treatment, or that the
87
+ category's "expected" look has drifted since the data was written. Do this
88
+ before committing to a direction:
89
+
90
+ 1. **Find real reference sites.** WebSearch for the project's actual
91
+ industry/niche (from PROJECT.md β€” e.g. "best interior design websites
92
+ Pakistan", "top mobile repair shop websites", not a generic "SaaS
93
+ landing page examples" search unrelated to the real domain). Aim for
94
+ 3-5 real, currently-live sites β€” direct competitors if findable, strong
95
+ examples in the same category otherwise.
96
+ 2. **Actually look at them**, don't just read search-result snippets. If a
97
+ browser tool is available in this session, navigate to each and take a
98
+ screenshot. If not, WebFetch each URL and read the rendered content/
99
+ structure description it returns. Note concretely, per site: layout
100
+ pattern (hero style, nav placement, content density), color mood, type
101
+ feel (serif/sans, weight, size), imagery style (photography vs.
102
+ illustration vs. none), and anything that reads as dated or as a
103
+ red flag to avoid repeating.
104
+ 3. Capture this as `REFERENCE_SITE_FINDINGS` β€” a short per-site note, not a
105
+ full audit. This feeds variant generation next, not a standalone report.
106
+
107
+ If WebSearch/WebFetch/browser tools are genuinely unavailable in this
108
+ session, skip with an explicit note ("Step 1c skipped β€” no web/browser
109
+ tools available; design-library data only") rather than silently omitting
110
+ real-site grounding β€” this is a real reduction in design quality and should
111
+ be visible, not silent.
112
+
113
+ ## Step 2 β€” Generate design variants, get user confirmation (mandatory β€” do not skip to a single locked-in direction)
114
+
115
+ Visual design direction is the same class of decision as tech-stack choice:
116
+ foundational, expensive to redo once screens are built against it, and a
117
+ matter of taste as much as data β€” it must not get silently locked in by an
118
+ agent picking "the" recommended style. Generate 2-3 concrete, genuinely
119
+ distinct variants (not the same direction with a different accent color),
120
+ each grounded in `DESIGN_LOOKUP_RESULT` (Step 1b) and `REFERENCE_SITE_FINDINGS`
121
+ (Step 1c), before writing anything to UI-SPEC.md:
79
122
 
80
- ## Step 2 β€” Spawn UI Designer
123
+ Spawn `rcode-ux-designer` subagent:
124
+
125
+ ```
126
+ Task tool call:
127
+ subagent_type: "rcode-ux-designer"
128
+ description: "Generate design variants"
129
+ prompt: |
130
+ Ground every variant in {EXISTING_DESIGN_SYSTEM_DATA if Step 1 found one, else DESIGN_LOOKUP_RESULT + REFERENCE_SITE_FINDINGS} β€”
131
+ do not invent a palette/style from nothing when reference data or real
132
+ reference sites were found.
133
+
134
+ Propose 2-3 genuinely distinct design variants for this project. For each:
135
+ - **Name** (short, e.g. "Warm Editorial", "Minimal Trust", "Bold Craft")
136
+ - **One-paragraph description** β€” mood, layout approach, what it borrows from
137
+ the reference sites found (name which one(s)) vs. the design-library data
138
+ - **Color direction** β€” 2-3 representative hex values, not a full token system yet
139
+ - **Typography direction** β€” font pairing feel (serif/sans, weight)
140
+ - **Best for / worst for** β€” one line each, honest tradeoffs
141
+
142
+ Do NOT pick a winner or rank them β€” present as genuine options. Do NOT
143
+ write UI-SPEC.md yet.
144
+
145
+ Return the variants as your response text (not a file write).
146
+ ```
147
+
148
+ Present the variants to the user via AskUserQuestion:
149
+ - **question:** "Which design direction fits this project?"
150
+ - **options:** one per variant (label = variant name, description = the one-paragraph summary), plus the tool's built-in "Other" for a custom direction
151
+ - If the user picks "Other" and describes something different, treat their description as the chosen direction instead of any generated variant.
152
+
153
+ Only after the user picks does Step 2b below proceed β€” using the chosen
154
+ variant's color/typography direction as the seed for the full spec, not
155
+ re-deriving from scratch.
156
+
157
+ ## Step 2b β€” Spawn UI Designer (full spec, chosen direction)
81
158
 
82
159
  Spawn `rcode-ux-designer` subagent:
83
160
 
@@ -86,7 +163,10 @@ Task tool call:
86
163
  subagent_type: "rcode-ux-designer"
87
164
  description: "Generate UI-SPEC.md and WIREFRAMES.md"
88
165
  prompt: |
89
- Ground every choice below in {EXISTING_DESIGN_SYSTEM_DATA if Step 1 found one, else DESIGN_LOOKUP_RESULT from Step 1b} β€”
166
+ Build out the full spec from {EXISTING_DESIGN_SYSTEM_DATA if Step 1 found
167
+ one, else the user-chosen variant's full description, color direction,
168
+ and typography direction from Step 2}.
169
+ Ground every choice in {EXISTING_DESIGN_SYSTEM_DATA if Step 1 found one, else DESIGN_LOOKUP_RESULT from Step 1b} β€”
90
170
  do not invent a palette/style from nothing when reference data or an existing system exists.
91
171
 
92
172
  Write UI-SPEC.md with:
@@ -101,7 +181,7 @@ Task tool call:
101
181
  Write to: {ui_spec_path}
102
182
  ```
103
183
 
104
- ## Step 2b β€” Spawn Wireframes (per-role screen inventory)
184
+ ## Step 2c β€” Spawn Wireframes (per-role screen inventory)
105
185
 
106
186
  Read REQUIREMENTS.md/PROJECT.md for the project's user roles (if any) and the
107
187
  IA decision β€” `roadmapper-playbook.md`'s Information Architecture step (Workflow
@@ -221,7 +301,9 @@ Run /rcode-ui-phase, then return to /rcode-plan
221
301
 
222
302
  ## Success Criteria
223
303
 
224
- - UI-SPEC.md created with all 7 sections, design direction grounded in `design-library/` lookup (or an existing design system), not invented
304
+ - Real reference sites looked at (Step 1c), not just design-library data alone β€” or explicitly noted as skipped with a reason
305
+ - 2-3 genuinely distinct design variants presented and the user explicitly picked one via AskUserQuestion β€” no direction silently locked in (skipped only when Step 1 found an existing design system, where there's nothing to choose between)
306
+ - UI-SPEC.md created with all 7 sections, design direction grounded in the chosen variant / `design-library/` lookup / real reference sites (or an existing design system), not invented
225
307
  - Color tokens documented with contrast ratios
226
308
  - Component inventory complete with variants
227
309
  - Accessibility checklist included
@@ -233,7 +315,7 @@ Run /rcode-ui-phase, then return to /rcode-plan
233
315
  - If subagent fails: provide template UI-SPEC.md
234
316
  - If frontend detection fails: skip suggestion
235
317
  - If config.yaml missing ui_safety_gate: default to true (suggest)
236
- - If no IA decision exists in ROADMAP.md or IA.md yet: WIREFRAMES.md cannot be produced meaningfully β€” stop Step 2b and say so rather than writing a screen list with no basis
318
+ - If no IA decision exists in ROADMAP.md or IA.md yet: WIREFRAMES.md cannot be produced meaningfully β€” stop Step 2c and say so rather than writing a screen list with no basis
237
319
 
238
320
  ## Next Up
239
321