@ulysses-ai/create-workspace 0.17.0-beta.0 → 0.18.0-beta.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/README.md +3 -3
  2. package/package.json +1 -1
  3. package/template/.claude/hooks/_utils.mjs +1 -1
  4. package/template/.claude/hooks/repo-write-detection.mjs +161 -64
  5. package/template/.claude/hooks/session-start.mjs +35 -1
  6. package/template/.claude/hooks/subagent-start.mjs +89 -22
  7. package/template/.claude/lib/session-frontmatter.mjs +28 -0
  8. package/template/.claude/rules/coherent-revisions.md +1 -1
  9. package/template/.claude/rules/forge-operations.md +37 -93
  10. package/template/.claude/rules/git-conventions.md +16 -11
  11. package/template/.claude/rules/goal-driven-work.md +8 -416
  12. package/template/.claude/rules/honest-pushback.md +37 -37
  13. package/template/.claude/rules/memory-guidance.md +43 -90
  14. package/template/.claude/rules/superpowers-workflow.md.skip +1 -1
  15. package/template/.claude/rules/work-item-tracking.md +30 -72
  16. package/template/.claude/rules/workspace-structure.md +36 -94
  17. package/template/.claude/scripts/build-workspace-context.mjs +61 -16
  18. package/template/.claude/scripts/chat-record.mjs +282 -0
  19. package/template/.claude/scripts/cleanup-work-session.mjs +257 -68
  20. package/template/.claude/scripts/context-footprint.mjs +282 -0
  21. package/template/.claude/scripts/forges/github.mjs +45 -0
  22. package/template/.claude/scripts/forges/gitlab.mjs +3 -2
  23. package/template/.claude/scripts/forges/interface.mjs +12 -0
  24. package/template/.claude/scripts/generate-claude-local.mjs +21 -2
  25. package/template/.claude/scripts/migrate-sessions.mjs +1571 -0
  26. package/template/.claude/scripts/migrate-to-workspace-context.mjs +7 -2
  27. package/template/.claude/scripts/task-worktree.mjs +525 -0
  28. package/template/.claude/scripts/workspace-diagnostics.mjs +654 -0
  29. package/template/.claude/skills/braindump/SKILL.md +11 -4
  30. package/template/.claude/skills/build-docs-site/SKILL.md +5 -5
  31. package/template/.claude/skills/build-docs-site/templates/spec.md.tmpl +1 -1
  32. package/template/.claude/skills/complete-work/SKILL.md +229 -219
  33. package/template/.claude/skills/context-placement/SKILL.md +199 -0
  34. package/template/.claude/skills/goal-driven-work/SKILL.md +459 -0
  35. package/template/.claude/skills/handoff/SKILL.md +11 -4
  36. package/template/.claude/skills/maintenance/SKILL.md +7 -0
  37. package/template/.claude/skills/migrate-sessions/SKILL.md +70 -0
  38. package/template/.claude/skills/pause-work/SKILL.md +9 -1
  39. package/template/.claude/skills/release/SKILL.md +44 -108
  40. package/template/.claude/skills/start-work/SKILL.md +89 -7
  41. package/template/.claude/skills/workspace-init/SKILL.md +3 -1
  42. package/template/.claude/skills/workspace-update/SKILL.md +4 -0
  43. package/template/CLAUDE.md.tmpl +19 -2
  44. package/template/_gitignore +9 -0
  45. package/template/workspace.json.tmpl +3 -2
@@ -0,0 +1,199 @@
1
+ ---
2
+ name: context-placement
3
+ description: Use when deciding where a durable fact, instruction, or procedure belongs — a rule, canonical context, shared or personal context, auto-memory, or a skill — and when writing it once the destination is settled. Routes the decision, states the always-loaded cost before anything is written, and carries the frontmatter schema, generator invocations, and the canonical admission test.
4
+ ---
5
+
6
+ # Context Placement
7
+
8
+ Something is worth keeping. This skill decides where it goes, and makes the cost of that
9
+ choice visible before you write it.
10
+
11
+ The hard part is routing, not prose. Most placement mistakes are not badly written files —
12
+ they are correctly written files in a destination that charges every future session for
13
+ them. Two of this workspace's own bugs (gh:136, gh:138) were exactly that: rules that had
14
+ quietly absorbed reference material and worked examples until the always-loaded directory
15
+ cost more than the conversation.
16
+
17
+ ## The destinations
18
+
19
+ Exactly one of these. The right-hand column is what every session pays, forever, for the
20
+ choice.
21
+
22
+ | destination | route here when | always-loaded cost |
23
+ |---|---|---|
24
+ | **nowhere** | it is already covered somewhere, or it is true only today | zero |
25
+ | `.claude/rules/{name}.md` with `paths:` | an instruction that applies only when specific files are touched | zero until a matching file is read |
26
+ | `.claude/skills/{name}/SKILL.md` | a procedure with steps, invoked on demand | the `description` line, ~200 B |
27
+ | auto-memory | a machine-local preference or correction; not shared, not reviewable | one `MEMORY.md` line, ~100 B |
28
+ | `workspace-context/team-member/{user}/` | one person's working context | one index line, ~120 B, that user only |
29
+ | `workspace-context/shared/` | team-visible reference to look up when relevant | one index line, ~120 B |
30
+ | `workspace-context/shared/locked/` | a team-wide fact or constraint that passes the canonical test | **the whole file, every session** |
31
+ | `.claude/rules/{name}.md` (no `paths:`) | an instruction that must shape every session, everywhere in the repo | **the whole file, every session** |
32
+
33
+ The table is ordered by cost, cheapest first. Work down it and stop at the first
34
+ destination that genuinely fits. The two bold rows are the only ones that tax every
35
+ session; reach them only after the cheaper ones have actually been ruled out, not skipped.
36
+
37
+ **`nowhere` is the most common correct answer.** Before anything else, check whether the
38
+ thing is already stated. A second copy in a second location is worse than no copy: they
39
+ drift, and the reader cannot tell which one is current.
40
+
41
+ ## Step 1 — classify what you have
42
+
43
+ Three kinds. The kind determines which destinations are even eligible.
44
+
45
+ - **Instruction** — changes what Claude *does*. "Never force push." "Rebase before opening
46
+ a PR." Imperative, applies without being asked for. Eligible: rules (scoped or not).
47
+ - **Reference** — something Claude needs to *know* when a particular topic comes up. API
48
+ shapes, schemas, architecture facts, why an approach was rejected. Eligible: canonical,
49
+ shared, team-member, auto-memory.
50
+ - **Procedure** — a sequence of steps with a beginning and an end, run on request.
51
+ Eligible: a skill.
52
+
53
+ Most drift comes from putting reference or procedure content in a rule. A rule that
54
+ contains an API table or a worked example is carrying reference material at instruction
55
+ prices. Split it: the instruction stays in the rule, the detail moves to a skill or a
56
+ context file, and the rule points at it in one line.
57
+
58
+ ## Step 2 — scope it
59
+
60
+ For an instruction, the only question that matters is *when must this be true?*
61
+
62
+ - Every session, every file → unconditional rule.
63
+ - Only when certain files are involved → rule with `paths:`.
64
+
65
+ `paths:` is a real Claude Code feature and is almost always the better answer for anything
66
+ domain-specific. A rule about migration scripts, or about a single repo's test conventions,
67
+ does not need to be in context while you are editing documentation.
68
+
69
+ ```yaml
70
+ ---
71
+ paths:
72
+ - ".claude/scripts/**/*.mjs"
73
+ - "repos/*/template/.claude/scripts/**/*.mjs"
74
+ ---
75
+ ```
76
+
77
+ Rules with `paths:` load only when Claude reads a matching file. Rules without it load at
78
+ launch, at the same priority as `CLAUDE.md`, in every session — and, for anything shipped
79
+ in the template, in every downstream workspace too.
80
+
81
+ For reference content, scope is about reach. Pick the narrowest that works: auto-memory
82
+ (this machine only) → `team-member/{user}/` (one person) → `shared/` (the team) →
83
+ `shared/locked/` (the team, always loaded).
84
+
85
+ ## Step 3 — state the cost, then write
86
+
87
+ Before writing, run the footprint tool and report the projection to the user in one line.
88
+ This step is not optional. Making the cost visible at the decision point is the whole
89
+ reason this skill exists.
90
+
91
+ ```bash
92
+ node .claude/scripts/context-footprint.mjs --root .
93
+ node .claude/scripts/context-footprint.mjs --root . --add <bytes> --as <destination>
94
+ ```
95
+
96
+ Valid `--as` values: `rule`, `rule-scoped`, `locked`, `shared`, `team-member`, `memory`,
97
+ `skill`, `nowhere`.
98
+
99
+ Say it plainly — "this adds 1.4 KB to every session, taking always-loaded context from
100
+ 7.8% to 8.0% of the window" — and then write the file. If the projection looks
101
+ disproportionate to the value, go back to the table and take a cheaper row.
102
+
103
+ ## The canonical test
104
+
105
+ Canonical content is loaded verbatim into every session prompt, so it frames how Claude
106
+ reads the rest of the conversation. The principle: canonical describes what *is* and what
107
+ *to do*, never what *to think*.
108
+
109
+ > If Claude read this for the first time during a session about an unrelated topic, would
110
+ > it (a) help frame the problem correctly, or (b) push Claude toward a particular answer to
111
+ > a question that hasn't been asked yet?
112
+
113
+ (a) is canonical. (b) is `shared/` at most, more often `team-member/{user}/`.
114
+
115
+ **Belongs in `shared/locked/`:**
116
+
117
+ - **Facts about the system** — architecture, supported targets, naming conventions, where
118
+ things live, what is published.
119
+ - **Hard constraints** — "must work on Windows and macOS," "PRs only, no direct push to
120
+ main." Constraints scope the solution space without prejudging the solution.
121
+ - **Process rules** — workflow discipline that applies regardless of task.
122
+ - **Settled-rejection guardrails** — "we evaluated X, rejected it because Y, do not propose
123
+ it again." Must include the *why*, so a genuine edge case can still be recognised.
124
+ - **Meta-principles for debiasing** — explicit reminders to widen evaluation.
125
+
126
+ **Does not:**
127
+
128
+ - **Opinions on open technical questions.** "Library X beats Y here" makes Claude start
129
+ from the conclusion instead of reasoning toward it. Write it as a constraint with its
130
+ reason, or leave it in `shared/`.
131
+ - **Conclusions Claude might be asked to question.** Locking a conclusion on a topic still
132
+ under design biases the discussion before it starts.
133
+ - **Personal preferences.** Those are `team-member/{user}/`.
134
+ - **Status snapshots that age fast.** `project-status.md` is the bounded exception; finer
135
+ detail lives in the tracker.
136
+
137
+ Pre-loaded conclusions do not read as opinions to Claude — they read as ground truth. A
138
+ reference doc Claude *finds* while researching is weighed against the question; a canonical
139
+ doc loaded before the question is asked frames what Claude considers at all.
140
+
141
+ ## Writing a workspace-context file
142
+
143
+ Frontmatter fields — conventions, not all required on every file:
144
+
145
+ - `state` — `locked` (team truth, lives under `shared/locked/`) or `ephemeral`.
146
+ - `lifecycle` — for ephemeral files: `active` or `resolved`.
147
+ - `type` — `reference`, `braindump`, `handoff`, `research`, `design`, `index`, `canonical`,
148
+ `promoted`.
149
+ - `priority` — locked files only: `critical` (always in canonical) or `reference` (eligible
150
+ for trim or stub under budget pressure). Absent defaults to `critical`.
151
+ - `topic` — kebab-case slug matching the filename after any type prefix.
152
+ - `author` — required for `team-member/{user}/` files.
153
+ - `updated` — ISO date of last meaningful edit. `/maintenance` flags stale `active` files.
154
+ - `description` — one line, used verbatim by the generated indexes. Without it the index
155
+ falls back to the first sentence, then the filename slug. Adding one to a file with a
156
+ weak fallback is the cheapest possible index improvement.
157
+ - `confidence` — `high` | `medium` | `low`. Use on research, design, and exploration where
158
+ conclusions may still shift. Skip on locked files and on handoffs and braindumps.
159
+
160
+ ```yaml
161
+ ---
162
+ state: ephemeral
163
+ lifecycle: active
164
+ type: research
165
+ topic: vector-search-evaluation
166
+ description: Evaluation of FAISS for workspace-context — concluded the NL index is sufficient at our scale.
167
+ author: alex
168
+ confidence: medium
169
+ updated: 2026-04-25
170
+ ---
171
+ ```
172
+
173
+ ## Regenerating the indexes
174
+
175
+ One generator produces all three artifacts in a single pass:
176
+
177
+ - `workspace-context/index.md` — catalog of `shared/` (locked first), imported by `CLAUDE.md`.
178
+ - `workspace-context/canonical.md` — verbatim concatenation of `shared/locked/*.md`, also
179
+ imported by `CLAUDE.md`.
180
+ - `workspace-context/team-member/{user}/index.md` — per-user catalog, imported by each
181
+ user's gitignored `CLAUDE.local.md`.
182
+
183
+ ```bash
184
+ node .claude/scripts/build-workspace-context.mjs --check --root . # exits 1 if stale
185
+ node .claude/scripts/build-workspace-context.mjs --write --root . # regenerate
186
+ ```
187
+
188
+ Gitignored files (anything matching `local-only-*`) are excluded automatically, and
189
+ `workspace-context/.indexignore` adds path-prefix excludes for tracked files that should
190
+ not appear in the shared index.
191
+
192
+ When `canonical.md` exceeds `workspace.canonicalBudgetBytes` (default 40960), the builder
193
+ honours per-file `priority` and section-level `<!-- canonical:trim --> ... <!-- canonical:end-trim -->`
194
+ markers to fit: `priority: reference` files are trimmed, then stubbed; `priority: critical`
195
+ files are always included in full. `/maintenance` audits the budget and offers triage when
196
+ over.
197
+
198
+ Hand edits to `index.md`, `canonical.md`, or any per-user index are overwritten. Change the
199
+ source file or its `description:` instead.
@@ -0,0 +1,459 @@
1
+ ---
2
+ name: goal-driven-work
3
+ description: Use when setting up or running multi-phase autonomous work under Claude Code's built-in /goal command in a Ulysses workspace — drafting a goal-{topic}.md artifact, defining phases, dispatching phase agent teams, or completing a goal-driven session. Covers the frontmatter schema, the three phase types, gate conventions, the integration-branch model with per-phase sub-PRs, and model tiering.
4
+ ---
5
+
6
+ # Goal-Driven Work
7
+
8
+ The convention layer over Claude Code's built-in `/goal` command. `/goal` is the autonomy
9
+ loop; this skill gives the main agent durable phase state and a consistent dispatch pattern
10
+ across turns and resumes.
11
+
12
+ Read `.claude/rules/goal-driven-work.md` first for whether `/goal` is the right shape for
13
+ the work at hand. This skill covers everything after that decision.
14
+
15
+ ## File layout
16
+
17
+ - One `goal-{topic}.md` artifact per effort. Session model: at the top of the session worktree, alongside `session.md`. Task model: in the chat drawer at `workspace-scratchpad/chats/{chat}/`, alongside its phase outputs. One goal per worktree or task either way.
18
+ - The artifact's frontmatter holds machine state; its body holds the human-readable goal statement, per-phase intent, and a mandatory `## Start command` section (see "Kicking off the goal") with the literal `/goal "..."` invocation the user runs to start the loop.
19
+ - Phase output artifacts live as siblings. `research-*.md` and `crossref-*.md` are goal-native (produced by `parallel-research` and `crossref` phase types). `design-*.md` and `plan-*.md` are pre-existing session-artifact patterns that `type: skill` phases reuse when the wrapped skill is `superpowers:brainstorming` or `superpowers:writing-plans`; they are not goal-specific.
20
+ - Session model: the artifact is tracked on the session branch and lives there until `/complete-work` runs, which strips it before the final PR. Task model: the drawer is machine-local and untracked; `/complete-work` routes the artifact (promote into `workspace-context/` or discard) at completion.
21
+
22
+ ## Frontmatter schema
23
+
24
+ ```yaml
25
+ ---
26
+ type: goal
27
+ topic: <kebab-case-topic>
28
+ status: active # pending | active | complete | cancelled
29
+ # pending: artifact written but /goal not yet invoked
30
+ # active: /goal loop is running
31
+ # complete: all phases complete and condition met
32
+ # cancelled: goal abandoned without completion
33
+ current_phase: <phase-name> # the phase currently in progress or next up
34
+ completion_condition: > # mirrored from the `/goal` text so it survives `/goal clear` and is auditable
35
+ <multi-line condition string>
36
+ turn_budget: <int> # backstop; the condition itself should reference it
37
+ phases:
38
+ - name: <phase-name>
39
+ type: parallel-research | crossref | skill
40
+ status: pending # pending | in_progress | complete | failed
41
+ artifact: <path or null> # output artifact path, relative to worktree top; null for phases that commit to repo
42
+ gate: review | auto # default: review
43
+
44
+ # for type: parallel-research
45
+ team:
46
+ agents:
47
+ - subagent_type: <type>
48
+ brief: |
49
+ <multi-line brief>
50
+ - subagent_type: <type>
51
+ brief: |
52
+ <multi-line brief>
53
+ synthesizer:
54
+ subagent_type: <type>
55
+ brief: |
56
+ <multi-line brief: reads sibling agent outputs, writes the phase artifact>
57
+
58
+ # for type: crossref
59
+ inputs:
60
+ independent: <path to the just-produced independent artifact>
61
+ against: <path or list of paths to compare against>
62
+ brief: |
63
+ <multi-line brief for the crossref agent>
64
+
65
+ # for type: skill
66
+ skill: <plugin:skill-name> # e.g. superpowers:brainstorming
67
+ ---
68
+ ```
69
+
70
+ Body of the file is human-readable: the goal statement, success criteria, and per-phase intent in prose. The frontmatter is the source of truth for machine state; the body explains it to a reader.
71
+
72
+ ## Phase types
73
+
74
+ Three types in v1. Prefer wrapping existing skills (`type: skill`) when a skill fits. Only invent a phase type when no skill covers the work.
75
+
76
+ - **`parallel-research`**: dispatch N researcher-style agents in parallel with independent briefs. Optionally run a synthesizer agent over their outputs to write one consolidated artifact. Use for tool surveys, option-space exploration, comparative research where work can be partitioned.
77
+ - **`crossref`**: given source A (the independent output the team just produced) and source B (existing material to validate against), dispatch an agent to produce a gap-and-overlap matrix plus a ranked list of validations and concerns. Use to compare independent work against prior research, canonical workspace context, or third-party material.
78
+ - **`skill`**: invoke an existing skill by name. The phase's `artifact:` path is the expected output location (or `null` for phases that commit to repo, like `executing-plans`). Use this for `brainstorming`, `writing-plans`, `executing-plans`, `test-driven-development`, or any other skill that fits the phase's intent.
79
+
80
+ If a new phase type appears to be needed, justify why no existing skill fits before adding it. New phase types are a maintenance cost. Prefer keeping `parallel-research` and `crossref` inline until cross-goal reuse pressure makes the extraction earn its keep.
81
+
82
+ ## Agent-team dispatch pattern (v1: stateless)
83
+
84
+ A "team" is just a list of `Agent`-tool dispatch configs (`subagent_type` + `brief`). Spawned fresh each phase. No persistent identity, no memory across phases except what's written into artifacts.
85
+
86
+ The main agent's responsibility per phase:
87
+
88
+ 1. Mark the phase `in_progress` in `goal-{topic}.md` frontmatter.
89
+ 2. Dispatch the team via the `Agent` tool, in parallel where independent. The brief for each agent is exactly the `brief:` field from the phase config, prefixed with any context the brief itself doesn't already carry (typically a pointer to relevant sibling artifacts).
90
+ 3. Collect agent outputs.
91
+ 4. If a `synthesizer` is declared, dispatch it with the sibling outputs and have it write the `artifact:` file.
92
+ 5. If no synthesizer, the main agent writes the `artifact:` file directly from the collected outputs.
93
+ 6. Mark the phase `complete`, update `current_phase` to the next pending phase.
94
+ 7. End the turn with a short status note (for the `/goal` evaluator) and, if `gate: review`, an explicit question to the user.
95
+
96
+ The artifact format supports adding a `persistent: true` flag on a phase team later if a future goal needs team continuity across phases. Don't add the field until something asks for it.
97
+
98
+ ## Gate convention
99
+
100
+ `gate: review` is the default for every phase. At the gate, the main agent ends its turn by asking the user to approve, reject with feedback, or pause.
101
+
102
+ - **approve** → phase stays `complete`, advance to next phase next turn.
103
+ - **reject with feedback** → flip phase back to `pending` and re-run it with the feedback prepended to the team brief.
104
+ - **pause** → state is durable in `goal-{topic}.md`; resume with `claude --resume`.
105
+
106
+ The evaluator's "no, awaiting review" response after a gated phase does not bypass the wait for user input. It just keeps the session running until the user replies.
107
+
108
+ `gate: auto` is allowed in the format but discouraged in v1. Earn it after the workflow has been exercised at least once on the work in question.
109
+
110
+ ## Gate-budget interaction and turn-budget sizing
111
+
112
+ A goal's `turn_budget` is a backstop that counts every orchestrator turn — including the short "awaiting your gate decision" exchanges that happen at every `gate: review` phase. The `/goal` evaluator re-pings the session whenever it tries to settle without the completion condition being met, and each ping consumes a turn. Concretely: a multi-phase gated goal whose author is away for stretches can spend a meaningful fraction of its budget *idling at gates* rather than advancing work. A run with six `gate: review` phases burned roughly half of a 150-turn budget on gate-idle pings before the actual work completed.
113
+
114
+ This is a structural property of `/goal` + review gates, not a per-goal accident. Plan for it:
115
+
116
+ - **Drop `gate: review` from early phases when the human is expected to be away for stretches.** Research, crossref, spec, and plan phases produce artifacts that the final session→main PR consumes anyway. The integration-branch model (`main` untouched until `/complete-work` opens the final PR for human review) already provides one strong human review point; piling per-phase gates on top of that, with no human watching, just burns budget. Set those phases to `gate: auto` when the run will be unattended.
117
+ - **Keep `gate: review` for phases with irreversible side effects or for phases whose outcome reshapes subsequent phases.** A spec gate that lets the human reprioritize the ranked execution list before `executing-plans` walks it is worth its cost; a research-synthesis gate that just rubber-stamps a matrix the human will see again in the final PR is not.
118
+ - **Size `turn_budget` for the worst case you actually expect.** If every phase is `gate: auto`: budget ≈ (estimated work-turns) × 1.2 (small headroom for retries). If any phases are `gate: review` and the human may be unavailable for hours: multiply the work-turn estimate by **2–3×** to absorb gate-idle pings, or raise the budget mid-goal by editing the goal artifact's `completion_condition` and `turn_budget` fields. The Stop-hook condition string is fixed from the original `/goal` invocation; raising `turn_budget` in the artifact keeps the auditable source-of-truth correct for resume but does not change the running evaluator's text — the user can re-run the (updated) `## Start command` to refresh it after `/goal clear`.
119
+ - **If the goal stops on the backstop mid-work because of gate-idle waste, that is the system working as designed.** `claude --resume` plus re-running the `## Start command` continues from durable phase state; phase artifacts and per-phase `status: complete` markers survive. Do not treat backstop-stop as a failure of the goal.
120
+
121
+ This guidance is workspace-side mitigation only. The underlying friction — the evaluator pinging during gate-idle — is `/goal` harness behavior, not workspace code. A cleaner fix (suspend the evaluator at review gates so it does not re-ping until a new user message arrives) is tracked separately and would obsolete the multiplier above when it lands.
122
+
123
+ ## Integration branch and per-phase sub-PRs
124
+
125
+ While a `/goal`-driven effort is running, its work branch (`feature/{session-name}` for a session, the task's branch for a task) acts as the goal's integration branch. Main is untouched until `/complete-work` opens the final PR for human review. This is the key autonomy boundary: phase agents can merge their own work, repeatedly, throughout the goal — but only into the integration branch, never into main. The integration-branch and sub-branch guidance in this section applies to task branches unchanged.
126
+
127
+ Two merge strategies, picked per phase:
128
+
129
+ - **Direct commit (default for artifact-only phases).** Phases of type `parallel-research` and `crossref`, and skill phases that wrap artifact-producing skills (`brainstorming`, `writing-plans`), commit their outputs directly to the session branch. These artifacts are stripped before the final PR anyway, so sub-PR ceremony adds nothing.
130
+ - **Sub-branch with self-merged PR (default for code phases).** Skill phases that wrap code-producing skills (`executing-plans`, `test-driven-development`, `subagent-driven-development`) work on a per-phase sub-branch and open a PR back to the session branch. After the phase's gate is satisfied, the main agent self-merges the sub-PR. Each sub-PR is a discrete reviewable unit — failures can be retried by rebuilding just the one sub-branch.
131
+
132
+ Sub-branch naming follows the existing `git-conventions.md` rule (kebab-case after prefix, no nesting): `feature/{session-name}-{phase-name}`. Examples for a session named `ulysses-goals`:
133
+
134
+ - `feature/ulysses-goals` — session/integration branch
135
+ - `feature/ulysses-goals-strategy-research` — phase sub-branch (if a research phase were promoted to sub-branch strategy)
136
+ - `feature/ulysses-goals-implement` — phase sub-branch for the implement phase
137
+
138
+ A phase declares its merge strategy in frontmatter via an optional `integration:` block:
139
+
140
+ ```yaml
141
+ phases:
142
+ - name: implement
143
+ type: skill
144
+ skill: superpowers:executing-plans
145
+ integration:
146
+ strategy: sub-branch # direct | sub-branch
147
+ branch: feature/{session-name}-implement # explicit, or derived if omitted
148
+ self_merge: true # main agent merges the sub-PR after gate is satisfied; default true
149
+ ```
150
+
151
+ If `integration:` is omitted, defaults are applied by phase type per the list above.
152
+
153
+ Multi-repo sessions: each project repo has its own session branch (same name across repos per existing convention). Phase sub-branches are created per-repo where the phase commits, with matching names. The sub-PR target in each repo is that repo's session branch.
154
+
155
+ `/complete-work` verifies before the final PR that every phase sub-branch has been merged into the session branch. If any are outstanding, it fails loudly with a list of unmerged sub-branches. The user resolves them (merge or close) before re-running.
156
+
157
+ ## Model tiering
158
+
159
+ `/goal`-driven work uses a three-tier model assignment, biased toward the right tool for each shape of work:
160
+
161
+ - **Opus** for the main orchestration thread. It reads the goal artifact, dispatches phase teams, holds the long-running session context, and makes the per-phase advance/retry decisions. Reasoning-heavy and context-heavy work.
162
+ - **Sonnet** for phase team agents (researchers, crossref agents, skill subagents) and for synthesizers. Heavy lifting that runs in parallel and writes substantial artifacts. The default model for any agent in a `team:` block.
163
+ - **Haiku** for quick checks: file-existence verification, frontmatter validation, status sweeps, tiny lookups. Use when the work is bounded and obvious, and speed matters more than reasoning depth.
164
+
165
+ Phase frontmatter sets the model per agent via the `model:` field, which maps to the `Agent` tool's `model` parameter:
166
+
167
+ ```yaml
168
+ team:
169
+ agents:
170
+ - subagent_type: researcher
171
+ model: sonnet # default for parallel-research; rarely overridden
172
+ brief: |
173
+ ...
174
+ synthesizer:
175
+ subagent_type: researcher
176
+ model: sonnet
177
+ brief: |
178
+ ...
179
+ ```
180
+
181
+ For artifact-only or tagging work where Haiku is enough, declare it explicitly:
182
+
183
+ ```yaml
184
+ - subagent_type: researcher
185
+ model: haiku
186
+ brief: |
187
+ Scan {paths} for files matching {pattern} and return a tagged list.
188
+ ```
189
+
190
+ The main agent's model is determined by the harness (the user's `claude` invocation), not the goal artifact. Goal artifacts assume the main agent runs on Opus.
191
+
192
+ ## Writing a good completion condition
193
+
194
+ The `/goal` evaluator runs after every turn against the conversation transcript. It does not call tools, so it can only judge what the main turn has surfaced. A good condition is:
195
+
196
+ - **Specific.** Names the artifacts that must exist (file paths, commit references, PR URLs) rather than vague outcomes.
197
+ - **Demonstrable from transcript.** The main agent's own output must be able to evidence completion. "The PR URL was reported in the transcript and `git status` showed clean" rather than "the work feels done."
198
+ - **Bounded.** Includes a turn budget as a backstop (e.g., "or stop after 60 turns") so the loop can't run away if something goes wrong.
199
+ - **Within the 4000-char limit.** Up to four kilobytes of condition text are accepted.
200
+ - **Reachable by the agent.** This is the one that gets written wrong. If any phase carries
201
+ `gate: review`, or the goal depends on a merge, a deploy, or anything else only the
202
+ operator can authorise, then a condition demanding those be *done* can never be satisfied
203
+ by the agent — and the evaluator will re-ping indefinitely against work that is correctly
204
+ waiting. Either the condition names the blocked state as terminal, or the goal cannot end
205
+ without the backstop.
206
+
207
+ **The failure this prevents, observed.** A goal was started with
208
+
209
+ > All 6 phases show status: complete. […] The seven PRs are merged or closed.
210
+
211
+ while three of its phases were `gate: review` and the PRs depended on two prerequisite PRs
212
+ merging to `main`. Everything the agent could do was done inside about forty turns; the
213
+ remaining eighty were spent re-reporting the same three blockers, because the condition had
214
+ no terminal state for "built, verified, awaiting authorisation."
215
+
216
+ Write the disjunction in from the start:
217
+
218
+ ```
219
+ Every phase shows status: complete, or status: awaiting-review with its artifact
220
+ written. […] Either <external step> is done, or it is still pending and that is
221
+ stated in the transcript. Or stop after <N> turns.
222
+ ```
223
+
224
+ Editing the artifact afterwards does not rescue a running goal: the evaluator's text is
225
+ fixed from the original `/goal` invocation. The corrected condition only takes effect after
226
+ `/goal clear` and a re-run of the `## Start command`.
227
+
228
+ A reasonable template:
229
+
230
+ ```
231
+ Every phase in goal-<topic>.md shows status: complete, or status: awaiting-review with its artifact written. Phase artifacts exist at: <list paths>. Either the /complete-work skill has opened the final PR with the URL in the transcript, or the goal is blocked on a named operator step and that is stated. Or stop after <N> turns.
232
+ ```
233
+
234
+ Fill in `<topic>`, paths, and `<N>` per goal. Anchor on artifacts and committed state, not on feelings.
235
+
236
+ ## Kicking off the goal
237
+
238
+ `/goal` is a Claude Code built-in that the **user** types — the agent cannot invoke it. So the moment the artifact is ready is the load-bearing hand-off, and the artifact itself carries the instruction rather than relying on an agent chat message that vanishes on the next compaction or resume.
239
+
240
+ Every `goal-{topic}.md` body MUST include a `## Start command` section containing the literal, copy-paste-ready invocation:
241
+
242
+ ````markdown
243
+ ## Start command
244
+
245
+ ```
246
+ /goal "All phases in goal-<topic>.md show status: complete. Phase artifacts exist at: <paths>. The /complete-work skill has opened the final PR; the PR URL appeared in the transcript. Or stop after <N> turns."
247
+ ```
248
+ ````
249
+
250
+ Rules for the `## Start command`:
251
+
252
+ - It is the `completion_condition` flattened to a **single line** and wrapped in `/goal "..."`. The frontmatter `completion_condition:` (a multi-line folded scalar) is the auditable source of truth; the start command is its runnable rendering. The two must express the same condition — if you edit one, re-derive the other.
253
+ - Flatten by collapsing the folded scalar's newlines to single spaces. Escape any embedded double quotes. Keep it within the 4000-character `/goal` limit.
254
+ - It lives in the body, not the frontmatter, because it is for a human to copy, not for machine parsing.
255
+
256
+ When the artifact is drafted and the user has reviewed it, the agent's hand-off is: point the user at the `## Start command` block and let them run it. Running it flips the goal from `status: pending` to `status: active` (the first `/goal` turn updates the frontmatter per the dispatch pattern). The agent never types `/goal` itself.
257
+
258
+ ## Lifecycle integration
259
+
260
+ - `/goal` runs inside started work. It does NOT replace `/start-work`. Session model: the session is created the normal way and the goal artifact is drafted at the worktree top (including its `## Start command` block). Task model: the task is started the normal way and the artifact is drafted into the chat drawer. Either way the user runs the `## Start command` block's `/goal "..."` to kick off the loop.
261
+ - Session model: the goal artifact lives on the session branch and travels with `git push`, surviving across machines and `--resume`. Task model: the drawer is machine-local — push the task branch for the code, and `/promote` the artifact if it must travel.
262
+ - Session model: `session.md`'s `## Tasks` should mirror the phase list at coarse grain (one task per phase) so `TodoWrite` shows high-level progress; the main agent updates `## Tasks` at phase transitions via the helper specified by the `task-list-mirroring` rule, in addition to updating `goal-{topic}.md`. Task model: TodoWrite is the live mirror and there is no durable `## Tasks` — the goal artifact's own `phases:` list is the durable state.
263
+ - `/pause-work` works without special handling. The goal-evaluator state resets on resume per the Claude Code docs; phase state is durable in the artifact.
264
+ - `/complete-work` reads `goal-*.md`, `research-*.md`, and `crossref-*.md` for release-note synthesis and strips them from the branch before the final PR, alongside the existing `design-*.md` and `plan-*.md` handling. When a goal artifact is present, it also runs a pre-flight check that every declared sub-branch (from phases with `integration.strategy: sub-branch`) has been merged into the session branch. Unmerged sub-branches abort completion with a clear list to resolve.
265
+
266
+ If a research or crossref artifact deserves to outlive the branch, the user runs `/promote` on it before `/complete-work`. `/promote` accepts arbitrary paths and routes them into `workspace-context/`.
267
+
268
+ ## Tracker integration
269
+
270
+ Goals do not replace work items. A goal is execution shape; a work item is the unit the team tracks. See `work-item-tracking.md` for how `workItem:` in session frontmatter links a session to its tracker issue. A goal lives inside a session and therefore inherits the session's `workItem:`.
271
+
272
+ When the goal artifact is itself the deliverable for a future session to execute (i.e., this session's job was to *write* the goal, and another session will *run* it), the strip rule needs an escape hatch. The pattern: preserve the artifact in the tracker issue body (as a fenced code block) before `/complete-work` runs. The future session picks up the issue via `/start-work`, copies the artifact text into its worktree top, and runs `/goal`. The strip rule stays clean and consistent; the deliverable is preserved through the tracker.
273
+
274
+ ## Out of scope
275
+
276
+ - Nested or sub-goals. v1 is flat.
277
+ - DAG between phases. Sequential at the phase level; parallelism only inside a phase via the team agents.
278
+ - Persistent named teams (`TeamCreate`). v1 is stateless dispatch.
279
+ - Frontmatter linter for `goal-*.md`. Manual review is fine for v1; revisit if workspaces using the template hit consistent shape errors.
280
+ - `/start-work` seeding of a `goal-{slug}.md` skeleton. Manual drafting is the v1 path. The drafting itself is high-leverage thinking; a template skeleton would risk skipping that.
281
+
282
+ ## Appendix: worked example
283
+
284
+ A complete `goal-evaluate-rate-limiting.md` illustrating all three phase types. The topic is intentionally generic — the example is reference material, not prescriptive. (The appendix is fenced with four backticks so the example's own `## Start command` code block renders intact.)
285
+
286
+ ````yaml
287
+ ---
288
+ type: goal
289
+ topic: evaluate-rate-limiting
290
+ status: pending
291
+ current_phase: strategy-research
292
+ completion_condition: >
293
+ All 5 phases in goal-evaluate-rate-limiting.md show status: complete.
294
+ Phase artifacts exist at: research-rate-limiting-strategies.md,
295
+ crossref-existing-infrastructure.md, design-rate-limiting.md,
296
+ plan-rate-limiting.md, and the implementation commits land on the session
297
+ branch (visible in git log). The /complete-work skill has produced release
298
+ notes and opened the final PR; the PR URL appeared in the transcript.
299
+ Or stop after 60 turns.
300
+ turn_budget: 60
301
+ phases:
302
+ - name: strategy-research
303
+ type: parallel-research
304
+ status: pending
305
+ artifact: research-rate-limiting-strategies.md
306
+ gate: review
307
+ integration:
308
+ strategy: direct # artifact-only phase; commits straight to session branch
309
+ team:
310
+ agents:
311
+ - subagent_type: researcher
312
+ model: sonnet
313
+ brief: |
314
+ Research the token-bucket rate-limiting algorithm. Cover the
315
+ mechanics, parameter trade-offs (capacity, refill rate), edge
316
+ cases (burst behavior, clock skew), reference implementations in
317
+ popular libraries, and known production failure modes.
318
+
319
+ Output: a markdown report under 1,200 words. Return the content
320
+ in your response.
321
+ - subagent_type: researcher
322
+ model: sonnet
323
+ brief: |
324
+ Research the sliding-window rate-limiting algorithm. Cover both
325
+ the sliding log and sliding counter variants, accuracy
326
+ trade-offs, memory cost at scale, reference implementations, and
327
+ known production failure modes.
328
+
329
+ Output: a markdown report under 1,200 words. Return the content
330
+ in your response.
331
+ - subagent_type: researcher
332
+ model: sonnet
333
+ brief: |
334
+ Research the leaky-bucket rate-limiting algorithm. Cover the
335
+ queue-based and meter-based variants, smoothing behavior under
336
+ burst, comparison to token bucket, reference implementations,
337
+ and known production failure modes.
338
+
339
+ Output: a markdown report under 1,200 words. Return the content
340
+ in your response.
341
+ synthesizer:
342
+ subagent_type: researcher
343
+ model: sonnet
344
+ brief: |
345
+ Synthesize the three algorithm reports into a single
346
+ recommendation document at research-rate-limiting-strategies.md
347
+ (top of the active worktree).
348
+
349
+ Frontmatter: type: research, topic: rate-limiting-strategies,
350
+ state: ephemeral, lifecycle: active, confidence: medium,
351
+ updated: <today's date>.
352
+
353
+ Document structure:
354
+ - One-line recommendation up front
355
+ - Comparison table across criteria (accuracy, memory cost, burst
356
+ behavior, implementation complexity, operational debuggability)
357
+ - Per-algorithm summary with strengths and weaknesses
358
+ - Risks and unknowns
359
+ - References
360
+
361
+ Maximum 2,000 words.
362
+
363
+ - name: crossref-existing-infrastructure
364
+ type: crossref
365
+ status: pending
366
+ artifact: crossref-existing-infrastructure.md
367
+ inputs:
368
+ independent: research-rate-limiting-strategies.md
369
+ against:
370
+ - workspace-context/canonical.md
371
+ - repos/api-gateway/
372
+ gate: review
373
+ integration:
374
+ strategy: direct
375
+ agent:
376
+ subagent_type: researcher
377
+ model: sonnet
378
+ brief: |
379
+ Compare the rate-limiting strategy recommendation against the
380
+ existing infrastructure (the repos/api-gateway/ codebase and any
381
+ relevant canonical workspace context).
382
+
383
+ Produce a gap-and-overlap matrix: where does the recommended
384
+ strategy align with current patterns, where does it diverge, and
385
+ what migration friction is implied. Rank concerns by severity.
386
+
387
+ - name: spec
388
+ type: skill
389
+ status: pending
390
+ skill: superpowers:brainstorming
391
+ artifact: design-rate-limiting.md
392
+ gate: review
393
+ integration:
394
+ strategy: direct # spec lands as design-*.md at worktree top; stripped before final PR
395
+
396
+ - name: plan
397
+ type: skill
398
+ status: pending
399
+ skill: superpowers:writing-plans
400
+ artifact: plan-rate-limiting.md
401
+ gate: review
402
+ integration:
403
+ strategy: direct
404
+
405
+ - name: implement
406
+ type: skill
407
+ status: pending
408
+ skill: superpowers:executing-plans
409
+ artifact: null
410
+ gate: review
411
+ integration:
412
+ strategy: sub-branch # code phase: sub-branch + self-merged PR
413
+ branch: feature/{session-name}-implement
414
+ self_merge: true
415
+ ---
416
+
417
+ # Goal: Evaluate and ship a rate-limiting strategy
418
+
419
+ The api-gateway needs rate limiting before the next traffic step-up. This
420
+ goal runs the full arc from candidate-algorithm research through implementation
421
+ on the session branch.
422
+
423
+ ## Start command
424
+
425
+ ```
426
+ /goal "All 5 phases in goal-evaluate-rate-limiting.md show status: complete. Phase artifacts exist at: research-rate-limiting-strategies.md, crossref-existing-infrastructure.md, design-rate-limiting.md, plan-rate-limiting.md, and the implementation commits land on the session branch (visible in git log). The /complete-work skill has opened the final PR; the PR URL appeared in the transcript. Or stop after 60 turns."
427
+ ```
428
+
429
+ This is the frontmatter `completion_condition` flattened to one line. Run it after reviewing the artifact; it flips the goal to `status: active`.
430
+
431
+ ## Per-phase intent
432
+
433
+ 1. **strategy-research** runs three researchers in parallel, one per
434
+ candidate algorithm, plus a synthesizer that writes the comparative
435
+ recommendation.
436
+
437
+ 2. **crossref-existing-infrastructure** validates the recommendation
438
+ against the current api-gateway codebase and canonical workspace
439
+ context, surfacing migration friction and divergences.
440
+
441
+ 3. **spec** wraps `superpowers:brainstorming` to produce the design doc
442
+ from the validated recommendation.
443
+
444
+ 4. **plan** wraps `superpowers:writing-plans` to produce the
445
+ implementation checklist.
446
+
447
+ 5. **implement** wraps `superpowers:executing-plans` on a per-phase
448
+ sub-branch (`feature/{session-name}-implement`). The main agent opens
449
+ a PR from the sub-branch back into the session/integration branch and
450
+ self-merges once the gate is satisfied. Sub-branch is deleted post-merge.
451
+
452
+ Each gate is `review`: the user approves, rejects with feedback, or
453
+ pauses at the end of every phase. No auto-advance in v1.
454
+
455
+ All phase agents and synthesizers run on Sonnet (the default for team work).
456
+ The main agent reading and orchestrating this goal runs on Opus. Haiku is
457
+ unused in this example; it would be appropriate for a phase whose sole job
458
+ is, e.g., scanning a tree and returning a tagged file list.
459
+ ````