@ulysses-ai/create-workspace 0.17.0-beta.0 → 0.18.0-beta.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +3 -3
- package/package.json +1 -1
- package/template/.claude/hooks/_utils.mjs +1 -1
- package/template/.claude/hooks/repo-write-detection.mjs +161 -64
- package/template/.claude/hooks/session-start.mjs +35 -1
- package/template/.claude/hooks/subagent-start.mjs +89 -22
- package/template/.claude/lib/session-frontmatter.mjs +28 -0
- package/template/.claude/rules/coherent-revisions.md +1 -1
- package/template/.claude/rules/forge-operations.md +37 -93
- package/template/.claude/rules/git-conventions.md +16 -11
- package/template/.claude/rules/goal-driven-work.md +8 -416
- package/template/.claude/rules/honest-pushback.md +37 -37
- package/template/.claude/rules/memory-guidance.md +43 -90
- package/template/.claude/rules/superpowers-workflow.md.skip +1 -1
- package/template/.claude/rules/work-item-tracking.md +30 -72
- package/template/.claude/rules/workspace-structure.md +36 -94
- package/template/.claude/scripts/build-workspace-context.mjs +61 -16
- package/template/.claude/scripts/chat-record.mjs +282 -0
- package/template/.claude/scripts/cleanup-work-session.mjs +257 -68
- package/template/.claude/scripts/context-footprint.mjs +282 -0
- package/template/.claude/scripts/forges/github.mjs +45 -0
- package/template/.claude/scripts/forges/gitlab.mjs +3 -2
- package/template/.claude/scripts/forges/interface.mjs +12 -0
- package/template/.claude/scripts/generate-claude-local.mjs +21 -2
- package/template/.claude/scripts/migrate-sessions.mjs +1571 -0
- package/template/.claude/scripts/migrate-to-workspace-context.mjs +7 -2
- package/template/.claude/scripts/task-worktree.mjs +525 -0
- package/template/.claude/scripts/workspace-diagnostics.mjs +654 -0
- package/template/.claude/skills/braindump/SKILL.md +11 -4
- package/template/.claude/skills/build-docs-site/SKILL.md +5 -5
- package/template/.claude/skills/build-docs-site/templates/spec.md.tmpl +1 -1
- package/template/.claude/skills/complete-work/SKILL.md +229 -219
- package/template/.claude/skills/context-placement/SKILL.md +199 -0
- package/template/.claude/skills/goal-driven-work/SKILL.md +459 -0
- package/template/.claude/skills/handoff/SKILL.md +11 -4
- package/template/.claude/skills/maintenance/SKILL.md +7 -0
- package/template/.claude/skills/migrate-sessions/SKILL.md +70 -0
- package/template/.claude/skills/pause-work/SKILL.md +9 -1
- package/template/.claude/skills/release/SKILL.md +44 -108
- package/template/.claude/skills/start-work/SKILL.md +89 -7
- package/template/.claude/skills/workspace-init/SKILL.md +3 -1
- package/template/.claude/skills/workspace-update/SKILL.md +4 -0
- package/template/CLAUDE.md.tmpl +19 -2
- package/template/_gitignore +9 -0
- package/template/workspace.json.tmpl +3 -2
|
@@ -1,107 +1,51 @@
|
|
|
1
|
-
Activate this rule if the workspace creates PRs, watches CI runs, or interacts with releases from skills. Sibling to `work-item-tracking.md` (which covers issues); together they cover everything a workspace
|
|
1
|
+
Activate this rule if the workspace creates PRs, watches CI runs, or interacts with releases from skills. Sibling to `work-item-tracking.md` (which covers issues); together they cover everything a workspace does against a code-hosting forge.
|
|
2
2
|
|
|
3
3
|
# Forge Operations
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
**Skills never call `gh` (or `glab`, or any forge CLI) inline for pull-request, release, or
|
|
6
|
+
workflow-run operations.** They go through the adapter at `.claude/scripts/forges/{type}.mjs`,
|
|
7
|
+
reached via `createForge()` from `.claude/scripts/forges/interface.mjs`.
|
|
6
8
|
|
|
7
|
-
|
|
9
|
+
The interface module is the source of truth for the available methods and their shapes —
|
|
10
|
+
read it when you need the API. `/complete-work` and `/pause-work` each carry the exact calls
|
|
11
|
+
they make, so in practice you rarely need to look.
|
|
8
12
|
|
|
9
|
-
|
|
10
|
-
- **Testable.** The adapter takes an injectable `spawnFn`; unit tests mock subprocess calls instead of running them.
|
|
11
|
-
- **One vocabulary.** Skills reason about `forge.prCreate`, `forge.prMerge`, `forge.workflowRunFind`, `forge.workflowRunWatch`, `forge.releaseView` regardless of backend. Failure modes share typed errors (`PrNotFound`, `MergeRejected`, `WorkflowNotFound`, `ReleaseNotFound`) instead of every callsite parsing stderr.
|
|
13
|
+
## Why
|
|
12
14
|
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
`
|
|
16
|
-
|
|
17
|
-
```json
|
|
18
|
-
{
|
|
19
|
-
"workspace": {
|
|
20
|
-
"forge": {
|
|
21
|
-
"type": "github"
|
|
22
|
-
}
|
|
23
|
-
}
|
|
24
|
-
}
|
|
25
|
-
```
|
|
26
|
-
|
|
27
|
-
- `type` — identifies the adapter module at `.claude/scripts/forges/{type}.mjs`. `github` is the default and the only fully-implemented adapter today. `gitlab.mjs` ships as a stub that throws `NOT_IMPLEMENTED` with a contribution pointer.
|
|
28
|
-
- `repo` (optional) — adapter-specific. For `github`, the `owner/name` slug to target. When unset or `"auto"`, the adapter resolves the repo from the local git `origin` remote.
|
|
29
|
-
|
|
30
|
-
Absence of `workspace.forge` is treated as `{ type: 'github' }` (back-compat for workspaces that predate the field). Setting `workspace.forge: false` explicitly disables forge operations — every adapter method then throws `FORGE_DISABLED`.
|
|
31
|
-
|
|
32
|
-
## Adapter interface (for Claude)
|
|
33
|
-
|
|
34
|
-
Import from `.claude/scripts/forges/interface.mjs`:
|
|
35
|
-
|
|
36
|
-
```javascript
|
|
37
|
-
import { createForge, PrNotFound, MergeRejected, WorkflowNotFound, ReleaseNotFound } from '.claude/scripts/forges/interface.mjs';
|
|
38
|
-
import { readFileSync } from 'node:fs';
|
|
39
|
-
|
|
40
|
-
const ws = JSON.parse(readFileSync('workspace.json', 'utf-8'));
|
|
41
|
-
const forge = createForge(ws.workspace?.forge);
|
|
42
|
-
|
|
43
|
-
// Pull request lifecycle
|
|
44
|
-
const pr = await forge.prCreate({
|
|
45
|
-
title: 'feat: add forge adapter',
|
|
46
|
-
body: 'long-form body',
|
|
47
|
-
draft: false, // omit or false for normal PRs; true for /pause-work drafts
|
|
48
|
-
base: 'main', // optional; defaults to repo default branch
|
|
49
|
-
head: 'feature/forge', // optional; defaults to current branch
|
|
50
|
-
}); // → { id: 'owner/repo#42', url, number }
|
|
15
|
+
Switching code hosts becomes a `workspace.json` field plus one adapter file instead of a
|
|
16
|
+
sweep across six skills. Failure modes share typed errors — `PrNotFound`, `MergeRejected`,
|
|
17
|
+
`WorkflowNotFound`, `ReleaseNotFound` — instead of every callsite parsing stderr. And the
|
|
18
|
+
adapter takes an injectable `spawnFn`, so tests mock subprocesses rather than running them.
|
|
51
19
|
|
|
52
|
-
|
|
53
|
-
id: pr.id,
|
|
54
|
-
strategy: 'squash', // 'merge' | 'squash' | 'rebase'
|
|
55
|
-
deleteBranch: true,
|
|
56
|
-
});
|
|
57
|
-
|
|
58
|
-
const view = await forge.prView({ id: pr.id });
|
|
59
|
-
// → { state, mergeable, mergeStateStatus, reviewDecision, ... }
|
|
60
|
-
|
|
61
|
-
// Releases (lookup only — release creation lives in the workflow tag-push)
|
|
62
|
-
const release = await forge.releaseView({ tag: 'v1.2.3' });
|
|
63
|
-
// → { tag, name, url, publishedAt, isDraft, isPrerelease }
|
|
64
|
-
// throws ReleaseNotFound if the tag has no release
|
|
65
|
-
|
|
66
|
-
// Workflow runs (used by /complete-work to follow the publish workflow)
|
|
67
|
-
const run = await forge.workflowRunFind({
|
|
68
|
-
workflow: 'publish.yml',
|
|
69
|
-
branch: 'v1.2.3',
|
|
70
|
-
limit: 1,
|
|
71
|
-
}); // → { runId, status, conclusion, url } | null
|
|
72
|
-
const result = await forge.workflowRunWatch({
|
|
73
|
-
runId: run.runId,
|
|
74
|
-
exitStatus: true, // true: exit non-zero on workflow failure
|
|
75
|
-
}); // → { exitCode } — does NOT throw on workflow failure
|
|
76
|
-
```
|
|
77
|
-
|
|
78
|
-
All methods are async. Adapter-detectable failures throw typed errors; raw spawn failures throw `Error`.
|
|
79
|
-
|
|
80
|
-
## Skill behavior
|
|
81
|
-
|
|
82
|
-
Skills that interact with forge operations:
|
|
83
|
-
|
|
84
|
-
- **`/pause-work`** — creates draft PRs via `forge.prCreate({ draft: true })`.
|
|
85
|
-
- **`/complete-work`** — creates PRs, merges them with `strategy: 'squash'`, and (release sessions only) finds + watches the publish workflow via `workflowRunFind` + `workflowRunWatch`. Uses `releaseView` to investigate existing tag conflicts before re-tagging.
|
|
86
|
-
|
|
87
|
-
Skills outside that list do not call the forge adapter; they either don't touch the forge or they touch it for tracker-setup-specific operations that are deliberately scoped out (see below).
|
|
20
|
+
## Configuration
|
|
88
21
|
|
|
89
|
-
|
|
22
|
+
`workspace.json` → `workspace.forge`: `{ "type": "github" }`. `type` names the adapter module;
|
|
23
|
+
`github` is the default and the only complete one, `gitlab.mjs` is a stub that throws
|
|
24
|
+
`NOT_IMPLEMENTED`. Optional `repo` is an `owner/name` slug; unset or `"auto"` resolves from the
|
|
25
|
+
git `origin` remote. An absent `workspace.forge` is treated as `{ type: 'github' }`; setting it
|
|
26
|
+
to `false` makes every adapter method throw `FORGE_DISABLED`.
|
|
90
27
|
|
|
91
|
-
|
|
92
|
-
- **Tracker-setup repo configuration.** `/setup-tracker` uses `gh repo view --json hasIssuesEnabled` and `gh api repos/{slug} -X PATCH -f has_issues=true` to inspect and enable the Issues feature on a GitHub repo. Those are GitHub-API-specific setup operations, not the cross-cutting PR/release ops the forge abstraction targets. A GitLab user running `/setup-tracker` would follow a different setup flow entirely, so wrapping these in the forge adapter would create a leaky abstraction. They remain direct `gh` calls.
|
|
93
|
-
- **`gh repo view` as a remote-type probe.** `/complete-work` uses `gh repo view` to detect whether a remote is a GitHub remote (separately from any PR operation that follows). This is a one-line capability check, not an operation that benefits from forge wrapping. It stays direct.
|
|
94
|
-
- **Repo creation.** `gh repo create` (in `/sync-work` and `/workspace-init` setup narratives) is an interactive one-off used when a workspace lacks a remote. No forge adapter method for it — pointing users at a wrapped form when none exists would be worse than the current direct mention.
|
|
95
|
-
- **Manual operator recovery.** `gh run rerun`, `gh run view`, `gh release view` referenced in `/release` recovery guidance are documented for an operator at a terminal investigating a failed publish. The forge adapter is for *skill code*, not the manual recovery prose.
|
|
28
|
+
## Deliberate exceptions — do not "fix" these
|
|
96
29
|
|
|
97
|
-
|
|
30
|
+
These stay as direct `gh` calls. Wrapping them would create a leaky abstraction, so leave them
|
|
31
|
+
alone:
|
|
98
32
|
|
|
99
|
-
|
|
33
|
+
- **Issue lifecycle** — issues, comments, labels and milestones belong to the tracker adapter.
|
|
34
|
+
See `work-item-tracking.md`. The two abstractions are intentionally separate.
|
|
35
|
+
- **`/setup-tracker` repo configuration** — `gh repo view --json hasIssuesEnabled` and
|
|
36
|
+
`gh api repos/{slug} -X PATCH -f has_issues=true` are GitHub-API-specific setup, not
|
|
37
|
+
cross-cutting operations. A GitLab user's setup flow differs entirely.
|
|
38
|
+
- **`gh repo view` as a remote-type probe** — a one-line capability check, not an operation.
|
|
39
|
+
- **`gh repo create`** — an interactive one-off when a workspace has no remote.
|
|
40
|
+
- **Manual recovery prose** — `gh run rerun`, `gh run view`, `gh release view` in `/release`
|
|
41
|
+
guidance are for an operator at a terminal, not for skill code.
|
|
100
42
|
|
|
101
|
-
|
|
43
|
+
## Boundaries
|
|
102
44
|
|
|
103
|
-
|
|
45
|
+
The adapter covers the operations the template's skills actually perform, not every `gh`
|
|
46
|
+
capability. New operations land as additive interface methods, never by a skill going around
|
|
47
|
+
the adapter. Forge-native features — PR comments, review webhooks, branch protection — remain
|
|
48
|
+
UI and direct-CLI territory.
|
|
104
49
|
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
- Does not promise that every `gh` capability is wrapped. The adapter covers the operations the template's skills actually perform. New operations land via additive interface methods, not by skills going around the adapter.
|
|
50
|
+
Workspaces predating `workspace.forge` keep working, since `createForge(undefined)` defaults to
|
|
51
|
+
GitHub. `/maintenance` surfaces a notice suggesting the explicit value.
|
|
@@ -10,17 +10,22 @@
|
|
|
10
10
|
|
|
11
11
|
## Worktrees
|
|
12
12
|
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
-
|
|
13
|
+
Both lifecycles are built on worktrees; `workspace.sessionModel` in `workspace.json` routes new work to one of them.
|
|
14
|
+
|
|
15
|
+
**Task model** (`"task"`): one worktree per repo the task touches, created and removed with `.claude/scripts/task-worktree.mjs` (the chat stays at the workspace root):
|
|
16
|
+
|
|
17
|
+
- Project repos: `repos/{repo}/.claude/worktrees/{slug}/`, where `{slug}` is the branch with `/` replaced by `-`
|
|
18
|
+
- The workspace repo itself, addressed as `.`: `.claude/worktrees/{slug}/` — Claude Code's native worktree location
|
|
19
|
+
- Source clones at `repos/{repo}/` stay on their default branch; `/complete-work` tears each worktree down with `task-worktree.mjs --remove` (worktree first, then the branch)
|
|
20
|
+
|
|
21
|
+
**Session model** (default): N+1 worktrees in one self-contained folder at `work-sessions/{session-name}/`:
|
|
22
|
+
|
|
23
|
+
- `work-sessions/{session-name}/workspace/` — workspace worktree
|
|
24
|
+
- `work-sessions/{session-name}/workspace/repos/{repo-name}/` — project worktrees nested inside it (no symlink)
|
|
25
|
+
- Example, session `fix-auth` on `bugfix/fix-auth` touching `my-app` and `my-api`: `work-sessions/fix-auth/workspace/` plus `workspace/repos/my-app/` and `workspace/repos/my-api/`
|
|
26
|
+
- Teardown order is mandatory: project worktrees first, then the workspace worktree, then prune — the cleanup helper enforces it
|
|
27
|
+
|
|
28
|
+
The workspace `.gitignore` covers all of it: `repos` (no trailing slash) matches the root's `repos/` and every session worktree's nested `repos/`; `.claude/worktrees/` matches task worktrees of the workspace repo itself.
|
|
24
29
|
|
|
25
30
|
## Branch Maintenance
|
|
26
31
|
|
|
@@ -12,421 +12,13 @@ Use `/goal` when the work meets all three:
|
|
|
12
12
|
|
|
13
13
|
If any of the three fails, prefer plain session work or a single skill invocation. `/goal` is overhead. Pay it only when the work is long enough to earn it.
|
|
14
14
|
|
|
15
|
-
##
|
|
15
|
+
## The convention lives in a skill
|
|
16
16
|
|
|
17
|
-
|
|
18
|
-
-
|
|
19
|
-
-
|
|
20
|
-
- The artifact is tracked on the session branch and lives there until `/complete-work` runs. It is removed from the branch before the final PR alongside other session artifacts.
|
|
17
|
+
Everything past this decision — the `goal-{topic}.md` frontmatter schema, the three phase
|
|
18
|
+
types, agent-team dispatch, gate conventions, the integration-branch model with per-phase
|
|
19
|
+
sub-PRs, model tiering, and the worked example — is in the `goal-driven-work` skill.
|
|
21
20
|
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
type: goal
|
|
27
|
-
topic: <kebab-case-topic>
|
|
28
|
-
status: active # pending | active | complete | cancelled
|
|
29
|
-
# pending: artifact written but /goal not yet invoked
|
|
30
|
-
# active: /goal loop is running
|
|
31
|
-
# complete: all phases complete and condition met
|
|
32
|
-
# cancelled: goal abandoned without completion
|
|
33
|
-
current_phase: <phase-name> # the phase currently in progress or next up
|
|
34
|
-
completion_condition: > # mirrored from the `/goal` text so it survives `/goal clear` and is auditable
|
|
35
|
-
<multi-line condition string>
|
|
36
|
-
turn_budget: <int> # backstop; the condition itself should reference it
|
|
37
|
-
phases:
|
|
38
|
-
- name: <phase-name>
|
|
39
|
-
type: parallel-research | crossref | skill
|
|
40
|
-
status: pending # pending | in_progress | complete | failed
|
|
41
|
-
artifact: <path or null> # output artifact path, relative to worktree top; null for phases that commit to repo
|
|
42
|
-
gate: review | auto # default: review
|
|
43
|
-
|
|
44
|
-
# for type: parallel-research
|
|
45
|
-
team:
|
|
46
|
-
agents:
|
|
47
|
-
- subagent_type: <type>
|
|
48
|
-
brief: |
|
|
49
|
-
<multi-line brief>
|
|
50
|
-
- subagent_type: <type>
|
|
51
|
-
brief: |
|
|
52
|
-
<multi-line brief>
|
|
53
|
-
synthesizer:
|
|
54
|
-
subagent_type: <type>
|
|
55
|
-
brief: |
|
|
56
|
-
<multi-line brief: reads sibling agent outputs, writes the phase artifact>
|
|
57
|
-
|
|
58
|
-
# for type: crossref
|
|
59
|
-
inputs:
|
|
60
|
-
independent: <path to the just-produced independent artifact>
|
|
61
|
-
against: <path or list of paths to compare against>
|
|
62
|
-
brief: |
|
|
63
|
-
<multi-line brief for the crossref agent>
|
|
64
|
-
|
|
65
|
-
# for type: skill
|
|
66
|
-
skill: <plugin:skill-name> # e.g. superpowers:brainstorming
|
|
67
|
-
---
|
|
68
|
-
```
|
|
69
|
-
|
|
70
|
-
Body of the file is human-readable: the goal statement, success criteria, and per-phase intent in prose. The frontmatter is the source of truth for machine state; the body explains it to a reader.
|
|
71
|
-
|
|
72
|
-
## Phase types
|
|
73
|
-
|
|
74
|
-
Three types in v1. Prefer wrapping existing skills (`type: skill`) when a skill fits. Only invent a phase type when no skill covers the work.
|
|
75
|
-
|
|
76
|
-
- **`parallel-research`**: dispatch N researcher-style agents in parallel with independent briefs. Optionally run a synthesizer agent over their outputs to write one consolidated artifact. Use for tool surveys, option-space exploration, comparative research where work can be partitioned.
|
|
77
|
-
- **`crossref`**: given source A (the independent output the team just produced) and source B (existing material to validate against), dispatch an agent to produce a gap-and-overlap matrix plus a ranked list of validations and concerns. Use to compare independent work against prior research, canonical workspace context, or third-party material.
|
|
78
|
-
- **`skill`**: invoke an existing skill by name. The phase's `artifact:` path is the expected output location (or `null` for phases that commit to repo, like `executing-plans`). Use this for `brainstorming`, `writing-plans`, `executing-plans`, `test-driven-development`, or any other skill that fits the phase's intent.
|
|
79
|
-
|
|
80
|
-
If a new phase type appears to be needed, justify why no existing skill fits before adding it. New phase types are a maintenance cost. Prefer keeping `parallel-research` and `crossref` inline until cross-goal reuse pressure makes the extraction earn its keep.
|
|
81
|
-
|
|
82
|
-
## Agent-team dispatch pattern (v1: stateless)
|
|
83
|
-
|
|
84
|
-
A "team" is just a list of `Agent`-tool dispatch configs (`subagent_type` + `brief`). Spawned fresh each phase. No persistent identity, no memory across phases except what's written into artifacts.
|
|
85
|
-
|
|
86
|
-
The main agent's responsibility per phase:
|
|
87
|
-
|
|
88
|
-
1. Mark the phase `in_progress` in `goal-{topic}.md` frontmatter.
|
|
89
|
-
2. Dispatch the team via the `Agent` tool, in parallel where independent. The brief for each agent is exactly the `brief:` field from the phase config, prefixed with any context the brief itself doesn't already carry (typically a pointer to relevant sibling artifacts).
|
|
90
|
-
3. Collect agent outputs.
|
|
91
|
-
4. If a `synthesizer` is declared, dispatch it with the sibling outputs and have it write the `artifact:` file.
|
|
92
|
-
5. If no synthesizer, the main agent writes the `artifact:` file directly from the collected outputs.
|
|
93
|
-
6. Mark the phase `complete`, update `current_phase` to the next pending phase.
|
|
94
|
-
7. End the turn with a short status note (for the `/goal` evaluator) and, if `gate: review`, an explicit question to the user.
|
|
95
|
-
|
|
96
|
-
The artifact format supports adding a `persistent: true` flag on a phase team later if a future goal needs team continuity across phases. Don't add the field until something asks for it.
|
|
97
|
-
|
|
98
|
-
## Gate convention
|
|
99
|
-
|
|
100
|
-
`gate: review` is the default for every phase. At the gate, the main agent ends its turn by asking the user to approve, reject with feedback, or pause.
|
|
101
|
-
|
|
102
|
-
- **approve** → phase stays `complete`, advance to next phase next turn.
|
|
103
|
-
- **reject with feedback** → flip phase back to `pending` and re-run it with the feedback prepended to the team brief.
|
|
104
|
-
- **pause** → state is durable in `goal-{topic}.md`; resume with `claude --resume`.
|
|
105
|
-
|
|
106
|
-
The evaluator's "no, awaiting review" response after a gated phase does not bypass the wait for user input. It just keeps the session running until the user replies.
|
|
107
|
-
|
|
108
|
-
`gate: auto` is allowed in the format but discouraged in v1. Earn it after the workflow has been exercised at least once on the work in question.
|
|
109
|
-
|
|
110
|
-
## Gate-budget interaction and turn-budget sizing
|
|
111
|
-
|
|
112
|
-
A goal's `turn_budget` is a backstop that counts every orchestrator turn — including the short "awaiting your gate decision" exchanges that happen at every `gate: review` phase. The `/goal` evaluator re-pings the session whenever it tries to settle without the completion condition being met, and each ping consumes a turn. Concretely: a multi-phase gated goal whose author is away for stretches can spend a meaningful fraction of its budget *idling at gates* rather than advancing work. A run with six `gate: review` phases burned roughly half of a 150-turn budget on gate-idle pings before the actual work completed.
|
|
113
|
-
|
|
114
|
-
This is a structural property of `/goal` + review gates, not a per-goal accident. Plan for it:
|
|
115
|
-
|
|
116
|
-
- **Drop `gate: review` from early phases when the human is expected to be away for stretches.** Research, crossref, spec, and plan phases produce artifacts that the final session→main PR consumes anyway. The integration-branch model (`main` untouched until `/complete-work` opens the final PR for human review) already provides one strong human review point; piling per-phase gates on top of that, with no human watching, just burns budget. Set those phases to `gate: auto` when the run will be unattended.
|
|
117
|
-
- **Keep `gate: review` for phases with irreversible side effects or for phases whose outcome reshapes subsequent phases.** A spec gate that lets the human reprioritize the ranked execution list before `executing-plans` walks it is worth its cost; a research-synthesis gate that just rubber-stamps a matrix the human will see again in the final PR is not.
|
|
118
|
-
- **Size `turn_budget` for the worst case you actually expect.** If every phase is `gate: auto`: budget ≈ (estimated work-turns) × 1.2 (small headroom for retries). If any phases are `gate: review` and the human may be unavailable for hours: multiply the work-turn estimate by **2–3×** to absorb gate-idle pings, or raise the budget mid-goal by editing the goal artifact's `completion_condition` and `turn_budget` fields. The Stop-hook condition string is fixed from the original `/goal` invocation; raising `turn_budget` in the artifact keeps the auditable source-of-truth correct for resume but does not change the running evaluator's text — the user can re-run the (updated) `## Start command` to refresh it after `/goal clear`.
|
|
119
|
-
- **If the goal stops on the backstop mid-work because of gate-idle waste, that is the system working as designed.** `claude --resume` plus re-running the `## Start command` continues from durable phase state; phase artifacts and per-phase `status: complete` markers survive. Do not treat backstop-stop as a failure of the goal.
|
|
120
|
-
|
|
121
|
-
This guidance is workspace-side mitigation only. The underlying friction — the evaluator pinging during gate-idle — is `/goal` harness behavior, not workspace code. A cleaner fix (suspend the evaluator at review gates so it does not re-ping until a new user message arrives) is tracked separately and would obsolete the multiplier above when it lands.
|
|
122
|
-
|
|
123
|
-
## Integration branch and per-phase sub-PRs
|
|
124
|
-
|
|
125
|
-
While a `/goal`-driven session is running, the session branch (`feature/{session-name}`) acts as the goal's integration branch. Main is untouched until `/complete-work` opens the final session→main PR for human review. This is the key autonomy boundary: phase agents can merge their own work, repeatedly, throughout the goal — but only into the integration branch, never into main.
|
|
126
|
-
|
|
127
|
-
Two merge strategies, picked per phase:
|
|
128
|
-
|
|
129
|
-
- **Direct commit (default for artifact-only phases).** Phases of type `parallel-research` and `crossref`, and skill phases that wrap artifact-producing skills (`brainstorming`, `writing-plans`), commit their outputs directly to the session branch. These artifacts are stripped before the final PR anyway, so sub-PR ceremony adds nothing.
|
|
130
|
-
- **Sub-branch with self-merged PR (default for code phases).** Skill phases that wrap code-producing skills (`executing-plans`, `test-driven-development`, `subagent-driven-development`) work on a per-phase sub-branch and open a PR back to the session branch. After the phase's gate is satisfied, the main agent self-merges the sub-PR. Each sub-PR is a discrete reviewable unit — failures can be retried by rebuilding just the one sub-branch.
|
|
131
|
-
|
|
132
|
-
Sub-branch naming follows the existing `git-conventions.md` rule (kebab-case after prefix, no nesting): `feature/{session-name}-{phase-name}`. Examples for a session named `ulysses-goals`:
|
|
133
|
-
|
|
134
|
-
- `feature/ulysses-goals` — session/integration branch
|
|
135
|
-
- `feature/ulysses-goals-strategy-research` — phase sub-branch (if a research phase were promoted to sub-branch strategy)
|
|
136
|
-
- `feature/ulysses-goals-implement` — phase sub-branch for the implement phase
|
|
137
|
-
|
|
138
|
-
A phase declares its merge strategy in frontmatter via an optional `integration:` block:
|
|
139
|
-
|
|
140
|
-
```yaml
|
|
141
|
-
phases:
|
|
142
|
-
- name: implement
|
|
143
|
-
type: skill
|
|
144
|
-
skill: superpowers:executing-plans
|
|
145
|
-
integration:
|
|
146
|
-
strategy: sub-branch # direct | sub-branch
|
|
147
|
-
branch: feature/{session-name}-implement # explicit, or derived if omitted
|
|
148
|
-
self_merge: true # main agent merges the sub-PR after gate is satisfied; default true
|
|
149
|
-
```
|
|
150
|
-
|
|
151
|
-
If `integration:` is omitted, defaults are applied by phase type per the list above.
|
|
152
|
-
|
|
153
|
-
Multi-repo sessions: each project repo has its own session branch (same name across repos per existing convention). Phase sub-branches are created per-repo where the phase commits, with matching names. The sub-PR target in each repo is that repo's session branch.
|
|
154
|
-
|
|
155
|
-
`/complete-work` verifies before the final PR that every phase sub-branch has been merged into the session branch. If any are outstanding, it fails loudly with a list of unmerged sub-branches. The user resolves them (merge or close) before re-running.
|
|
156
|
-
|
|
157
|
-
## Model tiering
|
|
158
|
-
|
|
159
|
-
`/goal`-driven work uses a three-tier model assignment, biased toward the right tool for each shape of work:
|
|
160
|
-
|
|
161
|
-
- **Opus** for the main orchestration thread. It reads the goal artifact, dispatches phase teams, holds the long-running session context, and makes the per-phase advance/retry decisions. Reasoning-heavy and context-heavy work.
|
|
162
|
-
- **Sonnet** for phase team agents (researchers, crossref agents, skill subagents) and for synthesizers. Heavy lifting that runs in parallel and writes substantial artifacts. The default model for any agent in a `team:` block.
|
|
163
|
-
- **Haiku** for quick checks: file-existence verification, frontmatter validation, status sweeps, tiny lookups. Use when the work is bounded and obvious, and speed matters more than reasoning depth.
|
|
164
|
-
|
|
165
|
-
Phase frontmatter sets the model per agent via the `model:` field, which maps to the `Agent` tool's `model` parameter:
|
|
166
|
-
|
|
167
|
-
```yaml
|
|
168
|
-
team:
|
|
169
|
-
agents:
|
|
170
|
-
- subagent_type: researcher
|
|
171
|
-
model: sonnet # default for parallel-research; rarely overridden
|
|
172
|
-
brief: |
|
|
173
|
-
...
|
|
174
|
-
synthesizer:
|
|
175
|
-
subagent_type: researcher
|
|
176
|
-
model: sonnet
|
|
177
|
-
brief: |
|
|
178
|
-
...
|
|
179
|
-
```
|
|
180
|
-
|
|
181
|
-
For artifact-only or tagging work where Haiku is enough, declare it explicitly:
|
|
182
|
-
|
|
183
|
-
```yaml
|
|
184
|
-
- subagent_type: researcher
|
|
185
|
-
model: haiku
|
|
186
|
-
brief: |
|
|
187
|
-
Scan {paths} for files matching {pattern} and return a tagged list.
|
|
188
|
-
```
|
|
189
|
-
|
|
190
|
-
The main agent's model is determined by the harness (the user's `claude` invocation), not the goal artifact. Goal artifacts assume the main agent runs on Opus.
|
|
191
|
-
|
|
192
|
-
## Writing a good completion condition
|
|
193
|
-
|
|
194
|
-
The `/goal` evaluator runs after every turn against the conversation transcript. It does not call tools, so it can only judge what the main turn has surfaced. A good condition is:
|
|
195
|
-
|
|
196
|
-
- **Specific.** Names the artifacts that must exist (file paths, commit references, PR URLs) rather than vague outcomes.
|
|
197
|
-
- **Demonstrable from transcript.** The main agent's own output must be able to evidence completion. "The PR URL was reported in the transcript and `git status` showed clean" rather than "the work feels done."
|
|
198
|
-
- **Bounded.** Includes a turn budget as a backstop (e.g., "or stop after 60 turns") so the loop can't run away if something goes wrong.
|
|
199
|
-
- **Within the 4000-char limit.** Up to four kilobytes of condition text are accepted.
|
|
200
|
-
|
|
201
|
-
A reasonable template:
|
|
202
|
-
|
|
203
|
-
```
|
|
204
|
-
All phases in goal-<topic>.md show status: complete. Phase artifacts exist at: <list paths>. The /complete-work skill has produced release notes and opened the final PR; the PR URL appeared in the transcript. Or stop after <N> turns.
|
|
205
|
-
```
|
|
206
|
-
|
|
207
|
-
Fill in `<topic>`, paths, and `<N>` per goal. Anchor on artifacts and committed state, not on feelings.
|
|
208
|
-
|
|
209
|
-
## Kicking off the goal
|
|
210
|
-
|
|
211
|
-
`/goal` is a Claude Code built-in that the **user** types — the agent cannot invoke it. So the moment the artifact is ready is the load-bearing hand-off, and the artifact itself carries the instruction rather than relying on an agent chat message that vanishes on the next compaction or resume.
|
|
212
|
-
|
|
213
|
-
Every `goal-{topic}.md` body MUST include a `## Start command` section containing the literal, copy-paste-ready invocation:
|
|
214
|
-
|
|
215
|
-
````markdown
|
|
216
|
-
## Start command
|
|
217
|
-
|
|
218
|
-
```
|
|
219
|
-
/goal "All phases in goal-<topic>.md show status: complete. Phase artifacts exist at: <paths>. The /complete-work skill has produced release notes and opened the final PR; the PR URL appeared in the transcript. Or stop after <N> turns."
|
|
220
|
-
```
|
|
221
|
-
````
|
|
222
|
-
|
|
223
|
-
Rules for the `## Start command`:
|
|
224
|
-
|
|
225
|
-
- It is the `completion_condition` flattened to a **single line** and wrapped in `/goal "..."`. The frontmatter `completion_condition:` (a multi-line folded scalar) is the auditable source of truth; the start command is its runnable rendering. The two must express the same condition — if you edit one, re-derive the other.
|
|
226
|
-
- Flatten by collapsing the folded scalar's newlines to single spaces. Escape any embedded double quotes. Keep it within the 4000-character `/goal` limit.
|
|
227
|
-
- It lives in the body, not the frontmatter, because it is for a human to copy, not for machine parsing.
|
|
228
|
-
|
|
229
|
-
When the artifact is drafted and the user has reviewed it, the agent's hand-off is: point the user at the `## Start command` block and let them run it. Running it flips the goal from `status: pending` to `status: active` (the first `/goal` turn updates the frontmatter per the dispatch pattern). The agent never types `/goal` itself.
|
|
230
|
-
|
|
231
|
-
## Lifecycle integration
|
|
232
|
-
|
|
233
|
-
- `/goal` runs inside an active work session. It does NOT replace `/start-work`. The session is created the normal way, the goal artifact is drafted at the worktree top (including its `## Start command` block), and the user runs that block's `/goal "..."` command to kick off the loop.
|
|
234
|
-
- The goal artifact lives on the session branch and travels with `git push`. It survives across machines and `--resume`.
|
|
235
|
-
- `session.md`'s `## Tasks` should mirror the phase list at coarse grain (one task per phase) so `TodoWrite` shows high-level progress. The main agent updates `## Tasks` at phase transitions via the helper specified by the `task-list-mirroring` rule, in addition to updating `goal-{topic}.md`.
|
|
236
|
-
- `/pause-work` works without special handling. The goal-evaluator state resets on resume per the Claude Code docs; phase state is durable in the artifact.
|
|
237
|
-
- `/complete-work` reads `goal-*.md`, `research-*.md`, and `crossref-*.md` for release-note synthesis and strips them from the branch before the final PR, alongside the existing `design-*.md` and `plan-*.md` handling. When a goal artifact is present, it also runs a pre-flight check that every declared sub-branch (from phases with `integration.strategy: sub-branch`) has been merged into the session branch. Unmerged sub-branches abort completion with a clear list to resolve.
|
|
238
|
-
|
|
239
|
-
If a research or crossref artifact deserves to outlive the branch, the user runs `/promote` on it before `/complete-work`. `/promote` accepts arbitrary paths and routes them into `workspace-context/`.
|
|
240
|
-
|
|
241
|
-
## Tracker integration
|
|
242
|
-
|
|
243
|
-
Goals do not replace work items. A goal is execution shape; a work item is the unit the team tracks. See `work-item-tracking.md` for how `workItem:` in session frontmatter links a session to its tracker issue. A goal lives inside a session and therefore inherits the session's `workItem:`.
|
|
244
|
-
|
|
245
|
-
When the goal artifact is itself the deliverable for a future session to execute (i.e., this session's job was to *write* the goal, and another session will *run* it), the strip rule needs an escape hatch. The pattern: preserve the artifact in the tracker issue body (as a fenced code block) before `/complete-work` runs. The future session picks up the issue via `/start-work`, copies the artifact text into its worktree top, and runs `/goal`. The strip rule stays clean and consistent; the deliverable is preserved through the tracker.
|
|
246
|
-
|
|
247
|
-
## Out of scope
|
|
248
|
-
|
|
249
|
-
- Nested or sub-goals. v1 is flat.
|
|
250
|
-
- DAG between phases. Sequential at the phase level; parallelism only inside a phase via the team agents.
|
|
251
|
-
- Persistent named teams (`TeamCreate`). v1 is stateless dispatch.
|
|
252
|
-
- Frontmatter linter for `goal-*.md`. Manual review is fine for v1; revisit if workspaces using the template hit consistent shape errors.
|
|
253
|
-
- `/start-work` seeding of a `goal-{slug}.md` skeleton. Manual drafting is the v1 path. The drafting itself is high-leverage thinking; a template skeleton would risk skipping that.
|
|
254
|
-
|
|
255
|
-
## Appendix: worked example
|
|
256
|
-
|
|
257
|
-
A complete `goal-evaluate-rate-limiting.md` illustrating all three phase types. The topic is intentionally generic — the example is reference material, not prescriptive. (The appendix is fenced with four backticks so the example's own `## Start command` code block renders intact.)
|
|
258
|
-
|
|
259
|
-
````yaml
|
|
260
|
-
---
|
|
261
|
-
type: goal
|
|
262
|
-
topic: evaluate-rate-limiting
|
|
263
|
-
status: pending
|
|
264
|
-
current_phase: strategy-research
|
|
265
|
-
completion_condition: >
|
|
266
|
-
All 5 phases in goal-evaluate-rate-limiting.md show status: complete.
|
|
267
|
-
Phase artifacts exist at: research-rate-limiting-strategies.md,
|
|
268
|
-
crossref-existing-infrastructure.md, design-rate-limiting.md,
|
|
269
|
-
plan-rate-limiting.md, and the implementation commits land on the session
|
|
270
|
-
branch (visible in git log). The /complete-work skill has produced release
|
|
271
|
-
notes and opened the final PR; the PR URL appeared in the transcript.
|
|
272
|
-
Or stop after 60 turns.
|
|
273
|
-
turn_budget: 60
|
|
274
|
-
phases:
|
|
275
|
-
- name: strategy-research
|
|
276
|
-
type: parallel-research
|
|
277
|
-
status: pending
|
|
278
|
-
artifact: research-rate-limiting-strategies.md
|
|
279
|
-
gate: review
|
|
280
|
-
integration:
|
|
281
|
-
strategy: direct # artifact-only phase; commits straight to session branch
|
|
282
|
-
team:
|
|
283
|
-
agents:
|
|
284
|
-
- subagent_type: researcher
|
|
285
|
-
model: sonnet
|
|
286
|
-
brief: |
|
|
287
|
-
Research the token-bucket rate-limiting algorithm. Cover the
|
|
288
|
-
mechanics, parameter trade-offs (capacity, refill rate), edge
|
|
289
|
-
cases (burst behavior, clock skew), reference implementations in
|
|
290
|
-
popular libraries, and known production failure modes.
|
|
291
|
-
|
|
292
|
-
Output: a markdown report under 1,200 words. Return the content
|
|
293
|
-
in your response.
|
|
294
|
-
- subagent_type: researcher
|
|
295
|
-
model: sonnet
|
|
296
|
-
brief: |
|
|
297
|
-
Research the sliding-window rate-limiting algorithm. Cover both
|
|
298
|
-
the sliding log and sliding counter variants, accuracy
|
|
299
|
-
trade-offs, memory cost at scale, reference implementations, and
|
|
300
|
-
known production failure modes.
|
|
301
|
-
|
|
302
|
-
Output: a markdown report under 1,200 words. Return the content
|
|
303
|
-
in your response.
|
|
304
|
-
- subagent_type: researcher
|
|
305
|
-
model: sonnet
|
|
306
|
-
brief: |
|
|
307
|
-
Research the leaky-bucket rate-limiting algorithm. Cover the
|
|
308
|
-
queue-based and meter-based variants, smoothing behavior under
|
|
309
|
-
burst, comparison to token bucket, reference implementations,
|
|
310
|
-
and known production failure modes.
|
|
311
|
-
|
|
312
|
-
Output: a markdown report under 1,200 words. Return the content
|
|
313
|
-
in your response.
|
|
314
|
-
synthesizer:
|
|
315
|
-
subagent_type: researcher
|
|
316
|
-
model: sonnet
|
|
317
|
-
brief: |
|
|
318
|
-
Synthesize the three algorithm reports into a single
|
|
319
|
-
recommendation document at research-rate-limiting-strategies.md
|
|
320
|
-
(top of the active worktree).
|
|
321
|
-
|
|
322
|
-
Frontmatter: type: research, topic: rate-limiting-strategies,
|
|
323
|
-
state: ephemeral, lifecycle: active, confidence: medium,
|
|
324
|
-
updated: <today's date>.
|
|
325
|
-
|
|
326
|
-
Document structure:
|
|
327
|
-
- One-line recommendation up front
|
|
328
|
-
- Comparison table across criteria (accuracy, memory cost, burst
|
|
329
|
-
behavior, implementation complexity, operational debuggability)
|
|
330
|
-
- Per-algorithm summary with strengths and weaknesses
|
|
331
|
-
- Risks and unknowns
|
|
332
|
-
- References
|
|
333
|
-
|
|
334
|
-
Maximum 2,000 words.
|
|
335
|
-
|
|
336
|
-
- name: crossref-existing-infrastructure
|
|
337
|
-
type: crossref
|
|
338
|
-
status: pending
|
|
339
|
-
artifact: crossref-existing-infrastructure.md
|
|
340
|
-
inputs:
|
|
341
|
-
independent: research-rate-limiting-strategies.md
|
|
342
|
-
against:
|
|
343
|
-
- workspace-context/canonical.md
|
|
344
|
-
- repos/api-gateway/
|
|
345
|
-
gate: review
|
|
346
|
-
integration:
|
|
347
|
-
strategy: direct
|
|
348
|
-
agent:
|
|
349
|
-
subagent_type: researcher
|
|
350
|
-
model: sonnet
|
|
351
|
-
brief: |
|
|
352
|
-
Compare the rate-limiting strategy recommendation against the
|
|
353
|
-
existing infrastructure (the repos/api-gateway/ codebase and any
|
|
354
|
-
relevant canonical workspace context).
|
|
355
|
-
|
|
356
|
-
Produce a gap-and-overlap matrix: where does the recommended
|
|
357
|
-
strategy align with current patterns, where does it diverge, and
|
|
358
|
-
what migration friction is implied. Rank concerns by severity.
|
|
359
|
-
|
|
360
|
-
- name: spec
|
|
361
|
-
type: skill
|
|
362
|
-
status: pending
|
|
363
|
-
skill: superpowers:brainstorming
|
|
364
|
-
artifact: design-rate-limiting.md
|
|
365
|
-
gate: review
|
|
366
|
-
integration:
|
|
367
|
-
strategy: direct # spec lands as design-*.md at worktree top; stripped before final PR
|
|
368
|
-
|
|
369
|
-
- name: plan
|
|
370
|
-
type: skill
|
|
371
|
-
status: pending
|
|
372
|
-
skill: superpowers:writing-plans
|
|
373
|
-
artifact: plan-rate-limiting.md
|
|
374
|
-
gate: review
|
|
375
|
-
integration:
|
|
376
|
-
strategy: direct
|
|
377
|
-
|
|
378
|
-
- name: implement
|
|
379
|
-
type: skill
|
|
380
|
-
status: pending
|
|
381
|
-
skill: superpowers:executing-plans
|
|
382
|
-
artifact: null
|
|
383
|
-
gate: review
|
|
384
|
-
integration:
|
|
385
|
-
strategy: sub-branch # code phase: sub-branch + self-merged PR
|
|
386
|
-
branch: feature/{session-name}-implement
|
|
387
|
-
self_merge: true
|
|
388
|
-
---
|
|
389
|
-
|
|
390
|
-
# Goal: Evaluate and ship a rate-limiting strategy
|
|
391
|
-
|
|
392
|
-
The api-gateway needs rate limiting before the next traffic step-up. This
|
|
393
|
-
goal runs the full arc from candidate-algorithm research through implementation
|
|
394
|
-
on the session branch.
|
|
395
|
-
|
|
396
|
-
## Start command
|
|
397
|
-
|
|
398
|
-
```
|
|
399
|
-
/goal "All 5 phases in goal-evaluate-rate-limiting.md show status: complete. Phase artifacts exist at: research-rate-limiting-strategies.md, crossref-existing-infrastructure.md, design-rate-limiting.md, plan-rate-limiting.md, and the implementation commits land on the session branch (visible in git log). The /complete-work skill has produced release notes and opened the final PR; the PR URL appeared in the transcript. Or stop after 60 turns."
|
|
400
|
-
```
|
|
401
|
-
|
|
402
|
-
This is the frontmatter `completion_condition` flattened to one line. Run it after reviewing the artifact; it flips the goal to `status: active`.
|
|
403
|
-
|
|
404
|
-
## Per-phase intent
|
|
405
|
-
|
|
406
|
-
1. **strategy-research** runs three researchers in parallel, one per
|
|
407
|
-
candidate algorithm, plus a synthesizer that writes the comparative
|
|
408
|
-
recommendation.
|
|
409
|
-
|
|
410
|
-
2. **crossref-existing-infrastructure** validates the recommendation
|
|
411
|
-
against the current api-gateway codebase and canonical workspace
|
|
412
|
-
context, surfacing migration friction and divergences.
|
|
413
|
-
|
|
414
|
-
3. **spec** wraps `superpowers:brainstorming` to produce the design doc
|
|
415
|
-
from the validated recommendation.
|
|
416
|
-
|
|
417
|
-
4. **plan** wraps `superpowers:writing-plans` to produce the
|
|
418
|
-
implementation checklist.
|
|
419
|
-
|
|
420
|
-
5. **implement** wraps `superpowers:executing-plans` on a per-phase
|
|
421
|
-
sub-branch (`feature/{session-name}-implement`). The main agent opens
|
|
422
|
-
a PR from the sub-branch back into the session/integration branch and
|
|
423
|
-
self-merges once the gate is satisfied. Sub-branch is deleted post-merge.
|
|
424
|
-
|
|
425
|
-
Each gate is `review`: the user approves, rejects with feedback, or
|
|
426
|
-
pauses at the end of every phase. No auto-advance in v1.
|
|
427
|
-
|
|
428
|
-
All phase agents and synthesizers run on Sonnet (the default for team work).
|
|
429
|
-
The main agent reading and orchestrating this goal runs on Opus. Haiku is
|
|
430
|
-
unused in this example; it would be appropriate for a phase whose sole job
|
|
431
|
-
is, e.g., scanning a tree and returning a tagged file list.
|
|
432
|
-
````
|
|
21
|
+
Invoke it before drafting a goal artifact or running a goal phase. It is a skill rather than
|
|
22
|
+
a rule because it is a procedure needed by the small number of sessions that run `/goal`,
|
|
23
|
+
not a constraint every session must carry. Loading it on demand keeps roughly 27 KB out of
|
|
24
|
+
every session's always-loaded context.
|