@ulysses-ai/create-workspace 0.16.0-beta.1 → 0.18.0-beta.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +5 -5
- package/lib/init.mjs +19 -0
- package/package.json +1 -1
- package/template/.claude/hooks/_utils.mjs +1 -1
- package/template/.claude/hooks/repo-write-detection.mjs +161 -64
- package/template/.claude/hooks/session-end.mjs +68 -2
- package/template/.claude/hooks/session-start.mjs +35 -1
- package/template/.claude/hooks/subagent-start.mjs +89 -22
- package/template/.claude/lib/session-frontmatter.mjs +28 -0
- package/template/.claude/rules/coherent-revisions.md +1 -1
- package/template/.claude/rules/config-review.md.skip +29 -0
- package/template/.claude/rules/forge-operations.md +51 -0
- package/template/.claude/rules/git-conventions.md +16 -11
- package/template/.claude/rules/goal-driven-work.md +8 -403
- package/template/.claude/rules/honest-pushback.md +37 -37
- package/template/.claude/rules/memory-guidance.md +43 -90
- package/template/.claude/rules/superpowers-workflow.md.skip +1 -1
- package/template/.claude/rules/work-item-tracking.md +30 -72
- package/template/.claude/rules/workspace-structure.md +49 -69
- package/template/.claude/scripts/build-workspace-context.mjs +61 -16
- package/template/.claude/scripts/chat-record.mjs +282 -0
- package/template/.claude/scripts/cleanup-work-session.mjs +363 -36
- package/template/.claude/scripts/context-footprint.mjs +282 -0
- package/template/.claude/scripts/forges/github.mjs +255 -0
- package/template/.claude/scripts/forges/gitlab.mjs +20 -0
- package/template/.claude/scripts/forges/interface.mjs +125 -0
- package/template/.claude/scripts/generate-claude-local.mjs +21 -2
- package/template/.claude/scripts/migrate-sessions.mjs +1571 -0
- package/template/.claude/scripts/migrate-to-workspace-context.mjs +7 -2
- package/template/.claude/scripts/task-worktree.mjs +525 -0
- package/template/.claude/scripts/workspace-diagnostics.mjs +654 -0
- package/template/.claude/settings.json +5 -13
- package/template/.claude/skills/braindump/SKILL.md +11 -4
- package/template/.claude/skills/build-docs-site/SKILL.md +5 -5
- package/template/.claude/skills/build-docs-site/templates/spec.md.tmpl +1 -1
- package/template/.claude/skills/complete-work/SKILL.md +255 -215
- package/template/.claude/skills/context-placement/SKILL.md +199 -0
- package/template/.claude/skills/goal-driven-work/SKILL.md +459 -0
- package/template/.claude/skills/handoff/SKILL.md +11 -4
- package/template/.claude/skills/maintenance/SKILL.md +39 -6
- package/template/.claude/skills/migrate-sessions/SKILL.md +70 -0
- package/template/.claude/skills/pause-work/SKILL.md +33 -8
- package/template/.claude/skills/release/SKILL.md +44 -108
- package/template/.claude/skills/start-work/SKILL.md +89 -7
- package/template/.claude/skills/workspace-init/SKILL.md +34 -0
- package/template/.claude/skills/workspace-update/SKILL.md +4 -0
- package/template/.claudeignore +3 -0
- package/template/CLAUDE.md.tmpl +20 -2
- package/template/CODEBASE.md.tmpl +13 -0
- package/template/_gitignore +9 -0
- package/template/repo-claude.md.tmpl +10 -0
- package/template/workspace.json.tmpl +5 -3
- package/template/.claude/hooks/worktree-create.mjs +0 -53
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Config Review Cadence
|
|
2
|
+
|
|
3
|
+
Opt-in reminder to review `.claude` component files on a regular schedule. Activate by removing the `.skip` extension (`mv config-review.md.skip config-review.md`). Add it back to deactivate.
|
|
4
|
+
|
|
5
|
+
## When to review
|
|
6
|
+
|
|
7
|
+
Review the files in `.claude/rules/`, `.claude/skills/*/SKILL.md`, `.claude/agents/*.md`, and `.claude/hooks/*.mjs` after each major model release and at least every 180 days. Claude Code evolves quickly — conventions that matched platform behavior six months ago may no longer be accurate, optimal, or even meaningful.
|
|
8
|
+
|
|
9
|
+
The 180-day threshold is not arbitrary: it corresponds to roughly two major Claude model generations. Conventions written for an older model generation can silently mis-steer the newer one without anyone noticing, because the file still runs without error.
|
|
10
|
+
|
|
11
|
+
## Connection to /maintenance
|
|
12
|
+
|
|
13
|
+
The `/maintenance` skill's **Component age check** (step 7 in the Cleanup section) surfaces which files have drifted past 180 days by reading their frontmatter `updated:` field. Files without an `updated:` field are skipped — the check is incremental and only flags files that have opted in by carrying the field.
|
|
14
|
+
|
|
15
|
+
When `/maintenance` reports stale components, the recommended action is to open each flagged file, read it against the current Claude Code documentation and platform behavior, update the content where needed, and bump `updated:` to today.
|
|
16
|
+
|
|
17
|
+
## Why this ships as .skip
|
|
18
|
+
|
|
19
|
+
Review cadence is org-specific. A solo maintainer doing active weekly development may want a shorter cadence; a team with a slower release cycle may need a longer one; some workspaces may defer review entirely to a dedicated maintenance session. Making this rule mandatory by default would encode one team's preference as a universal constraint — exactly the failure mode described in `product-bias-risk.md`.
|
|
20
|
+
|
|
21
|
+
Activate the rule only if your team wants the reminder to appear in every session. If 180 days is wrong for your cadence, edit the threshold in the rule body after activating it.
|
|
22
|
+
|
|
23
|
+
## Activating
|
|
24
|
+
|
|
25
|
+
```
|
|
26
|
+
mv .claude/rules/config-review.md.skip .claude/rules/config-review.md
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Once active, this reminder loads into every session context. To deactivate, rename it back.
|
|
@@ -0,0 +1,51 @@
|
|
|
1
|
+
Activate this rule if the workspace creates PRs, watches CI runs, or interacts with releases from skills. Sibling to `work-item-tracking.md` (which covers issues); together they cover everything a workspace does against a code-hosting forge.
|
|
2
|
+
|
|
3
|
+
# Forge Operations
|
|
4
|
+
|
|
5
|
+
**Skills never call `gh` (or `glab`, or any forge CLI) inline for pull-request, release, or
|
|
6
|
+
workflow-run operations.** They go through the adapter at `.claude/scripts/forges/{type}.mjs`,
|
|
7
|
+
reached via `createForge()` from `.claude/scripts/forges/interface.mjs`.
|
|
8
|
+
|
|
9
|
+
The interface module is the source of truth for the available methods and their shapes —
|
|
10
|
+
read it when you need the API. `/complete-work` and `/pause-work` each carry the exact calls
|
|
11
|
+
they make, so in practice you rarely need to look.
|
|
12
|
+
|
|
13
|
+
## Why
|
|
14
|
+
|
|
15
|
+
Switching code hosts becomes a `workspace.json` field plus one adapter file instead of a
|
|
16
|
+
sweep across six skills. Failure modes share typed errors — `PrNotFound`, `MergeRejected`,
|
|
17
|
+
`WorkflowNotFound`, `ReleaseNotFound` — instead of every callsite parsing stderr. And the
|
|
18
|
+
adapter takes an injectable `spawnFn`, so tests mock subprocesses rather than running them.
|
|
19
|
+
|
|
20
|
+
## Configuration
|
|
21
|
+
|
|
22
|
+
`workspace.json` → `workspace.forge`: `{ "type": "github" }`. `type` names the adapter module;
|
|
23
|
+
`github` is the default and the only complete one, `gitlab.mjs` is a stub that throws
|
|
24
|
+
`NOT_IMPLEMENTED`. Optional `repo` is an `owner/name` slug; unset or `"auto"` resolves from the
|
|
25
|
+
git `origin` remote. An absent `workspace.forge` is treated as `{ type: 'github' }`; setting it
|
|
26
|
+
to `false` makes every adapter method throw `FORGE_DISABLED`.
|
|
27
|
+
|
|
28
|
+
## Deliberate exceptions — do not "fix" these
|
|
29
|
+
|
|
30
|
+
These stay as direct `gh` calls. Wrapping them would create a leaky abstraction, so leave them
|
|
31
|
+
alone:
|
|
32
|
+
|
|
33
|
+
- **Issue lifecycle** — issues, comments, labels and milestones belong to the tracker adapter.
|
|
34
|
+
See `work-item-tracking.md`. The two abstractions are intentionally separate.
|
|
35
|
+
- **`/setup-tracker` repo configuration** — `gh repo view --json hasIssuesEnabled` and
|
|
36
|
+
`gh api repos/{slug} -X PATCH -f has_issues=true` are GitHub-API-specific setup, not
|
|
37
|
+
cross-cutting operations. A GitLab user's setup flow differs entirely.
|
|
38
|
+
- **`gh repo view` as a remote-type probe** — a one-line capability check, not an operation.
|
|
39
|
+
- **`gh repo create`** — an interactive one-off when a workspace has no remote.
|
|
40
|
+
- **Manual recovery prose** — `gh run rerun`, `gh run view`, `gh release view` in `/release`
|
|
41
|
+
guidance are for an operator at a terminal, not for skill code.
|
|
42
|
+
|
|
43
|
+
## Boundaries
|
|
44
|
+
|
|
45
|
+
The adapter covers the operations the template's skills actually perform, not every `gh`
|
|
46
|
+
capability. New operations land as additive interface methods, never by a skill going around
|
|
47
|
+
the adapter. Forge-native features — PR comments, review webhooks, branch protection — remain
|
|
48
|
+
UI and direct-CLI territory.
|
|
49
|
+
|
|
50
|
+
Workspaces predating `workspace.forge` keep working, since `createForge(undefined)` defaults to
|
|
51
|
+
GitHub. `/maintenance` surfaces a notice suggesting the explicit value.
|
|
@@ -10,17 +10,22 @@
|
|
|
10
10
|
|
|
11
11
|
## Worktrees
|
|
12
12
|
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
-
|
|
13
|
+
Both lifecycles are built on worktrees; `workspace.sessionModel` in `workspace.json` routes new work to one of them.
|
|
14
|
+
|
|
15
|
+
**Task model** (`"task"`): one worktree per repo the task touches, created and removed with `.claude/scripts/task-worktree.mjs` (the chat stays at the workspace root):
|
|
16
|
+
|
|
17
|
+
- Project repos: `repos/{repo}/.claude/worktrees/{slug}/`, where `{slug}` is the branch with `/` replaced by `-`
|
|
18
|
+
- The workspace repo itself, addressed as `.`: `.claude/worktrees/{slug}/` — Claude Code's native worktree location
|
|
19
|
+
- Source clones at `repos/{repo}/` stay on their default branch; `/complete-work` tears each worktree down with `task-worktree.mjs --remove` (worktree first, then the branch)
|
|
20
|
+
|
|
21
|
+
**Session model** (default): N+1 worktrees in one self-contained folder at `work-sessions/{session-name}/`:
|
|
22
|
+
|
|
23
|
+
- `work-sessions/{session-name}/workspace/` — workspace worktree
|
|
24
|
+
- `work-sessions/{session-name}/workspace/repos/{repo-name}/` — project worktrees nested inside it (no symlink)
|
|
25
|
+
- Example, session `fix-auth` on `bugfix/fix-auth` touching `my-app` and `my-api`: `work-sessions/fix-auth/workspace/` plus `workspace/repos/my-app/` and `workspace/repos/my-api/`
|
|
26
|
+
- Teardown order is mandatory: project worktrees first, then the workspace worktree, then prune — the cleanup helper enforces it
|
|
27
|
+
|
|
28
|
+
The workspace `.gitignore` covers all of it: `repos` (no trailing slash) matches the root's `repos/` and every session worktree's nested `repos/`; `.claude/worktrees/` matches task worktrees of the workspace repo itself.
|
|
24
29
|
|
|
25
30
|
## Branch Maintenance
|
|
26
31
|
|
|
@@ -12,408 +12,13 @@ Use `/goal` when the work meets all three:
|
|
|
12
12
|
|
|
13
13
|
If any of the three fails, prefer plain session work or a single skill invocation. `/goal` is overhead. Pay it only when the work is long enough to earn it.
|
|
14
14
|
|
|
15
|
-
##
|
|
15
|
+
## The convention lives in a skill
|
|
16
16
|
|
|
17
|
-
|
|
18
|
-
-
|
|
19
|
-
-
|
|
20
|
-
- The artifact is tracked on the session branch and lives there until `/complete-work` runs. It is removed from the branch before the final PR alongside other session artifacts.
|
|
17
|
+
Everything past this decision — the `goal-{topic}.md` frontmatter schema, the three phase
|
|
18
|
+
types, agent-team dispatch, gate conventions, the integration-branch model with per-phase
|
|
19
|
+
sub-PRs, model tiering, and the worked example — is in the `goal-driven-work` skill.
|
|
21
20
|
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
type: goal
|
|
27
|
-
topic: <kebab-case-topic>
|
|
28
|
-
status: active # pending | active | complete | cancelled
|
|
29
|
-
# pending: artifact written but /goal not yet invoked
|
|
30
|
-
# active: /goal loop is running
|
|
31
|
-
# complete: all phases complete and condition met
|
|
32
|
-
# cancelled: goal abandoned without completion
|
|
33
|
-
current_phase: <phase-name> # the phase currently in progress or next up
|
|
34
|
-
completion_condition: > # mirrored from the `/goal` text so it survives `/goal clear` and is auditable
|
|
35
|
-
<multi-line condition string>
|
|
36
|
-
turn_budget: <int> # backstop; the condition itself should reference it
|
|
37
|
-
phases:
|
|
38
|
-
- name: <phase-name>
|
|
39
|
-
type: parallel-research | crossref | skill
|
|
40
|
-
status: pending # pending | in_progress | complete | failed
|
|
41
|
-
artifact: <path or null> # output artifact path, relative to worktree top; null for phases that commit to repo
|
|
42
|
-
gate: review | auto # default: review
|
|
43
|
-
|
|
44
|
-
# for type: parallel-research
|
|
45
|
-
team:
|
|
46
|
-
agents:
|
|
47
|
-
- subagent_type: <type>
|
|
48
|
-
brief: |
|
|
49
|
-
<multi-line brief>
|
|
50
|
-
- subagent_type: <type>
|
|
51
|
-
brief: |
|
|
52
|
-
<multi-line brief>
|
|
53
|
-
synthesizer:
|
|
54
|
-
subagent_type: <type>
|
|
55
|
-
brief: |
|
|
56
|
-
<multi-line brief: reads sibling agent outputs, writes the phase artifact>
|
|
57
|
-
|
|
58
|
-
# for type: crossref
|
|
59
|
-
inputs:
|
|
60
|
-
independent: <path to the just-produced independent artifact>
|
|
61
|
-
against: <path or list of paths to compare against>
|
|
62
|
-
brief: |
|
|
63
|
-
<multi-line brief for the crossref agent>
|
|
64
|
-
|
|
65
|
-
# for type: skill
|
|
66
|
-
skill: <plugin:skill-name> # e.g. superpowers:brainstorming
|
|
67
|
-
---
|
|
68
|
-
```
|
|
69
|
-
|
|
70
|
-
Body of the file is human-readable: the goal statement, success criteria, and per-phase intent in prose. The frontmatter is the source of truth for machine state; the body explains it to a reader.
|
|
71
|
-
|
|
72
|
-
## Phase types
|
|
73
|
-
|
|
74
|
-
Three types in v1. Prefer wrapping existing skills (`type: skill`) when a skill fits. Only invent a phase type when no skill covers the work.
|
|
75
|
-
|
|
76
|
-
- **`parallel-research`**: dispatch N researcher-style agents in parallel with independent briefs. Optionally run a synthesizer agent over their outputs to write one consolidated artifact. Use for tool surveys, option-space exploration, comparative research where work can be partitioned.
|
|
77
|
-
- **`crossref`**: given source A (the independent output the team just produced) and source B (existing material to validate against), dispatch an agent to produce a gap-and-overlap matrix plus a ranked list of validations and concerns. Use to compare independent work against prior research, canonical workspace context, or third-party material.
|
|
78
|
-
- **`skill`**: invoke an existing skill by name. The phase's `artifact:` path is the expected output location (or `null` for phases that commit to repo, like `executing-plans`). Use this for `brainstorming`, `writing-plans`, `executing-plans`, `test-driven-development`, or any other skill that fits the phase's intent.
|
|
79
|
-
|
|
80
|
-
If a new phase type appears to be needed, justify why no existing skill fits before adding it. New phase types are a maintenance cost. Prefer keeping `parallel-research` and `crossref` inline until cross-goal reuse pressure makes the extraction earn its keep.
|
|
81
|
-
|
|
82
|
-
## Agent-team dispatch pattern (v1: stateless)
|
|
83
|
-
|
|
84
|
-
A "team" is just a list of `Agent`-tool dispatch configs (`subagent_type` + `brief`). Spawned fresh each phase. No persistent identity, no memory across phases except what's written into artifacts.
|
|
85
|
-
|
|
86
|
-
The main agent's responsibility per phase:
|
|
87
|
-
|
|
88
|
-
1. Mark the phase `in_progress` in `goal-{topic}.md` frontmatter.
|
|
89
|
-
2. Dispatch the team via the `Agent` tool, in parallel where independent. The brief for each agent is exactly the `brief:` field from the phase config, prefixed with any context the brief itself doesn't already carry (typically a pointer to relevant sibling artifacts).
|
|
90
|
-
3. Collect agent outputs.
|
|
91
|
-
4. If a `synthesizer` is declared, dispatch it with the sibling outputs and have it write the `artifact:` file.
|
|
92
|
-
5. If no synthesizer, the main agent writes the `artifact:` file directly from the collected outputs.
|
|
93
|
-
6. Mark the phase `complete`, update `current_phase` to the next pending phase.
|
|
94
|
-
7. End the turn with a short status note (for the `/goal` evaluator) and, if `gate: review`, an explicit question to the user.
|
|
95
|
-
|
|
96
|
-
The artifact format supports adding a `persistent: true` flag on a phase team later if a future goal needs team continuity across phases. Don't add the field until something asks for it.
|
|
97
|
-
|
|
98
|
-
## Gate convention
|
|
99
|
-
|
|
100
|
-
`gate: review` is the default for every phase. At the gate, the main agent ends its turn by asking the user to approve, reject with feedback, or pause.
|
|
101
|
-
|
|
102
|
-
- **approve** → phase stays `complete`, advance to next phase next turn.
|
|
103
|
-
- **reject with feedback** → flip phase back to `pending` and re-run it with the feedback prepended to the team brief.
|
|
104
|
-
- **pause** → state is durable in `goal-{topic}.md`; resume with `claude --resume`.
|
|
105
|
-
|
|
106
|
-
The evaluator's "no, awaiting review" response after a gated phase does not bypass the wait for user input. It just keeps the session running until the user replies.
|
|
107
|
-
|
|
108
|
-
`gate: auto` is allowed in the format but discouraged in v1. Earn it after the workflow has been exercised at least once on the work in question.
|
|
109
|
-
|
|
110
|
-
## Integration branch and per-phase sub-PRs
|
|
111
|
-
|
|
112
|
-
While a `/goal`-driven session is running, the session branch (`feature/{session-name}`) acts as the goal's integration branch. Main is untouched until `/complete-work` opens the final session→main PR for human review. This is the key autonomy boundary: phase agents can merge their own work, repeatedly, throughout the goal — but only into the integration branch, never into main.
|
|
113
|
-
|
|
114
|
-
Two merge strategies, picked per phase:
|
|
115
|
-
|
|
116
|
-
- **Direct commit (default for artifact-only phases).** Phases of type `parallel-research` and `crossref`, and skill phases that wrap artifact-producing skills (`brainstorming`, `writing-plans`), commit their outputs directly to the session branch. These artifacts are stripped before the final PR anyway, so sub-PR ceremony adds nothing.
|
|
117
|
-
- **Sub-branch with self-merged PR (default for code phases).** Skill phases that wrap code-producing skills (`executing-plans`, `test-driven-development`, `subagent-driven-development`) work on a per-phase sub-branch and open a PR back to the session branch. After the phase's gate is satisfied, the main agent self-merges the sub-PR. Each sub-PR is a discrete reviewable unit — failures can be retried by rebuilding just the one sub-branch.
|
|
118
|
-
|
|
119
|
-
Sub-branch naming follows the existing `git-conventions.md` rule (kebab-case after prefix, no nesting): `feature/{session-name}-{phase-name}`. Examples for a session named `ulysses-goals`:
|
|
120
|
-
|
|
121
|
-
- `feature/ulysses-goals` — session/integration branch
|
|
122
|
-
- `feature/ulysses-goals-strategy-research` — phase sub-branch (if a research phase were promoted to sub-branch strategy)
|
|
123
|
-
- `feature/ulysses-goals-implement` — phase sub-branch for the implement phase
|
|
124
|
-
|
|
125
|
-
A phase declares its merge strategy in frontmatter via an optional `integration:` block:
|
|
126
|
-
|
|
127
|
-
```yaml
|
|
128
|
-
phases:
|
|
129
|
-
- name: implement
|
|
130
|
-
type: skill
|
|
131
|
-
skill: superpowers:executing-plans
|
|
132
|
-
integration:
|
|
133
|
-
strategy: sub-branch # direct | sub-branch
|
|
134
|
-
branch: feature/{session-name}-implement # explicit, or derived if omitted
|
|
135
|
-
self_merge: true # main agent merges the sub-PR after gate is satisfied; default true
|
|
136
|
-
```
|
|
137
|
-
|
|
138
|
-
If `integration:` is omitted, defaults are applied by phase type per the list above.
|
|
139
|
-
|
|
140
|
-
Multi-repo sessions: each project repo has its own session branch (same name across repos per existing convention). Phase sub-branches are created per-repo where the phase commits, with matching names. The sub-PR target in each repo is that repo's session branch.
|
|
141
|
-
|
|
142
|
-
`/complete-work` verifies before the final PR that every phase sub-branch has been merged into the session branch. If any are outstanding, it fails loudly with a list of unmerged sub-branches. The user resolves them (merge or close) before re-running.
|
|
143
|
-
|
|
144
|
-
## Model tiering
|
|
145
|
-
|
|
146
|
-
`/goal`-driven work uses a three-tier model assignment, biased toward the right tool for each shape of work:
|
|
147
|
-
|
|
148
|
-
- **Opus** for the main orchestration thread. It reads the goal artifact, dispatches phase teams, holds the long-running session context, and makes the per-phase advance/retry decisions. Reasoning-heavy and context-heavy work.
|
|
149
|
-
- **Sonnet** for phase team agents (researchers, crossref agents, skill subagents) and for synthesizers. Heavy lifting that runs in parallel and writes substantial artifacts. The default model for any agent in a `team:` block.
|
|
150
|
-
- **Haiku** for quick checks: file-existence verification, frontmatter validation, status sweeps, tiny lookups. Use when the work is bounded and obvious, and speed matters more than reasoning depth.
|
|
151
|
-
|
|
152
|
-
Phase frontmatter sets the model per agent via the `model:` field, which maps to the `Agent` tool's `model` parameter:
|
|
153
|
-
|
|
154
|
-
```yaml
|
|
155
|
-
team:
|
|
156
|
-
agents:
|
|
157
|
-
- subagent_type: researcher
|
|
158
|
-
model: sonnet # default for parallel-research; rarely overridden
|
|
159
|
-
brief: |
|
|
160
|
-
...
|
|
161
|
-
synthesizer:
|
|
162
|
-
subagent_type: researcher
|
|
163
|
-
model: sonnet
|
|
164
|
-
brief: |
|
|
165
|
-
...
|
|
166
|
-
```
|
|
167
|
-
|
|
168
|
-
For artifact-only or tagging work where Haiku is enough, declare it explicitly:
|
|
169
|
-
|
|
170
|
-
```yaml
|
|
171
|
-
- subagent_type: researcher
|
|
172
|
-
model: haiku
|
|
173
|
-
brief: |
|
|
174
|
-
Scan {paths} for files matching {pattern} and return a tagged list.
|
|
175
|
-
```
|
|
176
|
-
|
|
177
|
-
The main agent's model is determined by the harness (the user's `claude` invocation), not the goal artifact. Goal artifacts assume the main agent runs on Opus.
|
|
178
|
-
|
|
179
|
-
## Writing a good completion condition
|
|
180
|
-
|
|
181
|
-
The `/goal` evaluator runs after every turn against the conversation transcript. It does not call tools, so it can only judge what the main turn has surfaced. A good condition is:
|
|
182
|
-
|
|
183
|
-
- **Specific.** Names the artifacts that must exist (file paths, commit references, PR URLs) rather than vague outcomes.
|
|
184
|
-
- **Demonstrable from transcript.** The main agent's own output must be able to evidence completion. "The PR URL was reported in the transcript and `git status` showed clean" rather than "the work feels done."
|
|
185
|
-
- **Bounded.** Includes a turn budget as a backstop (e.g., "or stop after 60 turns") so the loop can't run away if something goes wrong.
|
|
186
|
-
- **Within the 4000-char limit.** Up to four kilobytes of condition text are accepted.
|
|
187
|
-
|
|
188
|
-
A reasonable template:
|
|
189
|
-
|
|
190
|
-
```
|
|
191
|
-
All phases in goal-<topic>.md show status: complete. Phase artifacts exist at: <list paths>. The /complete-work skill has produced release notes and opened the final PR; the PR URL appeared in the transcript. Or stop after <N> turns.
|
|
192
|
-
```
|
|
193
|
-
|
|
194
|
-
Fill in `<topic>`, paths, and `<N>` per goal. Anchor on artifacts and committed state, not on feelings.
|
|
195
|
-
|
|
196
|
-
## Kicking off the goal
|
|
197
|
-
|
|
198
|
-
`/goal` is a Claude Code built-in that the **user** types — the agent cannot invoke it. So the moment the artifact is ready is the load-bearing hand-off, and the artifact itself carries the instruction rather than relying on an agent chat message that vanishes on the next compaction or resume.
|
|
199
|
-
|
|
200
|
-
Every `goal-{topic}.md` body MUST include a `## Start command` section containing the literal, copy-paste-ready invocation:
|
|
201
|
-
|
|
202
|
-
````markdown
|
|
203
|
-
## Start command
|
|
204
|
-
|
|
205
|
-
```
|
|
206
|
-
/goal "All phases in goal-<topic>.md show status: complete. Phase artifacts exist at: <paths>. The /complete-work skill has produced release notes and opened the final PR; the PR URL appeared in the transcript. Or stop after <N> turns."
|
|
207
|
-
```
|
|
208
|
-
````
|
|
209
|
-
|
|
210
|
-
Rules for the `## Start command`:
|
|
211
|
-
|
|
212
|
-
- It is the `completion_condition` flattened to a **single line** and wrapped in `/goal "..."`. The frontmatter `completion_condition:` (a multi-line folded scalar) is the auditable source of truth; the start command is its runnable rendering. The two must express the same condition — if you edit one, re-derive the other.
|
|
213
|
-
- Flatten by collapsing the folded scalar's newlines to single spaces. Escape any embedded double quotes. Keep it within the 4000-character `/goal` limit.
|
|
214
|
-
- It lives in the body, not the frontmatter, because it is for a human to copy, not for machine parsing.
|
|
215
|
-
|
|
216
|
-
When the artifact is drafted and the user has reviewed it, the agent's hand-off is: point the user at the `## Start command` block and let them run it. Running it flips the goal from `status: pending` to `status: active` (the first `/goal` turn updates the frontmatter per the dispatch pattern). The agent never types `/goal` itself.
|
|
217
|
-
|
|
218
|
-
## Lifecycle integration
|
|
219
|
-
|
|
220
|
-
- `/goal` runs inside an active work session. It does NOT replace `/start-work`. The session is created the normal way, the goal artifact is drafted at the worktree top (including its `## Start command` block), and the user runs that block's `/goal "..."` command to kick off the loop.
|
|
221
|
-
- The goal artifact lives on the session branch and travels with `git push`. It survives across machines and `--resume`.
|
|
222
|
-
- `session.md`'s `## Tasks` should mirror the phase list at coarse grain (one task per phase) so `TodoWrite` shows high-level progress. The main agent updates `## Tasks` at phase transitions via the helper specified by the `task-list-mirroring` rule, in addition to updating `goal-{topic}.md`.
|
|
223
|
-
- `/pause-work` works without special handling. The goal-evaluator state resets on resume per the Claude Code docs; phase state is durable in the artifact.
|
|
224
|
-
- `/complete-work` reads `goal-*.md`, `research-*.md`, and `crossref-*.md` for release-note synthesis and strips them from the branch before the final PR, alongside the existing `design-*.md` and `plan-*.md` handling. When a goal artifact is present, it also runs a pre-flight check that every declared sub-branch (from phases with `integration.strategy: sub-branch`) has been merged into the session branch. Unmerged sub-branches abort completion with a clear list to resolve.
|
|
225
|
-
|
|
226
|
-
If a research or crossref artifact deserves to outlive the branch, the user runs `/promote` on it before `/complete-work`. `/promote` accepts arbitrary paths and routes them into `workspace-context/`.
|
|
227
|
-
|
|
228
|
-
## Tracker integration
|
|
229
|
-
|
|
230
|
-
Goals do not replace work items. A goal is execution shape; a work item is the unit the team tracks. See `work-item-tracking.md` for how `workItem:` in session frontmatter links a session to its tracker issue. A goal lives inside a session and therefore inherits the session's `workItem:`.
|
|
231
|
-
|
|
232
|
-
When the goal artifact is itself the deliverable for a future session to execute (i.e., this session's job was to *write* the goal, and another session will *run* it), the strip rule needs an escape hatch. The pattern: preserve the artifact in the tracker issue body (as a fenced code block) before `/complete-work` runs. The future session picks up the issue via `/start-work`, copies the artifact text into its worktree top, and runs `/goal`. The strip rule stays clean and consistent; the deliverable is preserved through the tracker.
|
|
233
|
-
|
|
234
|
-
## Out of scope
|
|
235
|
-
|
|
236
|
-
- Nested or sub-goals. v1 is flat.
|
|
237
|
-
- DAG between phases. Sequential at the phase level; parallelism only inside a phase via the team agents.
|
|
238
|
-
- Persistent named teams (`TeamCreate`). v1 is stateless dispatch.
|
|
239
|
-
- Frontmatter linter for `goal-*.md`. Manual review is fine for v1; revisit if workspaces using the template hit consistent shape errors.
|
|
240
|
-
- `/start-work` seeding of a `goal-{slug}.md` skeleton. Manual drafting is the v1 path. The drafting itself is high-leverage thinking; a template skeleton would risk skipping that.
|
|
241
|
-
|
|
242
|
-
## Appendix: worked example
|
|
243
|
-
|
|
244
|
-
A complete `goal-evaluate-rate-limiting.md` illustrating all three phase types. The topic is intentionally generic — the example is reference material, not prescriptive. (The appendix is fenced with four backticks so the example's own `## Start command` code block renders intact.)
|
|
245
|
-
|
|
246
|
-
````yaml
|
|
247
|
-
---
|
|
248
|
-
type: goal
|
|
249
|
-
topic: evaluate-rate-limiting
|
|
250
|
-
status: pending
|
|
251
|
-
current_phase: strategy-research
|
|
252
|
-
completion_condition: >
|
|
253
|
-
All 5 phases in goal-evaluate-rate-limiting.md show status: complete.
|
|
254
|
-
Phase artifacts exist at: research-rate-limiting-strategies.md,
|
|
255
|
-
crossref-existing-infrastructure.md, design-rate-limiting.md,
|
|
256
|
-
plan-rate-limiting.md, and the implementation commits land on the session
|
|
257
|
-
branch (visible in git log). The /complete-work skill has produced release
|
|
258
|
-
notes and opened the final PR; the PR URL appeared in the transcript.
|
|
259
|
-
Or stop after 60 turns.
|
|
260
|
-
turn_budget: 60
|
|
261
|
-
phases:
|
|
262
|
-
- name: strategy-research
|
|
263
|
-
type: parallel-research
|
|
264
|
-
status: pending
|
|
265
|
-
artifact: research-rate-limiting-strategies.md
|
|
266
|
-
gate: review
|
|
267
|
-
integration:
|
|
268
|
-
strategy: direct # artifact-only phase; commits straight to session branch
|
|
269
|
-
team:
|
|
270
|
-
agents:
|
|
271
|
-
- subagent_type: researcher
|
|
272
|
-
model: sonnet
|
|
273
|
-
brief: |
|
|
274
|
-
Research the token-bucket rate-limiting algorithm. Cover the
|
|
275
|
-
mechanics, parameter trade-offs (capacity, refill rate), edge
|
|
276
|
-
cases (burst behavior, clock skew), reference implementations in
|
|
277
|
-
popular libraries, and known production failure modes.
|
|
278
|
-
|
|
279
|
-
Output: a markdown report under 1,200 words. Return the content
|
|
280
|
-
in your response.
|
|
281
|
-
- subagent_type: researcher
|
|
282
|
-
model: sonnet
|
|
283
|
-
brief: |
|
|
284
|
-
Research the sliding-window rate-limiting algorithm. Cover both
|
|
285
|
-
the sliding log and sliding counter variants, accuracy
|
|
286
|
-
trade-offs, memory cost at scale, reference implementations, and
|
|
287
|
-
known production failure modes.
|
|
288
|
-
|
|
289
|
-
Output: a markdown report under 1,200 words. Return the content
|
|
290
|
-
in your response.
|
|
291
|
-
- subagent_type: researcher
|
|
292
|
-
model: sonnet
|
|
293
|
-
brief: |
|
|
294
|
-
Research the leaky-bucket rate-limiting algorithm. Cover the
|
|
295
|
-
queue-based and meter-based variants, smoothing behavior under
|
|
296
|
-
burst, comparison to token bucket, reference implementations,
|
|
297
|
-
and known production failure modes.
|
|
298
|
-
|
|
299
|
-
Output: a markdown report under 1,200 words. Return the content
|
|
300
|
-
in your response.
|
|
301
|
-
synthesizer:
|
|
302
|
-
subagent_type: researcher
|
|
303
|
-
model: sonnet
|
|
304
|
-
brief: |
|
|
305
|
-
Synthesize the three algorithm reports into a single
|
|
306
|
-
recommendation document at research-rate-limiting-strategies.md
|
|
307
|
-
(top of the active worktree).
|
|
308
|
-
|
|
309
|
-
Frontmatter: type: research, topic: rate-limiting-strategies,
|
|
310
|
-
state: ephemeral, lifecycle: active, confidence: medium,
|
|
311
|
-
updated: <today's date>.
|
|
312
|
-
|
|
313
|
-
Document structure:
|
|
314
|
-
- One-line recommendation up front
|
|
315
|
-
- Comparison table across criteria (accuracy, memory cost, burst
|
|
316
|
-
behavior, implementation complexity, operational debuggability)
|
|
317
|
-
- Per-algorithm summary with strengths and weaknesses
|
|
318
|
-
- Risks and unknowns
|
|
319
|
-
- References
|
|
320
|
-
|
|
321
|
-
Maximum 2,000 words.
|
|
322
|
-
|
|
323
|
-
- name: crossref-existing-infrastructure
|
|
324
|
-
type: crossref
|
|
325
|
-
status: pending
|
|
326
|
-
artifact: crossref-existing-infrastructure.md
|
|
327
|
-
inputs:
|
|
328
|
-
independent: research-rate-limiting-strategies.md
|
|
329
|
-
against:
|
|
330
|
-
- workspace-context/canonical.md
|
|
331
|
-
- repos/api-gateway/
|
|
332
|
-
gate: review
|
|
333
|
-
integration:
|
|
334
|
-
strategy: direct
|
|
335
|
-
agent:
|
|
336
|
-
subagent_type: researcher
|
|
337
|
-
model: sonnet
|
|
338
|
-
brief: |
|
|
339
|
-
Compare the rate-limiting strategy recommendation against the
|
|
340
|
-
existing infrastructure (the repos/api-gateway/ codebase and any
|
|
341
|
-
relevant canonical workspace context).
|
|
342
|
-
|
|
343
|
-
Produce a gap-and-overlap matrix: where does the recommended
|
|
344
|
-
strategy align with current patterns, where does it diverge, and
|
|
345
|
-
what migration friction is implied. Rank concerns by severity.
|
|
346
|
-
|
|
347
|
-
- name: spec
|
|
348
|
-
type: skill
|
|
349
|
-
status: pending
|
|
350
|
-
skill: superpowers:brainstorming
|
|
351
|
-
artifact: design-rate-limiting.md
|
|
352
|
-
gate: review
|
|
353
|
-
integration:
|
|
354
|
-
strategy: direct # spec lands as design-*.md at worktree top; stripped before final PR
|
|
355
|
-
|
|
356
|
-
- name: plan
|
|
357
|
-
type: skill
|
|
358
|
-
status: pending
|
|
359
|
-
skill: superpowers:writing-plans
|
|
360
|
-
artifact: plan-rate-limiting.md
|
|
361
|
-
gate: review
|
|
362
|
-
integration:
|
|
363
|
-
strategy: direct
|
|
364
|
-
|
|
365
|
-
- name: implement
|
|
366
|
-
type: skill
|
|
367
|
-
status: pending
|
|
368
|
-
skill: superpowers:executing-plans
|
|
369
|
-
artifact: null
|
|
370
|
-
gate: review
|
|
371
|
-
integration:
|
|
372
|
-
strategy: sub-branch # code phase: sub-branch + self-merged PR
|
|
373
|
-
branch: feature/{session-name}-implement
|
|
374
|
-
self_merge: true
|
|
375
|
-
---
|
|
376
|
-
|
|
377
|
-
# Goal: Evaluate and ship a rate-limiting strategy
|
|
378
|
-
|
|
379
|
-
The api-gateway needs rate limiting before the next traffic step-up. This
|
|
380
|
-
goal runs the full arc from candidate-algorithm research through implementation
|
|
381
|
-
on the session branch.
|
|
382
|
-
|
|
383
|
-
## Start command
|
|
384
|
-
|
|
385
|
-
```
|
|
386
|
-
/goal "All 5 phases in goal-evaluate-rate-limiting.md show status: complete. Phase artifacts exist at: research-rate-limiting-strategies.md, crossref-existing-infrastructure.md, design-rate-limiting.md, plan-rate-limiting.md, and the implementation commits land on the session branch (visible in git log). The /complete-work skill has produced release notes and opened the final PR; the PR URL appeared in the transcript. Or stop after 60 turns."
|
|
387
|
-
```
|
|
388
|
-
|
|
389
|
-
This is the frontmatter `completion_condition` flattened to one line. Run it after reviewing the artifact; it flips the goal to `status: active`.
|
|
390
|
-
|
|
391
|
-
## Per-phase intent
|
|
392
|
-
|
|
393
|
-
1. **strategy-research** runs three researchers in parallel, one per
|
|
394
|
-
candidate algorithm, plus a synthesizer that writes the comparative
|
|
395
|
-
recommendation.
|
|
396
|
-
|
|
397
|
-
2. **crossref-existing-infrastructure** validates the recommendation
|
|
398
|
-
against the current api-gateway codebase and canonical workspace
|
|
399
|
-
context, surfacing migration friction and divergences.
|
|
400
|
-
|
|
401
|
-
3. **spec** wraps `superpowers:brainstorming` to produce the design doc
|
|
402
|
-
from the validated recommendation.
|
|
403
|
-
|
|
404
|
-
4. **plan** wraps `superpowers:writing-plans` to produce the
|
|
405
|
-
implementation checklist.
|
|
406
|
-
|
|
407
|
-
5. **implement** wraps `superpowers:executing-plans` on a per-phase
|
|
408
|
-
sub-branch (`feature/{session-name}-implement`). The main agent opens
|
|
409
|
-
a PR from the sub-branch back into the session/integration branch and
|
|
410
|
-
self-merges once the gate is satisfied. Sub-branch is deleted post-merge.
|
|
411
|
-
|
|
412
|
-
Each gate is `review`: the user approves, rejects with feedback, or
|
|
413
|
-
pauses at the end of every phase. No auto-advance in v1.
|
|
414
|
-
|
|
415
|
-
All phase agents and synthesizers run on Sonnet (the default for team work).
|
|
416
|
-
The main agent reading and orchestrating this goal runs on Opus. Haiku is
|
|
417
|
-
unused in this example; it would be appropriate for a phase whose sole job
|
|
418
|
-
is, e.g., scanning a tree and returning a tagged file list.
|
|
419
|
-
````
|
|
21
|
+
Invoke it before drafting a goal artifact or running a goal phase. It is a skill rather than
|
|
22
|
+
a rule because it is a procedure needed by the small number of sessions that run `/goal`,
|
|
23
|
+
not a constraint every session must carry. Loading it on demand keeps roughly 27 KB out of
|
|
24
|
+
every session's always-loaded context.
|
|
@@ -1,56 +1,56 @@
|
|
|
1
1
|
# Honest Pushback
|
|
2
2
|
|
|
3
|
-
Do not agree
|
|
3
|
+
Do not agree to be agreeable. Do not keep trying things that aren't working. Do not assume
|
|
4
|
+
when you can verify. Challenge assumptions and flag concerns even when the user is
|
|
5
|
+
enthusiastic.
|
|
4
6
|
|
|
5
|
-
## What
|
|
7
|
+
## What this means
|
|
6
8
|
|
|
7
|
-
- If an approach has obvious downsides, say so before implementing
|
|
8
|
-
- If a
|
|
9
|
+
- If an approach has obvious downsides, say so before implementing.
|
|
10
|
+
- If a decision contradicts an earlier one, name the contradiction.
|
|
9
11
|
- If scope is creeping, name it: "This started as X but is becoming Y. Split?"
|
|
10
|
-
- If you don't know
|
|
11
|
-
- If the
|
|
12
|
-
- If you made a mistake, own it plainly
|
|
12
|
+
- If you don't know, say so — don't fabricate confidence.
|
|
13
|
+
- If the idea is good, "that works" is enough. No embellishment.
|
|
14
|
+
- If you made a mistake, own it plainly. No hedging.
|
|
13
15
|
|
|
14
|
-
## No
|
|
16
|
+
## No retry loops
|
|
15
17
|
|
|
16
|
-
|
|
18
|
+
If a fix produced the same error or an unexpected result twice, stop. Do not try a variation
|
|
19
|
+
of the same approach. Instead:
|
|
17
20
|
|
|
18
|
-
1. **State
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
21
|
+
1. **State expected vs actual**, specifically — not "it didn't work" but "expected 200, got
|
|
22
|
+
403 with message X".
|
|
23
|
+
2. **Name the failing assumption.** What is surprising, and why?
|
|
24
|
+
3. **Research it.** Read the docs, search the error, read the source. Use web search if local
|
|
25
|
+
sources don't explain it.
|
|
26
|
+
4. **Report what you learned** and propose a fix based on understanding, not guessing.
|
|
22
27
|
|
|
23
|
-
This
|
|
28
|
+
This stops the cycle of trying broken variations when reading the docs would take one turn.
|
|
24
29
|
|
|
25
|
-
## Verify,
|
|
30
|
+
## Verify, don't assume
|
|
26
31
|
|
|
27
|
-
When evidence is available
|
|
32
|
+
When evidence is available, check it before proceeding. Logs, the database, the actual UI,
|
|
33
|
+
runtime state, a real API call — whichever settles the question. If the logs aren't verbose
|
|
34
|
+
enough, add instrumentation, run it, read the output, remove it.
|
|
28
35
|
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
- **UI/browser** — test the actual behavior instead of predicting what the user will see. Use browser tools, take screenshots, inspect network requests.
|
|
33
|
-
- **Runtime state** — add a console.log, print statement, or debugger breakpoint. Run it. Read the output.
|
|
34
|
-
- **API responses** — make the actual call instead of assuming the response shape.
|
|
36
|
+
The tell is reaching for "I think the issue is…", "probably", or "likely" about something
|
|
37
|
+
you could check in one step. Reasoning about what a function returns when you could call it
|
|
38
|
+
is the same mistake.
|
|
35
39
|
|
|
36
|
-
**
|
|
40
|
+
**Ask once**, then stop asking: "I want to verify {what} by {how}. Go ahead, or should I
|
|
41
|
+
just check without asking each time?" If the user says just check, verify proactively for
|
|
42
|
+
the rest of the session. Asking once is polite; asking every time is friction.
|
|
37
43
|
|
|
38
|
-
|
|
44
|
+
Use judgment — don't over-verify the trivial.
|
|
39
45
|
|
|
40
|
-
|
|
41
|
-
- You're about to say "I think the issue is..." when you could check
|
|
42
|
-
- You're reasoning about what a function returns when you could call it
|
|
43
|
-
- You're guessing at database state when you could query it
|
|
44
|
-
- You're predicting UI behavior when you could test it
|
|
45
|
-
- You catch yourself writing "probably" or "likely" about something verifiable
|
|
46
|
+
## What this does not mean
|
|
46
47
|
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
- Don't refuse to execute — voice the concern, then follow the user's decision
|
|
51
|
-
- Don't lecture — state the issue once, clearly, and move on
|
|
52
|
-
- Don't over-verify trivial things — use judgment about what's worth checking
|
|
48
|
+
Don't be contrarian as a personality trait; push back where there is substance. Don't refuse
|
|
49
|
+
to execute — voice the concern, then follow the user's decision. Don't lecture: state it
|
|
50
|
+
once and move on.
|
|
53
51
|
|
|
54
52
|
## Why
|
|
55
53
|
|
|
56
|
-
|
|
54
|
+
Sycophancy wastes time and lets bad decisions through. Retry loops burn tokens. Assumptions
|
|
55
|
+
that could have been checked cascade into wrong decisions. A useful collaborator says when
|
|
56
|
+
something is off, stops when it isn't working, and finds out why before trying again.
|