@crewbie/cli 0.1.0-alpha.49
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +253 -0
- package/dist/cli.d.ts +2 -0
- package/dist/cli.js +506 -0
- package/dist/cli.js.map +1 -0
- package/dist/config.d.ts +84 -0
- package/dist/config.js +178 -0
- package/dist/config.js.map +1 -0
- package/dist/core.d.ts +31 -0
- package/dist/core.js +186 -0
- package/dist/core.js.map +1 -0
- package/dist/execution/attribution.d.ts +10 -0
- package/dist/execution/attribution.js +102 -0
- package/dist/execution/attribution.js.map +1 -0
- package/dist/execution/cancel.d.ts +3 -0
- package/dist/execution/cancel.js +59 -0
- package/dist/execution/cancel.js.map +1 -0
- package/dist/execution/capabilities.d.ts +22 -0
- package/dist/execution/capabilities.js +101 -0
- package/dist/execution/capabilities.js.map +1 -0
- package/dist/execution/controls.d.ts +51 -0
- package/dist/execution/controls.js +257 -0
- package/dist/execution/controls.js.map +1 -0
- package/dist/execution/dispatch.d.ts +69 -0
- package/dist/execution/dispatch.js +831 -0
- package/dist/execution/dispatch.js.map +1 -0
- package/dist/execution/fix.d.ts +3 -0
- package/dist/execution/fix.js +267 -0
- package/dist/execution/fix.js.map +1 -0
- package/dist/execution/github.d.ts +15 -0
- package/dist/execution/github.js +43 -0
- package/dist/execution/github.js.map +1 -0
- package/dist/execution/merge.d.ts +15 -0
- package/dist/execution/merge.js +133 -0
- package/dist/execution/merge.js.map +1 -0
- package/dist/execution/planning-approval.d.ts +33 -0
- package/dist/execution/planning-approval.js +227 -0
- package/dist/execution/planning-approval.js.map +1 -0
- package/dist/execution/pr-review.d.ts +86 -0
- package/dist/execution/pr-review.js +339 -0
- package/dist/execution/pr-review.js.map +1 -0
- package/dist/execution/review-loop.d.ts +63 -0
- package/dist/execution/review-loop.js +463 -0
- package/dist/execution/review-loop.js.map +1 -0
- package/dist/execution/test-feature.d.ts +29 -0
- package/dist/execution/test-feature.js +240 -0
- package/dist/execution/test-feature.js.map +1 -0
- package/dist/execution/watch.d.ts +12 -0
- package/dist/execution/watch.js +34 -0
- package/dist/execution/watch.js.map +1 -0
- package/dist/memory/context.d.ts +14 -0
- package/dist/memory/context.js +82 -0
- package/dist/memory/context.js.map +1 -0
- package/dist/memory/improvement.d.ts +25 -0
- package/dist/memory/improvement.js +168 -0
- package/dist/memory/improvement.js.map +1 -0
- package/dist/memory/launch.d.ts +9 -0
- package/dist/memory/launch.js +65 -0
- package/dist/memory/launch.js.map +1 -0
- package/dist/memory/runner.d.ts +4 -0
- package/dist/memory/runner.js +74 -0
- package/dist/memory/runner.js.map +1 -0
- package/dist/memory/state.d.ts +10 -0
- package/dist/memory/state.js +42 -0
- package/dist/memory/state.js.map +1 -0
- package/dist/presentation.d.ts +19 -0
- package/dist/presentation.js +128 -0
- package/dist/presentation.js.map +1 -0
- package/dist/reporting/dashboard.d.ts +2 -0
- package/dist/reporting/dashboard.js +106 -0
- package/dist/reporting/dashboard.js.map +1 -0
- package/dist/reporting/pr-usage.d.ts +24 -0
- package/dist/reporting/pr-usage.js +179 -0
- package/dist/reporting/pr-usage.js.map +1 -0
- package/dist/reporting/records.d.ts +28 -0
- package/dist/reporting/records.js +152 -0
- package/dist/reporting/records.js.map +1 -0
- package/dist/setup/agents.d.ts +8 -0
- package/dist/setup/agents.js +83 -0
- package/dist/setup/agents.js.map +1 -0
- package/dist/setup/assessment.d.ts +31 -0
- package/dist/setup/assessment.js +75 -0
- package/dist/setup/assessment.js.map +1 -0
- package/dist/setup/auto-loaded.d.ts +17 -0
- package/dist/setup/auto-loaded.js +49 -0
- package/dist/setup/auto-loaded.js.map +1 -0
- package/dist/setup/copilot-version.d.ts +7 -0
- package/dist/setup/copilot-version.js +39 -0
- package/dist/setup/copilot-version.js.map +1 -0
- package/dist/setup/copilot.d.ts +22 -0
- package/dist/setup/copilot.js +164 -0
- package/dist/setup/copilot.js.map +1 -0
- package/dist/setup/init.d.ts +41 -0
- package/dist/setup/init.js +315 -0
- package/dist/setup/init.js.map +1 -0
- package/dist/setup/install.d.ts +12 -0
- package/dist/setup/install.js +211 -0
- package/dist/setup/install.js.map +1 -0
- package/dist/setup/instruction-quality.d.ts +25 -0
- package/dist/setup/instruction-quality.js +258 -0
- package/dist/setup/instruction-quality.js.map +1 -0
- package/dist/setup/inventory.d.ts +26 -0
- package/dist/setup/inventory.js +78 -0
- package/dist/setup/inventory.js.map +1 -0
- package/dist/setup/onboarding.d.ts +40 -0
- package/dist/setup/onboarding.js +472 -0
- package/dist/setup/onboarding.js.map +1 -0
- package/dist/setup/package.d.ts +2 -0
- package/dist/setup/package.js +8 -0
- package/dist/setup/package.js.map +1 -0
- package/dist/setup/planning-workflow.d.ts +3 -0
- package/dist/setup/planning-workflow.js +250 -0
- package/dist/setup/planning-workflow.js.map +1 -0
- package/dist/setup/repository-map.d.ts +34 -0
- package/dist/setup/repository-map.js +89 -0
- package/dist/setup/repository-map.js.map +1 -0
- package/dist/setup/review.d.ts +8 -0
- package/dist/setup/review.js +137 -0
- package/dist/setup/review.js.map +1 -0
- package/dist/setup/team.d.ts +18 -0
- package/dist/setup/team.js +172 -0
- package/dist/setup/team.js.map +1 -0
- package/dist/setup/templates.d.ts +11 -0
- package/dist/setup/templates.js +494 -0
- package/dist/setup/templates.js.map +1 -0
- package/dist/setup/terminal.d.ts +12 -0
- package/dist/setup/terminal.js +11 -0
- package/dist/setup/terminal.js.map +1 -0
- package/dist/setup/update.d.ts +6 -0
- package/dist/setup/update.js +71 -0
- package/dist/setup/update.js.map +1 -0
- package/dist/specification/batch.d.ts +50 -0
- package/dist/specification/batch.js +127 -0
- package/dist/specification/batch.js.map +1 -0
- package/dist/specification/plan-checks.d.ts +6 -0
- package/dist/specification/plan-checks.js +73 -0
- package/dist/specification/plan-checks.js.map +1 -0
- package/dist/specification/planning.d.ts +26 -0
- package/dist/specification/planning.js +507 -0
- package/dist/specification/planning.js.map +1 -0
- package/dist/specification/prose.d.ts +1 -0
- package/dist/specification/prose.js +15 -0
- package/dist/specification/prose.js.map +1 -0
- package/dist/tracking/ado.d.ts +19 -0
- package/dist/tracking/ado.js +128 -0
- package/dist/tracking/ado.js.map +1 -0
- package/dist/tracking/github.d.ts +8 -0
- package/dist/tracking/github.js +112 -0
- package/dist/tracking/github.js.map +1 -0
- package/dist/tracking/issues.d.ts +28 -0
- package/dist/tracking/issues.js +180 -0
- package/dist/tracking/issues.js.map +1 -0
- package/dist/tracking/native.d.ts +5 -0
- package/dist/tracking/native.js +30 -0
- package/dist/tracking/native.js.map +1 -0
- package/dist/tracking/refs.d.ts +14 -0
- package/dist/tracking/refs.js +56 -0
- package/dist/tracking/refs.js.map +1 -0
- package/dist/tracking/sources.d.ts +5 -0
- package/dist/tracking/sources.js +29 -0
- package/dist/tracking/sources.js.map +1 -0
- package/docs/operations.md +1114 -0
- package/examples/batch.json +27 -0
- package/examples/runs.json +22 -0
- package/package.json +50 -0
|
@@ -0,0 +1,1114 @@
|
|
|
1
|
+
# Operating Crewbie
|
|
2
|
+
|
|
3
|
+
## Human approval and credentials
|
|
4
|
+
|
|
5
|
+
Commit reviewed setup files to the repository's default branch before dispatch.
|
|
6
|
+
Crewbie authorizes human actions from users with write access to the configured
|
|
7
|
+
repository (`admin`, `maintain` or `write`). Publication checks the authenticated
|
|
8
|
+
user and writes exact-content approval comments; bots, read-only users and edited
|
|
9
|
+
approval comments are not accepted as human approval.
|
|
10
|
+
|
|
11
|
+
Local GitHub access uses `GH_TOKEN`, `GITHUB_TOKEN`, or `gh auth login`.
|
|
12
|
+
Cloud assignment needs a supported user-authorized credential, not an ordinary
|
|
13
|
+
GitHub App installation token. Follow GitHub's
|
|
14
|
+
[current assignment permissions](https://docs.github.com/en/copilot/how-tos/use-copilot-agents/cloud-agent/use-cloud-agent-via-the-api).
|
|
15
|
+
Do not print tokens, put them in command arguments, or commit them.
|
|
16
|
+
|
|
17
|
+
For a repository-restricted fine-grained PAT, issue assignment requires Metadata
|
|
18
|
+
read and Actions, Contents, Issues and Pull requests read/write. Crewbie also
|
|
19
|
+
needs **Agent tasks: read** to correlate sessions and release completed capacity;
|
|
20
|
+
assignment permissions alone do not grant telemetry access. The optional
|
|
21
|
+
same-PR correction loop needs **Agent tasks: read/write**. See GitHub's
|
|
22
|
+
[Agent Tasks permissions](https://docs.github.com/en/rest/agent-tasks/agent-tasks).
|
|
23
|
+
Unavailable telemetry keeps capacity reserved rather than guessing completion.
|
|
24
|
+
Merging a task PR into its feature branch reads the PR's check runs and commit statuses with the dispatch job's
|
|
25
|
+
own `GITHUB_TOKEN` (`checks: read`, `statuses: read`), so the user credential
|
|
26
|
+
needs no Checks or Commit statuses access; the merge itself uses the user
|
|
27
|
+
credential's Contents and Pull requests write access.
|
|
28
|
+
|
|
29
|
+
Read-only dashboard collection and PR checks use the job-scoped `GITHUB_TOKEN`
|
|
30
|
+
with explicit read permissions. They do not require copying a user's saved
|
|
31
|
+
credential into Actions secrets. Native assignment still requires its documented
|
|
32
|
+
user authentication. Maintenance uses separate built-in job tokens for read-only
|
|
33
|
+
preparation, Copilot requests, and guarded branch/PR publication.
|
|
34
|
+
|
|
35
|
+
The installed workflows use:
|
|
36
|
+
|
|
37
|
+
| Setting | Purpose |
|
|
38
|
+
|---|---|
|
|
39
|
+
| Secret `CREWBIE_USER_TOKEN` | Unattended native assignment using a supported user credential; unnecessary for `publish --dispatch-local` |
|
|
40
|
+
| Secret `CREWBIE_ADO_TOKEN` | Optional ADO work-item access |
|
|
41
|
+
| Variable `CREWBIE_PACKAGE` | Optional approved pinned package override; defaults to the exact version that generated the workflows |
|
|
42
|
+
| Variable `CREWBIE_COPILOT_VERSION` | An approved exact Copilot CLI package version |
|
|
43
|
+
| Variable `CREWBIE_MAINTENANCE_MODEL` | Explicit approved maintenance model; not `auto` |
|
|
44
|
+
| Variable `CREWBIE_PAGES_MODE` | Leave unset for artifact-only reports; opt into `private` or `public` |
|
|
45
|
+
| Config `planning.enabled` / `planning.model` | Opt into ready-label coordinator planning with an explicit model |
|
|
46
|
+
| Config `planning.executeOnMerge` | Opt into paid task execution after a verified human approval and merge |
|
|
47
|
+
| Config `review.enabled` / `review.role` | Have a configured role review each task PR whose CI ran against its own acceptance criteria before auto-merge, and every head of a plan's feature PR, in a tool-free Copilot CLI job (`review.model` overrides the role's model). Init enables it with the proposed reviewer (preferring a review or verification specialist) and preserves the installed choice, including `enabled: false`, in which case task PRs merge on passing checks alone |
|
|
48
|
+
| Config `merge.method` | How task PRs merge into the feature branch: `merge` (default), `squash` or `rebase`. Legacy `merge.mode` and `merge.minConfidence` are ignored |
|
|
49
|
+
| Config `local.start` | Optional local app start command used by `crewbie test` after it checks out a feature branch, for example `npm run dev` |
|
|
50
|
+
|
|
51
|
+
Generated workflows embed the exact installed version's GitHub release tarball
|
|
52
|
+
URL. A missing
|
|
53
|
+
`CREWBIE_PACKAGE` variable no longer blocks ready-label dispatch. The embedded
|
|
54
|
+
version must have its package asset attached to its GitHub release; no npm
|
|
55
|
+
registry publication is required. For an unpublished/custom build, distribute an
|
|
56
|
+
approved tarball and override the package source:
|
|
57
|
+
|
|
58
|
+
```powershell
|
|
59
|
+
gh variable set CREWBIE_PACKAGE --repo OWNER/REPO --body "https://YOUR-RELEASE-HOST/crewbie-cli-VERSION.tgz"
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Installing the package uses `--ignore-scripts`; the packed `dist` is prebuilt and
|
|
63
|
+
its JSONC parser is installed by npm. Keep credentials out of package URLs.
|
|
64
|
+
Existing repositories need to reapply reviewed setup with the new CLI and commit
|
|
65
|
+
the changed workflows; upgrading a local package alone cannot change hosted YAML.
|
|
66
|
+
|
|
67
|
+
The nightly analysis job grants only `copilot-requests: write` and uses a recent,
|
|
68
|
+
pinned CLI with the built-in token. Personal repositories bill the owner's Copilot
|
|
69
|
+
seat; organization repositories require the Copilot CLI organization-billing
|
|
70
|
+
policy. See [Copilot CLI Actions authentication](https://docs.github.com/en/copilot/concepts/agents/copilot-cli/copilot-cli-in-github-actions).
|
|
71
|
+
The separate publisher needs `contents: write`, `pull-requests: write`, and the
|
|
72
|
+
repository setting allowing Actions to create PRs. It never approves or merges a
|
|
73
|
+
PR. A missing entitlement or permission is an error, not a token/runtime fallback.
|
|
74
|
+
|
|
75
|
+
## npm publishing
|
|
76
|
+
|
|
77
|
+
`package.json` declares a public scoped package with a `crewbie` binary.
|
|
78
|
+
`npm pack` builds and includes `dist`, docs, examples and the license. Installing
|
|
79
|
+
the package does not run setup, create labels, or start agents.
|
|
80
|
+
|
|
81
|
+
GitHub releases run the package-artifact job, which uploads the npm-installable
|
|
82
|
+
tarball and `SHA256SUMS`. npm registry publication is **disabled by default**:
|
|
83
|
+
the separate publish job requires the repository Actions variable
|
|
84
|
+
`CREWBIE_NPM_PUBLISH_ENABLED` to equal `true`. Leave it unset while npm account
|
|
85
|
+
setup is deferred.
|
|
86
|
+
|
|
87
|
+
Before the first publication, the maintainer must have an npm account with 2FA
|
|
88
|
+
and publish permission under `@crewbie`. If available, create the `crewbie`
|
|
89
|
+
organization through the npm profile menu's **Add an Organization**, selecting
|
|
90
|
+
the free **Unlimited public packages** plan. If another party owns the scope,
|
|
91
|
+
obtain permission or explicitly choose a different package name.
|
|
92
|
+
|
|
93
|
+
Run `npm login --auth-type=web --registry=https://registry.npmjs.org` and complete
|
|
94
|
+
authentication directly with npm. `npm whoami --registry=https://registry.npmjs.org`
|
|
95
|
+
checks the authenticated account; it does not prove namespace permission.
|
|
96
|
+
Keep credentials and recovery codes out of Git and chat. A registry TLS failure
|
|
97
|
+
requires approved network/proxy/CA configuration, not `strict-ssl=false`.
|
|
98
|
+
|
|
99
|
+
Review the package contents with `npm pack --dry-run`, then bootstrap with
|
|
100
|
+
`npm publish --access public --tag next`. Complete any 2FA challenge directly
|
|
101
|
+
with npm. Verify the published version using
|
|
102
|
+
`npm view @crewbie/cli@0.1.0-alpha.49 version --registry=https://registry.npmjs.org`.
|
|
103
|
+
|
|
104
|
+
After the package exists, open its npm **Settings > Trusted publishing**, choose
|
|
105
|
+
GitHub Actions, and configure:
|
|
106
|
+
|
|
107
|
+
| Field | Value |
|
|
108
|
+
|---|---|
|
|
109
|
+
| Organization or user | `mvanderbend-msoft` |
|
|
110
|
+
| Repository | `crewbie` |
|
|
111
|
+
| Workflow filename | `publish.yml` |
|
|
112
|
+
| Environment | `npm` |
|
|
113
|
+
|
|
114
|
+
Configure the `npm` GitHub environment's release approvals and commit the
|
|
115
|
+
workflow before creating a release. After trusted publishing is ready, explicitly
|
|
116
|
+
enable `CREWBIE_NPM_PUBLISH_ENABLED=true`. No long-lived npm token is needed in
|
|
117
|
+
GitHub secrets. Publishing a GitHub release then runs the checks, verifies
|
|
118
|
+
the release tag matches `package.json`, and publishes with provenance using OIDC.
|
|
119
|
+
Prereleases use `next`; stable releases use `latest`. Merely changing this
|
|
120
|
+
repository does not publish a package or configure the npm account.
|
|
121
|
+
|
|
122
|
+
## Account and runtime capability matrix
|
|
123
|
+
|
|
124
|
+
| Capability | Personal repository | Organization repository | Verified here |
|
|
125
|
+
|---|---|---|---|
|
|
126
|
+
| Local onboarding/task decomposition | Supported | Supported | Local CLI and fixtures; LLM response handling uses deterministic fixtures |
|
|
127
|
+
| Ready-label issue planning | Opt-in Actions CLI; eligible Copilot seat | Opt-in Actions CLI; organization billing policy | Live private Java/React intake passed prepare/analyze/publish with the coordinator charter/history, proposing five tasks across four specialists |
|
|
128
|
+
| Approve-and-merge execution | Supported user-authorized assignment credential required | Credential plus repository/organization policy | Exact-head review/merge provenance, team materialization, publication and recovery fixtures; live webshop validation pending |
|
|
129
|
+
| Native custom-agent assignment | Requires eligible account/repo and user auth | Requires eligible account/policy and user auth | Named backend, frontend, tester and reviewer sessions in a private Java/React repository; native IDs confirmed |
|
|
130
|
+
| Model selection | Requested explicitly; entitlement varies | Requested explicitly; policy varies | `gpt-5.4` confirmed in native session metadata; no universal model guarantee |
|
|
131
|
+
| Same-PR review corrections | Agent Tasks API requires an eligible Business/Enterprise seat and user auth | Requires eligible seat/policy and user auth | Two correction rounds reused the original frontend PR, followed by tester refresh and independent re-review |
|
|
132
|
+
| Nightly CLI | Built-in token, billed to owner's Copilot seat | Built-in token with organization-billing policy | Private personal repository: preparation, paid analysis, proposal publication and no-new-evidence skip passed |
|
|
133
|
+
| Private Pages | Do not assume available | Requires appropriate Enterprise Cloud setup | Visibility guard fixtures; no deployment |
|
|
134
|
+
| Exact specialist cloud tokens/cost | Not guaranteed | Not guaranteed; org totals are not specialist totals | Null/provenance handling and interactive report |
|
|
135
|
+
|
|
136
|
+
Run `crewbie doctor --repo owner/name --agent crewbie-developer --model MODEL`
|
|
137
|
+
for read-only discovery. It does not prove live assignment, model selection, or
|
|
138
|
+
memory use. Keep those limitations visible when evaluating an account.
|
|
139
|
+
|
|
140
|
+
The ready-label trial exposed two separate permission requirements: planning
|
|
141
|
+
preparation/publication need `actions: read` to verify their run provenance, and
|
|
142
|
+
the user credential needs Agent tasks read access to reconcile native sessions.
|
|
143
|
+
After both were corrected, hosted planning published a reviewable PR and hosted
|
|
144
|
+
dispatch verified all four historical completed sessions without deleting claims.
|
|
145
|
+
The planning PR remains subject to human approval and merge; that result alone
|
|
146
|
+
does not establish live merge-triggered assignment.
|
|
147
|
+
|
|
148
|
+
The 2026-09-22 live smoke test used one cloud task in a private synthetic repository.
|
|
149
|
+
The completed task reported the selected specialist and one session. Its PR
|
|
150
|
+
included unique tags present only in the charter, hot memory and active cold
|
|
151
|
+
topic, and its implementation passed four tests locally and on hosted Linux.
|
|
152
|
+
These tags are read evidence, not proof of the model's internal reasoning.
|
|
153
|
+
That first smoke test did not exercise the rest of the team or nightly learning.
|
|
154
|
+
The subsequent team test ran the local coordinator, three named native cloud
|
|
155
|
+
specialists (implementation, testing and review), and the hosted improver.
|
|
156
|
+
Specialist PRs changed their own bounded history; the implementation PR also
|
|
157
|
+
proposed a shared decision. The reviewer ran both branches and identified a real
|
|
158
|
+
reporting gap. The three-job improvement workflow published its own four-line
|
|
159
|
+
hot-memory proposal with built-in tokens, and a second run skipped AI/publication
|
|
160
|
+
when no evidence changed. All learning proposals remained unmerged.
|
|
161
|
+
|
|
162
|
+
The test verified two execution slots, completed-session release for explicit
|
|
163
|
+
review tasks, and a blocked implementation merge-gate probe that never acquired
|
|
164
|
+
a launch claim. Organization entitlement, ADO and Pages remain unverified live.
|
|
165
|
+
|
|
166
|
+
The later Java/React run used a realistic inventory/order application, then normal
|
|
167
|
+
onboarding without context canaries or forced learning. An initial tester session
|
|
168
|
+
timed out and its failed verification was retained. The bounded review loop
|
|
169
|
+
published actual GitHub reviews, routed two rounds of frontend corrections to the
|
|
170
|
+
original named specialist, refreshed tester evidence and ended with clean reviews
|
|
171
|
+
on the exact final heads. Independent combined verification passed focused Java
|
|
172
|
+
checks, 17 frontend checks, the build and five real-browser journeys. Regression
|
|
173
|
+
probes also confirmed that observer-triggered retries stop after a page error and
|
|
174
|
+
that valid empty beyond-end pages remain accepted.
|
|
175
|
+
|
|
176
|
+
GitHub regenerated or retained stale PR descriptions during these sessions.
|
|
177
|
+
Verified attribution and closing links are maintained by reconciliation; the
|
|
178
|
+
local coordinator used guarded description finalization for concise, current
|
|
179
|
+
handoffs. The hosted improver proposed a small guard against inferring policy from
|
|
180
|
+
incomplete evidence, and the private dashboard workflow succeeded. These are
|
|
181
|
+
account-specific alpha results, not a claim of universal or unattended production
|
|
182
|
+
readiness. Application and learning PRs remained unmerged.
|
|
183
|
+
|
|
184
|
+
## Evolving the team
|
|
185
|
+
|
|
186
|
+
Onboarding prefers adopting existing frontend, backend, testing and review
|
|
187
|
+
specialists rather than replacing their ownership with invented combined roles.
|
|
188
|
+
Each eligible original gets an explicit adopt/retain decision and rationale.
|
|
189
|
+
`maxActive` is a concurrency limit, not a roster-size target.
|
|
190
|
+
|
|
191
|
+
An adopted role records `sourceAgent` in configuration. Its original moves from
|
|
192
|
+
`.github/agents/NAME.agent.md` to
|
|
193
|
+
`.crewbie/agent-archive/github/agents/NAME.agent.md` (or the corresponding `claude`
|
|
194
|
+
archive). The new active profile is `.github/agents/crewbie-ROLE.agent.md`.
|
|
195
|
+
The archive preserves the complete original bytes as backup provenance. The active
|
|
196
|
+
profile embeds the complete original instruction body, retains the description,
|
|
197
|
+
persona and frontmatter tool restrictions, and adds Crewbie memory, identity and
|
|
198
|
+
handoff rules. The archive is not a substitute for the active instructions.
|
|
199
|
+
The only lines left out are those that just tell the agent to read guidance Copilot
|
|
200
|
+
already attaches (`.github/copilot-instructions.md`, `AGENTS.md`, matching
|
|
201
|
+
`.github/instructions/*.instructions.md`); lines with other content stay and are flagged.
|
|
202
|
+
Known adopted-agent handoffs are
|
|
203
|
+
retargeted. The selected model governs the active profile. Unsupported tool
|
|
204
|
+
metadata requires manual review rather than silently widening permissions.
|
|
205
|
+
Charters have no word budget. The whole active prompt (preserved body plus Crewbie
|
|
206
|
+
additions) must fit GitHub's documented 30,000-character custom agent maximum. If
|
|
207
|
+
it does not, init stops before adoption and asks the user to shorten the original;
|
|
208
|
+
it never truncates instructions. Available from alpha.12; earlier releases stopped at 400 words.
|
|
209
|
+
|
|
210
|
+
Archival requires the exact inspected source hash, rejects conflicting archives
|
|
211
|
+
or edited originals, and is idempotent. Archives are written before originals
|
|
212
|
+
are removed. Installation previews name both actions. Adding these adoptions is
|
|
213
|
+
an init operation, not a planning PR's authority to retire original agents.
|
|
214
|
+
|
|
215
|
+
The approved configuration is the current roster, not a permanent template.
|
|
216
|
+
`init` on an installed repository and `init --update --model MODEL --out team-review.json`
|
|
217
|
+
reassess the current project without resetting its models, limits,
|
|
218
|
+
constitution, integrations or learning permissions.
|
|
219
|
+
|
|
220
|
+
The `review` report contains the LLM assessment; `team` contains static hints,
|
|
221
|
+
not a final roster. Discovery uses non-ignored production paths and bounded
|
|
222
|
+
manifest inspection: at most 20 manifests, 64 KB each and 512 KB total. Omitted
|
|
223
|
+
or malformed manifests are disclosed. Fixture, example and generated paths do
|
|
224
|
+
not automatically grow the team. No project script is executed.
|
|
225
|
+
|
|
226
|
+
Built-in signals are not a role enum. During init, the coordinator must also consider the
|
|
227
|
+
project's domain and can propose custom roles, splits, specialization or
|
|
228
|
+
retirement. Planning runs never change the team; they only list suggestions.
|
|
229
|
+
Review proposed purpose, checks, non-negotiables and models before
|
|
230
|
+
applying. New roles use the explicitly selected init model, subject to setup review.
|
|
231
|
+
Existing roles and their domain guidance remain intact unless explicitly edited.
|
|
232
|
+
|
|
233
|
+
Init writes a readable Markdown assessment next to its setup JSON. The terminal
|
|
234
|
+
keeps the overview short; the report contains detailed findings, adoption
|
|
235
|
+
decisions, coverage limits and concrete replacement text for guidance edits.
|
|
236
|
+
Advisory recommendations and unresolved policy decisions are not executable
|
|
237
|
+
changes. If `instructions` is empty and no constitution is proposed, choosing
|
|
238
|
+
guidance application cannot modify existing instruction files; init states this
|
|
239
|
+
explicitly. Skipping guidance preserves those proposals for later review.
|
|
240
|
+
|
|
241
|
+
Interactive terminals use arrow-key selectors for models and **Team / All / Save**,
|
|
242
|
+
with Save selected by default. Hosted planning and final installation have separate
|
|
243
|
+
confirmations defaulting to no. Ctrl+C cancels; a proposal already saved remains
|
|
244
|
+
available. Noninteractive automation keeps the existing explicit model, proposal,
|
|
245
|
+
apply and guidance flags; it never waits for a selection menu. Previews group
|
|
246
|
+
creation, updates and archival separately.
|
|
247
|
+
|
|
248
|
+
These selector and active-charter preservation changes, along with the stricter
|
|
249
|
+
review below, are available from alpha.11.
|
|
250
|
+
|
|
251
|
+
The shared CLI presentation layer also formats other commands: grouped help,
|
|
252
|
+
command headings, status-coloured tables, stacked nested records and separated
|
|
253
|
+
errors. It wraps to terminal width (up to 120 columns), measures Unicode display
|
|
254
|
+
width and falls back to stacked fields when a table would be too narrow.
|
|
255
|
+
`NO_COLOR`, `FORCE_COLOR=0` and `TERM=dumb` disable colours without removing the
|
|
256
|
+
readable layout. Redirected streams retain their previous output shapes; supported
|
|
257
|
+
`--json` views skip terminal decoration. Internal workflow commands remain plain,
|
|
258
|
+
including their unchanged Actions summaries and output files. No formatter changes
|
|
259
|
+
approval, execution or exit-code semantics.
|
|
260
|
+
Use `init --proposal FILE --json` for the machine-readable installation preview.
|
|
261
|
+
|
|
262
|
+
Init includes every inventoried instruction, custom-agent and MCP configuration
|
|
263
|
+
path in its assessment. The LLM receives only AI guidance: instructions, custom
|
|
264
|
+
agents (including archives), the constitution and sanitized MCP metadata.
|
|
265
|
+
Application code, READMEs, manifests, ADRs and other project files are not sent,
|
|
266
|
+
so the assessment reviews and edits AI guidance only. Proposed edits outside the
|
|
267
|
+
supported guidance paths are shown as deferred recommendations rather than
|
|
268
|
+
failing installation (from alpha.14; earlier releases also sent
|
|
269
|
+
READMEs, manifests and sampled source files). The static team hints above still use file paths and
|
|
270
|
+
bounded manifest detection. Text inspection has a 256 KB total/64 KB per-file
|
|
271
|
+
budget. Omissions are explicit, not a claim
|
|
272
|
+
of complete semantic coverage. MCP JSON/JSONC files expose only server names,
|
|
273
|
+
transport, executable basename and environment-variable names; values, arguments,
|
|
274
|
+
headers and URLs are withheld. Servers are not started or connectivity-tested.
|
|
275
|
+
Personal/global and ignored settings are not read. MCP configuration changes
|
|
276
|
+
remain recommendations for manual review, not automatic credential-bearing edits.
|
|
277
|
+
|
|
278
|
+
From alpha.44 the assessment and hosted planning also receive a names-only
|
|
279
|
+
repository map: directories to depth 4 and workspaces (directories with their own
|
|
280
|
+
package or project manifest). No file contents are added. Crewbie then checks the
|
|
281
|
+
proposal locally against the full list of file names and reports advisory
|
|
282
|
+
**Repository checks**: role text, adopted agents, proposed guidance or a
|
|
283
|
+
constitution that cite missing paths, and workspaces no role names. Plans get a
|
|
284
|
+
**Plan checks** section in `plan.md` and the planning PR: HTTP status codes and
|
|
285
|
+
quantities a task states that the source issue does not (and issue values no task
|
|
286
|
+
carries), overlapping acceptance criteria between implementation tasks, and task
|
|
287
|
+
paths not in the repository.
|
|
288
|
+
|
|
289
|
+
Model answers are parsed tolerantly: a prose preamble or fenced block around the
|
|
290
|
+
JSON is accepted, finding dispositions accept harmless synonyms (keep, update), and
|
|
291
|
+
an unknown disposition or an edit without replacement text becomes an explicit
|
|
292
|
+
deferral instead of discarding the paid assessment. Other validation failures save
|
|
293
|
+
the redacted answer to a temporary file and, in interactive init only, offer at
|
|
294
|
+
most two disclosed paid repair requests; nothing is retried without a yes.
|
|
295
|
+
|
|
296
|
+
Interactive init discovers the account's enabled model catalogue through the
|
|
297
|
+
official Copilot SDK and presents numbered choices with IDs and available billing
|
|
298
|
+
multipliers. Invalid selections reprompt; `q` cancels. Discovery failure stops
|
|
299
|
+
with an error rather than inventing model choices. An explicit `--model MODEL`
|
|
300
|
+
skips the assessment picker. New roles default to cost-aware choices from the
|
|
301
|
+
account catalog, with reported token prices/capabilities and a reviewed complexity
|
|
302
|
+
rationale. Installed choices are preserved. `--model-policy fixed` or
|
|
303
|
+
`--specialist-model MODEL` skips specialist discovery and uses an explicit override;
|
|
304
|
+
the account still needs entitlement. Legacy multipliers are not token prices.
|
|
305
|
+
|
|
306
|
+
The CLI catalogue can list models the Copilot cloud agent rejects. Issue assignment
|
|
307
|
+
then fails with a misleading "repository ruleset violation" comment. After the
|
|
308
|
+
assessment, init checks each chosen model with the cloud-agent tasks API. A new
|
|
309
|
+
role whose model is rejected triggers one reassessment without that model. An
|
|
310
|
+
installed role gets a warning. Before any launch, dispatch checks each model the
|
|
311
|
+
same way; tasks using a rejected model stay `crewbie:blocked` without a claim or
|
|
312
|
+
attempt, while the rest of the batch launches. The check needs Agent tasks **read
|
|
313
|
+
and write** on `CREWBIE_USER_TOKEN`; with read-only access GitHub answers 403, and
|
|
314
|
+
dispatch warns and launches without the check. The check uses a
|
|
315
|
+
branch that never exists, so no session starts. GitHub still lists a failed task
|
|
316
|
+
in the Agents tab for each check. Preflight skips it to stay read-only.
|
|
317
|
+
|
|
318
|
+
The LLM runs through the SDK's bundled runtime over shell-free stdio, tool-free
|
|
319
|
+
in a temporary working directory and isolated `COPILOT_HOME`, using environment
|
|
320
|
+
credentials or authenticated GitHub CLI. Runtime configuration discovery, skills,
|
|
321
|
+
file hooks, git context and shared session storage are disabled; tool permissions
|
|
322
|
+
are denied. Init waits for a completed assistant response rather than parsing
|
|
323
|
+
process stdout. Startup/model discovery have 30-second timeouts. Assessment has
|
|
324
|
+
no timeout, because no evidence supports a particular duration; it ends when the
|
|
325
|
+
session completes or reports an error, or when you press Ctrl+C. Empty or malformed JSON produces an explicit error;
|
|
326
|
+
check authentication/model access for runtime errors, or retry/narrow the context
|
|
327
|
+
for incomplete output.
|
|
328
|
+
While waiting, init reports elapsed time every 15 seconds, plus when the model
|
|
329
|
+
last streamed reasoning or output and how many characters it has produced. Steady
|
|
330
|
+
activity means the model is still working; a long silence points to a stalled
|
|
331
|
+
session. This heartbeat does not predict completion time. Init also shows the
|
|
332
|
+
prompt size before sending it. Available from alpha.13; earlier releases stopped after five minutes.
|
|
333
|
+
|
|
334
|
+
`contextPaths` is a reusable-guidance list, not a list of implementation targets.
|
|
335
|
+
The prompt and validator share the eligible inspected, unredacted Markdown paths.
|
|
336
|
+
Spaces in document names and `.MD` extensions are supported. Onboarding normalizes
|
|
337
|
+
Windows separators and leading `./` only when the result matches an eligible
|
|
338
|
+
file; traversal, absolute paths, source files, globs and uninspected links remain
|
|
339
|
+
rejected. Errors identify the specialist and invalid paths.
|
|
340
|
+
|
|
341
|
+
Interactive init can repair those links in the existing response: choose numbered
|
|
342
|
+
replacement documents, `none` to explicitly remove only the rejected links, or
|
|
343
|
+
`cancel`. Already-valid links remain. This does not invoke the model again; the
|
|
344
|
+
whole proposal still passes validation and human review before installation.
|
|
345
|
+
Noninteractive runs fail explicitly on invalid links rather than guessing fixes.
|
|
346
|
+
|
|
347
|
+
No fallback roster is installed if analysis fails. `--assessment-only` keeps the
|
|
348
|
+
offline inventory path explicit. Greenfield setup requires a description or
|
|
349
|
+
requirements in the repository; interactive clarification repeats, while
|
|
350
|
+
noninteractive runs persist questions and stop without installing a team.
|
|
351
|
+
|
|
352
|
+
`init --proposal FILE --apply --guidance skip` installs the team without proposed
|
|
353
|
+
guidance changes; `--guidance apply` includes them. Both create all workflow and
|
|
354
|
+
owner labels. Existing labels are preserved; retries add only missing labels.
|
|
355
|
+
Label failure reports that local setup succeeded and remote setup is incomplete.
|
|
356
|
+
Rerun the same command to repair it. Use `--skip-labels` only for deliberate
|
|
357
|
+
offline setup; rerun without it before hosted intake.
|
|
358
|
+
|
|
359
|
+
Reassessment proposals bind to the existing configuration fingerprint. A stale
|
|
360
|
+
proposal cannot overwrite intervening policy changes; LF/CRLF checkout differences
|
|
361
|
+
are tolerated. Installation retains historical profile and memory files when a
|
|
362
|
+
role is explicitly removed from the active config. Crewbie stops routing new
|
|
363
|
+
work to that role, but never reassigns existing tasks automatically. Review/drain
|
|
364
|
+
open work and reapprove any changed task ownership before retiring a role.
|
|
365
|
+
New roles do not silently expand nightly instruction-edit permissions.
|
|
366
|
+
|
|
367
|
+
## Ready-label issue intake
|
|
368
|
+
|
|
369
|
+
Interactive init asks for explicit hosted-planning opt-in. Existing enabled
|
|
370
|
+
planning settings are preserved; automatic execution on merge is not enabled by
|
|
371
|
+
this prompt. Commit both the approved configuration and regenerated planning
|
|
372
|
+
workflow before using the label.
|
|
373
|
+
|
|
374
|
+
Dispatch lists only open managed issues, reads claim refs once, and revisits a
|
|
375
|
+
closed issue only when it still holds a claim and closed within the last 24 hours
|
|
376
|
+
(so an outliving session keeps its slot) or when an open task depends on it.
|
|
377
|
+
Older closed tasks are not scanned and no longer appear in `crewbie status`.
|
|
378
|
+
|
|
379
|
+
**A successful dispatch run is not evidence that an agent started.** Dispatch
|
|
380
|
+
reconciles published managed implementation tasks, while
|
|
381
|
+
`crewbie:ready-for-planning` belongs to the separate planning workflow. With no
|
|
382
|
+
managed tasks, dispatch reports that no agents started and explains whether
|
|
383
|
+
planning is enabled, in both logs and the Actions job summary. If planning is
|
|
384
|
+
disabled, explicitly enable it through reviewed init, commit the resulting setup,
|
|
385
|
+
then remove/reapply the ready-for-planning label to request planning. Do not add
|
|
386
|
+
managed labels or bypass human execution approval to force a launch.
|
|
387
|
+
|
|
388
|
+
The label **`crewbie:ready-for-planning`** is separate from the execution-state
|
|
389
|
+
label `crewbie:ready`. It authorizes coordinator planning, not application work.
|
|
390
|
+
The coordinator aims for a summary of at most 100 words. If it writes more,
|
|
391
|
+
the complete summary stays in `plan.md`; the planning PR shows only complete
|
|
392
|
+
sentences that fit within 100 words and points to the full plan. If no complete
|
|
393
|
+
sentence fits, the PR points directly to `plan.md`.
|
|
394
|
+
|
|
395
|
+
Enable it in a reviewed setup proposal:
|
|
396
|
+
|
|
397
|
+
```json
|
|
398
|
+
{
|
|
399
|
+
"planning": {
|
|
400
|
+
"enabled": true,
|
|
401
|
+
"model": "gpt-5.4"
|
|
402
|
+
}
|
|
403
|
+
}
|
|
404
|
+
```
|
|
405
|
+
|
|
406
|
+
This is the `config.planning` fragment, not a complete setup file. Choose a model
|
|
407
|
+
your account supports; the example is not an entitlement guarantee. Apply the
|
|
408
|
+
reviewed proposal and commit the generated configuration, profiles, memory and
|
|
409
|
+
`crewbie-plan.yml` workflow to the default branch.
|
|
410
|
+
|
|
411
|
+
`crewbie update` previews repository integration changes without AI reassessment;
|
|
412
|
+
`crewbie update --apply` applies them and refreshes an existing `CREWBIE_PACKAGE`
|
|
413
|
+
override to the installed CLI release. Upgrade the CLI separately with npm.
|
|
414
|
+
Commit the generated file changes. Edit Crewbie agents and instructions at any
|
|
415
|
+
time: in an edited `crewbie-*.agent.md` update refreshes only the block between
|
|
416
|
+
`<!-- crewbie:managed:start -->` and `<!-- crewbie:managed:end -->` and keeps
|
|
417
|
+
everything else (a charter without those markers, `.crewbie/instructions.md`,
|
|
418
|
+
the Crewbie skill and the PR template are kept as you wrote them and listed as
|
|
419
|
+
kept). Only edited Crewbie workflows and colliding files you own still block
|
|
420
|
+
application; resolve those rather than changing ownership hashes blindly. `--offline`
|
|
421
|
+
leaves remote variables unchecked. Existing execution opt-outs and models remain
|
|
422
|
+
unchanged.
|
|
423
|
+
|
|
424
|
+
Hosted planning needs an approved exact `CREWBIE_COPILOT_VERSION`. When planning
|
|
425
|
+
is enabled and the variable is unset, interactive init asks for it (Enter accepts
|
|
426
|
+
the latest `@github/copilot` release) and creates it on apply; scripted
|
|
427
|
+
`init --proposal ... --apply` takes `--copilot-version X.Y.Z`. An existing value
|
|
428
|
+
is never overwritten, and `crewbie update` reports a missing one. Also allow
|
|
429
|
+
Actions to create pull requests. Copilot billing/organization policy still
|
|
430
|
+
applies. No saved user token is required for planning. To set it manually:
|
|
431
|
+
|
|
432
|
+
```powershell
|
|
433
|
+
gh variable set CREWBIE_COPILOT_VERSION --repo OWNER/REPO --body "1.0.88"
|
|
434
|
+
```
|
|
435
|
+
|
|
436
|
+
Init creates the ready-for-planning label with the other workflow labels.
|
|
437
|
+
Put the user-authored PRD/spec in the issue body, then have a write-access user apply the
|
|
438
|
+
label. The workflow checks the actual label-event actor and current issue
|
|
439
|
+
content before analysis. Bots, read-only users, closed issues, generated
|
|
440
|
+
execution issues and unrelated labels cannot start planning. Creating an issue
|
|
441
|
+
alone is not a trigger. After source changes, review the text and remove/reapply
|
|
442
|
+
the label; it does not continuously analyze every edit.
|
|
443
|
+
|
|
444
|
+
The prepare job uses read-only repository access and loads the coordinator
|
|
445
|
+
charter, bounded hot/index memory, relevant indexed history, shared guidance and
|
|
446
|
+
the repository assessment. The separate model job has Copilot-request permission
|
|
447
|
+
but no repository write permission or available tools. This is a named-context
|
|
448
|
+
Copilot CLI planning run in Actions, not a native Agent Tasks implementation
|
|
449
|
+
session. Model selection is requested explicitly; runtime model/billing
|
|
450
|
+
measurements are not inferred.
|
|
451
|
+
|
|
452
|
+
The publisher rechecks the source, label approval, policy and default-branch
|
|
453
|
+
revision. It creates a non-draft PR with `.crewbie/plans/<feature>-issue-N/` files: a concise
|
|
454
|
+
human-facing plan, an unchanged configuration copy and an unapproved task batch when requirements
|
|
455
|
+
are sufficient. Merge-enabled plans also include an execution manifest, as described
|
|
456
|
+
below. Each task names an existing owner, its model and dependencies; missing
|
|
457
|
+
expertise appears as team suggestions, not new roles. Missing requirements
|
|
458
|
+
produce questions rather than fabricated acceptance criteria; those PRs remain drafts.
|
|
459
|
+
Non-draft means ready for review, not permission to bypass branch protection.
|
|
460
|
+
Already-generated legacy `issue-N` directories remain executable.
|
|
461
|
+
Crewbie does not create PRDs/specs. The legacy batch `spec` field remains for
|
|
462
|
+
compatibility and holds a deterministic source reference in hosted plans;
|
|
463
|
+
model-authored specification text is discarded. The constitution remains in use.
|
|
464
|
+
|
|
465
|
+
With merge execution disabled, review the proposal on its branch. Preview and apply `setup.json` through `init`,
|
|
466
|
+
then review/merge the resulting configuration and profiles onto the default
|
|
467
|
+
branch. Resolve questions and inspect `batch.json` before `approve --batch ...
|
|
468
|
+
--yes --execute` and `publish --batch ... --apply --dispatch-local --watch`.
|
|
469
|
+
In this manual mode, merging the planning PR alone neither installs its
|
|
470
|
+
nested setup proposal nor approves execution. The coordinator never approves
|
|
471
|
+
its own task graph.
|
|
472
|
+
|
|
473
|
+
The same source/base/configuration snapshot is deduplicated, including a closed
|
|
474
|
+
planning PR. Existing branches without a matching PR indicate interrupted
|
|
475
|
+
publication and stop visibly; inspect them rather than deleting state or blindly
|
|
476
|
+
retrying. For an open plan, reply on the PR: while it has open questions (posted
|
|
477
|
+
as a Crewbie comment, PR in draft) any new write-access user comment is taken as the answers;
|
|
478
|
+
otherwise a comment starting with `/crewbie revise` carries the feedback. Each such
|
|
479
|
+
comment requests one paid run (the planning workflow listens to `issue_comment`).
|
|
480
|
+
The comment job prefilters bots and associations other than OWNER, MEMBER or
|
|
481
|
+
COLLABORATOR; plain-text answers remain possible while questions are open.
|
|
482
|
+
Locally, `crewbie revise-plan --pr N --feedback-file feedback.txt` previews the same
|
|
483
|
+
revision and `--apply` requests it.
|
|
484
|
+
The workflow accepts an explicit human request, reuses prior setup/plan/batch
|
|
485
|
+
context, skips the full assessment and regenerates the execution manifest.
|
|
486
|
+
Only users with write access may request it: the API permission check remains
|
|
487
|
+
authoritative because COLLABORATOR does not guarantee write access. It checks
|
|
488
|
+
source, policy and the exact prior head, and advances the branch
|
|
489
|
+
without force; a branch behind the default branch gets it merged into the revision
|
|
490
|
+
commit, so no manual branch update is needed. Previous approvals are stale after
|
|
491
|
+
revision; close/relabel is not needed for ordinary plan feedback. Re-running an already
|
|
492
|
+
published revision skips further analysis. If publication is interrupted, inspect
|
|
493
|
+
the existing branch and metadata before another paid request.
|
|
494
|
+
Planning never merges PRs.
|
|
495
|
+
|
|
496
|
+
Inputs are bounded to a 50 KB issue body and 100 KB total prompt/output. Plans
|
|
497
|
+
have at most eight tasks, five questions and four additional roles per proposal.
|
|
498
|
+
The three jobs have 3/9/3-minute limits; different issues can plan concurrently.
|
|
499
|
+
These limits are not spending caps. Links, attachments, Word/PDF files and
|
|
500
|
+
external URLs are **not fetched**: paste the relevant text into the issue.
|
|
501
|
+
Generated task issues are explicitly excluded, preventing recursive planning.
|
|
502
|
+
|
|
503
|
+
### Approve and merge to execute
|
|
504
|
+
|
|
505
|
+
For a GitHub-only per-feature handoff, `config.planning.executeOnMerge` defaults to
|
|
506
|
+
`true` when a new assessment is installed with hosted planning enabled. An explicit
|
|
507
|
+
`false` remains an opt-out. Existing installations (including legacy missing flags)
|
|
508
|
+
are not silently opted in. To enable one, explicitly set it true in the installed
|
|
509
|
+
configuration and preview/apply `crewbie update`. Install the updated workflows and commit
|
|
510
|
+
them to the default branch. Configure `CREWBIE_USER_TOKEN` in repository Actions
|
|
511
|
+
secrets using a supported user-authorized credential belonging to a user with
|
|
512
|
+
write access. It needs the documented native-assignment access, issue publication,
|
|
513
|
+
claim-ref writes and workflow-dispatch permissions. Follow the linked GitHub
|
|
514
|
+
permission guidance rather than assuming an installation token can assign agents.
|
|
515
|
+
Store credentials through the approved secret store, never issues or commits.
|
|
516
|
+
Enable `crewbie-execute-plan.yml` and `crewbie-dispatch.yml` if previously disabled.
|
|
517
|
+
The built-in job token is sufficient for planning, **not native assignment**.
|
|
518
|
+
|
|
519
|
+
For subsequent features, your only handoff is to review the planning PR, approve
|
|
520
|
+
its exact final commit, and merge it into the default branch. Ready plans only add
|
|
521
|
+
files under `.crewbie/plans/<issue>/`:
|
|
522
|
+
|
|
523
|
+
- The concise plan, task batch and any team suggestions (never applied).
|
|
524
|
+
- A copy of the unchanged configuration used for revisions (`setup.json`).
|
|
525
|
+
- A bounded execution manifest identifying the planning run and reviewed files.
|
|
526
|
+
|
|
527
|
+
The planning PR never changes roles, agent charters, role memory, configuration,
|
|
528
|
+
workflows or secrets; it may add or update `.crewbie/decisions/hot.md`, which is
|
|
529
|
+
reviewed with the plan but is not part of the execution manifest. Any other file
|
|
530
|
+
blocks execution. Change the team
|
|
531
|
+
through reviewed `crewbie init --update`. It includes no application changes. Clarification-only
|
|
532
|
+
plans do not contain an executable manifest and cannot start work when merged.
|
|
533
|
+
If a generated plan needs edits, regenerate it and review the new commit; changing
|
|
534
|
+
files without refreshing its fingerprints blocks execution rather than accepting
|
|
535
|
+
an ambiguous plan.
|
|
536
|
+
|
|
537
|
+
The merged-PR workflow runs trusted package code from the default branch, never
|
|
538
|
+
untrusted PR-head code with assignment credentials. It independently verifies:
|
|
539
|
+
|
|
540
|
+
- Prior opt-in policy from the recorded default-branch planning workflow run,
|
|
541
|
+
which must have completed successfully.
|
|
542
|
+
- A configured human's `APPROVED` review on the exact final head before merge,
|
|
543
|
+
and a configured human merger. Stale/bot/dismissed approvals and unresolved
|
|
544
|
+
human change requests do not qualify.
|
|
545
|
+
- The complete allowed file set, unchanged contents across reviewed head, merge
|
|
546
|
+
commit and current default branch, and the unchanged source-issue requirements.
|
|
547
|
+
A regenerated plan can reuse identical setup, plan or batch files: these remain
|
|
548
|
+
in the manifest even when GitHub omits them from the PR diff. Every omitted file
|
|
549
|
+
must also match the recorded planning base; it is not exempt from content checks.
|
|
550
|
+
|
|
551
|
+
After authorization, it publishes task issues with specialist/model ownership,
|
|
552
|
+
records their planning-PR approval provenance and explicitly requests the normal
|
|
553
|
+
dispatcher. No local `init`, `approve` or `publish` command is needed for that
|
|
554
|
+
feature. Shared locking and persistent launch claims retain concurrency and
|
|
555
|
+
duplicate-launch protections. The dispatcher reconciles on configured issue/PR
|
|
556
|
+
events and its hourly recovery schedule; GitHub can delay scheduled runs. Implementation
|
|
557
|
+
dependencies still require merged application PRs; independent work and explicit
|
|
558
|
+
review tasks progress under their existing rules. Application PRs are never
|
|
559
|
+
automatically merged.
|
|
560
|
+
|
|
561
|
+
Missing credentials, changed policy, lost planning-run evidence, changed files
|
|
562
|
+
or failed publication stop visibly. For recovery, use **Run workflow** on
|
|
563
|
+
`Crewbie execute approved plan` with the merged planning PR number. All approval
|
|
564
|
+
checks run again; matching issues and approvals are reused. Do not erase launch
|
|
565
|
+
claims or overwrite branches to force another paid session. Keep the planning
|
|
566
|
+
run record until execution/recovery finishes. A merge by an unconfigured bot or
|
|
567
|
+
merge-queue identity does not substitute for the required human merger.
|
|
568
|
+
|
|
569
|
+
## ADO-authoritative work
|
|
570
|
+
|
|
571
|
+
Configure `ado` with `organization`, `project`, and `workItemType`.
|
|
572
|
+
|
|
573
|
+
```powershell
|
|
574
|
+
crewbie status --ado-id 123
|
|
575
|
+
```
|
|
576
|
+
|
|
577
|
+
Preserve the returned URI, revision, and content fingerprint in `batch.sources`.
|
|
578
|
+
Use a task's optional `adoWorkItem` to link an existing item. Add `--ado-create`
|
|
579
|
+
to a reviewed `publish --apply` command to create missing ADO items instead.
|
|
580
|
+
The preview identifies that choice.
|
|
581
|
+
|
|
582
|
+
ADO creation uses a deterministic tag to recover a mapping after partial
|
|
583
|
+
failure. Source material changes stop publication/dispatch for reconciliation.
|
|
584
|
+
Fingerprints distinguish material changes from bookkeeping-only revisions.
|
|
585
|
+
An existing item is never silently rewritten to fit a changed batch.
|
|
586
|
+
|
|
587
|
+
GitHub issues/PRs and brief status comments are written back idempotently.
|
|
588
|
+
No ADO state transition, description replacement, or two-way acceptance-criteria
|
|
589
|
+
sync occurs. A failure after GitHub writes is a partial operation, not a rollback:
|
|
590
|
+
re-run after inspection to reuse existing issue/tag mappings. Verify uncertain
|
|
591
|
+
ADO creation outcomes before retrying because server-side indexing can lag.
|
|
592
|
+
|
|
593
|
+
## Instruction quality: evidence and limits
|
|
594
|
+
|
|
595
|
+
The read-only assessment cites
|
|
596
|
+
[Gloaguen et al., *Evaluating AGENTS.md*](https://arxiv.org/abs/2602.11988).
|
|
597
|
+
In the evaluated settings, context files did not generally improve task success
|
|
598
|
+
and increased inference cost; repository overviews were not helpful. The authors
|
|
599
|
+
still identify value in non-standard coding practices. These are study-specific
|
|
600
|
+
observations, not a universal instruction ban or proof that a particular number
|
|
601
|
+
of words is harmful.
|
|
602
|
+
|
|
603
|
+
Crewbie's documentation/profile overlap, generic-advice (including generic
|
|
604
|
+
agent charters), unverified-link, npm-script/package-manifest, broad-root-scope
|
|
605
|
+
and unconditional full-suite signals are **engineering heuristics**,
|
|
606
|
+
not validated causal rules from the paper. They identify concrete material for
|
|
607
|
+
human review. Necessary standalone context and explicit merge/compliance gates
|
|
608
|
+
should be retained. Contradictions, domain relevance and actual benefit still
|
|
609
|
+
need semantic review and representative before/after task evidence.
|
|
610
|
+
The 600-word root-guidance review threshold is advisory, not a gate. Scoped
|
|
611
|
+
Copilot instructions need valid YAML `applyTo` globs. When relocating domain
|
|
612
|
+
guidance, review the source reduction and destination together; preserve policy
|
|
613
|
+
coverage and host-specific instruction support. Every finding must explicitly
|
|
614
|
+
select `retain`, `edit` or `defer`. Concrete edits require complete replacements
|
|
615
|
+
listed in `editPaths`; a scoped move must include its source reduction.
|
|
616
|
+
Deferred recommendations require a nonempty `deferReason` identifying the blocker
|
|
617
|
+
and appear separately in the terminal and Markdown report. Routine approval alone
|
|
618
|
+
is not a reason to defer a safe proposal: approval already gates every write.
|
|
619
|
+
The model must assess every inspected file for all justified improvements, not
|
|
620
|
+
stop after one file or the static warning list. Retained guidance needs an
|
|
621
|
+
evidence-based rationale; nothing forces edits to already useful rules.
|
|
622
|
+
When supplied guidance or the project description identifies build, test or lint
|
|
623
|
+
tooling, proposed replacement guidance puts copy-pasteable commands (including
|
|
624
|
+
required flags) in an early `## Commands` section; it does not invent commands
|
|
625
|
+
without supporting evidence.
|
|
626
|
+
Copilot attaches repository-wide instructions, `AGENTS.md` and matching path-specific
|
|
627
|
+
instructions to every session, including custom agents. The `auto-loaded-reference`
|
|
628
|
+
warning flags lines in Copilot guidance and agents that tell the agent to read those
|
|
629
|
+
files, and notes targets that do not exist; the assessment proposes removing the pointer.
|
|
630
|
+
Claude and Gemini files are exempt because those hosts do not load `.github` guidance.
|
|
631
|
+
The `agent-only-context` advisory flags documents only one agent is told to read; when
|
|
632
|
+
they hold path conventions, a scoped `.github/instructions/<domain>.instructions.md` with
|
|
633
|
+
`applyTo` globs lets every session get them. The assessment also proposes new scoped files
|
|
634
|
+
when agents or root guidance carry path-specific conventions, and never invents rules for
|
|
635
|
+
documents it was not given. Generated profiles no longer list automatically loaded files
|
|
636
|
+
under "Reuse existing guidance".
|
|
637
|
+
For focused guidance the assessment prefers `.github/instructions/<domain>.instructions.md`:
|
|
638
|
+
the cloud agent, code review and VS Code, Visual Studio and JetBrains chat all load it,
|
|
639
|
+
and `applyTo` globs can target file types across directories. Nested `AGENTS.md` is
|
|
640
|
+
proposed only for directory scope in repositories that also use non-Copilot agents;
|
|
641
|
+
Visual Studio and JetBrains Copilot chat do not read it. Neither format has published
|
|
642
|
+
evidence of better task outcomes; the choice is about host coverage and scoping.
|
|
643
|
+
See the [README sources](../README.md#sources-behind-guidance-assessment).
|
|
644
|
+
|
|
645
|
+
The scanner reads visible non-ignored instruction files, README/CONTRIBUTING
|
|
646
|
+
documents and package manifests. It checks at most 32 instruction files, 12
|
|
647
|
+
reference documents and 20 manifests, with 64 KB per file and 512 KB total;
|
|
648
|
+
omissions are reported. Root guidance is prioritized before specialist profiles.
|
|
649
|
+
At most twelve signal details **per instruction file** are included, so a noisy
|
|
650
|
+
root file cannot exhaust the coverage for other files. `signalsOmitted` reports
|
|
651
|
+
remaining detected warnings rather than presenting the sample as complete.
|
|
652
|
+
These are inspection resource limits, **not quality thresholds**. No project
|
|
653
|
+
script or linked URL is executed. The remaining configurable word budgets are
|
|
654
|
+
readability constraints, not research-derived quality scores.
|
|
655
|
+
|
|
656
|
+
## Launch preflight, limits and stop controls
|
|
657
|
+
|
|
658
|
+
`crewbie preflight [--batch-id ID] [--json]` performs read-only inspection of the
|
|
659
|
+
managed tasks, current approval, dependencies, repository capacity, sources,
|
|
660
|
+
specialist files and live account models. It exposes requested models and profile
|
|
661
|
+
revisions, not proof of effective cloud-runtime selection. Missing catalog
|
|
662
|
+
entitlement stops launch; no model fallback or automatic paid retry is permitted.
|
|
663
|
+
Both issue assignment and review/correction launches repeat the guards.
|
|
664
|
+
|
|
665
|
+
`modelProfile` is `balanced` by default; `init --model-profile economy|balanced|quality`
|
|
666
|
+
sets the reviewed selection policy. Capability comes before price in every profile,
|
|
667
|
+
including non-code work. Existing models and explicit overrides remain unchanged.
|
|
668
|
+
An LLM's suitability rationale is a proposal, not a benchmark certification.
|
|
669
|
+
|
|
670
|
+
```json
|
|
671
|
+
{
|
|
672
|
+
"modelProfile": "balanced",
|
|
673
|
+
"execution": {
|
|
674
|
+
"maxLaunchesPerBatch": 20,
|
|
675
|
+
"maxAttemptsPerTask": 3
|
|
676
|
+
}
|
|
677
|
+
}
|
|
678
|
+
```
|
|
679
|
+
|
|
680
|
+
These configuration fields are optional for legacy configurations so parsing does
|
|
681
|
+
not change approved plan hashes. Absent limits use 20/3. Changing the profile alone
|
|
682
|
+
does not reassign models. Each Crewbie-initiated implementation/review launch
|
|
683
|
+
reserves one attempt under the existing repository dispatch lock **before** the
|
|
684
|
+
paid request. Initial launches, continuations and unknown request outcomes consume
|
|
685
|
+
allowance. Remote `crewbie/launches/<batch>/<task>/<issue>/...` tags retain the
|
|
686
|
+
ledger across local CLI sessions and Actions runs. Stable batch/task identities
|
|
687
|
+
keep the count across revisions. Do not delete or edit these refs to bypass limits.
|
|
688
|
+
When an approved re-plan replaces a task whose earlier issue is closed, Crewbie
|
|
689
|
+
publishes a new issue for it; that issue gets its own initial launch, which still
|
|
690
|
+
counts toward the shared task and batch allowances. An earlier issue that is still
|
|
691
|
+
open must be reconciled (closed) first.
|
|
692
|
+
The budget covers Crewbie requests, not the number of internal backend sessions,
|
|
693
|
+
tokens, monetary spend, manual PR follow-ups, onboarding, planning or nightly work.
|
|
694
|
+
|
|
695
|
+
**Upgrading an in-flight batch:** pre-alpha.10 launch claims have no trustworthy
|
|
696
|
+
complete attempt ledger. Further automatic launches in that batch stop visibly,
|
|
697
|
+
but existing sessions and completed work are not changed. Review each reported
|
|
698
|
+
issue's history (initial, correction, retry and uncertain requests), then record
|
|
699
|
+
the count without rebuilding the setup or republishing the tasks:
|
|
700
|
+
|
|
701
|
+
```powershell
|
|
702
|
+
crewbie budget --issue 42 --historical-attempts 2
|
|
703
|
+
crewbie budget --issue 42 --historical-attempts 2 --apply
|
|
704
|
+
crewbie preflight
|
|
705
|
+
```
|
|
706
|
+
|
|
707
|
+
The count is an explicit **human attestation**, not observed billing. It adds a
|
|
708
|
+
one-time immutable baseline consuming allowance; it cannot overwrite existing
|
|
709
|
+
history or refund attempts. Repeat for each unaccounted legacy claim that
|
|
710
|
+
preflight identifies. If you cannot establish its history, keep that batch
|
|
711
|
+
blocked rather than inventing a count. This does not repeat paid analysis.
|
|
712
|
+
|
|
713
|
+
`crewbie pause` previews a repository-wide gate; `--apply` requires a user with
|
|
714
|
+
write access and uses the same lock as dispatch/review. If dispatch holds the
|
|
715
|
+
lock, pause fails visibly: do not assume the repository is paused; retry once the
|
|
716
|
+
current operation ends. Once confirmed, no new controlled implementation/review
|
|
717
|
+
launches pass the gate. Already-running sessions continue. Planning and nightly
|
|
718
|
+
analysis have their own opt-in controls and are not paused by this command.
|
|
719
|
+
`crewbie resume --apply` removes only the gate, preserving claims and counters;
|
|
720
|
+
it does not itself dispatch.
|
|
721
|
+
|
|
722
|
+
`crewbie reapprove --issue N[,N...]` previews moving open task issues to their
|
|
723
|
+
owner's current `model` in `.crewbie/config.json`. `--apply` requires a user with
|
|
724
|
+
write access, rewrites only the model in the issue's task metadata (scope,
|
|
725
|
+
owner and dependencies are unchanged) and posts an execution approval bound to
|
|
726
|
+
the exact new title and body. It starts nothing; unclaimed tasks launch on the
|
|
727
|
+
next dispatch, and a task that already tried to start needs `crewbie:restart`.
|
|
728
|
+
|
|
729
|
+
`crewbie cancel --issue N --run-id ID` previews cancellation after verifying an
|
|
730
|
+
unambiguous closing Copilot PR and matching repository, branch, PR identity and
|
|
731
|
+
`dynamic` Actions run. `--apply` requires a user with write access, rechecks the run and
|
|
732
|
+
requests the documented Actions cancellation operation. The result distinguishes
|
|
733
|
+
an accepted request from a confirmed cancelled run. Native termination and
|
|
734
|
+
capacity are not assumed from the request; saved commits and claims remain.
|
|
735
|
+
An inaccessible/unsupported cancellation is an explicit error with instructions
|
|
736
|
+
to use GitHub's **Stop session** control. There is no force-cancel or replacement
|
|
737
|
+
fallback.
|
|
738
|
+
|
|
739
|
+
Sources: [GitHub session management](https://docs.github.com/en/copilot/how-tos/copilot-on-github/use-copilot-agents/manage-and-track-agents#stop-a-session)
|
|
740
|
+
and [Actions cancellation](https://docs.github.com/en/rest/actions/workflow-runs#cancel-a-workflow-run).
|
|
741
|
+
The documented [Agent Tasks API](https://docs.github.com/en/rest/agent-tasks/agent-tasks)
|
|
742
|
+
does not itself expose a task-cancellation endpoint; Crewbie does not invent one.
|
|
743
|
+
|
|
744
|
+
## Scheduling and recovery
|
|
745
|
+
|
|
746
|
+
### Autonomous review and correction
|
|
747
|
+
|
|
748
|
+
`publish --review-loop FILE --apply --watch` explicitly authorizes a bounded
|
|
749
|
+
review/correction workflow. It is a separate approval from the initial task
|
|
750
|
+
graph: it can review saved failing QA evidence rather than pretending a failed
|
|
751
|
+
session completed. Requirements, target ownership/model and existing issue
|
|
752
|
+
approval must still match. Final merges remain human-controlled.
|
|
753
|
+
|
|
754
|
+
The plan shape is:
|
|
755
|
+
|
|
756
|
+
```json
|
|
757
|
+
{
|
|
758
|
+
"schemaVersion": 1,
|
|
759
|
+
"reviewer": { "issue": 5, "issueDigest": "<exact issueDigest>" },
|
|
760
|
+
"targets": [
|
|
761
|
+
{
|
|
762
|
+
"issue": 3,
|
|
763
|
+
"pr": 7,
|
|
764
|
+
"issueDigest": "<exact issueDigest>",
|
|
765
|
+
"allowedPaths": ["frontend/src/", ".crewbie/team/frontend/"]
|
|
766
|
+
}
|
|
767
|
+
],
|
|
768
|
+
"maxRounds": 2
|
|
769
|
+
}
|
|
770
|
+
```
|
|
771
|
+
|
|
772
|
+
Compute each digest with the same exported `issueDigest(title, body)` used for
|
|
773
|
+
issue approval: SHA-256 of `JSON.stringify({title, body})`. Changing target scope,
|
|
774
|
+
paths or the budget creates a new plan; inspect any old running loop before
|
|
775
|
+
starting it. Paths ending in `/` authorize that subtree. Workflow, Git and
|
|
776
|
+
general Crewbie policy roots are excluded.
|
|
777
|
+
|
|
778
|
+
The reviewer commits a small machine-readable report to its own unmerged PR.
|
|
779
|
+
The coordinator validates exact target heads, findings, scope and native
|
|
780
|
+
profile/model evidence before publishing GitHub `COMMENT` reviews. The review
|
|
781
|
+
body names `crewbie-reviewer`, its native task, report PR and verdict. GitHub's
|
|
782
|
+
posting identity is the authorized coordinator account, not a fabricated agent
|
|
783
|
+
account. A clean automated review does not count as human approval.
|
|
784
|
+
|
|
785
|
+
Corrections use the documented Agent Tasks `base_ref`/`head_ref` continuation,
|
|
786
|
+
with the same approved specialist/model. State lives on the isolated
|
|
787
|
+
`crewbie/review-state/<digest>` branch, not in hot memory or on the application
|
|
788
|
+
default branch. Launch intent is persisted before a paid request. Unknown
|
|
789
|
+
outcomes are never automatically repeated. Immutable write-access-user-authored receipts
|
|
790
|
+
let ordinary reconciliation recognize an explicit chain of native tasks on one
|
|
791
|
+
PR; extra or concurrent unrecorded tasks remain ambiguous.
|
|
792
|
+
|
|
793
|
+
When the plan includes a tester target, completed implementation corrections
|
|
794
|
+
automatically trigger that specialist to refresh combined-head evidence before
|
|
795
|
+
re-review. Missing tester evidence does not suppress independently actionable
|
|
796
|
+
implementation findings. This authorizes verification within the original test
|
|
797
|
+
scope, not a weaker assertion or a broader application change.
|
|
798
|
+
|
|
799
|
+
The shared dispatch lock, repository-wide native task count and unresolved
|
|
800
|
+
assignment reservations limit concurrency. Continuations update the same
|
|
801
|
+
running/review/failed issue labels as initial dispatch.
|
|
802
|
+
Each correction round receives a fresh independent review. Required findings
|
|
803
|
+
remaining after the budget, scope changes, unverified models, cloud failures or
|
|
804
|
+
missing evidence stop visibly. Resume the exact plan after inspecting the cause.
|
|
805
|
+
Do not erase launch intent or claims to force a retry.
|
|
806
|
+
|
|
807
|
+
### Initial task graph
|
|
808
|
+
|
|
809
|
+
Owner/status labels describe the approved task; prerequisite IDs remain the
|
|
810
|
+
source of truth. Polling recovers missed GitHub events. Bot-applied labels are
|
|
811
|
+
not assumed to trigger another workflow, so publication also dispatches it.
|
|
812
|
+
Issue/PR attribution uses GitHub's authoritative closing references, not ordinary
|
|
813
|
+
timeline mentions. A review PR can discuss another task without becoming that
|
|
814
|
+
task's implementation PR or releasing its dependencies.
|
|
815
|
+
|
|
816
|
+
GitHub's cloud agent runs the Crewbie specialist as a subagent, and its final
|
|
817
|
+
summary replaces the specialist's own PR description. During attribution,
|
|
818
|
+
reconciliation recovers the specialist's last description from the PR edit
|
|
819
|
+
history (Copilot edits containing `Specialist: crewbie-<role>`), restores it as
|
|
820
|
+
the PR body and keeps Copilot's summary as a PR comment. This keeps the
|
|
821
|
+
specialist's voice, handoff and required headings in the PR. If no such edit
|
|
822
|
+
exists, the body is only attributed. Adopted personas apply to everything a specialist
|
|
823
|
+
writes for humans; memory files stay neutral.
|
|
824
|
+
|
|
825
|
+
With `publish --dispatch-local`, orchestration is one-shot: re-run publication
|
|
826
|
+
with the same approved batch to release newly ready work. Existing claims prevent
|
|
827
|
+
duplicate sessions; approval does not need to be repeated when the batch is
|
|
828
|
+
unchanged. This mode does not require an always-on local scheduler.
|
|
829
|
+
|
|
830
|
+
Add `--watch` for a bounded foreground reconciliation loop using the same local
|
|
831
|
+
credentials. `--poll-seconds` defaults to 30 and `--timeout-seconds` to 3600.
|
|
832
|
+
The exact published batch and approval are checked before each dispatch. Other
|
|
833
|
+
batches count toward capacity but are not launched by this command. Changed scope,
|
|
834
|
+
missing approval and API failures stop the loop; uncertain sessions retain their
|
|
835
|
+
claims. It does not bypass GitHub workflow-approval policy.
|
|
836
|
+
Local publication carries its confirmed issue IDs into reconciliation. Missing
|
|
837
|
+
entries in GitHub's freshly updated label index are fetched directly; their
|
|
838
|
+
contents and human approval are still checked before any launch.
|
|
839
|
+
Transient HTTP 502/503/504 reads (GET and explicit GraphQL queries) receive at most
|
|
840
|
+
two retries with bounded backoff. A Retry-After longer than 30 seconds stops
|
|
841
|
+
instead of retrying too early. Assignment, lock writes/deletes and mutations are
|
|
842
|
+
never retried automatically: their outcomes may be ambiguous.
|
|
843
|
+
|
|
844
|
+
The loop exits successfully when all selected tasks have completed cloud sessions
|
|
845
|
+
with linked PRs, or merged PRs. This is a handoff, not a passing-check or review
|
|
846
|
+
verdict. Human merge gates and failed work stop with exit code 2, as does timeout.
|
|
847
|
+
Timeout is checked between reconciliations; in-flight API calls retain their
|
|
848
|
+
normal request timeout. Stopping the watcher does not cancel remote sessions.
|
|
849
|
+
Resume with the same approved command after inspecting the reported condition.
|
|
850
|
+
|
|
851
|
+
The dispatcher obtains an atomic `refs/crewbie/dispatch-lock` Git ref, and a
|
|
852
|
+
`refs/crewbie/claims/<issue-number>` ref before each assignment. When another Crewbie
|
|
853
|
+
run holds the lock (for example, the dispatch and plan-release workflows both fire
|
|
854
|
+
on a planning merge), the run waits up to five minutes. If the lock is still held
|
|
855
|
+
and no Crewbie workflow is running, delete it with
|
|
856
|
+
`gh api -X DELETE repos/OWNER/REPO/git/refs/crewbie/dispatch-lock` and rerun. These
|
|
857
|
+
hold no transcripts or secrets. Claims deliberately survive unknown network outcomes;
|
|
858
|
+
repeated runs do not blindly start another paid session. Claims also identify
|
|
859
|
+
already-running work when a workflow restarts.
|
|
860
|
+
|
|
861
|
+
These operational refs (lock, pause, claims and the launch ledger) live under
|
|
862
|
+
`refs/crewbie/`, which is neither a branch nor a tag, so they never trigger a
|
|
863
|
+
repository's `on: push` workflows. Releases up to 0.1.0-alpha.43 wrote them as
|
|
864
|
+
`refs/tags/crewbie/*` tags, which fired push CI; those legacy tags are still read,
|
|
865
|
+
so existing claims and launch allowances keep counting. Once no pre-upgrade claim
|
|
866
|
+
is open, you may delete the legacy `crewbie/*` tags.
|
|
867
|
+
|
|
868
|
+
If a runner dies while holding the dispatcher lock, an administrator must
|
|
869
|
+
confirm no dispatch is active and remove only that reserved lock ref.
|
|
870
|
+
For a stalled issue, inspect its Copilot session and linked PR before any retry.
|
|
871
|
+
Only after confirming there is no active session may an administrator remove
|
|
872
|
+
that issue's specific claim and reapprove the task as needed. Never bulk-delete
|
|
873
|
+
claims or treat issue closure as successful implementation.
|
|
874
|
+
|
|
875
|
+
Reconciliation preserves unrelated labels. Failed/unmerged work blocks its
|
|
876
|
+
descendants. Claims survive session completion to prevent relaunch. A uniquely
|
|
877
|
+
correlated completed cloud task frees execution capacity while its PR awaits
|
|
878
|
+
review or after it is closed without merging. Closed work whose correlated task
|
|
879
|
+
is verified failed, timed out or cancelled also frees capacity; while its PR is
|
|
880
|
+
open the slot stays reserved for an authorized continuation. When Copilot itself
|
|
881
|
+
comments that it was unable to start working on the issue, no session ran:
|
|
882
|
+
capacity is freed and that sole launch does not count toward attempt or batch
|
|
883
|
+
allowances. The claim remains; a user with write access relaunches it with the
|
|
884
|
+
`crewbie:restart` label (see below). Closed-unmerged work stays failed,
|
|
885
|
+
retains its claim, and never satisfies a prerequisite. Missing, ambiguous, active or
|
|
886
|
+
inaccessible task telemetry retains capacity and reports why. A draft PR alone
|
|
887
|
+
is not proof that a session has finished.
|
|
888
|
+
An explicitly approved `kind: "review"` task can depend on completed sessions
|
|
889
|
+
with linked PRs; default implementation tasks still require merged prerequisites.
|
|
890
|
+
Changing a task's kind invalidates its approval like other scope changes.
|
|
891
|
+
|
|
892
|
+
### Feature branch, restart, Crewbie review and merge
|
|
893
|
+
|
|
894
|
+
- **Feature branch.** Every approved plan revision gets `crewbie/<plan>-<revision>`
|
|
895
|
+
(the first 8 characters of its digest, so a replan never reuses a branch), created from the
|
|
896
|
+
default branch when its first task launches. Tasks start from it and their PRs
|
|
897
|
+
target it; a task that depends on another waits until that task's PR merged
|
|
898
|
+
into the branch, review tasks included. GitHub links closing keywords only on PRs
|
|
899
|
+
into the default branch, so Crewbie finds task PRs through the issue timeline,
|
|
900
|
+
and task issues stay open with `crewbie:done` until the feature PR merges. Issues
|
|
901
|
+
published before feature branches carry no branch and are ignored by dispatch
|
|
902
|
+
and launch allowances; finish them by hand.
|
|
903
|
+
- **Restart.** A user with write access adds `crewbie:restart` to a task issue whose
|
|
904
|
+
previous session verifiably ended: Copilot reported it could not start, or its
|
|
905
|
+
task failed, timed out or was cancelled and its PR is closed. Dispatch, under
|
|
906
|
+
the lock, reserves a new ledger entry that counts as an attempt, comments the
|
|
907
|
+
attempt number with a `crewbie-restart` marker, reassigns Copilot and removes
|
|
908
|
+
the label. It refuses (and explains in a comment) for other labellers, closed
|
|
909
|
+
or unapproved issues, open PRs (push fixes to them, or close them first), sessions
|
|
910
|
+
not verified as ended and exhausted allowances. Without a free slot it waits.
|
|
911
|
+
- **Ready for review.** Copilot requests your review when its session finishes;
|
|
912
|
+
that triggers dispatch, which restores the specialist's description and then
|
|
913
|
+
marks the draft PR ready, so the description check runs on the final body, and
|
|
914
|
+
withdraws Copilot's review request because task PRs merge without one. If
|
|
915
|
+
that trigger waits for workflow approval (Copilot-actor runs can require it; see
|
|
916
|
+
*Settings → Copilot → Cloud agent → Actions workflow approval*), the hourly
|
|
917
|
+
reconcile does it instead.
|
|
918
|
+
- **Merge into the feature branch.** Dispatch merges a task PR into its feature
|
|
919
|
+
branch once the session completed, every check that ran on the head passed (the
|
|
920
|
+
newest run of each check counts) and GitHub reports no conflict. When CI ran
|
|
921
|
+
and a reviewer is configured, a trusted review of that exact task head must
|
|
922
|
+
also find no blocking acceptance-criteria issues or omitted patches; a
|
|
923
|
+
changes-requested or partial review leaves the PR for a fix or a human merge.
|
|
924
|
+
Crewbie needs no CI: with no checks (no CI, a branch filter such as
|
|
925
|
+
`branches: [main]`, or runs held for approval in Actions) it merges without a
|
|
926
|
+
task review and says so, because the feature PR, which Crewbie reviews and a
|
|
927
|
+
human tests and merges, is the gate. Crewbie pins
|
|
928
|
+
the head SHA and never bypasses branch protection. Pending
|
|
929
|
+
checks are re-evaluated on the next dispatch run, including the hourly schedule.
|
|
930
|
+
A PR that changes `.github/workflows/` is left for a human merge, because
|
|
931
|
+
workflows on the feature branch run with repository secrets for PRs into it.
|
|
932
|
+
`merge.method` accepts `merge` (default), `squash` or `rebase`.
|
|
933
|
+
- **Feature PR.** When every task of the plan merged, dispatch opens one PR from
|
|
934
|
+
the feature branch into the default branch, listing the tasks and closing their
|
|
935
|
+
issues. Only a human merges it. If it is closed without merging, Crewbie reports
|
|
936
|
+
that and does not reopen it.
|
|
937
|
+
- **Local testing.** `crewbie test [feature]` discovers open feature PRs and
|
|
938
|
+
in-progress `crewbie/...` branches from task metadata, matches by issue/PR
|
|
939
|
+
number, branch or title words, refuses dirty working trees, fetches and switches
|
|
940
|
+
to the feature branch, and runs configured `local.start` unless `--no-start` is
|
|
941
|
+
supplied. Use `--list` to see choices and `--json` for a machine-readable
|
|
942
|
+
selection/list. Interactive init suggests `local.start` from package scripts
|
|
943
|
+
(`dev`, optionally after `install:all`, or `start`) and `init --start "COMMAND"`
|
|
944
|
+
records an explicit command in the reviewed setup.
|
|
945
|
+
- **Crewbie review.** With `"review": { "enabled": true, "role": "<role id>" }`,
|
|
946
|
+
dispatch starts `crewbie-review.yml` once for each task PR head whose CI ran
|
|
947
|
+
and passed, reviewing that task's own criteria and diff before auto-merge. It also
|
|
948
|
+
reviews each head of the combined feature PR. The
|
|
949
|
+
reviewer reads its own charter and memory from the default branch plus the PR's
|
|
950
|
+
API diff (the PR's code is never checked out), runs tool-free in Copilot CLI
|
|
951
|
+
with the role's model, and posts one PR comment: a verdict, a summary and
|
|
952
|
+
findings marked blocking or minor, in the reviewer's voice. Any blocking finding
|
|
953
|
+
makes the verdict "changes". Dispatch trusts only comments posted by that
|
|
954
|
+
default-branch workflow run, started by a user with write access, for the PR's current head.
|
|
955
|
+
On "changes", push fixes to the feature branch for a fresh review, or merge
|
|
956
|
+
anyway if you disagree.
|
|
957
|
+
A failed review run is reported with its link and not retried automatically;
|
|
958
|
+
re-run it from Actions. Patches that do not fit the Copilot CLI prompt are
|
|
959
|
+
listed as not reviewed.
|
|
960
|
+
- **Feature-PR fix comments.** A write-access human can comment `/crewbie fix`
|
|
961
|
+
on the feature PR, optionally followed by notes, after a Crewbie
|
|
962
|
+
changes-requested review. `/crewbie revise` has the same meaning on feature PRs
|
|
963
|
+
and keeps its planning-revision meaning on planning PRs. The fix job prefilters
|
|
964
|
+
comments for `/crewbie` and an OWNER, MEMBER or COLLABORATOR association;
|
|
965
|
+
Crewbie still checks the exact command, non-bot identity and actual write access.
|
|
966
|
+
Crewbie handles each comment ID once, first tries to merge the default
|
|
967
|
+
branch into the feature branch
|
|
968
|
+
through the GitHub merges API. Each finding goes to the specialist whose merged
|
|
969
|
+
task PR changed the affected file (falling back to directory overlap, then the
|
|
970
|
+
first non-review task owner), and each specialist gets one fix task with only
|
|
971
|
+
its own findings plus the notes. Notes-only requests get one task for the owner
|
|
972
|
+
with the most merged changes. When GitHub reports conflicts, the task of the
|
|
973
|
+
specialist with the most findings resolves them first and the other fix tasks
|
|
974
|
+
depend on it. Launches go through the normal dispatch and launch-allowance
|
|
975
|
+
checks. Crewbie adds the issues to the feature PR's `Closes #...` list and
|
|
976
|
+
dispatches the guarded workflow; when the fix PRs merge into the feature branch,
|
|
977
|
+
dispatch requests another Crewbie review of the new head. A fix PR adding or
|
|
978
|
+
changing a workflow file (other than matching the default branch) is left for a
|
|
979
|
+
human merge.
|
|
980
|
+
## Nightly learning and bounded history
|
|
981
|
+
|
|
982
|
+
Set `nightly.enabled` to true in a reviewed setup proposal and install it.
|
|
983
|
+
The enabled workflow defaults to 02:37 UTC; GitHub schedules can be delayed. Disabled
|
|
984
|
+
installations have only a manual maintenance trigger, not a nightly cron. The input
|
|
985
|
+
cap defaults to 20 new records; preparation, analysis and publication have 3/9/3
|
|
986
|
+
minute timeouts (15 minutes of job execution in total). Neither is a
|
|
987
|
+
guaranteed monetary ceiling.
|
|
988
|
+
|
|
989
|
+
Only `nightly.allowedPaths` can be proposed. Broader access to an existing
|
|
990
|
+
constitution, `AGENTS.md`, or scoped instructions requires explicit configuration.
|
|
991
|
+
The analysis receives bounded snapshots and has no available tools or GitHub
|
|
992
|
+
write credential. A separate deterministic step checks paths, hashes, evidence,
|
|
993
|
+
word limits and obvious secret patterns before publishing; human review remains
|
|
994
|
+
essential.
|
|
995
|
+
|
|
996
|
+
Specialists close each task by assessing **handoff knowledge and reusable lessons
|
|
997
|
+
separately**. New contracts, integration constraints, decisions and limitations
|
|
998
|
+
needed by dependent tasks belong in relevant hot/index or linked topic memory,
|
|
999
|
+
even without a general lesson. Link discoverable implementation details instead
|
|
1000
|
+
of copying them. A no-update handoff needs a reason or a pointer to the exact
|
|
1001
|
+
existing memory section; "no new durable lesson" alone is insufficient.
|
|
1002
|
+
Approved memory paths can be changed on the work branch for human review;
|
|
1003
|
+
existing accepted policy is preserved.
|
|
1004
|
+
Memory entries carry no "proposed" marker; merging the PR is the review. Out-of-scope
|
|
1005
|
+
lessons go in one PR comment starting with `<!-- crewbie-memory-proposal -->`,
|
|
1006
|
+
with target path, lesson, reason and evidence. The collector includes bounded
|
|
1007
|
+
proposals from the Copilot bot or users with write access in nightly input. It does
|
|
1008
|
+
not promote comments to policy or force repetitive activity logs for routine work.
|
|
1009
|
+
|
|
1010
|
+
On an implementation PR, mention `@copilot` with specific changes or a targeted
|
|
1011
|
+
handoff request. GitHub continues using the same custom agent on its PR; another
|
|
1012
|
+
session can consume AI credits. Review and merge the new changes. On a planning
|
|
1013
|
+
PR, use the bounded `revise-plan` flow instead so provenance and manifest hashes
|
|
1014
|
+
are regenerated.
|
|
1015
|
+
|
|
1016
|
+
Own hot/index history and shared decisions are included. Human-closed improvement
|
|
1017
|
+
PRs become feedback; the agent's own open-PR bookkeeping is not fresh evidence.
|
|
1018
|
+
Affected specialists' hot/index files and charters are included with
|
|
1019
|
+
LF-normalized text hashes, so Git's LF/CRLF conversion does not create false
|
|
1020
|
+
conflicts. Installer ownership and proposal validation use the same normalization
|
|
1021
|
+
and accept legacy LF/CRLF fingerprints; real content edits still stop updates.
|
|
1022
|
+
Approval/source digests and explicit instruction-adoption hashes remain unchanged.
|
|
1023
|
+
Up to five
|
|
1024
|
+
relevant cold/archive links per role are selected by index-label keyword matches.
|
|
1025
|
+
Unrelated history stays unloaded; missing context is a reason to defer a change,
|
|
1026
|
+
not to invent it. Shared links can use `../../decisions/cold/topic.md` from a
|
|
1027
|
+
role index, or `cold/topic.md` from `.crewbie/decisions/index.md`; up to five
|
|
1028
|
+
shared topics are selected the same way. `.crewbie/decisions/hot.md` is always
|
|
1029
|
+
loaded and bounded by the hot limit.
|
|
1030
|
+
Operational cursors live on the orphan `crewbie/runtime` branch, separately from
|
|
1031
|
+
human-facing memory. It records the latest reviewed fingerprint/outcome per work
|
|
1032
|
+
item, so no-change analysis can advance without opening a pointless PR.
|
|
1033
|
+
|
|
1034
|
+
The reserved proposal branch is `crewbie/improvements`. Concurrent file changes,
|
|
1035
|
+
a branch behind its base, or a leftover closed proposal branch stop updates for
|
|
1036
|
+
human reconciliation rather than force-pushing over work. Merge, close and
|
|
1037
|
+
curate the active proposal when its combined description exceeds the readability
|
|
1038
|
+
budget. After reviewing a closed proposal, remove its reserved branch before
|
|
1039
|
+
creating another; accepted/rejected evidence remains in PR history and the cursor.
|
|
1040
|
+
|
|
1041
|
+
Cold/archive topics are loaded explicitly through the index, not all at once.
|
|
1042
|
+
The `limits` configuration can record deliberate word-budget exceptions:
|
|
1043
|
+
`spec`, `hot`, `constitution`, `topic`, and `pr`.
|
|
1044
|
+
Charters are bounded only by GitHub's 30,000-character agent prompt maximum; role
|
|
1045
|
+
and shared indexes have no word limit. Legacy `charter`, `index` and
|
|
1046
|
+
`decisions` entries are ignored.
|
|
1047
|
+
There is no silent truncation. Oversized external PR feedback is explicitly
|
|
1048
|
+
omitted with a source link, not partially presented as a complete summary.
|
|
1049
|
+
|
|
1050
|
+
The report workflow also checks Copilot PR descriptions using trusted
|
|
1051
|
+
default-branch code: the body is not empty, telemetry is not duplicated and, when
|
|
1052
|
+
`limits.pr` is set, the length fits. Headings are guidance, not a gate: a missing
|
|
1053
|
+
heading would block auto-merge after the session ended, with nobody left to fix
|
|
1054
|
+
it. There is no default word limit because none is evidence-backed. Human review
|
|
1055
|
+
still evaluates rationale and evidence.
|
|
1056
|
+
|
|
1057
|
+
Copilot's final session summary replaces the specialist's own PR description.
|
|
1058
|
+
During attribution Crewbie restores the specialist's last description (its own
|
|
1059
|
+
Copilot edit containing `Specialist: crewbie-<role>`) as the PR body and saves
|
|
1060
|
+
Copilot's summary as a PR comment first. If that comment cannot be saved, the
|
|
1061
|
+
summary stays in the body.
|
|
1062
|
+
|
|
1063
|
+
Installation proposes a short PR template only when no existing template is
|
|
1064
|
+
found in GitHub's supported repository locations. Existing templates are preserved.
|
|
1065
|
+
Agents do not always follow the heading guidance (a release review used Findings
|
|
1066
|
+
and "Established here" instead of Why and Checks), which is why headings are not
|
|
1067
|
+
enforced.
|
|
1068
|
+
|
|
1069
|
+
The team test confirmed that a specialist can produce a correctly structured
|
|
1070
|
+
final handoff while GitHub independently generates a different PR body. Treat
|
|
1071
|
+
persisted PR metadata as a coordinator finalization step, not proof of what the
|
|
1072
|
+
specialist reported. Inspect the native session handoff (for example,
|
|
1073
|
+
`gh agent-task view SESSION_ID --repo owner/repo --log`) and actual CI.
|
|
1074
|
+
Use `publish --pr N --proposal handoff.json` to preview a concise correction;
|
|
1075
|
+
`--apply` requires a user with write access and matching `headSha`/`beforeHash`.
|
|
1076
|
+
The proposal's `body` must pass the normal format/length check. This changes only
|
|
1077
|
+
PR metadata, not code, approvals or merge state, and avoids another paid
|
|
1078
|
+
implementation run merely to repair prose.
|
|
1079
|
+
|
|
1080
|
+
Cloud hosts inject the active charter and may protect its file path from agent
|
|
1081
|
+
tools. Record injection separately from file-read attestations; respect those
|
|
1082
|
+
restrictions. Native Copilot Memory is a separate platform feature, not Crewbie's
|
|
1083
|
+
reviewed role history or approval of proposed shared decisions.
|
|
1084
|
+
|
|
1085
|
+
## Reports and privacy
|
|
1086
|
+
|
|
1087
|
+
The report workflow uploads an Actions artifact. That requires repository access
|
|
1088
|
+
to download; it is not an in-browser hosted site. Check repository and artifact
|
|
1089
|
+
retention settings before treating it as a historical ledger.
|
|
1090
|
+
|
|
1091
|
+
To use Pages, configure the site and explicitly set `CREWBIE_PAGES_MODE`.
|
|
1092
|
+
The publishing job checks the actual Pages `public` flag and fails closed on a
|
|
1093
|
+
mismatch. A private source repository alone does not make a Pages site private.
|
|
1094
|
+
Public Pages requires an explicit administrator choice even for open-source code.
|
|
1095
|
+
|
|
1096
|
+
Imported records must include evidence for non-null usage/observed-model fields.
|
|
1097
|
+
Token counts, credits and currency are separate measurements. Example records
|
|
1098
|
+
are synthetic, not real billing. The report escapes all data and works without
|
|
1099
|
+
external scripts, fonts, network access or storage.
|
|
1100
|
+
|
|
1101
|
+
## Verification and release
|
|
1102
|
+
|
|
1103
|
+
`npm test` builds strict TypeScript and exercises CLI installation/approval,
|
|
1104
|
+
request contracts, dependencies, retry claims, memory, ADO, maintenance safety,
|
|
1105
|
+
workflow YAML, and the dashboard's interactive DOM. CI is configured for current
|
|
1106
|
+
Node 22.12+ on Windows, macOS and Linux. Onboarding transport also exercises the
|
|
1107
|
+
bundled SDK runtime against a loopback-only synthetic provider without paid model
|
|
1108
|
+
calls.
|
|
1109
|
+
|
|
1110
|
+
Before broader release, run a consenting personal/organization account matrix:
|
|
1111
|
+
select the actual profile/model, inspect memory-read attestations and issue/PR
|
|
1112
|
+
linkage, exercise maintenance authentication, and confirm private publishing.
|
|
1113
|
+
No live test is implied by a fixture passing. Registry publishing, trademark
|
|
1114
|
+
clearance, and paid cloud runs require the project owner's separate decision.
|