@rallycry/conveyor-skills 1.0.10 → 1.0.12
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -0
- package/package.json +1 -1
- package/skills/conveyor-build/SKILL.md +14 -5
- package/skills/conveyor-release-review/SKILL.md +372 -0
- package/skills/conveyor-release-review/references/dimensions.md +289 -0
- package/skills/conveyor-release-review/references/report-and-cards.md +183 -0
- package/skills/conveyor-workflows/SKILL.md +12 -0
package/README.md
CHANGED
|
@@ -17,6 +17,7 @@ normal dependency bumps: no git submodules, no manual syncing.
|
|
|
17
17
|
| `conveyor-start` | `/conveyor-start <card>` | Hand a planned card to a cloud pod, confirm the environment came up, report where to watch it (local surface only) |
|
|
18
18
|
| `conveyor-build` | `/conveyor-build <card>` | Execute a planned card to a PR — one task or a whole feature-branch pack |
|
|
19
19
|
| `conveyor-review` | `/conveyor-review <card>` | Review the PR against its plan and render one verdict, with risk |
|
|
20
|
+
| `conveyor-release-review` | `/conveyor-release-review [release]` | Audit everything in a pending release, keep only findings backed by evidence that survive an adversarial pass, and file them by priority as a blocker pack and a suggestions pack (local surface only) |
|
|
20
21
|
| `conveyor-consensus` | `/conveyor-consensus <question>` | Sweep the sources a project actually has, count what happened, score the proposals against those counts, and attach an HTML verdict to the card |
|
|
21
22
|
| `conveyor-local-loop` | `/loop /conveyor-local-loop` | Run this machine as a serial local agent: claim Open cards and packs, build to PR, repeat |
|
|
22
23
|
| `conveyor-prune` | `/conveyor-prune [project\|all]` | Sweep every Planning/Open card, give each one disposition (cancel, park on hold, merge into a pack, keep), apply only the confirmed rows (local surface only) |
|
package/package.json
CHANGED
|
@@ -204,7 +204,14 @@ runner: the pod can sleep mid-run and nothing will wake you.
|
|
|
204
204
|
**A UI-visible change needs visual proof before the PR**, in both environments:
|
|
205
205
|
capture a screenshot (static) or a short recording (interaction) with the host
|
|
206
206
|
repo's own tooling and attach it with `mcp__conveyor__upload_attachment`, then
|
|
207
|
-
embed the returned URL in the PR body. A UI PR without it is incomplete.
|
|
207
|
+
embed the returned URL in the PR body. A UI PR without it is incomplete. Pass
|
|
208
|
+
`showcase: true` for the capture that shows the finished result (the after
|
|
209
|
+
shot, the final recording); leave it off for before shots and scoping captures.
|
|
210
|
+
Showcase files post to chat, render first on the card, and land in the PR body
|
|
211
|
+
under `## Showcase`. A re-capture of something already on the card (a new cut
|
|
212
|
+
of the same recording) passes `supersedes: "<old fileId>"` from
|
|
213
|
+
`list_task_files`: the new file takes the old one's star and tags, and the old
|
|
214
|
+
one collapses under it as history instead of showing twice.
|
|
208
215
|
|
|
209
216
|
**Manual tests are for human eyes, and few.** Before the PR, record with
|
|
210
217
|
`mcp__conveyor__set_manual_tests` only what a person must see in the running
|
|
@@ -262,10 +269,12 @@ Instead: post the answer, config, or findings with
|
|
|
262
269
|
`mcp__conveyor__upload_attachment` (any file type, up to 25MB), and complete
|
|
263
270
|
the card directly with `force_update_task_status("Complete")` — there is no PR
|
|
264
271
|
or review step for a no-code task. Never publish a deliverable as an off-card
|
|
265
|
-
link; the card is where it belongs. When a revision
|
|
266
|
-
upload,
|
|
267
|
-
`
|
|
268
|
-
|
|
272
|
+
link; the card is where it belongs. When a revision replaces an earlier
|
|
273
|
+
upload, upload it with `supersedes: "<old fileId>"` (the id from
|
|
274
|
+
`list_task_files`) so the card shows one current version of each deliverable
|
|
275
|
+
with the history collapsed under it. Reach for
|
|
276
|
+
`mcp__conveyor__delete_attachment` (permanent) only for a copy that should not
|
|
277
|
+
survive at all.
|
|
269
278
|
|
|
270
279
|
When unsure, check the diff: a real diff means open a PR, no diff means finish
|
|
271
280
|
in chat and mark it Complete.
|
|
@@ -0,0 +1,372 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: conveyor-release-review
|
|
3
|
+
description: Audit a pending release before it ships — evaluate every change in it, back each finding with evidence (production or dev logs, a local reproduction, or a deterministic proof), grade findings by the project's own Priority levels, strike what an adversarial pass finds over-built, and file the survivors as two packs, release blockers and release suggestions. Use when the user says "/conveyor-release-review [release]", "review this release", "is this release safe to ship", "audit what's going out", or asks for a pre-release check of the dev-to-main gap. Local/MCP surface only; inside a pod the skill says so instead of failing. Never fixes, never starts a build, never merges.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Conveyor Release Review
|
|
7
|
+
|
|
8
|
+
Every change in a release already passed its own review. This audit exists for
|
|
9
|
+
what that process cannot see: changes that are fine alone and break **in
|
|
10
|
+
combination**, risk that only exists at release scale, and what happens at
|
|
11
|
+
deploy time.
|
|
12
|
+
|
|
13
|
+
It has two ways to fail and treats them as equal. Missing a real break ships
|
|
14
|
+
it. Manufacturing work from speculation buries the real break under cards
|
|
15
|
+
nobody will build, and teaches the team to ignore the next review. So nothing
|
|
16
|
+
here is a finding until evidence says so, and nothing is filed until a
|
|
17
|
+
reviewer whose job is to strike it has tried.
|
|
18
|
+
|
|
19
|
+
## Goal and finish line
|
|
20
|
+
|
|
21
|
+
**Goal:** every change in the release evaluated; every surviving finding
|
|
22
|
+
carrying its evidence and one of the project's Priority levels, filed in at
|
|
23
|
+
most two packs.
|
|
24
|
+
|
|
25
|
+
**Finish line:** the report — verdict on its first line — posted to the
|
|
26
|
+
conversation, and to the release card's chat when a release card exists.
|
|
27
|
+
|
|
28
|
+
**A turn may end only when:** (a) the finish line is true, or (b) an
|
|
29
|
+
`AskUserQuestion` is pending because the release scope is genuinely ambiguous.
|
|
30
|
+
Anything else — a ledger with blank rows, findings listed but not filed, packs
|
|
31
|
+
created with no report — is a stalled run, not a paused one.
|
|
32
|
+
|
|
33
|
+
**Runtime:** *Codex CLI* — before step 0, create a thread goal with your goal
|
|
34
|
+
tool: "Post a release verdict for <release> with every change evaluated and
|
|
35
|
+
every surviving finding filed." Mark it complete only when the finish line is
|
|
36
|
+
true, never because the budget is low. *Claude Code* — there is no goal tool;
|
|
37
|
+
the finish line above is the loop.
|
|
38
|
+
|
|
39
|
+
> **Environment — this skill does not apply inside a pod.** Filing the packs
|
|
40
|
+
> needs `create_task`, `manage_priorities`, `add_dependency` against another
|
|
41
|
+
> card, and `update_task` with `priorityValue` — all of which exist only on
|
|
42
|
+
> the external `conveyor-mcp` surface. If
|
|
43
|
+
> `mcp__conveyor__get_connection_context` reports a task binding and no user
|
|
44
|
+
> account, you are in a pod: say so and stop. Do not hunt for an equivalent.
|
|
45
|
+
|
|
46
|
+
## Ground rules
|
|
47
|
+
|
|
48
|
+
- **All Conveyor tools fully-qualified** — `mcp__conveyor__get_task`, not
|
|
49
|
+
`get_task`. Most are deferred; load every one you expect to need in a single
|
|
50
|
+
ToolSearch `select:` up front.
|
|
51
|
+
- **Pass `projectId` explicitly on every call.** The connection's default
|
|
52
|
+
project may not be the repo you are standing in.
|
|
53
|
+
- **Findings become cards; fixes get their own PRs.** Never fix anything inside
|
|
54
|
+
the review, however small. A fix made here ships unreviewed in the release
|
|
55
|
+
you were asked to judge.
|
|
56
|
+
- **This skill never moves a release.** No `start_task`, `create_release`,
|
|
57
|
+
`add_to_release`, `approve_and_merge_pr`, or verdict tool. Holding or
|
|
58
|
+
shipping is a human's call; your job is to make the right call obvious.
|
|
59
|
+
- **Production is read-only.** Logs, and SELECT-only data reads through the
|
|
60
|
+
project's own sanctioned path. Never a write, never a hand-exported
|
|
61
|
+
credential.
|
|
62
|
+
- **The project's rules outrank this skill.** Gate commands, log queries, data
|
|
63
|
+
access and any release checklist live in the host repo's CLAUDE.md. A wrapper
|
|
64
|
+
skill may add dimensions (see the end of this file); it cannot lower the
|
|
65
|
+
evidence bar.
|
|
66
|
+
|
|
67
|
+
## 0 — Discover the project
|
|
68
|
+
|
|
69
|
+
Which levels, branches and evidence sources exist is a per-project fact.
|
|
70
|
+
Guessing wastes the review either way.
|
|
71
|
+
|
|
72
|
+
1. `mcp__conveyor__get_connection_context`, then match the project to the
|
|
73
|
+
repo's `git remote` (`mcp__conveyor__list_projects` when there is no
|
|
74
|
+
default). Copy the project id exactly.
|
|
75
|
+
2. `mcp__conveyor__manage_priorities { action: "list" }` — the priority
|
|
76
|
+
ladder. The level with the **lowest `value`** is the **blocker level**;
|
|
77
|
+
every other level is a **suggestion level**. Use the project's own names in
|
|
78
|
+
every card and in the report. The seeded defaults are High, Medium and Low,
|
|
79
|
+
and the rest of this skill uses those names for readability. A project with
|
|
80
|
+
no priorities configured gets High / Medium / Low as text labels and no
|
|
81
|
+
priority writes.
|
|
82
|
+
3. The dev branch and the default branch, from the repo and its CLAUDE.md.
|
|
83
|
+
4. **Evidence sources — record what exists AND what does not:**
|
|
84
|
+
`mcp__conveyor__list_project_integrations`, whether
|
|
85
|
+
`mcp__conveyor__query_gcp_logs` / `mcp__conveyor__query_grafana_logs`
|
|
86
|
+
answer, the repo's local stack or mock scripts (`package.json` scripts such
|
|
87
|
+
as `dev:mock` or `dev`, or a repo skill that stands the app up), its scoped
|
|
88
|
+
test commands, and its sanctioned read-only data path.
|
|
89
|
+
5. `mcp__conveyor__list_tags` once — the vocabulary, and whether a `release`
|
|
90
|
+
tag exists (step 6 uses it, and never creates it).
|
|
91
|
+
|
|
92
|
+
Announce one line naming the release, the priority ladder, and the evidence
|
|
93
|
+
sources you have and lack. Then start.
|
|
94
|
+
|
|
95
|
+
## 1 — Scope the release
|
|
96
|
+
|
|
97
|
+
Anchor to a concrete diff before reading any code.
|
|
98
|
+
|
|
99
|
+
1. **Find the release.** An explicit version, branch or range in the argument
|
|
100
|
+
wins. Otherwise `mcp__conveyor__search_tasks` for the release card
|
|
101
|
+
(`Release <version>`, or the `release` tag, in ReviewPR or InProgress);
|
|
102
|
+
`mcp__conveyor__get_task` on it gives the branch and PR number. No release
|
|
103
|
+
cut → review the dev-to-default gap and label the run **pre-release**.
|
|
104
|
+
2. **Fetch with explicit refspecs.** Many checkouts only track the dev branch,
|
|
105
|
+
so an unfetched ref silently reviews the wrong range.
|
|
106
|
+
|
|
107
|
+
```bash
|
|
108
|
+
git ls-remote --heads origin 'release/*'
|
|
109
|
+
git fetch origin <default>:refs/remotes/origin/<default> \
|
|
110
|
+
<dev>:refs/remotes/origin/<dev> \
|
|
111
|
+
'refs/heads/release/*:refs/remotes/origin/release/*'
|
|
112
|
+
TIP=origin/release/<version> # or origin/<dev> when pre-release
|
|
113
|
+
git rev-parse "$TIP" # the SHA this verdict is for
|
|
114
|
+
git log --merges --oneline origin/<default>.."$TIP"
|
|
115
|
+
git diff origin/<default>..."$TIP" --name-only
|
|
116
|
+
gh pr list --base <dev> --state merged \
|
|
117
|
+
--search "merged:>$(git log -1 --format=%cs origin/<default>)"
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
The `gh pr list` cross-check catches squash and rebase merges the merge log
|
|
121
|
+
misses. Use `--name-only`, never `--stat`: truncated paths hide files.
|
|
122
|
+
3. **Record the tip SHA.** A full release keeps absorbing the dev branch until
|
|
123
|
+
it merges, so the tip moves under you. The verdict is for the SHA you
|
|
124
|
+
recorded. If it moved by report time, review the delta or list the commits
|
|
125
|
+
you did not review.
|
|
126
|
+
4. **Map every PR to its card.** `conveyor/<slug>-<id>` branch names embed the
|
|
127
|
+
card slug; otherwise `search_tasks` by PR title. `get_task` gives the plan,
|
|
128
|
+
which is the intent. Reviewing a diff without its intent is how a
|
|
129
|
+
single-interpretation solution gets waved through twice. A PR with no card
|
|
130
|
+
is reviewed on its diff alone and noted as such.
|
|
131
|
+
5. **Build the coverage ledger** — one row per PR or card, one row per
|
|
132
|
+
dimension in step 2. Every row ends `clean`, `finding`, or
|
|
133
|
+
`not reviewable (why)`. "Evaluate everything" means the ledger has no blank
|
|
134
|
+
rows when you finish. On a large release, fan subagents out per area and
|
|
135
|
+
give each the ground rules and the evidence bar verbatim; the ledger is
|
|
136
|
+
still the completeness check, and it stays with you.
|
|
137
|
+
|
|
138
|
+
## 2 — Review every dimension
|
|
139
|
+
|
|
140
|
+
Work all of them. Each gets a coverage note in the report even when the note
|
|
141
|
+
is "nothing found", so coverage is auditable. What to look for, the cheapest
|
|
142
|
+
evidence, and a worked example of a real finding next to one that should be
|
|
143
|
+
struck are in [references/dimensions.md](references/dimensions.md).
|
|
144
|
+
|
|
145
|
+
| # | Dimension | The question |
|
|
146
|
+
| --- | --- | --- |
|
|
147
|
+
| 1 | **Cross-PR interaction** | Which files, symbols, tables and shared packages did two or more PRs touch, and what does their combined diff do? This is the headline check. |
|
|
148
|
+
| 2 | **Data and migrations** | Destructive steps, locks on large tables, ordering between migrations, and the window where old code runs against the new schema. |
|
|
149
|
+
| 3 | **Rollback and off-switches** | Which behavior changes have no flag, and what is the rollback for each risky change? |
|
|
150
|
+
| 4 | **Deploy and config** | New env vars or secrets production lacks, CI or workflow changes, infra manifests, dependency major bumps. |
|
|
151
|
+
| 5 | **Contracts and compatibility** | API, socket, MCP or published-package contract changes, against clients that are already deployed. |
|
|
152
|
+
| 6 | **Security and access** | New methods or endpoints without an access check, secrets or personal data in logs, unbounded queries. |
|
|
153
|
+
| 7 | **Stubs and cleanup** | In the diff only: TODO / FIXME / HACK, debug logging, `.only` or skipped tests, commented-out code, placeholder copy, orphaned flags. |
|
|
154
|
+
| 8 | **Documented-contract drift** | `mcp__conveyor__get_tag` on each touched subsystem: did the release make an overview or a pinned invariant false? |
|
|
155
|
+
| 9 | **Verification gaps** | Risky changed paths with no test, manual tests on member cards left unapproved or rejected (`mcp__conveyor__list_manual_tests`), CI state on the release PR. |
|
|
156
|
+
| 10 | **Open signals** | Follow-ups promised in member-card chat that never became a card; incidents filed since the merge that implicate a release change. |
|
|
157
|
+
| + | **Project dimensions** | Whatever a wrapper skill or the invoking prompt added. Same bar. |
|
|
158
|
+
|
|
159
|
+
What comes out of this step is a list of **candidates**. None of them is a
|
|
160
|
+
finding yet.
|
|
161
|
+
|
|
162
|
+
## 3 — Evidence, or it is not a finding
|
|
163
|
+
|
|
164
|
+
A candidate becomes a finding only with at least one of these. Record the
|
|
165
|
+
exact query or command beside the result, so anyone can re-run it.
|
|
166
|
+
|
|
167
|
+
| Grade | What it is |
|
|
168
|
+
| --- | --- |
|
|
169
|
+
| **Observed** | Logs or data. `env: "dev"` logs show the release's own code running, because its changes already merged to the dev branch. `env: "prod"` logs and read-only data show current traffic, data shape and row counts — what the release will meet. |
|
|
170
|
+
| **Reproduced** | The project's local stack or mocks booted and the failure shown, or a scoped test or script that demonstrates the break. The steps go on the card. Nothing is committed. |
|
|
171
|
+
| **Proven** | A deterministic static proof with both ends cited as `file:line` — a migration drops a column that code still reads. Allowed only when the failure needs no runtime assumption. "If a user happens to…" is not a proof. |
|
|
172
|
+
|
|
173
|
+
- **Cheapest first:** logs, then a static proof, then a scoped test, then the
|
|
174
|
+
full local stack.
|
|
175
|
+
- **Evidence downgrades too.** A path with no traffic in thirty days is a
|
|
176
|
+
reason to lower a finding or strike it.
|
|
177
|
+
- **Time-box it.** A candidate that has cost fifteen minutes with no evidence
|
|
178
|
+
either way is Unverified. Move on.
|
|
179
|
+
- **Unverified candidates are never filed.** They go in the report's
|
|
180
|
+
Unverified list, each with the evidence that would settle it. One that would
|
|
181
|
+
be blocker-level if true is named in the verdict line as an open question,
|
|
182
|
+
so a human decides with their eyes open.
|
|
183
|
+
- **A missing source is reported, not substituted.** "No log access on this
|
|
184
|
+
project, so dimensions 4 and 10 rest on static reading" is a caveat the
|
|
185
|
+
reader needs.
|
|
186
|
+
|
|
187
|
+
## 4 — Priority
|
|
188
|
+
|
|
189
|
+
Grade by **consequence and evidence**, never by which dimension surfaced it.
|
|
190
|
+
|
|
191
|
+
| Level | Test | Filed in |
|
|
192
|
+
| --- | --- | --- |
|
|
193
|
+
| **High** (the blocker level) | Shipping as-is causes data loss or corruption, a security or access hole, a broken live flow with no workaround, or a deploy that fails or cannot be rolled back. Observed, Reproduced or Proven only. | Release blockers |
|
|
194
|
+
| **Medium** | A real user-facing defect or regression with a workaround or limited reach, or a gap that will bite within a release or two. | Release suggestions |
|
|
195
|
+
| **Low** | Real but minor: cleanup, polish, small debt this release introduced. | Release suggestions |
|
|
196
|
+
|
|
197
|
+
**Priority is not risk.** Priority is how urgently the fix is needed. Risk is
|
|
198
|
+
how much important surface a change touches; identification and review set
|
|
199
|
+
it, and this skill never does.
|
|
200
|
+
|
|
201
|
+
The test for High is one question: *would I actually hold the release for
|
|
202
|
+
this?* If the honest answer is "no, but it should be fixed soon", it is
|
|
203
|
+
Medium. A blocker pack of twelve gets ignored, and then so does the one card
|
|
204
|
+
in it that mattered.
|
|
205
|
+
|
|
206
|
+
## 5 — The adversarial pass
|
|
207
|
+
|
|
208
|
+
Every candidate that reached step 4, **together with the fix you would
|
|
209
|
+
propose**, now faces a reviewer whose only job is to strike it. Use a fresh
|
|
210
|
+
subagent given the finding, its evidence and the tests below, told to argue
|
|
211
|
+
for striking. Where subagents are unavailable, switch stance explicitly and
|
|
212
|
+
write the case against each finding before deciding.
|
|
213
|
+
|
|
214
|
+
Strike when:
|
|
215
|
+
|
|
216
|
+
1. **Over-engineering.** The fix is bigger than the problem: a new
|
|
217
|
+
abstraction, option or framework for something a few lines solve, or a
|
|
218
|
+
defense against a future nobody has asked for.
|
|
219
|
+
2. **A niche path that already fails gracefully.** The path is rare — the
|
|
220
|
+
evidence shows little or no traffic — and its failure is already handled:
|
|
221
|
+
the error surfaces, no data is damaged, a retry works.
|
|
222
|
+
3. **A redundant code path.** The fix adds a second guard, fallback, retry or
|
|
223
|
+
validation for something already guaranteed upstream. Cite the guard that
|
|
224
|
+
already exists.
|
|
225
|
+
4. **The evidence does not support the claim.** It shows something nearby, or
|
|
226
|
+
something smaller.
|
|
227
|
+
5. **The release neither caused nor worsened it.** Pre-existing problems are
|
|
228
|
+
out of scope. Worth recording → `mcp__conveyor__create_suggestion`, never a
|
|
229
|
+
release-pack card.
|
|
230
|
+
6. **An open card already tracks it.** `mcp__conveyor__search_tasks`, two or
|
|
231
|
+
three keyword variants, all statuses. Name that card in the report instead
|
|
232
|
+
of filing a second one.
|
|
233
|
+
|
|
234
|
+
Four outcomes: **Keep**, **Downgrade**, **Shrink** — keep the finding, replace
|
|
235
|
+
the fix with the smallest change that removes the consequence — or **Strike**.
|
|
236
|
+
Every struck finding is listed in the report with its reason. Nothing
|
|
237
|
+
disappears silently, and a reader who disagrees with a strike can say so.
|
|
238
|
+
|
|
239
|
+
## 6 — File the packs
|
|
240
|
+
|
|
241
|
+
Survivors only. Created without asking, never started. Templates for the pack
|
|
242
|
+
parent and the finding card are in
|
|
243
|
+
[references/report-and-cards.md](references/report-and-cards.md).
|
|
244
|
+
|
|
245
|
+
**The bar for every card is the `/conveyor-plan` plan format**
|
|
246
|
+
([../conveyor-plan/references/plan-format.md](../conveyor-plan/references/plan-format.md)),
|
|
247
|
+
Builder briefing included, because `/conveyor-build` is what runs these cards.
|
|
248
|
+
Cite that format; do **not** invoke `/conveyor-plan`. You have already done
|
|
249
|
+
the research, and its create-and-promote steps would file each card twice.
|
|
250
|
+
|
|
251
|
+
`<label>` is `Release <version>` when a release is cut, and
|
|
252
|
+
`Pre-release <YYYY-MM-DD>` otherwise.
|
|
253
|
+
|
|
254
|
+
| | Release blockers | Release suggestions |
|
|
255
|
+
| --- | --- | --- |
|
|
256
|
+
| Holds | High findings | Medium and Low findings |
|
|
257
|
+
| Parent title | `<label> blockers` | `<label> suggestions` |
|
|
258
|
+
| Parent priority | the blocker level | the highest level among its children |
|
|
259
|
+
| Lands in | **Open** — the release is waiting on it | **Planning** — a human promotes what they want built |
|
|
260
|
+
|
|
261
|
+
In this order:
|
|
262
|
+
|
|
263
|
+
1. `mcp__conveyor__create_task` the parent in Planning, with `priorityValue`,
|
|
264
|
+
and `tags: ["release"]` when step 0 found that tag.
|
|
265
|
+
2. `mcp__conveyor__create_subtask` once per finding, full plan included, with
|
|
266
|
+
`dependsOn` only where one fix genuinely needs another to land first.
|
|
267
|
+
3. `mcp__conveyor__update_task { taskId: <child>, priorityValue }` per child.
|
|
268
|
+
`create_subtask` carries no priority field; this is the sanctioned route.
|
|
269
|
+
4. `mcp__conveyor__post_to_chat` on the parent — the sizing recommendation
|
|
270
|
+
and what the pack is for. Chat goes **before** any status move, because
|
|
271
|
+
identification reads it when the card leaves Planning.
|
|
272
|
+
5. Blockers only: `mcp__conveyor__update_subtask { status: "Open" }` per
|
|
273
|
+
child, then `mcp__conveyor__update_task { status: "Open" }` on the parent.
|
|
274
|
+
|
|
275
|
+
Then:
|
|
276
|
+
|
|
277
|
+
- **A pack of one is a card.** One survivor at a level → a single card titled
|
|
278
|
+
`<label> blocker: <finding>` or `<label> suggestion: <finding>`.
|
|
279
|
+
- **No survivors at a level → no pack at that level.** A clean release creates
|
|
280
|
+
nothing, and saying so is a complete result.
|
|
281
|
+
- **Blocking is advisory.** Nothing in Conveyor stops a release PR from
|
|
282
|
+
merging. With a release card present, make the hold impossible to miss:
|
|
283
|
+
`mcp__conveyor__add_dependency { taskId: <release card>, dependsOnSlugOrId: <blocker pack or card> }`,
|
|
284
|
+
which shows the release as blocked until the fixes merge, and a
|
|
285
|
+
`post_to_chat` on the release card whose **first line** is the hold notice.
|
|
286
|
+
If `add_dependency` is refused, report the error once and carry on.
|
|
287
|
+
- **Say how the fixes reach the release.** A full release absorbs whatever
|
|
288
|
+
merges to the dev branch, so merged fixes ship with it. A cherry-pick
|
|
289
|
+
release does not: a human adds them with Add to Release. Put the right
|
|
290
|
+
sentence in the blocker pack's plan.
|
|
291
|
+
- **No `priorityValue` on `update_task`?** The connected MCP server predates
|
|
292
|
+
the field. Prefix each title with its level — `[High] …` — and say in the
|
|
293
|
+
report that priorities need setting by hand.
|
|
294
|
+
- **On a tool error, stop and report the card rather than retrying it.** A
|
|
295
|
+
create that "failed" with a timeout may have landed; check with
|
|
296
|
+
`search_tasks` before creating it again.
|
|
297
|
+
|
|
298
|
+
## 7 — Verdict and report
|
|
299
|
+
|
|
300
|
+
The first line is the verdict, and it names the SHA:
|
|
301
|
+
|
|
302
|
+
- **`HOLD`** — one or more High findings.
|
|
303
|
+
- **`SHIP`** — none.
|
|
304
|
+
- **`SHIP WITH OPEN QUESTIONS`** — no confirmed High, but at least one
|
|
305
|
+
Unverified candidate that would be High if true.
|
|
306
|
+
|
|
307
|
+
Below it, in this order: the reviewed range and tip SHA; the packs, linked,
|
|
308
|
+
with a findings table; the **deploy checklist**; the ledger summary with a
|
|
309
|
+
note per dimension; struck findings with reasons; the Unverified list;
|
|
310
|
+
evidence sources used and unavailable; anything not reviewed.
|
|
311
|
+
|
|
312
|
+
The **deploy checklist** is every step a human must take at deploy — set a
|
|
313
|
+
secret, run a backfill, flip a flag, apply a manifest — whether or not it is a
|
|
314
|
+
finding. A release that is fine in code and fails because nobody knew to set a
|
|
315
|
+
variable is the most avoidable outage there is.
|
|
316
|
+
|
|
317
|
+
Post the report to the conversation, and with
|
|
318
|
+
`mcp__conveyor__post_to_chat` to the release card when one exists. Never
|
|
319
|
+
create a release card to hold it.
|
|
320
|
+
|
|
321
|
+
## Running it again
|
|
322
|
+
|
|
323
|
+
A release gets reviewed more than once: after the fixes land, or after the tip
|
|
324
|
+
moves.
|
|
325
|
+
|
|
326
|
+
- Find the existing packs by title (`search_tasks "<label> blockers"`) and add
|
|
327
|
+
new findings to them. Never create a second pack for the same release.
|
|
328
|
+
- Read the last reviewed SHA from the release card's chat and review the
|
|
329
|
+
delta, not the whole release again.
|
|
330
|
+
- Post the new verdict with the new SHA. The dependency on the release card
|
|
331
|
+
clears itself once the blocker pack merges to the dev branch.
|
|
332
|
+
|
|
333
|
+
## Common mistakes
|
|
334
|
+
|
|
335
|
+
- Reviewing the hunks. Combination bugs live in the overlap map, not in any
|
|
336
|
+
one diff.
|
|
337
|
+
- Skipping migrations because CI passed. CI ran them against a schema no
|
|
338
|
+
production data ever stressed.
|
|
339
|
+
- Grading by dimension. A cleanup item is not High because it touched a
|
|
340
|
+
migration, and a broken flow is not Low because it surfaced under "stubs".
|
|
341
|
+
- One mega-card of findings. One card per finding, or nothing is
|
|
342
|
+
independently schedulable.
|
|
343
|
+
- Re-litigating a merged PR. Per-PR review is `/conveyor-review`, and it
|
|
344
|
+
already happened.
|
|
345
|
+
- Diagnosing an incident here. An incident the release implicates is
|
|
346
|
+
evidence; its root cause belongs to `/conveyor-triage`.
|
|
347
|
+
|
|
348
|
+
## Wrapping this skill for one project
|
|
349
|
+
|
|
350
|
+
A repo with its own release concerns keeps them in a thin project skill that
|
|
351
|
+
invokes this one and names its extra dimensions and evidence tools, rather
|
|
352
|
+
than in a fork. The wrapper's shape is at the end of
|
|
353
|
+
[references/dimensions.md](references/dimensions.md). Project dimensions
|
|
354
|
+
become extra ledger rows, graded by the same rubric and held to the same
|
|
355
|
+
evidence bar and adversarial pass.
|
|
356
|
+
|
|
357
|
+
## Environment notes
|
|
358
|
+
|
|
359
|
+
- All Conveyor tools are **fully qualified** and mostly **deferred** — one
|
|
360
|
+
ToolSearch `select:` with every name you expect to need, up front.
|
|
361
|
+
- `search_tasks` returns 20 rows by default — pass `limit`.
|
|
362
|
+
- `gh` reads its own auth. An empty result with exit 0 from `gh pr list` or
|
|
363
|
+
`gh pr checks` can mean an expired token rather than "nothing there";
|
|
364
|
+
re-check before concluding.
|
|
365
|
+
- The log tools appear only for a project with the matching integration. If
|
|
366
|
+
they are absent, the project has none; say so instead of hunting.
|
|
367
|
+
|
|
368
|
+
## Improve This Skill
|
|
369
|
+
|
|
370
|
+
If this skill was insufficient or slowed the work down, file it with
|
|
371
|
+
`mcp__conveyor__create_suggestion` on the Conveyor project: the issue,
|
|
372
|
+
evidence, and proposed fix.
|
|
@@ -0,0 +1,289 @@
|
|
|
1
|
+
# Review dimensions — what to look for, and what to strike
|
|
2
|
+
|
|
3
|
+
The per-dimension checklist behind step 2 of
|
|
4
|
+
[conveyor-release-review](../SKILL.md). Read that first: the evidence bar, the
|
|
5
|
+
priority rubric and the adversarial pass all live there.
|
|
6
|
+
|
|
7
|
+
Each dimension below has four parts:
|
|
8
|
+
|
|
9
|
+
- **Look for** — the candidates worth raising.
|
|
10
|
+
- **Cheapest evidence** — the first thing to try, because step 3 works
|
|
11
|
+
cheapest-first.
|
|
12
|
+
- **A finding** — what a real one looks like once the evidence is in.
|
|
13
|
+
- **A strike** — a candidate from the same dimension that the adversarial pass
|
|
14
|
+
should kill, and why. These matter as much as the findings: most of what a
|
|
15
|
+
release review turns up belongs here.
|
|
16
|
+
|
|
17
|
+
The examples are illustrations, not a list to match against. File paths in
|
|
18
|
+
them are invented.
|
|
19
|
+
|
|
20
|
+
## 1. Cross-PR interaction
|
|
21
|
+
|
|
22
|
+
The headline check, and the one no single-card review was positioned to make.
|
|
23
|
+
|
|
24
|
+
**Look for**
|
|
25
|
+
|
|
26
|
+
- Build the **overlap map** first: from `git diff --name-only`, group files by
|
|
27
|
+
the PRs that touched them, then widen to symbols, database tables, caches,
|
|
28
|
+
queues and shared packages. Two PRs that never touch the same file can still
|
|
29
|
+
meet at a table or a cache key.
|
|
30
|
+
- Two green PRs on one surface: read their combined diff as a single change.
|
|
31
|
+
- A shared-package change rippling into a caller that merged **earlier** in
|
|
32
|
+
the release, written against the old behavior.
|
|
33
|
+
- Two PRs that each added a default, a retry, or a debounce to the same path.
|
|
34
|
+
|
|
35
|
+
**Cheapest evidence** — a scoped test that exercises both changes together;
|
|
36
|
+
the dev environment's logs for the shared path since the later PR merged.
|
|
37
|
+
|
|
38
|
+
**A finding** — PR A renames a status value in the shared package; PR B,
|
|
39
|
+
merged two days earlier, filters a query on the old name. Proven:
|
|
40
|
+
`shared/status.ts:41` no longer exports the value `board/query.ts:118`
|
|
41
|
+
compares against, so the filter matches nothing. High if the board is a live
|
|
42
|
+
flow.
|
|
43
|
+
|
|
44
|
+
**A strike** — two PRs both edited the same settings page, in different
|
|
45
|
+
sections, with no shared state. Overlap on a file is a reason to look, not a
|
|
46
|
+
finding. Struck: the evidence does not support the claim.
|
|
47
|
+
|
|
48
|
+
## 2. Data and migrations
|
|
49
|
+
|
|
50
|
+
Inherently untoggleable. CI runs migrations against a schema no production
|
|
51
|
+
data ever stressed.
|
|
52
|
+
|
|
53
|
+
**Look for**
|
|
54
|
+
|
|
55
|
+
- Destructive steps: dropped columns or tables, narrowed types, new `NOT NULL`
|
|
56
|
+
without a default, unique indexes on columns that may hold duplicates.
|
|
57
|
+
- Lock-heavy operations on large tables: index builds, rewrites, backfills
|
|
58
|
+
inside the migration transaction.
|
|
59
|
+
- **Ordering between migrations** from different PRs on the same table. They
|
|
60
|
+
merge cleanly and still break in sequence.
|
|
61
|
+
- The deploy window: old code running against the new schema, or new code
|
|
62
|
+
against the old one, depending on the project's deploy order.
|
|
63
|
+
- A backfill the code assumes has run.
|
|
64
|
+
|
|
65
|
+
**Cheapest evidence** — production row counts and a duplicate check through
|
|
66
|
+
the project's sanctioned read-only path; the migration SQL itself.
|
|
67
|
+
|
|
68
|
+
**A finding** — a migration adds a unique index on `(project_id, slug)`.
|
|
69
|
+
Observed: a read-only count on production returns 14 duplicate pairs. The
|
|
70
|
+
migration fails at deploy. High.
|
|
71
|
+
|
|
72
|
+
**A strike** — "the index build could lock the table". Observed: the table
|
|
73
|
+
holds 900 rows. Struck: niche path, and the evidence shows the cost is
|
|
74
|
+
milliseconds.
|
|
75
|
+
|
|
76
|
+
## 3. Rollback and off-switches
|
|
77
|
+
|
|
78
|
+
**Look for**
|
|
79
|
+
|
|
80
|
+
- Large behavior changes with no flag, setting or config to turn them off.
|
|
81
|
+
With a switch, mitigation is a config change; without one it is a rollback
|
|
82
|
+
or a hotfix.
|
|
83
|
+
- Changes whose rollback is unsafe: a migration with no down path that the
|
|
84
|
+
previous code cannot run against, a data format the old code cannot read.
|
|
85
|
+
- A flag that exists but defaults on for everyone at once.
|
|
86
|
+
|
|
87
|
+
**Cheapest evidence** — a static read for the gate; the project's flag or
|
|
88
|
+
settings catalog.
|
|
89
|
+
|
|
90
|
+
**A finding** — a rewritten notification fan-out ships ungated, and the
|
|
91
|
+
migration beside it drops the column the old implementation read. Proven:
|
|
92
|
+
reverting the code alone leaves it reading a column that no longer exists.
|
|
93
|
+
Rollback needs a restore. High.
|
|
94
|
+
|
|
95
|
+
**A strike** — "this refactor has no feature flag". The refactor preserves
|
|
96
|
+
behavior and is covered by the existing suite. Struck: over-engineering — a
|
|
97
|
+
flag for a change with no behavior difference is a second code path to
|
|
98
|
+
maintain for nothing.
|
|
99
|
+
|
|
100
|
+
## 4. Deploy and config
|
|
101
|
+
|
|
102
|
+
**Look for**
|
|
103
|
+
|
|
104
|
+
- Environment variables or secrets the code now requires, checked against
|
|
105
|
+
what the deploy configuration actually provides.
|
|
106
|
+
- CI, workflow and infra-manifest changes that only take effect at release.
|
|
107
|
+
- Dependency major bumps, runtime version requirements, new system packages.
|
|
108
|
+
- Steps a human has to take. These go in the report's **deploy checklist**
|
|
109
|
+
whether or not they are findings.
|
|
110
|
+
|
|
111
|
+
**Cheapest evidence** — a grep of the diff for new environment reads, against
|
|
112
|
+
the deploy manifests; startup validation output from a local boot.
|
|
113
|
+
|
|
114
|
+
**A finding** — the API validates a new required variable at startup, and the
|
|
115
|
+
production deploy manifest does not set it. Proven from the two files:
|
|
116
|
+
the service refuses to boot. High.
|
|
117
|
+
|
|
118
|
+
**A strike** — a new optional variable with a working default. Not a finding.
|
|
119
|
+
It belongs in the deploy checklist as a note, and nowhere else.
|
|
120
|
+
|
|
121
|
+
## 5. Contracts and compatibility
|
|
122
|
+
|
|
123
|
+
**Look for**
|
|
124
|
+
|
|
125
|
+
- API, socket, webhook or tool contract changes — renamed fields, newly
|
|
126
|
+
required fields, removed methods — against clients that deploy on a
|
|
127
|
+
different schedule: a web client cached in a browser, a mobile build, an
|
|
128
|
+
agent baked into an image, a published package pinned downstream.
|
|
129
|
+
- Strict schemas that now reject what yesterday's client still sends.
|
|
130
|
+
- Published-package changes that are breaking under a patch release.
|
|
131
|
+
|
|
132
|
+
**Cheapest evidence** — production logs for the old shape still arriving;
|
|
133
|
+
the consumer's pinned version.
|
|
134
|
+
|
|
135
|
+
**A finding** — a request schema gained `.strict()`. Observed: production
|
|
136
|
+
logs show a client still sending the field it now rejects, 2,100 requests in
|
|
137
|
+
the last day. Those requests fail after the release. High or Medium by what
|
|
138
|
+
the requests are.
|
|
139
|
+
|
|
140
|
+
**A strike** — an added optional response field. Additive, ignored by old
|
|
141
|
+
clients. Struck: no consequence.
|
|
142
|
+
|
|
143
|
+
## 6. Security and access
|
|
144
|
+
|
|
145
|
+
**Look for**
|
|
146
|
+
|
|
147
|
+
- New methods, endpoints or tools without an access check, or with a check
|
|
148
|
+
that resolves no entity and so passes for any signed-in user.
|
|
149
|
+
- Secrets, tokens or personal data written to logs.
|
|
150
|
+
- Unbounded queries and unpaginated lists on user-controlled input.
|
|
151
|
+
- Injection: string-built queries, shell commands, unescaped rendering.
|
|
152
|
+
|
|
153
|
+
**Cheapest evidence** — a local call to the new method as a user who should
|
|
154
|
+
be refused.
|
|
155
|
+
|
|
156
|
+
**A finding** — a new read method declares member-level access but resolves
|
|
157
|
+
no entry id. Reproduced locally: a user outside the project receives the
|
|
158
|
+
project's data. High.
|
|
159
|
+
|
|
160
|
+
**A strike** — "this admin-only method should also rate-limit". It sits
|
|
161
|
+
behind an admin check and an existing global limiter, cited. Struck:
|
|
162
|
+
redundant code path.
|
|
163
|
+
|
|
164
|
+
## 7. Stubs and cleanup
|
|
165
|
+
|
|
166
|
+
Search the **diff**, not the tree. Pre-existing debt is out of scope.
|
|
167
|
+
|
|
168
|
+
**Look for** — TODO / FIXME / HACK, debug logging, `.only` and skipped tests,
|
|
169
|
+
commented-out code, placeholder copy, temporary endpoints, flags and
|
|
170
|
+
translation keys the release added and never referenced, or whose last
|
|
171
|
+
reference it removed.
|
|
172
|
+
|
|
173
|
+
**Cheapest evidence** — `git diff origin/<default>..."$TIP"` piped through a
|
|
174
|
+
search for the markers, added lines only.
|
|
175
|
+
|
|
176
|
+
**A finding** — a test file gained `.only`. Proven: the rest of that suite
|
|
177
|
+
has not run in CI since the PR merged, and two of its cases fail when run.
|
|
178
|
+
Medium, or High if those cases guard a live flow.
|
|
179
|
+
|
|
180
|
+
**A strike** — a TODO noting a possible future optimization. Struck: no
|
|
181
|
+
consequence. Twelve of these in a report are how a real finding gets skimmed
|
|
182
|
+
past.
|
|
183
|
+
|
|
184
|
+
## 8. Documented-contract drift
|
|
185
|
+
|
|
186
|
+
**Look for**
|
|
187
|
+
|
|
188
|
+
- For each subsystem the release touched, `mcp__conveyor__get_tag` and read
|
|
189
|
+
the overview and its linked files. Three findings hide here: a change that
|
|
190
|
+
silently makes a documented claim false, a change that should have updated
|
|
191
|
+
the overview and did not, and a renamed or deleted file that a tag still
|
|
192
|
+
links.
|
|
193
|
+
- Rules and docs in the repo that agents load as context. A rule that now
|
|
194
|
+
describes removed behavior misleads every future session.
|
|
195
|
+
|
|
196
|
+
**Cheapest evidence** — a static comparison of the claim against the code;
|
|
197
|
+
the tag's link-verification status.
|
|
198
|
+
|
|
199
|
+
**A finding** — a tag overview states that one process owns a given write,
|
|
200
|
+
and the release added a second writer. Proven with both `file:line` cites.
|
|
201
|
+
Medium: nothing breaks today, and every future change will be planned on a
|
|
202
|
+
false premise.
|
|
203
|
+
|
|
204
|
+
**A strike** — an overview that is merely less detailed than the new code.
|
|
205
|
+
Struck: not caused by the release, and nothing in it is false.
|
|
206
|
+
|
|
207
|
+
## 9. Verification gaps
|
|
208
|
+
|
|
209
|
+
**Look for**
|
|
210
|
+
|
|
211
|
+
- Changed paths that are risky and have no test.
|
|
212
|
+
- Manual tests on member cards left unapproved or rejected:
|
|
213
|
+
`mcp__conveyor__list_manual_tests`.
|
|
214
|
+
- CI state on the release PR, and checks that were skipped or never ran.
|
|
215
|
+
- Suites weakened in the release: assertions removed, a filter added.
|
|
216
|
+
|
|
217
|
+
**Cheapest evidence** — run the scoped test that should exist; the manual
|
|
218
|
+
test list.
|
|
219
|
+
|
|
220
|
+
**A finding** — a payment path changed and its only manual test was
|
|
221
|
+
rejected with a note that was never addressed. Observed on the card. Medium,
|
|
222
|
+
or High if a reproduction confirms the rejection.
|
|
223
|
+
|
|
224
|
+
**A strike** — "this helper has no unit test". It is covered by the
|
|
225
|
+
integration suite, cited. Struck: redundant.
|
|
226
|
+
|
|
227
|
+
## 10. Open signals
|
|
228
|
+
|
|
229
|
+
**Look for**
|
|
230
|
+
|
|
231
|
+
- `mcp__conveyor__read_task_chat` on member cards: a follow-up promised
|
|
232
|
+
("I'll handle the empty state in a later card") with no card to show for
|
|
233
|
+
it, or a reviewer concern that was acknowledged and never resolved.
|
|
234
|
+
- Incidents filed since the changes merged to the dev branch that implicate
|
|
235
|
+
a release change: `search_tasks` with `typeFilters: ["incident"]`.
|
|
236
|
+
- Errors in the dev environment's logs that began at a merge in this release.
|
|
237
|
+
|
|
238
|
+
**Cheapest evidence** — dev logs with a start time at the merge; the
|
|
239
|
+
incident's own context.
|
|
240
|
+
|
|
241
|
+
**A finding** — an error appears in dev logs 410 times, first seen four
|
|
242
|
+
minutes after a release PR merged, with a stack in a file that PR changed.
|
|
243
|
+
Observed. Priority by what the failing call does.
|
|
244
|
+
|
|
245
|
+
**A strike** — an incident filed during the window whose stack is in code
|
|
246
|
+
the release never touched. Struck: not caused by the release. It already has
|
|
247
|
+
a card.
|
|
248
|
+
|
|
249
|
+
## Project dimensions
|
|
250
|
+
|
|
251
|
+
A wrapper skill adds dimensions in the same four-part shape, and each becomes
|
|
252
|
+
a ledger row. The shape to follow:
|
|
253
|
+
|
|
254
|
+
```markdown
|
|
255
|
+
## <Name>
|
|
256
|
+
|
|
257
|
+
**Look for** — <the candidates, and where the project's rules for them live>
|
|
258
|
+
**Cheapest evidence** — <the project tool that settles it fastest>
|
|
259
|
+
**A finding** — <one real example>
|
|
260
|
+
**A strike** — <one that should be killed, and which strike test kills it>
|
|
261
|
+
```
|
|
262
|
+
|
|
263
|
+
Dimensions that tend to be project-specific: multi-tenant assumptions (one
|
|
264
|
+
customer's usage treated as the contract), a project's own authorization
|
|
265
|
+
model, domain invariants, regulated data. They face the same evidence bar and
|
|
266
|
+
the same adversarial pass as everything above.
|
|
267
|
+
|
|
268
|
+
### The wrapper skill
|
|
269
|
+
|
|
270
|
+
The project's own skill stays thin. It invokes this one and adds what only
|
|
271
|
+
that project knows:
|
|
272
|
+
|
|
273
|
+
```markdown
|
|
274
|
+
---
|
|
275
|
+
name: release-review
|
|
276
|
+
description: Review a pending release of <project>.
|
|
277
|
+
---
|
|
278
|
+
Run `/conveyor-release-review` with these additions.
|
|
279
|
+
|
|
280
|
+
Project dimensions:
|
|
281
|
+
- <name>: <the question>, <where the rules live>, <cheapest evidence>
|
|
282
|
+
|
|
283
|
+
Evidence tools:
|
|
284
|
+
- Local reproduction: `<the repo's mock or dev stack command>`
|
|
285
|
+
- Read-only production data: `<the repo's sanctioned script>`
|
|
286
|
+
```
|
|
287
|
+
|
|
288
|
+
A wrapper may add dimensions and name tools. It may not lower the evidence
|
|
289
|
+
bar, skip the adversarial pass, or change what the priority levels mean.
|
|
@@ -0,0 +1,183 @@
|
|
|
1
|
+
# Report and card templates
|
|
2
|
+
|
|
3
|
+
The shapes behind steps 6 and 7 of
|
|
4
|
+
[conveyor-release-review](../SKILL.md). Fill them in; do not pad them. A
|
|
5
|
+
section with nothing to say is written as one line saying so, never dropped,
|
|
6
|
+
because a missing section reads as a skipped check.
|
|
7
|
+
|
|
8
|
+
`<label>` is `Release <version>` when a release is cut, and
|
|
9
|
+
`Pre-release <YYYY-MM-DD>` otherwise.
|
|
10
|
+
|
|
11
|
+
## The report
|
|
12
|
+
|
|
13
|
+
Posted to the conversation, and to the release card's chat when one exists.
|
|
14
|
+
|
|
15
|
+
```markdown
|
|
16
|
+
**HOLD — <label> @ <tip sha, 7 chars>.** 2 High findings. Blockers: <pack link>
|
|
17
|
+
|
|
18
|
+
Reviewed `origin/<default>...<tip>`: 38 PRs, 214 files. Tip recorded
|
|
19
|
+
<timestamp>; <"unchanged at report time" | "moved to <sha>, N commits not reviewed">.
|
|
20
|
+
|
|
21
|
+
### Findings
|
|
22
|
+
|
|
23
|
+
| Priority | Finding | Evidence | Introduced by | Card |
|
|
24
|
+
| --- | --- | --- | --- | --- |
|
|
25
|
+
| High | Unique index fails on 14 duplicate rows | Observed | #4712 | <link> |
|
|
26
|
+
| Medium | Board filter matches nothing after status rename | Proven | #4698 + #4705 | <link> |
|
|
27
|
+
|
|
28
|
+
Release blockers: <pack link> (2 cards, Open)
|
|
29
|
+
Release suggestions: <pack link> (3 cards, Planning)
|
|
30
|
+
|
|
31
|
+
### Deploy checklist
|
|
32
|
+
|
|
33
|
+
- [ ] Set `<VARIABLE>` in production before the API deploys (#4720)
|
|
34
|
+
- [ ] Nothing else requires a human step.
|
|
35
|
+
|
|
36
|
+
### Coverage
|
|
37
|
+
|
|
38
|
+
38 of 38 PRs evaluated. 2 without a card, reviewed on the diff alone: #4701, #4733.
|
|
39
|
+
|
|
40
|
+
| Dimension | Result |
|
|
41
|
+
| --- | --- |
|
|
42
|
+
| Cross-PR interaction | 1 finding. 6 overlapping surfaces read as combined diffs. |
|
|
43
|
+
| Data and migrations | 1 finding. 4 migrations, none lock-heavy at production size. |
|
|
44
|
+
| Rollback and off-switches | Clean. |
|
|
45
|
+
| ... | ... |
|
|
46
|
+
|
|
47
|
+
### Struck
|
|
48
|
+
|
|
49
|
+
| Candidate | Struck because |
|
|
50
|
+
| --- | --- |
|
|
51
|
+
| Add a retry to the export job | Redundant: the queue already retries 3 times (`jobs/queue.ts:88`) |
|
|
52
|
+
| Guard against a malformed legacy payload | Niche and graceful: 0 occurrences in 30 days, error already surfaces to the user |
|
|
53
|
+
|
|
54
|
+
### Unverified
|
|
55
|
+
|
|
56
|
+
| Candidate | Would be | What would settle it |
|
|
57
|
+
| --- | --- | --- |
|
|
58
|
+
| Webhook handler may double-fire on redelivery | High | A redelivery against the dev environment, or the provider's delivery log |
|
|
59
|
+
|
|
60
|
+
### Method
|
|
61
|
+
|
|
62
|
+
Evidence sources used: dev and prod logs, local stack, read-only production data.
|
|
63
|
+
Unavailable: <source>, so <which dimensions rest on static reading>.
|
|
64
|
+
Not reviewed: <anything skipped, and why>.
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
Rules for the verdict line:
|
|
68
|
+
|
|
69
|
+
- It is the **first line**, in both destinations. Someone who reads nothing
|
|
70
|
+
else must still get the verdict.
|
|
71
|
+
- It names the SHA. A verdict without one cannot be checked against what
|
|
72
|
+
actually shipped.
|
|
73
|
+
- `SHIP WITH OPEN QUESTIONS` names each open question on that same line or the
|
|
74
|
+
next one. It is not a softer `SHIP`.
|
|
75
|
+
|
|
76
|
+
## The hold notice
|
|
77
|
+
|
|
78
|
+
Posted to the release card when the verdict is `HOLD`. The first line carries
|
|
79
|
+
the whole message; the report follows beneath it.
|
|
80
|
+
|
|
81
|
+
```markdown
|
|
82
|
+
**HOLD — do not ship <label> until the blockers merge.** 2 High findings: <pack link>
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
## The pack parent
|
|
86
|
+
|
|
87
|
+
```markdown
|
|
88
|
+
## Objective
|
|
89
|
+
Clear the findings that block <label>, found by the release review of
|
|
90
|
+
`<tip sha>` on <date>.
|
|
91
|
+
|
|
92
|
+
## Approach
|
|
93
|
+
Each child is one finding with its own evidence and its own smallest fix.
|
|
94
|
+
They are independent unless a `dependsOn` edge says otherwise.
|
|
95
|
+
|
|
96
|
+
## Implementation Steps
|
|
97
|
+
1. <child title> — <one line>
|
|
98
|
+
2. <child title> — <one line>
|
|
99
|
+
|
|
100
|
+
## Testing
|
|
101
|
+
Each child carries its own gates and acceptance criteria.
|
|
102
|
+
|
|
103
|
+
## Notes
|
|
104
|
+
- Release card: <link>. Reviewed SHA: `<tip sha>`.
|
|
105
|
+
- <One of:> This is a full release, so fixes merged to `<dev>` ship with it.
|
|
106
|
+
<or> This is a cherry-pick release: once these merge, a human adds them with
|
|
107
|
+
Add to Release.
|
|
108
|
+
- Priority is the urgency of the fix. Risk is set by identification and review.
|
|
109
|
+
|
|
110
|
+
## Builder briefing
|
|
111
|
+
- **Start here:** the child with the earliest deploy-time consequence.
|
|
112
|
+
- **Already decided:** each fix is the smallest change that removes the
|
|
113
|
+
consequence. Do not widen it.
|
|
114
|
+
- **Traps:** <what looks right and is not>
|
|
115
|
+
- **Verify in this order:** re-run each child's evidence command first; a
|
|
116
|
+
finding that no longer reproduces is a finding someone already fixed.
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
The suggestions pack uses the same shape. Its Objective says these are worth
|
|
120
|
+
doing and do not hold the release, and its Notes say the pack sits in Planning
|
|
121
|
+
until a human promotes it.
|
|
122
|
+
|
|
123
|
+
## The finding card
|
|
124
|
+
|
|
125
|
+
One per finding, as a pack child or as a standalone card. It follows the
|
|
126
|
+
`/conveyor-plan` plan format
|
|
127
|
+
([plan-format.md](../../conveyor-plan/references/plan-format.md)) with two
|
|
128
|
+
sections added at the top, because a builder who cannot re-run the evidence
|
|
129
|
+
cannot tell a fixed finding from a wrong one.
|
|
130
|
+
|
|
131
|
+
```markdown
|
|
132
|
+
## Finding
|
|
133
|
+
<What breaks, for whom, and when. One short paragraph.>
|
|
134
|
+
|
|
135
|
+
Priority: <level>. Found in <label> @ `<tip sha>`. Release card: <link>.
|
|
136
|
+
Introduced by: #<pr> (<card slug>), #<pr> (<card slug>).
|
|
137
|
+
|
|
138
|
+
## Evidence
|
|
139
|
+
Grade: <Observed | Reproduced | Proven>
|
|
140
|
+
|
|
141
|
+
<the exact query or command>
|
|
142
|
+
|
|
143
|
+
Result: <what it returned, with the numbers>
|
|
144
|
+
|
|
145
|
+
## Objective
|
|
146
|
+
<One sentence: the consequence this card removes.>
|
|
147
|
+
|
|
148
|
+
## Approach
|
|
149
|
+
<The smallest change that removes it, and why nothing larger is needed.>
|
|
150
|
+
|
|
151
|
+
## Implementation Steps
|
|
152
|
+
1. `path/to/file.ts:120` — `symbolName`: <the change>
|
|
153
|
+
2. ...
|
|
154
|
+
|
|
155
|
+
## Testing
|
|
156
|
+
- Gates: <the repo's scoped commands for this diff>
|
|
157
|
+
- The evidence command above no longer shows the failure.
|
|
158
|
+
- Manual tests (0-3): <plain sentences, user path only>
|
|
159
|
+
|
|
160
|
+
## Notes
|
|
161
|
+
- What the adversarial pass considered and rejected: <larger fixes, and why>
|
|
162
|
+
- Blast radius: <callers and consumers of what this touches>
|
|
163
|
+
|
|
164
|
+
## Builder briefing
|
|
165
|
+
- **Start here:** <file>
|
|
166
|
+
- **Already decided:** <what not to reopen, and why>
|
|
167
|
+
- **Traps:** <what looks right and is not>
|
|
168
|
+
- **Verify in this order:** <cheapest disqualifying check first>
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
Rules for the card:
|
|
172
|
+
|
|
173
|
+
- **The description is for a non-engineer**: one or two plain sentences on
|
|
174
|
+
what goes wrong and for whom. Technical detail goes in the plan.
|
|
175
|
+
- **The title names the consequence, not the dimension.** "Deploy fails on
|
|
176
|
+
duplicate slugs", not "Migration issue".
|
|
177
|
+
- **Evidence is reproducible or it is not evidence.** An exact command, not
|
|
178
|
+
"checked the logs".
|
|
179
|
+
- **Never paste a secret, a token or personal data** into a card. Quote the
|
|
180
|
+
shape of the log line, not the user's email in it.
|
|
181
|
+
- **The fix is the smallest one.** Anything larger that the adversarial pass
|
|
182
|
+
rejected is recorded under Notes, so the builder does not rediscover it and
|
|
183
|
+
build it anyway.
|
|
@@ -132,6 +132,7 @@ labels. Each tag carries a `description` (≤255 — the summary), an `overview`
|
|
|
132
132
|
| Hand it to the cloud | `conveyor-start` → starts a pod and confirms it came up (local surface only — `start_task` does not exist in a pod) |
|
|
133
133
|
| Do the work here | `conveyor-build` → follows the plan to a PR |
|
|
134
134
|
| Judge the work | `conveyor-review` → one verdict, with risk |
|
|
135
|
+
| Audit a release before it ships | `conveyor-release-review` → evidence-backed findings, filed by priority as a blocker pack and a suggestions pack (local surface only — it needs `create_task`, `manage_priorities`, and `add_dependency` on another card) |
|
|
135
136
|
| Work a whole queue locally | `conveyor-local-loop` → selection and pacing over the above |
|
|
136
137
|
| Keep the board honest | `conveyor-prune` → one disposition per open card, applied only after confirmation (local surface only — it needs `set_task_parent`, `create_task`, and the on-hold flag) |
|
|
137
138
|
|
|
@@ -151,6 +152,17 @@ labels. Each tag carries a `description` (≤255 — the summary), an `overview`
|
|
|
151
152
|
work. The loop idles on `conveyor-wait`, a CLI in `@rallycry/conveyor-mcp`
|
|
152
153
|
that blocks until a card becomes claimable, so an idle loop wakes on a board
|
|
153
154
|
event instead of a timer.
|
|
155
|
+
- **From an app, with a project token (no MCP, no user)**: an external service
|
|
156
|
+
holding an incident-scoped project token files AND starts a card in one
|
|
157
|
+
call — `POST /api/incidents/report` with `type: "task"`, a `plan`, and
|
|
158
|
+
`start: true` (optionally `tags: [<glossary names>]`; an unknown name is a
|
|
159
|
+
400). The build starts through the same path as `start_task`, as the token's
|
|
160
|
+
owner, and the response carries `{ taskId, slug, url, status, buildStarted,
|
|
161
|
+
buildError? }`. A failed start still returns 201 with `buildStarted: false`
|
|
162
|
+
and leaves the card for a human — do not retry the report, that files a
|
|
163
|
+
second card. Poll `GET /api/incidents/tasks/<taskId>` with the same token for
|
|
164
|
+
`{ status, githubPRUrl, githubPRNumber, agentRunnerStatus }`. `start` is
|
|
165
|
+
rate-limited to 30 per token per 10 minutes.
|
|
154
166
|
- **Reserved-branch trap**: a card with an assigned agent may have a reserved
|
|
155
167
|
`githubBranch`. Check `get_task` before pushing: if set, push to THAT
|
|
156
168
|
branch; a PR from any other branch gets auto-closed and unlinked.
|