@rallycry/conveyor-skills 1.0.11 → 1.0.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,399 @@
1
+ ---
2
+ name: conveyor-release-review
3
+ description: Audit a pending release before it ships — evaluate every change in it, back each finding with evidence (production or dev logs, a local reproduction, or a deterministic proof), grade findings by the project's own Priority levels, strike what an adversarial pass finds over-built, and file the survivors as two packs, release blockers and release suggestions. Use when the user says "/conveyor-release-review [release]", "review this release", "is this release safe to ship", "audit what's going out", or asks for a pre-release check of the dev-to-main gap. Local/MCP surface only; inside a pod the skill says so instead of failing. Never fixes, never starts a build, never merges.
4
+ ---
5
+
6
+ # Conveyor Release Review
7
+
8
+ Every change in a release already passed its own review. This audit exists for
9
+ what that process cannot see: changes that are fine alone and break **in
10
+ combination**, risk that only exists at release scale, and what happens at
11
+ deploy time.
12
+
13
+ It has two ways to fail and treats them as equal. Missing a real break ships
14
+ it. Manufacturing work from speculation buries the real break under cards
15
+ nobody will build, and teaches the team to ignore the next review. So nothing
16
+ here is a finding until evidence says so, and nothing is filed until a
17
+ reviewer whose job is to strike it has tried.
18
+
19
+ ## Goal and finish line
20
+
21
+ **Goal:** every change in the release evaluated; every surviving finding
22
+ carrying its evidence and one of the project's Priority levels, filed in at
23
+ most two packs.
24
+
25
+ **Finish line:** the report — verdict on its first line — posted to the
26
+ conversation, and to the release card's chat when a release card exists.
27
+
28
+ **A turn may end only when:** (a) the finish line is true, (b) an
29
+ `AskUserQuestion` is pending because the release scope is genuinely ambiguous,
30
+ or (c) agents you launched are still running; their completion notices resume
31
+ the run. Anything else — a ledger with blank rows, findings listed but not
32
+ filed, packs created with no report — is a stalled run, not a paused one.
33
+
34
+ **Runtime:** *Codex CLI* — before step 0, create a thread goal with your goal
35
+ tool: "Post a release verdict for <release> with every change evaluated and
36
+ every surviving finding filed." Mark it complete only when the finish line is
37
+ true, never because the budget is low. *Claude Code* — there is no goal tool;
38
+ the finish line above is the loop.
39
+
40
+ > **Environment — this skill does not apply inside a pod.** Filing the packs
41
+ > needs `create_task`, `manage_priorities`, `add_dependency` against another
42
+ > card, and `update_task` with `priorityValue` — all of which exist only on
43
+ > the external `conveyor-mcp` surface. If
44
+ > `mcp__conveyor__get_connection_context` reports a task binding and no user
45
+ > account, you are in a pod: say so and stop. Do not hunt for an equivalent.
46
+
47
+ ## Ground rules
48
+
49
+ - **All Conveyor tools fully-qualified** — `mcp__conveyor__get_task`, not
50
+ `get_task`. Most are deferred; load every one you expect to need in a single
51
+ ToolSearch `select:` up front.
52
+ - **Pass `projectId` explicitly on every call.** The connection's default
53
+ project may not be the repo you are standing in.
54
+ - **Findings become cards; fixes get their own PRs.** Never fix anything inside
55
+ the review, however small. A fix made here ships unreviewed in the release
56
+ you were asked to judge.
57
+ - **This skill never moves a release.** No `start_task`, `create_release`,
58
+ `add_to_release`, `approve_and_merge_pr`, or verdict tool. Holding or
59
+ shipping is a human's call; your job is to make the right call obvious.
60
+ - **Production is read-only.** Logs, and SELECT-only data reads through the
61
+ project's own sanctioned path. Never a write, never a hand-exported
62
+ credential.
63
+ - **The project's rules outrank this skill.** Gate commands, log queries, data
64
+ access and any release checklist live in the host repo's CLAUDE.md. A wrapper
65
+ skill may add dimensions (see the end of this file); it cannot lower the
66
+ evidence bar.
67
+ - **The fan-out rules are not project policy.** They follow from how the
68
+ harness behaves and from account limits, so no project can relax them. A
69
+ wrapper may LOWER the concurrency cap; it may never raise it or lift a rule.
70
+
71
+ ## 0 — Discover the project
72
+
73
+ Which levels, branches and evidence sources exist is a per-project fact.
74
+ Guessing wastes the review either way.
75
+
76
+ 1. `mcp__conveyor__get_connection_context`, then match the project to the
77
+ repo's `git remote` (`mcp__conveyor__list_projects` when there is no
78
+ default). Copy the project id exactly.
79
+ 2. `mcp__conveyor__manage_priorities { action: "list" }` — the priority
80
+ ladder. The level with the **lowest `value`** is the **blocker level**;
81
+ every other level is a **suggestion level**. Use the project's own names in
82
+ every card and in the report. The seeded defaults are High, Medium and Low,
83
+ and the rest of this skill uses those names for readability. A project with
84
+ no priorities configured gets High / Medium / Low as text labels and no
85
+ priority writes.
86
+ 3. The dev branch and the default branch, from the repo and its CLAUDE.md.
87
+ 4. **Evidence sources — record what exists AND what does not:**
88
+ `mcp__conveyor__list_project_integrations`, whether
89
+ `mcp__conveyor__query_gcp_logs` / `mcp__conveyor__query_grafana_logs`
90
+ answer, the repo's local stack or mock scripts (`package.json` scripts such
91
+ as `dev:mock` or `dev`, or a repo skill that stands the app up), its scoped
92
+ test commands, and its sanctioned read-only data path.
93
+ 5. `mcp__conveyor__list_tags` once — the vocabulary, and whether a `release`
94
+ tag exists (step 6 uses it, and never creates it).
95
+
96
+ Announce one line naming the release, the priority ladder, and the evidence
97
+ sources you have and lack. Then start.
98
+
99
+ ## 1 — Scope the release
100
+
101
+ Anchor to a concrete diff before reading any code.
102
+
103
+ 1. **Find the release.** An explicit version, branch or range in the argument
104
+ wins. Otherwise `mcp__conveyor__search_tasks` for the release card
105
+ (`Release <version>`, or the `release` tag, in ReviewPR or InProgress);
106
+ `mcp__conveyor__get_task` on it gives the branch and PR number. No release
107
+ cut → review the dev-to-default gap and label the run **pre-release**.
108
+ 2. **Fetch with explicit refspecs.** Many checkouts only track the dev branch,
109
+ so an unfetched ref silently reviews the wrong range.
110
+
111
+ ```bash
112
+ git ls-remote --heads origin 'release/*'
113
+ git fetch origin <default>:refs/remotes/origin/<default> \
114
+ <dev>:refs/remotes/origin/<dev> \
115
+ 'refs/heads/release/*:refs/remotes/origin/release/*'
116
+ TIP=origin/release/<version> # or origin/<dev> when pre-release
117
+ git rev-parse "$TIP" # the SHA this verdict is for
118
+ git log --merges --oneline origin/<default>.."$TIP"
119
+ git diff origin/<default>..."$TIP" --name-only
120
+ gh pr list --base <dev> --state merged \
121
+ --search "merged:>$(git log -1 --format=%cs origin/<default>)"
122
+ ```
123
+
124
+ The `gh pr list` cross-check catches squash and rebase merges the merge log
125
+ misses. Use `--name-only`, never `--stat`: truncated paths hide files.
126
+ 3. **Record the tip SHA.** A full release keeps absorbing the dev branch until
127
+ it merges, so the tip moves under you. The verdict is for the SHA you
128
+ recorded. If it moved by report time, review the delta or list the commits
129
+ you did not review.
130
+ 4. **Map every PR to its card.** `conveyor/<slug>-<id>` branch names embed the
131
+ card slug; otherwise `search_tasks` by PR title. `get_task` gives the plan,
132
+ which is the intent. Reviewing a diff without its intent is how a
133
+ single-interpretation solution gets waved through twice. A PR with no card
134
+ is reviewed on its diff alone and noted as such.
135
+ 5. **Build the coverage ledger** — one row per PR or card, one row per
136
+ dimension in step 2. Every row ends `clean`, `finding`, or
137
+ `not reviewable (why)`. "Evaluate everything" means the ledger has no blank
138
+ rows when you finish.
139
+
140
+ **Fanning out.** On a large release, split the ledger across subagents per
141
+ area. Give each the ground rules and the evidence bar verbatim. Three rules
142
+ hold, and the mechanics are in [references/fan-out.md](references/fan-out.md):
143
+
144
+ 1. **The hand-back is the only report channel.** The harness refuses report
145
+ files written by subagents. Each agent's final message is its report, and
146
+ you save each one to the run directory the moment it arrives.
147
+ 2. **Only you spawn agents.** Every agent prompt forbids the Agent, Task and
148
+ Workflow tools. Nested agents share your session limit.
149
+ 3. **At most four agents in flight**, the adversarial agent included. Launch
150
+ the next slice when one hands back, never in fixed waves.
151
+
152
+ The ledger is still the completeness check, and it stays with you.
153
+
154
+ ## 2 — Review every dimension
155
+
156
+ Work all of them. Each gets a coverage note in the report even when the note
157
+ is "nothing found", so coverage is auditable. What to look for, the cheapest
158
+ evidence, and a worked example of a real finding next to one that should be
159
+ struck are in [references/dimensions.md](references/dimensions.md).
160
+
161
+ | # | Dimension | The question |
162
+ | --- | --- | --- |
163
+ | 1 | **Cross-PR interaction** | Which files, symbols, tables and shared packages did two or more PRs touch, and what does their combined diff do? This is the headline check. |
164
+ | 2 | **Data and migrations** | Destructive steps, locks on large tables, ordering between migrations, and the window where old code runs against the new schema. |
165
+ | 3 | **Rollback and off-switches** | Which behavior changes have no flag, and what is the rollback for each risky change? |
166
+ | 4 | **Deploy and config** | New env vars or secrets production lacks, CI or workflow changes, infra manifests, dependency major bumps. |
167
+ | 5 | **Contracts and compatibility** | API, socket, MCP or published-package contract changes, against clients that are already deployed. |
168
+ | 6 | **Security and access** | New methods or endpoints without an access check, secrets or personal data in logs, unbounded queries. |
169
+ | 7 | **Stubs and cleanup** | In the diff only: TODO / FIXME / HACK, debug logging, `.only` or skipped tests, commented-out code, placeholder copy, orphaned flags. |
170
+ | 8 | **Documented-contract drift** | `mcp__conveyor__get_tag` on each touched subsystem: did the release make an overview or a pinned invariant false? |
171
+ | 9 | **Verification gaps** | Risky changed paths with no test, manual tests on member cards left unapproved or rejected (`mcp__conveyor__list_manual_tests`), CI state on the release PR. |
172
+ | 10 | **Open signals** | Follow-ups promised in member-card chat that never became a card; incidents filed since the merge that implicate a release change. |
173
+ | + | **Project dimensions** | Whatever a wrapper skill or the invoking prompt added. Same bar. |
174
+
175
+ What comes out of this step is a list of **candidates**. None of them is a
176
+ finding yet.
177
+
178
+ ## 3 — Evidence, or it is not a finding
179
+
180
+ A candidate becomes a finding only with at least one of these. Record the
181
+ exact query or command beside the result, so anyone can re-run it.
182
+
183
+ | Grade | What it is |
184
+ | --- | --- |
185
+ | **Observed** | Logs or data. `env: "dev"` logs show the release's own code running, because its changes already merged to the dev branch. `env: "prod"` logs and read-only data show current traffic, data shape and row counts — what the release will meet. |
186
+ | **Reproduced** | The project's local stack or mocks booted and the failure shown, or a scoped test or script that demonstrates the break. The steps go on the card. Nothing is committed. |
187
+ | **Proven** | A deterministic static proof with both ends cited as `file:line` — a migration drops a column that code still reads. Allowed only when the failure needs no runtime assumption. "If a user happens to…" is not a proof. |
188
+
189
+ - **Cheapest first:** logs, then a static proof, then a scoped test, then the
190
+ full local stack.
191
+ - **Evidence downgrades too.** A path with no traffic in thirty days is a
192
+ reason to lower a finding or strike it.
193
+ - **Time-box it.** A candidate that has cost fifteen minutes with no evidence
194
+ either way is Unverified. Move on.
195
+ - **Unverified candidates are never filed.** They go in the report's
196
+ Unverified list, each with the evidence that would settle it. One that would
197
+ be blocker-level if true is named in the verdict line as an open question,
198
+ so a human decides with their eyes open.
199
+ - **A missing source is reported, not substituted.** "No log access on this
200
+ project, so dimensions 4 and 10 rest on static reading" is a caveat the
201
+ reader needs.
202
+
203
+ ## 4 — Priority
204
+
205
+ Grade by **consequence and evidence**, never by which dimension surfaced it.
206
+
207
+ | Level | Test | Filed in |
208
+ | --- | --- | --- |
209
+ | **High** (the blocker level) | Shipping as-is causes data loss or corruption, a security or access hole, a broken live flow with no workaround, or a deploy that fails or cannot be rolled back. Observed, Reproduced or Proven only. | Release blockers |
210
+ | **Medium** | A real user-facing defect or regression with a workaround or limited reach, or a gap that will bite within a release or two. | Release suggestions |
211
+ | **Low** | Real but minor: cleanup, polish, small debt this release introduced. | Release suggestions |
212
+
213
+ **Priority is not risk.** Priority is how urgently the fix is needed. Risk is
214
+ how much important surface a change touches; identification and review set
215
+ it, and this skill never does.
216
+
217
+ The test for High is one question: *would I actually hold the release for
218
+ this?* If the honest answer is "no, but it should be fixed soon", it is
219
+ Medium. A blocker pack of twelve gets ignored, and then so does the one card
220
+ in it that mattered.
221
+
222
+ ## 5 — The adversarial pass
223
+
224
+ Every candidate that reached step 4, **together with the fix you would
225
+ propose**, now faces a reviewer whose only job is to strike it. Use one fresh
226
+ subagent for all candidates, under the fan-out rules. Write the candidates,
227
+ with their evidence, priority and proposed fix, to `candidates.md` in the run
228
+ directory; with no fan-out, put them in the prompt instead. The agent argues
229
+ for striking each against the tests below and hands back one verdict per
230
+ candidate. Where subagents are unavailable, switch stance explicitly and
231
+ write the case against each finding before deciding.
232
+
233
+ Strike when:
234
+
235
+ 1. **Over-engineering.** The fix is bigger than the problem: a new
236
+ abstraction, option or framework for something a few lines solve, or a
237
+ defense against a future nobody has asked for.
238
+ 2. **A niche path that already fails gracefully.** The path is rare — the
239
+ evidence shows little or no traffic — and its failure is already handled:
240
+ the error surfaces, no data is damaged, a retry works.
241
+ 3. **A redundant code path.** The fix adds a second guard, fallback, retry or
242
+ validation for something already guaranteed upstream. Cite the guard that
243
+ already exists.
244
+ 4. **The evidence does not support the claim.** It shows something nearby, or
245
+ something smaller.
246
+ 5. **The release neither caused nor worsened it.** Pre-existing problems are
247
+ out of scope. Worth recording → `mcp__conveyor__create_suggestion`, never a
248
+ release-pack card.
249
+ 6. **An open card already tracks it.** `mcp__conveyor__search_tasks`, two or
250
+ three keyword variants, all statuses. Name that card in the report instead
251
+ of filing a second one.
252
+
253
+ Four outcomes: **Keep**, **Downgrade**, **Shrink** — keep the finding, replace
254
+ the fix with the smallest change that removes the consequence — or **Strike**.
255
+ Every struck finding is listed in the report with its reason. Nothing
256
+ disappears silently, and a reader who disagrees with a strike can say so.
257
+
258
+ ## 6 — File the packs
259
+
260
+ Survivors only. Created without asking, never started. Templates for the pack
261
+ parent and the finding card are in
262
+ [references/report-and-cards.md](references/report-and-cards.md).
263
+
264
+ **The bar for every card is the `/conveyor-plan` plan format**
265
+ ([../conveyor-plan/references/plan-format.md](../conveyor-plan/references/plan-format.md)),
266
+ Builder briefing included, because `/conveyor-build` is what runs these cards.
267
+ Cite that format; do **not** invoke `/conveyor-plan`. You have already done
268
+ the research, and its create-and-promote steps would file each card twice.
269
+
270
+ `<label>` is `Release <version>` when a release is cut, and
271
+ `Pre-release <YYYY-MM-DD>` otherwise.
272
+
273
+ | | Release blockers | Release suggestions |
274
+ | --- | --- | --- |
275
+ | Holds | High findings | Medium and Low findings |
276
+ | Parent title | `<label> blockers` | `<label> suggestions` |
277
+ | Parent priority | the blocker level | the highest level among its children |
278
+ | Lands in | **Open** — the release is waiting on it | **Planning** — a human promotes what they want built |
279
+
280
+ In this order:
281
+
282
+ 1. `mcp__conveyor__create_task` the parent in Planning, with `priorityValue`,
283
+ and `tags: ["release"]` when step 0 found that tag.
284
+ 2. `mcp__conveyor__create_subtask` once per finding, full plan included, with
285
+ `dependsOn` only where one fix genuinely needs another to land first.
286
+ 3. `mcp__conveyor__update_task { taskId: <child>, priorityValue }` per child.
287
+ `create_subtask` carries no priority field; this is the sanctioned route.
288
+ 4. `mcp__conveyor__post_to_chat` on the parent — the sizing recommendation
289
+ and what the pack is for. Chat goes **before** any status move, because
290
+ identification reads it when the card leaves Planning.
291
+ 5. Blockers only: `mcp__conveyor__update_subtask { status: "Open" }` per
292
+ child, then `mcp__conveyor__update_task { status: "Open" }` on the parent.
293
+
294
+ Then:
295
+
296
+ - **A pack of one is a card.** One survivor at a level → a single card titled
297
+ `<label> blocker: <finding>` or `<label> suggestion: <finding>`.
298
+ - **No survivors at a level → no pack at that level.** A clean release creates
299
+ nothing, and saying so is a complete result.
300
+ - **Blocking is advisory.** Nothing in Conveyor stops a release PR from
301
+ merging. With a release card present, make the hold impossible to miss:
302
+ `mcp__conveyor__add_dependency { taskId: <release card>, dependsOnSlugOrId: <blocker pack or card> }`,
303
+ which shows the release as blocked until the fixes merge, and a
304
+ `post_to_chat` on the release card whose **first line** is the hold notice.
305
+ If `add_dependency` is refused, report the error once and carry on.
306
+ - **Say how the fixes reach the release.** A full release absorbs whatever
307
+ merges to the dev branch, so merged fixes ship with it. A cherry-pick
308
+ release does not: a human adds them with Add to Release. Put the right
309
+ sentence in the blocker pack's plan.
310
+ - **No `priorityValue` on `update_task`?** The connected MCP server predates
311
+ the field. Prefix each title with its level — `[High] …` — and say in the
312
+ report that priorities need setting by hand.
313
+ - **On a tool error, stop and report the card rather than retrying it.** A
314
+ create that "failed" with a timeout may have landed; check with
315
+ `search_tasks` before creating it again.
316
+
317
+ ## 7 — Verdict and report
318
+
319
+ The first line is the verdict, and it names the SHA:
320
+
321
+ - **`HOLD`** — one or more High findings.
322
+ - **`SHIP`** — none.
323
+ - **`SHIP WITH OPEN QUESTIONS`** — no confirmed High, but at least one
324
+ Unverified candidate that would be High if true.
325
+
326
+ Below it, in this order: the reviewed range and tip SHA; the packs, linked,
327
+ with a findings table; the **deploy checklist**; the ledger summary with a
328
+ note per dimension; struck findings with reasons; the Unverified list;
329
+ evidence sources used and unavailable; anything not reviewed.
330
+
331
+ The **deploy checklist** is every step a human must take at deploy — set a
332
+ secret, run a backfill, flip a flag, apply a manifest — whether or not it is a
333
+ finding. A release that is fine in code and fails because nobody knew to set a
334
+ variable is the most avoidable outage there is.
335
+
336
+ Post the report to the conversation, and with
337
+ `mcp__conveyor__post_to_chat` to the release card when one exists. Never
338
+ create a release card to hold it.
339
+
340
+ ## Running it again
341
+
342
+ A release gets reviewed more than once: after the fixes land, or after the tip
343
+ moves.
344
+
345
+ - Find the existing packs by title (`search_tasks "<label> blockers"`) and add
346
+ new findings to them. Never create a second pack for the same release.
347
+ - Read the last reviewed SHA from the release card's chat and review the
348
+ delta, not the whole release again.
349
+ - Post the new verdict with the new SHA. The dependency on the release card
350
+ clears itself once the blocker pack merges to the dev branch.
351
+
352
+ ## Common mistakes
353
+
354
+ - Reviewing the hunks. Combination bugs live in the overlap map, not in any
355
+ one diff.
356
+ - Skipping migrations because CI passed. CI ran them against a schema no
357
+ production data ever stressed.
358
+ - Grading by dimension. A cleanup item is not High because it touched a
359
+ migration, and a broken flow is not Low because it surfaced under "stubs".
360
+ - One mega-card of findings. One card per finding, or nothing is
361
+ independently schedulable.
362
+ - Re-litigating a merged PR. Per-PR review is `/conveyor-review`, and it
363
+ already happened.
364
+ - Diagnosing an incident here. An incident the release implicates is
365
+ evidence; its root cause belongs to `/conveyor-triage`.
366
+ - Telling subagents to write findings files. The harness refuses them. The
367
+ hand-back is the report, and saving it is your job.
368
+ - Letting a reviewer fan out. Its agents share your session limit, and a limit
369
+ hit kills every agent in flight with its unsaved work.
370
+
371
+ ## Wrapping this skill for one project
372
+
373
+ A repo with its own release concerns keeps them in a thin project skill that
374
+ invokes this one and names its extra dimensions and evidence tools, rather
375
+ than in a fork. The wrapper's shape is at the end of
376
+ [references/dimensions.md](references/dimensions.md). Project dimensions
377
+ become extra ledger rows, graded by the same rubric and held to the same
378
+ evidence bar and adversarial pass.
379
+
380
+ ## Environment notes
381
+
382
+ - All Conveyor tools are **fully qualified** and mostly **deferred** — one
383
+ ToolSearch `select:` with every name you expect to need, up front.
384
+ - `search_tasks` returns 20 rows by default — pass `limit`.
385
+ - `gh` reads its own auth. An empty result with exit 0 from `gh pr list` or
386
+ `gh pr checks` can mean an expired token rather than "nothing there";
387
+ re-check before concluding.
388
+ - The log tools appear only for a project with the matching integration. If
389
+ they are absent, the project has none; say so instead of hunting.
390
+ - `You've hit your session limit · resets <time>` (HTTP 429) is the account's
391
+ five-hour usage window, shared by every agent and by you. Nothing runs until
392
+ the reset, so do not retry. Resume per
393
+ [references/fan-out.md](references/fan-out.md).
394
+
395
+ ## Improve This Skill
396
+
397
+ If this skill was insufficient or slowed the work down, file it with
398
+ `mcp__conveyor__create_suggestion` on the Conveyor project: the issue,
399
+ evidence, and proposed fix.