@hecer/yoke 1.11.0 → 1.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +13 -13
- package/.codex-plugin/plugin.json +7 -7
- package/CHANGELOG.md +416 -398
- package/README.md +931 -915
- package/TODOS.md +5 -5
- package/agents/docs.toml +6 -6
- package/agents/implementer.toml +6 -6
- package/agents/reviewer.toml +6 -6
- package/agents/security.toml +6 -6
- package/bench/README.md +86 -86
- package/bench/RESULTS.md +35 -35
- package/bench/output-compaction.mjs +65 -65
- package/bench/result-schema.mjs +12 -12
- package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
- package/bench/results/codex-unavailable-1785175418318.json +15 -15
- package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
- package/bench/run-matrix.mjs +26 -26
- package/bench/run.mjs +106 -106
- package/canon/AGENTS.md +30 -30
- package/canon/context/DECISIONS.md +4 -4
- package/canon/context/GLOSSARY.md +11 -11
- package/canon/context/KNOWLEDGE.md +4 -4
- package/canon/context/PROJECT.md +15 -15
- package/canon/loop/loop-spec.md +65 -65
- package/canon/loop/prd.schema.md +41 -41
- package/canon/manifest.yaml +59 -59
- package/canon/policy/gates.md +7 -7
- package/canon/policy/roles.md +9 -9
- package/canon/skills/ATTRIBUTION.md +99 -99
- package/canon/skills/authoring-prd/SKILL.md +56 -56
- package/canon/skills/brainstorming/SKILL.md +164 -164
- package/canon/skills/codebase-design/DEEPENING.md +15 -15
- package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
- package/canon/skills/codebase-design/SKILL.md +39 -39
- package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
- package/canon/skills/document-release/SKILL.md +302 -302
- package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
- package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
- package/canon/skills/domain-modeling/SKILL.md +35 -35
- package/canon/skills/executing-plans/SKILL.md +70 -70
- package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
- package/canon/skills/health/SKILL.md +177 -177
- package/canon/skills/maintaining-context/SKILL.md +34 -34
- package/canon/skills/minimal-code/SKILL.md +21 -21
- package/canon/skills/no-ai-slop/SKILL.md +103 -103
- package/canon/skills/no-ai-slop/eval.md +43 -43
- package/canon/skills/plan-ceo-review/SKILL.md +541 -541
- package/canon/skills/plan-eng-review/SKILL.md +362 -362
- package/canon/skills/receiving-code-review/SKILL.md +213 -213
- package/canon/skills/requesting-code-review/SKILL.md +105 -105
- package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
- package/canon/skills/retro/SKILL.md +397 -397
- package/canon/skills/review/SKILL.md +246 -246
- package/canon/skills/ship/SKILL.md +691 -691
- package/canon/skills/subagent-driven-development/SKILL.md +277 -277
- package/canon/skills/systematic-debugging/SKILL.md +296 -296
- package/canon/skills/tdd/SKILL.md +371 -371
- package/canon/skills/unslop-ui/SKILL.md +34 -34
- package/canon/skills/using-git-worktrees/SKILL.md +218 -218
- package/canon/skills/verification-before-completion/SKILL.md +139 -139
- package/canon/skills/visual-verification/SKILL.md +54 -54
- package/canon/skills/workflow/SKILL.md +22 -22
- package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
- package/canon/skills/writing-for-agents/SKILL.md +42 -42
- package/canon/skills/writing-plans/SKILL.md +152 -152
- package/canon/skills/writing-skills/SKILL.md +655 -655
- package/canon/skills/yoke-retrofit/SKILL.md +26 -26
- package/canon/skills/yoke-workflow/SKILL.md +20 -20
- package/canon/tools/codex-rtk-hook.mjs +35 -35
- package/canon/tools/gemini-rtk-hook.mjs +25 -25
- package/canon/tools/graphify.md +3 -3
- package/canon/tools/playwright-mcp.md +3 -3
- package/canon/tools/qwen-rtk-hook.mjs +25 -0
- package/canon/tools/rtk.md +7 -7
- package/canon/tools/serena.md +6 -6
- package/dist/agents/host.js +1 -1
- package/dist/agents/providers.js +18 -5
- package/dist/agents/telemetry.js +35 -36
- package/dist/cli.js +18 -10
- package/dist/dashboard/page.js +122 -122
- package/dist/dashboard/panels.js +91 -91
- package/dist/loop/run-command.js +3 -3
- package/dist/prd/command.js +17 -17
- package/dist/retrofit/apply.js +8 -1
- package/dist/retrofit/config.js +1 -1
- package/dist/retrofit/detect.js +2 -0
- package/dist/retrofit/planners/claude.js +14 -14
- package/dist/retrofit/planners/qwen.js +3 -3
- package/dist/retrofit/preserve.js +2 -2
- package/dist/retrofit/qwen-settings.js +17 -0
- package/dist/retrofit/skill-actions.js +1 -1
- package/dist/setup/command.js +22 -8
- package/dist/setup/model-presets.js +48 -0
- package/docs/CAPABILITY-ROUTING.md +51 -51
- package/docs/DASHBOARD-EVOLUTION.md +33 -33
- package/docs/MIGRATING-TO-1.0.md +33 -33
- package/docs/MIGRATING-TO-1.1.md +27 -27
- package/docs/MIGRATING-TO-1.4.md +70 -70
- package/docs/PRODUCT-DIRECTION-2026-09-05.md +210 -210
- package/docs/PUBLISHING.md +114 -114
- package/docs/QWEN-MODEL-SUPPORT.md +142 -0
- package/docs/VERIFIED-PROJECTS-VALIDATION.md +29 -29
- package/docs/VERIFIED-PROJECTS.md +167 -167
- package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
- package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
- package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
- package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
- package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
- package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
- package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
- package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
- package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
- package/docs/superpowers/plans/2026-09-05-verified-projects.md +83 -83
- package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
- package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
- package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
- package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
- package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
- package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
- package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
- package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
- package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
- package/gemini-extension.json +6 -6
- package/hooks/hooks.json +19 -19
- package/package.json +87 -87
- package/dist/dashboard/discovery.js +0 -73
- package/docs/community-outreach-2026-08-20.md +0 -85
- package/docs/launch-copy-2026-08-21.md +0 -193
|
@@ -1,246 +1,246 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: review
|
|
3
|
-
description: |
|
|
4
|
-
Pre-merge code review — the single canonical review of a change before it lands. Covers BOTH
|
|
5
|
-
diff safety/structure (SQL safety, LLM trust-boundary violations, conditional side effects)
|
|
6
|
-
AND engineering quality (architecture fit, edge cases, test coverage, performance). Use when
|
|
7
|
-
asked to "review this PR", "code review", "pre-landing review", "check my diff", or before
|
|
8
|
-
merging. (For plan-time review use plan-eng-review or plan-ceo-review instead.)
|
|
9
|
-
triggers:
|
|
10
|
-
- review this pr
|
|
11
|
-
- code review
|
|
12
|
-
- check my diff
|
|
13
|
-
- pre-landing review
|
|
14
|
-
---
|
|
15
|
-
|
|
16
|
-
# Pre-Landing PR Review
|
|
17
|
-
|
|
18
|
-
You are running the `review` workflow. Analyze the current branch's diff against the base branch for structural issues that tests don't catch.
|
|
19
|
-
|
|
20
|
-
---
|
|
21
|
-
|
|
22
|
-
## Step 0: Detect platform and base branch
|
|
23
|
-
|
|
24
|
-
Detect the git hosting platform from the remote URL:
|
|
25
|
-
|
|
26
|
-
```bash
|
|
27
|
-
git remote get-url origin 2>/dev/null
|
|
28
|
-
```
|
|
29
|
-
|
|
30
|
-
- URL contains "github.com" → platform is **GitHub**
|
|
31
|
-
- URL contains "gitlab" → platform is **GitLab**
|
|
32
|
-
- Otherwise check: `gh auth status 2>/dev/null` → GitHub; `glab auth status 2>/dev/null` → GitLab; neither → unknown
|
|
33
|
-
|
|
34
|
-
Determine the base branch (target of the PR, or the repo's default):
|
|
35
|
-
|
|
36
|
-
- GitHub: `gh pr view --json baseRefName -q .baseRefName` or `gh repo view --json defaultBranchRef -q .defaultBranchRef.name`
|
|
37
|
-
- GitLab: `glab mr view -F json 2>/dev/null` → extract `target_branch` or `default_branch`
|
|
38
|
-
- Fallback: `git symbolic-ref refs/remotes/origin/HEAD`, then `origin/main`, then `origin/master`, then `main`
|
|
39
|
-
|
|
40
|
-
Print the detected base branch. Use it as `<base>` in all subsequent commands.
|
|
41
|
-
|
|
42
|
-
---
|
|
43
|
-
|
|
44
|
-
## Step 1: Check branch
|
|
45
|
-
|
|
46
|
-
1. Run `git branch --show-current`.
|
|
47
|
-
2. If on the base branch: output **"Nothing to review — you're on the base branch or have no changes against it."** and stop.
|
|
48
|
-
3. Run `git fetch origin <base> --quiet && git diff origin/<base> --stat`. If no diff, output the same message and stop.
|
|
49
|
-
|
|
50
|
-
---
|
|
51
|
-
|
|
52
|
-
## Step 1.5: Scope Drift Detection
|
|
53
|
-
|
|
54
|
-
Check whether the diff matches what was requested.
|
|
55
|
-
|
|
56
|
-
1. Read `TODOS.md` (if it exists). Read PR description (`gh pr view --json body --jq .body 2>/dev/null || true`). Read commit messages (`git log origin/<base>..HEAD --oneline`).
|
|
57
|
-
2. Identify the **stated intent** — what was this branch supposed to accomplish?
|
|
58
|
-
3. Run `git diff origin/<base>...HEAD --stat` and compare against the stated intent.
|
|
59
|
-
|
|
60
|
-
Evaluate for:
|
|
61
|
-
- **SCOPE CREEP** — files changed that are unrelated to stated intent; "while I was in there" changes
|
|
62
|
-
- **MISSING REQUIREMENTS** — requirements from TODOS.md/PR description not in the diff; partial implementations
|
|
63
|
-
|
|
64
|
-
Output (before the main review begins):
|
|
65
|
-
```
|
|
66
|
-
Scope Check: [CLEAN / DRIFT DETECTED / REQUIREMENTS MISSING]
|
|
67
|
-
Intent: <1-line summary of what was requested>
|
|
68
|
-
Delivered: <1-line summary of what the diff actually does>
|
|
69
|
-
[If drift: list each out-of-scope change]
|
|
70
|
-
[If missing: list each unaddressed requirement]
|
|
71
|
-
```
|
|
72
|
-
|
|
73
|
-
This is **INFORMATIONAL** — does not block the review.
|
|
74
|
-
|
|
75
|
-
---
|
|
76
|
-
|
|
77
|
-
## Step 1.6: Plan Completion Audit (optional)
|
|
78
|
-
|
|
79
|
-
Check if there is a plan file referenced in the conversation context or a recent `.md` file in common plan locations (e.g., `~/.claude/plans/`, `.claude/plans/`). If found and relevant to the current branch:
|
|
80
|
-
|
|
81
|
-
Extract actionable items (checkboxes, numbered steps, imperative statements, file-level specs, test requirements). Cross-reference each item against the diff:
|
|
82
|
-
- **DONE** — clear evidence in diff
|
|
83
|
-
- **PARTIAL** — some work started but incomplete
|
|
84
|
-
- **NOT DONE** — no evidence in diff
|
|
85
|
-
- **CHANGED** — goal met by different means
|
|
86
|
-
|
|
87
|
-
For `PARTIAL` or `NOT DONE`, investigate why and rate impact (HIGH/MEDIUM/LOW). For HIGH-impact gaps, use AskUserQuestion:
|
|
88
|
-
- A) Stop and implement missing items
|
|
89
|
-
- B) Ship anyway + create P1 TODOs
|
|
90
|
-
- C) Intentionally dropped
|
|
91
|
-
|
|
92
|
-
Output format:
|
|
93
|
-
```
|
|
94
|
-
PLAN COMPLETION AUDIT
|
|
95
|
-
═══════════════════════
|
|
96
|
-
Plan: {path}
|
|
97
|
-
[DONE] Create UserService — src/services/user_service.rb
|
|
98
|
-
[NOT DONE] Add caching layer — no cache-related changes in diff
|
|
99
|
-
COMPLETION: N/M DONE
|
|
100
|
-
```
|
|
101
|
-
|
|
102
|
-
---
|
|
103
|
-
|
|
104
|
-
## Step 2: Read the checklist
|
|
105
|
-
|
|
106
|
-
Read `.claude/skills/review/checklist.md` (if it exists). If the file cannot be read, continue with the built-in checks below.
|
|
107
|
-
|
|
108
|
-
---
|
|
109
|
-
|
|
110
|
-
## Step 3: Get the diff
|
|
111
|
-
|
|
112
|
-
```bash
|
|
113
|
-
git fetch origin <base> --quiet
|
|
114
|
-
git diff origin/<base>
|
|
115
|
-
```
|
|
116
|
-
|
|
117
|
-
---
|
|
118
|
-
|
|
119
|
-
## Step 4: Critical review pass
|
|
120
|
-
|
|
121
|
-
Apply these categories against the diff:
|
|
122
|
-
|
|
123
|
-
**CRITICAL:**
|
|
124
|
-
- **SQL & Data Safety** — string interpolation in queries, missing parameterization, N+1 patterns
|
|
125
|
-
- **Race Conditions & Concurrency** — shared mutable state, missing locks, idempotency violations
|
|
126
|
-
- **LLM Output Trust Boundary** — LLM output used in SQL, shell commands, or DB writes without validation
|
|
127
|
-
- **Shell Injection** — user input passed to shell commands unsanitized
|
|
128
|
-
- **Enum & Value Completeness** — new enum values/types not handled in all switch/case branches
|
|
129
|
-
|
|
130
|
-
For Enum & Value Completeness: use Grep to find all files referencing sibling values, then Read those files. This requires looking outside the diff.
|
|
131
|
-
|
|
132
|
-
**INFORMATIONAL:**
|
|
133
|
-
- Async/sync mixing, column/field name safety, LLM prompt issues, type coercion, frontend/view issues, time window safety, completeness gaps, distribution/CI gaps
|
|
134
|
-
|
|
135
|
-
**Finding format:**
|
|
136
|
-
```
|
|
137
|
-
[SEVERITY] (confidence: N/10) file:line — description
|
|
138
|
-
```
|
|
139
|
-
|
|
140
|
-
Confidence scale:
|
|
141
|
-
- 9-10: Verified by reading specific code, concrete bug demonstrated
|
|
142
|
-
- 7-8: High confidence pattern match
|
|
143
|
-
- 5-6: Moderate — show with caveat "Medium confidence, verify this is actually an issue"
|
|
144
|
-
- 3-4: Low — include in appendix only
|
|
145
|
-
- 1-2: Speculation — only report if P0
|
|
146
|
-
|
|
147
|
-
---
|
|
148
|
-
|
|
149
|
-
## Step 4.5: Adversarial review (always-on)
|
|
150
|
-
|
|
151
|
-
Dispatch an independent subagent via the Agent tool to review the diff with fresh context. Subagent prompt:
|
|
152
|
-
|
|
153
|
-
> "Run `git diff origin/<base>` to get the diff. Think like an attacker and a chaos engineer. Find ways this code will fail in production: edge cases, race conditions, security holes, resource leaks, failure modes, silent data corruption, logic errors, error handling that swallows failures, trust boundary violations. For each finding, classify as FIXABLE (you know how to fix it) or INVESTIGATE (needs human judgment)."
|
|
154
|
-
|
|
155
|
-
Present findings under `ADVERSARIAL REVIEW (subagent):`. FIXABLE findings flow into the Fix-First pipeline. INVESTIGATE findings are informational.
|
|
156
|
-
|
|
157
|
-
---
|
|
158
|
-
|
|
159
|
-
## Step 5: Fix-First Review
|
|
160
|
-
|
|
161
|
-
### 5a: Classify each finding
|
|
162
|
-
|
|
163
|
-
For each finding from Steps 4 and 4.5:
|
|
164
|
-
- **AUTO-FIX** — mechanical, low-risk, single-file changes (dead code, stale comments, obvious formatting)
|
|
165
|
-
- **ASK** — architectural, security-sensitive, ambiguous scope, or user preference
|
|
166
|
-
|
|
167
|
-
### 5b: Apply AUTO-FIX items
|
|
168
|
-
|
|
169
|
-
Apply each fix directly. For each:
|
|
170
|
-
`[AUTO-FIXED] [file:line] Problem → what you did`
|
|
171
|
-
|
|
172
|
-
### 5c: Batch-ask about ASK items
|
|
173
|
-
|
|
174
|
-
Present all ASK items in one AskUserQuestion:
|
|
175
|
-
```
|
|
176
|
-
I auto-fixed N issues. M need your input:
|
|
177
|
-
|
|
178
|
-
1. [CRITICAL] file:line — description
|
|
179
|
-
Fix: recommended fix
|
|
180
|
-
→ A) Fix B) Skip
|
|
181
|
-
|
|
182
|
-
RECOMMENDATION: Fix all — [reason].
|
|
183
|
-
```
|
|
184
|
-
|
|
185
|
-
### 5d: Apply approved fixes
|
|
186
|
-
|
|
187
|
-
Apply fixes for items where the user chose "Fix."
|
|
188
|
-
|
|
189
|
-
---
|
|
190
|
-
|
|
191
|
-
## Step 5.5: TODOS cross-reference
|
|
192
|
-
|
|
193
|
-
Read `TODOS.md` (if it exists). Cross-reference the PR:
|
|
194
|
-
- Does this PR close any open TODOs? Note: "This PR addresses TODO: <title>"
|
|
195
|
-
- Does this PR create work that should become a TODO? Flag as informational.
|
|
196
|
-
|
|
197
|
-
---
|
|
198
|
-
|
|
199
|
-
## Step 5.6: Documentation staleness check
|
|
200
|
-
|
|
201
|
-
For each `.md` file in the repo root: if the code it describes was changed but the doc was NOT updated in this branch, flag as informational:
|
|
202
|
-
"Documentation may be stale: [file] describes [feature] but code changed. Consider running the `document-release` skill."
|
|
203
|
-
|
|
204
|
-
---
|
|
205
|
-
|
|
206
|
-
## Completion Status
|
|
207
|
-
|
|
208
|
-
Report one of:
|
|
209
|
-
- **DONE** — All steps completed, no blocking issues.
|
|
210
|
-
- **DONE_WITH_CONCERNS** — Completed with issues the user should know about.
|
|
211
|
-
- **BLOCKED** — Cannot proceed. State what is blocking.
|
|
212
|
-
- **NEEDS_CONTEXT** — Missing information required to continue.
|
|
213
|
-
|
|
214
|
-
## Important Rules
|
|
215
|
-
|
|
216
|
-
- **Read the FULL diff before commenting.** Do not flag issues already addressed in the diff.
|
|
217
|
-
- **Fix-first, not read-only.** AUTO-FIX items are applied directly; ASK items only after user approval.
|
|
218
|
-
- **Never commit, push, or create PRs** — that is the `ship` skill's job.
|
|
219
|
-
- **Be terse.** One line problem, one line fix. No preamble.
|
|
220
|
-
- **Only flag real problems.** Skip anything that is fine.
|
|
221
|
-
|
|
222
|
-
## Engineering-manager checklist
|
|
223
|
-
|
|
224
|
-
Beyond the structural/safety scan above, also review the change as an engineering manager would —
|
|
225
|
-
this is the angle the old `eng-review` skill covered, now folded in here so there is one
|
|
226
|
-
pre-merge review:
|
|
227
|
-
|
|
228
|
-
- **Architecture fit & data flow:** does the change follow the project's established patterns and
|
|
229
|
-
data flow, or does it drift? Flag architectural drift.
|
|
230
|
-
- **Edge cases & error paths:** are unhandled inputs, failure modes, and boundary conditions
|
|
231
|
-
covered?
|
|
232
|
-
- **Test coverage:** is the changed behavior covered by tests that verify behavior (not just
|
|
233
|
-
mocks)? Missing or weak tests for changed behavior is a blocking issue.
|
|
234
|
-
- **Performance:** any obvious regressions (N+1, unbounded growth, needless work in hot paths)?
|
|
235
|
-
|
|
236
|
-
Output a pass/block verdict with specific, actionable findings. A reviewer never reviews their
|
|
237
|
-
own implementation (see `policy/roles.md`).
|
|
238
|
-
|
|
239
|
-
## Interactive cross-model review
|
|
240
|
-
|
|
241
|
-
Outside the loop, run `yoke review` to have a *second* model review your current diff
|
|
242
|
-
before you commit or push. It resolves to the first available of codex → gemini → claude
|
|
243
|
-
(preferring a model other than the one you are driving), reviews the uncommitted working
|
|
244
|
-
tree by default (or `--base=<ref>` for a branch range), and exits non-zero if it finds a
|
|
245
|
-
blocking issue — so it chains as a gate (`... && yoke review`) or a pre-push hook.
|
|
246
|
-
This is the interactive counterpart to the loop's `--review`/`--reviewer`.
|
|
1
|
+
---
|
|
2
|
+
name: review
|
|
3
|
+
description: |
|
|
4
|
+
Pre-merge code review — the single canonical review of a change before it lands. Covers BOTH
|
|
5
|
+
diff safety/structure (SQL safety, LLM trust-boundary violations, conditional side effects)
|
|
6
|
+
AND engineering quality (architecture fit, edge cases, test coverage, performance). Use when
|
|
7
|
+
asked to "review this PR", "code review", "pre-landing review", "check my diff", or before
|
|
8
|
+
merging. (For plan-time review use plan-eng-review or plan-ceo-review instead.)
|
|
9
|
+
triggers:
|
|
10
|
+
- review this pr
|
|
11
|
+
- code review
|
|
12
|
+
- check my diff
|
|
13
|
+
- pre-landing review
|
|
14
|
+
---
|
|
15
|
+
|
|
16
|
+
# Pre-Landing PR Review
|
|
17
|
+
|
|
18
|
+
You are running the `review` workflow. Analyze the current branch's diff against the base branch for structural issues that tests don't catch.
|
|
19
|
+
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## Step 0: Detect platform and base branch
|
|
23
|
+
|
|
24
|
+
Detect the git hosting platform from the remote URL:
|
|
25
|
+
|
|
26
|
+
```bash
|
|
27
|
+
git remote get-url origin 2>/dev/null
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
- URL contains "github.com" → platform is **GitHub**
|
|
31
|
+
- URL contains "gitlab" → platform is **GitLab**
|
|
32
|
+
- Otherwise check: `gh auth status 2>/dev/null` → GitHub; `glab auth status 2>/dev/null` → GitLab; neither → unknown
|
|
33
|
+
|
|
34
|
+
Determine the base branch (target of the PR, or the repo's default):
|
|
35
|
+
|
|
36
|
+
- GitHub: `gh pr view --json baseRefName -q .baseRefName` or `gh repo view --json defaultBranchRef -q .defaultBranchRef.name`
|
|
37
|
+
- GitLab: `glab mr view -F json 2>/dev/null` → extract `target_branch` or `default_branch`
|
|
38
|
+
- Fallback: `git symbolic-ref refs/remotes/origin/HEAD`, then `origin/main`, then `origin/master`, then `main`
|
|
39
|
+
|
|
40
|
+
Print the detected base branch. Use it as `<base>` in all subsequent commands.
|
|
41
|
+
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
## Step 1: Check branch
|
|
45
|
+
|
|
46
|
+
1. Run `git branch --show-current`.
|
|
47
|
+
2. If on the base branch: output **"Nothing to review — you're on the base branch or have no changes against it."** and stop.
|
|
48
|
+
3. Run `git fetch origin <base> --quiet && git diff origin/<base> --stat`. If no diff, output the same message and stop.
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## Step 1.5: Scope Drift Detection
|
|
53
|
+
|
|
54
|
+
Check whether the diff matches what was requested.
|
|
55
|
+
|
|
56
|
+
1. Read `TODOS.md` (if it exists). Read PR description (`gh pr view --json body --jq .body 2>/dev/null || true`). Read commit messages (`git log origin/<base>..HEAD --oneline`).
|
|
57
|
+
2. Identify the **stated intent** — what was this branch supposed to accomplish?
|
|
58
|
+
3. Run `git diff origin/<base>...HEAD --stat` and compare against the stated intent.
|
|
59
|
+
|
|
60
|
+
Evaluate for:
|
|
61
|
+
- **SCOPE CREEP** — files changed that are unrelated to stated intent; "while I was in there" changes
|
|
62
|
+
- **MISSING REQUIREMENTS** — requirements from TODOS.md/PR description not in the diff; partial implementations
|
|
63
|
+
|
|
64
|
+
Output (before the main review begins):
|
|
65
|
+
```
|
|
66
|
+
Scope Check: [CLEAN / DRIFT DETECTED / REQUIREMENTS MISSING]
|
|
67
|
+
Intent: <1-line summary of what was requested>
|
|
68
|
+
Delivered: <1-line summary of what the diff actually does>
|
|
69
|
+
[If drift: list each out-of-scope change]
|
|
70
|
+
[If missing: list each unaddressed requirement]
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
This is **INFORMATIONAL** — does not block the review.
|
|
74
|
+
|
|
75
|
+
---
|
|
76
|
+
|
|
77
|
+
## Step 1.6: Plan Completion Audit (optional)
|
|
78
|
+
|
|
79
|
+
Check if there is a plan file referenced in the conversation context or a recent `.md` file in common plan locations (e.g., `~/.claude/plans/`, `.claude/plans/`). If found and relevant to the current branch:
|
|
80
|
+
|
|
81
|
+
Extract actionable items (checkboxes, numbered steps, imperative statements, file-level specs, test requirements). Cross-reference each item against the diff:
|
|
82
|
+
- **DONE** — clear evidence in diff
|
|
83
|
+
- **PARTIAL** — some work started but incomplete
|
|
84
|
+
- **NOT DONE** — no evidence in diff
|
|
85
|
+
- **CHANGED** — goal met by different means
|
|
86
|
+
|
|
87
|
+
For `PARTIAL` or `NOT DONE`, investigate why and rate impact (HIGH/MEDIUM/LOW). For HIGH-impact gaps, use AskUserQuestion:
|
|
88
|
+
- A) Stop and implement missing items
|
|
89
|
+
- B) Ship anyway + create P1 TODOs
|
|
90
|
+
- C) Intentionally dropped
|
|
91
|
+
|
|
92
|
+
Output format:
|
|
93
|
+
```
|
|
94
|
+
PLAN COMPLETION AUDIT
|
|
95
|
+
═══════════════════════
|
|
96
|
+
Plan: {path}
|
|
97
|
+
[DONE] Create UserService — src/services/user_service.rb
|
|
98
|
+
[NOT DONE] Add caching layer — no cache-related changes in diff
|
|
99
|
+
COMPLETION: N/M DONE
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
---
|
|
103
|
+
|
|
104
|
+
## Step 2: Read the checklist
|
|
105
|
+
|
|
106
|
+
Read `.claude/skills/review/checklist.md` (if it exists). If the file cannot be read, continue with the built-in checks below.
|
|
107
|
+
|
|
108
|
+
---
|
|
109
|
+
|
|
110
|
+
## Step 3: Get the diff
|
|
111
|
+
|
|
112
|
+
```bash
|
|
113
|
+
git fetch origin <base> --quiet
|
|
114
|
+
git diff origin/<base>
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
---
|
|
118
|
+
|
|
119
|
+
## Step 4: Critical review pass
|
|
120
|
+
|
|
121
|
+
Apply these categories against the diff:
|
|
122
|
+
|
|
123
|
+
**CRITICAL:**
|
|
124
|
+
- **SQL & Data Safety** — string interpolation in queries, missing parameterization, N+1 patterns
|
|
125
|
+
- **Race Conditions & Concurrency** — shared mutable state, missing locks, idempotency violations
|
|
126
|
+
- **LLM Output Trust Boundary** — LLM output used in SQL, shell commands, or DB writes without validation
|
|
127
|
+
- **Shell Injection** — user input passed to shell commands unsanitized
|
|
128
|
+
- **Enum & Value Completeness** — new enum values/types not handled in all switch/case branches
|
|
129
|
+
|
|
130
|
+
For Enum & Value Completeness: use Grep to find all files referencing sibling values, then Read those files. This requires looking outside the diff.
|
|
131
|
+
|
|
132
|
+
**INFORMATIONAL:**
|
|
133
|
+
- Async/sync mixing, column/field name safety, LLM prompt issues, type coercion, frontend/view issues, time window safety, completeness gaps, distribution/CI gaps
|
|
134
|
+
|
|
135
|
+
**Finding format:**
|
|
136
|
+
```
|
|
137
|
+
[SEVERITY] (confidence: N/10) file:line — description
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
Confidence scale:
|
|
141
|
+
- 9-10: Verified by reading specific code, concrete bug demonstrated
|
|
142
|
+
- 7-8: High confidence pattern match
|
|
143
|
+
- 5-6: Moderate — show with caveat "Medium confidence, verify this is actually an issue"
|
|
144
|
+
- 3-4: Low — include in appendix only
|
|
145
|
+
- 1-2: Speculation — only report if P0
|
|
146
|
+
|
|
147
|
+
---
|
|
148
|
+
|
|
149
|
+
## Step 4.5: Adversarial review (always-on)
|
|
150
|
+
|
|
151
|
+
Dispatch an independent subagent via the Agent tool to review the diff with fresh context. Subagent prompt:
|
|
152
|
+
|
|
153
|
+
> "Run `git diff origin/<base>` to get the diff. Think like an attacker and a chaos engineer. Find ways this code will fail in production: edge cases, race conditions, security holes, resource leaks, failure modes, silent data corruption, logic errors, error handling that swallows failures, trust boundary violations. For each finding, classify as FIXABLE (you know how to fix it) or INVESTIGATE (needs human judgment)."
|
|
154
|
+
|
|
155
|
+
Present findings under `ADVERSARIAL REVIEW (subagent):`. FIXABLE findings flow into the Fix-First pipeline. INVESTIGATE findings are informational.
|
|
156
|
+
|
|
157
|
+
---
|
|
158
|
+
|
|
159
|
+
## Step 5: Fix-First Review
|
|
160
|
+
|
|
161
|
+
### 5a: Classify each finding
|
|
162
|
+
|
|
163
|
+
For each finding from Steps 4 and 4.5:
|
|
164
|
+
- **AUTO-FIX** — mechanical, low-risk, single-file changes (dead code, stale comments, obvious formatting)
|
|
165
|
+
- **ASK** — architectural, security-sensitive, ambiguous scope, or user preference
|
|
166
|
+
|
|
167
|
+
### 5b: Apply AUTO-FIX items
|
|
168
|
+
|
|
169
|
+
Apply each fix directly. For each:
|
|
170
|
+
`[AUTO-FIXED] [file:line] Problem → what you did`
|
|
171
|
+
|
|
172
|
+
### 5c: Batch-ask about ASK items
|
|
173
|
+
|
|
174
|
+
Present all ASK items in one AskUserQuestion:
|
|
175
|
+
```
|
|
176
|
+
I auto-fixed N issues. M need your input:
|
|
177
|
+
|
|
178
|
+
1. [CRITICAL] file:line — description
|
|
179
|
+
Fix: recommended fix
|
|
180
|
+
→ A) Fix B) Skip
|
|
181
|
+
|
|
182
|
+
RECOMMENDATION: Fix all — [reason].
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
### 5d: Apply approved fixes
|
|
186
|
+
|
|
187
|
+
Apply fixes for items where the user chose "Fix."
|
|
188
|
+
|
|
189
|
+
---
|
|
190
|
+
|
|
191
|
+
## Step 5.5: TODOS cross-reference
|
|
192
|
+
|
|
193
|
+
Read `TODOS.md` (if it exists). Cross-reference the PR:
|
|
194
|
+
- Does this PR close any open TODOs? Note: "This PR addresses TODO: <title>"
|
|
195
|
+
- Does this PR create work that should become a TODO? Flag as informational.
|
|
196
|
+
|
|
197
|
+
---
|
|
198
|
+
|
|
199
|
+
## Step 5.6: Documentation staleness check
|
|
200
|
+
|
|
201
|
+
For each `.md` file in the repo root: if the code it describes was changed but the doc was NOT updated in this branch, flag as informational:
|
|
202
|
+
"Documentation may be stale: [file] describes [feature] but code changed. Consider running the `document-release` skill."
|
|
203
|
+
|
|
204
|
+
---
|
|
205
|
+
|
|
206
|
+
## Completion Status
|
|
207
|
+
|
|
208
|
+
Report one of:
|
|
209
|
+
- **DONE** — All steps completed, no blocking issues.
|
|
210
|
+
- **DONE_WITH_CONCERNS** — Completed with issues the user should know about.
|
|
211
|
+
- **BLOCKED** — Cannot proceed. State what is blocking.
|
|
212
|
+
- **NEEDS_CONTEXT** — Missing information required to continue.
|
|
213
|
+
|
|
214
|
+
## Important Rules
|
|
215
|
+
|
|
216
|
+
- **Read the FULL diff before commenting.** Do not flag issues already addressed in the diff.
|
|
217
|
+
- **Fix-first, not read-only.** AUTO-FIX items are applied directly; ASK items only after user approval.
|
|
218
|
+
- **Never commit, push, or create PRs** — that is the `ship` skill's job.
|
|
219
|
+
- **Be terse.** One line problem, one line fix. No preamble.
|
|
220
|
+
- **Only flag real problems.** Skip anything that is fine.
|
|
221
|
+
|
|
222
|
+
## Engineering-manager checklist
|
|
223
|
+
|
|
224
|
+
Beyond the structural/safety scan above, also review the change as an engineering manager would —
|
|
225
|
+
this is the angle the old `eng-review` skill covered, now folded in here so there is one
|
|
226
|
+
pre-merge review:
|
|
227
|
+
|
|
228
|
+
- **Architecture fit & data flow:** does the change follow the project's established patterns and
|
|
229
|
+
data flow, or does it drift? Flag architectural drift.
|
|
230
|
+
- **Edge cases & error paths:** are unhandled inputs, failure modes, and boundary conditions
|
|
231
|
+
covered?
|
|
232
|
+
- **Test coverage:** is the changed behavior covered by tests that verify behavior (not just
|
|
233
|
+
mocks)? Missing or weak tests for changed behavior is a blocking issue.
|
|
234
|
+
- **Performance:** any obvious regressions (N+1, unbounded growth, needless work in hot paths)?
|
|
235
|
+
|
|
236
|
+
Output a pass/block verdict with specific, actionable findings. A reviewer never reviews their
|
|
237
|
+
own implementation (see `policy/roles.md`).
|
|
238
|
+
|
|
239
|
+
## Interactive cross-model review
|
|
240
|
+
|
|
241
|
+
Outside the loop, run `yoke review` to have a *second* model review your current diff
|
|
242
|
+
before you commit or push. It resolves to the first available of codex → gemini → claude
|
|
243
|
+
(preferring a model other than the one you are driving), reviews the uncommitted working
|
|
244
|
+
tree by default (or `--base=<ref>` for a branch range), and exits non-zero if it finds a
|
|
245
|
+
blocking issue — so it chains as a gate (`... && yoke review`) or a pre-push hook.
|
|
246
|
+
This is the interactive counterpart to the loop's `--review`/`--reviewer`.
|