@hecer/yoke 1.6.0 → 1.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (103) hide show
  1. package/.claude-plugin/plugin.json +13 -13
  2. package/.codex-plugin/plugin.json +7 -7
  3. package/CHANGELOG.md +294 -288
  4. package/README.md +874 -874
  5. package/TODOS.md +5 -5
  6. package/agents/docs.toml +6 -6
  7. package/agents/implementer.toml +6 -6
  8. package/agents/reviewer.toml +6 -6
  9. package/agents/security.toml +6 -6
  10. package/bench/README.md +86 -86
  11. package/bench/RESULTS.md +35 -35
  12. package/bench/output-compaction.mjs +65 -65
  13. package/bench/result-schema.mjs +12 -12
  14. package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
  15. package/bench/results/codex-unavailable-1785175418318.json +15 -15
  16. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
  17. package/bench/run-matrix.mjs +26 -26
  18. package/bench/run.mjs +106 -106
  19. package/canon/AGENTS.md +30 -30
  20. package/canon/context/DECISIONS.md +4 -4
  21. package/canon/context/GLOSSARY.md +11 -11
  22. package/canon/context/KNOWLEDGE.md +4 -4
  23. package/canon/context/PROJECT.md +15 -15
  24. package/canon/loop/loop-spec.md +65 -65
  25. package/canon/loop/prd.schema.md +43 -43
  26. package/canon/manifest.yaml +59 -59
  27. package/canon/policy/gates.md +7 -7
  28. package/canon/policy/roles.md +9 -9
  29. package/canon/skills/ATTRIBUTION.md +99 -99
  30. package/canon/skills/authoring-prd/SKILL.md +58 -58
  31. package/canon/skills/brainstorming/SKILL.md +164 -164
  32. package/canon/skills/codebase-design/DEEPENING.md +15 -15
  33. package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
  34. package/canon/skills/codebase-design/SKILL.md +39 -39
  35. package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
  36. package/canon/skills/document-release/SKILL.md +302 -302
  37. package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
  38. package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
  39. package/canon/skills/domain-modeling/SKILL.md +35 -35
  40. package/canon/skills/executing-plans/SKILL.md +70 -70
  41. package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
  42. package/canon/skills/health/SKILL.md +177 -177
  43. package/canon/skills/maintaining-context/SKILL.md +34 -34
  44. package/canon/skills/minimal-code/SKILL.md +21 -21
  45. package/canon/skills/no-ai-slop/SKILL.md +103 -103
  46. package/canon/skills/no-ai-slop/eval.md +43 -43
  47. package/canon/skills/plan-ceo-review/SKILL.md +541 -541
  48. package/canon/skills/plan-eng-review/SKILL.md +362 -362
  49. package/canon/skills/receiving-code-review/SKILL.md +213 -213
  50. package/canon/skills/requesting-code-review/SKILL.md +105 -105
  51. package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
  52. package/canon/skills/retro/SKILL.md +397 -397
  53. package/canon/skills/review/SKILL.md +246 -246
  54. package/canon/skills/ship/SKILL.md +691 -691
  55. package/canon/skills/subagent-driven-development/SKILL.md +277 -277
  56. package/canon/skills/systematic-debugging/SKILL.md +296 -296
  57. package/canon/skills/tdd/SKILL.md +371 -371
  58. package/canon/skills/unslop-ui/SKILL.md +34 -34
  59. package/canon/skills/using-git-worktrees/SKILL.md +218 -218
  60. package/canon/skills/verification-before-completion/SKILL.md +139 -139
  61. package/canon/skills/visual-verification/SKILL.md +54 -54
  62. package/canon/skills/workflow/SKILL.md +22 -22
  63. package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
  64. package/canon/skills/writing-for-agents/SKILL.md +42 -42
  65. package/canon/skills/writing-plans/SKILL.md +152 -152
  66. package/canon/skills/writing-skills/SKILL.md +655 -655
  67. package/canon/skills/yoke-retrofit/SKILL.md +26 -26
  68. package/canon/skills/yoke-workflow/SKILL.md +20 -20
  69. package/canon/tools/codex-rtk-hook.mjs +35 -35
  70. package/canon/tools/graphify.md +3 -3
  71. package/canon/tools/playwright-mcp.md +3 -3
  72. package/canon/tools/rtk.md +7 -7
  73. package/canon/tools/serena.md +6 -6
  74. package/dist/agents/process.js +3 -0
  75. package/dist/loop/watchdog.js +1 -1
  76. package/dist/prd/command.js +17 -17
  77. package/dist/retrofit/planners/claude.js +14 -14
  78. package/dist/retrofit/preserve.js +2 -2
  79. package/docs/MIGRATING-TO-1.0.md +33 -33
  80. package/docs/MIGRATING-TO-1.1.md +27 -27
  81. package/docs/MIGRATING-TO-1.4.md +70 -70
  82. package/docs/PUBLISHING.md +91 -91
  83. package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
  84. package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
  85. package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
  86. package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
  87. package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
  88. package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
  89. package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
  90. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
  91. package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
  92. package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
  93. package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
  94. package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
  95. package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
  96. package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
  97. package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
  98. package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
  99. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
  100. package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
  101. package/gemini-extension.json +6 -6
  102. package/hooks/hooks.json +19 -19
  103. package/package.json +87 -87
@@ -1,246 +1,246 @@
1
- ---
2
- name: review
3
- description: |
4
- Pre-merge code review — the single canonical review of a change before it lands. Covers BOTH
5
- diff safety/structure (SQL safety, LLM trust-boundary violations, conditional side effects)
6
- AND engineering quality (architecture fit, edge cases, test coverage, performance). Use when
7
- asked to "review this PR", "code review", "pre-landing review", "check my diff", or before
8
- merging. (For plan-time review use plan-eng-review or plan-ceo-review instead.)
9
- triggers:
10
- - review this pr
11
- - code review
12
- - check my diff
13
- - pre-landing review
14
- ---
15
-
16
- # Pre-Landing PR Review
17
-
18
- You are running the `review` workflow. Analyze the current branch's diff against the base branch for structural issues that tests don't catch.
19
-
20
- ---
21
-
22
- ## Step 0: Detect platform and base branch
23
-
24
- Detect the git hosting platform from the remote URL:
25
-
26
- ```bash
27
- git remote get-url origin 2>/dev/null
28
- ```
29
-
30
- - URL contains "github.com" → platform is **GitHub**
31
- - URL contains "gitlab" → platform is **GitLab**
32
- - Otherwise check: `gh auth status 2>/dev/null` → GitHub; `glab auth status 2>/dev/null` → GitLab; neither → unknown
33
-
34
- Determine the base branch (target of the PR, or the repo's default):
35
-
36
- - GitHub: `gh pr view --json baseRefName -q .baseRefName` or `gh repo view --json defaultBranchRef -q .defaultBranchRef.name`
37
- - GitLab: `glab mr view -F json 2>/dev/null` → extract `target_branch` or `default_branch`
38
- - Fallback: `git symbolic-ref refs/remotes/origin/HEAD`, then `origin/main`, then `origin/master`, then `main`
39
-
40
- Print the detected base branch. Use it as `<base>` in all subsequent commands.
41
-
42
- ---
43
-
44
- ## Step 1: Check branch
45
-
46
- 1. Run `git branch --show-current`.
47
- 2. If on the base branch: output **"Nothing to review — you're on the base branch or have no changes against it."** and stop.
48
- 3. Run `git fetch origin <base> --quiet && git diff origin/<base> --stat`. If no diff, output the same message and stop.
49
-
50
- ---
51
-
52
- ## Step 1.5: Scope Drift Detection
53
-
54
- Check whether the diff matches what was requested.
55
-
56
- 1. Read `TODOS.md` (if it exists). Read PR description (`gh pr view --json body --jq .body 2>/dev/null || true`). Read commit messages (`git log origin/<base>..HEAD --oneline`).
57
- 2. Identify the **stated intent** — what was this branch supposed to accomplish?
58
- 3. Run `git diff origin/<base>...HEAD --stat` and compare against the stated intent.
59
-
60
- Evaluate for:
61
- - **SCOPE CREEP** — files changed that are unrelated to stated intent; "while I was in there" changes
62
- - **MISSING REQUIREMENTS** — requirements from TODOS.md/PR description not in the diff; partial implementations
63
-
64
- Output (before the main review begins):
65
- ```
66
- Scope Check: [CLEAN / DRIFT DETECTED / REQUIREMENTS MISSING]
67
- Intent: <1-line summary of what was requested>
68
- Delivered: <1-line summary of what the diff actually does>
69
- [If drift: list each out-of-scope change]
70
- [If missing: list each unaddressed requirement]
71
- ```
72
-
73
- This is **INFORMATIONAL** — does not block the review.
74
-
75
- ---
76
-
77
- ## Step 1.6: Plan Completion Audit (optional)
78
-
79
- Check if there is a plan file referenced in the conversation context or a recent `.md` file in common plan locations (e.g., `~/.claude/plans/`, `.claude/plans/`). If found and relevant to the current branch:
80
-
81
- Extract actionable items (checkboxes, numbered steps, imperative statements, file-level specs, test requirements). Cross-reference each item against the diff:
82
- - **DONE** — clear evidence in diff
83
- - **PARTIAL** — some work started but incomplete
84
- - **NOT DONE** — no evidence in diff
85
- - **CHANGED** — goal met by different means
86
-
87
- For `PARTIAL` or `NOT DONE`, investigate why and rate impact (HIGH/MEDIUM/LOW). For HIGH-impact gaps, use AskUserQuestion:
88
- - A) Stop and implement missing items
89
- - B) Ship anyway + create P1 TODOs
90
- - C) Intentionally dropped
91
-
92
- Output format:
93
- ```
94
- PLAN COMPLETION AUDIT
95
- ═══════════════════════
96
- Plan: {path}
97
- [DONE] Create UserService — src/services/user_service.rb
98
- [NOT DONE] Add caching layer — no cache-related changes in diff
99
- COMPLETION: N/M DONE
100
- ```
101
-
102
- ---
103
-
104
- ## Step 2: Read the checklist
105
-
106
- Read `.claude/skills/review/checklist.md` (if it exists). If the file cannot be read, continue with the built-in checks below.
107
-
108
- ---
109
-
110
- ## Step 3: Get the diff
111
-
112
- ```bash
113
- git fetch origin <base> --quiet
114
- git diff origin/<base>
115
- ```
116
-
117
- ---
118
-
119
- ## Step 4: Critical review pass
120
-
121
- Apply these categories against the diff:
122
-
123
- **CRITICAL:**
124
- - **SQL & Data Safety** — string interpolation in queries, missing parameterization, N+1 patterns
125
- - **Race Conditions & Concurrency** — shared mutable state, missing locks, idempotency violations
126
- - **LLM Output Trust Boundary** — LLM output used in SQL, shell commands, or DB writes without validation
127
- - **Shell Injection** — user input passed to shell commands unsanitized
128
- - **Enum & Value Completeness** — new enum values/types not handled in all switch/case branches
129
-
130
- For Enum & Value Completeness: use Grep to find all files referencing sibling values, then Read those files. This requires looking outside the diff.
131
-
132
- **INFORMATIONAL:**
133
- - Async/sync mixing, column/field name safety, LLM prompt issues, type coercion, frontend/view issues, time window safety, completeness gaps, distribution/CI gaps
134
-
135
- **Finding format:**
136
- ```
137
- [SEVERITY] (confidence: N/10) file:line — description
138
- ```
139
-
140
- Confidence scale:
141
- - 9-10: Verified by reading specific code, concrete bug demonstrated
142
- - 7-8: High confidence pattern match
143
- - 5-6: Moderate — show with caveat "Medium confidence, verify this is actually an issue"
144
- - 3-4: Low — include in appendix only
145
- - 1-2: Speculation — only report if P0
146
-
147
- ---
148
-
149
- ## Step 4.5: Adversarial review (always-on)
150
-
151
- Dispatch an independent subagent via the Agent tool to review the diff with fresh context. Subagent prompt:
152
-
153
- > "Run `git diff origin/<base>` to get the diff. Think like an attacker and a chaos engineer. Find ways this code will fail in production: edge cases, race conditions, security holes, resource leaks, failure modes, silent data corruption, logic errors, error handling that swallows failures, trust boundary violations. For each finding, classify as FIXABLE (you know how to fix it) or INVESTIGATE (needs human judgment)."
154
-
155
- Present findings under `ADVERSARIAL REVIEW (subagent):`. FIXABLE findings flow into the Fix-First pipeline. INVESTIGATE findings are informational.
156
-
157
- ---
158
-
159
- ## Step 5: Fix-First Review
160
-
161
- ### 5a: Classify each finding
162
-
163
- For each finding from Steps 4 and 4.5:
164
- - **AUTO-FIX** — mechanical, low-risk, single-file changes (dead code, stale comments, obvious formatting)
165
- - **ASK** — architectural, security-sensitive, ambiguous scope, or user preference
166
-
167
- ### 5b: Apply AUTO-FIX items
168
-
169
- Apply each fix directly. For each:
170
- `[AUTO-FIXED] [file:line] Problem → what you did`
171
-
172
- ### 5c: Batch-ask about ASK items
173
-
174
- Present all ASK items in one AskUserQuestion:
175
- ```
176
- I auto-fixed N issues. M need your input:
177
-
178
- 1. [CRITICAL] file:line — description
179
- Fix: recommended fix
180
- → A) Fix B) Skip
181
-
182
- RECOMMENDATION: Fix all — [reason].
183
- ```
184
-
185
- ### 5d: Apply approved fixes
186
-
187
- Apply fixes for items where the user chose "Fix."
188
-
189
- ---
190
-
191
- ## Step 5.5: TODOS cross-reference
192
-
193
- Read `TODOS.md` (if it exists). Cross-reference the PR:
194
- - Does this PR close any open TODOs? Note: "This PR addresses TODO: <title>"
195
- - Does this PR create work that should become a TODO? Flag as informational.
196
-
197
- ---
198
-
199
- ## Step 5.6: Documentation staleness check
200
-
201
- For each `.md` file in the repo root: if the code it describes was changed but the doc was NOT updated in this branch, flag as informational:
202
- "Documentation may be stale: [file] describes [feature] but code changed. Consider running the `document-release` skill."
203
-
204
- ---
205
-
206
- ## Completion Status
207
-
208
- Report one of:
209
- - **DONE** — All steps completed, no blocking issues.
210
- - **DONE_WITH_CONCERNS** — Completed with issues the user should know about.
211
- - **BLOCKED** — Cannot proceed. State what is blocking.
212
- - **NEEDS_CONTEXT** — Missing information required to continue.
213
-
214
- ## Important Rules
215
-
216
- - **Read the FULL diff before commenting.** Do not flag issues already addressed in the diff.
217
- - **Fix-first, not read-only.** AUTO-FIX items are applied directly; ASK items only after user approval.
218
- - **Never commit, push, or create PRs** — that is the `ship` skill's job.
219
- - **Be terse.** One line problem, one line fix. No preamble.
220
- - **Only flag real problems.** Skip anything that is fine.
221
-
222
- ## Engineering-manager checklist
223
-
224
- Beyond the structural/safety scan above, also review the change as an engineering manager would —
225
- this is the angle the old `eng-review` skill covered, now folded in here so there is one
226
- pre-merge review:
227
-
228
- - **Architecture fit & data flow:** does the change follow the project's established patterns and
229
- data flow, or does it drift? Flag architectural drift.
230
- - **Edge cases & error paths:** are unhandled inputs, failure modes, and boundary conditions
231
- covered?
232
- - **Test coverage:** is the changed behavior covered by tests that verify behavior (not just
233
- mocks)? Missing or weak tests for changed behavior is a blocking issue.
234
- - **Performance:** any obvious regressions (N+1, unbounded growth, needless work in hot paths)?
235
-
236
- Output a pass/block verdict with specific, actionable findings. A reviewer never reviews their
237
- own implementation (see `policy/roles.md`).
238
-
239
- ## Interactive cross-model review
240
-
241
- Outside the loop, run `yoke review` to have a *second* model review your current diff
242
- before you commit or push. It resolves to the first available of codex → gemini → claude
243
- (preferring a model other than the one you are driving), reviews the uncommitted working
244
- tree by default (or `--base=<ref>` for a branch range), and exits non-zero if it finds a
245
- blocking issue — so it chains as a gate (`... && yoke review`) or a pre-push hook.
246
- This is the interactive counterpart to the loop's `--review`/`--reviewer`.
1
+ ---
2
+ name: review
3
+ description: |
4
+ Pre-merge code review — the single canonical review of a change before it lands. Covers BOTH
5
+ diff safety/structure (SQL safety, LLM trust-boundary violations, conditional side effects)
6
+ AND engineering quality (architecture fit, edge cases, test coverage, performance). Use when
7
+ asked to "review this PR", "code review", "pre-landing review", "check my diff", or before
8
+ merging. (For plan-time review use plan-eng-review or plan-ceo-review instead.)
9
+ triggers:
10
+ - review this pr
11
+ - code review
12
+ - check my diff
13
+ - pre-landing review
14
+ ---
15
+
16
+ # Pre-Landing PR Review
17
+
18
+ You are running the `review` workflow. Analyze the current branch's diff against the base branch for structural issues that tests don't catch.
19
+
20
+ ---
21
+
22
+ ## Step 0: Detect platform and base branch
23
+
24
+ Detect the git hosting platform from the remote URL:
25
+
26
+ ```bash
27
+ git remote get-url origin 2>/dev/null
28
+ ```
29
+
30
+ - URL contains "github.com" → platform is **GitHub**
31
+ - URL contains "gitlab" → platform is **GitLab**
32
+ - Otherwise check: `gh auth status 2>/dev/null` → GitHub; `glab auth status 2>/dev/null` → GitLab; neither → unknown
33
+
34
+ Determine the base branch (target of the PR, or the repo's default):
35
+
36
+ - GitHub: `gh pr view --json baseRefName -q .baseRefName` or `gh repo view --json defaultBranchRef -q .defaultBranchRef.name`
37
+ - GitLab: `glab mr view -F json 2>/dev/null` → extract `target_branch` or `default_branch`
38
+ - Fallback: `git symbolic-ref refs/remotes/origin/HEAD`, then `origin/main`, then `origin/master`, then `main`
39
+
40
+ Print the detected base branch. Use it as `<base>` in all subsequent commands.
41
+
42
+ ---
43
+
44
+ ## Step 1: Check branch
45
+
46
+ 1. Run `git branch --show-current`.
47
+ 2. If on the base branch: output **"Nothing to review — you're on the base branch or have no changes against it."** and stop.
48
+ 3. Run `git fetch origin <base> --quiet && git diff origin/<base> --stat`. If no diff, output the same message and stop.
49
+
50
+ ---
51
+
52
+ ## Step 1.5: Scope Drift Detection
53
+
54
+ Check whether the diff matches what was requested.
55
+
56
+ 1. Read `TODOS.md` (if it exists). Read PR description (`gh pr view --json body --jq .body 2>/dev/null || true`). Read commit messages (`git log origin/<base>..HEAD --oneline`).
57
+ 2. Identify the **stated intent** — what was this branch supposed to accomplish?
58
+ 3. Run `git diff origin/<base>...HEAD --stat` and compare against the stated intent.
59
+
60
+ Evaluate for:
61
+ - **SCOPE CREEP** — files changed that are unrelated to stated intent; "while I was in there" changes
62
+ - **MISSING REQUIREMENTS** — requirements from TODOS.md/PR description not in the diff; partial implementations
63
+
64
+ Output (before the main review begins):
65
+ ```
66
+ Scope Check: [CLEAN / DRIFT DETECTED / REQUIREMENTS MISSING]
67
+ Intent: <1-line summary of what was requested>
68
+ Delivered: <1-line summary of what the diff actually does>
69
+ [If drift: list each out-of-scope change]
70
+ [If missing: list each unaddressed requirement]
71
+ ```
72
+
73
+ This is **INFORMATIONAL** — does not block the review.
74
+
75
+ ---
76
+
77
+ ## Step 1.6: Plan Completion Audit (optional)
78
+
79
+ Check if there is a plan file referenced in the conversation context or a recent `.md` file in common plan locations (e.g., `~/.claude/plans/`, `.claude/plans/`). If found and relevant to the current branch:
80
+
81
+ Extract actionable items (checkboxes, numbered steps, imperative statements, file-level specs, test requirements). Cross-reference each item against the diff:
82
+ - **DONE** — clear evidence in diff
83
+ - **PARTIAL** — some work started but incomplete
84
+ - **NOT DONE** — no evidence in diff
85
+ - **CHANGED** — goal met by different means
86
+
87
+ For `PARTIAL` or `NOT DONE`, investigate why and rate impact (HIGH/MEDIUM/LOW). For HIGH-impact gaps, use AskUserQuestion:
88
+ - A) Stop and implement missing items
89
+ - B) Ship anyway + create P1 TODOs
90
+ - C) Intentionally dropped
91
+
92
+ Output format:
93
+ ```
94
+ PLAN COMPLETION AUDIT
95
+ ═══════════════════════
96
+ Plan: {path}
97
+ [DONE] Create UserService — src/services/user_service.rb
98
+ [NOT DONE] Add caching layer — no cache-related changes in diff
99
+ COMPLETION: N/M DONE
100
+ ```
101
+
102
+ ---
103
+
104
+ ## Step 2: Read the checklist
105
+
106
+ Read `.claude/skills/review/checklist.md` (if it exists). If the file cannot be read, continue with the built-in checks below.
107
+
108
+ ---
109
+
110
+ ## Step 3: Get the diff
111
+
112
+ ```bash
113
+ git fetch origin <base> --quiet
114
+ git diff origin/<base>
115
+ ```
116
+
117
+ ---
118
+
119
+ ## Step 4: Critical review pass
120
+
121
+ Apply these categories against the diff:
122
+
123
+ **CRITICAL:**
124
+ - **SQL & Data Safety** — string interpolation in queries, missing parameterization, N+1 patterns
125
+ - **Race Conditions & Concurrency** — shared mutable state, missing locks, idempotency violations
126
+ - **LLM Output Trust Boundary** — LLM output used in SQL, shell commands, or DB writes without validation
127
+ - **Shell Injection** — user input passed to shell commands unsanitized
128
+ - **Enum & Value Completeness** — new enum values/types not handled in all switch/case branches
129
+
130
+ For Enum & Value Completeness: use Grep to find all files referencing sibling values, then Read those files. This requires looking outside the diff.
131
+
132
+ **INFORMATIONAL:**
133
+ - Async/sync mixing, column/field name safety, LLM prompt issues, type coercion, frontend/view issues, time window safety, completeness gaps, distribution/CI gaps
134
+
135
+ **Finding format:**
136
+ ```
137
+ [SEVERITY] (confidence: N/10) file:line — description
138
+ ```
139
+
140
+ Confidence scale:
141
+ - 9-10: Verified by reading specific code, concrete bug demonstrated
142
+ - 7-8: High confidence pattern match
143
+ - 5-6: Moderate — show with caveat "Medium confidence, verify this is actually an issue"
144
+ - 3-4: Low — include in appendix only
145
+ - 1-2: Speculation — only report if P0
146
+
147
+ ---
148
+
149
+ ## Step 4.5: Adversarial review (always-on)
150
+
151
+ Dispatch an independent subagent via the Agent tool to review the diff with fresh context. Subagent prompt:
152
+
153
+ > "Run `git diff origin/<base>` to get the diff. Think like an attacker and a chaos engineer. Find ways this code will fail in production: edge cases, race conditions, security holes, resource leaks, failure modes, silent data corruption, logic errors, error handling that swallows failures, trust boundary violations. For each finding, classify as FIXABLE (you know how to fix it) or INVESTIGATE (needs human judgment)."
154
+
155
+ Present findings under `ADVERSARIAL REVIEW (subagent):`. FIXABLE findings flow into the Fix-First pipeline. INVESTIGATE findings are informational.
156
+
157
+ ---
158
+
159
+ ## Step 5: Fix-First Review
160
+
161
+ ### 5a: Classify each finding
162
+
163
+ For each finding from Steps 4 and 4.5:
164
+ - **AUTO-FIX** — mechanical, low-risk, single-file changes (dead code, stale comments, obvious formatting)
165
+ - **ASK** — architectural, security-sensitive, ambiguous scope, or user preference
166
+
167
+ ### 5b: Apply AUTO-FIX items
168
+
169
+ Apply each fix directly. For each:
170
+ `[AUTO-FIXED] [file:line] Problem → what you did`
171
+
172
+ ### 5c: Batch-ask about ASK items
173
+
174
+ Present all ASK items in one AskUserQuestion:
175
+ ```
176
+ I auto-fixed N issues. M need your input:
177
+
178
+ 1. [CRITICAL] file:line — description
179
+ Fix: recommended fix
180
+ → A) Fix B) Skip
181
+
182
+ RECOMMENDATION: Fix all — [reason].
183
+ ```
184
+
185
+ ### 5d: Apply approved fixes
186
+
187
+ Apply fixes for items where the user chose "Fix."
188
+
189
+ ---
190
+
191
+ ## Step 5.5: TODOS cross-reference
192
+
193
+ Read `TODOS.md` (if it exists). Cross-reference the PR:
194
+ - Does this PR close any open TODOs? Note: "This PR addresses TODO: <title>"
195
+ - Does this PR create work that should become a TODO? Flag as informational.
196
+
197
+ ---
198
+
199
+ ## Step 5.6: Documentation staleness check
200
+
201
+ For each `.md` file in the repo root: if the code it describes was changed but the doc was NOT updated in this branch, flag as informational:
202
+ "Documentation may be stale: [file] describes [feature] but code changed. Consider running the `document-release` skill."
203
+
204
+ ---
205
+
206
+ ## Completion Status
207
+
208
+ Report one of:
209
+ - **DONE** — All steps completed, no blocking issues.
210
+ - **DONE_WITH_CONCERNS** — Completed with issues the user should know about.
211
+ - **BLOCKED** — Cannot proceed. State what is blocking.
212
+ - **NEEDS_CONTEXT** — Missing information required to continue.
213
+
214
+ ## Important Rules
215
+
216
+ - **Read the FULL diff before commenting.** Do not flag issues already addressed in the diff.
217
+ - **Fix-first, not read-only.** AUTO-FIX items are applied directly; ASK items only after user approval.
218
+ - **Never commit, push, or create PRs** — that is the `ship` skill's job.
219
+ - **Be terse.** One line problem, one line fix. No preamble.
220
+ - **Only flag real problems.** Skip anything that is fine.
221
+
222
+ ## Engineering-manager checklist
223
+
224
+ Beyond the structural/safety scan above, also review the change as an engineering manager would —
225
+ this is the angle the old `eng-review` skill covered, now folded in here so there is one
226
+ pre-merge review:
227
+
228
+ - **Architecture fit & data flow:** does the change follow the project's established patterns and
229
+ data flow, or does it drift? Flag architectural drift.
230
+ - **Edge cases & error paths:** are unhandled inputs, failure modes, and boundary conditions
231
+ covered?
232
+ - **Test coverage:** is the changed behavior covered by tests that verify behavior (not just
233
+ mocks)? Missing or weak tests for changed behavior is a blocking issue.
234
+ - **Performance:** any obvious regressions (N+1, unbounded growth, needless work in hot paths)?
235
+
236
+ Output a pass/block verdict with specific, actionable findings. A reviewer never reviews their
237
+ own implementation (see `policy/roles.md`).
238
+
239
+ ## Interactive cross-model review
240
+
241
+ Outside the loop, run `yoke review` to have a *second* model review your current diff
242
+ before you commit or push. It resolves to the first available of codex → gemini → claude
243
+ (preferring a model other than the one you are driving), reviews the uncommitted working
244
+ tree by default (or `--base=<ref>` for a branch range), and exits non-zero if it finds a
245
+ blocking issue — so it chains as a gate (`... && yoke review`) or a pre-push hook.
246
+ This is the interactive counterpart to the loop's `--review`/`--reviewer`.