@mrciphersmith/keryx 0.2.70 → 0.2.71

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (177) hide show
  1. package/dist/cli.js +10357 -4676
  2. package/docs/README.md +54 -0
  3. package/docs/requirements/shared-agent-context/README.md +104 -0
  4. package/package.json +3 -2
  5. package/src/gdskills/bundled/rules/core/model-selection.mdc +184 -31
  6. package/src/gdskills/bundled/rules/core/skills-storage-workflow.mdc +36 -0
  7. package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.codex.md +1 -1
  8. package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.cursor.md +1 -1
  9. package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.md +2 -1
  10. package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.opencode.md +1 -1
  11. package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.zed.md +1 -1
  12. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.codex.md +1 -1
  13. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.cursor.md +1 -1
  14. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.md +1 -1
  15. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.opencode.md +1 -1
  16. package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.zed.md +1 -1
  17. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.codex.md +1 -1
  18. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.cursor.md +1 -1
  19. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.md +1 -1
  20. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.opencode.md +1 -1
  21. package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.zed.md +1 -1
  22. package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.codex.md +1 -1
  23. package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.cursor.md +1 -1
  24. package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.md +1 -1
  25. package/src/gdskills/bundled/skills/orchestration/flow-orchestrator/SKILL.md +78 -19
  26. package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.codex.md +1 -1
  27. package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.cursor.md +1 -1
  28. package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.md +1 -1
  29. package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.opencode.md +1 -1
  30. package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.zed.md +1 -1
  31. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.codex.md +1 -1
  32. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.cursor.md +1 -1
  33. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.md +2 -1
  34. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.opencode.md +1 -1
  35. package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.zed.md +1 -1
  36. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.codex.md +28 -3
  37. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.cursor.md +28 -3
  38. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +28 -3
  39. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.opencode.md +28 -3
  40. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.zed.md +28 -3
  41. package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.codex.md +20 -2
  42. package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.cursor.md +20 -2
  43. package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.md +21 -2
  44. package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.opencode.md +20 -2
  45. package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.zed.md +20 -2
  46. package/src/gdskills/bundled/skills/planning/autodoc-analyst/SKILL.md +2 -1
  47. package/src/gdskills/bundled/skills/planning/autodoc-architect/SKILL.md +3 -1
  48. package/src/gdskills/bundled/skills/planning/autodoc-assembler/SKILL.md +2 -1
  49. package/src/gdskills/bundled/skills/planning/autodoc-orchestrator/SKILL.md +2 -1
  50. package/src/gdskills/bundled/skills/planning/autodoc-scanner/SKILL.md +2 -1
  51. package/src/gdskills/bundled/skills/planning/autodoc-writer/SKILL.md +2 -1
  52. package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.codex.md +1 -1
  53. package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.cursor.md +1 -1
  54. package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
  55. package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.codex.md +1 -1
  56. package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.cursor.md +1 -1
  57. package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.md +2 -1
  58. package/src/gdskills/bundled/skills/planning/docpack-orchestrator/SKILL.md +1 -1
  59. package/src/gdskills/bundled/skills/planning/docpack-review/SKILL.md +1 -1
  60. package/src/gdskills/bundled/skills/planning/interview/SKILL.codex.md +1 -1
  61. package/src/gdskills/bundled/skills/planning/interview/SKILL.cursor.md +1 -1
  62. package/src/gdskills/bundled/skills/planning/interview/SKILL.md +1 -1
  63. package/src/gdskills/bundled/skills/planning/interviewer/SKILL.codex.md +1 -1
  64. package/src/gdskills/bundled/skills/planning/interviewer/SKILL.cursor.md +1 -1
  65. package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
  66. package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.codex.md +1 -1
  67. package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.cursor.md +1 -1
  68. package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.md +2 -1
  69. package/src/gdskills/bundled/skills/planning/planner/SKILL.codex.md +1 -1
  70. package/src/gdskills/bundled/skills/planning/planner/SKILL.cursor.md +1 -1
  71. package/src/gdskills/bundled/skills/planning/planner/SKILL.md +2 -1
  72. package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.codex.md +1 -1
  73. package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.cursor.md +1 -1
  74. package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.md +1 -1
  75. package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.opencode.md +1 -1
  76. package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.zed.md +1 -1
  77. package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.codex.md +1 -1
  78. package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.cursor.md +1 -1
  79. package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.md +2 -1
  80. package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.codex.md +1 -1
  81. package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.cursor.md +1 -1
  82. package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.md +2 -1
  83. package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.codex.md +1 -1
  84. package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.cursor.md +1 -1
  85. package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.md +2 -1
  86. package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.codex.md +1 -1
  87. package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.cursor.md +1 -1
  88. package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.md +2 -1
  89. package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.codex.md +1 -1
  90. package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.cursor.md +1 -1
  91. package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.md +1 -1
  92. package/src/gdskills/bundled/skills/platform/hookify/SKILL.codex.md +1 -1
  93. package/src/gdskills/bundled/skills/platform/hookify/SKILL.cursor.md +1 -1
  94. package/src/gdskills/bundled/skills/platform/hookify/SKILL.md +1 -1
  95. package/src/gdskills/bundled/skills/quality/changelog/SKILL.codex.md +1 -1
  96. package/src/gdskills/bundled/skills/quality/changelog/SKILL.cursor.md +1 -1
  97. package/src/gdskills/bundled/skills/quality/changelog/SKILL.md +1 -1
  98. package/src/gdskills/bundled/skills/quality/commit/SKILL.codex.md +1 -1
  99. package/src/gdskills/bundled/skills/quality/commit/SKILL.cursor.md +1 -1
  100. package/src/gdskills/bundled/skills/quality/commit/SKILL.md +1 -1
  101. package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.codex.md +1 -1
  102. package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.cursor.md +1 -1
  103. package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.md +1 -1
  104. package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.codex.md +1 -1
  105. package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.cursor.md +1 -1
  106. package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.md +1 -1
  107. package/src/gdskills/bundled/skills/quality/deploy/SKILL.codex.md +1 -1
  108. package/src/gdskills/bundled/skills/quality/deploy/SKILL.cursor.md +1 -1
  109. package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
  110. package/src/gdskills/bundled/skills/quality/metaproject-security/SKILL.md +1 -1
  111. package/src/gdskills/bundled/skills/quality/perf-check/SKILL.codex.md +1 -1
  112. package/src/gdskills/bundled/skills/quality/perf-check/SKILL.cursor.md +1 -1
  113. package/src/gdskills/bundled/skills/quality/perf-check/SKILL.md +1 -1
  114. package/src/gdskills/bundled/skills/quality/pr/SKILL.codex.md +1 -1
  115. package/src/gdskills/bundled/skills/quality/pr/SKILL.cursor.md +1 -1
  116. package/src/gdskills/bundled/skills/quality/pr/SKILL.md +1 -1
  117. package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.codex.md +1 -1
  118. package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.cursor.md +1 -1
  119. package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.md +1 -1
  120. package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.opencode.md +1 -1
  121. package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.zed.md +1 -1
  122. package/src/gdskills/bundled/skills/quality/push/SKILL.codex.md +1 -1
  123. package/src/gdskills/bundled/skills/quality/push/SKILL.cursor.md +1 -1
  124. package/src/gdskills/bundled/skills/quality/push/SKILL.md +1 -1
  125. package/src/gdskills/bundled/skills/quality/security-audit/SKILL.codex.md +1 -1
  126. package/src/gdskills/bundled/skills/quality/security-audit/SKILL.cursor.md +1 -1
  127. package/src/gdskills/bundled/skills/quality/security-audit/SKILL.md +1 -1
  128. package/src/gdskills/bundled/skills/quality/test-gen/SKILL.codex.md +1 -1
  129. package/src/gdskills/bundled/skills/quality/test-gen/SKILL.cursor.md +1 -1
  130. package/src/gdskills/bundled/skills/quality/test-gen/SKILL.md +1 -1
  131. package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.codex.md +1 -1
  132. package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.cursor.md +1 -1
  133. package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.md +1 -1
  134. package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.opencode.md +1 -1
  135. package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.zed.md +1 -1
  136. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.codex.md +1 -1
  137. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.cursor.md +1 -1
  138. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.md +1 -1
  139. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.opencode.md +1 -1
  140. package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.zed.md +1 -1
  141. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.codex.md +1 -1
  142. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.cursor.md +1 -1
  143. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.md +1 -1
  144. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.opencode.md +1 -1
  145. package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.zed.md +1 -1
  146. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.codex.md +1 -1
  147. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.cursor.md +1 -1
  148. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.md +2 -1
  149. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.opencode.md +1 -1
  150. package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.zed.md +1 -1
  151. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.codex.md +1 -1
  152. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.cursor.md +1 -1
  153. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.md +1 -1
  154. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.opencode.md +1 -1
  155. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.zed.md +1 -1
  156. package/src/gdskills/bundled/skills/review/review-architecture/SKILL.md +37 -10
  157. package/src/gdskills/bundled/skills/review/review-backend/SKILL.md +48 -14
  158. package/src/gdskills/bundled/skills/review/review-clean-code/SKILL.md +49 -12
  159. package/src/gdskills/bundled/skills/review/review-core-boundaries/SKILL.md +34 -2
  160. package/src/gdskills/bundled/skills/review/review-flow-graph/SKILL.md +33 -2
  161. package/src/gdskills/bundled/skills/review/review-frontend/SKILL.md +70 -29
  162. package/src/gdskills/bundled/skills/review/review-frontend-conventions/SKILL.md +34 -3
  163. package/src/gdskills/bundled/skills/review/review-highload/SKILL.md +49 -15
  164. package/src/gdskills/bundled/skills/review/review-logic/SKILL.md +39 -11
  165. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +590 -21
  166. package/src/gdskills/bundled/skills/review/review-orchestrator/reviewer-finding.schema.json +7 -0
  167. package/src/gdskills/bundled/skills/review/review-orchestrator/verification-claim.schema.json +78 -0
  168. package/src/gdskills/bundled/skills/review/review-performance/SKILL.md +43 -13
  169. package/src/gdskills/bundled/skills/review/review-pr-feedback/SKILL.md +8 -2
  170. package/src/gdskills/bundled/skills/review/review-regression/SKILL.md +185 -0
  171. package/src/gdskills/bundled/skills/review/review-security-code/SKILL.md +44 -13
  172. package/src/gdskills/bundled/skills/review/review-style/SKILL.md +26 -6
  173. package/src/gdskills/bundled/skills/review/review-testing-practices/SKILL.md +35 -3
  174. package/src/gdskills/bundled/skills/review/review-verifier/SKILL.md +276 -0
  175. package/src/gdskills/contracts/review-finding.schema.json +119 -1
  176. package/src/gdskills/contracts/subagent-dispatch.schema.json +59 -3
  177. package/src/gdskills/bundled/skills/review/review-strict/SKILL.md +0 -328
@@ -4,10 +4,9 @@ description: |
4
4
  Use when: a code review is requested and the user does not explicitly name a specialized reviewer.
5
5
  Handles "review", "code review", "review PR", "review --frontend", "review --backend",
6
6
  "review --architecture", "review --security", "review --performance", "review --style",
7
- "review --strict", "review --project-conventions", "review --legacy-profiles", "review --all". Routes to specialized reviewers in parallel and
7
+ "review --verify", "review --project-conventions", "review --legacy-profiles", "review --all". Routes to specialized reviewers in parallel and
8
8
  consolidates findings into one unified report.
9
9
  NOT for: running a single specialized reviewer — invoke it directly by name instead.
10
- version: "1.6.0"
11
10
  triggers:
12
11
  - "review"
13
12
  - "code review"
@@ -18,7 +17,7 @@ triggers:
18
17
  - "review --security"
19
18
  - "review --performance"
20
19
  - "review --style"
21
- - "review --strict"
20
+ - "review --verify"
22
21
  - "review --all"
23
22
  - "review --clean-code"
24
23
  - "review --highload"
@@ -34,10 +33,10 @@ triggers:
34
33
  - "review --mobx-store"
35
34
  metadata:
36
35
  author: "MrCipherSmith"
37
- version: "1.6.0"
36
+ version: "1.8.0"
38
37
  category: "review"
38
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
39
39
  license: "MIT"
40
- compatibility: "cursor,codex,zed,opencode,claude"
41
40
  ---
42
41
 
43
42
  # Review Orchestrator
@@ -52,20 +51,47 @@ unified report sorted by severity. It does not perform any review logic itself.
52
51
 
53
52
  ```
54
53
  Review Orchestrator Progress:
54
+ - [ ] Step 0: On a PR target, collect external comments — `keryx review comments collect`
55
55
  - [ ] Step 1: Build Review Context Pack (PR metadata, scope, rules, context_doc summary)
56
56
  - [ ] Step 2: Detect review mode (diff mode vs. path mode)
57
- - [ ] Step 3: Collect bounded scope - git diff OR file list from path
57
+ - [ ] Step 3: Build the bounded scope with `keryx review scope` — never by hand
58
+ - [ ] Step 3b: On a deep round, compute scope B with `keryx review blast-radius` — never by browsing — and KEEP the `--json` file; `review ingest --blast-radius <file>` is refused without it
58
59
  - [ ] Step 4: Parse flags / auto-detect domain from scope
59
60
  - [ ] Step 5: Ask user to confirm optional convention reviewers (legacy/profile reviewers are flag-only, never prompted)
60
- - [ ] Step 6: Plan sub-agent dispatch, token budgets, and model strategy
61
+ - [ ] Step 6: Plan sub-agent dispatch and token budgets, and compute each dispatch's model with `keryx review tier` — never by hand
61
62
  - [ ] Step 7: Stage 1 gate - spec compliance check (if issue/task provided)
62
63
  - [ ] Step 8: Dispatch selected reviewers in PARALLEL with reviewer-input schema
63
64
  - [ ] Step 9: Collect reviewer-finding schema results and handle NEEDS_CONTEXT
64
- - [ ] Step 10: Run strict synthesis when blockers/majors exist or --strict is set
65
+ - [ ] Step 10: Wave C — dispatch `review-verifier` over the consolidated findings
65
66
  - [ ] Step 11: Sort by severity, deduplicate, emit unified report
66
67
  - [ ] Step 12: Emit the machine-readable `keryx:findings` block alongside the report
68
+ - [ ] Step 13: Report the stage counts: dropped by pre-filter, refuted by the verifier, retained
69
+ - [ ] Step 14: AFTER THE FINAL ROUND ONLY — answer every external comment once, `keryx review comments reply --final`
67
70
  ```
68
71
 
72
+ Step 0 runs on **every** round. Step 14 runs **once**, after the last one. They are
73
+ two commands for that reason: a caller that already runs collection per round
74
+ would carry the posting along with it, and the reviewer would get six replies to
75
+ one comment.
76
+
77
+ ### Step 6 — the model is computed, not chosen
78
+
79
+ Before dispatching each reviewer, run `keryx review tier` with the signals you
80
+ already hold (`--scope`, `--findings`, `--diff-lines`, `--fix-attempt`,
81
+ `--verifier`, `--security`, `--forced-strategy-change`) and paste the `model`
82
+ block it prints into that dispatch.
83
+
84
+ Do NOT assign the tier by reading the table in `rules/core/model-selection.mdc`.
85
+ Working it out in your head is exactly the mechanical step that rule moves into
86
+ code — and it is the step that was documented as running for a whole release
87
+ while nothing called it.
88
+
89
+ The command names no model. It ranks whatever your provider reports at runtime
90
+ and places the tiers relative to your own model; when it cannot rank anything it
91
+ prints `inherit: true` and exits 0, which means the dispatch runs on the session
92
+ model. That is a correct answer, not a failure — never "fix" it by writing a
93
+ model id into a dispatch.
94
+
69
95
  ---
70
96
 
71
97
  ## Step 12 — the `keryx:findings` block
@@ -127,7 +153,7 @@ which of its inputs were recovered and which were never written down.
127
153
 
128
154
  | Field | Type | Required | Description |
129
155
  |-------|------|----------|-------------|
130
- | `flags` | string[] | no | One or more of: `--frontend`, `--backend`, `--architecture`, `--security`, `--performance`, `--style`, `--clean-code`, `--highload`, `--project-conventions`, `--frontend-conventions`, `--testing-practices`, `--core-boundaries`, `--flow-graph`, `--legacy-profiles`, `--code-ai`, `--b091`, `--code-style`, `--mobx-store`, `--strict`, `--all` |
156
+ | `flags` | string[] | no | One or more of: `--frontend`, `--backend`, `--architecture`, `--security`, `--performance`, `--style`, `--clean-code`, `--highload`, `--project-conventions`, `--frontend-conventions`, `--testing-practices`, `--core-boundaries`, `--flow-graph`, `--legacy-profiles`, `--code-ai`, `--b091`, `--code-style`, `--mobx-store`, `--verify`, `--all` |
131
157
  | `path` | string | no | File or directory path to review (e.g., `src/stores/`, `src/components/UserCard.tsx`). Activates **path mode** — reviews the files at this path directly, not a git diff. |
132
158
  | `commit_range` | string | no | Explicit commit hash or range (e.g., `abc123..HEAD`). Overrides merge-base detection. Ignored in path mode. |
133
159
  | `issue_url` | string | no | GitHub issue or task URL. If provided, Stage 1 gate checks spec compliance before dispatching reviewers. |
@@ -136,6 +162,8 @@ which of its inputs were recovered and which were never written down.
136
162
  | `token_budget` | object | no | Optional budget controls: `{total, per_reviewer, diff_max_chars, file_max_chars}`. |
137
163
  | `model_strategy` | string | no | `current`, `ask`, or `adaptive`. Default: `current`; do not switch models unless user or automation allows it. |
138
164
  | `managed_review` | object | no | Optional managed review mode: `{mode, target, target_ref, flow_id, reviewers}` where mode is `lightweight`, `attach-review`, `review-flow`, or `ingest`. |
165
+ | `verification_mode` | string | no | `off`, `annotate`, or `filter`. Default `annotate` — verdicts are recorded and nothing is removed. See Wave C. |
166
+ | `pr_comments` | object | no | `{enabled, max_replies_total, max_sentences_per_reply}`. Defaults: enabled when a PR exists, `30`, `2`. Collect every round, reply once at the end. See External PR comments. |
139
167
 
140
168
  ---
141
169
 
@@ -166,10 +194,66 @@ Runtime CLI surface:
166
194
  keryx review attach --flow <id> --target <kind> --ref <ref>
167
195
  keryx review start --target <kind> --ref <ref>
168
196
  keryx review ingest --report <path> [--flow <id>] --ref <ref>
197
+ [--verifications <file>] [--verification-mode off|annotate|filter]
198
+ [--scope <scope.json>] [--blast-radius <blast-radius.json>]
199
+ [--refuted <file>]
200
+ # --blast-radius is REQUIRED whenever this round dispatched a
201
+ # scope-B reviewer (review-regression). See below.
169
202
  keryx review status <review-id-or-path>
170
203
  keryx review complete <review-id-or-path>
204
+ [--finding <id> --disposition <state> --evidence <ref>]...
171
205
  ```
172
206
 
207
+ **An unrecognised option is refused, not ignored.** A misspelling used to be
208
+ accepted with exit 0, so `review complete --disposition ...` printed
209
+ `status: closed` and wrote nothing at all.
210
+
211
+ `--verifications` takes what `review-verifier` returned. `--scope` takes the
212
+ whole `--json` output of `keryx review scope`, so the package records what the
213
+ pre-filter dropped, **with a reason per drop**, as well as what the verifier
214
+ refuted. Without it — and with no `## Pre-filter scope` block already in the
215
+ package — the record says **`not recorded`** for that stage rather than `0`;
216
+ "dropped nothing" and "never ran" are different facts.
217
+
218
+ `--blast-radius` is **not optional on any round that dispatched
219
+ `review-regression`** — which is every `recommended` and `full` round, because
220
+ `review-regression` is in both profiles. AC3 is enforced in code: an ingest
221
+ carrying a scope-B finding with no blast-radius record is **refused**, and the
222
+ round is not recordable until the set is supplied. Pass the `--json` output of
223
+ `keryx review blast-radius` — the same record you computed at dispatch time to
224
+ bound the round.
225
+
226
+ Two channels, and both work:
227
+
228
+ - `--blast-radius <file>` — always correct, and the one to use.
229
+ - `.metaproject/reviews/blast-radius.json` (or
230
+ `.metaproject/flows/<flow>/reviews/blast-radius.json` for a flow-attached
231
+ round) — the handoff slot, read when no flag is passed. The ingest **consumes**
232
+ it: it is copied into the package it screened and removed from the slot, so a
233
+ later round cannot be screened against a set nobody recomputed.
234
+
235
+ Do **not** write the record inside the package directory before the ingest. A
236
+ round that names no `--review-id` has not allocated that directory yet, and
237
+ creating it makes the ingest take the next free name instead.
238
+
239
+ `--refuted` takes the findings this round **raised and then dismissed**, in the
240
+ same finding shape, each carrying `disposition: {state, evidence}` with one of
241
+ the `dismissed-*` states. Without it a package keeps only the survivors of an
242
+ unlogged triage — which is why precision measured over the recorded corpus
243
+ returns 100% whatever the reviewers actually got right.
244
+
245
+ `review complete --finding <id> --disposition <state> --evidence <ref>` records
246
+ what became of a finding when the fix round closes; the triple is repeatable, one
247
+ group per finding. States: `unknown`, `acted-on`, `dismissed-incorrect`,
248
+ `dismissed-wont-fix`, `dismissed-out-of-scope`, `dismissed-deprioritised`.
249
+ Everything except `unknown` must cite where the outcome is written down — the
250
+ commit, the test, the decision. A recorded state and its citation cannot be
251
+ overwritten by a later close; record a correction as a new round.
252
+
253
+ **Closing a fix round without dispositions leaves every finding reading
254
+ `unknown`.** That is not "the reviewers were right"; it is "nobody wrote down
255
+ what happened", and it is the single reason the precision figure cannot be read.
256
+
173
257
  Managed modes:
174
258
 
175
259
  - `lightweight`: report-only; no flow or managed review artifacts are created.
@@ -193,6 +277,147 @@ When attaching to a flow, resolve the flow by explicit `flow_id`, PR URL, issue
193
277
  URL, or branch metadata. Never mutate `.metaproject/flows/*/flow.json` from
194
278
  review code; Task Manager state changes remain owned by `keryx flow`.
195
279
 
280
+ ---
281
+
282
+ ## External PR comments
283
+
284
+ A human or a bot reviews our pull request. Before this existed, nothing collected
285
+ it, nothing fixed it and nothing answered it — silence was the behaviour for every
286
+ comment, and to the person who wrote it silence is indistinguishable from
287
+ disagreement.
288
+
289
+ **Collect every round. Answer once, at the end.** Both halves are mechanical and
290
+ both live in the CLI; the judgement — is this comment right, what do we say — stays
291
+ here.
292
+
293
+ ```text
294
+ keryx review comments collect --repo <owner/repo> --pr <n> --sha <head-sha>
295
+ [--self <login>] [--round <n>] [--out <findings.json>] [--json]
296
+ keryx review comments reply --repo <owner/repo> --pr <n> --outcomes <file|->
297
+ --sha <head-sha> --final [--dry-run]
298
+ [--max-replies <n>] [--max-sentences <n>] [--max-chars <n>]
299
+ [--flow-link <url>]
300
+ ```
301
+
302
+ Add `--fixtures <dir>` to either to run the whole loop against JSON on disk —
303
+ no token, no network, nothing posted. Use it to see what a reply pass would say
304
+ before it says it.
305
+
306
+ `--sha` is the commit you collected against, and it is required. The completion
307
+ gate compares it to the pull request's head: a collection that ran before the
308
+ comments arrived is **stale**, and a gate that could not tell the difference
309
+ would pass a flow with unanswered reviewers on it while printing
310
+ `0 outstanding`. A record with no SHA reads as "cannot be shown current", never
311
+ as "fresh".
312
+
313
+ ### What collection does, so you do not do it by hand
314
+
315
+ - Reads **all three** sources: inline review comments, review submissions and
316
+ their bodies, and PR-level discussion.
317
+ - **A bot reviewer is a reviewer.** CodeRabbit, Greptile and Copilot comments go
318
+ down exactly the same path as a human's. The bot flag is recorded so a report
319
+ can say who spoke; nothing filters on it.
320
+ - Excludes our own identity, and comments already answered — **unless** the thread
321
+ has a newer reply from somebody else, which makes the comment new again.
322
+ - Everything filtered is listed with its reason. A filter that removes silently
323
+ reads as "nobody commented".
324
+
325
+ ### Severity is classified, never invented
326
+
327
+ A comment on a review whose state is `CHANGES_REQUESTED` starts at **`major`**.
328
+ Everything else starts at **`minor`**. There is no third rule and no model call.
329
+
330
+ When the classifying fact is missing — an inline comment whose parent review was
331
+ not returned, or a review state GitHub does not document — the comment is **not
332
+ dropped and the severity is not guessed**: it takes the `minor` floor and carries
333
+ `basis: unclassified` naming what was missing. A derived `minor` and a defaulted
334
+ one are different claims, and a record that cannot tell them apart is the
335
+ `dismissed-out-of-scope: 0` failure in a new field.
336
+
337
+ You may **lower** a severity only by assigning a terminal disposition with a
338
+ reason. You may never silently drop an external comment.
339
+
340
+ ### The verifier cannot refute an external comment
341
+
342
+ An external finding enters the same fix loop as an internal one with one
343
+ exception: **a `refuted` verdict does not remove it and does not dismiss it.** A
344
+ human asked a question; a machine deciding the question was invalid is not an
345
+ answer. `keryx review ingest` turns that verdict into the disposition
346
+ `answered-disagree`, keeps the finding, and records the reclaim in `scope.md`.
347
+ `answered-disagree` still owes a reply explaining why.
348
+
349
+ The per-reviewer findings cap does not truncate external comments either, for the
350
+ same reason: the cap drops silently, and an external comment may not be dropped
351
+ silently.
352
+
353
+ ### Replying — once, at the end, briefly
354
+
355
+ The reply pass runs **after the final round and before the completion gate**, so
356
+ every reply states a settled outcome rather than an intention. `keryx review
357
+ comments reply` refuses without `--final`; it is not a reminder you can skip.
358
+
359
+ | Outcome | Reply is |
360
+ |---|---|
361
+ | `acted-on` | one sentence naming what changed, plus the commit SHA |
362
+ | `answered-disagree` | one or two sentences on why not, and a link to the flow's journal entry |
363
+ | `dismissed-out-of-scope` / `dismissed-deprioritised` | one sentence, and where it was recorded instead |
364
+
365
+ Rules, all of them enforced in code rather than asked for here:
366
+
367
+ - **At most two sentences per comment.** A longer reply is CUT to two and the
368
+ remainder is replaced by a link — the long version is not reachable from the
369
+ command's output. A truncation with no link to point at is refused outright: the
370
+ conclusion posted and the explanation nowhere is worse than either alternative.
371
+ - A fenced code block in a reply is refused. Link, do not paste.
372
+ - Replies go **in the thread**. A review submission body and a PR-level comment
373
+ have no thread — GitHub offers no reply endpoint for either — so those become one
374
+ top-level comment that names what it answers.
375
+ - **Never resolve or hide a thread we did not open.** Replying is ours; resolving
376
+ is the reviewer's call, and auto-resolving is how a bot silences a human. The
377
+ resolve, hide, minimise and dismiss endpoints are unreachable through the port
378
+ this command uses, GraphQL included.
379
+ - Exactly **one** reply per comment, and one disposition. A round that changed
380
+ nothing for a comment still gets a reply saying so, with a terminal disposition
381
+ — `unknown` is refused, because it is what an unanswered comment already reads
382
+ as.
383
+ - Capped at **30** replies. Beyond it, one summary comment and a backlog reported
384
+ by id.
385
+ - Handling is durable: `.metaproject/reviews/pr-comments/<owner>__<repo>__<n>.json`
386
+ records id, thread, author, url, first-seen round, handled-at, sha, disposition
387
+ and reply url, written after **every** post. A resumed session answers nobody
388
+ twice.
389
+
390
+ **The trade-off, stated rather than hidden:** a reviewer who comments early waits
391
+ until the end. That is deliberate — answering with a work-in-progress state that
392
+ later changes is worse. If a comment **blocks** progress rather than reporting a
393
+ problem, mark its outcome `escalate: true`: it leaves the reply queue, is reported
394
+ to the operator immediately, and the command exits non-zero. Answering a blocking
395
+ question at the end answers the wrong question late.
396
+
397
+ ---
398
+
399
+ ## Everything written to GitHub is brief
400
+
401
+ One rule, applied to every outward surface: **PR bodies, PR comments, review
402
+ replies, issue comments, and commit messages going to a PR.**
403
+
404
+ - Lead with the conclusion. No preamble, no restating the question, no apology,
405
+ no summary of the flow.
406
+ - Say what changed and where. **Link, do not paste.**
407
+ - The reasoning, the evidence, the rejected alternatives and the round history live
408
+ in the flow package — `journal.md`, `context.md`, the review artifacts — which is
409
+ durable, searchable, and costs a reader nothing to skip.
410
+ - A GitHub artifact that needs more than a short paragraph is a signal that the
411
+ detail belongs in the flow with a link out, **not** that the paragraph should
412
+ grow.
413
+ - No orchestrator-written PR comment or reply exceeds two sentences without
414
+ carrying a link to the artifact holding the detail. The reply pass enforces
415
+ this; for anything else you write outward, hold yourself to it.
416
+
417
+ This is deliberately asymmetric: **verbose in the flow, terse on GitHub.** The
418
+ flow is written for whoever resumes the work; GitHub is read by someone who did
419
+ not ask for our reasoning and is reading between other tasks.
420
+
196
421
  ## Review Context Pack
197
422
 
198
423
  Before routing reviewers, build a compact `review_context` object. This is the shared source of truth for all sub-agents and must follow `skills/review-orchestrator/review-context.schema.json`.
@@ -352,34 +577,152 @@ Before anything else, determine whether the request is **diff mode** or **path m
352
577
 
353
578
  See shared script: `skills/shared/git-merge-base.md`
354
579
 
355
- Run the script to determine `BASE_SHA`, then:
580
+ Run the script to determine `BASE_SHA`, then let the pre-filter build the scope:
356
581
 
357
582
  ```bash
358
- git diff --name-only "${BASE_SHA}" # changed files for auto-detection
359
- git diff "${BASE_SHA}" # full diff passed to reviewers
583
+ keryx review scope --ref "${BASE_SHA}" --json > scope.json # KEEP THIS FILE
584
+ keryx review scope --ref "${BASE_SHA}" --scoped-diff # what reviewers get
360
585
  ```
361
586
 
587
+ **Keep `scope.json` until the round is ingested, and pass it as `--scope`.** That
588
+ file is how the drop list reaches the review record. `--append
589
+ "<review-package>/scope.md"` also writes it and is still supported — it now
590
+ REPLACES an existing `## Pre-filter scope` block rather than adding a second, and
591
+ `review ingest` carries any block it finds forward verbatim rather than
592
+ overwriting it — but `--scope scope.json` is the supported path, because it is
593
+ the one that does not depend on running two commands against the same file in the
594
+ right order.
595
+
596
+ **Do not run `git diff` yourself, and do not decide what to leave out.** The
597
+ pre-filter is deterministic code with no model call: it drops generated,
598
+ lockfile, snapshot, vendored and minified paths, drops whitespace-only and
599
+ comment-only change blocks, and bounds every retained change to ±20 lines of
600
+ context (`--context <n>`) instead of the whole file. Dropping a lockfile needs no
601
+ judgement, so it does not get one.
602
+
603
+ The record carries the retained scope **and every drop with its reason**. Both
604
+ halves are required: a scope that shrank without saying so reads afterwards as
605
+ "we reviewed everything". Note that `--scope` takes the WHOLE `--json` document —
606
+ handing over only its `counts` object is refused, because eight integers carry no
607
+ reason for any individual drop.
608
+
609
+ Use `.files` from `scope.json` for the auto-detection table below. A dropped path
610
+ must not select a reviewer, and **neither may a blast-radius path**: scope B is
611
+ under regression check, so a `.tsx` file that only appears there must not pull in
612
+ `review-frontend`. Reviewer selection is driven by the scope-A file list alone.
613
+
362
614
  Scope is limited to **changes introduced in the current branch since merge-base**.
363
615
 
364
616
  ---
365
617
 
618
+ ### Scope B — the blast radius (deep rounds)
619
+
620
+ Everything above is **scope A**: the change, bounded. It answers *is this change
621
+ correct?* It does not answer *did this change break something that was working*,
622
+ and those are different questions — only the first has ever been asked here.
623
+
624
+ A deep round dispatches under **both**. Scope B is computed, never browsed:
625
+
626
+ ```bash
627
+ keryx review blast-radius --ref "${BASE_SHA}" --json > blast-radius.json # KEEP THIS FILE
628
+ keryx review blast-radius --ref "${BASE_SHA}" --brief # what a scope-B reviewer is told
629
+ ```
630
+
631
+ **KEEP THIS FILE** is not advice. The ingest at Step 12 is **refused** if a
632
+ scope-B finding arrives without it — pass it back as
633
+ `review ingest ... --blast-radius blast-radius.json`.
634
+
635
+ It walks `gdgraph affected` outward from every changed file, ranks by edge
636
+ distance, keeps distance ≤ 2, cuts at 40 files closest-first, and adds a changed
637
+ file's naming-related tests when the graph did not already reach them. Requires a
638
+ built graph — run `keryx gdgraph build` if it refuses.
639
+
640
+ **Do not pick the files yourself, and do not widen it.** "Review the
641
+ functionality so nothing breaks" naively means "review the whole repository every
642
+ round", which is unaffordable *and* actively harmful: review quality decays as
643
+ context grows — measured F1 0.65 at round 2 falling to 0.29 at round 10. An
644
+ unbounded scope B makes later rounds worse than earlier ones.
645
+
646
+ The bounds are measured on this repository, not guessed: at depth 2 the set is a
647
+ median of 19 files (p90 65); depth 3 buys eight more in the median and doubles
648
+ the p90. The 40-file cap fires on 25% of commits and removes only hop-2 entries
649
+ on all but 2 of 80, so it almost never costs a direct dependent — and when it
650
+ does, it says so.
651
+
652
+ **Record the whole thing.** `--out "<review-package>/blast-radius.md"` writes the
653
+ set, the depth, and **every file the cap removed**. A truncation nobody can see
654
+ reads afterwards as "we checked everything", which is the claim this pipeline
655
+ exists to stop making. An empty radius is reported as `unresolved`, not as clean:
656
+ the graph indexes code, so a change to a skill, a rule or a schema has no blast
657
+ radius at all and that is a different fact from "nothing depends on it".
658
+
659
+ #### The scope-B question, and what is rejected
660
+
661
+ > Does this change break an existing behaviour **at these sites**?
662
+
663
+ Nothing else. The blast-radius set is **under regression check, not under
664
+ review**. A finding about style, naming or architecture in code the change did
665
+ not touch is refused **by the orchestrator in code** — not discouraged here —
666
+ under three rules, every one of them a fact about the claim rather than about who
667
+ made it:
668
+
669
+ | Rule | Refused because |
670
+ |---|---|
671
+ | `outside-set` | the file is neither in the computed set nor in the changed set; the reviewer went browsing |
672
+ | `non-regression-severity` | below `major`. Under the canonical rubric `minor` states the code behaves correctly and `info` names neither trigger nor outcome; neither can be a claim that something broke |
673
+ | `no-link-to-change` | nothing in the finding names a changed file, module or symbol. A regression claim says THE CHANGE broke this site |
674
+
675
+ Rejections are **recorded, not deleted** — raise the observation under scope A or
676
+ as a separate review. Pass `--brief` output verbatim into the scope-B dispatch:
677
+ the code rejection is the enforcement, but a reviewer told afterwards has already
678
+ spent the round producing findings that will all be refused.
679
+
680
+ `class_scope` on a scope-B finding names the **caller that breaks**, not the
681
+ changed line, because that is the site a human has to look at.
682
+
683
+ #### When it is recomputed
684
+
685
+ | Round | Scope A | Scope B |
686
+ |---|---|---|
687
+ | 1 (first after the draft PR) | yes | yes |
688
+ | 2..N | yes | recomputed only if the changed-file set moved |
689
+ | final | yes | **yes, always** |
690
+
691
+ Do not decide this by memory:
692
+
693
+ ```bash
694
+ keryx review blast-radius --ref "${BASE_SHA}" --previous blast-radius.json [--final]
695
+ ```
696
+
697
+ It prints the decision and the reason, and reuses the previous record when
698
+ nothing moved. The final round recomputes whatever the file set did — otherwise a
699
+ fix introduced in round 3 gets no regression check at all, and the round that
700
+ certifies the flow is the one that checked the least.
701
+
702
+ ---
703
+
366
704
  ### Path Mode
367
705
 
368
- When a path or target is named, collect the files to review:
706
+ When a path or target is named, collect the candidate files:
369
707
 
370
708
  ```bash
371
709
  # If a directory path is given:
372
710
  find <path> -type f \( -name "*.ts" -o -name "*.tsx" -o -name "*.js" -o -name "*.jsx" \) | sort
373
711
 
374
- # If a file path is given:
375
- cat <file>
376
-
377
712
  # If a module name is given (e.g. "UserStore", "pipelines module"):
378
713
  find . -type f -name "*<name>*" \( -name "*.ts" -o -name "*.tsx" \)
379
714
  # Also check common locations: src/stores/, src/modules/, src/components/
380
715
  ```
381
716
 
382
- Pass the full **file contents** (not a diff) to sub-reviewers. Set `SCOPE_MODE: path`.
717
+ Then put the list through the same exclusions before reading anything:
718
+
719
+ ```bash
720
+ keryx review scope --path "src/a.ts,src/b.ts" --json > scope.json
721
+ ```
722
+
723
+ Pass the full **file contents** of the paths it **retained** to sub-reviewers, and
724
+ read none of the ones it dropped. Set `SCOPE_MODE: path`. Path mode has no hunks
725
+ and therefore no context window; the drop list is recorded exactly the same way.
383
726
 
384
727
  **Reviewer behavior in path mode:** reviewers check the entire file content — not just added lines. All findings apply to the current state of the code, not only to changes.
385
728
 
@@ -414,6 +757,33 @@ If the repository has local convention docs such as `CLAUDE.md`, `AGENTS.md`,
414
757
  These convention reviewers are additive: keep the generic reviewers selected by normal detection,
415
758
  then add the matching convention pass. Deduplicate reviewer names before dispatch.
416
759
 
760
+ ### Stack scoping — run it after detection, before dispatch
761
+
762
+ The tables above select reviewers by **file shape**. A `.ts` file looks the same
763
+ whether or not the repository has React in it, so those tables will happily
764
+ dispatch a React/MobX conventions reviewer at a Bun CLI with no frontend — which
765
+ is exactly what happened here, on every review, for months.
766
+
767
+ So the selected set is filtered once more, by what the repository actually
768
+ declares:
769
+
770
+ ```bash
771
+ keryx review stack --json
772
+ ```
773
+
774
+ It reads `package.json` and reports, per reviewer, `include` or `exclude` with a
775
+ reason. A reviewer carrying `metadata.stack_requires` is dispatched when **any**
776
+ tag it names is present — matching what `keryx review stack` actually computes,
777
+ and failing toward inclusion rather than away from it.
778
+
779
+ **Its failure mode is to include, never to skip.** A missing, unparsable or
780
+ unexpected manifest sets `uncertain`, and an uncertain detection marks every tag
781
+ present, so every reviewer runs. A reviewer that runs needlessly costs tokens; a
782
+ reviewer wrongly skipped hides a real defect, and that asymmetry is not close.
783
+
784
+ Record the exclusions with their reasons alongside the pre-filter drops. A
785
+ reviewer silently absent from a report reads as "it had nothing to say".
786
+
417
787
  ### Convention Reviewer Confirmation
418
788
 
419
789
  When convention reviewers are auto-detected and the user did not explicitly pass
@@ -498,7 +868,7 @@ Skipped reviewers:
498
868
  | `--core-boundaries` | `review-core-boundaries` |
499
869
  | `--flow-graph` | `review-flow-graph` |
500
870
  | `--all` | all reviewers above (including `review-clean-code`, `review-highload`, applicable legacy/profile reviewers, and project convention reviewers when local convention docs exist) |
501
- | `--strict` | runs AFTER all others; adds a strict commentary pass on consolidated findings |
871
+ | `--verify` | `review-verifier`, AFTER all others; checks the consolidated findings by running something. Delete-only. |
502
872
  | (auto) | detected from diff file extensions — see Auto-detection table |
503
873
 
504
874
  Multiple flags may be combined. Example: `review --backend --security` dispatches
@@ -526,7 +896,85 @@ Dispatch selected reviewers in parallel when independent. Use waves when token b
526
896
 
527
897
  1. Wave A - core correctness/risk reviewers: logic, architecture, security/highload when selected.
528
898
  2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files.
529
- 3. Wave C - synthesis: strict pass when blockers/majors exist, `--strict` is set, or PR is high-risk.
899
+ 3. Wave C - **verification**: `review-verifier` over the consolidated findings, when blockers/majors
900
+ exist, `--verify` is set, or the PR is high-risk. See below.
901
+
902
+ ### Wave C — verification, and what it replaced
903
+
904
+ Wave C used to run `review-strict`: a meta-pass that re-read the consolidated
905
+ findings and **adjusted their severity with no new evidence**, under an elevation
906
+ table biased 3:1 toward escalation. It was **removed, not improved**, and the
907
+ reason is measured rather than stylistic:
908
+
909
+ - **GPT-4 on GSM8K across self-correction rounds: 95.5 → 91.5 → 89.0.**
910
+ **GPT-3.5 on CommonSenseQA: 75.8 → 38.1.** Among the answers that changed,
911
+ correct → incorrect exceeded incorrect → correct (Huang et al., *Large Language
912
+ Models Cannot Self-Correct Reasoning Yet*, ICLR 2024, arXiv:2310.01798).
913
+ - **Self-Refine (arXiv:2303.17651): +49.2 on dialogue response generation, +0.2
914
+ on maths.** Self-refinement gains are on subjective tasks and vanish on
915
+ verifiable reasoning. Judging whether a null-guard is missing is verifiable
916
+ reasoning.
917
+
918
+ Re-scoring a finding by re-reading it is therefore not a rigour pass; it is a
919
+ coin flip weighted toward more findings. **Do not restore it because it looks
920
+ obviously useful — it looked obviously useful the first time.**
921
+
922
+ `review-verifier` occupies the slot and differs in exactly one way that matters:
923
+ **it runs something.** Verification that executes rejects 85–96% of false reports
924
+ against 4–15% unaided while finding 30–44% more true bugs (AnyPoC,
925
+ arXiv:2604.11950); Meta's TestGen-LLM funnel discards 75% of its own output
926
+ (75% build → 57% build and pass → 25% improve coverage) and the surviving quarter
927
+ reaches 73% human acceptance (arXiv:2402.09171).
928
+
929
+ It also **never votes.** 80+ agents unanimously endorsed a padding-oracle
930
+ vulnerability that did not exist, and a single empirical test killed it: consensus
931
+ cannot detect a hallucination its members share, so agreement between reviewers is
932
+ not evidence and must never be recorded as verification.
933
+
934
+ That rule is about agreement *standing in for* evidence. It is not a rule that
935
+ two verifiers may not both check the same finding: each claim is admitted on its
936
+ own — a named non-author, a real method, real evidence, with `reasoning` already
937
+ capped — and when two such claims reach the **same** verdict the merge records it,
938
+ naming both verifiers and carrying both pieces of evidence. Claims that
939
+ **disagree** still cancel, because there the only thing deciding the outcome
940
+ would be claim order.
941
+
942
+ Dispatch rules:
943
+
944
+ - Pass the consolidated findings, each carrying `global_id` and the **real**
945
+ originating `reviewer`. A finding whose `reviewer` is the orchestrator cannot be
946
+ routed away from its author, so the never-self-verify rule silently stops
947
+ applying — that field was hardcoded to `review-orchestrator` on all 83 recorded
948
+ findings and is fixed only from 0.2.70 onward.
949
+ - **A finding is never verified by the reviewer that raised it.** When only one
950
+ reviewer ran, its findings are simply left unverified; verifying them yourself
951
+ is worse than not verifying them. The merge compares the two names after
952
+ normalising case, surrounding whitespace, `_`/`-`, and a trailing `(model)`
953
+ annotation, so `review-logic `, `Review-Logic` and `review-logic (sonnet)` are
954
+ all the same actor. Do not try to route around it by respelling the name — the
955
+ comparison deliberately over-matches, because a refused claim only ever costs a
956
+ verdict while a missed self-verification costs the finding.
957
+ - The verifier returns `verification-claim.schema.json`. Merge it with
958
+ `keryx review ingest --verifications <file>`; do not apply verdicts by hand.
959
+ - **The verifier can only delete.** If it returns a severity, a new finding, or a
960
+ rewritten finding, the merge discards that whole claim and records the attempt.
961
+ Do not "help" by applying it.
962
+
963
+ ### `verification_mode`
964
+
965
+ `off` | `annotate` | `filter`. **Default `annotate`, and it stays `annotate` for
966
+ one release.**
967
+
968
+ | Mode | What happens |
969
+ |---|---|
970
+ | `off` | No verification. Claims are refused rather than silently ignored. |
971
+ | `annotate` | Verdicts are recorded on the findings. **Nothing is removed.** A `refuted` finding is still reported, marked refuted. |
972
+ | `filter` | An applied `refuted` verdict removes the finding from the reported set and records it as `dismissed-incorrect`, with the verification evidence. |
973
+
974
+ `annotate` is the default so the drop rate is a **measured number** before it
975
+ costs a real finding. The risk is named rather than assumed away: SWE-agent keeps
976
+ its equivalent step opt-in because it sometimes rejects correct patches. Do not
977
+ switch a project to `filter` on the strength of one round.
530
978
 
531
979
  ### Agent Runtime Compatibility
532
980
 
@@ -579,6 +1027,7 @@ Each reviewer must return a `REVIEW_RESULT` object matching `skills/review-orche
579
1027
  | Security vulnerabilities | NO | `review-security-code` |
580
1028
  | Performance anti-patterns | NO | `review-performance` |
581
1029
  | Style / naming / import order | NO | `review-style` |
1030
+ | Checking whether a reported finding is real | NO | `review-verifier` |
582
1031
  | Clean Code principles + SOLID at code level | NO | `review-clean-code` |
583
1032
  | Concurrency, resource pools, caching, queues, idempotency | NO | `review-highload` |
584
1033
  | Frontend repository conventions | NO | `review-frontend-conventions` |
@@ -602,6 +1051,93 @@ Before consolidation, validate every reviewer result:
602
1051
 
603
1052
  ---
604
1053
 
1054
+ ## Severity (canonical)
1055
+
1056
+ **This is the only severity rubric in the review domain.** Reviewers do not carry
1057
+ their own. Ten private rubrics feeding one sort produce a ranking that means ten
1058
+ different things at once, and ranking is what an operator uses to decide what to
1059
+ read first. A reviewer may state which of *its* conditions land where; it may not
1060
+ redefine the levels.
1061
+
1062
+ ### `blocker` — merge-blocking, and nothing else
1063
+
1064
+ Exactly four shapes. Nothing outside this list is a `blocker`, however strongly
1065
+ the reviewer feels about it:
1066
+
1067
+ 1. **A crash** — the process, request, or render dies on an input the change
1068
+ admits.
1069
+ 2. **Data loss or corruption** — something persisted, transmitted, or returned is
1070
+ destroyed or silently wrong.
1071
+ 3. **An exploitable vulnerability** — an attacker action with a named entry point
1072
+ and a named impact.
1073
+ 4. **An unimplemented acceptance criterion** — the change claims work the diff
1074
+ does not contain.
1075
+
1076
+ Everything else is at most `major`. "This will definitely cause problems later"
1077
+ is not one of the four. Neither is "this violates the architecture", "this fails
1078
+ the linter", or "this is how the last outage started".
1079
+
1080
+ ### `major` / `minor` / `info` — the boundary test
1081
+
1082
+ Ask one question, and ask it of the **finding**, not of the code:
1083
+
1084
+ > **Does it name a trigger, and the observable outcome that trigger produces?**
1085
+
1086
+ - **`major`** — it does. There is an input, a call, a render, or a load level, and
1087
+ a resulting behaviour a user or a caller would call wrong: a wrong value, a lost
1088
+ update, a leak, a hang, a cost stated together with the frequency that makes it
1089
+ a cost. Not one of the four shapes above, so not merge-blocking — but the code
1090
+ does the wrong thing.
1091
+ - **`minor`** — it does not, and does not claim to. The code behaves correctly;
1092
+ the cost lands on whoever reads or edits it next, and the finding names that
1093
+ cost at a named site.
1094
+ - **`info`** — it names neither. An observation, a preference, or a risk with no
1095
+ demonstrated path.
1096
+
1097
+ The test is procedural on purpose. It is applied by reading the finding, so
1098
+ someone who did not write it — and has not read the code — reaches the same
1099
+ answer: look for the trigger and the outcome. Present → `major`. Absent, but a
1100
+ concrete maintenance cost is named → `minor`. Neither → `info`.
1101
+
1102
+ Two consequences, both previously decided differently in different files:
1103
+
1104
+ - A finding that **claims** runtime harm and cannot name the trigger is `info`,
1105
+ not `major`. It is not demoted to `minor`: `minor` is for findings that never
1106
+ claimed runtime harm at all. The two are different failures and stay
1107
+ distinguishable.
1108
+ - Severity is a property of the demonstrated outcome, never of the reviewer that
1109
+ found it. A security reviewer's unproven concern is `info` under the same test
1110
+ that puts a style reviewer's unproven concern there.
1111
+ - **And never of how crisply the finding is worded.** An outcome that costs a
1112
+ user, a caller or persisted state nothing is `minor` however precisely its
1113
+ trigger is named. Without this clause the test above rates prose quality: a
1114
+ cosmetic wording nit stated as "trigger X produces output Y" reads as `major`,
1115
+ while a real defect stated tersely reads as `info`. That is not academic — the
1116
+ findings cap truncates by severity, so the well-written typo would survive and
1117
+ the terse real defect would be cut.
1118
+
1119
+ The boundary this rubric does **not** draw is `major` against `major`. Two
1120
+ findings that both name a trigger and an outcome are the same severity even when
1121
+ one is obviously worse; the ordering inside a severity is the operator's, and
1122
+ inventing a fifth level to express it would put us back where we started.
1123
+
1124
+ ### Shared laws (every reviewer)
1125
+
1126
+ 1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
1127
+ name the input, call, or condition that reaches the code, you have an
1128
+ observation, not a finding. Report it as `info` and say what would settle it.
1129
+ 2. **Never flag the theoretical.** The path you describe must exist in the code
1130
+ under review. Do not report a safe API because it could be misused, or a
1131
+ pattern because it is often wrong elsewhere.
1132
+ 3. **One finding per class, not one per occurrence.** When the same shape appears
1133
+ at several sites, report it once and list every site. Ten findings that are one
1134
+ finding hide the other nine problems.
1135
+
1136
+ `review-security-code` carries a fourth — every security finding states its attack
1137
+ vector — which does not generalise and stays there.
1138
+
1139
+ ---
1140
+
605
1141
  ## Finding Format
606
1142
 
607
1143
  ### Class scope — required for `blocker` and `major`
@@ -702,6 +1238,18 @@ STATUS: DONE | DONE_WITH_CONCERNS
702
1238
  - minor: N
703
1239
  - info: N
704
1240
 
1241
+ ## Stage counts
1242
+ <!-- Required. State what each stage REMOVED, and never state it as a precision
1243
+ improvement: no precision baseline exists to improve on. The one measured
1244
+ from the review packages on disk was 53/53 = 100% — pinned there by
1245
+ construction, because nothing in that corpus could record a finding as
1246
+ wrong. Copy these from `scope.md`; do not re-count by hand. -->
1247
+ - dropped by pre-filter: <files>, <blocks>, <changed lines> (or `not recorded` if no scope was built)
1248
+ - verification mode: `<off | annotate | filter>`
1249
+ - verdicts: confirmed N, refuted N, unverifiable N, unverified N
1250
+ - refuted by the verifier: N (removed: N — always 0 outside `filter`)
1251
+ - retained: N
1252
+
705
1253
  ## Blockers (must fix before merge)
706
1254
  <[F-NNN] findings with severity=blocker, sorted by file>
707
1255
 
@@ -770,13 +1318,22 @@ Publish this review report to the PR?
770
1318
 
771
1319
  ### Concise PR Comment
772
1320
 
773
- The visible PR comment is for humans. It must be written in English only and stay concise.
1321
+ The visible PR comment is for humans. It must be written in English only and stay
1322
+ concise, under the brevity rule above: **the summary is at most two sentences and
1323
+ carries a link to the artifact holding the detail.**
1324
+
1325
+ The finding rows below are a bounded exception, not a licence: they exist because
1326
+ a reviewer scanning a PR needs the blockers in front of them. Keep them to the
1327
+ `blocker` and `major` rows; everything at `minor` or below goes behind the
1328
+ `<details>` fold or, better, into the AI artifact and is linked. The full findings
1329
+ set, the round history and the reasoning belong in the flow package — pasting them
1330
+ here is the failure this rule names.
774
1331
 
775
1332
  ```markdown
776
1333
  ## AI Review Report
777
1334
 
778
1335
  **Verdict:** REQUEST_CHANGES
779
- **Summary:** 2-3 concise sentences with overall risk and the main merge blocker.
1336
+ **Summary:** At most two sentences: the overall risk and the main merge blocker. Detail: <link to the AI artifact or the flow package>.
780
1337
 
781
1338
  | Severity | Area | Finding | Suggested Fix | Owner |
782
1339
  |---|---|---|---|---|
@@ -948,6 +1505,18 @@ If absent, proceed normally — context is optional and non-blocking.
948
1505
  | "Spec compliance can wait until after quality review" | Stage 1 gate exists because unimplemented requirements invalidate quality work |
949
1506
  | "I'll deduplicate findings manually in my head" | Always normalize to [F-NNN] format before consolidation to avoid losing findings |
950
1507
  | "Minor findings from one reviewer cancel out the major from another" | Each finding stands independently; severity is per-finding, not averaged |
1508
+ | "The reviewer's own severity table said blocker" | There are no reviewer tables. One rubric, in **Severity (canonical)** above; a reviewer that ships one is the defect this replaced |
1509
+ | "It's a security/architecture finding, so it's a blocker" | Severity is the demonstrated outcome, not the domain that found it. `blocker` is exactly the four shapes |
1510
+ | "It will definitely break something eventually, so blocker" | Name the trigger and the outcome. Named → `major`. Unnamed → `info`. "Eventually" is neither |
1511
+ | "A strict re-read of the findings will sharpen them" | That pass existed and was removed: self-correction without new evidence measured 95.5 → 91.5 → 89.0 on GSM8K and 75.8 → 38.1 on CommonSenseQA. Run something instead |
1512
+ | "Three reviewers agree, so the finding is verified" | Consensus is not evidence. 80+ agents unanimously endorsed a vulnerability that did not exist; one empirical test killed it |
1513
+ | "The verifier suggested a higher severity, I'll apply it" | It cannot suggest one. A claim carrying a severity is discarded whole and the attempt is recorded |
1514
+ | "This finding has no `verification`, so it can be dropped" | Absent means nobody checked. All 83 recorded findings are in that state; none of them is thereby wrong |
1515
+ | "Precision went up after the verifier landed" | There is no precision baseline to have gone up from. State stage counts: dropped, refuted, retained |
1516
+ | "I'll widen the blast radius, this change looks risky" | It is bounded because review quality decays with context: F1 0.65 at round 2 → 0.29 at round 10. Widening makes the later rounds worse, not safer |
1517
+ | "The blast radius came back empty, so nothing can break" | Empty and unresolved are different facts. The graph indexes code — a Markdown or JSON change has no radius at all, and the record says which one you got |
1518
+ | "The changed files are the same as last round, so scope B can be skipped on the final round" | The final round always recomputes. A fix landed in round 3 is the change; skipping means the certifying round checked the least |
1519
+ | "This scope-B file has an obvious naming problem, I'll report it" | Rejected in code: a naming problem is `minor` at best, and the floor is `major`. The set is under regression check, not under review — raise it under scope A |
951
1520
  | "No flags means no reviewers" | No flags → run auto-detection; never produce an empty review |
952
1521
  | "User named a module so I'll use diff mode" | Named module/component/store → path mode; diff mode is only for branch changes |
953
1522
  | "Path mode should only show lines I'd flag in diff mode" | Path mode reviews the entire file — all findings apply, not just added lines |