@mrciphersmith/keryx 0.2.70 → 0.2.72
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/cli.js +25237 -17483
- package/docs/README.md +54 -0
- package/docs/requirements/shared-agent-context/README.md +104 -0
- package/package.json +3 -2
- package/src/gdskills/bundled/rules/core/code-review-learned-profile.mdc +81 -0
- package/src/gdskills/bundled/rules/core/jobs-documentation.mdc +1 -1
- package/src/gdskills/bundled/rules/core/model-selection.mdc +184 -31
- package/src/gdskills/bundled/rules/core/review-strict-profile.mdc +8 -4
- package/src/gdskills/bundled/rules/core/skills-storage-workflow.mdc +36 -0
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/code-verifier/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.codex.md +3 -3
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.cursor.md +3 -3
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.md +3 -3
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.opencode.md +3 -3
- package/src/gdskills/bundled/skills/orchestration/context-collector/SKILL.zed.md +3 -3
- package/src/gdskills/bundled/skills/orchestration/context-collector/orchestrator-prompt.md +2 -2
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.codex.md +4 -4
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.cursor.md +4 -4
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.md +4 -4
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.opencode.md +4 -4
- package/src/gdskills/bundled/skills/orchestration/feature-analyzer/SKILL.zed.md +4 -4
- package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.codex.md +2 -2
- package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.cursor.md +2 -2
- package/src/gdskills/bundled/skills/orchestration/feature-dev/SKILL.md +3 -3
- package/src/gdskills/bundled/skills/orchestration/flow-orchestrator/SKILL.md +80 -21
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/issue-analyzer/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.codex.md +2 -2
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.cursor.md +2 -2
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.md +3 -2
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.opencode.md +2 -2
- package/src/gdskills/bundled/skills/orchestration/job-documenter/SKILL.zed.md +2 -2
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.codex.md +997 -509
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.cursor.md +997 -509
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +968 -513
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.opencode.md +997 -509
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.zed.md +997 -509
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/input-contract.schema.json +38 -28
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/orchestrator-prompt.md +98 -66
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/output-contract.schema.json +27 -5
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.codex.md +20 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.cursor.md +20 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.md +21 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.opencode.md +20 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/SKILL.zed.md +20 -2
- package/src/gdskills/bundled/skills/orchestration/task-implementer/input-contract.schema.json +1 -1
- package/src/gdskills/bundled/skills/planning/autodoc-analyst/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-architect/SKILL.md +3 -1
- package/src/gdskills/bundled/skills/planning/autodoc-assembler/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-orchestrator/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-scanner/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/autodoc-writer/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/consistency-checker/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/docpack-orchestrator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/docpack-review/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/interview/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/interview/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/interview/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/patterns-researcher/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/planner/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/planning/prd-creator/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/problem-definer/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/project-discovery/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/spec-writer/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/planning/stack-advisor/SKILL.md +2 -1
- package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/platform/claude-md-management/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/platform/hookify/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/platform/hookify/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/platform/hookify/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/changelog/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/changelog/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/changelog/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/commit/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/commit/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/commit/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/db-migrate/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/dependency-update/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/metaproject-security/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/perf-check/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/perf-check/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/perf-check/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/quality/pr-issue-documenter/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/quality/push/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/push/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/push/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/security-audit/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/security-audit/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/security-audit/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/test-gen/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/test-gen/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/test-gen/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/quality/tests-creator/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.md +3 -3
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-ai-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-learned-review/SKILL.codex.md +252 -0
- package/src/gdskills/bundled/skills/review/code-learned-review/SKILL.cursor.md +252 -0
- package/src/gdskills/bundled/skills/review/code-learned-review/SKILL.md +243 -0
- package/src/gdskills/bundled/skills/review/code-learned-review/SKILL.opencode.md +252 -0
- package/src/gdskills/bundled/skills/review/code-learned-review/SKILL.zed.md +252 -0
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.md +3 -2
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-mobx-store-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.codex.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.cursor.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.md +2 -2
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.opencode.md +1 -1
- package/src/gdskills/bundled/skills/review/code-style-review/SKILL.zed.md +1 -1
- package/src/gdskills/bundled/skills/review/review-architecture/SKILL.md +38 -11
- package/src/gdskills/bundled/skills/review/review-backend/SKILL.md +49 -15
- package/src/gdskills/bundled/skills/review/review-clean-code/SKILL.md +50 -13
- package/src/gdskills/bundled/skills/review/review-core-boundaries/SKILL.md +35 -3
- package/src/gdskills/bundled/skills/review/review-flow-graph/SKILL.md +34 -3
- package/src/gdskills/bundled/skills/review/review-frontend/SKILL.md +71 -30
- package/src/gdskills/bundled/skills/review/review-frontend-conventions/SKILL.md +35 -4
- package/src/gdskills/bundled/skills/review/review-highload/SKILL.md +50 -16
- package/src/gdskills/bundled/skills/review/review-logic/SKILL.md +41 -13
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +599 -30
- package/src/gdskills/bundled/skills/review/review-orchestrator/reviewer-finding.schema.json +7 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/verification-claim.schema.json +78 -0
- package/src/gdskills/bundled/skills/review/review-performance/SKILL.md +44 -14
- package/src/gdskills/bundled/skills/review/review-pr-feedback/SKILL.md +44 -19
- package/src/gdskills/bundled/skills/review/review-regression/SKILL.md +185 -0
- package/src/gdskills/bundled/skills/review/review-security-code/SKILL.md +45 -14
- package/src/gdskills/bundled/skills/review/review-style/SKILL.md +27 -7
- package/src/gdskills/bundled/skills/review/review-testing-practices/SKILL.md +36 -4
- package/src/gdskills/bundled/skills/review/review-verifier/SKILL.md +276 -0
- package/src/gdskills/bundled/skills/shared/git-merge-base.md +1 -1
- package/src/gdskills/contracts/review-finding.schema.json +119 -1
- package/src/gdskills/contracts/subagent-dispatch.schema.json +59 -3
- package/src/gdskills/bundled/rules/core/code-review-b091-profile.mdc +0 -48
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.codex.md +0 -209
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.cursor.md +0 -209
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.md +0 -208
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.opencode.md +0 -209
- package/src/gdskills/bundled/skills/review/code-b091-review/SKILL.zed.md +0 -209
- package/src/gdskills/bundled/skills/review/review-strict/SKILL.md +0 -328
|
@@ -4,10 +4,9 @@ description: |
|
|
|
4
4
|
Use when: a code review is requested and the user does not explicitly name a specialized reviewer.
|
|
5
5
|
Handles "review", "code review", "review PR", "review --frontend", "review --backend",
|
|
6
6
|
"review --architecture", "review --security", "review --performance", "review --style",
|
|
7
|
-
"review --
|
|
7
|
+
"review --verify", "review --project-conventions", "review --legacy-profiles", "review --all". Routes to specialized reviewers in parallel and
|
|
8
8
|
consolidates findings into one unified report.
|
|
9
9
|
NOT for: running a single specialized reviewer — invoke it directly by name instead.
|
|
10
|
-
version: "1.6.0"
|
|
11
10
|
triggers:
|
|
12
11
|
- "review"
|
|
13
12
|
- "code review"
|
|
@@ -18,7 +17,7 @@ triggers:
|
|
|
18
17
|
- "review --security"
|
|
19
18
|
- "review --performance"
|
|
20
19
|
- "review --style"
|
|
21
|
-
- "review --
|
|
20
|
+
- "review --verify"
|
|
22
21
|
- "review --all"
|
|
23
22
|
- "review --clean-code"
|
|
24
23
|
- "review --highload"
|
|
@@ -29,15 +28,15 @@ triggers:
|
|
|
29
28
|
- "review --flow-graph"
|
|
30
29
|
- "review --legacy-profiles"
|
|
31
30
|
- "review --code-ai"
|
|
32
|
-
- "review --
|
|
31
|
+
- "review --learned"
|
|
33
32
|
- "review --code-style"
|
|
34
33
|
- "review --mobx-store"
|
|
35
34
|
metadata:
|
|
36
35
|
author: "MrCipherSmith"
|
|
37
|
-
version: "1.
|
|
36
|
+
version: "1.8.0"
|
|
38
37
|
category: "review"
|
|
38
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
39
39
|
license: "MIT"
|
|
40
|
-
compatibility: "cursor,codex,zed,opencode,claude"
|
|
41
40
|
---
|
|
42
41
|
|
|
43
42
|
# Review Orchestrator
|
|
@@ -52,20 +51,47 @@ unified report sorted by severity. It does not perform any review logic itself.
|
|
|
52
51
|
|
|
53
52
|
```
|
|
54
53
|
Review Orchestrator Progress:
|
|
54
|
+
- [ ] Step 0: On a PR target, collect external comments — `keryx review comments collect`
|
|
55
55
|
- [ ] Step 1: Build Review Context Pack (PR metadata, scope, rules, context_doc summary)
|
|
56
56
|
- [ ] Step 2: Detect review mode (diff mode vs. path mode)
|
|
57
|
-
- [ ] Step 3:
|
|
57
|
+
- [ ] Step 3: Build the bounded scope with `keryx review scope` — never by hand
|
|
58
|
+
- [ ] Step 3b: On a deep round, compute scope B with `keryx review blast-radius` — never by browsing — and KEEP the `--json` file; `review ingest --blast-radius <file>` is refused without it
|
|
58
59
|
- [ ] Step 4: Parse flags / auto-detect domain from scope
|
|
59
60
|
- [ ] Step 5: Ask user to confirm optional convention reviewers (legacy/profile reviewers are flag-only, never prompted)
|
|
60
|
-
- [ ] Step 6: Plan sub-agent dispatch
|
|
61
|
+
- [ ] Step 6: Plan sub-agent dispatch and token budgets, and compute each dispatch's model with `keryx review tier` — never by hand
|
|
61
62
|
- [ ] Step 7: Stage 1 gate - spec compliance check (if issue/task provided)
|
|
62
63
|
- [ ] Step 8: Dispatch selected reviewers in PARALLEL with reviewer-input schema
|
|
63
64
|
- [ ] Step 9: Collect reviewer-finding schema results and handle NEEDS_CONTEXT
|
|
64
|
-
- [ ] Step 10:
|
|
65
|
+
- [ ] Step 10: Wave C — dispatch `review-verifier` over the consolidated findings
|
|
65
66
|
- [ ] Step 11: Sort by severity, deduplicate, emit unified report
|
|
66
67
|
- [ ] Step 12: Emit the machine-readable `keryx:findings` block alongside the report
|
|
68
|
+
- [ ] Step 13: Report the stage counts: dropped by pre-filter, refuted by the verifier, retained
|
|
69
|
+
- [ ] Step 14: AFTER THE FINAL ROUND ONLY — answer every external comment once, `keryx review comments reply --final`
|
|
67
70
|
```
|
|
68
71
|
|
|
72
|
+
Step 0 runs on **every** round. Step 14 runs **once**, after the last one. They are
|
|
73
|
+
two commands for that reason: a caller that already runs collection per round
|
|
74
|
+
would carry the posting along with it, and the reviewer would get six replies to
|
|
75
|
+
one comment.
|
|
76
|
+
|
|
77
|
+
### Step 6 — the model is computed, not chosen
|
|
78
|
+
|
|
79
|
+
Before dispatching each reviewer, run `keryx review tier` with the signals you
|
|
80
|
+
already hold (`--scope`, `--findings`, `--diff-lines`, `--fix-attempt`,
|
|
81
|
+
`--verifier`, `--security`, `--forced-strategy-change`) and paste the `model`
|
|
82
|
+
block it prints into that dispatch.
|
|
83
|
+
|
|
84
|
+
Do NOT assign the tier by reading the table in `rules/core/model-selection.mdc`.
|
|
85
|
+
Working it out in your head is exactly the mechanical step that rule moves into
|
|
86
|
+
code — and it is the step that was documented as running for a whole release
|
|
87
|
+
while nothing called it.
|
|
88
|
+
|
|
89
|
+
The command names no model. It ranks whatever your provider reports at runtime
|
|
90
|
+
and places the tiers relative to your own model; when it cannot rank anything it
|
|
91
|
+
prints `inherit: true` and exits 0, which means the dispatch runs on the session
|
|
92
|
+
model. That is a correct answer, not a failure — never "fix" it by writing a
|
|
93
|
+
model id into a dispatch.
|
|
94
|
+
|
|
69
95
|
---
|
|
70
96
|
|
|
71
97
|
## Step 12 — the `keryx:findings` block
|
|
@@ -127,7 +153,7 @@ which of its inputs were recovered and which were never written down.
|
|
|
127
153
|
|
|
128
154
|
| Field | Type | Required | Description |
|
|
129
155
|
|-------|------|----------|-------------|
|
|
130
|
-
| `flags` | string[] | no | One or more of: `--frontend`, `--backend`, `--architecture`, `--security`, `--performance`, `--style`, `--clean-code`, `--highload`, `--project-conventions`, `--frontend-conventions`, `--testing-practices`, `--core-boundaries`, `--flow-graph`, `--legacy-profiles`, `--code-ai`, `--
|
|
156
|
+
| `flags` | string[] | no | One or more of: `--frontend`, `--backend`, `--architecture`, `--security`, `--performance`, `--style`, `--clean-code`, `--highload`, `--project-conventions`, `--frontend-conventions`, `--testing-practices`, `--core-boundaries`, `--flow-graph`, `--legacy-profiles`, `--code-ai`, `--learned`, `--code-style`, `--mobx-store`, `--verify`, `--all` |
|
|
131
157
|
| `path` | string | no | File or directory path to review (e.g., `src/stores/`, `src/components/UserCard.tsx`). Activates **path mode** — reviews the files at this path directly, not a git diff. |
|
|
132
158
|
| `commit_range` | string | no | Explicit commit hash or range (e.g., `abc123..HEAD`). Overrides merge-base detection. Ignored in path mode. |
|
|
133
159
|
| `issue_url` | string | no | GitHub issue or task URL. If provided, Stage 1 gate checks spec compliance before dispatching reviewers. |
|
|
@@ -136,6 +162,8 @@ which of its inputs were recovered and which were never written down.
|
|
|
136
162
|
| `token_budget` | object | no | Optional budget controls: `{total, per_reviewer, diff_max_chars, file_max_chars}`. |
|
|
137
163
|
| `model_strategy` | string | no | `current`, `ask`, or `adaptive`. Default: `current`; do not switch models unless user or automation allows it. |
|
|
138
164
|
| `managed_review` | object | no | Optional managed review mode: `{mode, target, target_ref, flow_id, reviewers}` where mode is `lightweight`, `attach-review`, `review-flow`, or `ingest`. |
|
|
165
|
+
| `verification_mode` | string | no | `off`, `annotate`, or `filter`. Default `annotate` — verdicts are recorded and nothing is removed. See Wave C. |
|
|
166
|
+
| `pr_comments` | object | no | `{enabled, max_replies_total, max_sentences_per_reply}`. Defaults: enabled when a PR exists, `30`, `2`. Collect every round, reply once at the end. See External PR comments. |
|
|
139
167
|
|
|
140
168
|
---
|
|
141
169
|
|
|
@@ -166,10 +194,66 @@ Runtime CLI surface:
|
|
|
166
194
|
keryx review attach --flow <id> --target <kind> --ref <ref>
|
|
167
195
|
keryx review start --target <kind> --ref <ref>
|
|
168
196
|
keryx review ingest --report <path> [--flow <id>] --ref <ref>
|
|
197
|
+
[--verifications <file>] [--verification-mode off|annotate|filter]
|
|
198
|
+
[--scope <scope.json>] [--blast-radius <blast-radius.json>]
|
|
199
|
+
[--refuted <file>]
|
|
200
|
+
# --blast-radius is REQUIRED whenever this round dispatched a
|
|
201
|
+
# scope-B reviewer (review-regression). See below.
|
|
169
202
|
keryx review status <review-id-or-path>
|
|
170
203
|
keryx review complete <review-id-or-path>
|
|
204
|
+
[--finding <id> --disposition <state> --evidence <ref>]...
|
|
171
205
|
```
|
|
172
206
|
|
|
207
|
+
**An unrecognised option is refused, not ignored.** A misspelling used to be
|
|
208
|
+
accepted with exit 0, so `review complete --disposition ...` printed
|
|
209
|
+
`status: closed` and wrote nothing at all.
|
|
210
|
+
|
|
211
|
+
`--verifications` takes what `review-verifier` returned. `--scope` takes the
|
|
212
|
+
whole `--json` output of `keryx review scope`, so the package records what the
|
|
213
|
+
pre-filter dropped, **with a reason per drop**, as well as what the verifier
|
|
214
|
+
refuted. Without it — and with no `## Pre-filter scope` block already in the
|
|
215
|
+
package — the record says **`not recorded`** for that stage rather than `0`;
|
|
216
|
+
"dropped nothing" and "never ran" are different facts.
|
|
217
|
+
|
|
218
|
+
`--blast-radius` is **not optional on any round that dispatched
|
|
219
|
+
`review-regression`** — which is every `recommended` and `full` round, because
|
|
220
|
+
`review-regression` is in both profiles. AC3 is enforced in code: an ingest
|
|
221
|
+
carrying a scope-B finding with no blast-radius record is **refused**, and the
|
|
222
|
+
round is not recordable until the set is supplied. Pass the `--json` output of
|
|
223
|
+
`keryx review blast-radius` — the same record you computed at dispatch time to
|
|
224
|
+
bound the round.
|
|
225
|
+
|
|
226
|
+
Two channels, and both work:
|
|
227
|
+
|
|
228
|
+
- `--blast-radius <file>` — always correct, and the one to use.
|
|
229
|
+
- `.metaproject/reviews/blast-radius.json` (or
|
|
230
|
+
`.metaproject/flows/<flow>/reviews/blast-radius.json` for a flow-attached
|
|
231
|
+
round) — the handoff slot, read when no flag is passed. The ingest **consumes**
|
|
232
|
+
it: it is copied into the package it screened and removed from the slot, so a
|
|
233
|
+
later round cannot be screened against a set nobody recomputed.
|
|
234
|
+
|
|
235
|
+
Do **not** write the record inside the package directory before the ingest. A
|
|
236
|
+
round that names no `--review-id` has not allocated that directory yet, and
|
|
237
|
+
creating it makes the ingest take the next free name instead.
|
|
238
|
+
|
|
239
|
+
`--refuted` takes the findings this round **raised and then dismissed**, in the
|
|
240
|
+
same finding shape, each carrying `disposition: {state, evidence}` with one of
|
|
241
|
+
the `dismissed-*` states. Without it a package keeps only the survivors of an
|
|
242
|
+
unlogged triage — which is why precision measured over the recorded corpus
|
|
243
|
+
returns 100% whatever the reviewers actually got right.
|
|
244
|
+
|
|
245
|
+
`review complete --finding <id> --disposition <state> --evidence <ref>` records
|
|
246
|
+
what became of a finding when the fix round closes; the triple is repeatable, one
|
|
247
|
+
group per finding. States: `unknown`, `acted-on`, `dismissed-incorrect`,
|
|
248
|
+
`dismissed-wont-fix`, `dismissed-out-of-scope`, `dismissed-deprioritised`.
|
|
249
|
+
Everything except `unknown` must cite where the outcome is written down — the
|
|
250
|
+
commit, the test, the decision. A recorded state and its citation cannot be
|
|
251
|
+
overwritten by a later close; record a correction as a new round.
|
|
252
|
+
|
|
253
|
+
**Closing a fix round without dispositions leaves every finding reading
|
|
254
|
+
`unknown`.** That is not "the reviewers were right"; it is "nobody wrote down
|
|
255
|
+
what happened", and it is the single reason the precision figure cannot be read.
|
|
256
|
+
|
|
173
257
|
Managed modes:
|
|
174
258
|
|
|
175
259
|
- `lightweight`: report-only; no flow or managed review artifacts are created.
|
|
@@ -193,9 +277,150 @@ When attaching to a flow, resolve the flow by explicit `flow_id`, PR URL, issue
|
|
|
193
277
|
URL, or branch metadata. Never mutate `.metaproject/flows/*/flow.json` from
|
|
194
278
|
review code; Task Manager state changes remain owned by `keryx flow`.
|
|
195
279
|
|
|
280
|
+
---
|
|
281
|
+
|
|
282
|
+
## External PR comments
|
|
283
|
+
|
|
284
|
+
A human or a bot reviews our pull request. Before this existed, nothing collected
|
|
285
|
+
it, nothing fixed it and nothing answered it — silence was the behaviour for every
|
|
286
|
+
comment, and to the person who wrote it silence is indistinguishable from
|
|
287
|
+
disagreement.
|
|
288
|
+
|
|
289
|
+
**Collect every round. Answer once, at the end.** Both halves are mechanical and
|
|
290
|
+
both live in the CLI; the judgement — is this comment right, what do we say — stays
|
|
291
|
+
here.
|
|
292
|
+
|
|
293
|
+
```text
|
|
294
|
+
keryx review comments collect --repo <owner/repo> --pr <n> --sha <head-sha>
|
|
295
|
+
[--self <login>] [--round <n>] [--out <findings.json>] [--json]
|
|
296
|
+
keryx review comments reply --repo <owner/repo> --pr <n> --outcomes <file|->
|
|
297
|
+
--sha <head-sha> --final [--dry-run]
|
|
298
|
+
[--max-replies <n>] [--max-sentences <n>] [--max-chars <n>]
|
|
299
|
+
[--flow-link <url>]
|
|
300
|
+
```
|
|
301
|
+
|
|
302
|
+
Add `--fixtures <dir>` to either to run the whole loop against JSON on disk —
|
|
303
|
+
no token, no network, nothing posted. Use it to see what a reply pass would say
|
|
304
|
+
before it says it.
|
|
305
|
+
|
|
306
|
+
`--sha` is the commit you collected against, and it is required. The completion
|
|
307
|
+
gate compares it to the pull request's head: a collection that ran before the
|
|
308
|
+
comments arrived is **stale**, and a gate that could not tell the difference
|
|
309
|
+
would pass a flow with unanswered reviewers on it while printing
|
|
310
|
+
`0 outstanding`. A record with no SHA reads as "cannot be shown current", never
|
|
311
|
+
as "fresh".
|
|
312
|
+
|
|
313
|
+
### What collection does, so you do not do it by hand
|
|
314
|
+
|
|
315
|
+
- Reads **all three** sources: inline review comments, review submissions and
|
|
316
|
+
their bodies, and PR-level discussion.
|
|
317
|
+
- **A bot reviewer is a reviewer.** CodeRabbit, Greptile and Copilot comments go
|
|
318
|
+
down exactly the same path as a human's. The bot flag is recorded so a report
|
|
319
|
+
can say who spoke; nothing filters on it.
|
|
320
|
+
- Excludes our own identity, and comments already answered — **unless** the thread
|
|
321
|
+
has a newer reply from somebody else, which makes the comment new again.
|
|
322
|
+
- Everything filtered is listed with its reason. A filter that removes silently
|
|
323
|
+
reads as "nobody commented".
|
|
324
|
+
|
|
325
|
+
### Severity is classified, never invented
|
|
326
|
+
|
|
327
|
+
A comment on a review whose state is `CHANGES_REQUESTED` starts at **`major`**.
|
|
328
|
+
Everything else starts at **`minor`**. There is no third rule and no model call.
|
|
329
|
+
|
|
330
|
+
When the classifying fact is missing — an inline comment whose parent review was
|
|
331
|
+
not returned, or a review state GitHub does not document — the comment is **not
|
|
332
|
+
dropped and the severity is not guessed**: it takes the `minor` floor and carries
|
|
333
|
+
`basis: unclassified` naming what was missing. A derived `minor` and a defaulted
|
|
334
|
+
one are different claims, and a record that cannot tell them apart is the
|
|
335
|
+
`dismissed-out-of-scope: 0` failure in a new field.
|
|
336
|
+
|
|
337
|
+
You may **lower** a severity only by assigning a terminal disposition with a
|
|
338
|
+
reason. You may never silently drop an external comment.
|
|
339
|
+
|
|
340
|
+
### The verifier cannot refute an external comment
|
|
341
|
+
|
|
342
|
+
An external finding enters the same fix loop as an internal one with one
|
|
343
|
+
exception: **a `refuted` verdict does not remove it and does not dismiss it.** A
|
|
344
|
+
human asked a question; a machine deciding the question was invalid is not an
|
|
345
|
+
answer. `keryx review ingest` turns that verdict into the disposition
|
|
346
|
+
`answered-disagree`, keeps the finding, and records the reclaim in `scope.md`.
|
|
347
|
+
`answered-disagree` still owes a reply explaining why.
|
|
348
|
+
|
|
349
|
+
The per-reviewer findings cap does not truncate external comments either, for the
|
|
350
|
+
same reason: the cap drops silently, and an external comment may not be dropped
|
|
351
|
+
silently.
|
|
352
|
+
|
|
353
|
+
### Replying — once, at the end, briefly
|
|
354
|
+
|
|
355
|
+
The reply pass runs **after the final round and before the completion gate**, so
|
|
356
|
+
every reply states a settled outcome rather than an intention. `keryx review
|
|
357
|
+
comments reply` refuses without `--final`; it is not a reminder you can skip.
|
|
358
|
+
|
|
359
|
+
| Outcome | Reply is |
|
|
360
|
+
|---|---|
|
|
361
|
+
| `acted-on` | one sentence naming what changed, plus the commit SHA |
|
|
362
|
+
| `answered-disagree` | one or two sentences on why not, and a link to the flow's journal entry |
|
|
363
|
+
| `dismissed-out-of-scope` / `dismissed-deprioritised` | one sentence, and where it was recorded instead |
|
|
364
|
+
|
|
365
|
+
Rules, all of them enforced in code rather than asked for here:
|
|
366
|
+
|
|
367
|
+
- **At most two sentences per comment.** A longer reply is CUT to two and the
|
|
368
|
+
remainder is replaced by a link — the long version is not reachable from the
|
|
369
|
+
command's output. A truncation with no link to point at is refused outright: the
|
|
370
|
+
conclusion posted and the explanation nowhere is worse than either alternative.
|
|
371
|
+
- A fenced code block in a reply is refused. Link, do not paste.
|
|
372
|
+
- Replies go **in the thread**. A review submission body and a PR-level comment
|
|
373
|
+
have no thread — GitHub offers no reply endpoint for either — so those become one
|
|
374
|
+
top-level comment that names what it answers.
|
|
375
|
+
- **Never resolve or hide a thread we did not open.** Replying is ours; resolving
|
|
376
|
+
is the reviewer's call, and auto-resolving is how a bot silences a human. The
|
|
377
|
+
resolve, hide, minimise and dismiss endpoints are unreachable through the port
|
|
378
|
+
this command uses, GraphQL included.
|
|
379
|
+
- Exactly **one** reply per comment, and one disposition. A round that changed
|
|
380
|
+
nothing for a comment still gets a reply saying so, with a terminal disposition
|
|
381
|
+
— `unknown` is refused, because it is what an unanswered comment already reads
|
|
382
|
+
as.
|
|
383
|
+
- Capped at **30** replies. Beyond it, one summary comment and a backlog reported
|
|
384
|
+
by id.
|
|
385
|
+
- Handling is durable: `.metaproject/reviews/pr-comments/<owner>__<repo>__<n>.json`
|
|
386
|
+
records id, thread, author, url, first-seen round, handled-at, sha, disposition
|
|
387
|
+
and reply url, written after **every** post. A resumed session answers nobody
|
|
388
|
+
twice.
|
|
389
|
+
|
|
390
|
+
**The trade-off, stated rather than hidden:** a reviewer who comments early waits
|
|
391
|
+
until the end. That is deliberate — answering with a work-in-progress state that
|
|
392
|
+
later changes is worse. If a comment **blocks** progress rather than reporting a
|
|
393
|
+
problem, mark its outcome `escalate: true`: it leaves the reply queue, is reported
|
|
394
|
+
to the operator immediately, and the command exits non-zero. Answering a blocking
|
|
395
|
+
question at the end answers the wrong question late.
|
|
396
|
+
|
|
397
|
+
---
|
|
398
|
+
|
|
399
|
+
## Everything written to GitHub is brief
|
|
400
|
+
|
|
401
|
+
One rule, applied to every outward surface: **PR bodies, PR comments, review
|
|
402
|
+
replies, issue comments, and commit messages going to a PR.**
|
|
403
|
+
|
|
404
|
+
- Lead with the conclusion. No preamble, no restating the question, no apology,
|
|
405
|
+
no summary of the flow.
|
|
406
|
+
- Say what changed and where. **Link, do not paste.**
|
|
407
|
+
- The reasoning, the evidence, the rejected alternatives and the round history live
|
|
408
|
+
in the flow package — `journal.md`, `context.md`, the review artifacts — which is
|
|
409
|
+
durable, searchable, and costs a reader nothing to skip.
|
|
410
|
+
- A GitHub artifact that needs more than a short paragraph is a signal that the
|
|
411
|
+
detail belongs in the flow with a link out, **not** that the paragraph should
|
|
412
|
+
grow.
|
|
413
|
+
- No orchestrator-written PR comment or reply exceeds two sentences without
|
|
414
|
+
carrying a link to the artifact holding the detail. The reply pass enforces
|
|
415
|
+
this; for anything else you write outward, hold yourself to it.
|
|
416
|
+
|
|
417
|
+
This is deliberately asymmetric: **verbose in the flow, terse on GitHub.** The
|
|
418
|
+
flow is written for whoever resumes the work; GitHub is read by someone who did
|
|
419
|
+
not ask for our reasoning and is reading between other tasks.
|
|
420
|
+
|
|
196
421
|
## Review Context Pack
|
|
197
422
|
|
|
198
|
-
Before routing reviewers, build a compact `review_context` object. This is the shared source of truth for all sub-agents and must follow `skills/review-orchestrator/review-context.schema.json`.
|
|
423
|
+
Before routing reviewers, build a compact `review_context` object. This is the shared source of truth for all sub-agents and must follow `skills/review/review-orchestrator/review-context.schema.json`.
|
|
199
424
|
|
|
200
425
|
Required content:
|
|
201
426
|
- Request: raw user request, flags, review mode, explicit paths or commit range.
|
|
@@ -352,34 +577,152 @@ Before anything else, determine whether the request is **diff mode** or **path m
|
|
|
352
577
|
|
|
353
578
|
See shared script: `skills/shared/git-merge-base.md`
|
|
354
579
|
|
|
355
|
-
Run the script to determine `BASE_SHA`, then:
|
|
580
|
+
Run the script to determine `BASE_SHA`, then let the pre-filter build the scope:
|
|
356
581
|
|
|
357
582
|
```bash
|
|
358
|
-
|
|
359
|
-
|
|
583
|
+
keryx review scope --ref "${BASE_SHA}" --json > scope.json # KEEP THIS FILE
|
|
584
|
+
keryx review scope --ref "${BASE_SHA}" --scoped-diff # what reviewers get
|
|
360
585
|
```
|
|
361
586
|
|
|
587
|
+
**Keep `scope.json` until the round is ingested, and pass it as `--scope`.** That
|
|
588
|
+
file is how the drop list reaches the review record. `--append
|
|
589
|
+
"<review-package>/scope.md"` also writes it and is still supported — it now
|
|
590
|
+
REPLACES an existing `## Pre-filter scope` block rather than adding a second, and
|
|
591
|
+
`review ingest` carries any block it finds forward verbatim rather than
|
|
592
|
+
overwriting it — but `--scope scope.json` is the supported path, because it is
|
|
593
|
+
the one that does not depend on running two commands against the same file in the
|
|
594
|
+
right order.
|
|
595
|
+
|
|
596
|
+
**Do not run `git diff` yourself, and do not decide what to leave out.** The
|
|
597
|
+
pre-filter is deterministic code with no model call: it drops generated,
|
|
598
|
+
lockfile, snapshot, vendored and minified paths, drops whitespace-only and
|
|
599
|
+
comment-only change blocks, and bounds every retained change to ±20 lines of
|
|
600
|
+
context (`--context <n>`) instead of the whole file. Dropping a lockfile needs no
|
|
601
|
+
judgement, so it does not get one.
|
|
602
|
+
|
|
603
|
+
The record carries the retained scope **and every drop with its reason**. Both
|
|
604
|
+
halves are required: a scope that shrank without saying so reads afterwards as
|
|
605
|
+
"we reviewed everything". Note that `--scope` takes the WHOLE `--json` document —
|
|
606
|
+
handing over only its `counts` object is refused, because eight integers carry no
|
|
607
|
+
reason for any individual drop.
|
|
608
|
+
|
|
609
|
+
Use `.files` from `scope.json` for the auto-detection table below. A dropped path
|
|
610
|
+
must not select a reviewer, and **neither may a blast-radius path**: scope B is
|
|
611
|
+
under regression check, so a `.tsx` file that only appears there must not pull in
|
|
612
|
+
`review-frontend`. Reviewer selection is driven by the scope-A file list alone.
|
|
613
|
+
|
|
362
614
|
Scope is limited to **changes introduced in the current branch since merge-base**.
|
|
363
615
|
|
|
364
616
|
---
|
|
365
617
|
|
|
618
|
+
### Scope B — the blast radius (deep rounds)
|
|
619
|
+
|
|
620
|
+
Everything above is **scope A**: the change, bounded. It answers *is this change
|
|
621
|
+
correct?* It does not answer *did this change break something that was working*,
|
|
622
|
+
and those are different questions — only the first has ever been asked here.
|
|
623
|
+
|
|
624
|
+
A deep round dispatches under **both**. Scope B is computed, never browsed:
|
|
625
|
+
|
|
626
|
+
```bash
|
|
627
|
+
keryx review blast-radius --ref "${BASE_SHA}" --json > blast-radius.json # KEEP THIS FILE
|
|
628
|
+
keryx review blast-radius --ref "${BASE_SHA}" --brief # what a scope-B reviewer is told
|
|
629
|
+
```
|
|
630
|
+
|
|
631
|
+
**KEEP THIS FILE** is not advice. The ingest at Step 12 is **refused** if a
|
|
632
|
+
scope-B finding arrives without it — pass it back as
|
|
633
|
+
`review ingest ... --blast-radius blast-radius.json`.
|
|
634
|
+
|
|
635
|
+
It walks `gdgraph affected` outward from every changed file, ranks by edge
|
|
636
|
+
distance, keeps distance ≤ 2, cuts at 40 files closest-first, and adds a changed
|
|
637
|
+
file's naming-related tests when the graph did not already reach them. Requires a
|
|
638
|
+
built graph — run `keryx gdgraph build` if it refuses.
|
|
639
|
+
|
|
640
|
+
**Do not pick the files yourself, and do not widen it.** "Review the
|
|
641
|
+
functionality so nothing breaks" naively means "review the whole repository every
|
|
642
|
+
round", which is unaffordable *and* actively harmful: review quality decays as
|
|
643
|
+
context grows — measured F1 0.65 at round 2 falling to 0.29 at round 10. An
|
|
644
|
+
unbounded scope B makes later rounds worse than earlier ones.
|
|
645
|
+
|
|
646
|
+
The bounds are measured on this repository, not guessed: at depth 2 the set is a
|
|
647
|
+
median of 19 files (p90 65); depth 3 buys eight more in the median and doubles
|
|
648
|
+
the p90. The 40-file cap fires on 25% of commits and removes only hop-2 entries
|
|
649
|
+
on all but 2 of 80, so it almost never costs a direct dependent — and when it
|
|
650
|
+
does, it says so.
|
|
651
|
+
|
|
652
|
+
**Record the whole thing.** `--out "<review-package>/blast-radius.md"` writes the
|
|
653
|
+
set, the depth, and **every file the cap removed**. A truncation nobody can see
|
|
654
|
+
reads afterwards as "we checked everything", which is the claim this pipeline
|
|
655
|
+
exists to stop making. An empty radius is reported as `unresolved`, not as clean:
|
|
656
|
+
the graph indexes code, so a change to a skill, a rule or a schema has no blast
|
|
657
|
+
radius at all and that is a different fact from "nothing depends on it".
|
|
658
|
+
|
|
659
|
+
#### The scope-B question, and what is rejected
|
|
660
|
+
|
|
661
|
+
> Does this change break an existing behaviour **at these sites**?
|
|
662
|
+
|
|
663
|
+
Nothing else. The blast-radius set is **under regression check, not under
|
|
664
|
+
review**. A finding about style, naming or architecture in code the change did
|
|
665
|
+
not touch is refused **by the orchestrator in code** — not discouraged here —
|
|
666
|
+
under three rules, every one of them a fact about the claim rather than about who
|
|
667
|
+
made it:
|
|
668
|
+
|
|
669
|
+
| Rule | Refused because |
|
|
670
|
+
|---|---|
|
|
671
|
+
| `outside-set` | the file is neither in the computed set nor in the changed set; the reviewer went browsing |
|
|
672
|
+
| `non-regression-severity` | below `major`. Under the canonical rubric `minor` states the code behaves correctly and `info` names neither trigger nor outcome; neither can be a claim that something broke |
|
|
673
|
+
| `no-link-to-change` | nothing in the finding names a changed file, module or symbol. A regression claim says THE CHANGE broke this site |
|
|
674
|
+
|
|
675
|
+
Rejections are **recorded, not deleted** — raise the observation under scope A or
|
|
676
|
+
as a separate review. Pass `--brief` output verbatim into the scope-B dispatch:
|
|
677
|
+
the code rejection is the enforcement, but a reviewer told afterwards has already
|
|
678
|
+
spent the round producing findings that will all be refused.
|
|
679
|
+
|
|
680
|
+
`class_scope` on a scope-B finding names the **caller that breaks**, not the
|
|
681
|
+
changed line, because that is the site a human has to look at.
|
|
682
|
+
|
|
683
|
+
#### When it is recomputed
|
|
684
|
+
|
|
685
|
+
| Round | Scope A | Scope B |
|
|
686
|
+
|---|---|---|
|
|
687
|
+
| 1 (first after the draft PR) | yes | yes |
|
|
688
|
+
| 2..N | yes | recomputed only if the changed-file set moved |
|
|
689
|
+
| final | yes | **yes, always** |
|
|
690
|
+
|
|
691
|
+
Do not decide this by memory:
|
|
692
|
+
|
|
693
|
+
```bash
|
|
694
|
+
keryx review blast-radius --ref "${BASE_SHA}" --previous blast-radius.json [--final]
|
|
695
|
+
```
|
|
696
|
+
|
|
697
|
+
It prints the decision and the reason, and reuses the previous record when
|
|
698
|
+
nothing moved. The final round recomputes whatever the file set did — otherwise a
|
|
699
|
+
fix introduced in round 3 gets no regression check at all, and the round that
|
|
700
|
+
certifies the flow is the one that checked the least.
|
|
701
|
+
|
|
702
|
+
---
|
|
703
|
+
|
|
366
704
|
### Path Mode
|
|
367
705
|
|
|
368
|
-
When a path or target is named, collect the files
|
|
706
|
+
When a path or target is named, collect the candidate files:
|
|
369
707
|
|
|
370
708
|
```bash
|
|
371
709
|
# If a directory path is given:
|
|
372
710
|
find <path> -type f \( -name "*.ts" -o -name "*.tsx" -o -name "*.js" -o -name "*.jsx" \) | sort
|
|
373
711
|
|
|
374
|
-
# If a file path is given:
|
|
375
|
-
cat <file>
|
|
376
|
-
|
|
377
712
|
# If a module name is given (e.g. "UserStore", "pipelines module"):
|
|
378
713
|
find . -type f -name "*<name>*" \( -name "*.ts" -o -name "*.tsx" \)
|
|
379
714
|
# Also check common locations: src/stores/, src/modules/, src/components/
|
|
380
715
|
```
|
|
381
716
|
|
|
382
|
-
|
|
717
|
+
Then put the list through the same exclusions before reading anything:
|
|
718
|
+
|
|
719
|
+
```bash
|
|
720
|
+
keryx review scope --path "src/a.ts,src/b.ts" --json > scope.json
|
|
721
|
+
```
|
|
722
|
+
|
|
723
|
+
Pass the full **file contents** of the paths it **retained** to sub-reviewers, and
|
|
724
|
+
read none of the ones it dropped. Set `SCOPE_MODE: path`. Path mode has no hunks
|
|
725
|
+
and therefore no context window; the drop list is recorded exactly the same way.
|
|
383
726
|
|
|
384
727
|
**Reviewer behavior in path mode:** reviewers check the entire file content — not just added lines. All findings apply to the current state of the code, not only to changes.
|
|
385
728
|
|
|
@@ -414,6 +757,33 @@ If the repository has local convention docs such as `CLAUDE.md`, `AGENTS.md`,
|
|
|
414
757
|
These convention reviewers are additive: keep the generic reviewers selected by normal detection,
|
|
415
758
|
then add the matching convention pass. Deduplicate reviewer names before dispatch.
|
|
416
759
|
|
|
760
|
+
### Stack scoping — run it after detection, before dispatch
|
|
761
|
+
|
|
762
|
+
The tables above select reviewers by **file shape**. A `.ts` file looks the same
|
|
763
|
+
whether or not the repository has React in it, so those tables will happily
|
|
764
|
+
dispatch a React/MobX conventions reviewer at a Bun CLI with no frontend — which
|
|
765
|
+
is exactly what happened here, on every review, for months.
|
|
766
|
+
|
|
767
|
+
So the selected set is filtered once more, by what the repository actually
|
|
768
|
+
declares:
|
|
769
|
+
|
|
770
|
+
```bash
|
|
771
|
+
keryx review stack --json
|
|
772
|
+
```
|
|
773
|
+
|
|
774
|
+
It reads `package.json` and reports, per reviewer, `include` or `exclude` with a
|
|
775
|
+
reason. A reviewer carrying `metadata.stack_requires` is dispatched when **any**
|
|
776
|
+
tag it names is present — matching what `keryx review stack` actually computes,
|
|
777
|
+
and failing toward inclusion rather than away from it.
|
|
778
|
+
|
|
779
|
+
**Its failure mode is to include, never to skip.** A missing, unparsable or
|
|
780
|
+
unexpected manifest sets `uncertain`, and an uncertain detection marks every tag
|
|
781
|
+
present, so every reviewer runs. A reviewer that runs needlessly costs tokens; a
|
|
782
|
+
reviewer wrongly skipped hides a real defect, and that asymmetry is not close.
|
|
783
|
+
|
|
784
|
+
Record the exclusions with their reasons alongside the pre-filter drops. A
|
|
785
|
+
reviewer silently absent from a report reads as "it had nothing to say".
|
|
786
|
+
|
|
417
787
|
### Convention Reviewer Confirmation
|
|
418
788
|
|
|
419
789
|
When convention reviewers are auto-detected and the user did not explicitly pass
|
|
@@ -447,9 +817,9 @@ Legacy/profile reviewers are specialized review profiles that predate the review
|
|
|
447
817
|
|
|
448
818
|
| Trigger | Reviewers appended |
|
|
449
819
|
|---|---|
|
|
450
|
-
| `--legacy-profiles` | `code-ai-review` + `code-
|
|
820
|
+
| `--legacy-profiles` | `code-ai-review` + `code-learned-review` + `code-style-review` + `code-mobx-store-review` when MobX/store files are present |
|
|
451
821
|
| `--code-ai` | `code-ai-review` |
|
|
452
|
-
| `--
|
|
822
|
+
| `--learned` | `code-learned-review` |
|
|
453
823
|
| `--code-style` | `code-style-review` |
|
|
454
824
|
| `--mobx-store` | `code-mobx-store-review` |
|
|
455
825
|
| `*.store.ts`, `makeObservable`, `observable`, `computed`, `action.bound` | suggest `code-mobx-store-review` as optional profile reviewer |
|
|
@@ -469,13 +839,13 @@ Review Plan Preview must include an `Optional legacy/profile reviewers` group an
|
|
|
469
839
|
```text
|
|
470
840
|
Optional legacy/profile reviewers:
|
|
471
841
|
- code-ai-review: available via --code-ai or --legacy-profiles
|
|
472
|
-
- code-
|
|
842
|
+
- code-learned-review: available via --learned or --legacy-profiles
|
|
473
843
|
- code-style-review: available via --code-style or --legacy-profiles
|
|
474
844
|
- code-mobx-store-review: auto-suggest when *.store.ts or MobX patterns are present; available via --mobx-store or --legacy-profiles
|
|
475
845
|
|
|
476
846
|
Skipped reviewers:
|
|
477
847
|
- code-ai-review: profile reviewer, not selected unless --code-ai/--legacy-profiles
|
|
478
|
-
- code-
|
|
848
|
+
- code-learned-review: profile reviewer, not selected unless --learned/--legacy-profiles
|
|
479
849
|
- code-style-review: legacy style profile, not selected unless --code-style/--legacy-profiles
|
|
480
850
|
- code-mobx-store-review: not selected unless --mobx-store/--legacy-profiles or MobX store files are detected
|
|
481
851
|
```
|
|
@@ -498,7 +868,7 @@ Skipped reviewers:
|
|
|
498
868
|
| `--core-boundaries` | `review-core-boundaries` |
|
|
499
869
|
| `--flow-graph` | `review-flow-graph` |
|
|
500
870
|
| `--all` | all reviewers above (including `review-clean-code`, `review-highload`, applicable legacy/profile reviewers, and project convention reviewers when local convention docs exist) |
|
|
501
|
-
| `--
|
|
871
|
+
| `--verify` | `review-verifier`, AFTER all others; checks the consolidated findings by running something. Delete-only. |
|
|
502
872
|
| (auto) | detected from diff file extensions — see Auto-detection table |
|
|
503
873
|
|
|
504
874
|
Multiple flags may be combined. Example: `review --backend --security` dispatches
|
|
@@ -526,7 +896,85 @@ Dispatch selected reviewers in parallel when independent. Use waves when token b
|
|
|
526
896
|
|
|
527
897
|
1. Wave A - core correctness/risk reviewers: logic, architecture, security/highload when selected.
|
|
528
898
|
2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files.
|
|
529
|
-
3. Wave C -
|
|
899
|
+
3. Wave C - **verification**: `review-verifier` over the consolidated findings, when blockers/majors
|
|
900
|
+
exist, `--verify` is set, or the PR is high-risk. See below.
|
|
901
|
+
|
|
902
|
+
### Wave C — verification, and what it replaced
|
|
903
|
+
|
|
904
|
+
Wave C used to run `review-strict`: a meta-pass that re-read the consolidated
|
|
905
|
+
findings and **adjusted their severity with no new evidence**, under an elevation
|
|
906
|
+
table biased 3:1 toward escalation. It was **removed, not improved**, and the
|
|
907
|
+
reason is measured rather than stylistic:
|
|
908
|
+
|
|
909
|
+
- **GPT-4 on GSM8K across self-correction rounds: 95.5 → 91.5 → 89.0.**
|
|
910
|
+
**GPT-3.5 on CommonSenseQA: 75.8 → 38.1.** Among the answers that changed,
|
|
911
|
+
correct → incorrect exceeded incorrect → correct (Huang et al., *Large Language
|
|
912
|
+
Models Cannot Self-Correct Reasoning Yet*, ICLR 2024, arXiv:2310.01798).
|
|
913
|
+
- **Self-Refine (arXiv:2303.17651): +49.2 on dialogue response generation, +0.2
|
|
914
|
+
on maths.** Self-refinement gains are on subjective tasks and vanish on
|
|
915
|
+
verifiable reasoning. Judging whether a null-guard is missing is verifiable
|
|
916
|
+
reasoning.
|
|
917
|
+
|
|
918
|
+
Re-scoring a finding by re-reading it is therefore not a rigour pass; it is a
|
|
919
|
+
coin flip weighted toward more findings. **Do not restore it because it looks
|
|
920
|
+
obviously useful — it looked obviously useful the first time.**
|
|
921
|
+
|
|
922
|
+
`review-verifier` occupies the slot and differs in exactly one way that matters:
|
|
923
|
+
**it runs something.** Verification that executes rejects 85–96% of false reports
|
|
924
|
+
against 4–15% unaided while finding 30–44% more true bugs (AnyPoC,
|
|
925
|
+
arXiv:2604.11950); Meta's TestGen-LLM funnel discards 75% of its own output
|
|
926
|
+
(75% build → 57% build and pass → 25% improve coverage) and the surviving quarter
|
|
927
|
+
reaches 73% human acceptance (arXiv:2402.09171).
|
|
928
|
+
|
|
929
|
+
It also **never votes.** 80+ agents unanimously endorsed a padding-oracle
|
|
930
|
+
vulnerability that did not exist, and a single empirical test killed it: consensus
|
|
931
|
+
cannot detect a hallucination its members share, so agreement between reviewers is
|
|
932
|
+
not evidence and must never be recorded as verification.
|
|
933
|
+
|
|
934
|
+
That rule is about agreement *standing in for* evidence. It is not a rule that
|
|
935
|
+
two verifiers may not both check the same finding: each claim is admitted on its
|
|
936
|
+
own — a named non-author, a real method, real evidence, with `reasoning` already
|
|
937
|
+
capped — and when two such claims reach the **same** verdict the merge records it,
|
|
938
|
+
naming both verifiers and carrying both pieces of evidence. Claims that
|
|
939
|
+
**disagree** still cancel, because there the only thing deciding the outcome
|
|
940
|
+
would be claim order.
|
|
941
|
+
|
|
942
|
+
Dispatch rules:
|
|
943
|
+
|
|
944
|
+
- Pass the consolidated findings, each carrying `global_id` and the **real**
|
|
945
|
+
originating `reviewer`. A finding whose `reviewer` is the orchestrator cannot be
|
|
946
|
+
routed away from its author, so the never-self-verify rule silently stops
|
|
947
|
+
applying — that field was hardcoded to `review-orchestrator` on all 83 recorded
|
|
948
|
+
findings and is fixed only from 0.2.70 onward.
|
|
949
|
+
- **A finding is never verified by the reviewer that raised it.** When only one
|
|
950
|
+
reviewer ran, its findings are simply left unverified; verifying them yourself
|
|
951
|
+
is worse than not verifying them. The merge compares the two names after
|
|
952
|
+
normalising case, surrounding whitespace, `_`/`-`, and a trailing `(model)`
|
|
953
|
+
annotation, so `review-logic `, `Review-Logic` and `review-logic (sonnet)` are
|
|
954
|
+
all the same actor. Do not try to route around it by respelling the name — the
|
|
955
|
+
comparison deliberately over-matches, because a refused claim only ever costs a
|
|
956
|
+
verdict while a missed self-verification costs the finding.
|
|
957
|
+
- The verifier returns `verification-claim.schema.json`. Merge it with
|
|
958
|
+
`keryx review ingest --verifications <file>`; do not apply verdicts by hand.
|
|
959
|
+
- **The verifier can only delete.** If it returns a severity, a new finding, or a
|
|
960
|
+
rewritten finding, the merge discards that whole claim and records the attempt.
|
|
961
|
+
Do not "help" by applying it.
|
|
962
|
+
|
|
963
|
+
### `verification_mode`
|
|
964
|
+
|
|
965
|
+
`off` | `annotate` | `filter`. **Default `annotate`, and it stays `annotate` for
|
|
966
|
+
one release.**
|
|
967
|
+
|
|
968
|
+
| Mode | What happens |
|
|
969
|
+
|---|---|
|
|
970
|
+
| `off` | No verification. Claims are refused rather than silently ignored. |
|
|
971
|
+
| `annotate` | Verdicts are recorded on the findings. **Nothing is removed.** A `refuted` finding is still reported, marked refuted. |
|
|
972
|
+
| `filter` | An applied `refuted` verdict removes the finding from the reported set and records it as `dismissed-incorrect`, with the verification evidence. |
|
|
973
|
+
|
|
974
|
+
`annotate` is the default so the drop rate is a **measured number** before it
|
|
975
|
+
costs a real finding. The risk is named rather than assumed away: SWE-agent keeps
|
|
976
|
+
its equivalent step opt-in because it sometimes rejects correct patches. Do not
|
|
977
|
+
switch a project to `filter` on the strength of one round.
|
|
530
978
|
|
|
531
979
|
### Agent Runtime Compatibility
|
|
532
980
|
|
|
@@ -541,7 +989,7 @@ Runtime rules:
|
|
|
541
989
|
|
|
542
990
|
Do not use vague fallback messages such as "running through available agent types" without naming which reviewers used fallback and why.
|
|
543
991
|
|
|
544
|
-
Pass each sub-reviewer a payload matching `skills/review-orchestrator/reviewer-input.schema.json`:
|
|
992
|
+
Pass each sub-reviewer a payload matching `skills/review/review-orchestrator/reviewer-input.schema.json`:
|
|
545
993
|
|
|
546
994
|
```yaml
|
|
547
995
|
review_context: <bounded context pack>
|
|
@@ -564,7 +1012,7 @@ target_path: <resolved path or file list>
|
|
|
564
1012
|
file_contents: <bounded file contents relevant to this reviewer>
|
|
565
1013
|
```
|
|
566
1014
|
|
|
567
|
-
Each reviewer must return a `REVIEW_RESULT` object matching `skills/review-orchestrator/reviewer-finding.schema.json`, followed by a concise markdown summary. The orchestrator must reject or normalize free-form reports before consolidation.
|
|
1015
|
+
Each reviewer must return a `REVIEW_RESULT` object matching `skills/review/review-orchestrator/reviewer-finding.schema.json`, followed by a concise markdown summary. The orchestrator must reject or normalize free-form reports before consolidation.
|
|
568
1016
|
|
|
569
1017
|
**Important for path mode:** instruct each reviewer to check the **entire file**, not just changes. The scope report should say "Path: `<TARGET_PATH>`" instead of a branch/merge-base.
|
|
570
1018
|
|
|
@@ -579,13 +1027,14 @@ Each reviewer must return a `REVIEW_RESULT` object matching `skills/review-orche
|
|
|
579
1027
|
| Security vulnerabilities | NO | `review-security-code` |
|
|
580
1028
|
| Performance anti-patterns | NO | `review-performance` |
|
|
581
1029
|
| Style / naming / import order | NO | `review-style` |
|
|
1030
|
+
| Checking whether a reported finding is real | NO | `review-verifier` |
|
|
582
1031
|
| Clean Code principles + SOLID at code level | NO | `review-clean-code` |
|
|
583
1032
|
| Concurrency, resource pools, caching, queues, idempotency | NO | `review-highload` |
|
|
584
1033
|
| Frontend repository conventions | NO | `review-frontend-conventions` |
|
|
585
1034
|
| Test / e2e conventions | NO | `review-testing-practices` |
|
|
586
1035
|
| Shared core boundary rules | NO | `review-core-boundaries` |
|
|
587
1036
|
| Shared flow/graph abstraction contracts | NO | `review-flow-graph` |
|
|
588
|
-
| Legacy/profile review profiles | NO | `code-ai-review`, `code-
|
|
1037
|
+
| Legacy/profile review profiles | NO | `code-ai-review`, `code-learned-review`, `code-style-review`, `code-mobx-store-review` |
|
|
589
1038
|
|
|
590
1039
|
---
|
|
591
1040
|
|
|
@@ -602,6 +1051,93 @@ Before consolidation, validate every reviewer result:
|
|
|
602
1051
|
|
|
603
1052
|
---
|
|
604
1053
|
|
|
1054
|
+
## Severity (canonical)
|
|
1055
|
+
|
|
1056
|
+
**This is the only severity rubric in the review domain.** Reviewers do not carry
|
|
1057
|
+
their own. Ten private rubrics feeding one sort produce a ranking that means ten
|
|
1058
|
+
different things at once, and ranking is what an operator uses to decide what to
|
|
1059
|
+
read first. A reviewer may state which of *its* conditions land where; it may not
|
|
1060
|
+
redefine the levels.
|
|
1061
|
+
|
|
1062
|
+
### `blocker` — merge-blocking, and nothing else
|
|
1063
|
+
|
|
1064
|
+
Exactly four shapes. Nothing outside this list is a `blocker`, however strongly
|
|
1065
|
+
the reviewer feels about it:
|
|
1066
|
+
|
|
1067
|
+
1. **A crash** — the process, request, or render dies on an input the change
|
|
1068
|
+
admits.
|
|
1069
|
+
2. **Data loss or corruption** — something persisted, transmitted, or returned is
|
|
1070
|
+
destroyed or silently wrong.
|
|
1071
|
+
3. **An exploitable vulnerability** — an attacker action with a named entry point
|
|
1072
|
+
and a named impact.
|
|
1073
|
+
4. **An unimplemented acceptance criterion** — the change claims work the diff
|
|
1074
|
+
does not contain.
|
|
1075
|
+
|
|
1076
|
+
Everything else is at most `major`. "This will definitely cause problems later"
|
|
1077
|
+
is not one of the four. Neither is "this violates the architecture", "this fails
|
|
1078
|
+
the linter", or "this is how the last outage started".
|
|
1079
|
+
|
|
1080
|
+
### `major` / `minor` / `info` — the boundary test
|
|
1081
|
+
|
|
1082
|
+
Ask one question, and ask it of the **finding**, not of the code:
|
|
1083
|
+
|
|
1084
|
+
> **Does it name a trigger, and the observable outcome that trigger produces?**
|
|
1085
|
+
|
|
1086
|
+
- **`major`** — it does. There is an input, a call, a render, or a load level, and
|
|
1087
|
+
a resulting behaviour a user or a caller would call wrong: a wrong value, a lost
|
|
1088
|
+
update, a leak, a hang, a cost stated together with the frequency that makes it
|
|
1089
|
+
a cost. Not one of the four shapes above, so not merge-blocking — but the code
|
|
1090
|
+
does the wrong thing.
|
|
1091
|
+
- **`minor`** — it does not, and does not claim to. The code behaves correctly;
|
|
1092
|
+
the cost lands on whoever reads or edits it next, and the finding names that
|
|
1093
|
+
cost at a named site.
|
|
1094
|
+
- **`info`** — it names neither. An observation, a preference, or a risk with no
|
|
1095
|
+
demonstrated path.
|
|
1096
|
+
|
|
1097
|
+
The test is procedural on purpose. It is applied by reading the finding, so
|
|
1098
|
+
someone who did not write it — and has not read the code — reaches the same
|
|
1099
|
+
answer: look for the trigger and the outcome. Present → `major`. Absent, but a
|
|
1100
|
+
concrete maintenance cost is named → `minor`. Neither → `info`.
|
|
1101
|
+
|
|
1102
|
+
Two consequences, both previously decided differently in different files:
|
|
1103
|
+
|
|
1104
|
+
- A finding that **claims** runtime harm and cannot name the trigger is `info`,
|
|
1105
|
+
not `major`. It is not demoted to `minor`: `minor` is for findings that never
|
|
1106
|
+
claimed runtime harm at all. The two are different failures and stay
|
|
1107
|
+
distinguishable.
|
|
1108
|
+
- Severity is a property of the demonstrated outcome, never of the reviewer that
|
|
1109
|
+
found it. A security reviewer's unproven concern is `info` under the same test
|
|
1110
|
+
that puts a style reviewer's unproven concern there.
|
|
1111
|
+
- **And never of how crisply the finding is worded.** An outcome that costs a
|
|
1112
|
+
user, a caller or persisted state nothing is `minor` however precisely its
|
|
1113
|
+
trigger is named. Without this clause the test above rates prose quality: a
|
|
1114
|
+
cosmetic wording nit stated as "trigger X produces output Y" reads as `major`,
|
|
1115
|
+
while a real defect stated tersely reads as `info`. That is not academic — the
|
|
1116
|
+
findings cap truncates by severity, so the well-written typo would survive and
|
|
1117
|
+
the terse real defect would be cut.
|
|
1118
|
+
|
|
1119
|
+
The boundary this rubric does **not** draw is `major` against `major`. Two
|
|
1120
|
+
findings that both name a trigger and an outcome are the same severity even when
|
|
1121
|
+
one is obviously worse; the ordering inside a severity is the operator's, and
|
|
1122
|
+
inventing a fifth level to express it would put us back where we started.
|
|
1123
|
+
|
|
1124
|
+
### Shared laws (every reviewer)
|
|
1125
|
+
|
|
1126
|
+
1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
|
|
1127
|
+
name the input, call, or condition that reaches the code, you have an
|
|
1128
|
+
observation, not a finding. Report it as `info` and say what would settle it.
|
|
1129
|
+
2. **Never flag the theoretical.** The path you describe must exist in the code
|
|
1130
|
+
under review. Do not report a safe API because it could be misused, or a
|
|
1131
|
+
pattern because it is often wrong elsewhere.
|
|
1132
|
+
3. **One finding per class, not one per occurrence.** When the same shape appears
|
|
1133
|
+
at several sites, report it once and list every site. Ten findings that are one
|
|
1134
|
+
finding hide the other nine problems.
|
|
1135
|
+
|
|
1136
|
+
`review-security-code` carries a fourth — every security finding states its attack
|
|
1137
|
+
vector — which does not generalise and stays there.
|
|
1138
|
+
|
|
1139
|
+
---
|
|
1140
|
+
|
|
605
1141
|
## Finding Format
|
|
606
1142
|
|
|
607
1143
|
### Class scope — required for `blocker` and `major`
|
|
@@ -702,6 +1238,18 @@ STATUS: DONE | DONE_WITH_CONCERNS
|
|
|
702
1238
|
- minor: N
|
|
703
1239
|
- info: N
|
|
704
1240
|
|
|
1241
|
+
## Stage counts
|
|
1242
|
+
<!-- Required. State what each stage REMOVED, and never state it as a precision
|
|
1243
|
+
improvement: no precision baseline exists to improve on. The one measured
|
|
1244
|
+
from the review packages on disk was 53/53 = 100% — pinned there by
|
|
1245
|
+
construction, because nothing in that corpus could record a finding as
|
|
1246
|
+
wrong. Copy these from `scope.md`; do not re-count by hand. -->
|
|
1247
|
+
- dropped by pre-filter: <files>, <blocks>, <changed lines> (or `not recorded` if no scope was built)
|
|
1248
|
+
- verification mode: `<off | annotate | filter>`
|
|
1249
|
+
- verdicts: confirmed N, refuted N, unverifiable N, unverified N
|
|
1250
|
+
- refuted by the verifier: N (removed: N — always 0 outside `filter`)
|
|
1251
|
+
- retained: N
|
|
1252
|
+
|
|
705
1253
|
## Blockers (must fix before merge)
|
|
706
1254
|
<[F-NNN] findings with severity=blocker, sorted by file>
|
|
707
1255
|
|
|
@@ -770,13 +1318,22 @@ Publish this review report to the PR?
|
|
|
770
1318
|
|
|
771
1319
|
### Concise PR Comment
|
|
772
1320
|
|
|
773
|
-
The visible PR comment is for humans. It must be written in English only and stay
|
|
1321
|
+
The visible PR comment is for humans. It must be written in English only and stay
|
|
1322
|
+
concise, under the brevity rule above: **the summary is at most two sentences and
|
|
1323
|
+
carries a link to the artifact holding the detail.**
|
|
1324
|
+
|
|
1325
|
+
The finding rows below are a bounded exception, not a licence: they exist because
|
|
1326
|
+
a reviewer scanning a PR needs the blockers in front of them. Keep them to the
|
|
1327
|
+
`blocker` and `major` rows; everything at `minor` or below goes behind the
|
|
1328
|
+
`<details>` fold or, better, into the AI artifact and is linked. The full findings
|
|
1329
|
+
set, the round history and the reasoning belong in the flow package — pasting them
|
|
1330
|
+
here is the failure this rule names.
|
|
774
1331
|
|
|
775
1332
|
```markdown
|
|
776
1333
|
## AI Review Report
|
|
777
1334
|
|
|
778
1335
|
**Verdict:** REQUEST_CHANGES
|
|
779
|
-
**Summary:**
|
|
1336
|
+
**Summary:** At most two sentences: the overall risk and the main merge blocker. Detail: <link to the AI artifact or the flow package>.
|
|
780
1337
|
|
|
781
1338
|
| Severity | Area | Finding | Suggested Fix | Owner |
|
|
782
1339
|
|---|---|---|---|---|
|
|
@@ -948,6 +1505,18 @@ If absent, proceed normally — context is optional and non-blocking.
|
|
|
948
1505
|
| "Spec compliance can wait until after quality review" | Stage 1 gate exists because unimplemented requirements invalidate quality work |
|
|
949
1506
|
| "I'll deduplicate findings manually in my head" | Always normalize to [F-NNN] format before consolidation to avoid losing findings |
|
|
950
1507
|
| "Minor findings from one reviewer cancel out the major from another" | Each finding stands independently; severity is per-finding, not averaged |
|
|
1508
|
+
| "The reviewer's own severity table said blocker" | There are no reviewer tables. One rubric, in **Severity (canonical)** above; a reviewer that ships one is the defect this replaced |
|
|
1509
|
+
| "It's a security/architecture finding, so it's a blocker" | Severity is the demonstrated outcome, not the domain that found it. `blocker` is exactly the four shapes |
|
|
1510
|
+
| "It will definitely break something eventually, so blocker" | Name the trigger and the outcome. Named → `major`. Unnamed → `info`. "Eventually" is neither |
|
|
1511
|
+
| "A strict re-read of the findings will sharpen them" | That pass existed and was removed: self-correction without new evidence measured 95.5 → 91.5 → 89.0 on GSM8K and 75.8 → 38.1 on CommonSenseQA. Run something instead |
|
|
1512
|
+
| "Three reviewers agree, so the finding is verified" | Consensus is not evidence. 80+ agents unanimously endorsed a vulnerability that did not exist; one empirical test killed it |
|
|
1513
|
+
| "The verifier suggested a higher severity, I'll apply it" | It cannot suggest one. A claim carrying a severity is discarded whole and the attempt is recorded |
|
|
1514
|
+
| "This finding has no `verification`, so it can be dropped" | Absent means nobody checked. All 83 recorded findings are in that state; none of them is thereby wrong |
|
|
1515
|
+
| "Precision went up after the verifier landed" | There is no precision baseline to have gone up from. State stage counts: dropped, refuted, retained |
|
|
1516
|
+
| "I'll widen the blast radius, this change looks risky" | It is bounded because review quality decays with context: F1 0.65 at round 2 → 0.29 at round 10. Widening makes the later rounds worse, not safer |
|
|
1517
|
+
| "The blast radius came back empty, so nothing can break" | Empty and unresolved are different facts. The graph indexes code — a Markdown or JSON change has no radius at all, and the record says which one you got |
|
|
1518
|
+
| "The changed files are the same as last round, so scope B can be skipped on the final round" | The final round always recomputes. A fix landed in round 3 is the change; skipping means the certifying round checked the least |
|
|
1519
|
+
| "This scope-B file has an obvious naming problem, I'll report it" | Rejected in code: a naming problem is `minor` at best, and the floor is `major`. The set is under regression check, not under review — raise it under scope A |
|
|
951
1520
|
| "No flags means no reviewers" | No flags → run auto-detection; never produce an empty review |
|
|
952
1521
|
| "User named a module so I'll use diff mode" | Named module/component/store → path mode; diff mode is only for branch changes |
|
|
953
1522
|
| "Path mode should only show lines I'd flag in diff mode" | Path mode reviews the entire file — all findings apply, not just added lines |
|