@mrciphersmith/keryx 0.3.0 → 0.3.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +4 -1
- package/dist/cli.js +13362 -7078
- package/dist/core.js +11706 -11330
- package/package.json +1 -1
- package/src/gdskills/bundled/agents/go-code-auditor.md +1 -1
- package/src/gdskills/bundled/agents/python-code-auditor.md +1 -1
- package/src/gdskills/bundled/install-manifest.json +271 -4
- package/src/gdskills/bundled/rules/core/model-selection.mdc +51 -0
- package/src/gdskills/bundled/skills/review/review-jev-comments/SKILL.md +184 -0
- package/src/gdskills/bundled/skills/review/review-jev-docs/SKILL.md +189 -0
- package/src/gdskills/bundled/skills/review/review-jev-risk/SKILL.md +190 -0
- package/src/gdskills/bundled/skills/review/review-jev-rules/SKILL.md +267 -0
- package/src/gdskills/bundled/skills/review/review-jev-scenarios/SKILL.md +187 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +39 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +1 -1
- package/src/gdskills/bundled/stacks/angular/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/angular/governance/eval.json +1751 -0
- package/src/gdskills/bundled/stacks/angular/governance/scout.json +32 -0
- package/src/gdskills/bundled/stacks/angular/pack.json +55 -0
- package/src/gdskills/bundled/stacks/angular/rules/coding-style.mdc +82 -0
- package/src/gdskills/bundled/stacks/angular/rules/patterns.mdc +84 -0
- package/src/gdskills/bundled/stacks/angular/rules/security.mdc +70 -0
- package/src/gdskills/bundled/stacks/angular/rules/testing.mdc +73 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-build-fix/SKILL.md +127 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-build-fix/evals.json +72 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-code-review/SKILL.md +98 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-implementation/SKILL.md +112 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-testing/SKILL.md +102 -0
- package/src/gdskills/bundled/stacks/angular/skills/angular-testing/evals.json +71 -0
- package/src/gdskills/bundled/stacks/mobx/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/mobx/governance/eval.json +904 -0
- package/src/gdskills/bundled/stacks/mobx/governance/scout.json +18 -0
- package/src/gdskills/bundled/stacks/mobx/pack.json +28 -0
- package/src/gdskills/bundled/stacks/mobx/rules/coding-style.mdc +91 -0
- package/src/gdskills/bundled/stacks/mobx/rules/patterns.mdc +122 -0
- package/src/gdskills/bundled/stacks/mobx/rules/security.mdc +56 -0
- package/src/gdskills/bundled/stacks/mobx/rules/testing.mdc +63 -0
- package/src/gdskills/bundled/stacks/mobx/skills/mobx-observable-testing/SKILL.md +124 -0
- package/src/gdskills/bundled/stacks/mobx/skills/mobx-observable-testing/evals.json +73 -0
- package/src/gdskills/bundled/stacks/mobx/skills/mobx-store-implementation/SKILL.md +149 -0
- package/src/gdskills/bundled/stacks/mobx/skills/mobx-store-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/nestjs/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/nestjs/governance/eval.json +1308 -0
- package/src/gdskills/bundled/stacks/nestjs/governance/scout.json +34 -0
- package/src/gdskills/bundled/stacks/nestjs/pack.json +53 -0
- package/src/gdskills/bundled/stacks/nestjs/rules/coding-style.mdc +70 -0
- package/src/gdskills/bundled/stacks/nestjs/rules/patterns.mdc +83 -0
- package/src/gdskills/bundled/stacks/nestjs/rules/security.mdc +73 -0
- package/src/gdskills/bundled/stacks/nestjs/rules/testing.mdc +69 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-build-fix/SKILL.md +157 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-build-fix/evals.json +70 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-implementation/SKILL.md +129 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-implementation/evals.json +71 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-testing/SKILL.md +143 -0
- package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-testing/evals.json +69 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/governance/eval.json +2413 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/governance/scout.json +42 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/pack.json +42 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/coding-style.mdc +69 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/patterns.mdc +88 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/security.mdc +72 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/testing.mdc +64 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-build-fix/SKILL.md +147 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-build-fix/evals.json +75 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-code-review/SKILL.md +118 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-code-review/evals.json +76 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-implementation/SKILL.md +135 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-implementation/evals.json +78 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-testing/SKILL.md +116 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-testing/evals.json +75 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-upgrade-migration/SKILL.md +134 -0
- package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-upgrade-migration/evals.json +76 -0
- package/src/gdskills/bundled/stacks/vue/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/vue/governance/eval.json +2215 -0
- package/src/gdskills/bundled/stacks/vue/governance/scout.json +42 -0
- package/src/gdskills/bundled/stacks/vue/pack.json +42 -0
- package/src/gdskills/bundled/stacks/vue/rules/coding-style.mdc +73 -0
- package/src/gdskills/bundled/stacks/vue/rules/patterns.mdc +84 -0
- package/src/gdskills/bundled/stacks/vue/rules/security.mdc +60 -0
- package/src/gdskills/bundled/stacks/vue/rules/testing.mdc +69 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-build-fix/SKILL.md +137 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-build-fix/evals.json +72 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-code-review/SKILL.md +120 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-code-review/evals.json +71 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-implementation/SKILL.md +122 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-implementation/evals.json +72 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-testing/SKILL.md +115 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue-testing/evals.json +72 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue2-to-vue3-migration/SKILL.md +135 -0
- package/src/gdskills/bundled/stacks/vue/skills/vue2-to-vue3-migration/evals.json +71 -0
|
@@ -0,0 +1,189 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: review-jev-docs
|
|
3
|
+
model_tier: light
|
|
4
|
+
description: |
|
|
5
|
+
Use when: an ADDITIONAL, machine-scored pass is wanted over documentation sections that
|
|
6
|
+
may have gone STALE because of a diff — never replacing any other reviewer. Dispatched
|
|
7
|
+
by review-orchestrator in Wave B, via the CLI (`keryx review jev-docs`), when
|
|
8
|
+
`review.jev.docs: true` in .metaproject/tasks.config.json and a Jev/OpenRouter
|
|
9
|
+
credential is resolvable. It is not an LLM sub-agent: there is nothing to dispatch
|
|
10
|
+
through a platform-native agent mechanism, and no prose is generated by a model —
|
|
11
|
+
Jev only answers a `noul` staleness probability per doc section; keryx writes every
|
|
12
|
+
word of every finding, and one class of finding (a removed/renamed CLI flag still
|
|
13
|
+
documented) is found with no Jev call at all.
|
|
14
|
+
NOT for: judging whether documentation is well-written, complete, or accurate about
|
|
15
|
+
something the diff never touched — this reviewer only ever looks at sections
|
|
16
|
+
deterministically linked to code the diff changed.
|
|
17
|
+
triggers:
|
|
18
|
+
- "jev docs"
|
|
19
|
+
- "stale docs"
|
|
20
|
+
- "review --jev-docs"
|
|
21
|
+
metadata:
|
|
22
|
+
author: "MrCipherSmith"
|
|
23
|
+
version: "1.0.0"
|
|
24
|
+
category: "review"
|
|
25
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
26
|
+
engine: "jev"
|
|
27
|
+
origin: "authored"
|
|
28
|
+
license: "MIT"
|
|
29
|
+
---
|
|
30
|
+
|
|
31
|
+
# Review — Jev Docs (stale-documentation detection)
|
|
32
|
+
|
|
33
|
+
An ADDITIONAL orchestrator reviewer, flow 333. It is a **keryx program**, not an
|
|
34
|
+
LLM sub-agent: `review-orchestrator` never dispatches a platform-native agent
|
|
35
|
+
for it, it runs `keryx review jev-docs` and reads the `--json` output. Every
|
|
36
|
+
finding is composed by keryx; Jev ("System One" on OpenRouter) supplies only a
|
|
37
|
+
`noul` staleness probability per doc section — it never writes prose, and
|
|
38
|
+
nothing it returns is quoted verbatim into a finding.
|
|
39
|
+
|
|
40
|
+
---
|
|
41
|
+
|
|
42
|
+
## What it does
|
|
43
|
+
|
|
44
|
+
1. Splits every discovered documentation file — the default corpus is
|
|
45
|
+
USER-FACING docs only: `docs/**`, the root `README*` (never `CHANGELOG*`),
|
|
46
|
+
and gdwiki pages (`.metaproject/wiki/**`) — into sections by heading,
|
|
47
|
+
deterministically (`src/review/jev-docs.ts`'s `extractDocSections` — no
|
|
48
|
+
model call). `--include <glob>` (repeatable) widens the corpus back out
|
|
49
|
+
(skills, rules, anything project-specific); see Input Contract.
|
|
50
|
+
2. Links each section to code deterministically — an explicit repo-relative
|
|
51
|
+
path it quotes, a backtick-quoted symbol that appears verbatim in one of
|
|
52
|
+
the diff's own hunks, or a `keryx <verb>` invocation whose command file the
|
|
53
|
+
diff changed. Only sections linked to code the diff CHANGES, and that the
|
|
54
|
+
diff does NOT itself edit, are candidates.
|
|
55
|
+
3. Ranks linked sections by link strength — path mention > symbol mention >
|
|
56
|
+
verb mention; more distinct links to changed code within the same kind
|
|
57
|
+
rank higher — and asks Jev exactly ONE `noul` question per selected
|
|
58
|
+
section, up to `--max-calls` (default 30) and 8 per doc file, batched
|
|
59
|
+
under the vendor's 64k token budget, with the ranking basis and every drop
|
|
60
|
+
reported (`selection`), never silent, never alphabetical.
|
|
61
|
+
4. Separately, and with NO Jev call: a CLI flag that disappears from a
|
|
62
|
+
changed file's diff (present on a removed line, absent from every added
|
|
63
|
+
line of the same file) and is still mentioned by an untouched doc section
|
|
64
|
+
is flagged deterministically.
|
|
65
|
+
5. Synthesizes findings deterministically: `problem` names the section
|
|
66
|
+
(file + heading + line) and the code change that likely outdated it,
|
|
67
|
+
`suggested_fix` names the update target, `evidence` carries Jev's
|
|
68
|
+
probability (or, for the flag check, the deterministic reasoning).
|
|
69
|
+
Severity is always `minor`.
|
|
70
|
+
|
|
71
|
+
---
|
|
72
|
+
|
|
73
|
+
## Input Contract
|
|
74
|
+
|
|
75
|
+
Not dispatched with a prompt — invoked as a CLI command:
|
|
76
|
+
|
|
77
|
+
```text
|
|
78
|
+
keryx review jev-docs (--diff <ref> | --pr <n>) [--max-calls <n>] [--threshold <0..1>]
|
|
79
|
+
[--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
|
|
80
|
+
[--fixtures <dir>] [--include <glob>]... [--json]
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
`review-orchestrator` passes `--diff <ref>` (or `--pr <n>`) matching the same
|
|
84
|
+
target every other reviewer's dispatch checks.
|
|
85
|
+
|
|
86
|
+
---
|
|
87
|
+
|
|
88
|
+
## Opt-in and privacy
|
|
89
|
+
|
|
90
|
+
Refuses before any doc read or network call unless
|
|
91
|
+
`.metaproject/tasks.config.json` declares:
|
|
92
|
+
|
|
93
|
+
```json
|
|
94
|
+
{ "review": { "jev": { "docs": true } } }
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
Every doc-section excerpt and hunk sent to Jev is redacted first through
|
|
98
|
+
`src/security/service.ts` — the same floor `review conform`/`review jev-rules`
|
|
99
|
+
already apply. No cache or output file stores a credential.
|
|
100
|
+
|
|
101
|
+
---
|
|
102
|
+
|
|
103
|
+
## Output Contract
|
|
104
|
+
|
|
105
|
+
Emits a `REVIEW_RESULT`-shaped object matching
|
|
106
|
+
`.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
|
|
107
|
+
— `status`, `reviewer: "review-jev-docs"`, `summary`, `findings`, `stats` —
|
|
108
|
+
under `--json`. Its findings merge into the consolidated array exactly like
|
|
109
|
+
any other reviewer's: same Quality Gate, same dedup, same Wave C
|
|
110
|
+
verification.
|
|
111
|
+
|
|
112
|
+
### Class scope — required for `blocker` and `major`
|
|
113
|
+
|
|
114
|
+
Every finding this reviewer emits is `minor` (`docsFindingStats` caps it by
|
|
115
|
+
construction), so `class_scope` never applies here — `blocker`/`major`
|
|
116
|
+
findings, and the `class_scope` contract that comes with them, belong to the
|
|
117
|
+
reviewers named in `review-orchestrator/SKILL.md`'s own Finding Format
|
|
118
|
+
section.
|
|
119
|
+
|
|
120
|
+
---
|
|
121
|
+
|
|
122
|
+
## Orchestrator integration
|
|
123
|
+
|
|
124
|
+
- `keryx review reviewers --json` lists it under `bundled` with
|
|
125
|
+
`"engine": "jev"`.
|
|
126
|
+
- `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
|
|
127
|
+
convention reviewers, when the opt-in is on and a Jev/OpenRouter credential
|
|
128
|
+
resolves (`resolveJevApiKeyResolution`). When either is false it is
|
|
129
|
+
**skipped with the stated reason**, recorded in `Skipped reviewers` —
|
|
130
|
+
never silently absent.
|
|
131
|
+
- It is **ADDITIONAL**: it never replaces any documentation-adjacent finding
|
|
132
|
+
another reviewer already raises; the dedup pass merges rather than
|
|
133
|
+
double-counts.
|
|
134
|
+
|
|
135
|
+
---
|
|
136
|
+
|
|
137
|
+
## Red Flags
|
|
138
|
+
|
|
139
|
+
| Rationalization | Why it is wrong |
|
|
140
|
+
|---|---|
|
|
141
|
+
| "Jev said 0.9, so the doc is definitely wrong now" | A `noul` score is a probability, not a verdict — read the linked hunk yourself before editing the doc |
|
|
142
|
+
| "No findings means the docs are current" | The default corpus is `docs/**`/`README*`/gdwiki only — a project's `.metaproject/skills/**`/`rules/**` need `--include` to be scored at all; `--max-calls`/8-per-file also bound what gets scored (`selection.droppedSections`) |
|
|
143
|
+
| "This section wasn't linked, so it's fine" | Linking is deliberately narrow (explicit paths/symbols/verbs) — a section describing behaviour with no explicit code reference is out of this reviewer's reach by design, not proven current |
|
|
144
|
+
|
|
145
|
+
---
|
|
146
|
+
|
|
147
|
+
## Verification
|
|
148
|
+
|
|
149
|
+
Before trusting a run's findings:
|
|
150
|
+
|
|
151
|
+
1. Read `selection.droppedSections` in the `--json` output — a capped run
|
|
152
|
+
covered less than every linked section.
|
|
153
|
+
2. Spot-check a finding against the cited hunk: does the section's own text
|
|
154
|
+
actually conflict with what the diff changed?
|
|
155
|
+
3. Confirm every finding carries `reviewer: "review-jev-docs"` and a
|
|
156
|
+
`dedupe_key` — both are required for the orchestrator's Quality Gate and
|
|
157
|
+
Wave C verification to route it correctly.
|
|
158
|
+
|
|
159
|
+
---
|
|
160
|
+
|
|
161
|
+
## Iron Laws
|
|
162
|
+
|
|
163
|
+
### Shared laws (every reviewer)
|
|
164
|
+
|
|
165
|
+
1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
|
|
166
|
+
name the input, call, or condition that reaches the code, you have an
|
|
167
|
+
observation, not a finding. Report it as `info` and say what would settle it.
|
|
168
|
+
2. **Never flag the theoretical.** The path you describe must exist in the code
|
|
169
|
+
under review. Do not report a safe API because it could be misused, or a
|
|
170
|
+
pattern because it is often wrong elsewhere.
|
|
171
|
+
3. **One finding per class, not one per occurrence.** When the same shape appears
|
|
172
|
+
at several sites, report it once and list every site. Ten findings that are one
|
|
173
|
+
finding hide the other nine problems.
|
|
174
|
+
|
|
175
|
+
Severity levels are defined once, in `review-orchestrator/SKILL.md` →
|
|
176
|
+
**Severity (canonical)**. This reviewer does not restate them — it only ever
|
|
177
|
+
emits `minor`, capped by construction (`docsFindingStats`), so the
|
|
178
|
+
`major`/`blocker` shapes never apply here.
|
|
179
|
+
|
|
180
|
+
---
|
|
181
|
+
|
|
182
|
+
## Scope Boundaries
|
|
183
|
+
|
|
184
|
+
| Concern | This skill | Use instead |
|
|
185
|
+
|---------|------------|-------------|
|
|
186
|
+
| A doc section linked to changed code likely went stale | YES | — |
|
|
187
|
+
| Documentation quality/completeness unrelated to this diff | NO | a human, or a dedicated docs review |
|
|
188
|
+
| Judging whether a reported finding is real | NO | `review-verifier` |
|
|
189
|
+
| Editing the documentation itself | NO | a human, following `suggested_fix` |
|
|
@@ -0,0 +1,190 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: review-jev-risk
|
|
3
|
+
model_tier: light
|
|
4
|
+
description: |
|
|
5
|
+
Use when: an ADDITIONAL, machine-scored RISK MAP is wanted over every changed hunk —
|
|
6
|
+
never replacing any other reviewer. Dispatched by review-orchestrator in Wave B, via
|
|
7
|
+
the CLI (`keryx review jev-risk`), when `review.jev.risk: true` in
|
|
8
|
+
.metaproject/tasks.config.json and a Jev/OpenRouter credential is resolvable. It is not
|
|
9
|
+
an LLM sub-agent: there is nothing to dispatch through a platform-native agent
|
|
10
|
+
mechanism, and no prose is generated by a model — Jev only answers one `noul`
|
|
11
|
+
probability per (hunk, risk dimension) pair; keryx writes every word of every finding.
|
|
12
|
+
NOT for: judgement calls a risk dimension does not state (style, architecture beyond
|
|
13
|
+
risk-flagging — those stay with the reviewers that already cover them), and not a
|
|
14
|
+
substitute for a human reviewer reading the ranked map it produces.
|
|
15
|
+
triggers:
|
|
16
|
+
- "jev risk"
|
|
17
|
+
- "risk map"
|
|
18
|
+
- "review --jev-risk"
|
|
19
|
+
metadata:
|
|
20
|
+
author: "MrCipherSmith"
|
|
21
|
+
version: "1.0.0"
|
|
22
|
+
category: "review"
|
|
23
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
24
|
+
engine: "jev"
|
|
25
|
+
origin: "authored"
|
|
26
|
+
license: "MIT"
|
|
27
|
+
---
|
|
28
|
+
|
|
29
|
+
# Review — Jev Risk (deterministic risk map of the diff)
|
|
30
|
+
|
|
31
|
+
An ADDITIONAL orchestrator reviewer, flow 332. It is a **keryx program**, not
|
|
32
|
+
an LLM sub-agent: `review-orchestrator` never dispatches a platform-native
|
|
33
|
+
agent for it, it runs `keryx review jev-risk` and reads the `--json` output.
|
|
34
|
+
Every finding is composed by keryx; Jev ("System One" on OpenRouter) supplies
|
|
35
|
+
only five `noul` probabilities per hunk — one per risk dimension — never
|
|
36
|
+
prose.
|
|
37
|
+
|
|
38
|
+
---
|
|
39
|
+
|
|
40
|
+
## What it does
|
|
41
|
+
|
|
42
|
+
1. Takes every changed hunk from `keryx review scope`/`buildReviewScope`
|
|
43
|
+
(mechanical bulk already dropped).
|
|
44
|
+
2. Computes deterministic facts FIRST, per hunk: a path class (auth/
|
|
45
|
+
permissions, crypto, migrations, schema, public API, config, concurrency
|
|
46
|
+
primitives, IO), the exported symbols it touches, its changed-line count,
|
|
47
|
+
and whether a test file elsewhere in the diff touches the same module.
|
|
48
|
+
3. Asks Jev **one `noul` per risk dimension** per hunk — security-sensitive,
|
|
49
|
+
data/migration, public-API/contract change, concurrency, error-handling —
|
|
50
|
+
batched under the vendor's 64k token budget, capped at `--max-calls`
|
|
51
|
+
(default 150, counted as hunk x dimension pairs), with every drop
|
|
52
|
+
reported.
|
|
53
|
+
4. Ranks hunks by combined risk (the MAX across its five dimensions — one
|
|
54
|
+
high-risk dimension is enough to draw attention) so a human reviewer knows
|
|
55
|
+
where to look first.
|
|
56
|
+
5. Emits a finding ONLY for a hunk above threshold (default 0.7) AND with no
|
|
57
|
+
test touched nearby in the same diff (a fact) — severity capped at
|
|
58
|
+
`info`/`minor`, never higher: this flags attention, it does not assert a
|
|
59
|
+
defect.
|
|
60
|
+
6. Additionally emits a **routing hint**: hunks above threshold with a
|
|
61
|
+
security or concurrency dimension are listed with a suggested reviewer
|
|
62
|
+
(`review-security-code`/`review-highload`) for the orchestrator to
|
|
63
|
+
consider dispatching — see "Orchestrator integration" below.
|
|
64
|
+
|
|
65
|
+
---
|
|
66
|
+
|
|
67
|
+
## Input Contract
|
|
68
|
+
|
|
69
|
+
Not dispatched with a prompt — invoked as a CLI command:
|
|
70
|
+
|
|
71
|
+
```text
|
|
72
|
+
keryx review jev-risk (--diff <ref> | --pr <n> | --scope <scope.json>)
|
|
73
|
+
[--max-calls <n>] [--threshold <0..1>]
|
|
74
|
+
[--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
|
|
75
|
+
[--fixtures <dir>] [--json]
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
`review-orchestrator` passes `--scope <scope.json>` — the same `keryx review
|
|
79
|
+
scope --json` file every other reviewer's dispatch already reads — so this
|
|
80
|
+
reviewer checks exactly the same hunks as everyone else in the round.
|
|
81
|
+
|
|
82
|
+
---
|
|
83
|
+
|
|
84
|
+
## Opt-in and privacy
|
|
85
|
+
|
|
86
|
+
Refuses before any read or network call unless `.metaproject/tasks.config.json`
|
|
87
|
+
declares:
|
|
88
|
+
|
|
89
|
+
```json
|
|
90
|
+
{ "review": { "jev": { "risk": true } } }
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
Every hunk sent to Jev is redacted first through `src/security/service.ts` —
|
|
94
|
+
the same floor `review conform`/`review ci-triage`/`review jev-rules` already
|
|
95
|
+
apply. No cache or output file stores a credential.
|
|
96
|
+
|
|
97
|
+
---
|
|
98
|
+
|
|
99
|
+
## Output Contract
|
|
100
|
+
|
|
101
|
+
Emits a `REVIEW_RESULT`-shaped object matching
|
|
102
|
+
`.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
|
|
103
|
+
— `status`, `reviewer: "review-jev-risk"`, `summary`, `findings`, `stats` —
|
|
104
|
+
under `--json`, plus two orchestrator-facing extras: `ranked` (every scored
|
|
105
|
+
hunk, highest risk first) and `routingHints` (see above). Its findings merge
|
|
106
|
+
into the consolidated array exactly like any other reviewer's: same Quality
|
|
107
|
+
Gate, same dedup, same Wave C verification.
|
|
108
|
+
|
|
109
|
+
---
|
|
110
|
+
|
|
111
|
+
## Orchestrator integration
|
|
112
|
+
|
|
113
|
+
- `keryx review reviewers --json` lists it under `bundled` with
|
|
114
|
+
`"engine": "jev"`.
|
|
115
|
+
- `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
|
|
116
|
+
convention reviewers and flow 330's `review-jev-rules`, when the opt-in is
|
|
117
|
+
on and a Jev/OpenRouter credential resolves. When either is false it is
|
|
118
|
+
**skipped with the stated reason**, recorded in `Skipped reviewers` — never
|
|
119
|
+
silently absent.
|
|
120
|
+
- It is **ADDITIONAL**: it never replaces `review-security-code`,
|
|
121
|
+
`review-highload`, or any other pass.
|
|
122
|
+
- **Routing hint**: read `routingHints` from its `--json` output. Each entry
|
|
123
|
+
names a hunk location, the dimension that crossed threshold
|
|
124
|
+
(`security`/`concurrency`) and a `suggestedReviewer`
|
|
125
|
+
(`review-security-code`/`review-highload`). When that reviewer was not
|
|
126
|
+
already selected for the round, dispatch it too — the hint is advisory,
|
|
127
|
+
not a hard requirement, and the orchestrator's own path-based selection
|
|
128
|
+
always takes precedence when it already covers the same hunk.
|
|
129
|
+
|
|
130
|
+
---
|
|
131
|
+
|
|
132
|
+
### Shared laws (every reviewer)
|
|
133
|
+
|
|
134
|
+
1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
|
|
135
|
+
name the input, call, or condition that reaches the code, you have an
|
|
136
|
+
observation, not a finding. Report it as `info` and say what would settle it.
|
|
137
|
+
2. **Never flag the theoretical.** The path you describe must exist in the code
|
|
138
|
+
under review. Do not report a safe API because it could be misused, or a
|
|
139
|
+
pattern because it is often wrong elsewhere.
|
|
140
|
+
3. **One finding per class, not one per occurrence.** When the same shape appears
|
|
141
|
+
at several sites, report it once and list every site. Ten findings that are one
|
|
142
|
+
finding hide the other nine problems.
|
|
143
|
+
|
|
144
|
+
This reviewer satisfies all three by construction: `evidence` always names the
|
|
145
|
+
hunk location, quotes the changed line, and states every dimension's
|
|
146
|
+
probability (never an unreproducible claim); a finding is synthesized only
|
|
147
|
+
against the hunk actually scored, never a hypothetical pattern; and each
|
|
148
|
+
hunk is a distinct site with its own probabilities, so there is no repeated
|
|
149
|
+
class across sites for this reviewer to collapse — the routing hint follows
|
|
150
|
+
the same discipline, naming the exact hunk and dimension that crossed
|
|
151
|
+
threshold rather than a general warning.
|
|
152
|
+
|
|
153
|
+
---
|
|
154
|
+
|
|
155
|
+
## Red Flags
|
|
156
|
+
|
|
157
|
+
| Rationalization | Why it is wrong |
|
|
158
|
+
|---|---|
|
|
159
|
+
| "Jev said 0.9, so this hunk is definitely dangerous" | A `noul` score is a probability across one dimension, not a verdict — severity stays capped at `info`/`minor` precisely because Jev alone is not authoritative |
|
|
160
|
+
| "No nearby test means nobody tested this at all" | `hasNearbyTest` requires the nearby test's OWN diff text to mention a touched symbol or import the hunk's module (`testHunkEvidence`), tightened after a live-check false negative — still a regex heuristic, not proof of coverage through an indirection it cannot see |
|
|
161
|
+
| "This hunk touches a .md file, so 'no nearby test' is a meaningful finding" | Docs (`.md`/`.txt`) hunks are excluded before scoring (`isNonCodeHunk`) after a live check where prose scored `public-api` — this cannot happen; `selection.hunksNotCode` reports the count |
|
|
162
|
+
| "No findings means the diff is low risk" | `--max-calls` bounds how many hunks are scored; a capped run reports `selection.hunksSkipped` for exactly this reason — read the selection stats before treating silence as clean |
|
|
163
|
+
|
|
164
|
+
---
|
|
165
|
+
|
|
166
|
+
## Verification
|
|
167
|
+
|
|
168
|
+
Before trusting a run's findings:
|
|
169
|
+
|
|
170
|
+
1. Read `selection.hunksSkipped` in the `--json` output — a capped run
|
|
171
|
+
covered fewer than "every retained hunk" and the report should say so.
|
|
172
|
+
2. Spot-check a handful of `ranked` entries against the actual hunk: does the
|
|
173
|
+
file/line range and the top dimension make sense for what actually
|
|
174
|
+
changed there? Docs (`.md`/`.txt`) hunks never reach `ranked` at all —
|
|
175
|
+
`selection.hunksNotCode` should account for every one of them in the
|
|
176
|
+
diff.
|
|
177
|
+
3. Confirm every finding carries `reviewer: "review-jev-risk"` and a
|
|
178
|
+
`dedupe_key` — both are required for the orchestrator's Quality Gate and
|
|
179
|
+
Wave C verification to route it correctly.
|
|
180
|
+
|
|
181
|
+
---
|
|
182
|
+
|
|
183
|
+
## Scope Boundaries
|
|
184
|
+
|
|
185
|
+
| Concern | This skill | Use instead |
|
|
186
|
+
|---------|------------|-------------|
|
|
187
|
+
| A ranked risk map of the diff's hunks | YES | — |
|
|
188
|
+
| An actual security/concurrency finding | NO (routing hint only) | `review-security-code` / `review-highload` |
|
|
189
|
+
| Judging whether a reported finding is real | NO | `review-verifier` |
|
|
190
|
+
| Which user scenarios changed | NO | `review-jev-scenarios` |
|
|
@@ -0,0 +1,267 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: review-jev-rules
|
|
3
|
+
model_tier: light
|
|
4
|
+
description: |
|
|
5
|
+
Use when: an ADDITIONAL, machine-scored pass is wanted over every changed hunk against
|
|
6
|
+
every applicable project rule clause — never replacing any other reviewer. Dispatched
|
|
7
|
+
by review-orchestrator in Wave B, via the CLI (`keryx review jev-rules`), when
|
|
8
|
+
`review.jev.rules: true` in .metaproject/tasks.config.json and a Jev/OpenRouter
|
|
9
|
+
credential is resolvable. It is not an LLM sub-agent: there is nothing to dispatch
|
|
10
|
+
through a platform-native agent mechanism, and no prose is generated by a model —
|
|
11
|
+
Jev only answers a `noul` violation probability per (hunk, rule clause) pair; keryx
|
|
12
|
+
writes every word of every finding.
|
|
13
|
+
NOT for: judgement calls a rule clause does not state (logic bugs, architecture,
|
|
14
|
+
security, performance — those stay with the reviewers that already cover them), and
|
|
15
|
+
not a substitute for reading the rules yourself. A finding here says a hunk likely
|
|
16
|
+
contradicts a clause of a rule this project already wrote down; it says nothing about
|
|
17
|
+
code that violates no documented rule.
|
|
18
|
+
triggers:
|
|
19
|
+
- "jev rules"
|
|
20
|
+
- "rule check"
|
|
21
|
+
- "review --jev-rules"
|
|
22
|
+
metadata:
|
|
23
|
+
author: "MrCipherSmith"
|
|
24
|
+
version: "1.0.0"
|
|
25
|
+
category: "review"
|
|
26
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
27
|
+
engine: "jev"
|
|
28
|
+
origin: "authored"
|
|
29
|
+
license: "MIT"
|
|
30
|
+
---
|
|
31
|
+
|
|
32
|
+
# Review — Jev Rules (deterministic project-rule conformance)
|
|
33
|
+
|
|
34
|
+
An ADDITIONAL orchestrator reviewer, flow 330. It is a **keryx program**, not an
|
|
35
|
+
LLM sub-agent: `review-orchestrator` never dispatches a platform-native agent
|
|
36
|
+
for it, it runs `keryx review jev-rules` and reads the `--json` output. Every
|
|
37
|
+
finding is composed by keryx; Jev ("System One" on OpenRouter) supplies only a
|
|
38
|
+
`noul` probability per `(hunk, rule clause)` pair — it never writes prose, and
|
|
39
|
+
nothing it returns is quoted verbatim into a finding except the clause text
|
|
40
|
+
keryx already had.
|
|
41
|
+
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
## What it does
|
|
45
|
+
|
|
46
|
+
1. Discovers rule sources deterministically (`src/review/jev-rules.ts`'s
|
|
47
|
+
module header names the exact scope): every file under
|
|
48
|
+
`.metaproject/rules/**` and `rules/**`, every project-skill or installed
|
|
49
|
+
gdskill whose name or `metadata.category` marks it a coding convention, and
|
|
50
|
+
anything named by `--rules <paths>`. **`--rules` bypass is scoped to ONE
|
|
51
|
+
FILE**: an entry naming a single document bypasses the category filter
|
|
52
|
+
below outright; an entry naming a DIRECTORY is walked and its files are
|
|
53
|
+
filtered exactly like auto-discovery — asking for a whole directory is not
|
|
54
|
+
the same as naming one document, and the whole point of the filter is lost
|
|
55
|
+
if it is.
|
|
56
|
+
2. **Category filter, before clause extraction.** Each discovered source
|
|
57
|
+
(auto-discovered, or a file found by walking a `--rules` DIRECTORY — never
|
|
58
|
+
a `--rules` entry naming that ONE file explicitly) is classified
|
|
59
|
+
`code`/`process`/`docs` — explicit frontmatter first (`applies_to:
|
|
60
|
+
code|process|docs`, or `metadata.category`), else a documented
|
|
61
|
+
filename/title heuristic (`PROCESS_RULE_HEURISTIC_TERMS`: `commit`, `git`,
|
|
62
|
+
`tdd`, `workflow`, `definition-of-done`, `documentation`, `requirements`,
|
|
63
|
+
`plan`, `prompting`, `subagent`, `skill`, `jobs`, `orchestrat`,
|
|
64
|
+
`review-process`, `release`). A `process`/`docs` source never reaches
|
|
65
|
+
clause extraction at all — excluded, and reported with its reason under
|
|
66
|
+
`excludedSources`.
|
|
67
|
+
3. Splits each remaining rule document into clauses with
|
|
68
|
+
`extractReferenceClauses` (the same deterministic splitter `review
|
|
69
|
+
conform` uses), drops authoring-template scaffolding
|
|
70
|
+
(`isPlaceholderClauseText`: an unfilled `[x] <criterion> — verified by
|
|
71
|
+
<test>` checklist line, text dominated by `<...>` placeholders, or a bare
|
|
72
|
+
code-fence line — never a real clause, dropped BEFORE tagging so it costs
|
|
73
|
+
nothing), then TAGS every remaining clause `state_kind:
|
|
74
|
+
"pr"|"report"|"hunk"` + `checkable` by REUSING `review conform`'s own
|
|
75
|
+
`applyClauseTags`/`buildClauseTagQuestions`/`clauseTagFromChoice` — an
|
|
76
|
+
explicit `[state:hunk]`/`[not-checkable: ...]` marker on the clause text
|
|
77
|
+
when the rule author wrote one, else one Jev `choice` call per doc's
|
|
78
|
+
untagged clauses (never per clause), cached by the doc's content hash at
|
|
79
|
+
`.metaproject/data/review-jev-rules/clause-tags.json` (a jev-rules-
|
|
80
|
+
specific cache file — `review conform` and `review-jev-rules` never race
|
|
81
|
+
on the same one). Only a clause tagged `state_kind: "hunk"` and
|
|
82
|
+
`checkable` is ever paired against a hunk; anything else is dropped and
|
|
83
|
+
reported under `droppedClauses`.
|
|
84
|
+
4. Decides, per `(rule, changed file)` pair, whether an applicable clause
|
|
85
|
+
applies — by declared path globs (`metadata.paths`) and/or
|
|
86
|
+
`metadata.stack_requires`, failing toward inclusion exactly like `keryx
|
|
87
|
+
review stack`/`scope.ts` already do, THEN by file kind: a docs hunk
|
|
88
|
+
(`.md`/`.mdx`/`.txt`) pairs only with a `docs`-categorised source or a
|
|
89
|
+
clause explicitly tagged `[docs-applicable]`; a code source never pairs
|
|
90
|
+
with a docs hunk, and vice versa (`clauseFileKindApplicability`).
|
|
91
|
+
5. Takes every changed hunk from `keryx review scope` (mechanical bulk
|
|
92
|
+
already dropped) and asks Jev, per applicable `(hunk, hunk-checkable
|
|
93
|
+
clause)` pair, **one question**: "does this hunk VIOLATE this clause?" —
|
|
94
|
+
batched under the vendor's 64k token budget, capped at `--max-calls`
|
|
95
|
+
(default 150). The budget is spread FAIRLY: hunks are ranked code, then
|
|
96
|
+
tests, then docs, and pairs are allocated round-robin across hunks rather
|
|
97
|
+
than draining the cap on the first hunk in diff order — every selection,
|
|
98
|
+
every drop, and per-hunk coverage (`selection.hunkCoverage`, how many
|
|
99
|
+
hunks were reached vs. never reached) is reported, never silent.
|
|
100
|
+
6. Synthesizes findings deterministically: `problem` quotes the clause,
|
|
101
|
+
`impact` is the rule's own stated rationale (if it has a `## Rationale`/
|
|
102
|
+
`## Why` section) or a fixed template, `suggested_fix` names the clause to
|
|
103
|
+
bring the hunk in line with, `evidence` is the hunk location(s) + Jev's
|
|
104
|
+
probability. Severity is capped at `minor` unless the rule itself declares
|
|
105
|
+
a higher one (`[severity: major]` on the clause). One finding per
|
|
106
|
+
`(clause, file)`, deduped across every hunk of that file with a hunk list.
|
|
107
|
+
|
|
108
|
+
**Why steps 2-5's filters exist:** a live check of this repository's own
|
|
109
|
+
41-doc `.metaproject/rules/**` corpus against a merged PR hand-labelled ~1/10
|
|
110
|
+
findings correct. Two causes, both fixed here: (a) process/agent-behaviour
|
|
111
|
+
rules (commit-message formatting, TDD workflow, an agent's own prompting
|
|
112
|
+
standard) paired against a code hunk they were never meant to describe,
|
|
113
|
+
because the corpus declares no `metadata.paths`/`stack_requires` and
|
|
114
|
+
applicability fails open to "applies everywhere" — the category and
|
|
115
|
+
file-kind filters (steps 2 and 4); (b) a re-measurement then found the
|
|
116
|
+
`--rules` bypass defeating the category filter for an entire directory, and
|
|
117
|
+
the whole `--max-calls` budget landing on a single docs hunk while no code
|
|
118
|
+
hunk was ever scored — the scoped bypass and fair round-robin budget (steps
|
|
119
|
+
1 and 5). All four are cheap, deterministic, and applied BEFORE any
|
|
120
|
+
violation-scoring Jev call, so they cut cost as well as noise. See
|
|
121
|
+
`.metaproject/flows/330-*/journal.md` for the full before/after numbers.
|
|
122
|
+
|
|
123
|
+
---
|
|
124
|
+
|
|
125
|
+
## Input Contract
|
|
126
|
+
|
|
127
|
+
Not dispatched with a prompt — invoked as a CLI command:
|
|
128
|
+
|
|
129
|
+
```text
|
|
130
|
+
keryx review jev-rules (--diff <ref> | --pr <n> | --scope <scope.json>)
|
|
131
|
+
[--rules <paths>] [--max-calls <n>] [--threshold <0..1>]
|
|
132
|
+
[--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
|
|
133
|
+
[--fixtures <dir>] [--json]
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
`review-orchestrator` passes `--scope <scope.json>` — the same `keryx review
|
|
137
|
+
scope --json` file every other reviewer's dispatch already reads — so this
|
|
138
|
+
reviewer checks exactly the same hunks, with the same mechanical-bulk drops,
|
|
139
|
+
as everyone else in the round.
|
|
140
|
+
|
|
141
|
+
---
|
|
142
|
+
|
|
143
|
+
## Opt-in and privacy
|
|
144
|
+
|
|
145
|
+
Refuses before any read or network call unless `.metaproject/tasks.config.json`
|
|
146
|
+
declares:
|
|
147
|
+
|
|
148
|
+
```json
|
|
149
|
+
{ "review": { "jev": { "rules": true } } }
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
Every hunk, every rule-clause sent for violation scoring, and every clause
|
|
153
|
+
sent for tagging is redacted first through `src/security/service.ts` — the
|
|
154
|
+
same floor `review conform` and `review ci-triage` already apply. No cache or
|
|
155
|
+
output file stores a credential. Two caches, both gitignored and mode `0600`
|
|
156
|
+
under `.metaproject/data/review-jev-rules/`:
|
|
157
|
+
|
|
158
|
+
- `violation-cache.json`, keyed on the clause text and hunk location, so a
|
|
159
|
+
re-run over unchanged hunks and unchanged rules costs zero additional Jev
|
|
160
|
+
calls.
|
|
161
|
+
- `clause-tags.json`, keyed on the rule doc's own content hash — REUSING
|
|
162
|
+
`review conform`'s own clause-tagging cache module
|
|
163
|
+
(`src/review/conform-tag-cache.ts`) at this jev-rules-specific path, so a
|
|
164
|
+
re-run against a DIFFERENT diff but the SAME rule corpus costs zero
|
|
165
|
+
additional tagging calls even when the hunks (and so the violation cache)
|
|
166
|
+
miss.
|
|
167
|
+
|
|
168
|
+
---
|
|
169
|
+
|
|
170
|
+
## Output Contract
|
|
171
|
+
|
|
172
|
+
Emits a `REVIEW_RESULT`-shaped object matching
|
|
173
|
+
`.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
|
|
174
|
+
— `status`, `reviewer: "review-jev-rules"`, `summary`, `findings`, `stats`,
|
|
175
|
+
plus `tokens` (with `taggingCalls`/`violationCalls` counted separately),
|
|
176
|
+
`selection` (`droppedClauses`'s count, and `hunkCoverage` — one
|
|
177
|
+
`{path, applicablePairs, selectedPairs}` per hunk with something to check,
|
|
178
|
+
plus `hunksWithPairs`/`hunksReached`/`hunksNeverReached`), `ruleSources`
|
|
179
|
+
(each with its `category`), `excludedSources` (category-filtered sources
|
|
180
|
+
with their reason), `droppedClauses` (tag-filtered `(ruleId, clauseId)`
|
|
181
|
+
pairs with their reason), and `droppedPlaceholderClauses`
|
|
182
|
+
(authoring-template clauses dropped before tagging, with their reason) —
|
|
183
|
+
under `--json`. `summary` names how many hunks the budget never reached in
|
|
184
|
+
plain text. Its findings merge into the consolidated array exactly like any
|
|
185
|
+
other reviewer's: same Quality Gate, same dedup, same Wave C verification.
|
|
186
|
+
|
|
187
|
+
---
|
|
188
|
+
|
|
189
|
+
## Orchestrator integration
|
|
190
|
+
|
|
191
|
+
- `keryx review reviewers --json` lists it under `bundled` with
|
|
192
|
+
`"engine": "jev"`.
|
|
193
|
+
- `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
|
|
194
|
+
convention reviewers, when the opt-in is on and a Jev/OpenRouter credential
|
|
195
|
+
resolves (`resolveJevApiKeyResolution`). When either is false it is
|
|
196
|
+
**skipped with the stated reason**, recorded in `Skipped reviewers` —
|
|
197
|
+
never silently absent.
|
|
198
|
+
- It is **ADDITIONAL**: it never replaces `review-style`, the convention
|
|
199
|
+
reviewers, or any other pass. A rule clause this reviewer flags may also be
|
|
200
|
+
the exact concern a domain reviewer raises independently; the dedup pass
|
|
201
|
+
(by `dedupe_key`/`(file, quote, problem)`) merges the two rather than
|
|
202
|
+
double-counting.
|
|
203
|
+
|
|
204
|
+
---
|
|
205
|
+
|
|
206
|
+
### Shared laws (every reviewer)
|
|
207
|
+
|
|
208
|
+
1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
|
|
209
|
+
name the input, call, or condition that reaches the code, you have an
|
|
210
|
+
observation, not a finding. Report it as `info` and say what would settle it.
|
|
211
|
+
2. **Never flag the theoretical.** The path you describe must exist in the code
|
|
212
|
+
under review. Do not report a safe API because it could be misused, or a
|
|
213
|
+
pattern because it is often wrong elsewhere.
|
|
214
|
+
3. **One finding per class, not one per occurrence.** When the same shape appears
|
|
215
|
+
at several sites, report it once and list every site. Ten findings that are one
|
|
216
|
+
finding hide the other nine problems.
|
|
217
|
+
|
|
218
|
+
This reviewer already satisfies all three by construction: `evidence` always
|
|
219
|
+
names the hunk and quotes it (never an unreproducible claim), a finding is
|
|
220
|
+
only synthesized when the hunk itself is checked against the clause (never a
|
|
221
|
+
theoretical pattern), and `synthesizeFindingsFromViolations` dedupes by
|
|
222
|
+
`(clause, file)` with every hunk listed, never one finding per hunk.
|
|
223
|
+
|
|
224
|
+
---
|
|
225
|
+
|
|
226
|
+
## Red Flags
|
|
227
|
+
|
|
228
|
+
| Rationalization | Why it is wrong |
|
|
229
|
+
|---|---|
|
|
230
|
+
| "Jev said 0.9, so this is definitely a violation" | A `noul` score is a probability, not a verdict — severity stays capped at `minor` unless the rule itself declares higher, precisely because Jev alone is not authoritative |
|
|
231
|
+
| "This rule clause doesn't really describe code, but the pair matched anyway" | Applicability defaults to "applies everywhere" when a rule declares no path/stack restriction — the clause-kind filter (step 3), the category filter (step 2), and the file-kind gate (step 4) catch most of this before scoring, but a `state_kind: "hunk"` clause on an off-topic rule can still pass through; that is a property of the rule corpus, not a defect in this reviewer |
|
|
232
|
+
| "No findings means the diff is clean" | `--max-calls` bounds how many pairs are checked; a capped run reports `selection.droppedPairs`/`hunksNeverReached` for exactly this reason — read the selection stats before treating silence as clean |
|
|
233
|
+
| "A rule wasn't checked and I don't know why" | Read `excludedSources` (category filter, at discovery), `droppedClauses` (tag filter, per clause), and `droppedPlaceholderClauses` (template scaffolding, per clause) before assuming a rule was silently skipped — each names the exact reason |
|
|
234
|
+
| "I passed `--rules` at a whole rule directory, so every file in it was checked" | Only a `--rules` entry naming ONE file bypasses the category filter; a directory's files are filtered exactly like auto-discovery — check `excludedSources` |
|
|
235
|
+
| "I'll skip the opt-in check since I trust this project" | The opt-in and credential gates exist because hunk and rule-clause text leaves the machine; skipping them is skipping consent, not a shortcut |
|
|
236
|
+
|
|
237
|
+
---
|
|
238
|
+
|
|
239
|
+
## Verification
|
|
240
|
+
|
|
241
|
+
Before trusting a run's findings:
|
|
242
|
+
|
|
243
|
+
1. Read `selection.droppedPairs`/`selection.notApplicable`/
|
|
244
|
+
`selection.droppedClauses`/`selection.hunkCoverage` in the `--json`
|
|
245
|
+
output — a capped, narrowly-applicable, or tag-filtered run covered less
|
|
246
|
+
than "every changed hunk against every applicable clause", and
|
|
247
|
+
`hunksNeverReached` says exactly how many hunks the budget never got to.
|
|
248
|
+
2. Read `excludedSources` — a process/meta rule doc excluded by the category
|
|
249
|
+
filter is reported by name and reason, never silently absent from
|
|
250
|
+
`ruleSources`.
|
|
251
|
+
3. Spot-check a handful of findings against the actual hunk: does the quoted
|
|
252
|
+
line and the cited clause text support the claim? Jev's probability is not
|
|
253
|
+
evidence a human can inspect; the quote and the clause text are.
|
|
254
|
+
4. Confirm every finding carries `reviewer: "review-jev-rules"` and a
|
|
255
|
+
`dedupe_key` — both are required for the orchestrator's Quality Gate and
|
|
256
|
+
Wave C verification to route it correctly.
|
|
257
|
+
|
|
258
|
+
---
|
|
259
|
+
|
|
260
|
+
## Scope Boundaries
|
|
261
|
+
|
|
262
|
+
| Concern | This skill | Use instead |
|
|
263
|
+
|---------|------------|-------------|
|
|
264
|
+
| A hunk contradicts a documented project rule clause | YES | — |
|
|
265
|
+
| Logic bugs, architecture, security, performance not named by any rule | NO | the matching domain reviewer |
|
|
266
|
+
| Judging whether a reported finding is real | NO | `review-verifier` |
|
|
267
|
+
| Writing or maintaining the rules themselves | NO | a human, or `rules/core/*.mdc` directly |
|