@mrciphersmith/keryx 0.3.2 → 0.3.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/cli.js +4634 -2445
- package/dist/core.js +66 -10
- package/package.json +1 -1
- package/src/gdskills/bundled/install-manifest.json +349 -2
- package/src/gdskills/bundled/rules/core/model-selection.mdc +18 -0
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/review-jev-contract/SKILL.md +193 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +81 -21
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +4 -4
- package/src/gdskills/bundled/stacks/c-cpp/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/c-cpp/governance/eval.json +1777 -0
- package/src/gdskills/bundled/stacks/c-cpp/governance/scout.json +31 -0
- package/src/gdskills/bundled/stacks/c-cpp/pack.json +42 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/coding-style.mdc +80 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/patterns.mdc +87 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/security.mdc +90 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/testing.mdc +83 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/SKILL.md +153 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/SKILL.md +132 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/SKILL.md +151 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/SKILL.md +152 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/evals.json +74 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/eval.json +1295 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/scout.json +26 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/pack.json +41 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/patterns.mdc +77 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/security.mdc +144 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/SKILL.md +121 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/evals.json +73 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/SKILL.md +139 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/SKILL.md +147 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/eval.json +865 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/scout.json +16 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/pack.json +46 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/coding-style.mdc +74 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/patterns.mdc +81 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/security.mdc +146 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/testing.mdc +61 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/SKILL.md +151 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/SKILL.md +135 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/evals.json +76 -0
- package/src/gdskills/bundled/stacks/php-laravel/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/php-laravel/governance/eval.json +1829 -0
- package/src/gdskills/bundled/stacks/php-laravel/governance/scout.json +33 -0
- package/src/gdskills/bundled/stacks/php-laravel/pack.json +41 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/coding-style.mdc +82 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/patterns.mdc +80 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/security.mdc +80 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/testing.mdc +82 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/SKILL.md +143 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/SKILL.md +126 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/evals.json +76 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/SKILL.md +140 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/evals.json +75 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/SKILL.md +124 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/evals.json +74 -0
- package/src/gdskills/bundled/stacks/ruby-rails/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/ruby-rails/governance/eval.json +1673 -0
- package/src/gdskills/bundled/stacks/ruby-rails/governance/scout.json +33 -0
- package/src/gdskills/bundled/stacks/ruby-rails/pack.json +42 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/coding-style.mdc +69 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/patterns.mdc +93 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/security.mdc +90 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/testing.mdc +89 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/SKILL.md +143 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/evals.json +73 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/SKILL.md +134 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/evals.json +71 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/SKILL.md +141 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/evals.json +72 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/SKILL.md +125 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/evals.json +72 -0
- package/src/gdskills/bundled/stacks/sql-db/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/sql-db/governance/eval.json +1829 -0
- package/src/gdskills/bundled/stacks/sql-db/governance/scout.json +30 -0
- package/src/gdskills/bundled/stacks/sql-db/pack.json +40 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/coding-style.mdc +69 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/patterns.mdc +134 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/security.mdc +74 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/testing.mdc +83 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/SKILL.md +147 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/evals.json +72 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/SKILL.md +132 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/SKILL.md +153 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/evals.json +77 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/SKILL.md +129 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/evals.json +73 -0
|
@@ -0,0 +1,193 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: review-jev-contract
|
|
3
|
+
model_tier: light
|
|
4
|
+
description: |
|
|
5
|
+
Use when: an ADDITIONAL, machine-scored check of a PR description's own CLAIMS —
|
|
6
|
+
and, when a flow is linked, that flow's frozen acceptance criteria — against the diff
|
|
7
|
+
is wanted, never replacing any other reviewer. Dispatched by review-orchestrator in
|
|
8
|
+
Wave B, via the CLI (`keryx review jev-contract`), when `review.jev.contract: true` in
|
|
9
|
+
.metaproject/tasks.config.json and a Jev/OpenRouter credential is resolvable. It is not
|
|
10
|
+
an LLM sub-agent: there is nothing to dispatch through a platform-native agent
|
|
11
|
+
mechanism, and no prose is generated by a model — Jev only answers one `noul`
|
|
12
|
+
probability per claim; keryx writes every word of every finding. Replaces the
|
|
13
|
+
orchestrator's by-eye Stage 1 "description vs diff" judgement with a scored input when
|
|
14
|
+
the opt-in is on; the by-eye check stays the fallback when it is off.
|
|
15
|
+
NOT for: judging a flow's frozen criteria from scratch (that machinery is flow 328's
|
|
16
|
+
`check-ac.ts`, reused, not duplicated), and not a substitute for a human reviewer
|
|
17
|
+
reading the claim list it produces.
|
|
18
|
+
triggers:
|
|
19
|
+
- "jev contract"
|
|
20
|
+
- "check the PR description against the diff"
|
|
21
|
+
- "review --jev-contract"
|
|
22
|
+
metadata:
|
|
23
|
+
author: "MrCipherSmith"
|
|
24
|
+
version: "1.0.0"
|
|
25
|
+
category: "review"
|
|
26
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
27
|
+
engine: "jev"
|
|
28
|
+
origin: "authored"
|
|
29
|
+
license: "MIT"
|
|
30
|
+
---
|
|
31
|
+
|
|
32
|
+
# Review — Jev Contract (PR-description claims and frozen acceptance criteria vs. the diff)
|
|
33
|
+
|
|
34
|
+
An ADDITIONAL orchestrator reviewer, flow 335. It is a **keryx program**, not
|
|
35
|
+
an LLM sub-agent: `review-orchestrator` never dispatches a platform-native
|
|
36
|
+
agent for it, it runs `keryx review jev-contract` and reads the `--json`
|
|
37
|
+
output. Every finding is composed by keryx; Jev ("System One" on OpenRouter)
|
|
38
|
+
supplies only one `noul` probability per claim — never prose.
|
|
39
|
+
|
|
40
|
+
---
|
|
41
|
+
|
|
42
|
+
## What it does
|
|
43
|
+
|
|
44
|
+
Two independent tracks, both advisory, both merged into one report:
|
|
45
|
+
|
|
46
|
+
1. **Claims** — extracted deterministically from the PR description: every
|
|
47
|
+
bullet/numbered-list item is a claim, plus every prose sentence carrying a
|
|
48
|
+
verb cue (adds, fixes, removes, "does not change", tests). For each claim,
|
|
49
|
+
keryx computes deterministic facts FIRST — named files/symbols/flags
|
|
50
|
+
present in the diff, whether a test file was touched (for a "tests added"
|
|
51
|
+
claim), whether an EXPORTED symbol changed anywhere in the diff (for a "no
|
|
52
|
+
API change" claim) — then asks Jev **one `noul`**: "does the diff support
|
|
53
|
+
this claim?". A claim the facts directly CONTRADICT (e.g. "no API change"
|
|
54
|
+
with an exported symbol touched) is `major`, with `class_scope`,
|
|
55
|
+
regardless of what Jev answered — a contradiction is a fact, not an
|
|
56
|
+
opinion Jev could outvote. A claim scoring below threshold with no
|
|
57
|
+
contradiction is `minor` ("unsupported"). A claim at/above threshold with
|
|
58
|
+
no contradiction produces no finding.
|
|
59
|
+
2. **Criteria** — the linked flow's frozen acceptance criteria, when
|
|
60
|
+
`--flow <id>` names one. This track reuses flow 328's
|
|
61
|
+
`src/flow/check-ac.ts` **wholesale**, through `runCheckAc`
|
|
62
|
+
(`src/commands/flow-check-ac.ts`) — the exact same diff acquisition, Jev
|
|
63
|
+
batching, cache and degrade path `keryx flow check-ac` already uses. No
|
|
64
|
+
criterion-checking logic is reimplemented for this reviewer.
|
|
65
|
+
|
|
66
|
+
---
|
|
67
|
+
|
|
68
|
+
## Input Contract
|
|
69
|
+
|
|
70
|
+
Not dispatched with a prompt — invoked as a CLI command:
|
|
71
|
+
|
|
72
|
+
```text
|
|
73
|
+
keryx review jev-contract (--diff <ref> | --pr <n>) [--flow <id>]
|
|
74
|
+
[--max-calls <n>] [--threshold <0..1>]
|
|
75
|
+
[--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
|
|
76
|
+
[--fixtures <dir>] [--json]
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
`--pr <n>` is required to check claims at all — a `--diff <ref>` target has
|
|
80
|
+
no PR description to extract claims from, so that track reports zero claims
|
|
81
|
+
(honestly, not a refusal). `--flow <id>` is independent of either target: it
|
|
82
|
+
names the flow whose frozen acceptance criteria are checked against the same
|
|
83
|
+
diff, when one is linked.
|
|
84
|
+
|
|
85
|
+
---
|
|
86
|
+
|
|
87
|
+
## Opt-in and privacy
|
|
88
|
+
|
|
89
|
+
Refuses before any read or network call unless `.metaproject/tasks.config.json`
|
|
90
|
+
declares:
|
|
91
|
+
|
|
92
|
+
```json
|
|
93
|
+
{ "review": { "jev": { "contract": true } } }
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
Every claim and every matched diff hunk sent to Jev is redacted first through
|
|
97
|
+
`src/security/service.ts` — the same floor `review conform`/`review
|
|
98
|
+
jev-risk`/`review jev-rules` already apply. No cache or output file stores a
|
|
99
|
+
credential.
|
|
100
|
+
|
|
101
|
+
---
|
|
102
|
+
|
|
103
|
+
## Output Contract
|
|
104
|
+
|
|
105
|
+
Emits a `REVIEW_RESULT`-shaped object matching
|
|
106
|
+
`.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
|
|
107
|
+
— `status`, `reviewer: "review-jev-contract"`, `summary`, `findings`, `stats`
|
|
108
|
+
— under `--json`, plus two extras: `claims` (every extracted claim, its
|
|
109
|
+
intent, and its `noul` score) and `budget` (`--max-calls`, claims scored,
|
|
110
|
+
claims skipped). When `--flow` names a flow, an `acCheck` object is merged in
|
|
111
|
+
(`flowId`, `verdicts`, `jevAsked`, `summary`) — the same shape `keryx flow
|
|
112
|
+
check-ac` reports. Its findings merge into the consolidated array exactly
|
|
113
|
+
like any other reviewer's: same Quality Gate, same dedup, same Wave C
|
|
114
|
+
verification. A `major` finding always carries a schema-valid `class_scope`
|
|
115
|
+
object (`sites`, `enumeration_method`) — never a bare label.
|
|
116
|
+
|
|
117
|
+
---
|
|
118
|
+
|
|
119
|
+
## Orchestrator integration
|
|
120
|
+
|
|
121
|
+
- `keryx review reviewers --json` lists it under `bundled` with
|
|
122
|
+
`"engine": "jev"`.
|
|
123
|
+
- `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
|
|
124
|
+
convention reviewers and the other CLI-engine reviewers, when the opt-in is
|
|
125
|
+
on and a Jev/OpenRouter credential resolves. When either is false it is
|
|
126
|
+
**skipped with the stated reason**, recorded in `Skipped reviewers` — never
|
|
127
|
+
silently absent.
|
|
128
|
+
- It is **ADDITIONAL**: it never replaces any other pass. **When the opt-in is
|
|
129
|
+
on, it REPLACES the orchestrator's by-eye Stage 1 "description vs diff"
|
|
130
|
+
judgement with this scored input** — see `review-orchestrator/SKILL.md`'s
|
|
131
|
+
"This gate owns the description-vs-diff comparison" and `SKILL.detail.md`'s
|
|
132
|
+
"CLI-engine reviewers" section. When the opt-in is off, the by-eye check
|
|
133
|
+
stays the fallback, unchanged.
|
|
134
|
+
|
|
135
|
+
---
|
|
136
|
+
|
|
137
|
+
### Shared laws (every reviewer)
|
|
138
|
+
|
|
139
|
+
1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
|
|
140
|
+
name the input, call, or condition that reaches the code, you have an
|
|
141
|
+
observation, not a finding. Report it as `info` and say what would settle it.
|
|
142
|
+
2. **Never flag the theoretical.** The path you describe must exist in the code
|
|
143
|
+
under review. Do not report a safe API because it could be misused, or a
|
|
144
|
+
pattern because it is often wrong elsewhere.
|
|
145
|
+
3. **One finding per class, not one per occurrence.** When the same shape appears
|
|
146
|
+
at several sites, report it once and list every site. Ten findings that are one
|
|
147
|
+
finding hide the other nine problems.
|
|
148
|
+
|
|
149
|
+
This reviewer satisfies all three by construction: `evidence` always names the
|
|
150
|
+
matched artefacts/facts and the claim text (never an unreproducible claim); a
|
|
151
|
+
finding is synthesized only against a claim actually extracted from the real
|
|
152
|
+
description, never a hypothetical one; and each claim is its own site, so
|
|
153
|
+
there is no repeated class across claims for this reviewer to collapse — a
|
|
154
|
+
`major` finding's `class_scope` still enumerates every matched region, not
|
|
155
|
+
just the first.
|
|
156
|
+
|
|
157
|
+
---
|
|
158
|
+
|
|
159
|
+
## Red Flags
|
|
160
|
+
|
|
161
|
+
| Rationalization | Why it is wrong |
|
|
162
|
+
|---|---|
|
|
163
|
+
| "Jev said 0.9, so the claim is definitely supported" | A `noul` score is a probability the claim is evidenced, not a verdict — a contradiction detected by the deterministic facts overrides it regardless of the score |
|
|
164
|
+
| "The description doesn't mention an API, so 'no change' can't be contradicted" | The contradiction check only runs when the claim's own text IS API-shaped (mentions api/public/export/interface/signature/contract) — a claim like "no change to behavior for CLI callers" is checked for tokens, but the exported-symbol contradiction path is intentionally narrower |
|
|
165
|
+
| "No claims extracted means the description said nothing" | `--diff` targets have no PR description at all — zero claims there is a target-shape fact, not a description-quality one; check `budget.claimsSkipped` too before reading silence as clean |
|
|
166
|
+
| "The AC track failed, so the whole review failed" | `acCheck` degrades to facts-only (never throws) exactly like `keryx flow check-ac` does — a Jev failure on the criteria track is advisory, same as on the claims track |
|
|
167
|
+
|
|
168
|
+
---
|
|
169
|
+
|
|
170
|
+
## Verification
|
|
171
|
+
|
|
172
|
+
Before trusting a run's findings:
|
|
173
|
+
|
|
174
|
+
1. Read `budget.claimsSkipped` in the `--json` output — a capped run checked
|
|
175
|
+
fewer than "every extracted claim" and the report should say so.
|
|
176
|
+
2. Spot-check a `major` finding's `class_scope.sites` against the actual
|
|
177
|
+
diff: does the named region really touch an exported symbol the claim
|
|
178
|
+
said would not change?
|
|
179
|
+
3. Confirm every finding carries `reviewer: "review-jev-contract"` and a
|
|
180
|
+
`dedupe_key` — both are required for the orchestrator's Quality Gate and
|
|
181
|
+
Wave C verification to route it correctly.
|
|
182
|
+
|
|
183
|
+
---
|
|
184
|
+
|
|
185
|
+
## Scope Boundaries
|
|
186
|
+
|
|
187
|
+
| Concern | This skill | Use instead |
|
|
188
|
+
|---------|------------|-------------|
|
|
189
|
+
| Whether the PR description's own claims match the diff | YES | — |
|
|
190
|
+
| Whether a flow's frozen acceptance criteria are likely met | YES (via `--flow`, reusing `check-ac.ts`) | `keryx flow check-ac` directly, for a standalone advisory run |
|
|
191
|
+
| An actual security/concurrency finding | NO | `review-security-code` / `review-highload` |
|
|
192
|
+
| Judging whether a reported finding is real | NO | `review-verifier` |
|
|
193
|
+
| A ranked risk map of the diff's hunks | NO | `review-jev-risk` |
|
|
@@ -47,34 +47,45 @@ review: it is a fact about the review, not about the diff.
|
|
|
47
47
|
## CLI-engine reviewers — dispatched as a command, not a sub-agent
|
|
48
48
|
|
|
49
49
|
`review-jev-rules` (flow 330), `review-jev-risk` and `review-jev-scenarios`
|
|
50
|
-
(both flow 332),
|
|
51
|
-
are ADDITIONAL reviewers, never replacing
|
|
52
|
-
mechanism differs from every reviewer named in
|
|
53
|
-
there is no platform-native agent to invoke,
|
|
54
|
-
**keryx program**. Run them with `keryx
|
|
55
|
-
<scope.json> --json`, `keryx review jev-risk
|
|
56
|
-
|
|
57
|
-
`scope.json` every other Wave A/B reviewer's
|
|
58
|
-
checks exactly the same hunks.
|
|
59
|
-
take a diff/PR target instead:
|
|
60
|
-
|
|
61
|
-
--
|
|
50
|
+
(both flow 332), `review-jev-docs` and `review-jev-comments` (flow 333), and
|
|
51
|
+
`review-jev-contract` (flow 335) are ADDITIONAL reviewers, never replacing
|
|
52
|
+
any other, and their dispatch mechanism differs from every reviewer named in
|
|
53
|
+
SKILL.md's Routing Table: there is no platform-native agent to invoke,
|
|
54
|
+
because each is a deterministic **keryx program**. Run them with `keryx
|
|
55
|
+
review jev-rules --scope <scope.json> --json`, `keryx review jev-risk
|
|
56
|
+
--scope <scope.json> --json`, and `keryx review jev-scenarios --scope
|
|
57
|
+
<scope.json> --json` — the SAME `scope.json` every other Wave A/B reviewer's
|
|
58
|
+
dispatch already reads, so each checks exactly the same hunks.
|
|
59
|
+
`review-jev-docs` and `review-jev-comments` take a diff/PR target instead:
|
|
60
|
+
`keryx review jev-docs (--diff <ref>|--pr <n>) --json` and `keryx review
|
|
61
|
+
jev-comments --pr <n> --repo <owner/repo> --json`. `review-jev-contract`
|
|
62
|
+
also takes a diff/PR target, plus an optional linked flow: `keryx review
|
|
63
|
+
jev-contract (--diff <ref>|--pr <n>) [--flow <id>] --json` — a `--pr` target
|
|
64
|
+
checks the description's own claims; `--flow <id>` additionally checks that
|
|
65
|
+
flow's frozen acceptance criteria (reusing flow 328's `check-ac.ts`, see its
|
|
66
|
+
own SKILL.md). Read each `--json` output as a `REVIEW_RESULT` and merge its
|
|
62
67
|
`findings` into the consolidated array exactly like a sub-agent reviewer's:
|
|
63
68
|
same Sub-Agent Report Quality Gate, same dedup, same Wave C verification.
|
|
64
69
|
`review-jev-risk` additionally emits `ranked` and `routingHints`;
|
|
65
|
-
`review-jev-scenarios` additionally emits `checklist`
|
|
66
|
-
|
|
70
|
+
`review-jev-scenarios` additionally emits `checklist`; `review-jev-contract`
|
|
71
|
+
additionally emits `claims`, `budget`, and (when `--flow` was given)
|
|
72
|
+
`acCheck` — see each reviewer's own SKILL.md for what to do with its extra.
|
|
67
73
|
|
|
68
74
|
Gate each BEFORE running its command, not after: skip it — recorded in
|
|
69
75
|
`Skipped reviewers` with the reason, never silently absent — when its own
|
|
70
76
|
opt-in (`review.jev.rules` / `review.jev.risk` / `review.jev.scenarios` /
|
|
71
|
-
`review.jev.docs` / `review.jev.comments`) is not
|
|
72
|
-
`.metaproject/tasks.config.json`, or when no Jev/OpenRouter
|
|
73
|
-
resolvable. Either gate failing means the command itself would
|
|
74
|
-
any read or network call, so checking first saves a doomed
|
|
75
|
-
|
|
76
|
-
additionally needs a comment ledger to already exist
|
|
77
|
-
collect` run first) — it refuses, before any read,
|
|
77
|
+
`review.jev.docs` / `review.jev.comments` / `review.jev.contract`) is not
|
|
78
|
+
`true` in `.metaproject/tasks.config.json`, or when no Jev/OpenRouter
|
|
79
|
+
credential is resolvable. Either gate failing means the command itself would
|
|
80
|
+
refuse before any read or network call, so checking first saves a doomed
|
|
81
|
+
dispatch. The six opt-ins are independent: any subset may be on.
|
|
82
|
+
`review-jev-comments` additionally needs a comment ledger to already exist
|
|
83
|
+
(`keryx review comments collect` run first) — it refuses, before any read,
|
|
84
|
+
when there is none. When `review.jev.contract` is on, dispatch
|
|
85
|
+
`review-jev-contract` with `--pr` (never `--diff`, which has no description
|
|
86
|
+
to check) and let its `findings` cover the description-vs-diff claim, in
|
|
87
|
+
place of the by-eye Stage 1 judgement described in SKILL.md's "This gate owns
|
|
88
|
+
the description-vs-diff comparison" — see that section's own note.
|
|
78
89
|
|
|
79
90
|
`keryx review reviewers --json` marks a CLI-engine reviewer with
|
|
80
91
|
`"engine": "jev"` on its `bundled` entry — the field's presence, not its
|
|
@@ -82,3 +93,52 @@ absence, is what distinguishes it from the default (an LLM sub-agent
|
|
|
82
93
|
dispatch). A future engine-backed reviewer follows the same pattern: gate on
|
|
83
94
|
its own opt-in and reachability, dispatch as a command, merge its `--json`
|
|
84
95
|
output the same way.
|
|
96
|
+
|
|
97
|
+
## `jev-triage` — advisory annotations, not a reviewer (Step 9b)
|
|
98
|
+
|
|
99
|
+
`review-jev-triage` (flow 340) is a DIFFERENT shape from every CLI-engine
|
|
100
|
+
reviewer above: it never produces a `findings` array of its own, and it is
|
|
101
|
+
never merged into the consolidated array. It runs AFTER the Sub-Agent Report
|
|
102
|
+
Quality Gate has already validated/merged/deduplicated the round's findings
|
|
103
|
+
(Step 9) and BEFORE Wave C verification (Step 10) — Step 9b in the Workflow
|
|
104
|
+
block — annotating the findings that already exist. It is advisory and
|
|
105
|
+
annotate-only by construction: nothing it returns drops or demotes a finding,
|
|
106
|
+
and no downstream step is permitted to treat its output as anything but a
|
|
107
|
+
hint.
|
|
108
|
+
|
|
109
|
+
Run it once per round, over the round's own consolidated findings:
|
|
110
|
+
|
|
111
|
+
```bash
|
|
112
|
+
keryx review jev-triage --report <this round's review package dir> --json
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
Gate it the same way as every other CLI-engine reviewer: skip it — recorded
|
|
116
|
+
in `Skipped reviewers` with the reason — when `review.jev.triage` is not
|
|
117
|
+
`true` in `.metaproject/tasks.config.json`, or when no Jev/OpenRouter
|
|
118
|
+
credential is resolvable. Either gate failing means the command itself would
|
|
119
|
+
refuse before any read or network call.
|
|
120
|
+
|
|
121
|
+
Its `--json` output is `{status, reviewer: "review-jev-triage", summary,
|
|
122
|
+
annotations, budget, tokens}` — never a `REVIEW_RESULT` merged into the
|
|
123
|
+
findings array. `annotations` carries three tracks, each read and used
|
|
124
|
+
differently in the report:
|
|
125
|
+
|
|
126
|
+
| Track | Shape | What to do with it |
|
|
127
|
+
|---|---|---|
|
|
128
|
+
| `severity_check` | `{id, p, flagged}` per blocker/major finding — `flagged` means `p < 0.4` | Show `flagged` findings as a soft note next to their existing severity ("Jev's own trigger+outcome check scored this low"). Never change the finding's `severity` field from this alone. |
|
|
129
|
+
| `merge_candidates` | `{id, a, b, reason, p}` per candidate pair | A high `p` is a suggestion to a human that `a` and `b` may be the same defect — surface it in the report as a note on both findings. Never merge, drop, or renumber either finding from this alone. |
|
|
130
|
+
| `verify_order` | `{id, p}`, sorted lowest-`p`-first | Feed this order into Wave C's own dispatch — verify the lowest-plausibility findings first — never as a reason to skip verifying any of them. |
|
|
131
|
+
|
|
132
|
+
`status` is `DONE_WITH_CONCERNS` — never `DONE` — when any `severity_check` is
|
|
133
|
+
`flagged` OR any `merge_candidates` pair scores `p >= 0.5`
|
|
134
|
+
(`LIKELY_DUPLICATE_THRESHOLD`, `src/commands/review-jev-triage.ts`). Live-check
|
|
135
|
+
calibration found same-file pairs that were NOT duplicates scoring around
|
|
136
|
+
0.6, so a merge candidate above the threshold is a prompt to look, never a
|
|
137
|
+
merge — the status flip is a nudge to read the pair, not a verdict that the
|
|
138
|
+
findings are duplicates.
|
|
139
|
+
|
|
140
|
+
Show its annotations in the report as their own subsection, next to (not
|
|
141
|
+
inside) the findings they annotate — same separation the finding schema
|
|
142
|
+
itself draws between a reviewer's claim (`severity`, `problem`, …) and what
|
|
143
|
+
became of it (`disposition`), which a reviewer never states and this pass
|
|
144
|
+
does not either.
|
|
@@ -42,6 +42,7 @@ Review Orchestrator Progress:
|
|
|
42
42
|
- [ ] Step 7: Stage 1 gate - spec compliance check (if issue/task provided)
|
|
43
43
|
- [ ] Step 8: Dispatch selected reviewers in PARALLEL with reviewer-input schema
|
|
44
44
|
- [ ] Step 9: Collect reviewer-finding schema results and handle NEEDS_CONTEXT
|
|
45
|
+
- [ ] Step 9b: `jev-triage` — advisory, annotate-only severity/duplicate/verify-order annotations over the consolidated findings (opt-in `review.jev.triage`)
|
|
45
46
|
- [ ] Step 10: Wave C — dispatch `review-verifier` over the consolidated findings
|
|
46
47
|
- [ ] Step 11: Sort by severity, deduplicate, emit unified report
|
|
47
48
|
- [ ] Step 12: Emit the machine-readable `keryx:findings` block alongside the report
|
|
@@ -1168,7 +1169,7 @@ that already holds intent and diff side by side is this one.
|
|
|
1168
1169
|
|
|
1169
1170
|
Run it on every round, not only the first. The drift the finding catches is
|
|
1170
1171
|
created BY the rounds: the code moves to answer findings, the body does not, and
|
|
1171
|
-
whoever reads the merge commit a year later reads the body.
|
|
1172
|
+
whoever reads the merge commit a year later reads the body. When `review.jev.contract` is on, dispatch `review-jev-contract --pr` and read its `findings` as this comparison's scored result instead of judging it by eye; the by-eye judgement is the fallback when that opt-in is off — `SKILL.detail.md` § "CLI-engine reviewers".
|
|
1172
1173
|
|
|
1173
1174
|
---
|
|
1174
1175
|
|
|
@@ -1177,7 +1178,7 @@ whoever reads the merge commit a year later reads the body.
|
|
|
1177
1178
|
Dispatch selected reviewers in parallel when independent. Use waves when token budget is tight or when one reviewer needs another result:
|
|
1178
1179
|
|
|
1179
1180
|
1. Wave A - core correctness/risk reviewers: logic, architecture, security/highload when selected.
|
|
1180
|
-
2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files. `review-jev-rules` (flow 330), `review-jev-risk`/`review-jev-scenarios` (flow 332), `review-jev-docs`/`review-jev-comments` (flow 333) also run here, CLI-engine not sub-agent, `"engine": "jev"` in `keryx review reviewers --json`, gated on their own opt-in and a resolvable Jev/OpenRouter credential — `SKILL.detail.md` § "CLI-engine reviewers".
|
|
1181
|
+
2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files. `review-jev-rules` (flow 330), `review-jev-risk`/`review-jev-scenarios` (flow 332), `review-jev-docs`/`review-jev-comments` (flow 333), `review-jev-contract` (flow 335) also run here, CLI-engine not sub-agent, `"engine": "jev"` in `keryx review reviewers --json`, gated on their own opt-in and a resolvable Jev/OpenRouter credential — `SKILL.detail.md` § "CLI-engine reviewers".
|
|
1181
1182
|
3. Wave C - **verification**: `review-verifier` over the consolidated findings, when blockers/majors
|
|
1182
1183
|
exist, `--verify` is set, or the PR is high-risk. See below.
|
|
1183
1184
|
|
|
@@ -1691,8 +1692,7 @@ CONTEXT_PATH: .metaproject/jobs/<job-name>/ai/context.md
|
|
|
1691
1692
|
If provided and the file exists, read the context document **before** running scope detection.
|
|
1692
1693
|
Use it to understand:
|
|
1693
1694
|
- Intentionally chosen libraries and patterns (do not flag as issues)
|
|
1694
|
-
- Architectural decisions already agreed upon
|
|
1695
|
-
- Acceptance criteria to drive the Stage 1 spec compliance gate
|
|
1695
|
+
- Architectural decisions already agreed upon, and acceptance criteria driving the Stage 1 spec compliance gate
|
|
1696
1696
|
|
|
1697
1697
|
If absent, proceed normally — context is optional and non-blocking.
|
|
1698
1698
|
|
|
@@ -0,0 +1,4 @@
|
|
|
1
|
+
{
|
|
2
|
+
"agents": [],
|
|
3
|
+
"note": "honest gate (DeepSeek deepseek-chat runner+judge, strictness high, trials 10, flow 337) ran and failed for all four skills -- trigger accuracy, not behavior content (every behavior scenario scored 0.9 or 1.0). No generated pair ships; pack stays stability: experimental. See governance/eval.json for the recorded reports and W1-stack-catalog.md's Wave 4 batch 5 implementation notes for the diagnosis."
|
|
4
|
+
}
|