@mrciphersmith/keryx 0.3.1 → 0.3.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/cli.js +7310 -2471
- package/dist/core.js +116 -10
- package/package.json +1 -1
- package/src/gdskills/bundled/install-manifest.json +349 -2
- package/src/gdskills/bundled/rules/core/model-selection.mdc +18 -0
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/review-jev-comments/SKILL.md +184 -0
- package/src/gdskills/bundled/skills/review/review-jev-contract/SKILL.md +193 -0
- package/src/gdskills/bundled/skills/review/review-jev-docs/SKILL.md +189 -0
- package/src/gdskills/bundled/skills/review/review-jev-risk/SKILL.md +190 -0
- package/src/gdskills/bundled/skills/review/review-jev-scenarios/SKILL.md +187 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +88 -15
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +4 -4
- package/src/gdskills/bundled/stacks/c-cpp/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/c-cpp/governance/eval.json +1777 -0
- package/src/gdskills/bundled/stacks/c-cpp/governance/scout.json +31 -0
- package/src/gdskills/bundled/stacks/c-cpp/pack.json +42 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/coding-style.mdc +80 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/patterns.mdc +87 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/security.mdc +90 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/testing.mdc +83 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/SKILL.md +153 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/SKILL.md +132 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/SKILL.md +151 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/SKILL.md +152 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/evals.json +74 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/eval.json +1295 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/scout.json +26 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/pack.json +41 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/patterns.mdc +77 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/security.mdc +144 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/SKILL.md +121 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/evals.json +73 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/SKILL.md +139 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/SKILL.md +147 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/eval.json +865 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/scout.json +16 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/pack.json +46 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/coding-style.mdc +74 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/patterns.mdc +81 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/security.mdc +146 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/testing.mdc +61 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/SKILL.md +151 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/SKILL.md +135 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/evals.json +76 -0
- package/src/gdskills/bundled/stacks/php-laravel/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/php-laravel/governance/eval.json +1829 -0
- package/src/gdskills/bundled/stacks/php-laravel/governance/scout.json +33 -0
- package/src/gdskills/bundled/stacks/php-laravel/pack.json +41 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/coding-style.mdc +82 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/patterns.mdc +80 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/security.mdc +80 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/testing.mdc +82 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/SKILL.md +143 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/SKILL.md +126 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/evals.json +76 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/SKILL.md +140 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/evals.json +75 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/SKILL.md +124 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/evals.json +74 -0
- package/src/gdskills/bundled/stacks/ruby-rails/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/ruby-rails/governance/eval.json +1673 -0
- package/src/gdskills/bundled/stacks/ruby-rails/governance/scout.json +33 -0
- package/src/gdskills/bundled/stacks/ruby-rails/pack.json +42 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/coding-style.mdc +69 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/patterns.mdc +93 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/security.mdc +90 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/testing.mdc +89 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/SKILL.md +143 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/evals.json +73 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/SKILL.md +134 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/evals.json +71 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/SKILL.md +141 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/evals.json +72 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/SKILL.md +125 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/evals.json +72 -0
- package/src/gdskills/bundled/stacks/sql-db/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/sql-db/governance/eval.json +1829 -0
- package/src/gdskills/bundled/stacks/sql-db/governance/scout.json +30 -0
- package/src/gdskills/bundled/stacks/sql-db/pack.json +40 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/coding-style.mdc +69 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/patterns.mdc +134 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/security.mdc +74 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/testing.mdc +83 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/SKILL.md +147 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/evals.json +72 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/SKILL.md +132 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/SKILL.md +153 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/evals.json +77 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/SKILL.md +129 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/evals.json +73 -0
|
@@ -0,0 +1,190 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: review-jev-risk
|
|
3
|
+
model_tier: light
|
|
4
|
+
description: |
|
|
5
|
+
Use when: an ADDITIONAL, machine-scored RISK MAP is wanted over every changed hunk —
|
|
6
|
+
never replacing any other reviewer. Dispatched by review-orchestrator in Wave B, via
|
|
7
|
+
the CLI (`keryx review jev-risk`), when `review.jev.risk: true` in
|
|
8
|
+
.metaproject/tasks.config.json and a Jev/OpenRouter credential is resolvable. It is not
|
|
9
|
+
an LLM sub-agent: there is nothing to dispatch through a platform-native agent
|
|
10
|
+
mechanism, and no prose is generated by a model — Jev only answers one `noul`
|
|
11
|
+
probability per (hunk, risk dimension) pair; keryx writes every word of every finding.
|
|
12
|
+
NOT for: judgement calls a risk dimension does not state (style, architecture beyond
|
|
13
|
+
risk-flagging — those stay with the reviewers that already cover them), and not a
|
|
14
|
+
substitute for a human reviewer reading the ranked map it produces.
|
|
15
|
+
triggers:
|
|
16
|
+
- "jev risk"
|
|
17
|
+
- "risk map"
|
|
18
|
+
- "review --jev-risk"
|
|
19
|
+
metadata:
|
|
20
|
+
author: "MrCipherSmith"
|
|
21
|
+
version: "1.0.0"
|
|
22
|
+
category: "review"
|
|
23
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
24
|
+
engine: "jev"
|
|
25
|
+
origin: "authored"
|
|
26
|
+
license: "MIT"
|
|
27
|
+
---
|
|
28
|
+
|
|
29
|
+
# Review — Jev Risk (deterministic risk map of the diff)
|
|
30
|
+
|
|
31
|
+
An ADDITIONAL orchestrator reviewer, flow 332. It is a **keryx program**, not
|
|
32
|
+
an LLM sub-agent: `review-orchestrator` never dispatches a platform-native
|
|
33
|
+
agent for it, it runs `keryx review jev-risk` and reads the `--json` output.
|
|
34
|
+
Every finding is composed by keryx; Jev ("System One" on OpenRouter) supplies
|
|
35
|
+
only five `noul` probabilities per hunk — one per risk dimension — never
|
|
36
|
+
prose.
|
|
37
|
+
|
|
38
|
+
---
|
|
39
|
+
|
|
40
|
+
## What it does
|
|
41
|
+
|
|
42
|
+
1. Takes every changed hunk from `keryx review scope`/`buildReviewScope`
|
|
43
|
+
(mechanical bulk already dropped).
|
|
44
|
+
2. Computes deterministic facts FIRST, per hunk: a path class (auth/
|
|
45
|
+
permissions, crypto, migrations, schema, public API, config, concurrency
|
|
46
|
+
primitives, IO), the exported symbols it touches, its changed-line count,
|
|
47
|
+
and whether a test file elsewhere in the diff touches the same module.
|
|
48
|
+
3. Asks Jev **one `noul` per risk dimension** per hunk — security-sensitive,
|
|
49
|
+
data/migration, public-API/contract change, concurrency, error-handling —
|
|
50
|
+
batched under the vendor's 64k token budget, capped at `--max-calls`
|
|
51
|
+
(default 150, counted as hunk x dimension pairs), with every drop
|
|
52
|
+
reported.
|
|
53
|
+
4. Ranks hunks by combined risk (the MAX across its five dimensions — one
|
|
54
|
+
high-risk dimension is enough to draw attention) so a human reviewer knows
|
|
55
|
+
where to look first.
|
|
56
|
+
5. Emits a finding ONLY for a hunk above threshold (default 0.7) AND with no
|
|
57
|
+
test touched nearby in the same diff (a fact) — severity capped at
|
|
58
|
+
`info`/`minor`, never higher: this flags attention, it does not assert a
|
|
59
|
+
defect.
|
|
60
|
+
6. Additionally emits a **routing hint**: hunks above threshold with a
|
|
61
|
+
security or concurrency dimension are listed with a suggested reviewer
|
|
62
|
+
(`review-security-code`/`review-highload`) for the orchestrator to
|
|
63
|
+
consider dispatching — see "Orchestrator integration" below.
|
|
64
|
+
|
|
65
|
+
---
|
|
66
|
+
|
|
67
|
+
## Input Contract
|
|
68
|
+
|
|
69
|
+
Not dispatched with a prompt — invoked as a CLI command:
|
|
70
|
+
|
|
71
|
+
```text
|
|
72
|
+
keryx review jev-risk (--diff <ref> | --pr <n> | --scope <scope.json>)
|
|
73
|
+
[--max-calls <n>] [--threshold <0..1>]
|
|
74
|
+
[--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
|
|
75
|
+
[--fixtures <dir>] [--json]
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
`review-orchestrator` passes `--scope <scope.json>` — the same `keryx review
|
|
79
|
+
scope --json` file every other reviewer's dispatch already reads — so this
|
|
80
|
+
reviewer checks exactly the same hunks as everyone else in the round.
|
|
81
|
+
|
|
82
|
+
---
|
|
83
|
+
|
|
84
|
+
## Opt-in and privacy
|
|
85
|
+
|
|
86
|
+
Refuses before any read or network call unless `.metaproject/tasks.config.json`
|
|
87
|
+
declares:
|
|
88
|
+
|
|
89
|
+
```json
|
|
90
|
+
{ "review": { "jev": { "risk": true } } }
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
Every hunk sent to Jev is redacted first through `src/security/service.ts` —
|
|
94
|
+
the same floor `review conform`/`review ci-triage`/`review jev-rules` already
|
|
95
|
+
apply. No cache or output file stores a credential.
|
|
96
|
+
|
|
97
|
+
---
|
|
98
|
+
|
|
99
|
+
## Output Contract
|
|
100
|
+
|
|
101
|
+
Emits a `REVIEW_RESULT`-shaped object matching
|
|
102
|
+
`.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
|
|
103
|
+
— `status`, `reviewer: "review-jev-risk"`, `summary`, `findings`, `stats` —
|
|
104
|
+
under `--json`, plus two orchestrator-facing extras: `ranked` (every scored
|
|
105
|
+
hunk, highest risk first) and `routingHints` (see above). Its findings merge
|
|
106
|
+
into the consolidated array exactly like any other reviewer's: same Quality
|
|
107
|
+
Gate, same dedup, same Wave C verification.
|
|
108
|
+
|
|
109
|
+
---
|
|
110
|
+
|
|
111
|
+
## Orchestrator integration
|
|
112
|
+
|
|
113
|
+
- `keryx review reviewers --json` lists it under `bundled` with
|
|
114
|
+
`"engine": "jev"`.
|
|
115
|
+
- `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
|
|
116
|
+
convention reviewers and flow 330's `review-jev-rules`, when the opt-in is
|
|
117
|
+
on and a Jev/OpenRouter credential resolves. When either is false it is
|
|
118
|
+
**skipped with the stated reason**, recorded in `Skipped reviewers` — never
|
|
119
|
+
silently absent.
|
|
120
|
+
- It is **ADDITIONAL**: it never replaces `review-security-code`,
|
|
121
|
+
`review-highload`, or any other pass.
|
|
122
|
+
- **Routing hint**: read `routingHints` from its `--json` output. Each entry
|
|
123
|
+
names a hunk location, the dimension that crossed threshold
|
|
124
|
+
(`security`/`concurrency`) and a `suggestedReviewer`
|
|
125
|
+
(`review-security-code`/`review-highload`). When that reviewer was not
|
|
126
|
+
already selected for the round, dispatch it too — the hint is advisory,
|
|
127
|
+
not a hard requirement, and the orchestrator's own path-based selection
|
|
128
|
+
always takes precedence when it already covers the same hunk.
|
|
129
|
+
|
|
130
|
+
---
|
|
131
|
+
|
|
132
|
+
### Shared laws (every reviewer)
|
|
133
|
+
|
|
134
|
+
1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
|
|
135
|
+
name the input, call, or condition that reaches the code, you have an
|
|
136
|
+
observation, not a finding. Report it as `info` and say what would settle it.
|
|
137
|
+
2. **Never flag the theoretical.** The path you describe must exist in the code
|
|
138
|
+
under review. Do not report a safe API because it could be misused, or a
|
|
139
|
+
pattern because it is often wrong elsewhere.
|
|
140
|
+
3. **One finding per class, not one per occurrence.** When the same shape appears
|
|
141
|
+
at several sites, report it once and list every site. Ten findings that are one
|
|
142
|
+
finding hide the other nine problems.
|
|
143
|
+
|
|
144
|
+
This reviewer satisfies all three by construction: `evidence` always names the
|
|
145
|
+
hunk location, quotes the changed line, and states every dimension's
|
|
146
|
+
probability (never an unreproducible claim); a finding is synthesized only
|
|
147
|
+
against the hunk actually scored, never a hypothetical pattern; and each
|
|
148
|
+
hunk is a distinct site with its own probabilities, so there is no repeated
|
|
149
|
+
class across sites for this reviewer to collapse — the routing hint follows
|
|
150
|
+
the same discipline, naming the exact hunk and dimension that crossed
|
|
151
|
+
threshold rather than a general warning.
|
|
152
|
+
|
|
153
|
+
---
|
|
154
|
+
|
|
155
|
+
## Red Flags
|
|
156
|
+
|
|
157
|
+
| Rationalization | Why it is wrong |
|
|
158
|
+
|---|---|
|
|
159
|
+
| "Jev said 0.9, so this hunk is definitely dangerous" | A `noul` score is a probability across one dimension, not a verdict — severity stays capped at `info`/`minor` precisely because Jev alone is not authoritative |
|
|
160
|
+
| "No nearby test means nobody tested this at all" | `hasNearbyTest` requires the nearby test's OWN diff text to mention a touched symbol or import the hunk's module (`testHunkEvidence`), tightened after a live-check false negative — still a regex heuristic, not proof of coverage through an indirection it cannot see |
|
|
161
|
+
| "This hunk touches a .md file, so 'no nearby test' is a meaningful finding" | Docs (`.md`/`.txt`) hunks are excluded before scoring (`isNonCodeHunk`) after a live check where prose scored `public-api` — this cannot happen; `selection.hunksNotCode` reports the count |
|
|
162
|
+
| "No findings means the diff is low risk" | `--max-calls` bounds how many hunks are scored; a capped run reports `selection.hunksSkipped` for exactly this reason — read the selection stats before treating silence as clean |
|
|
163
|
+
|
|
164
|
+
---
|
|
165
|
+
|
|
166
|
+
## Verification
|
|
167
|
+
|
|
168
|
+
Before trusting a run's findings:
|
|
169
|
+
|
|
170
|
+
1. Read `selection.hunksSkipped` in the `--json` output — a capped run
|
|
171
|
+
covered fewer than "every retained hunk" and the report should say so.
|
|
172
|
+
2. Spot-check a handful of `ranked` entries against the actual hunk: does the
|
|
173
|
+
file/line range and the top dimension make sense for what actually
|
|
174
|
+
changed there? Docs (`.md`/`.txt`) hunks never reach `ranked` at all —
|
|
175
|
+
`selection.hunksNotCode` should account for every one of them in the
|
|
176
|
+
diff.
|
|
177
|
+
3. Confirm every finding carries `reviewer: "review-jev-risk"` and a
|
|
178
|
+
`dedupe_key` — both are required for the orchestrator's Quality Gate and
|
|
179
|
+
Wave C verification to route it correctly.
|
|
180
|
+
|
|
181
|
+
---
|
|
182
|
+
|
|
183
|
+
## Scope Boundaries
|
|
184
|
+
|
|
185
|
+
| Concern | This skill | Use instead |
|
|
186
|
+
|---------|------------|-------------|
|
|
187
|
+
| A ranked risk map of the diff's hunks | YES | — |
|
|
188
|
+
| An actual security/concurrency finding | NO (routing hint only) | `review-security-code` / `review-highload` |
|
|
189
|
+
| Judging whether a reported finding is real | NO | `review-verifier` |
|
|
190
|
+
| Which user scenarios changed | NO | `review-jev-scenarios` |
|
|
@@ -0,0 +1,187 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: review-jev-scenarios
|
|
3
|
+
model_tier: light
|
|
4
|
+
description: |
|
|
5
|
+
Use when: an ADDITIONAL, machine-scored FUNCTIONAL review is wanted — which user
|
|
6
|
+
scenarios a PR likely changes — never replacing any other reviewer. Dispatched by
|
|
7
|
+
review-orchestrator in Wave B, via the CLI (`keryx review jev-scenarios`), when
|
|
8
|
+
`review.jev.scenarios: true` in .metaproject/tasks.config.json and a Jev/OpenRouter
|
|
9
|
+
credential is resolvable. It is not an LLM sub-agent: there is nothing to dispatch
|
|
10
|
+
through a platform-native agent mechanism, and no prose is generated by a model — Jev
|
|
11
|
+
only answers one `noul` "does this change alter this scenario's behaviour?" per
|
|
12
|
+
scenario whose linked code the diff touches; keryx writes every word of every finding.
|
|
13
|
+
NOT for: discovering scenarios nobody documented (it reads existing gdwiki
|
|
14
|
+
user-scenario pages, PRD requirement/scenario sections, and README/docs "how to"
|
|
15
|
+
sections — a scenario that exists only in someone's head is invisible to it), and not
|
|
16
|
+
a substitute for manually walking the checklist it produces.
|
|
17
|
+
triggers:
|
|
18
|
+
- "jev scenarios"
|
|
19
|
+
- "functional review"
|
|
20
|
+
- "review --jev-scenarios"
|
|
21
|
+
metadata:
|
|
22
|
+
author: "MrCipherSmith"
|
|
23
|
+
version: "1.0.0"
|
|
24
|
+
category: "review"
|
|
25
|
+
compatible_harnesses: "cursor,codex,zed,opencode,claude"
|
|
26
|
+
engine: "jev"
|
|
27
|
+
origin: "authored"
|
|
28
|
+
license: "MIT"
|
|
29
|
+
---
|
|
30
|
+
|
|
31
|
+
# Review — Jev Scenarios (deterministic functional review)
|
|
32
|
+
|
|
33
|
+
An ADDITIONAL orchestrator reviewer, flow 332. It is a **keryx program**, not
|
|
34
|
+
an LLM sub-agent: `review-orchestrator` never dispatches a platform-native
|
|
35
|
+
agent for it, it runs `keryx review jev-scenarios` and reads the `--json`
|
|
36
|
+
output. Every finding is composed by keryx; Jev ("System One" on OpenRouter)
|
|
37
|
+
supplies only one `noul` probability per scenario — never prose.
|
|
38
|
+
|
|
39
|
+
---
|
|
40
|
+
|
|
41
|
+
## What it does
|
|
42
|
+
|
|
43
|
+
1. Gathers user scenarios deterministically from three sources: gdwiki pages
|
|
44
|
+
of type `user-scenario` (`.metaproject/wiki/user-scenarios/**`), PRD
|
|
45
|
+
requirement/scenario sections (`docs/requirements/**`, any heading
|
|
46
|
+
matching scenario/requirement/user story/use case), and README/docs "how
|
|
47
|
+
to" sections (`README.md`, `docs/docs/**`).
|
|
48
|
+
2. Extracts each scenario's own links to code — markdown links and inline
|
|
49
|
+
code-span file references, the same convention `src/wiki/backlinks.ts`
|
|
50
|
+
already established for the wiki's own code edges.
|
|
51
|
+
3. Computes, per scenario, which of its linked files the diff actually
|
|
52
|
+
touches (a fact) and whether a test in the diff covers a touched link (a
|
|
53
|
+
fact).
|
|
54
|
+
4. For every scenario with at least one touched link, asks Jev **one
|
|
55
|
+
`noul`**: "does this change alter this scenario's behaviour?" — batched
|
|
56
|
+
under the vendor's 64k token budget, capped at `--max-calls` (default
|
|
57
|
+
150), with every drop reported.
|
|
58
|
+
5. Emits a **manual-check list**: every scenario at/above threshold (default
|
|
59
|
+
0.5), ranked, each with its touched links as evidence — for a human to
|
|
60
|
+
walk before merge.
|
|
61
|
+
6. Emits a `minor` finding for each likely-affected scenario with **no test
|
|
62
|
+
in the diff covering it** (a fact) — severity fixed at `minor`: this
|
|
63
|
+
flags attention over a functional scenario, it does not assert a defect.
|
|
64
|
+
|
|
65
|
+
---
|
|
66
|
+
|
|
67
|
+
## Input Contract
|
|
68
|
+
|
|
69
|
+
Not dispatched with a prompt — invoked as a CLI command:
|
|
70
|
+
|
|
71
|
+
```text
|
|
72
|
+
keryx review jev-scenarios (--diff <ref> | --pr <n> | --scope <scope.json>)
|
|
73
|
+
[--max-calls <n>] [--threshold <0..1>]
|
|
74
|
+
[--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
|
|
75
|
+
[--fixtures <dir>] [--json]
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
`review-orchestrator` passes `--scope <scope.json>` — the same file every
|
|
79
|
+
other reviewer's dispatch already reads — so this reviewer's "touched by this
|
|
80
|
+
diff" facts agree with the round's own scope.
|
|
81
|
+
|
|
82
|
+
---
|
|
83
|
+
|
|
84
|
+
## Opt-in and privacy
|
|
85
|
+
|
|
86
|
+
Refuses before any read or network call unless `.metaproject/tasks.config.json`
|
|
87
|
+
declares:
|
|
88
|
+
|
|
89
|
+
```json
|
|
90
|
+
{ "review": { "jev": { "scenarios": true } } }
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
Every scenario's text sent to Jev is redacted first through
|
|
94
|
+
`src/security/service.ts` — the same floor every other Jev-backed mode in
|
|
95
|
+
this repository already applies. No cache or output file stores a
|
|
96
|
+
credential.
|
|
97
|
+
|
|
98
|
+
---
|
|
99
|
+
|
|
100
|
+
## Output Contract
|
|
101
|
+
|
|
102
|
+
Emits a `REVIEW_RESULT`-shaped object matching
|
|
103
|
+
`.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
|
|
104
|
+
— `status`, `reviewer: "review-jev-scenarios"`, `summary`, `findings`,
|
|
105
|
+
`stats` — under `--json`, plus one orchestrator-facing extra: `checklist`
|
|
106
|
+
(the ranked manual-check list described above). Its findings merge into the
|
|
107
|
+
consolidated array exactly like any other reviewer's: same Quality Gate,
|
|
108
|
+
same dedup, same Wave C verification.
|
|
109
|
+
|
|
110
|
+
---
|
|
111
|
+
|
|
112
|
+
## Orchestrator integration
|
|
113
|
+
|
|
114
|
+
- `keryx review reviewers --json` lists it under `bundled` with
|
|
115
|
+
`"engine": "jev"`.
|
|
116
|
+
- `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
|
|
117
|
+
convention reviewers and flow 330's `review-jev-rules`/flow 332's
|
|
118
|
+
`review-jev-risk`, when the opt-in is on and a Jev/OpenRouter credential
|
|
119
|
+
resolves. When either is false it is **skipped with the stated reason**,
|
|
120
|
+
recorded in `Skipped reviewers` — never silently absent.
|
|
121
|
+
- It is **ADDITIONAL**: it never replaces any domain reviewer's own read of
|
|
122
|
+
the diff.
|
|
123
|
+
- The consolidated report's "Manual verification" section (or equivalent)
|
|
124
|
+
should surface `checklist` verbatim alongside any findings — a scenario
|
|
125
|
+
that scored below the finding threshold but above the checklist threshold
|
|
126
|
+
is still worth a human glance even though it produced no finding.
|
|
127
|
+
|
|
128
|
+
---
|
|
129
|
+
|
|
130
|
+
### Shared laws (every reviewer)
|
|
131
|
+
|
|
132
|
+
1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
|
|
133
|
+
name the input, call, or condition that reaches the code, you have an
|
|
134
|
+
observation, not a finding. Report it as `info` and say what would settle it.
|
|
135
|
+
2. **Never flag the theoretical.** The path you describe must exist in the code
|
|
136
|
+
under review. Do not report a safe API because it could be misused, or a
|
|
137
|
+
pattern because it is often wrong elsewhere.
|
|
138
|
+
3. **One finding per class, not one per occurrence.** When the same shape appears
|
|
139
|
+
at several sites, report it once and list every site. Ten findings that are one
|
|
140
|
+
finding hide the other nine problems.
|
|
141
|
+
|
|
142
|
+
This reviewer satisfies all three by construction: severity is fixed at
|
|
143
|
+
`minor` — never higher — and `evidence` always names the scenario's linked
|
|
144
|
+
code and the exact files this diff touched (never an unreproducible claim);
|
|
145
|
+
a scenario is only asked about when a real link of its own is in the diff,
|
|
146
|
+
never a hypothetical connection; and each scenario is its own class with its
|
|
147
|
+
own `dedupe_key`, so there is nothing to collapse across sites.
|
|
148
|
+
|
|
149
|
+
---
|
|
150
|
+
|
|
151
|
+
## Red Flags
|
|
152
|
+
|
|
153
|
+
| Rationalization | Why it is wrong |
|
|
154
|
+
|---|---|
|
|
155
|
+
| "Jev said 0.9, so this scenario is definitely broken" | A `noul` score is a probability that behaviour changed, not a verdict — it only ever produces a `minor` finding, and only when paired with the separate "no covering test" fact |
|
|
156
|
+
| "This PRD scenario links to a huge, frequently-touched file, so the match is meaningless" | Down-weighted after a live check (`SCENARIO_LINK_FANOUT_THRESHOLD`): a link to a file referenced by more scenarios than that only still counts as touched when the scenario's own text names one of the diff's touched exported symbols — read the scenario's own text before assuming a surviving entry is still spurious |
|
|
157
|
+
| "No entries in checklist means nothing changed functionally" | Discovery only covers three sources (gdwiki `user-scenario` pages, `docs/requirements/**`, README/docs "how to" sections); a scenario that exists only in someone's head, or in a doc outside those three locations, is invisible to this reviewer |
|
|
158
|
+
| "checklist is empty so scenarioSources must be empty too" | `scenarioSources` lists every discovered scenario; `checklist` lists only those with a touched link at/above threshold — a project can have hundreds of scenarios discovered and zero touched by a given diff |
|
|
159
|
+
|
|
160
|
+
---
|
|
161
|
+
|
|
162
|
+
## Verification
|
|
163
|
+
|
|
164
|
+
Before trusting a run's findings:
|
|
165
|
+
|
|
166
|
+
1. Read `selection.notApplicable`/`selection.scenariosSkipped` in the
|
|
167
|
+
`--json` output — a scenario with no touched link was never asked, and a
|
|
168
|
+
capped run may have skipped some that were.
|
|
169
|
+
2. Spot-check a `checklist` entry against the scenario's own source
|
|
170
|
+
(`id` names the file, and for a PRD/README section, the heading): does
|
|
171
|
+
the touched link actually sit inside the part of the diff that matters
|
|
172
|
+
for that scenario, or did a large, widely-linked file produce a spurious
|
|
173
|
+
match?
|
|
174
|
+
3. Confirm every finding carries `reviewer: "review-jev-scenarios"` and a
|
|
175
|
+
`dedupe_key` — both are required for the orchestrator's Quality Gate and
|
|
176
|
+
Wave C verification to route it correctly.
|
|
177
|
+
|
|
178
|
+
---
|
|
179
|
+
|
|
180
|
+
## Scope Boundaries
|
|
181
|
+
|
|
182
|
+
| Concern | This skill | Use instead |
|
|
183
|
+
|---------|------------|-------------|
|
|
184
|
+
| Which documented user scenarios a diff likely changes | YES | — |
|
|
185
|
+
| Discovering an undocumented scenario | NO | a human, or write the scenario down first |
|
|
186
|
+
| Judging whether a reported finding is real | NO | `review-verifier` |
|
|
187
|
+
| A risk map of the diff's hunks | NO | `review-jev-risk` |
|
|
@@ -46,22 +46,46 @@ review: it is a fact about the review, not about the diff.
|
|
|
46
46
|
|
|
47
47
|
## CLI-engine reviewers — dispatched as a command, not a sub-agent
|
|
48
48
|
|
|
49
|
-
`review-jev-rules` (flow 330)
|
|
50
|
-
|
|
49
|
+
`review-jev-rules` (flow 330), `review-jev-risk` and `review-jev-scenarios`
|
|
50
|
+
(both flow 332), `review-jev-docs` and `review-jev-comments` (flow 333), and
|
|
51
|
+
`review-jev-contract` (flow 335) are ADDITIONAL reviewers, never replacing
|
|
52
|
+
any other, and their dispatch mechanism differs from every reviewer named in
|
|
51
53
|
SKILL.md's Routing Table: there is no platform-native agent to invoke,
|
|
52
|
-
because
|
|
53
|
-
jev-rules --scope <scope.json> --json
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
a
|
|
54
|
+
because each is a deterministic **keryx program**. Run them with `keryx
|
|
55
|
+
review jev-rules --scope <scope.json> --json`, `keryx review jev-risk
|
|
56
|
+
--scope <scope.json> --json`, and `keryx review jev-scenarios --scope
|
|
57
|
+
<scope.json> --json` — the SAME `scope.json` every other Wave A/B reviewer's
|
|
58
|
+
dispatch already reads, so each checks exactly the same hunks.
|
|
59
|
+
`review-jev-docs` and `review-jev-comments` take a diff/PR target instead:
|
|
60
|
+
`keryx review jev-docs (--diff <ref>|--pr <n>) --json` and `keryx review
|
|
61
|
+
jev-comments --pr <n> --repo <owner/repo> --json`. `review-jev-contract`
|
|
62
|
+
also takes a diff/PR target, plus an optional linked flow: `keryx review
|
|
63
|
+
jev-contract (--diff <ref>|--pr <n>) [--flow <id>] --json` — a `--pr` target
|
|
64
|
+
checks the description's own claims; `--flow <id>` additionally checks that
|
|
65
|
+
flow's frozen acceptance criteria (reusing flow 328's `check-ac.ts`, see its
|
|
66
|
+
own SKILL.md). Read each `--json` output as a `REVIEW_RESULT` and merge its
|
|
67
|
+
`findings` into the consolidated array exactly like a sub-agent reviewer's:
|
|
68
|
+
same Sub-Agent Report Quality Gate, same dedup, same Wave C verification.
|
|
69
|
+
`review-jev-risk` additionally emits `ranked` and `routingHints`;
|
|
70
|
+
`review-jev-scenarios` additionally emits `checklist`; `review-jev-contract`
|
|
71
|
+
additionally emits `claims`, `budget`, and (when `--flow` was given)
|
|
72
|
+
`acCheck` — see each reviewer's own SKILL.md for what to do with its extra.
|
|
73
|
+
|
|
74
|
+
Gate each BEFORE running its command, not after: skip it — recorded in
|
|
75
|
+
`Skipped reviewers` with the reason, never silently absent — when its own
|
|
76
|
+
opt-in (`review.jev.rules` / `review.jev.risk` / `review.jev.scenarios` /
|
|
77
|
+
`review.jev.docs` / `review.jev.comments` / `review.jev.contract`) is not
|
|
78
|
+
`true` in `.metaproject/tasks.config.json`, or when no Jev/OpenRouter
|
|
79
|
+
credential is resolvable. Either gate failing means the command itself would
|
|
80
|
+
refuse before any read or network call, so checking first saves a doomed
|
|
81
|
+
dispatch. The six opt-ins are independent: any subset may be on.
|
|
82
|
+
`review-jev-comments` additionally needs a comment ledger to already exist
|
|
83
|
+
(`keryx review comments collect` run first) — it refuses, before any read,
|
|
84
|
+
when there is none. When `review.jev.contract` is on, dispatch
|
|
85
|
+
`review-jev-contract` with `--pr` (never `--diff`, which has no description
|
|
86
|
+
to check) and let its `findings` cover the description-vs-diff claim, in
|
|
87
|
+
place of the by-eye Stage 1 judgement described in SKILL.md's "This gate owns
|
|
88
|
+
the description-vs-diff comparison" — see that section's own note.
|
|
65
89
|
|
|
66
90
|
`keryx review reviewers --json` marks a CLI-engine reviewer with
|
|
67
91
|
`"engine": "jev"` on its `bundled` entry — the field's presence, not its
|
|
@@ -69,3 +93,52 @@ absence, is what distinguishes it from the default (an LLM sub-agent
|
|
|
69
93
|
dispatch). A future engine-backed reviewer follows the same pattern: gate on
|
|
70
94
|
its own opt-in and reachability, dispatch as a command, merge its `--json`
|
|
71
95
|
output the same way.
|
|
96
|
+
|
|
97
|
+
## `jev-triage` — advisory annotations, not a reviewer (Step 9b)
|
|
98
|
+
|
|
99
|
+
`review-jev-triage` (flow 340) is a DIFFERENT shape from every CLI-engine
|
|
100
|
+
reviewer above: it never produces a `findings` array of its own, and it is
|
|
101
|
+
never merged into the consolidated array. It runs AFTER the Sub-Agent Report
|
|
102
|
+
Quality Gate has already validated/merged/deduplicated the round's findings
|
|
103
|
+
(Step 9) and BEFORE Wave C verification (Step 10) — Step 9b in the Workflow
|
|
104
|
+
block — annotating the findings that already exist. It is advisory and
|
|
105
|
+
annotate-only by construction: nothing it returns drops or demotes a finding,
|
|
106
|
+
and no downstream step is permitted to treat its output as anything but a
|
|
107
|
+
hint.
|
|
108
|
+
|
|
109
|
+
Run it once per round, over the round's own consolidated findings:
|
|
110
|
+
|
|
111
|
+
```bash
|
|
112
|
+
keryx review jev-triage --report <this round's review package dir> --json
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
Gate it the same way as every other CLI-engine reviewer: skip it — recorded
|
|
116
|
+
in `Skipped reviewers` with the reason — when `review.jev.triage` is not
|
|
117
|
+
`true` in `.metaproject/tasks.config.json`, or when no Jev/OpenRouter
|
|
118
|
+
credential is resolvable. Either gate failing means the command itself would
|
|
119
|
+
refuse before any read or network call.
|
|
120
|
+
|
|
121
|
+
Its `--json` output is `{status, reviewer: "review-jev-triage", summary,
|
|
122
|
+
annotations, budget, tokens}` — never a `REVIEW_RESULT` merged into the
|
|
123
|
+
findings array. `annotations` carries three tracks, each read and used
|
|
124
|
+
differently in the report:
|
|
125
|
+
|
|
126
|
+
| Track | Shape | What to do with it |
|
|
127
|
+
|---|---|---|
|
|
128
|
+
| `severity_check` | `{id, p, flagged}` per blocker/major finding — `flagged` means `p < 0.4` | Show `flagged` findings as a soft note next to their existing severity ("Jev's own trigger+outcome check scored this low"). Never change the finding's `severity` field from this alone. |
|
|
129
|
+
| `merge_candidates` | `{id, a, b, reason, p}` per candidate pair | A high `p` is a suggestion to a human that `a` and `b` may be the same defect — surface it in the report as a note on both findings. Never merge, drop, or renumber either finding from this alone. |
|
|
130
|
+
| `verify_order` | `{id, p}`, sorted lowest-`p`-first | Feed this order into Wave C's own dispatch — verify the lowest-plausibility findings first — never as a reason to skip verifying any of them. |
|
|
131
|
+
|
|
132
|
+
`status` is `DONE_WITH_CONCERNS` — never `DONE` — when any `severity_check` is
|
|
133
|
+
`flagged` OR any `merge_candidates` pair scores `p >= 0.5`
|
|
134
|
+
(`LIKELY_DUPLICATE_THRESHOLD`, `src/commands/review-jev-triage.ts`). Live-check
|
|
135
|
+
calibration found same-file pairs that were NOT duplicates scoring around
|
|
136
|
+
0.6, so a merge candidate above the threshold is a prompt to look, never a
|
|
137
|
+
merge — the status flip is a nudge to read the pair, not a verdict that the
|
|
138
|
+
findings are duplicates.
|
|
139
|
+
|
|
140
|
+
Show its annotations in the report as their own subsection, next to (not
|
|
141
|
+
inside) the findings they annotate — same separation the finding schema
|
|
142
|
+
itself draws between a reviewer's claim (`severity`, `problem`, …) and what
|
|
143
|
+
became of it (`disposition`), which a reviewer never states and this pass
|
|
144
|
+
does not either.
|
|
@@ -42,6 +42,7 @@ Review Orchestrator Progress:
|
|
|
42
42
|
- [ ] Step 7: Stage 1 gate - spec compliance check (if issue/task provided)
|
|
43
43
|
- [ ] Step 8: Dispatch selected reviewers in PARALLEL with reviewer-input schema
|
|
44
44
|
- [ ] Step 9: Collect reviewer-finding schema results and handle NEEDS_CONTEXT
|
|
45
|
+
- [ ] Step 9b: `jev-triage` — advisory, annotate-only severity/duplicate/verify-order annotations over the consolidated findings (opt-in `review.jev.triage`)
|
|
45
46
|
- [ ] Step 10: Wave C — dispatch `review-verifier` over the consolidated findings
|
|
46
47
|
- [ ] Step 11: Sort by severity, deduplicate, emit unified report
|
|
47
48
|
- [ ] Step 12: Emit the machine-readable `keryx:findings` block alongside the report
|
|
@@ -1168,7 +1169,7 @@ that already holds intent and diff side by side is this one.
|
|
|
1168
1169
|
|
|
1169
1170
|
Run it on every round, not only the first. The drift the finding catches is
|
|
1170
1171
|
created BY the rounds: the code moves to answer findings, the body does not, and
|
|
1171
|
-
whoever reads the merge commit a year later reads the body.
|
|
1172
|
+
whoever reads the merge commit a year later reads the body. When `review.jev.contract` is on, dispatch `review-jev-contract --pr` and read its `findings` as this comparison's scored result instead of judging it by eye; the by-eye judgement is the fallback when that opt-in is off — `SKILL.detail.md` § "CLI-engine reviewers".
|
|
1172
1173
|
|
|
1173
1174
|
---
|
|
1174
1175
|
|
|
@@ -1177,7 +1178,7 @@ whoever reads the merge commit a year later reads the body.
|
|
|
1177
1178
|
Dispatch selected reviewers in parallel when independent. Use waves when token budget is tight or when one reviewer needs another result:
|
|
1178
1179
|
|
|
1179
1180
|
1. Wave A - core correctness/risk reviewers: logic, architecture, security/highload when selected.
|
|
1180
|
-
2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files. `review-jev-rules` (flow 330) also
|
|
1181
|
+
2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files. `review-jev-rules` (flow 330), `review-jev-risk`/`review-jev-scenarios` (flow 332), `review-jev-docs`/`review-jev-comments` (flow 333), `review-jev-contract` (flow 335) also run here, CLI-engine not sub-agent, `"engine": "jev"` in `keryx review reviewers --json`, gated on their own opt-in and a resolvable Jev/OpenRouter credential — `SKILL.detail.md` § "CLI-engine reviewers".
|
|
1181
1182
|
3. Wave C - **verification**: `review-verifier` over the consolidated findings, when blockers/majors
|
|
1182
1183
|
exist, `--verify` is set, or the PR is high-risk. See below.
|
|
1183
1184
|
|
|
@@ -1691,8 +1692,7 @@ CONTEXT_PATH: .metaproject/jobs/<job-name>/ai/context.md
|
|
|
1691
1692
|
If provided and the file exists, read the context document **before** running scope detection.
|
|
1692
1693
|
Use it to understand:
|
|
1693
1694
|
- Intentionally chosen libraries and patterns (do not flag as issues)
|
|
1694
|
-
- Architectural decisions already agreed upon
|
|
1695
|
-
- Acceptance criteria to drive the Stage 1 spec compliance gate
|
|
1695
|
+
- Architectural decisions already agreed upon, and acceptance criteria driving the Stage 1 spec compliance gate
|
|
1696
1696
|
|
|
1697
1697
|
If absent, proceed normally — context is optional and non-blocking.
|
|
1698
1698
|
|
|
@@ -0,0 +1,4 @@
|
|
|
1
|
+
{
|
|
2
|
+
"agents": [],
|
|
3
|
+
"note": "honest gate (DeepSeek deepseek-chat runner+judge, strictness high, trials 10, flow 337) ran and failed for all four skills -- trigger accuracy, not behavior content (every behavior scenario scored 0.9 or 1.0). No generated pair ships; pack stays stability: experimental. See governance/eval.json for the recorded reports and W1-stack-catalog.md's Wave 4 batch 5 implementation notes for the diagnosis."
|
|
4
|
+
}
|