@mrciphersmith/keryx 0.3.1 → 0.3.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (104) hide show
  1. package/dist/cli.js +7310 -2471
  2. package/dist/core.js +116 -10
  3. package/package.json +1 -1
  4. package/src/gdskills/bundled/install-manifest.json +349 -2
  5. package/src/gdskills/bundled/rules/core/model-selection.mdc +18 -0
  6. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +1 -1
  7. package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
  8. package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
  9. package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
  10. package/src/gdskills/bundled/skills/review/review-jev-comments/SKILL.md +184 -0
  11. package/src/gdskills/bundled/skills/review/review-jev-contract/SKILL.md +193 -0
  12. package/src/gdskills/bundled/skills/review/review-jev-docs/SKILL.md +189 -0
  13. package/src/gdskills/bundled/skills/review/review-jev-risk/SKILL.md +190 -0
  14. package/src/gdskills/bundled/skills/review/review-jev-scenarios/SKILL.md +187 -0
  15. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +88 -15
  16. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +4 -4
  17. package/src/gdskills/bundled/stacks/c-cpp/agent-refs.json +4 -0
  18. package/src/gdskills/bundled/stacks/c-cpp/governance/eval.json +1777 -0
  19. package/src/gdskills/bundled/stacks/c-cpp/governance/scout.json +31 -0
  20. package/src/gdskills/bundled/stacks/c-cpp/pack.json +42 -0
  21. package/src/gdskills/bundled/stacks/c-cpp/rules/coding-style.mdc +80 -0
  22. package/src/gdskills/bundled/stacks/c-cpp/rules/patterns.mdc +87 -0
  23. package/src/gdskills/bundled/stacks/c-cpp/rules/security.mdc +90 -0
  24. package/src/gdskills/bundled/stacks/c-cpp/rules/testing.mdc +83 -0
  25. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/SKILL.md +153 -0
  26. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/evals.json +74 -0
  27. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/SKILL.md +132 -0
  28. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/evals.json +73 -0
  29. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/SKILL.md +151 -0
  30. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/evals.json +74 -0
  31. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/SKILL.md +152 -0
  32. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/evals.json +74 -0
  33. package/src/gdskills/bundled/stacks/ci-github-gitlab/agent-refs.json +4 -0
  34. package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/eval.json +1295 -0
  35. package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/scout.json +26 -0
  36. package/src/gdskills/bundled/stacks/ci-github-gitlab/pack.json +41 -0
  37. package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/patterns.mdc +77 -0
  38. package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/security.mdc +144 -0
  39. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/SKILL.md +121 -0
  40. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/evals.json +73 -0
  41. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/SKILL.md +139 -0
  42. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/evals.json +73 -0
  43. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/SKILL.md +147 -0
  44. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/evals.json +74 -0
  45. package/src/gdskills/bundled/stacks/docker-k8s-terraform/agent-refs.json +4 -0
  46. package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/eval.json +865 -0
  47. package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/scout.json +16 -0
  48. package/src/gdskills/bundled/stacks/docker-k8s-terraform/pack.json +46 -0
  49. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/coding-style.mdc +74 -0
  50. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/patterns.mdc +81 -0
  51. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/security.mdc +146 -0
  52. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/testing.mdc +61 -0
  53. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/SKILL.md +151 -0
  54. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/evals.json +74 -0
  55. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/SKILL.md +135 -0
  56. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/evals.json +76 -0
  57. package/src/gdskills/bundled/stacks/php-laravel/agent-refs.json +4 -0
  58. package/src/gdskills/bundled/stacks/php-laravel/governance/eval.json +1829 -0
  59. package/src/gdskills/bundled/stacks/php-laravel/governance/scout.json +33 -0
  60. package/src/gdskills/bundled/stacks/php-laravel/pack.json +41 -0
  61. package/src/gdskills/bundled/stacks/php-laravel/rules/coding-style.mdc +82 -0
  62. package/src/gdskills/bundled/stacks/php-laravel/rules/patterns.mdc +80 -0
  63. package/src/gdskills/bundled/stacks/php-laravel/rules/security.mdc +80 -0
  64. package/src/gdskills/bundled/stacks/php-laravel/rules/testing.mdc +82 -0
  65. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/SKILL.md +143 -0
  66. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/evals.json +74 -0
  67. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/SKILL.md +126 -0
  68. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/evals.json +76 -0
  69. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/SKILL.md +140 -0
  70. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/evals.json +75 -0
  71. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/SKILL.md +124 -0
  72. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/evals.json +74 -0
  73. package/src/gdskills/bundled/stacks/ruby-rails/agent-refs.json +4 -0
  74. package/src/gdskills/bundled/stacks/ruby-rails/governance/eval.json +1673 -0
  75. package/src/gdskills/bundled/stacks/ruby-rails/governance/scout.json +33 -0
  76. package/src/gdskills/bundled/stacks/ruby-rails/pack.json +42 -0
  77. package/src/gdskills/bundled/stacks/ruby-rails/rules/coding-style.mdc +69 -0
  78. package/src/gdskills/bundled/stacks/ruby-rails/rules/patterns.mdc +93 -0
  79. package/src/gdskills/bundled/stacks/ruby-rails/rules/security.mdc +90 -0
  80. package/src/gdskills/bundled/stacks/ruby-rails/rules/testing.mdc +89 -0
  81. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/SKILL.md +143 -0
  82. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/evals.json +73 -0
  83. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/SKILL.md +134 -0
  84. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/evals.json +71 -0
  85. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/SKILL.md +141 -0
  86. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/evals.json +72 -0
  87. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/SKILL.md +125 -0
  88. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/evals.json +72 -0
  89. package/src/gdskills/bundled/stacks/sql-db/agent-refs.json +4 -0
  90. package/src/gdskills/bundled/stacks/sql-db/governance/eval.json +1829 -0
  91. package/src/gdskills/bundled/stacks/sql-db/governance/scout.json +30 -0
  92. package/src/gdskills/bundled/stacks/sql-db/pack.json +40 -0
  93. package/src/gdskills/bundled/stacks/sql-db/rules/coding-style.mdc +69 -0
  94. package/src/gdskills/bundled/stacks/sql-db/rules/patterns.mdc +134 -0
  95. package/src/gdskills/bundled/stacks/sql-db/rules/security.mdc +74 -0
  96. package/src/gdskills/bundled/stacks/sql-db/rules/testing.mdc +83 -0
  97. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/SKILL.md +147 -0
  98. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/evals.json +72 -0
  99. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/SKILL.md +132 -0
  100. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/evals.json +73 -0
  101. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/SKILL.md +153 -0
  102. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/evals.json +77 -0
  103. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/SKILL.md +129 -0
  104. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/evals.json +73 -0
@@ -0,0 +1,190 @@
1
+ ---
2
+ name: review-jev-risk
3
+ model_tier: light
4
+ description: |
5
+ Use when: an ADDITIONAL, machine-scored RISK MAP is wanted over every changed hunk —
6
+ never replacing any other reviewer. Dispatched by review-orchestrator in Wave B, via
7
+ the CLI (`keryx review jev-risk`), when `review.jev.risk: true` in
8
+ .metaproject/tasks.config.json and a Jev/OpenRouter credential is resolvable. It is not
9
+ an LLM sub-agent: there is nothing to dispatch through a platform-native agent
10
+ mechanism, and no prose is generated by a model — Jev only answers one `noul`
11
+ probability per (hunk, risk dimension) pair; keryx writes every word of every finding.
12
+ NOT for: judgement calls a risk dimension does not state (style, architecture beyond
13
+ risk-flagging — those stay with the reviewers that already cover them), and not a
14
+ substitute for a human reviewer reading the ranked map it produces.
15
+ triggers:
16
+ - "jev risk"
17
+ - "risk map"
18
+ - "review --jev-risk"
19
+ metadata:
20
+ author: "MrCipherSmith"
21
+ version: "1.0.0"
22
+ category: "review"
23
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
24
+ engine: "jev"
25
+ origin: "authored"
26
+ license: "MIT"
27
+ ---
28
+
29
+ # Review — Jev Risk (deterministic risk map of the diff)
30
+
31
+ An ADDITIONAL orchestrator reviewer, flow 332. It is a **keryx program**, not
32
+ an LLM sub-agent: `review-orchestrator` never dispatches a platform-native
33
+ agent for it, it runs `keryx review jev-risk` and reads the `--json` output.
34
+ Every finding is composed by keryx; Jev ("System One" on OpenRouter) supplies
35
+ only five `noul` probabilities per hunk — one per risk dimension — never
36
+ prose.
37
+
38
+ ---
39
+
40
+ ## What it does
41
+
42
+ 1. Takes every changed hunk from `keryx review scope`/`buildReviewScope`
43
+ (mechanical bulk already dropped).
44
+ 2. Computes deterministic facts FIRST, per hunk: a path class (auth/
45
+ permissions, crypto, migrations, schema, public API, config, concurrency
46
+ primitives, IO), the exported symbols it touches, its changed-line count,
47
+ and whether a test file elsewhere in the diff touches the same module.
48
+ 3. Asks Jev **one `noul` per risk dimension** per hunk — security-sensitive,
49
+ data/migration, public-API/contract change, concurrency, error-handling —
50
+ batched under the vendor's 64k token budget, capped at `--max-calls`
51
+ (default 150, counted as hunk x dimension pairs), with every drop
52
+ reported.
53
+ 4. Ranks hunks by combined risk (the MAX across its five dimensions — one
54
+ high-risk dimension is enough to draw attention) so a human reviewer knows
55
+ where to look first.
56
+ 5. Emits a finding ONLY for a hunk above threshold (default 0.7) AND with no
57
+ test touched nearby in the same diff (a fact) — severity capped at
58
+ `info`/`minor`, never higher: this flags attention, it does not assert a
59
+ defect.
60
+ 6. Additionally emits a **routing hint**: hunks above threshold with a
61
+ security or concurrency dimension are listed with a suggested reviewer
62
+ (`review-security-code`/`review-highload`) for the orchestrator to
63
+ consider dispatching — see "Orchestrator integration" below.
64
+
65
+ ---
66
+
67
+ ## Input Contract
68
+
69
+ Not dispatched with a prompt — invoked as a CLI command:
70
+
71
+ ```text
72
+ keryx review jev-risk (--diff <ref> | --pr <n> | --scope <scope.json>)
73
+ [--max-calls <n>] [--threshold <0..1>]
74
+ [--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
75
+ [--fixtures <dir>] [--json]
76
+ ```
77
+
78
+ `review-orchestrator` passes `--scope <scope.json>` — the same `keryx review
79
+ scope --json` file every other reviewer's dispatch already reads — so this
80
+ reviewer checks exactly the same hunks as everyone else in the round.
81
+
82
+ ---
83
+
84
+ ## Opt-in and privacy
85
+
86
+ Refuses before any read or network call unless `.metaproject/tasks.config.json`
87
+ declares:
88
+
89
+ ```json
90
+ { "review": { "jev": { "risk": true } } }
91
+ ```
92
+
93
+ Every hunk sent to Jev is redacted first through `src/security/service.ts` —
94
+ the same floor `review conform`/`review ci-triage`/`review jev-rules` already
95
+ apply. No cache or output file stores a credential.
96
+
97
+ ---
98
+
99
+ ## Output Contract
100
+
101
+ Emits a `REVIEW_RESULT`-shaped object matching
102
+ `.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
103
+ — `status`, `reviewer: "review-jev-risk"`, `summary`, `findings`, `stats` —
104
+ under `--json`, plus two orchestrator-facing extras: `ranked` (every scored
105
+ hunk, highest risk first) and `routingHints` (see above). Its findings merge
106
+ into the consolidated array exactly like any other reviewer's: same Quality
107
+ Gate, same dedup, same Wave C verification.
108
+
109
+ ---
110
+
111
+ ## Orchestrator integration
112
+
113
+ - `keryx review reviewers --json` lists it under `bundled` with
114
+ `"engine": "jev"`.
115
+ - `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
116
+ convention reviewers and flow 330's `review-jev-rules`, when the opt-in is
117
+ on and a Jev/OpenRouter credential resolves. When either is false it is
118
+ **skipped with the stated reason**, recorded in `Skipped reviewers` — never
119
+ silently absent.
120
+ - It is **ADDITIONAL**: it never replaces `review-security-code`,
121
+ `review-highload`, or any other pass.
122
+ - **Routing hint**: read `routingHints` from its `--json` output. Each entry
123
+ names a hunk location, the dimension that crossed threshold
124
+ (`security`/`concurrency`) and a `suggestedReviewer`
125
+ (`review-security-code`/`review-highload`). When that reviewer was not
126
+ already selected for the round, dispatch it too — the hint is advisory,
127
+ not a hard requirement, and the orchestrator's own path-based selection
128
+ always takes precedence when it already covers the same hunk.
129
+
130
+ ---
131
+
132
+ ### Shared laws (every reviewer)
133
+
134
+ 1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
135
+ name the input, call, or condition that reaches the code, you have an
136
+ observation, not a finding. Report it as `info` and say what would settle it.
137
+ 2. **Never flag the theoretical.** The path you describe must exist in the code
138
+ under review. Do not report a safe API because it could be misused, or a
139
+ pattern because it is often wrong elsewhere.
140
+ 3. **One finding per class, not one per occurrence.** When the same shape appears
141
+ at several sites, report it once and list every site. Ten findings that are one
142
+ finding hide the other nine problems.
143
+
144
+ This reviewer satisfies all three by construction: `evidence` always names the
145
+ hunk location, quotes the changed line, and states every dimension's
146
+ probability (never an unreproducible claim); a finding is synthesized only
147
+ against the hunk actually scored, never a hypothetical pattern; and each
148
+ hunk is a distinct site with its own probabilities, so there is no repeated
149
+ class across sites for this reviewer to collapse — the routing hint follows
150
+ the same discipline, naming the exact hunk and dimension that crossed
151
+ threshold rather than a general warning.
152
+
153
+ ---
154
+
155
+ ## Red Flags
156
+
157
+ | Rationalization | Why it is wrong |
158
+ |---|---|
159
+ | "Jev said 0.9, so this hunk is definitely dangerous" | A `noul` score is a probability across one dimension, not a verdict — severity stays capped at `info`/`minor` precisely because Jev alone is not authoritative |
160
+ | "No nearby test means nobody tested this at all" | `hasNearbyTest` requires the nearby test's OWN diff text to mention a touched symbol or import the hunk's module (`testHunkEvidence`), tightened after a live-check false negative — still a regex heuristic, not proof of coverage through an indirection it cannot see |
161
+ | "This hunk touches a .md file, so 'no nearby test' is a meaningful finding" | Docs (`.md`/`.txt`) hunks are excluded before scoring (`isNonCodeHunk`) after a live check where prose scored `public-api` — this cannot happen; `selection.hunksNotCode` reports the count |
162
+ | "No findings means the diff is low risk" | `--max-calls` bounds how many hunks are scored; a capped run reports `selection.hunksSkipped` for exactly this reason — read the selection stats before treating silence as clean |
163
+
164
+ ---
165
+
166
+ ## Verification
167
+
168
+ Before trusting a run's findings:
169
+
170
+ 1. Read `selection.hunksSkipped` in the `--json` output — a capped run
171
+ covered fewer than "every retained hunk" and the report should say so.
172
+ 2. Spot-check a handful of `ranked` entries against the actual hunk: does the
173
+ file/line range and the top dimension make sense for what actually
174
+ changed there? Docs (`.md`/`.txt`) hunks never reach `ranked` at all —
175
+ `selection.hunksNotCode` should account for every one of them in the
176
+ diff.
177
+ 3. Confirm every finding carries `reviewer: "review-jev-risk"` and a
178
+ `dedupe_key` — both are required for the orchestrator's Quality Gate and
179
+ Wave C verification to route it correctly.
180
+
181
+ ---
182
+
183
+ ## Scope Boundaries
184
+
185
+ | Concern | This skill | Use instead |
186
+ |---------|------------|-------------|
187
+ | A ranked risk map of the diff's hunks | YES | — |
188
+ | An actual security/concurrency finding | NO (routing hint only) | `review-security-code` / `review-highload` |
189
+ | Judging whether a reported finding is real | NO | `review-verifier` |
190
+ | Which user scenarios changed | NO | `review-jev-scenarios` |
@@ -0,0 +1,187 @@
1
+ ---
2
+ name: review-jev-scenarios
3
+ model_tier: light
4
+ description: |
5
+ Use when: an ADDITIONAL, machine-scored FUNCTIONAL review is wanted — which user
6
+ scenarios a PR likely changes — never replacing any other reviewer. Dispatched by
7
+ review-orchestrator in Wave B, via the CLI (`keryx review jev-scenarios`), when
8
+ `review.jev.scenarios: true` in .metaproject/tasks.config.json and a Jev/OpenRouter
9
+ credential is resolvable. It is not an LLM sub-agent: there is nothing to dispatch
10
+ through a platform-native agent mechanism, and no prose is generated by a model — Jev
11
+ only answers one `noul` "does this change alter this scenario's behaviour?" per
12
+ scenario whose linked code the diff touches; keryx writes every word of every finding.
13
+ NOT for: discovering scenarios nobody documented (it reads existing gdwiki
14
+ user-scenario pages, PRD requirement/scenario sections, and README/docs "how to"
15
+ sections — a scenario that exists only in someone's head is invisible to it), and not
16
+ a substitute for manually walking the checklist it produces.
17
+ triggers:
18
+ - "jev scenarios"
19
+ - "functional review"
20
+ - "review --jev-scenarios"
21
+ metadata:
22
+ author: "MrCipherSmith"
23
+ version: "1.0.0"
24
+ category: "review"
25
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
26
+ engine: "jev"
27
+ origin: "authored"
28
+ license: "MIT"
29
+ ---
30
+
31
+ # Review — Jev Scenarios (deterministic functional review)
32
+
33
+ An ADDITIONAL orchestrator reviewer, flow 332. It is a **keryx program**, not
34
+ an LLM sub-agent: `review-orchestrator` never dispatches a platform-native
35
+ agent for it, it runs `keryx review jev-scenarios` and reads the `--json`
36
+ output. Every finding is composed by keryx; Jev ("System One" on OpenRouter)
37
+ supplies only one `noul` probability per scenario — never prose.
38
+
39
+ ---
40
+
41
+ ## What it does
42
+
43
+ 1. Gathers user scenarios deterministically from three sources: gdwiki pages
44
+ of type `user-scenario` (`.metaproject/wiki/user-scenarios/**`), PRD
45
+ requirement/scenario sections (`docs/requirements/**`, any heading
46
+ matching scenario/requirement/user story/use case), and README/docs "how
47
+ to" sections (`README.md`, `docs/docs/**`).
48
+ 2. Extracts each scenario's own links to code — markdown links and inline
49
+ code-span file references, the same convention `src/wiki/backlinks.ts`
50
+ already established for the wiki's own code edges.
51
+ 3. Computes, per scenario, which of its linked files the diff actually
52
+ touches (a fact) and whether a test in the diff covers a touched link (a
53
+ fact).
54
+ 4. For every scenario with at least one touched link, asks Jev **one
55
+ `noul`**: "does this change alter this scenario's behaviour?" — batched
56
+ under the vendor's 64k token budget, capped at `--max-calls` (default
57
+ 150), with every drop reported.
58
+ 5. Emits a **manual-check list**: every scenario at/above threshold (default
59
+ 0.5), ranked, each with its touched links as evidence — for a human to
60
+ walk before merge.
61
+ 6. Emits a `minor` finding for each likely-affected scenario with **no test
62
+ in the diff covering it** (a fact) — severity fixed at `minor`: this
63
+ flags attention over a functional scenario, it does not assert a defect.
64
+
65
+ ---
66
+
67
+ ## Input Contract
68
+
69
+ Not dispatched with a prompt — invoked as a CLI command:
70
+
71
+ ```text
72
+ keryx review jev-scenarios (--diff <ref> | --pr <n> | --scope <scope.json>)
73
+ [--max-calls <n>] [--threshold <0..1>]
74
+ [--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
75
+ [--fixtures <dir>] [--json]
76
+ ```
77
+
78
+ `review-orchestrator` passes `--scope <scope.json>` — the same file every
79
+ other reviewer's dispatch already reads — so this reviewer's "touched by this
80
+ diff" facts agree with the round's own scope.
81
+
82
+ ---
83
+
84
+ ## Opt-in and privacy
85
+
86
+ Refuses before any read or network call unless `.metaproject/tasks.config.json`
87
+ declares:
88
+
89
+ ```json
90
+ { "review": { "jev": { "scenarios": true } } }
91
+ ```
92
+
93
+ Every scenario's text sent to Jev is redacted first through
94
+ `src/security/service.ts` — the same floor every other Jev-backed mode in
95
+ this repository already applies. No cache or output file stores a
96
+ credential.
97
+
98
+ ---
99
+
100
+ ## Output Contract
101
+
102
+ Emits a `REVIEW_RESULT`-shaped object matching
103
+ `.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
104
+ — `status`, `reviewer: "review-jev-scenarios"`, `summary`, `findings`,
105
+ `stats` — under `--json`, plus one orchestrator-facing extra: `checklist`
106
+ (the ranked manual-check list described above). Its findings merge into the
107
+ consolidated array exactly like any other reviewer's: same Quality Gate,
108
+ same dedup, same Wave C verification.
109
+
110
+ ---
111
+
112
+ ## Orchestrator integration
113
+
114
+ - `keryx review reviewers --json` lists it under `bundled` with
115
+ `"engine": "jev"`.
116
+ - `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
117
+ convention reviewers and flow 330's `review-jev-rules`/flow 332's
118
+ `review-jev-risk`, when the opt-in is on and a Jev/OpenRouter credential
119
+ resolves. When either is false it is **skipped with the stated reason**,
120
+ recorded in `Skipped reviewers` — never silently absent.
121
+ - It is **ADDITIONAL**: it never replaces any domain reviewer's own read of
122
+ the diff.
123
+ - The consolidated report's "Manual verification" section (or equivalent)
124
+ should surface `checklist` verbatim alongside any findings — a scenario
125
+ that scored below the finding threshold but above the checklist threshold
126
+ is still worth a human glance even though it produced no finding.
127
+
128
+ ---
129
+
130
+ ### Shared laws (every reviewer)
131
+
132
+ 1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
133
+ name the input, call, or condition that reaches the code, you have an
134
+ observation, not a finding. Report it as `info` and say what would settle it.
135
+ 2. **Never flag the theoretical.** The path you describe must exist in the code
136
+ under review. Do not report a safe API because it could be misused, or a
137
+ pattern because it is often wrong elsewhere.
138
+ 3. **One finding per class, not one per occurrence.** When the same shape appears
139
+ at several sites, report it once and list every site. Ten findings that are one
140
+ finding hide the other nine problems.
141
+
142
+ This reviewer satisfies all three by construction: severity is fixed at
143
+ `minor` — never higher — and `evidence` always names the scenario's linked
144
+ code and the exact files this diff touched (never an unreproducible claim);
145
+ a scenario is only asked about when a real link of its own is in the diff,
146
+ never a hypothetical connection; and each scenario is its own class with its
147
+ own `dedupe_key`, so there is nothing to collapse across sites.
148
+
149
+ ---
150
+
151
+ ## Red Flags
152
+
153
+ | Rationalization | Why it is wrong |
154
+ |---|---|
155
+ | "Jev said 0.9, so this scenario is definitely broken" | A `noul` score is a probability that behaviour changed, not a verdict — it only ever produces a `minor` finding, and only when paired with the separate "no covering test" fact |
156
+ | "This PRD scenario links to a huge, frequently-touched file, so the match is meaningless" | Down-weighted after a live check (`SCENARIO_LINK_FANOUT_THRESHOLD`): a link to a file referenced by more scenarios than that only still counts as touched when the scenario's own text names one of the diff's touched exported symbols — read the scenario's own text before assuming a surviving entry is still spurious |
157
+ | "No entries in checklist means nothing changed functionally" | Discovery only covers three sources (gdwiki `user-scenario` pages, `docs/requirements/**`, README/docs "how to" sections); a scenario that exists only in someone's head, or in a doc outside those three locations, is invisible to this reviewer |
158
+ | "checklist is empty so scenarioSources must be empty too" | `scenarioSources` lists every discovered scenario; `checklist` lists only those with a touched link at/above threshold — a project can have hundreds of scenarios discovered and zero touched by a given diff |
159
+
160
+ ---
161
+
162
+ ## Verification
163
+
164
+ Before trusting a run's findings:
165
+
166
+ 1. Read `selection.notApplicable`/`selection.scenariosSkipped` in the
167
+ `--json` output — a scenario with no touched link was never asked, and a
168
+ capped run may have skipped some that were.
169
+ 2. Spot-check a `checklist` entry against the scenario's own source
170
+ (`id` names the file, and for a PRD/README section, the heading): does
171
+ the touched link actually sit inside the part of the diff that matters
172
+ for that scenario, or did a large, widely-linked file produce a spurious
173
+ match?
174
+ 3. Confirm every finding carries `reviewer: "review-jev-scenarios"` and a
175
+ `dedupe_key` — both are required for the orchestrator's Quality Gate and
176
+ Wave C verification to route it correctly.
177
+
178
+ ---
179
+
180
+ ## Scope Boundaries
181
+
182
+ | Concern | This skill | Use instead |
183
+ |---------|------------|-------------|
184
+ | Which documented user scenarios a diff likely changes | YES | — |
185
+ | Discovering an undocumented scenario | NO | a human, or write the scenario down first |
186
+ | Judging whether a reported finding is real | NO | `review-verifier` |
187
+ | A risk map of the diff's hunks | NO | `review-jev-risk` |
@@ -46,22 +46,46 @@ review: it is a fact about the review, not about the diff.
46
46
 
47
47
  ## CLI-engine reviewers — dispatched as a command, not a sub-agent
48
48
 
49
- `review-jev-rules` (flow 330) is an ADDITIONAL reviewer, never replacing any
50
- other, and its dispatch mechanism differs from every reviewer named in
49
+ `review-jev-rules` (flow 330), `review-jev-risk` and `review-jev-scenarios`
50
+ (both flow 332), `review-jev-docs` and `review-jev-comments` (flow 333), and
51
+ `review-jev-contract` (flow 335) are ADDITIONAL reviewers, never replacing
52
+ any other, and their dispatch mechanism differs from every reviewer named in
51
53
  SKILL.md's Routing Table: there is no platform-native agent to invoke,
52
- because it is a deterministic **keryx program**. Run it with `keryx review
53
- jev-rules --scope <scope.json> --json` — the SAME `scope.json` every other
54
- Wave A/B reviewer's dispatch already reads, so it checks exactly the same
55
- hunks. Read its `--json` output as a `REVIEW_RESULT` and merge its `findings`
56
- into the consolidated array exactly like a sub-agent reviewer's: same
57
- Sub-Agent Report Quality Gate, same dedup, same Wave C verification.
58
-
59
- Gate it BEFORE running the command, not after: skip it — recorded in `Skipped
60
- reviewers` with the reason, never silently absent — when `review.jev.rules`
61
- is not `true` in `.metaproject/tasks.config.json`, or when no Jev/OpenRouter
62
- credential is resolvable. Either gate failing means `keryx review jev-rules`
63
- itself would refuse before any read or network call, so checking first saves
64
- a doomed dispatch.
54
+ because each is a deterministic **keryx program**. Run them with `keryx
55
+ review jev-rules --scope <scope.json> --json`, `keryx review jev-risk
56
+ --scope <scope.json> --json`, and `keryx review jev-scenarios --scope
57
+ <scope.json> --json` — the SAME `scope.json` every other Wave A/B reviewer's
58
+ dispatch already reads, so each checks exactly the same hunks.
59
+ `review-jev-docs` and `review-jev-comments` take a diff/PR target instead:
60
+ `keryx review jev-docs (--diff <ref>|--pr <n>) --json` and `keryx review
61
+ jev-comments --pr <n> --repo <owner/repo> --json`. `review-jev-contract`
62
+ also takes a diff/PR target, plus an optional linked flow: `keryx review
63
+ jev-contract (--diff <ref>|--pr <n>) [--flow <id>] --json` — a `--pr` target
64
+ checks the description's own claims; `--flow <id>` additionally checks that
65
+ flow's frozen acceptance criteria (reusing flow 328's `check-ac.ts`, see its
66
+ own SKILL.md). Read each `--json` output as a `REVIEW_RESULT` and merge its
67
+ `findings` into the consolidated array exactly like a sub-agent reviewer's:
68
+ same Sub-Agent Report Quality Gate, same dedup, same Wave C verification.
69
+ `review-jev-risk` additionally emits `ranked` and `routingHints`;
70
+ `review-jev-scenarios` additionally emits `checklist`; `review-jev-contract`
71
+ additionally emits `claims`, `budget`, and (when `--flow` was given)
72
+ `acCheck` — see each reviewer's own SKILL.md for what to do with its extra.
73
+
74
+ Gate each BEFORE running its command, not after: skip it — recorded in
75
+ `Skipped reviewers` with the reason, never silently absent — when its own
76
+ opt-in (`review.jev.rules` / `review.jev.risk` / `review.jev.scenarios` /
77
+ `review.jev.docs` / `review.jev.comments` / `review.jev.contract`) is not
78
+ `true` in `.metaproject/tasks.config.json`, or when no Jev/OpenRouter
79
+ credential is resolvable. Either gate failing means the command itself would
80
+ refuse before any read or network call, so checking first saves a doomed
81
+ dispatch. The six opt-ins are independent: any subset may be on.
82
+ `review-jev-comments` additionally needs a comment ledger to already exist
83
+ (`keryx review comments collect` run first) — it refuses, before any read,
84
+ when there is none. When `review.jev.contract` is on, dispatch
85
+ `review-jev-contract` with `--pr` (never `--diff`, which has no description
86
+ to check) and let its `findings` cover the description-vs-diff claim, in
87
+ place of the by-eye Stage 1 judgement described in SKILL.md's "This gate owns
88
+ the description-vs-diff comparison" — see that section's own note.
65
89
 
66
90
  `keryx review reviewers --json` marks a CLI-engine reviewer with
67
91
  `"engine": "jev"` on its `bundled` entry — the field's presence, not its
@@ -69,3 +93,52 @@ absence, is what distinguishes it from the default (an LLM sub-agent
69
93
  dispatch). A future engine-backed reviewer follows the same pattern: gate on
70
94
  its own opt-in and reachability, dispatch as a command, merge its `--json`
71
95
  output the same way.
96
+
97
+ ## `jev-triage` — advisory annotations, not a reviewer (Step 9b)
98
+
99
+ `review-jev-triage` (flow 340) is a DIFFERENT shape from every CLI-engine
100
+ reviewer above: it never produces a `findings` array of its own, and it is
101
+ never merged into the consolidated array. It runs AFTER the Sub-Agent Report
102
+ Quality Gate has already validated/merged/deduplicated the round's findings
103
+ (Step 9) and BEFORE Wave C verification (Step 10) — Step 9b in the Workflow
104
+ block — annotating the findings that already exist. It is advisory and
105
+ annotate-only by construction: nothing it returns drops or demotes a finding,
106
+ and no downstream step is permitted to treat its output as anything but a
107
+ hint.
108
+
109
+ Run it once per round, over the round's own consolidated findings:
110
+
111
+ ```bash
112
+ keryx review jev-triage --report <this round's review package dir> --json
113
+ ```
114
+
115
+ Gate it the same way as every other CLI-engine reviewer: skip it — recorded
116
+ in `Skipped reviewers` with the reason — when `review.jev.triage` is not
117
+ `true` in `.metaproject/tasks.config.json`, or when no Jev/OpenRouter
118
+ credential is resolvable. Either gate failing means the command itself would
119
+ refuse before any read or network call.
120
+
121
+ Its `--json` output is `{status, reviewer: "review-jev-triage", summary,
122
+ annotations, budget, tokens}` — never a `REVIEW_RESULT` merged into the
123
+ findings array. `annotations` carries three tracks, each read and used
124
+ differently in the report:
125
+
126
+ | Track | Shape | What to do with it |
127
+ |---|---|---|
128
+ | `severity_check` | `{id, p, flagged}` per blocker/major finding — `flagged` means `p < 0.4` | Show `flagged` findings as a soft note next to their existing severity ("Jev's own trigger+outcome check scored this low"). Never change the finding's `severity` field from this alone. |
129
+ | `merge_candidates` | `{id, a, b, reason, p}` per candidate pair | A high `p` is a suggestion to a human that `a` and `b` may be the same defect — surface it in the report as a note on both findings. Never merge, drop, or renumber either finding from this alone. |
130
+ | `verify_order` | `{id, p}`, sorted lowest-`p`-first | Feed this order into Wave C's own dispatch — verify the lowest-plausibility findings first — never as a reason to skip verifying any of them. |
131
+
132
+ `status` is `DONE_WITH_CONCERNS` — never `DONE` — when any `severity_check` is
133
+ `flagged` OR any `merge_candidates` pair scores `p >= 0.5`
134
+ (`LIKELY_DUPLICATE_THRESHOLD`, `src/commands/review-jev-triage.ts`). Live-check
135
+ calibration found same-file pairs that were NOT duplicates scoring around
136
+ 0.6, so a merge candidate above the threshold is a prompt to look, never a
137
+ merge — the status flip is a nudge to read the pair, not a verdict that the
138
+ findings are duplicates.
139
+
140
+ Show its annotations in the report as their own subsection, next to (not
141
+ inside) the findings they annotate — same separation the finding schema
142
+ itself draws between a reviewer's claim (`severity`, `problem`, …) and what
143
+ became of it (`disposition`), which a reviewer never states and this pass
144
+ does not either.
@@ -42,6 +42,7 @@ Review Orchestrator Progress:
42
42
  - [ ] Step 7: Stage 1 gate - spec compliance check (if issue/task provided)
43
43
  - [ ] Step 8: Dispatch selected reviewers in PARALLEL with reviewer-input schema
44
44
  - [ ] Step 9: Collect reviewer-finding schema results and handle NEEDS_CONTEXT
45
+ - [ ] Step 9b: `jev-triage` — advisory, annotate-only severity/duplicate/verify-order annotations over the consolidated findings (opt-in `review.jev.triage`)
45
46
  - [ ] Step 10: Wave C — dispatch `review-verifier` over the consolidated findings
46
47
  - [ ] Step 11: Sort by severity, deduplicate, emit unified report
47
48
  - [ ] Step 12: Emit the machine-readable `keryx:findings` block alongside the report
@@ -1168,7 +1169,7 @@ that already holds intent and diff side by side is this one.
1168
1169
 
1169
1170
  Run it on every round, not only the first. The drift the finding catches is
1170
1171
  created BY the rounds: the code moves to answer findings, the body does not, and
1171
- whoever reads the merge commit a year later reads the body.
1172
+ whoever reads the merge commit a year later reads the body. When `review.jev.contract` is on, dispatch `review-jev-contract --pr` and read its `findings` as this comparison's scored result instead of judging it by eye; the by-eye judgement is the fallback when that opt-in is off — `SKILL.detail.md` § "CLI-engine reviewers".
1172
1173
 
1173
1174
  ---
1174
1175
 
@@ -1177,7 +1178,7 @@ whoever reads the merge commit a year later reads the body.
1177
1178
  Dispatch selected reviewers in parallel when independent. Use waves when token budget is tight or when one reviewer needs another result:
1178
1179
 
1179
1180
  1. Wave A - core correctness/risk reviewers: logic, architecture, security/highload when selected.
1180
- 2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files. `review-jev-rules` (flow 330) also runs here, CLI-engine not sub-agent — `SKILL.detail.md` § "CLI-engine reviewers".
1181
+ 2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files. `review-jev-rules` (flow 330), `review-jev-risk`/`review-jev-scenarios` (flow 332), `review-jev-docs`/`review-jev-comments` (flow 333), `review-jev-contract` (flow 335) also run here, CLI-engine not sub-agent, `"engine": "jev"` in `keryx review reviewers --json`, gated on their own opt-in and a resolvable Jev/OpenRouter credential — `SKILL.detail.md` § "CLI-engine reviewers".
1181
1182
  3. Wave C - **verification**: `review-verifier` over the consolidated findings, when blockers/majors
1182
1183
  exist, `--verify` is set, or the PR is high-risk. See below.
1183
1184
 
@@ -1691,8 +1692,7 @@ CONTEXT_PATH: .metaproject/jobs/<job-name>/ai/context.md
1691
1692
  If provided and the file exists, read the context document **before** running scope detection.
1692
1693
  Use it to understand:
1693
1694
  - Intentionally chosen libraries and patterns (do not flag as issues)
1694
- - Architectural decisions already agreed upon
1695
- - Acceptance criteria to drive the Stage 1 spec compliance gate
1695
+ - Architectural decisions already agreed upon, and acceptance criteria driving the Stage 1 spec compliance gate
1696
1696
 
1697
1697
  If absent, proceed normally — context is optional and non-blocking.
1698
1698
 
@@ -0,0 +1,4 @@
1
+ {
2
+ "agents": [],
3
+ "note": "honest gate (DeepSeek deepseek-chat runner+judge, strictness high, trials 10, flow 337) ran and failed for all four skills -- trigger accuracy, not behavior content (every behavior scenario scored 0.9 or 1.0). No generated pair ships; pack stays stability: experimental. See governance/eval.json for the recorded reports and W1-stack-catalog.md's Wave 4 batch 5 implementation notes for the diagnosis."
4
+ }