@mrciphersmith/keryx 0.3.1 → 0.3.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (104) hide show
  1. package/dist/cli.js +7310 -2471
  2. package/dist/core.js +116 -10
  3. package/package.json +1 -1
  4. package/src/gdskills/bundled/install-manifest.json +349 -2
  5. package/src/gdskills/bundled/rules/core/model-selection.mdc +18 -0
  6. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +1 -1
  7. package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
  8. package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
  9. package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
  10. package/src/gdskills/bundled/skills/review/review-jev-comments/SKILL.md +184 -0
  11. package/src/gdskills/bundled/skills/review/review-jev-contract/SKILL.md +193 -0
  12. package/src/gdskills/bundled/skills/review/review-jev-docs/SKILL.md +189 -0
  13. package/src/gdskills/bundled/skills/review/review-jev-risk/SKILL.md +190 -0
  14. package/src/gdskills/bundled/skills/review/review-jev-scenarios/SKILL.md +187 -0
  15. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +88 -15
  16. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +4 -4
  17. package/src/gdskills/bundled/stacks/c-cpp/agent-refs.json +4 -0
  18. package/src/gdskills/bundled/stacks/c-cpp/governance/eval.json +1777 -0
  19. package/src/gdskills/bundled/stacks/c-cpp/governance/scout.json +31 -0
  20. package/src/gdskills/bundled/stacks/c-cpp/pack.json +42 -0
  21. package/src/gdskills/bundled/stacks/c-cpp/rules/coding-style.mdc +80 -0
  22. package/src/gdskills/bundled/stacks/c-cpp/rules/patterns.mdc +87 -0
  23. package/src/gdskills/bundled/stacks/c-cpp/rules/security.mdc +90 -0
  24. package/src/gdskills/bundled/stacks/c-cpp/rules/testing.mdc +83 -0
  25. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/SKILL.md +153 -0
  26. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/evals.json +74 -0
  27. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/SKILL.md +132 -0
  28. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/evals.json +73 -0
  29. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/SKILL.md +151 -0
  30. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/evals.json +74 -0
  31. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/SKILL.md +152 -0
  32. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/evals.json +74 -0
  33. package/src/gdskills/bundled/stacks/ci-github-gitlab/agent-refs.json +4 -0
  34. package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/eval.json +1295 -0
  35. package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/scout.json +26 -0
  36. package/src/gdskills/bundled/stacks/ci-github-gitlab/pack.json +41 -0
  37. package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/patterns.mdc +77 -0
  38. package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/security.mdc +144 -0
  39. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/SKILL.md +121 -0
  40. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/evals.json +73 -0
  41. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/SKILL.md +139 -0
  42. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/evals.json +73 -0
  43. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/SKILL.md +147 -0
  44. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/evals.json +74 -0
  45. package/src/gdskills/bundled/stacks/docker-k8s-terraform/agent-refs.json +4 -0
  46. package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/eval.json +865 -0
  47. package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/scout.json +16 -0
  48. package/src/gdskills/bundled/stacks/docker-k8s-terraform/pack.json +46 -0
  49. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/coding-style.mdc +74 -0
  50. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/patterns.mdc +81 -0
  51. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/security.mdc +146 -0
  52. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/testing.mdc +61 -0
  53. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/SKILL.md +151 -0
  54. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/evals.json +74 -0
  55. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/SKILL.md +135 -0
  56. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/evals.json +76 -0
  57. package/src/gdskills/bundled/stacks/php-laravel/agent-refs.json +4 -0
  58. package/src/gdskills/bundled/stacks/php-laravel/governance/eval.json +1829 -0
  59. package/src/gdskills/bundled/stacks/php-laravel/governance/scout.json +33 -0
  60. package/src/gdskills/bundled/stacks/php-laravel/pack.json +41 -0
  61. package/src/gdskills/bundled/stacks/php-laravel/rules/coding-style.mdc +82 -0
  62. package/src/gdskills/bundled/stacks/php-laravel/rules/patterns.mdc +80 -0
  63. package/src/gdskills/bundled/stacks/php-laravel/rules/security.mdc +80 -0
  64. package/src/gdskills/bundled/stacks/php-laravel/rules/testing.mdc +82 -0
  65. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/SKILL.md +143 -0
  66. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/evals.json +74 -0
  67. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/SKILL.md +126 -0
  68. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/evals.json +76 -0
  69. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/SKILL.md +140 -0
  70. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/evals.json +75 -0
  71. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/SKILL.md +124 -0
  72. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/evals.json +74 -0
  73. package/src/gdskills/bundled/stacks/ruby-rails/agent-refs.json +4 -0
  74. package/src/gdskills/bundled/stacks/ruby-rails/governance/eval.json +1673 -0
  75. package/src/gdskills/bundled/stacks/ruby-rails/governance/scout.json +33 -0
  76. package/src/gdskills/bundled/stacks/ruby-rails/pack.json +42 -0
  77. package/src/gdskills/bundled/stacks/ruby-rails/rules/coding-style.mdc +69 -0
  78. package/src/gdskills/bundled/stacks/ruby-rails/rules/patterns.mdc +93 -0
  79. package/src/gdskills/bundled/stacks/ruby-rails/rules/security.mdc +90 -0
  80. package/src/gdskills/bundled/stacks/ruby-rails/rules/testing.mdc +89 -0
  81. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/SKILL.md +143 -0
  82. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/evals.json +73 -0
  83. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/SKILL.md +134 -0
  84. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/evals.json +71 -0
  85. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/SKILL.md +141 -0
  86. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/evals.json +72 -0
  87. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/SKILL.md +125 -0
  88. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/evals.json +72 -0
  89. package/src/gdskills/bundled/stacks/sql-db/agent-refs.json +4 -0
  90. package/src/gdskills/bundled/stacks/sql-db/governance/eval.json +1829 -0
  91. package/src/gdskills/bundled/stacks/sql-db/governance/scout.json +30 -0
  92. package/src/gdskills/bundled/stacks/sql-db/pack.json +40 -0
  93. package/src/gdskills/bundled/stacks/sql-db/rules/coding-style.mdc +69 -0
  94. package/src/gdskills/bundled/stacks/sql-db/rules/patterns.mdc +134 -0
  95. package/src/gdskills/bundled/stacks/sql-db/rules/security.mdc +74 -0
  96. package/src/gdskills/bundled/stacks/sql-db/rules/testing.mdc +83 -0
  97. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/SKILL.md +147 -0
  98. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/evals.json +72 -0
  99. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/SKILL.md +132 -0
  100. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/evals.json +73 -0
  101. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/SKILL.md +153 -0
  102. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/evals.json +77 -0
  103. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/SKILL.md +129 -0
  104. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/evals.json +73 -0
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: brainstorm
3
- description: "Use when exploring architecture decisions, tech choices, feature ideas, or any open-ended problem that benefits from multiple perspectives. NOT for: writing the chosen option up as a formal requirements document (use prd-creator)."
3
+ description: "Use when exploring architecture decisions, tech choices, feature ideas, or any open-ended problem with several options that benefits from multiple perspectives. NOT for: writing the chosen option up as a formal requirements document (use prd-creator)."
4
4
  triggers:
5
5
  - "brainstorm"
6
6
  - "explore options"
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: interviewer
3
- description: "Use when a request is ambiguous and must be pinned down BEFORE any context is collected — the entry-point interview that turns a vague or expensive ask into a scoped brief. This is the `custom`-intent gate job-orchestrator runs at 0.1.5. NOT for: clarifying implementation specifics AFTER context is already collected — use `interview` instead."
3
+ description: "Use when a request is ambiguous and must be pinned down BEFORE any context is collected — the entry-point interview that clarifies the real requirements and turns a vague or expensive ask into a scoped brief. This is the `custom`-intent gate job-orchestrator runs at 0.1.5. NOT for: clarifying implementation specifics AFTER context is already collected — use `interview` instead."
4
4
  triggers:
5
5
  - "ask questions"
6
6
  - "clarify requirements"
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: deploy
3
- description: "Use when deploying to any environment (staging, production) or when a deployment pipeline needs to run. NOT for the database schema changes a release depends on (use `db-migrate`)."
3
+ description: "Use when deploying a release to any environment (staging, production) or when a deployment pipeline needs to run. NOT for the database schema changes a release depends on (use `db-migrate`)."
4
4
  triggers:
5
5
  - "deploy"
6
6
  - "deployment"
@@ -0,0 +1,184 @@
1
+ ---
2
+ name: review-jev-comments
3
+ model_tier: light
4
+ description: |
5
+ Use when: an ADDITIONAL, machine-scored pass is wanted over open PR review comments
6
+ already in the existing ledger — never replacing any other reviewer, and never
7
+ auto-replying or auto-resolving. Dispatched by review-orchestrator in Wave B, via the
8
+ CLI (`keryx review jev-comments`), when `review.jev.comments: true` in
9
+ .metaproject/tasks.config.json and a Jev/OpenRouter credential is resolvable. It is not
10
+ an LLM sub-agent: there is nothing to dispatch through a platform-native agent
11
+ mechanism, and no prose is generated by a model — Jev only answers a `choice` among
12
+ resolved-by-fix/still-open/not-actionable/needs-escalation per open comment; keryx
13
+ writes every word of every finding.
14
+ NOT for: posting a reply, resolving a thread, or deciding what `keryx review comments
15
+ reply` sends — this reviewer only ever reads the ledger and labels.
16
+ triggers:
17
+ - "jev comments"
18
+ - "open comments"
19
+ - "review --jev-comments"
20
+ metadata:
21
+ author: "MrCipherSmith"
22
+ version: "1.0.0"
23
+ category: "review"
24
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
25
+ engine: "jev"
26
+ origin: "authored"
27
+ license: "MIT"
28
+ ---
29
+
30
+ # Review — Jev Comments (open PR comment triage)
31
+
32
+ An ADDITIONAL orchestrator reviewer, flow 333. It is a **keryx program**, not an
33
+ LLM sub-agent: `review-orchestrator` never dispatches a platform-native agent
34
+ for it, it runs `keryx review jev-comments` and reads the `--json` output.
35
+ Every finding is composed by keryx; Jev ("System One" on OpenRouter) supplies
36
+ only a `choice` label per comment — it never writes prose, and nothing it
37
+ returns is quoted verbatim into a finding.
38
+
39
+ ---
40
+
41
+ ## What it does
42
+
43
+ 1. Reads open comments from the EXISTING ledger
44
+ (`.metaproject/reviews/pr-comments/*.json`, written by `keryx review
45
+ comments collect`) — never re-collects from GitHub itself.
46
+ 2. Computes deterministic facts per open comment: commits after the comment's
47
+ timestamp touching its file (local `git log`, capped), the review thread's
48
+ resolved flag (best-effort, via one read-only GraphQL query — degrades to
49
+ `"unknown"` on any failure, never a guess), and its replies already on the
50
+ ledger.
51
+ 3. Asks Jev exactly ONE `choice` question per open comment (batched under the
52
+ vendor's 64k token budget): `resolved-by-fix` / `still-open` /
53
+ `not-actionable` / `needs-escalation`, given the comment, its replies, and
54
+ the code's current hunk(s) at that location.
55
+ 4. Synthesizes findings for `still-open` (severity `minor`) and
56
+ `needs-escalation` (severity `major`) only — `resolved-by-fix` and
57
+ `not-actionable` produce no finding. It never posts a reply and never
58
+ resolves a thread; the choice for every open comment is also written to an
59
+ advisory-label cache
60
+ (`.metaproject/data/review-jev-comments/labels__<repo>__<n>.json`) that
61
+ `keryx review comments reply` reads back and prints, informationally, next
62
+ to its own output — it never changes what that command sends.
63
+
64
+ ---
65
+
66
+ ## Input Contract
67
+
68
+ Not dispatched with a prompt — invoked as a CLI command:
69
+
70
+ ```text
71
+ keryx review jev-comments --pr <n> --repo <owner/repo>
72
+ [--model <jev-1.13|jev-latest>] [--fixtures <dir>] [--json]
73
+ ```
74
+
75
+ Requires a comment ledger to already exist for `--repo`/`--pr` (`keryx review
76
+ comments collect` first) — this reviewer refuses, before any read, when it
77
+ does not.
78
+
79
+ ---
80
+
81
+ ## Opt-in and privacy
82
+
83
+ Refuses before any ledger read or network call unless
84
+ `.metaproject/tasks.config.json` declares:
85
+
86
+ ```json
87
+ { "review": { "jev": { "comments": true } } }
88
+ ```
89
+
90
+ Every comment body, reply, and hunk sent to Jev is redacted first through
91
+ `src/security/service.ts` — the same floor every other Jev-backed reviewer
92
+ applies. No cache or output file stores a credential.
93
+
94
+ ---
95
+
96
+ ## Output Contract
97
+
98
+ Emits a `REVIEW_RESULT`-shaped object matching
99
+ `.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
100
+ — `status`, `reviewer: "review-jev-comments"`, `summary`, `findings`, `stats`
101
+ — under `--json`. Its findings merge into the consolidated array exactly like
102
+ any other reviewer's: same Quality Gate, same dedup, same Wave C
103
+ verification.
104
+
105
+ ### Class scope — required for `blocker` and `major`
106
+
107
+ A `needs-escalation` finding is `major` and always carries `class_scope`:
108
+ `sites` is the single `file:line` the comment is anchored to, and
109
+ `enumeration_method` states it is exactly that one comment — there is no
110
+ "class" beyond it, one comment is one finding, never a set to enumerate
111
+ further. `still-open` is `minor` and carries none, matching the schema's own
112
+ rule (`class_scope` required only for `blocker`/`major`).
113
+
114
+ ---
115
+
116
+ ## Orchestrator integration
117
+
118
+ - `keryx review reviewers --json` lists it under `bundled` with
119
+ `"engine": "jev"`.
120
+ - `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
121
+ convention reviewers, when the opt-in is on and a Jev/OpenRouter credential
122
+ resolves (`resolveJevApiKeyResolution`). When either is false it is
123
+ **skipped with the stated reason**, recorded in `Skipped reviewers` —
124
+ never silently absent.
125
+ - It is **ADDITIONAL** and strictly advisory over the comment-reply flow: its
126
+ label is printed by `keryx review comments reply`, never acted on
127
+ automatically.
128
+
129
+ ---
130
+
131
+ ## Red Flags
132
+
133
+ | Rationalization | Why it is wrong |
134
+ |---|---|
135
+ | "Jev said still-open, so post a reply now" | This reviewer never posts anything — a `still-open`/`needs-escalation` finding is a signal for a human or `keryx review comments reply`'s own logic, not an instruction to act |
136
+ | "The thread's resolved flag came back unknown, so treat it as resolved" | `unknown` is a distinct, honest state — it means the read failed, not that the thread is closed |
137
+ | "No findings means every comment was addressed" | Only `still-open`/`needs-escalation` produce findings; `resolved-by-fix` and `not-actionable` are correctly silent, but check `openComments` in `--json` to see how many were actually judged |
138
+
139
+ ---
140
+
141
+ ## Verification
142
+
143
+ Before trusting a run's findings:
144
+
145
+ 1. Read `openComments` in the `--json` output against the ledger's own
146
+ `unanswered so far` count from the last `comments collect` — they should
147
+ agree.
148
+ 2. Spot-check a `needs-escalation` finding against the comment text: does it
149
+ actually block progress, or is it a `still-open` case Jev over-called?
150
+ 3. Confirm every finding carries `reviewer: "review-jev-comments"` and a
151
+ `dedupe_key` — both are required for the orchestrator's Quality Gate and
152
+ Wave C verification to route it correctly.
153
+
154
+ ---
155
+
156
+ ## Iron Laws
157
+
158
+ ### Shared laws (every reviewer)
159
+
160
+ 1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
161
+ name the input, call, or condition that reaches the code, you have an
162
+ observation, not a finding. Report it as `info` and say what would settle it.
163
+ 2. **Never flag the theoretical.** The path you describe must exist in the code
164
+ under review. Do not report a safe API because it could be misused, or a
165
+ pattern because it is often wrong elsewhere.
166
+ 3. **One finding per class, not one per occurrence.** When the same shape appears
167
+ at several sites, report it once and list every site. Ten findings that are one
168
+ finding hide the other nine problems.
169
+
170
+ Severity levels are defined once, in `review-orchestrator/SKILL.md` →
171
+ **Severity (canonical)**. This reviewer does not restate them: it emits only
172
+ `minor` (`still-open`) or `major` (`needs-escalation`, always with
173
+ `class_scope`) — `blocker` never applies to an unanswered comment.
174
+
175
+ ---
176
+
177
+ ## Scope Boundaries
178
+
179
+ | Concern | This skill | Use instead |
180
+ |---------|------------|-------------|
181
+ | An open PR comment looks unaddressed or blocking | YES | — |
182
+ | Posting a reply, resolving a thread | NO | `keryx review comments reply`, a human |
183
+ | Judging whether a reported finding is real | NO | `review-verifier` |
184
+ | Collecting comments from GitHub | NO | `keryx review comments collect` |
@@ -0,0 +1,193 @@
1
+ ---
2
+ name: review-jev-contract
3
+ model_tier: light
4
+ description: |
5
+ Use when: an ADDITIONAL, machine-scored check of a PR description's own CLAIMS —
6
+ and, when a flow is linked, that flow's frozen acceptance criteria — against the diff
7
+ is wanted, never replacing any other reviewer. Dispatched by review-orchestrator in
8
+ Wave B, via the CLI (`keryx review jev-contract`), when `review.jev.contract: true` in
9
+ .metaproject/tasks.config.json and a Jev/OpenRouter credential is resolvable. It is not
10
+ an LLM sub-agent: there is nothing to dispatch through a platform-native agent
11
+ mechanism, and no prose is generated by a model — Jev only answers one `noul`
12
+ probability per claim; keryx writes every word of every finding. Replaces the
13
+ orchestrator's by-eye Stage 1 "description vs diff" judgement with a scored input when
14
+ the opt-in is on; the by-eye check stays the fallback when it is off.
15
+ NOT for: judging a flow's frozen criteria from scratch (that machinery is flow 328's
16
+ `check-ac.ts`, reused, not duplicated), and not a substitute for a human reviewer
17
+ reading the claim list it produces.
18
+ triggers:
19
+ - "jev contract"
20
+ - "check the PR description against the diff"
21
+ - "review --jev-contract"
22
+ metadata:
23
+ author: "MrCipherSmith"
24
+ version: "1.0.0"
25
+ category: "review"
26
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
27
+ engine: "jev"
28
+ origin: "authored"
29
+ license: "MIT"
30
+ ---
31
+
32
+ # Review — Jev Contract (PR-description claims and frozen acceptance criteria vs. the diff)
33
+
34
+ An ADDITIONAL orchestrator reviewer, flow 335. It is a **keryx program**, not
35
+ an LLM sub-agent: `review-orchestrator` never dispatches a platform-native
36
+ agent for it, it runs `keryx review jev-contract` and reads the `--json`
37
+ output. Every finding is composed by keryx; Jev ("System One" on OpenRouter)
38
+ supplies only one `noul` probability per claim — never prose.
39
+
40
+ ---
41
+
42
+ ## What it does
43
+
44
+ Two independent tracks, both advisory, both merged into one report:
45
+
46
+ 1. **Claims** — extracted deterministically from the PR description: every
47
+ bullet/numbered-list item is a claim, plus every prose sentence carrying a
48
+ verb cue (adds, fixes, removes, "does not change", tests). For each claim,
49
+ keryx computes deterministic facts FIRST — named files/symbols/flags
50
+ present in the diff, whether a test file was touched (for a "tests added"
51
+ claim), whether an EXPORTED symbol changed anywhere in the diff (for a "no
52
+ API change" claim) — then asks Jev **one `noul`**: "does the diff support
53
+ this claim?". A claim the facts directly CONTRADICT (e.g. "no API change"
54
+ with an exported symbol touched) is `major`, with `class_scope`,
55
+ regardless of what Jev answered — a contradiction is a fact, not an
56
+ opinion Jev could outvote. A claim scoring below threshold with no
57
+ contradiction is `minor` ("unsupported"). A claim at/above threshold with
58
+ no contradiction produces no finding.
59
+ 2. **Criteria** — the linked flow's frozen acceptance criteria, when
60
+ `--flow <id>` names one. This track reuses flow 328's
61
+ `src/flow/check-ac.ts` **wholesale**, through `runCheckAc`
62
+ (`src/commands/flow-check-ac.ts`) — the exact same diff acquisition, Jev
63
+ batching, cache and degrade path `keryx flow check-ac` already uses. No
64
+ criterion-checking logic is reimplemented for this reviewer.
65
+
66
+ ---
67
+
68
+ ## Input Contract
69
+
70
+ Not dispatched with a prompt — invoked as a CLI command:
71
+
72
+ ```text
73
+ keryx review jev-contract (--diff <ref> | --pr <n>) [--flow <id>]
74
+ [--max-calls <n>] [--threshold <0..1>]
75
+ [--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
76
+ [--fixtures <dir>] [--json]
77
+ ```
78
+
79
+ `--pr <n>` is required to check claims at all — a `--diff <ref>` target has
80
+ no PR description to extract claims from, so that track reports zero claims
81
+ (honestly, not a refusal). `--flow <id>` is independent of either target: it
82
+ names the flow whose frozen acceptance criteria are checked against the same
83
+ diff, when one is linked.
84
+
85
+ ---
86
+
87
+ ## Opt-in and privacy
88
+
89
+ Refuses before any read or network call unless `.metaproject/tasks.config.json`
90
+ declares:
91
+
92
+ ```json
93
+ { "review": { "jev": { "contract": true } } }
94
+ ```
95
+
96
+ Every claim and every matched diff hunk sent to Jev is redacted first through
97
+ `src/security/service.ts` — the same floor `review conform`/`review
98
+ jev-risk`/`review jev-rules` already apply. No cache or output file stores a
99
+ credential.
100
+
101
+ ---
102
+
103
+ ## Output Contract
104
+
105
+ Emits a `REVIEW_RESULT`-shaped object matching
106
+ `.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
107
+ — `status`, `reviewer: "review-jev-contract"`, `summary`, `findings`, `stats`
108
+ — under `--json`, plus two extras: `claims` (every extracted claim, its
109
+ intent, and its `noul` score) and `budget` (`--max-calls`, claims scored,
110
+ claims skipped). When `--flow` names a flow, an `acCheck` object is merged in
111
+ (`flowId`, `verdicts`, `jevAsked`, `summary`) — the same shape `keryx flow
112
+ check-ac` reports. Its findings merge into the consolidated array exactly
113
+ like any other reviewer's: same Quality Gate, same dedup, same Wave C
114
+ verification. A `major` finding always carries a schema-valid `class_scope`
115
+ object (`sites`, `enumeration_method`) — never a bare label.
116
+
117
+ ---
118
+
119
+ ## Orchestrator integration
120
+
121
+ - `keryx review reviewers --json` lists it under `bundled` with
122
+ `"engine": "jev"`.
123
+ - `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
124
+ convention reviewers and the other CLI-engine reviewers, when the opt-in is
125
+ on and a Jev/OpenRouter credential resolves. When either is false it is
126
+ **skipped with the stated reason**, recorded in `Skipped reviewers` — never
127
+ silently absent.
128
+ - It is **ADDITIONAL**: it never replaces any other pass. **When the opt-in is
129
+ on, it REPLACES the orchestrator's by-eye Stage 1 "description vs diff"
130
+ judgement with this scored input** — see `review-orchestrator/SKILL.md`'s
131
+ "This gate owns the description-vs-diff comparison" and `SKILL.detail.md`'s
132
+ "CLI-engine reviewers" section. When the opt-in is off, the by-eye check
133
+ stays the fallback, unchanged.
134
+
135
+ ---
136
+
137
+ ### Shared laws (every reviewer)
138
+
139
+ 1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
140
+ name the input, call, or condition that reaches the code, you have an
141
+ observation, not a finding. Report it as `info` and say what would settle it.
142
+ 2. **Never flag the theoretical.** The path you describe must exist in the code
143
+ under review. Do not report a safe API because it could be misused, or a
144
+ pattern because it is often wrong elsewhere.
145
+ 3. **One finding per class, not one per occurrence.** When the same shape appears
146
+ at several sites, report it once and list every site. Ten findings that are one
147
+ finding hide the other nine problems.
148
+
149
+ This reviewer satisfies all three by construction: `evidence` always names the
150
+ matched artefacts/facts and the claim text (never an unreproducible claim); a
151
+ finding is synthesized only against a claim actually extracted from the real
152
+ description, never a hypothetical one; and each claim is its own site, so
153
+ there is no repeated class across claims for this reviewer to collapse — a
154
+ `major` finding's `class_scope` still enumerates every matched region, not
155
+ just the first.
156
+
157
+ ---
158
+
159
+ ## Red Flags
160
+
161
+ | Rationalization | Why it is wrong |
162
+ |---|---|
163
+ | "Jev said 0.9, so the claim is definitely supported" | A `noul` score is a probability the claim is evidenced, not a verdict — a contradiction detected by the deterministic facts overrides it regardless of the score |
164
+ | "The description doesn't mention an API, so 'no change' can't be contradicted" | The contradiction check only runs when the claim's own text IS API-shaped (mentions api/public/export/interface/signature/contract) — a claim like "no change to behavior for CLI callers" is checked for tokens, but the exported-symbol contradiction path is intentionally narrower |
165
+ | "No claims extracted means the description said nothing" | `--diff` targets have no PR description at all — zero claims there is a target-shape fact, not a description-quality one; check `budget.claimsSkipped` too before reading silence as clean |
166
+ | "The AC track failed, so the whole review failed" | `acCheck` degrades to facts-only (never throws) exactly like `keryx flow check-ac` does — a Jev failure on the criteria track is advisory, same as on the claims track |
167
+
168
+ ---
169
+
170
+ ## Verification
171
+
172
+ Before trusting a run's findings:
173
+
174
+ 1. Read `budget.claimsSkipped` in the `--json` output — a capped run checked
175
+ fewer than "every extracted claim" and the report should say so.
176
+ 2. Spot-check a `major` finding's `class_scope.sites` against the actual
177
+ diff: does the named region really touch an exported symbol the claim
178
+ said would not change?
179
+ 3. Confirm every finding carries `reviewer: "review-jev-contract"` and a
180
+ `dedupe_key` — both are required for the orchestrator's Quality Gate and
181
+ Wave C verification to route it correctly.
182
+
183
+ ---
184
+
185
+ ## Scope Boundaries
186
+
187
+ | Concern | This skill | Use instead |
188
+ |---------|------------|-------------|
189
+ | Whether the PR description's own claims match the diff | YES | — |
190
+ | Whether a flow's frozen acceptance criteria are likely met | YES (via `--flow`, reusing `check-ac.ts`) | `keryx flow check-ac` directly, for a standalone advisory run |
191
+ | An actual security/concurrency finding | NO | `review-security-code` / `review-highload` |
192
+ | Judging whether a reported finding is real | NO | `review-verifier` |
193
+ | A ranked risk map of the diff's hunks | NO | `review-jev-risk` |
@@ -0,0 +1,189 @@
1
+ ---
2
+ name: review-jev-docs
3
+ model_tier: light
4
+ description: |
5
+ Use when: an ADDITIONAL, machine-scored pass is wanted over documentation sections that
6
+ may have gone STALE because of a diff — never replacing any other reviewer. Dispatched
7
+ by review-orchestrator in Wave B, via the CLI (`keryx review jev-docs`), when
8
+ `review.jev.docs: true` in .metaproject/tasks.config.json and a Jev/OpenRouter
9
+ credential is resolvable. It is not an LLM sub-agent: there is nothing to dispatch
10
+ through a platform-native agent mechanism, and no prose is generated by a model —
11
+ Jev only answers a `noul` staleness probability per doc section; keryx writes every
12
+ word of every finding, and one class of finding (a removed/renamed CLI flag still
13
+ documented) is found with no Jev call at all.
14
+ NOT for: judging whether documentation is well-written, complete, or accurate about
15
+ something the diff never touched — this reviewer only ever looks at sections
16
+ deterministically linked to code the diff changed.
17
+ triggers:
18
+ - "jev docs"
19
+ - "stale docs"
20
+ - "review --jev-docs"
21
+ metadata:
22
+ author: "MrCipherSmith"
23
+ version: "1.0.0"
24
+ category: "review"
25
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
26
+ engine: "jev"
27
+ origin: "authored"
28
+ license: "MIT"
29
+ ---
30
+
31
+ # Review — Jev Docs (stale-documentation detection)
32
+
33
+ An ADDITIONAL orchestrator reviewer, flow 333. It is a **keryx program**, not an
34
+ LLM sub-agent: `review-orchestrator` never dispatches a platform-native agent
35
+ for it, it runs `keryx review jev-docs` and reads the `--json` output. Every
36
+ finding is composed by keryx; Jev ("System One" on OpenRouter) supplies only a
37
+ `noul` staleness probability per doc section — it never writes prose, and
38
+ nothing it returns is quoted verbatim into a finding.
39
+
40
+ ---
41
+
42
+ ## What it does
43
+
44
+ 1. Splits every discovered documentation file — the default corpus is
45
+ USER-FACING docs only: `docs/**`, the root `README*` (never `CHANGELOG*`),
46
+ and gdwiki pages (`.metaproject/wiki/**`) — into sections by heading,
47
+ deterministically (`src/review/jev-docs.ts`'s `extractDocSections` — no
48
+ model call). `--include <glob>` (repeatable) widens the corpus back out
49
+ (skills, rules, anything project-specific); see Input Contract.
50
+ 2. Links each section to code deterministically — an explicit repo-relative
51
+ path it quotes, a backtick-quoted symbol that appears verbatim in one of
52
+ the diff's own hunks, or a `keryx <verb>` invocation whose command file the
53
+ diff changed. Only sections linked to code the diff CHANGES, and that the
54
+ diff does NOT itself edit, are candidates.
55
+ 3. Ranks linked sections by link strength — path mention > symbol mention >
56
+ verb mention; more distinct links to changed code within the same kind
57
+ rank higher — and asks Jev exactly ONE `noul` question per selected
58
+ section, up to `--max-calls` (default 30) and 8 per doc file, batched
59
+ under the vendor's 64k token budget, with the ranking basis and every drop
60
+ reported (`selection`), never silent, never alphabetical.
61
+ 4. Separately, and with NO Jev call: a CLI flag that disappears from a
62
+ changed file's diff (present on a removed line, absent from every added
63
+ line of the same file) and is still mentioned by an untouched doc section
64
+ is flagged deterministically.
65
+ 5. Synthesizes findings deterministically: `problem` names the section
66
+ (file + heading + line) and the code change that likely outdated it,
67
+ `suggested_fix` names the update target, `evidence` carries Jev's
68
+ probability (or, for the flag check, the deterministic reasoning).
69
+ Severity is always `minor`.
70
+
71
+ ---
72
+
73
+ ## Input Contract
74
+
75
+ Not dispatched with a prompt — invoked as a CLI command:
76
+
77
+ ```text
78
+ keryx review jev-docs (--diff <ref> | --pr <n>) [--max-calls <n>] [--threshold <0..1>]
79
+ [--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
80
+ [--fixtures <dir>] [--include <glob>]... [--json]
81
+ ```
82
+
83
+ `review-orchestrator` passes `--diff <ref>` (or `--pr <n>`) matching the same
84
+ target every other reviewer's dispatch checks.
85
+
86
+ ---
87
+
88
+ ## Opt-in and privacy
89
+
90
+ Refuses before any doc read or network call unless
91
+ `.metaproject/tasks.config.json` declares:
92
+
93
+ ```json
94
+ { "review": { "jev": { "docs": true } } }
95
+ ```
96
+
97
+ Every doc-section excerpt and hunk sent to Jev is redacted first through
98
+ `src/security/service.ts` — the same floor `review conform`/`review jev-rules`
99
+ already apply. No cache or output file stores a credential.
100
+
101
+ ---
102
+
103
+ ## Output Contract
104
+
105
+ Emits a `REVIEW_RESULT`-shaped object matching
106
+ `.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
107
+ — `status`, `reviewer: "review-jev-docs"`, `summary`, `findings`, `stats` —
108
+ under `--json`. Its findings merge into the consolidated array exactly like
109
+ any other reviewer's: same Quality Gate, same dedup, same Wave C
110
+ verification.
111
+
112
+ ### Class scope — required for `blocker` and `major`
113
+
114
+ Every finding this reviewer emits is `minor` (`docsFindingStats` caps it by
115
+ construction), so `class_scope` never applies here — `blocker`/`major`
116
+ findings, and the `class_scope` contract that comes with them, belong to the
117
+ reviewers named in `review-orchestrator/SKILL.md`'s own Finding Format
118
+ section.
119
+
120
+ ---
121
+
122
+ ## Orchestrator integration
123
+
124
+ - `keryx review reviewers --json` lists it under `bundled` with
125
+ `"engine": "jev"`.
126
+ - `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
127
+ convention reviewers, when the opt-in is on and a Jev/OpenRouter credential
128
+ resolves (`resolveJevApiKeyResolution`). When either is false it is
129
+ **skipped with the stated reason**, recorded in `Skipped reviewers` —
130
+ never silently absent.
131
+ - It is **ADDITIONAL**: it never replaces any documentation-adjacent finding
132
+ another reviewer already raises; the dedup pass merges rather than
133
+ double-counts.
134
+
135
+ ---
136
+
137
+ ## Red Flags
138
+
139
+ | Rationalization | Why it is wrong |
140
+ |---|---|
141
+ | "Jev said 0.9, so the doc is definitely wrong now" | A `noul` score is a probability, not a verdict — read the linked hunk yourself before editing the doc |
142
+ | "No findings means the docs are current" | The default corpus is `docs/**`/`README*`/gdwiki only — a project's `.metaproject/skills/**`/`rules/**` need `--include` to be scored at all; `--max-calls`/8-per-file also bound what gets scored (`selection.droppedSections`) |
143
+ | "This section wasn't linked, so it's fine" | Linking is deliberately narrow (explicit paths/symbols/verbs) — a section describing behaviour with no explicit code reference is out of this reviewer's reach by design, not proven current |
144
+
145
+ ---
146
+
147
+ ## Verification
148
+
149
+ Before trusting a run's findings:
150
+
151
+ 1. Read `selection.droppedSections` in the `--json` output — a capped run
152
+ covered less than every linked section.
153
+ 2. Spot-check a finding against the cited hunk: does the section's own text
154
+ actually conflict with what the diff changed?
155
+ 3. Confirm every finding carries `reviewer: "review-jev-docs"` and a
156
+ `dedupe_key` — both are required for the orchestrator's Quality Gate and
157
+ Wave C verification to route it correctly.
158
+
159
+ ---
160
+
161
+ ## Iron Laws
162
+
163
+ ### Shared laws (every reviewer)
164
+
165
+ 1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
166
+ name the input, call, or condition that reaches the code, you have an
167
+ observation, not a finding. Report it as `info` and say what would settle it.
168
+ 2. **Never flag the theoretical.** The path you describe must exist in the code
169
+ under review. Do not report a safe API because it could be misused, or a
170
+ pattern because it is often wrong elsewhere.
171
+ 3. **One finding per class, not one per occurrence.** When the same shape appears
172
+ at several sites, report it once and list every site. Ten findings that are one
173
+ finding hide the other nine problems.
174
+
175
+ Severity levels are defined once, in `review-orchestrator/SKILL.md` →
176
+ **Severity (canonical)**. This reviewer does not restate them — it only ever
177
+ emits `minor`, capped by construction (`docsFindingStats`), so the
178
+ `major`/`blocker` shapes never apply here.
179
+
180
+ ---
181
+
182
+ ## Scope Boundaries
183
+
184
+ | Concern | This skill | Use instead |
185
+ |---------|------------|-------------|
186
+ | A doc section linked to changed code likely went stale | YES | — |
187
+ | Documentation quality/completeness unrelated to this diff | NO | a human, or a dedicated docs review |
188
+ | Judging whether a reported finding is real | NO | `review-verifier` |
189
+ | Editing the documentation itself | NO | a human, following `suggested_fix` |