@mrciphersmith/keryx 0.3.0 → 0.3.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (93) hide show
  1. package/README.md +4 -1
  2. package/dist/cli.js +13362 -7078
  3. package/dist/core.js +11706 -11330
  4. package/package.json +1 -1
  5. package/src/gdskills/bundled/agents/go-code-auditor.md +1 -1
  6. package/src/gdskills/bundled/agents/python-code-auditor.md +1 -1
  7. package/src/gdskills/bundled/install-manifest.json +271 -4
  8. package/src/gdskills/bundled/rules/core/model-selection.mdc +51 -0
  9. package/src/gdskills/bundled/skills/review/review-jev-comments/SKILL.md +184 -0
  10. package/src/gdskills/bundled/skills/review/review-jev-docs/SKILL.md +189 -0
  11. package/src/gdskills/bundled/skills/review/review-jev-risk/SKILL.md +190 -0
  12. package/src/gdskills/bundled/skills/review/review-jev-rules/SKILL.md +267 -0
  13. package/src/gdskills/bundled/skills/review/review-jev-scenarios/SKILL.md +187 -0
  14. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +39 -0
  15. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +1 -1
  16. package/src/gdskills/bundled/stacks/angular/agent-refs.json +4 -0
  17. package/src/gdskills/bundled/stacks/angular/governance/eval.json +1751 -0
  18. package/src/gdskills/bundled/stacks/angular/governance/scout.json +32 -0
  19. package/src/gdskills/bundled/stacks/angular/pack.json +55 -0
  20. package/src/gdskills/bundled/stacks/angular/rules/coding-style.mdc +82 -0
  21. package/src/gdskills/bundled/stacks/angular/rules/patterns.mdc +84 -0
  22. package/src/gdskills/bundled/stacks/angular/rules/security.mdc +70 -0
  23. package/src/gdskills/bundled/stacks/angular/rules/testing.mdc +73 -0
  24. package/src/gdskills/bundled/stacks/angular/skills/angular-build-fix/SKILL.md +127 -0
  25. package/src/gdskills/bundled/stacks/angular/skills/angular-build-fix/evals.json +72 -0
  26. package/src/gdskills/bundled/stacks/angular/skills/angular-code-review/SKILL.md +98 -0
  27. package/src/gdskills/bundled/stacks/angular/skills/angular-code-review/evals.json +73 -0
  28. package/src/gdskills/bundled/stacks/angular/skills/angular-implementation/SKILL.md +112 -0
  29. package/src/gdskills/bundled/stacks/angular/skills/angular-implementation/evals.json +74 -0
  30. package/src/gdskills/bundled/stacks/angular/skills/angular-testing/SKILL.md +102 -0
  31. package/src/gdskills/bundled/stacks/angular/skills/angular-testing/evals.json +71 -0
  32. package/src/gdskills/bundled/stacks/mobx/agent-refs.json +4 -0
  33. package/src/gdskills/bundled/stacks/mobx/governance/eval.json +904 -0
  34. package/src/gdskills/bundled/stacks/mobx/governance/scout.json +18 -0
  35. package/src/gdskills/bundled/stacks/mobx/pack.json +28 -0
  36. package/src/gdskills/bundled/stacks/mobx/rules/coding-style.mdc +91 -0
  37. package/src/gdskills/bundled/stacks/mobx/rules/patterns.mdc +122 -0
  38. package/src/gdskills/bundled/stacks/mobx/rules/security.mdc +56 -0
  39. package/src/gdskills/bundled/stacks/mobx/rules/testing.mdc +63 -0
  40. package/src/gdskills/bundled/stacks/mobx/skills/mobx-observable-testing/SKILL.md +124 -0
  41. package/src/gdskills/bundled/stacks/mobx/skills/mobx-observable-testing/evals.json +73 -0
  42. package/src/gdskills/bundled/stacks/mobx/skills/mobx-store-implementation/SKILL.md +149 -0
  43. package/src/gdskills/bundled/stacks/mobx/skills/mobx-store-implementation/evals.json +74 -0
  44. package/src/gdskills/bundled/stacks/nestjs/agent-refs.json +4 -0
  45. package/src/gdskills/bundled/stacks/nestjs/governance/eval.json +1308 -0
  46. package/src/gdskills/bundled/stacks/nestjs/governance/scout.json +34 -0
  47. package/src/gdskills/bundled/stacks/nestjs/pack.json +53 -0
  48. package/src/gdskills/bundled/stacks/nestjs/rules/coding-style.mdc +70 -0
  49. package/src/gdskills/bundled/stacks/nestjs/rules/patterns.mdc +83 -0
  50. package/src/gdskills/bundled/stacks/nestjs/rules/security.mdc +73 -0
  51. package/src/gdskills/bundled/stacks/nestjs/rules/testing.mdc +69 -0
  52. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-build-fix/SKILL.md +157 -0
  53. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-build-fix/evals.json +70 -0
  54. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-implementation/SKILL.md +129 -0
  55. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-implementation/evals.json +71 -0
  56. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-testing/SKILL.md +143 -0
  57. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-testing/evals.json +69 -0
  58. package/src/gdskills/bundled/stacks/nextjs-nuxt/agent-refs.json +4 -0
  59. package/src/gdskills/bundled/stacks/nextjs-nuxt/governance/eval.json +2413 -0
  60. package/src/gdskills/bundled/stacks/nextjs-nuxt/governance/scout.json +42 -0
  61. package/src/gdskills/bundled/stacks/nextjs-nuxt/pack.json +42 -0
  62. package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/coding-style.mdc +69 -0
  63. package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/patterns.mdc +88 -0
  64. package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/security.mdc +72 -0
  65. package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/testing.mdc +64 -0
  66. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-build-fix/SKILL.md +147 -0
  67. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-build-fix/evals.json +75 -0
  68. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-code-review/SKILL.md +118 -0
  69. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-code-review/evals.json +76 -0
  70. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-implementation/SKILL.md +135 -0
  71. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-implementation/evals.json +78 -0
  72. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-testing/SKILL.md +116 -0
  73. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-testing/evals.json +75 -0
  74. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-upgrade-migration/SKILL.md +134 -0
  75. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-upgrade-migration/evals.json +76 -0
  76. package/src/gdskills/bundled/stacks/vue/agent-refs.json +4 -0
  77. package/src/gdskills/bundled/stacks/vue/governance/eval.json +2215 -0
  78. package/src/gdskills/bundled/stacks/vue/governance/scout.json +42 -0
  79. package/src/gdskills/bundled/stacks/vue/pack.json +42 -0
  80. package/src/gdskills/bundled/stacks/vue/rules/coding-style.mdc +73 -0
  81. package/src/gdskills/bundled/stacks/vue/rules/patterns.mdc +84 -0
  82. package/src/gdskills/bundled/stacks/vue/rules/security.mdc +60 -0
  83. package/src/gdskills/bundled/stacks/vue/rules/testing.mdc +69 -0
  84. package/src/gdskills/bundled/stacks/vue/skills/vue-build-fix/SKILL.md +137 -0
  85. package/src/gdskills/bundled/stacks/vue/skills/vue-build-fix/evals.json +72 -0
  86. package/src/gdskills/bundled/stacks/vue/skills/vue-code-review/SKILL.md +120 -0
  87. package/src/gdskills/bundled/stacks/vue/skills/vue-code-review/evals.json +71 -0
  88. package/src/gdskills/bundled/stacks/vue/skills/vue-implementation/SKILL.md +122 -0
  89. package/src/gdskills/bundled/stacks/vue/skills/vue-implementation/evals.json +72 -0
  90. package/src/gdskills/bundled/stacks/vue/skills/vue-testing/SKILL.md +115 -0
  91. package/src/gdskills/bundled/stacks/vue/skills/vue-testing/evals.json +72 -0
  92. package/src/gdskills/bundled/stacks/vue/skills/vue2-to-vue3-migration/SKILL.md +135 -0
  93. package/src/gdskills/bundled/stacks/vue/skills/vue2-to-vue3-migration/evals.json +71 -0
@@ -0,0 +1,187 @@
1
+ ---
2
+ name: review-jev-scenarios
3
+ model_tier: light
4
+ description: |
5
+ Use when: an ADDITIONAL, machine-scored FUNCTIONAL review is wanted — which user
6
+ scenarios a PR likely changes — never replacing any other reviewer. Dispatched by
7
+ review-orchestrator in Wave B, via the CLI (`keryx review jev-scenarios`), when
8
+ `review.jev.scenarios: true` in .metaproject/tasks.config.json and a Jev/OpenRouter
9
+ credential is resolvable. It is not an LLM sub-agent: there is nothing to dispatch
10
+ through a platform-native agent mechanism, and no prose is generated by a model — Jev
11
+ only answers one `noul` "does this change alter this scenario's behaviour?" per
12
+ scenario whose linked code the diff touches; keryx writes every word of every finding.
13
+ NOT for: discovering scenarios nobody documented (it reads existing gdwiki
14
+ user-scenario pages, PRD requirement/scenario sections, and README/docs "how to"
15
+ sections — a scenario that exists only in someone's head is invisible to it), and not
16
+ a substitute for manually walking the checklist it produces.
17
+ triggers:
18
+ - "jev scenarios"
19
+ - "functional review"
20
+ - "review --jev-scenarios"
21
+ metadata:
22
+ author: "MrCipherSmith"
23
+ version: "1.0.0"
24
+ category: "review"
25
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
26
+ engine: "jev"
27
+ origin: "authored"
28
+ license: "MIT"
29
+ ---
30
+
31
+ # Review — Jev Scenarios (deterministic functional review)
32
+
33
+ An ADDITIONAL orchestrator reviewer, flow 332. It is a **keryx program**, not
34
+ an LLM sub-agent: `review-orchestrator` never dispatches a platform-native
35
+ agent for it, it runs `keryx review jev-scenarios` and reads the `--json`
36
+ output. Every finding is composed by keryx; Jev ("System One" on OpenRouter)
37
+ supplies only one `noul` probability per scenario — never prose.
38
+
39
+ ---
40
+
41
+ ## What it does
42
+
43
+ 1. Gathers user scenarios deterministically from three sources: gdwiki pages
44
+ of type `user-scenario` (`.metaproject/wiki/user-scenarios/**`), PRD
45
+ requirement/scenario sections (`docs/requirements/**`, any heading
46
+ matching scenario/requirement/user story/use case), and README/docs "how
47
+ to" sections (`README.md`, `docs/docs/**`).
48
+ 2. Extracts each scenario's own links to code — markdown links and inline
49
+ code-span file references, the same convention `src/wiki/backlinks.ts`
50
+ already established for the wiki's own code edges.
51
+ 3. Computes, per scenario, which of its linked files the diff actually
52
+ touches (a fact) and whether a test in the diff covers a touched link (a
53
+ fact).
54
+ 4. For every scenario with at least one touched link, asks Jev **one
55
+ `noul`**: "does this change alter this scenario's behaviour?" — batched
56
+ under the vendor's 64k token budget, capped at `--max-calls` (default
57
+ 150), with every drop reported.
58
+ 5. Emits a **manual-check list**: every scenario at/above threshold (default
59
+ 0.5), ranked, each with its touched links as evidence — for a human to
60
+ walk before merge.
61
+ 6. Emits a `minor` finding for each likely-affected scenario with **no test
62
+ in the diff covering it** (a fact) — severity fixed at `minor`: this
63
+ flags attention over a functional scenario, it does not assert a defect.
64
+
65
+ ---
66
+
67
+ ## Input Contract
68
+
69
+ Not dispatched with a prompt — invoked as a CLI command:
70
+
71
+ ```text
72
+ keryx review jev-scenarios (--diff <ref> | --pr <n> | --scope <scope.json>)
73
+ [--max-calls <n>] [--threshold <0..1>]
74
+ [--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
75
+ [--fixtures <dir>] [--json]
76
+ ```
77
+
78
+ `review-orchestrator` passes `--scope <scope.json>` — the same file every
79
+ other reviewer's dispatch already reads — so this reviewer's "touched by this
80
+ diff" facts agree with the round's own scope.
81
+
82
+ ---
83
+
84
+ ## Opt-in and privacy
85
+
86
+ Refuses before any read or network call unless `.metaproject/tasks.config.json`
87
+ declares:
88
+
89
+ ```json
90
+ { "review": { "jev": { "scenarios": true } } }
91
+ ```
92
+
93
+ Every scenario's text sent to Jev is redacted first through
94
+ `src/security/service.ts` — the same floor every other Jev-backed mode in
95
+ this repository already applies. No cache or output file stores a
96
+ credential.
97
+
98
+ ---
99
+
100
+ ## Output Contract
101
+
102
+ Emits a `REVIEW_RESULT`-shaped object matching
103
+ `.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
104
+ — `status`, `reviewer: "review-jev-scenarios"`, `summary`, `findings`,
105
+ `stats` — under `--json`, plus one orchestrator-facing extra: `checklist`
106
+ (the ranked manual-check list described above). Its findings merge into the
107
+ consolidated array exactly like any other reviewer's: same Quality Gate,
108
+ same dedup, same Wave C verification.
109
+
110
+ ---
111
+
112
+ ## Orchestrator integration
113
+
114
+ - `keryx review reviewers --json` lists it under `bundled` with
115
+ `"engine": "jev"`.
116
+ - `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
117
+ convention reviewers and flow 330's `review-jev-rules`/flow 332's
118
+ `review-jev-risk`, when the opt-in is on and a Jev/OpenRouter credential
119
+ resolves. When either is false it is **skipped with the stated reason**,
120
+ recorded in `Skipped reviewers` — never silently absent.
121
+ - It is **ADDITIONAL**: it never replaces any domain reviewer's own read of
122
+ the diff.
123
+ - The consolidated report's "Manual verification" section (or equivalent)
124
+ should surface `checklist` verbatim alongside any findings — a scenario
125
+ that scored below the finding threshold but above the checklist threshold
126
+ is still worth a human glance even though it produced no finding.
127
+
128
+ ---
129
+
130
+ ### Shared laws (every reviewer)
131
+
132
+ 1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
133
+ name the input, call, or condition that reaches the code, you have an
134
+ observation, not a finding. Report it as `info` and say what would settle it.
135
+ 2. **Never flag the theoretical.** The path you describe must exist in the code
136
+ under review. Do not report a safe API because it could be misused, or a
137
+ pattern because it is often wrong elsewhere.
138
+ 3. **One finding per class, not one per occurrence.** When the same shape appears
139
+ at several sites, report it once and list every site. Ten findings that are one
140
+ finding hide the other nine problems.
141
+
142
+ This reviewer satisfies all three by construction: severity is fixed at
143
+ `minor` — never higher — and `evidence` always names the scenario's linked
144
+ code and the exact files this diff touched (never an unreproducible claim);
145
+ a scenario is only asked about when a real link of its own is in the diff,
146
+ never a hypothetical connection; and each scenario is its own class with its
147
+ own `dedupe_key`, so there is nothing to collapse across sites.
148
+
149
+ ---
150
+
151
+ ## Red Flags
152
+
153
+ | Rationalization | Why it is wrong |
154
+ |---|---|
155
+ | "Jev said 0.9, so this scenario is definitely broken" | A `noul` score is a probability that behaviour changed, not a verdict — it only ever produces a `minor` finding, and only when paired with the separate "no covering test" fact |
156
+ | "This PRD scenario links to a huge, frequently-touched file, so the match is meaningless" | Down-weighted after a live check (`SCENARIO_LINK_FANOUT_THRESHOLD`): a link to a file referenced by more scenarios than that only still counts as touched when the scenario's own text names one of the diff's touched exported symbols — read the scenario's own text before assuming a surviving entry is still spurious |
157
+ | "No entries in checklist means nothing changed functionally" | Discovery only covers three sources (gdwiki `user-scenario` pages, `docs/requirements/**`, README/docs "how to" sections); a scenario that exists only in someone's head, or in a doc outside those three locations, is invisible to this reviewer |
158
+ | "checklist is empty so scenarioSources must be empty too" | `scenarioSources` lists every discovered scenario; `checklist` lists only those with a touched link at/above threshold — a project can have hundreds of scenarios discovered and zero touched by a given diff |
159
+
160
+ ---
161
+
162
+ ## Verification
163
+
164
+ Before trusting a run's findings:
165
+
166
+ 1. Read `selection.notApplicable`/`selection.scenariosSkipped` in the
167
+ `--json` output — a scenario with no touched link was never asked, and a
168
+ capped run may have skipped some that were.
169
+ 2. Spot-check a `checklist` entry against the scenario's own source
170
+ (`id` names the file, and for a PRD/README section, the heading): does
171
+ the touched link actually sit inside the part of the diff that matters
172
+ for that scenario, or did a large, widely-linked file produce a spurious
173
+ match?
174
+ 3. Confirm every finding carries `reviewer: "review-jev-scenarios"` and a
175
+ `dedupe_key` — both are required for the orchestrator's Quality Gate and
176
+ Wave C verification to route it correctly.
177
+
178
+ ---
179
+
180
+ ## Scope Boundaries
181
+
182
+ | Concern | This skill | Use instead |
183
+ |---------|------------|-------------|
184
+ | Which documented user scenarios a diff likely changes | YES | — |
185
+ | Discovering an undocumented scenario | NO | a human, or write the scenario down first |
186
+ | Judging whether a reported finding is real | NO | `review-verifier` |
187
+ | A risk map of the diff's hunks | NO | `review-jev-risk` |
@@ -43,3 +43,42 @@ suppressing it would trade real coverage for tidiness. Record the drift in
43
43
  `review_context` and name it once in the report, so the next person knows the
44
44
  profile is due a re-read. Never file it as a finding against the code under
45
45
  review: it is a fact about the review, not about the diff.
46
+
47
+ ## CLI-engine reviewers — dispatched as a command, not a sub-agent
48
+
49
+ `review-jev-rules` (flow 330), `review-jev-risk` and `review-jev-scenarios`
50
+ (both flow 332), and `review-jev-docs` and `review-jev-comments` (flow 333)
51
+ are ADDITIONAL reviewers, never replacing any other, and their dispatch
52
+ mechanism differs from every reviewer named in SKILL.md's Routing Table:
53
+ there is no platform-native agent to invoke, because each is a deterministic
54
+ **keryx program**. Run them with `keryx review jev-rules --scope
55
+ <scope.json> --json`, `keryx review jev-risk --scope <scope.json> --json`,
56
+ and `keryx review jev-scenarios --scope <scope.json> --json` — the SAME
57
+ `scope.json` every other Wave A/B reviewer's dispatch already reads, so each
58
+ checks exactly the same hunks. `review-jev-docs` and `review-jev-comments`
59
+ take a diff/PR target instead: `keryx review jev-docs (--diff <ref>|--pr
60
+ <n>) --json` and `keryx review jev-comments --pr <n> --repo <owner/repo>
61
+ --json`. Read each `--json` output as a `REVIEW_RESULT` and merge its
62
+ `findings` into the consolidated array exactly like a sub-agent reviewer's:
63
+ same Sub-Agent Report Quality Gate, same dedup, same Wave C verification.
64
+ `review-jev-risk` additionally emits `ranked` and `routingHints`;
65
+ `review-jev-scenarios` additionally emits `checklist` — see each reviewer's
66
+ own SKILL.md for what to do with its extra.
67
+
68
+ Gate each BEFORE running its command, not after: skip it — recorded in
69
+ `Skipped reviewers` with the reason, never silently absent — when its own
70
+ opt-in (`review.jev.rules` / `review.jev.risk` / `review.jev.scenarios` /
71
+ `review.jev.docs` / `review.jev.comments`) is not `true` in
72
+ `.metaproject/tasks.config.json`, or when no Jev/OpenRouter credential is
73
+ resolvable. Either gate failing means the command itself would refuse before
74
+ any read or network call, so checking first saves a doomed dispatch. The
75
+ five opt-ins are independent: any subset may be on. `review-jev-comments`
76
+ additionally needs a comment ledger to already exist (`keryx review comments
77
+ collect` run first) — it refuses, before any read, when there is none.
78
+
79
+ `keryx review reviewers --json` marks a CLI-engine reviewer with
80
+ `"engine": "jev"` on its `bundled` entry — the field's presence, not its
81
+ absence, is what distinguishes it from the default (an LLM sub-agent
82
+ dispatch). A future engine-backed reviewer follows the same pattern: gate on
83
+ its own opt-in and reachability, dispatch as a command, merge its `--json`
84
+ output the same way.
@@ -1177,7 +1177,7 @@ whoever reads the merge commit a year later reads the body.
1177
1177
  Dispatch selected reviewers in parallel when independent. Use waves when token budget is tight or when one reviewer needs another result:
1178
1178
 
1179
1179
  1. Wave A - core correctness/risk reviewers: logic, architecture, security/highload when selected.
1180
- 2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files.
1180
+ 2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files. `review-jev-rules` (flow 330), `review-jev-risk`/`review-jev-scenarios` (flow 332), `review-jev-docs`/`review-jev-comments` (flow 333) also run here, CLI-engine not sub-agent, `"engine": "jev"` in `keryx review reviewers --json`, gated on their own opt-in and a resolvable Jev/OpenRouter credential — `SKILL.detail.md` § "CLI-engine reviewers".
1181
1181
  3. Wave C - **verification**: `review-verifier` over the consolidated findings, when blockers/majors
1182
1182
  exist, `--verify` is set, or the PR is high-risk. See below.
1183
1183
 
@@ -0,0 +1,4 @@
1
+ {
2
+ "agents": [],
3
+ "note": "Demoted from stable (R1 review round 2, PR #719, N-M1): a realism sweep found angular-implementation's and angular-testing's positive/negative trigger prompts were trigger-prefix near-copies or scorer-avoidant phrasing rather than genuinely independent realistic requests. Rewritten with natural phrasing per the reviewer's own probes and NOT iterated against the router. Honest result: angular-implementation's trigger accuracy collapsed (TP 1/7) and angular-testing picked up a real false positive (an e2e/staging negative it should not select) -- both accepted as the true routing signal rather than polished back to a passing score. angular-build-fix and angular-code-review still pass cleanly. A pack needs every skill to pass, so angular reverts to stability: \"experimental\"; the generated agent-code-auditor/build-fixer pair is removed."
4
+ }