@mrciphersmith/keryx 0.3.2 → 0.3.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (164) hide show
  1. package/dist/cli.js +4745 -2482
  2. package/dist/core.js +66 -10
  3. package/package.json +1 -1
  4. package/src/gdskills/bundled/install-manifest.json +578 -2
  5. package/src/gdskills/bundled/rules/core/model-selection.mdc +18 -0
  6. package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +1 -1
  7. package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
  8. package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
  9. package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
  10. package/src/gdskills/bundled/skills/review/review-jev-contract/SKILL.md +193 -0
  11. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +81 -21
  12. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +4 -4
  13. package/src/gdskills/bundled/stacks/c-cpp/agent-refs.json +4 -0
  14. package/src/gdskills/bundled/stacks/c-cpp/governance/eval.json +1777 -0
  15. package/src/gdskills/bundled/stacks/c-cpp/governance/scout.json +31 -0
  16. package/src/gdskills/bundled/stacks/c-cpp/pack.json +42 -0
  17. package/src/gdskills/bundled/stacks/c-cpp/rules/coding-style.mdc +80 -0
  18. package/src/gdskills/bundled/stacks/c-cpp/rules/patterns.mdc +87 -0
  19. package/src/gdskills/bundled/stacks/c-cpp/rules/security.mdc +90 -0
  20. package/src/gdskills/bundled/stacks/c-cpp/rules/testing.mdc +83 -0
  21. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/SKILL.md +153 -0
  22. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/evals.json +74 -0
  23. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/SKILL.md +132 -0
  24. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/evals.json +73 -0
  25. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/SKILL.md +151 -0
  26. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/evals.json +74 -0
  27. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/SKILL.md +152 -0
  28. package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/evals.json +74 -0
  29. package/src/gdskills/bundled/stacks/ci-github-gitlab/agent-refs.json +4 -0
  30. package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/eval.json +1295 -0
  31. package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/scout.json +26 -0
  32. package/src/gdskills/bundled/stacks/ci-github-gitlab/pack.json +41 -0
  33. package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/patterns.mdc +77 -0
  34. package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/security.mdc +144 -0
  35. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/SKILL.md +121 -0
  36. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/evals.json +73 -0
  37. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/SKILL.md +139 -0
  38. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/evals.json +73 -0
  39. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/SKILL.md +147 -0
  40. package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/evals.json +74 -0
  41. package/src/gdskills/bundled/stacks/csharp-dotnet/agent-refs.json +4 -0
  42. package/src/gdskills/bundled/stacks/csharp-dotnet/governance/eval.json +1881 -0
  43. package/src/gdskills/bundled/stacks/csharp-dotnet/governance/scout.json +33 -0
  44. package/src/gdskills/bundled/stacks/csharp-dotnet/pack.json +38 -0
  45. package/src/gdskills/bundled/stacks/csharp-dotnet/rules/coding-style.mdc +100 -0
  46. package/src/gdskills/bundled/stacks/csharp-dotnet/rules/patterns.mdc +107 -0
  47. package/src/gdskills/bundled/stacks/csharp-dotnet/rules/security.mdc +86 -0
  48. package/src/gdskills/bundled/stacks/csharp-dotnet/rules/testing.mdc +89 -0
  49. package/src/gdskills/bundled/stacks/csharp-dotnet/skills/dotnet-build-fix/SKILL.md +143 -0
  50. package/src/gdskills/bundled/stacks/csharp-dotnet/skills/dotnet-build-fix/evals.json +77 -0
  51. package/src/gdskills/bundled/stacks/csharp-dotnet/skills/dotnet-code-review/SKILL.md +121 -0
  52. package/src/gdskills/bundled/stacks/csharp-dotnet/skills/dotnet-code-review/evals.json +77 -0
  53. package/src/gdskills/bundled/stacks/csharp-dotnet/skills/dotnet-implementation/SKILL.md +134 -0
  54. package/src/gdskills/bundled/stacks/csharp-dotnet/skills/dotnet-implementation/evals.json +76 -0
  55. package/src/gdskills/bundled/stacks/csharp-dotnet/skills/dotnet-testing/SKILL.md +130 -0
  56. package/src/gdskills/bundled/stacks/csharp-dotnet/skills/dotnet-testing/evals.json +77 -0
  57. package/src/gdskills/bundled/stacks/docker-k8s-terraform/agent-refs.json +4 -0
  58. package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/eval.json +865 -0
  59. package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/scout.json +16 -0
  60. package/src/gdskills/bundled/stacks/docker-k8s-terraform/pack.json +46 -0
  61. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/coding-style.mdc +74 -0
  62. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/patterns.mdc +81 -0
  63. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/security.mdc +146 -0
  64. package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/testing.mdc +61 -0
  65. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/SKILL.md +151 -0
  66. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/evals.json +74 -0
  67. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/SKILL.md +135 -0
  68. package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/evals.json +76 -0
  69. package/src/gdskills/bundled/stacks/flutter-dart/agent-refs.json +4 -0
  70. package/src/gdskills/bundled/stacks/flutter-dart/governance/eval.json +1849 -0
  71. package/src/gdskills/bundled/stacks/flutter-dart/governance/scout.json +33 -0
  72. package/src/gdskills/bundled/stacks/flutter-dart/pack.json +41 -0
  73. package/src/gdskills/bundled/stacks/flutter-dart/rules/coding-style.mdc +98 -0
  74. package/src/gdskills/bundled/stacks/flutter-dart/rules/patterns.mdc +88 -0
  75. package/src/gdskills/bundled/stacks/flutter-dart/rules/security.mdc +91 -0
  76. package/src/gdskills/bundled/stacks/flutter-dart/rules/testing.mdc +101 -0
  77. package/src/gdskills/bundled/stacks/flutter-dart/skills/flutter-build-fix/SKILL.md +134 -0
  78. package/src/gdskills/bundled/stacks/flutter-dart/skills/flutter-build-fix/evals.json +79 -0
  79. package/src/gdskills/bundled/stacks/flutter-dart/skills/flutter-code-review/SKILL.md +124 -0
  80. package/src/gdskills/bundled/stacks/flutter-dart/skills/flutter-code-review/evals.json +74 -0
  81. package/src/gdskills/bundled/stacks/flutter-dart/skills/flutter-implementation/SKILL.md +139 -0
  82. package/src/gdskills/bundled/stacks/flutter-dart/skills/flutter-implementation/evals.json +77 -0
  83. package/src/gdskills/bundled/stacks/flutter-dart/skills/flutter-testing/SKILL.md +134 -0
  84. package/src/gdskills/bundled/stacks/flutter-dart/skills/flutter-testing/evals.json +74 -0
  85. package/src/gdskills/bundled/stacks/kotlin-android/agent-refs.json +4 -0
  86. package/src/gdskills/bundled/stacks/kotlin-android/governance/eval.json +1889 -0
  87. package/src/gdskills/bundled/stacks/kotlin-android/governance/scout.json +34 -0
  88. package/src/gdskills/bundled/stacks/kotlin-android/pack.json +38 -0
  89. package/src/gdskills/bundled/stacks/kotlin-android/rules/coding-style.mdc +89 -0
  90. package/src/gdskills/bundled/stacks/kotlin-android/rules/patterns.mdc +96 -0
  91. package/src/gdskills/bundled/stacks/kotlin-android/rules/security.mdc +90 -0
  92. package/src/gdskills/bundled/stacks/kotlin-android/rules/testing.mdc +89 -0
  93. package/src/gdskills/bundled/stacks/kotlin-android/skills/compose-implementation/SKILL.md +150 -0
  94. package/src/gdskills/bundled/stacks/kotlin-android/skills/compose-implementation/evals.json +77 -0
  95. package/src/gdskills/bundled/stacks/kotlin-android/skills/kotlin-android-build-fix/SKILL.md +151 -0
  96. package/src/gdskills/bundled/stacks/kotlin-android/skills/kotlin-android-build-fix/evals.json +76 -0
  97. package/src/gdskills/bundled/stacks/kotlin-android/skills/kotlin-android-code-review/SKILL.md +139 -0
  98. package/src/gdskills/bundled/stacks/kotlin-android/skills/kotlin-android-code-review/evals.json +78 -0
  99. package/src/gdskills/bundled/stacks/kotlin-android/skills/kotlin-android-testing/SKILL.md +131 -0
  100. package/src/gdskills/bundled/stacks/kotlin-android/skills/kotlin-android-testing/evals.json +77 -0
  101. package/src/gdskills/bundled/stacks/php-laravel/agent-refs.json +4 -0
  102. package/src/gdskills/bundled/stacks/php-laravel/governance/eval.json +1829 -0
  103. package/src/gdskills/bundled/stacks/php-laravel/governance/scout.json +33 -0
  104. package/src/gdskills/bundled/stacks/php-laravel/pack.json +41 -0
  105. package/src/gdskills/bundled/stacks/php-laravel/rules/coding-style.mdc +82 -0
  106. package/src/gdskills/bundled/stacks/php-laravel/rules/patterns.mdc +80 -0
  107. package/src/gdskills/bundled/stacks/php-laravel/rules/security.mdc +80 -0
  108. package/src/gdskills/bundled/stacks/php-laravel/rules/testing.mdc +82 -0
  109. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/SKILL.md +143 -0
  110. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/evals.json +74 -0
  111. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/SKILL.md +126 -0
  112. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/evals.json +76 -0
  113. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/SKILL.md +140 -0
  114. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/evals.json +75 -0
  115. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/SKILL.md +124 -0
  116. package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/evals.json +74 -0
  117. package/src/gdskills/bundled/stacks/ruby-rails/agent-refs.json +4 -0
  118. package/src/gdskills/bundled/stacks/ruby-rails/governance/eval.json +1673 -0
  119. package/src/gdskills/bundled/stacks/ruby-rails/governance/scout.json +33 -0
  120. package/src/gdskills/bundled/stacks/ruby-rails/pack.json +42 -0
  121. package/src/gdskills/bundled/stacks/ruby-rails/rules/coding-style.mdc +69 -0
  122. package/src/gdskills/bundled/stacks/ruby-rails/rules/patterns.mdc +93 -0
  123. package/src/gdskills/bundled/stacks/ruby-rails/rules/security.mdc +90 -0
  124. package/src/gdskills/bundled/stacks/ruby-rails/rules/testing.mdc +89 -0
  125. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/SKILL.md +143 -0
  126. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/evals.json +73 -0
  127. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/SKILL.md +134 -0
  128. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/evals.json +71 -0
  129. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/SKILL.md +141 -0
  130. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/evals.json +72 -0
  131. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/SKILL.md +125 -0
  132. package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/evals.json +72 -0
  133. package/src/gdskills/bundled/stacks/sql-db/agent-refs.json +4 -0
  134. package/src/gdskills/bundled/stacks/sql-db/governance/eval.json +1829 -0
  135. package/src/gdskills/bundled/stacks/sql-db/governance/scout.json +30 -0
  136. package/src/gdskills/bundled/stacks/sql-db/pack.json +40 -0
  137. package/src/gdskills/bundled/stacks/sql-db/rules/coding-style.mdc +69 -0
  138. package/src/gdskills/bundled/stacks/sql-db/rules/patterns.mdc +134 -0
  139. package/src/gdskills/bundled/stacks/sql-db/rules/security.mdc +74 -0
  140. package/src/gdskills/bundled/stacks/sql-db/rules/testing.mdc +83 -0
  141. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/SKILL.md +147 -0
  142. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/evals.json +72 -0
  143. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/SKILL.md +132 -0
  144. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/evals.json +73 -0
  145. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/SKILL.md +153 -0
  146. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/evals.json +77 -0
  147. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/SKILL.md +129 -0
  148. package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/evals.json +73 -0
  149. package/src/gdskills/bundled/stacks/swift-ios/agent-refs.json +4 -0
  150. package/src/gdskills/bundled/stacks/swift-ios/governance/eval.json +1803 -0
  151. package/src/gdskills/bundled/stacks/swift-ios/governance/scout.json +32 -0
  152. package/src/gdskills/bundled/stacks/swift-ios/pack.json +38 -0
  153. package/src/gdskills/bundled/stacks/swift-ios/rules/coding-style.mdc +92 -0
  154. package/src/gdskills/bundled/stacks/swift-ios/rules/patterns.mdc +112 -0
  155. package/src/gdskills/bundled/stacks/swift-ios/rules/security.mdc +78 -0
  156. package/src/gdskills/bundled/stacks/swift-ios/rules/testing.mdc +90 -0
  157. package/src/gdskills/bundled/stacks/swift-ios/skills/swift-build-fix/SKILL.md +144 -0
  158. package/src/gdskills/bundled/stacks/swift-ios/skills/swift-build-fix/evals.json +75 -0
  159. package/src/gdskills/bundled/stacks/swift-ios/skills/swift-code-review/SKILL.md +122 -0
  160. package/src/gdskills/bundled/stacks/swift-ios/skills/swift-code-review/evals.json +75 -0
  161. package/src/gdskills/bundled/stacks/swift-ios/skills/swift-testing/SKILL.md +131 -0
  162. package/src/gdskills/bundled/stacks/swift-ios/skills/swift-testing/evals.json +75 -0
  163. package/src/gdskills/bundled/stacks/swift-ios/skills/swiftui-implementation/SKILL.md +149 -0
  164. package/src/gdskills/bundled/stacks/swift-ios/skills/swiftui-implementation/evals.json +76 -0
@@ -0,0 +1,193 @@
1
+ ---
2
+ name: review-jev-contract
3
+ model_tier: light
4
+ description: |
5
+ Use when: an ADDITIONAL, machine-scored check of a PR description's own CLAIMS —
6
+ and, when a flow is linked, that flow's frozen acceptance criteria — against the diff
7
+ is wanted, never replacing any other reviewer. Dispatched by review-orchestrator in
8
+ Wave B, via the CLI (`keryx review jev-contract`), when `review.jev.contract: true` in
9
+ .metaproject/tasks.config.json and a Jev/OpenRouter credential is resolvable. It is not
10
+ an LLM sub-agent: there is nothing to dispatch through a platform-native agent
11
+ mechanism, and no prose is generated by a model — Jev only answers one `noul`
12
+ probability per claim; keryx writes every word of every finding. Replaces the
13
+ orchestrator's by-eye Stage 1 "description vs diff" judgement with a scored input when
14
+ the opt-in is on; the by-eye check stays the fallback when it is off.
15
+ NOT for: judging a flow's frozen criteria from scratch (that machinery is flow 328's
16
+ `check-ac.ts`, reused, not duplicated), and not a substitute for a human reviewer
17
+ reading the claim list it produces.
18
+ triggers:
19
+ - "jev contract"
20
+ - "check the PR description against the diff"
21
+ - "review --jev-contract"
22
+ metadata:
23
+ author: "MrCipherSmith"
24
+ version: "1.0.0"
25
+ category: "review"
26
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
27
+ engine: "jev"
28
+ origin: "authored"
29
+ license: "MIT"
30
+ ---
31
+
32
+ # Review — Jev Contract (PR-description claims and frozen acceptance criteria vs. the diff)
33
+
34
+ An ADDITIONAL orchestrator reviewer, flow 335. It is a **keryx program**, not
35
+ an LLM sub-agent: `review-orchestrator` never dispatches a platform-native
36
+ agent for it, it runs `keryx review jev-contract` and reads the `--json`
37
+ output. Every finding is composed by keryx; Jev ("System One" on OpenRouter)
38
+ supplies only one `noul` probability per claim — never prose.
39
+
40
+ ---
41
+
42
+ ## What it does
43
+
44
+ Two independent tracks, both advisory, both merged into one report:
45
+
46
+ 1. **Claims** — extracted deterministically from the PR description: every
47
+ bullet/numbered-list item is a claim, plus every prose sentence carrying a
48
+ verb cue (adds, fixes, removes, "does not change", tests). For each claim,
49
+ keryx computes deterministic facts FIRST — named files/symbols/flags
50
+ present in the diff, whether a test file was touched (for a "tests added"
51
+ claim), whether an EXPORTED symbol changed anywhere in the diff (for a "no
52
+ API change" claim) — then asks Jev **one `noul`**: "does the diff support
53
+ this claim?". A claim the facts directly CONTRADICT (e.g. "no API change"
54
+ with an exported symbol touched) is `major`, with `class_scope`,
55
+ regardless of what Jev answered — a contradiction is a fact, not an
56
+ opinion Jev could outvote. A claim scoring below threshold with no
57
+ contradiction is `minor` ("unsupported"). A claim at/above threshold with
58
+ no contradiction produces no finding.
59
+ 2. **Criteria** — the linked flow's frozen acceptance criteria, when
60
+ `--flow <id>` names one. This track reuses flow 328's
61
+ `src/flow/check-ac.ts` **wholesale**, through `runCheckAc`
62
+ (`src/commands/flow-check-ac.ts`) — the exact same diff acquisition, Jev
63
+ batching, cache and degrade path `keryx flow check-ac` already uses. No
64
+ criterion-checking logic is reimplemented for this reviewer.
65
+
66
+ ---
67
+
68
+ ## Input Contract
69
+
70
+ Not dispatched with a prompt — invoked as a CLI command:
71
+
72
+ ```text
73
+ keryx review jev-contract (--diff <ref> | --pr <n>) [--flow <id>]
74
+ [--max-calls <n>] [--threshold <0..1>]
75
+ [--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
76
+ [--fixtures <dir>] [--json]
77
+ ```
78
+
79
+ `--pr <n>` is required to check claims at all — a `--diff <ref>` target has
80
+ no PR description to extract claims from, so that track reports zero claims
81
+ (honestly, not a refusal). `--flow <id>` is independent of either target: it
82
+ names the flow whose frozen acceptance criteria are checked against the same
83
+ diff, when one is linked.
84
+
85
+ ---
86
+
87
+ ## Opt-in and privacy
88
+
89
+ Refuses before any read or network call unless `.metaproject/tasks.config.json`
90
+ declares:
91
+
92
+ ```json
93
+ { "review": { "jev": { "contract": true } } }
94
+ ```
95
+
96
+ Every claim and every matched diff hunk sent to Jev is redacted first through
97
+ `src/security/service.ts` — the same floor `review conform`/`review
98
+ jev-risk`/`review jev-rules` already apply. No cache or output file stores a
99
+ credential.
100
+
101
+ ---
102
+
103
+ ## Output Contract
104
+
105
+ Emits a `REVIEW_RESULT`-shaped object matching
106
+ `.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
107
+ — `status`, `reviewer: "review-jev-contract"`, `summary`, `findings`, `stats`
108
+ — under `--json`, plus two extras: `claims` (every extracted claim, its
109
+ intent, and its `noul` score) and `budget` (`--max-calls`, claims scored,
110
+ claims skipped). When `--flow` names a flow, an `acCheck` object is merged in
111
+ (`flowId`, `verdicts`, `jevAsked`, `summary`) — the same shape `keryx flow
112
+ check-ac` reports. Its findings merge into the consolidated array exactly
113
+ like any other reviewer's: same Quality Gate, same dedup, same Wave C
114
+ verification. A `major` finding always carries a schema-valid `class_scope`
115
+ object (`sites`, `enumeration_method`) — never a bare label.
116
+
117
+ ---
118
+
119
+ ## Orchestrator integration
120
+
121
+ - `keryx review reviewers --json` lists it under `bundled` with
122
+ `"engine": "jev"`.
123
+ - `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
124
+ convention reviewers and the other CLI-engine reviewers, when the opt-in is
125
+ on and a Jev/OpenRouter credential resolves. When either is false it is
126
+ **skipped with the stated reason**, recorded in `Skipped reviewers` — never
127
+ silently absent.
128
+ - It is **ADDITIONAL**: it never replaces any other pass. **When the opt-in is
129
+ on, it REPLACES the orchestrator's by-eye Stage 1 "description vs diff"
130
+ judgement with this scored input** — see `review-orchestrator/SKILL.md`'s
131
+ "This gate owns the description-vs-diff comparison" and `SKILL.detail.md`'s
132
+ "CLI-engine reviewers" section. When the opt-in is off, the by-eye check
133
+ stays the fallback, unchanged.
134
+
135
+ ---
136
+
137
+ ### Shared laws (every reviewer)
138
+
139
+ 1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
140
+ name the input, call, or condition that reaches the code, you have an
141
+ observation, not a finding. Report it as `info` and say what would settle it.
142
+ 2. **Never flag the theoretical.** The path you describe must exist in the code
143
+ under review. Do not report a safe API because it could be misused, or a
144
+ pattern because it is often wrong elsewhere.
145
+ 3. **One finding per class, not one per occurrence.** When the same shape appears
146
+ at several sites, report it once and list every site. Ten findings that are one
147
+ finding hide the other nine problems.
148
+
149
+ This reviewer satisfies all three by construction: `evidence` always names the
150
+ matched artefacts/facts and the claim text (never an unreproducible claim); a
151
+ finding is synthesized only against a claim actually extracted from the real
152
+ description, never a hypothetical one; and each claim is its own site, so
153
+ there is no repeated class across claims for this reviewer to collapse — a
154
+ `major` finding's `class_scope` still enumerates every matched region, not
155
+ just the first.
156
+
157
+ ---
158
+
159
+ ## Red Flags
160
+
161
+ | Rationalization | Why it is wrong |
162
+ |---|---|
163
+ | "Jev said 0.9, so the claim is definitely supported" | A `noul` score is a probability the claim is evidenced, not a verdict — a contradiction detected by the deterministic facts overrides it regardless of the score |
164
+ | "The description doesn't mention an API, so 'no change' can't be contradicted" | The contradiction check only runs when the claim's own text IS API-shaped (mentions api/public/export/interface/signature/contract) — a claim like "no change to behavior for CLI callers" is checked for tokens, but the exported-symbol contradiction path is intentionally narrower |
165
+ | "No claims extracted means the description said nothing" | `--diff` targets have no PR description at all — zero claims there is a target-shape fact, not a description-quality one; check `budget.claimsSkipped` too before reading silence as clean |
166
+ | "The AC track failed, so the whole review failed" | `acCheck` degrades to facts-only (never throws) exactly like `keryx flow check-ac` does — a Jev failure on the criteria track is advisory, same as on the claims track |
167
+
168
+ ---
169
+
170
+ ## Verification
171
+
172
+ Before trusting a run's findings:
173
+
174
+ 1. Read `budget.claimsSkipped` in the `--json` output — a capped run checked
175
+ fewer than "every extracted claim" and the report should say so.
176
+ 2. Spot-check a `major` finding's `class_scope.sites` against the actual
177
+ diff: does the named region really touch an exported symbol the claim
178
+ said would not change?
179
+ 3. Confirm every finding carries `reviewer: "review-jev-contract"` and a
180
+ `dedupe_key` — both are required for the orchestrator's Quality Gate and
181
+ Wave C verification to route it correctly.
182
+
183
+ ---
184
+
185
+ ## Scope Boundaries
186
+
187
+ | Concern | This skill | Use instead |
188
+ |---------|------------|-------------|
189
+ | Whether the PR description's own claims match the diff | YES | — |
190
+ | Whether a flow's frozen acceptance criteria are likely met | YES (via `--flow`, reusing `check-ac.ts`) | `keryx flow check-ac` directly, for a standalone advisory run |
191
+ | An actual security/concurrency finding | NO | `review-security-code` / `review-highload` |
192
+ | Judging whether a reported finding is real | NO | `review-verifier` |
193
+ | A ranked risk map of the diff's hunks | NO | `review-jev-risk` |
@@ -47,34 +47,45 @@ review: it is a fact about the review, not about the diff.
47
47
  ## CLI-engine reviewers — dispatched as a command, not a sub-agent
48
48
 
49
49
  `review-jev-rules` (flow 330), `review-jev-risk` and `review-jev-scenarios`
50
- (both flow 332), and `review-jev-docs` and `review-jev-comments` (flow 333)
51
- are ADDITIONAL reviewers, never replacing any other, and their dispatch
52
- mechanism differs from every reviewer named in SKILL.md's Routing Table:
53
- there is no platform-native agent to invoke, because each is a deterministic
54
- **keryx program**. Run them with `keryx review jev-rules --scope
55
- <scope.json> --json`, `keryx review jev-risk --scope <scope.json> --json`,
56
- and `keryx review jev-scenarios --scope <scope.json> --json` — the SAME
57
- `scope.json` every other Wave A/B reviewer's dispatch already reads, so each
58
- checks exactly the same hunks. `review-jev-docs` and `review-jev-comments`
59
- take a diff/PR target instead: `keryx review jev-docs (--diff <ref>|--pr
60
- <n>) --json` and `keryx review jev-comments --pr <n> --repo <owner/repo>
61
- --json`. Read each `--json` output as a `REVIEW_RESULT` and merge its
50
+ (both flow 332), `review-jev-docs` and `review-jev-comments` (flow 333), and
51
+ `review-jev-contract` (flow 335) are ADDITIONAL reviewers, never replacing
52
+ any other, and their dispatch mechanism differs from every reviewer named in
53
+ SKILL.md's Routing Table: there is no platform-native agent to invoke,
54
+ because each is a deterministic **keryx program**. Run them with `keryx
55
+ review jev-rules --scope <scope.json> --json`, `keryx review jev-risk
56
+ --scope <scope.json> --json`, and `keryx review jev-scenarios --scope
57
+ <scope.json> --json` — the SAME `scope.json` every other Wave A/B reviewer's
58
+ dispatch already reads, so each checks exactly the same hunks.
59
+ `review-jev-docs` and `review-jev-comments` take a diff/PR target instead:
60
+ `keryx review jev-docs (--diff <ref>|--pr <n>) --json` and `keryx review
61
+ jev-comments --pr <n> --repo <owner/repo> --json`. `review-jev-contract`
62
+ also takes a diff/PR target, plus an optional linked flow: `keryx review
63
+ jev-contract (--diff <ref>|--pr <n>) [--flow <id>] --json` — a `--pr` target
64
+ checks the description's own claims; `--flow <id>` additionally checks that
65
+ flow's frozen acceptance criteria (reusing flow 328's `check-ac.ts`, see its
66
+ own SKILL.md). Read each `--json` output as a `REVIEW_RESULT` and merge its
62
67
  `findings` into the consolidated array exactly like a sub-agent reviewer's:
63
68
  same Sub-Agent Report Quality Gate, same dedup, same Wave C verification.
64
69
  `review-jev-risk` additionally emits `ranked` and `routingHints`;
65
- `review-jev-scenarios` additionally emits `checklist` — see each reviewer's
66
- own SKILL.md for what to do with its extra.
70
+ `review-jev-scenarios` additionally emits `checklist`; `review-jev-contract`
71
+ additionally emits `claims`, `budget`, and (when `--flow` was given)
72
+ `acCheck` — see each reviewer's own SKILL.md for what to do with its extra.
67
73
 
68
74
  Gate each BEFORE running its command, not after: skip it — recorded in
69
75
  `Skipped reviewers` with the reason, never silently absent — when its own
70
76
  opt-in (`review.jev.rules` / `review.jev.risk` / `review.jev.scenarios` /
71
- `review.jev.docs` / `review.jev.comments`) is not `true` in
72
- `.metaproject/tasks.config.json`, or when no Jev/OpenRouter credential is
73
- resolvable. Either gate failing means the command itself would refuse before
74
- any read or network call, so checking first saves a doomed dispatch. The
75
- five opt-ins are independent: any subset may be on. `review-jev-comments`
76
- additionally needs a comment ledger to already exist (`keryx review comments
77
- collect` run first) — it refuses, before any read, when there is none.
77
+ `review.jev.docs` / `review.jev.comments` / `review.jev.contract`) is not
78
+ `true` in `.metaproject/tasks.config.json`, or when no Jev/OpenRouter
79
+ credential is resolvable. Either gate failing means the command itself would
80
+ refuse before any read or network call, so checking first saves a doomed
81
+ dispatch. The six opt-ins are independent: any subset may be on.
82
+ `review-jev-comments` additionally needs a comment ledger to already exist
83
+ (`keryx review comments collect` run first) — it refuses, before any read,
84
+ when there is none. When `review.jev.contract` is on, dispatch
85
+ `review-jev-contract` with `--pr` (never `--diff`, which has no description
86
+ to check) and let its `findings` cover the description-vs-diff claim, in
87
+ place of the by-eye Stage 1 judgement described in SKILL.md's "This gate owns
88
+ the description-vs-diff comparison" — see that section's own note.
78
89
 
79
90
  `keryx review reviewers --json` marks a CLI-engine reviewer with
80
91
  `"engine": "jev"` on its `bundled` entry — the field's presence, not its
@@ -82,3 +93,52 @@ absence, is what distinguishes it from the default (an LLM sub-agent
82
93
  dispatch). A future engine-backed reviewer follows the same pattern: gate on
83
94
  its own opt-in and reachability, dispatch as a command, merge its `--json`
84
95
  output the same way.
96
+
97
+ ## `jev-triage` — advisory annotations, not a reviewer (Step 9b)
98
+
99
+ `review-jev-triage` (flow 340) is a DIFFERENT shape from every CLI-engine
100
+ reviewer above: it never produces a `findings` array of its own, and it is
101
+ never merged into the consolidated array. It runs AFTER the Sub-Agent Report
102
+ Quality Gate has already validated/merged/deduplicated the round's findings
103
+ (Step 9) and BEFORE Wave C verification (Step 10) — Step 9b in the Workflow
104
+ block — annotating the findings that already exist. It is advisory and
105
+ annotate-only by construction: nothing it returns drops or demotes a finding,
106
+ and no downstream step is permitted to treat its output as anything but a
107
+ hint.
108
+
109
+ Run it once per round, over the round's own consolidated findings:
110
+
111
+ ```bash
112
+ keryx review jev-triage --report <this round's review package dir> --json
113
+ ```
114
+
115
+ Gate it the same way as every other CLI-engine reviewer: skip it — recorded
116
+ in `Skipped reviewers` with the reason — when `review.jev.triage` is not
117
+ `true` in `.metaproject/tasks.config.json`, or when no Jev/OpenRouter
118
+ credential is resolvable. Either gate failing means the command itself would
119
+ refuse before any read or network call.
120
+
121
+ Its `--json` output is `{status, reviewer: "review-jev-triage", summary,
122
+ annotations, budget, tokens}` — never a `REVIEW_RESULT` merged into the
123
+ findings array. `annotations` carries three tracks, each read and used
124
+ differently in the report:
125
+
126
+ | Track | Shape | What to do with it |
127
+ |---|---|---|
128
+ | `severity_check` | `{id, p, flagged}` per blocker/major finding — `flagged` means `p < 0.4` | Show `flagged` findings as a soft note next to their existing severity ("Jev's own trigger+outcome check scored this low"). Never change the finding's `severity` field from this alone. |
129
+ | `merge_candidates` | `{id, a, b, reason, p}` per candidate pair | A high `p` is a suggestion to a human that `a` and `b` may be the same defect — surface it in the report as a note on both findings. Never merge, drop, or renumber either finding from this alone. |
130
+ | `verify_order` | `{id, p}`, sorted lowest-`p`-first | Feed this order into Wave C's own dispatch — verify the lowest-plausibility findings first — never as a reason to skip verifying any of them. |
131
+
132
+ `status` is `DONE_WITH_CONCERNS` — never `DONE` — when any `severity_check` is
133
+ `flagged` OR any `merge_candidates` pair scores `p >= 0.5`
134
+ (`LIKELY_DUPLICATE_THRESHOLD`, `src/commands/review-jev-triage.ts`). Live-check
135
+ calibration found same-file pairs that were NOT duplicates scoring around
136
+ 0.6, so a merge candidate above the threshold is a prompt to look, never a
137
+ merge — the status flip is a nudge to read the pair, not a verdict that the
138
+ findings are duplicates.
139
+
140
+ Show its annotations in the report as their own subsection, next to (not
141
+ inside) the findings they annotate — same separation the finding schema
142
+ itself draws between a reviewer's claim (`severity`, `problem`, …) and what
143
+ became of it (`disposition`), which a reviewer never states and this pass
144
+ does not either.
@@ -42,6 +42,7 @@ Review Orchestrator Progress:
42
42
  - [ ] Step 7: Stage 1 gate - spec compliance check (if issue/task provided)
43
43
  - [ ] Step 8: Dispatch selected reviewers in PARALLEL with reviewer-input schema
44
44
  - [ ] Step 9: Collect reviewer-finding schema results and handle NEEDS_CONTEXT
45
+ - [ ] Step 9b: `jev-triage` — advisory, annotate-only severity/duplicate/verify-order annotations over the consolidated findings (opt-in `review.jev.triage`)
45
46
  - [ ] Step 10: Wave C — dispatch `review-verifier` over the consolidated findings
46
47
  - [ ] Step 11: Sort by severity, deduplicate, emit unified report
47
48
  - [ ] Step 12: Emit the machine-readable `keryx:findings` block alongside the report
@@ -1168,7 +1169,7 @@ that already holds intent and diff side by side is this one.
1168
1169
 
1169
1170
  Run it on every round, not only the first. The drift the finding catches is
1170
1171
  created BY the rounds: the code moves to answer findings, the body does not, and
1171
- whoever reads the merge commit a year later reads the body.
1172
+ whoever reads the merge commit a year later reads the body. When `review.jev.contract` is on, dispatch `review-jev-contract --pr` and read its `findings` as this comparison's scored result instead of judging it by eye; the by-eye judgement is the fallback when that opt-in is off — `SKILL.detail.md` § "CLI-engine reviewers".
1172
1173
 
1173
1174
  ---
1174
1175
 
@@ -1177,7 +1178,7 @@ whoever reads the merge commit a year later reads the body.
1177
1178
  Dispatch selected reviewers in parallel when independent. Use waves when token budget is tight or when one reviewer needs another result:
1178
1179
 
1179
1180
  1. Wave A - core correctness/risk reviewers: logic, architecture, security/highload when selected.
1180
- 2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files. `review-jev-rules` (flow 330), `review-jev-risk`/`review-jev-scenarios` (flow 332), `review-jev-docs`/`review-jev-comments` (flow 333) also run here, CLI-engine not sub-agent, `"engine": "jev"` in `keryx review reviewers --json`, gated on their own opt-in and a resolvable Jev/OpenRouter credential — `SKILL.detail.md` § "CLI-engine reviewers".
1181
+ 2. Wave B - domain reviewers: frontend/backend/testing/convention reviewers filtered to relevant files. `review-jev-rules` (flow 330), `review-jev-risk`/`review-jev-scenarios` (flow 332), `review-jev-docs`/`review-jev-comments` (flow 333), `review-jev-contract` (flow 335) also run here, CLI-engine not sub-agent, `"engine": "jev"` in `keryx review reviewers --json`, gated on their own opt-in and a resolvable Jev/OpenRouter credential — `SKILL.detail.md` § "CLI-engine reviewers".
1181
1182
  3. Wave C - **verification**: `review-verifier` over the consolidated findings, when blockers/majors
1182
1183
  exist, `--verify` is set, or the PR is high-risk. See below.
1183
1184
 
@@ -1691,8 +1692,7 @@ CONTEXT_PATH: .metaproject/jobs/<job-name>/ai/context.md
1691
1692
  If provided and the file exists, read the context document **before** running scope detection.
1692
1693
  Use it to understand:
1693
1694
  - Intentionally chosen libraries and patterns (do not flag as issues)
1694
- - Architectural decisions already agreed upon
1695
- - Acceptance criteria to drive the Stage 1 spec compliance gate
1695
+ - Architectural decisions already agreed upon, and acceptance criteria driving the Stage 1 spec compliance gate
1696
1696
 
1697
1697
  If absent, proceed normally — context is optional and non-blocking.
1698
1698
 
@@ -0,0 +1,4 @@
1
+ {
2
+ "agents": [],
3
+ "note": "honest gate (DeepSeek deepseek-chat runner+judge, strictness high, trials 10, flow 337) ran and failed for all four skills -- trigger accuracy, not behavior content (every behavior scenario scored 0.9 or 1.0). No generated pair ships; pack stays stability: experimental. See governance/eval.json for the recorded reports and W1-stack-catalog.md's Wave 4 batch 5 implementation notes for the diagnosis."
4
+ }