@mrciphersmith/keryx 0.2.164 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (182) hide show
  1. package/README.md +4 -1
  2. package/dist/cli.js +82540 -50300
  3. package/dist/core.js +28967 -18937
  4. package/package.json +2 -2
  5. package/src/gdgraph/affected-report.ts +141 -0
  6. package/src/gdgraph/build.ts +170 -23
  7. package/src/gdgraph/service.ts +6 -0
  8. package/src/gdgraph/staleness.ts +253 -45
  9. package/src/gdskills/bundled/agents/codebase-navigator.md +55 -0
  10. package/src/gdskills/bundled/agents/design-advisor.md +64 -0
  11. package/src/gdskills/bundled/agents/docs-maintainer.md +56 -0
  12. package/src/gdskills/bundled/agents/end-to-end-tester.md +56 -0
  13. package/src/gdskills/bundled/agents/error-path-auditor.md +57 -0
  14. package/src/gdskills/bundled/agents/go-build-fixer.md +52 -0
  15. package/src/gdskills/bundled/agents/go-code-auditor.md +49 -0
  16. package/src/gdskills/bundled/agents/performance-auditor.md +63 -0
  17. package/src/gdskills/bundled/agents/python-build-fixer.md +52 -0
  18. package/src/gdskills/bundled/agents/python-code-auditor.md +49 -0
  19. package/src/gdskills/bundled/agents/refactoring-steward.md +61 -0
  20. package/src/gdskills/bundled/agents/security-auditor.md +62 -0
  21. package/src/gdskills/bundled/agents/test-first-driver.md +61 -0
  22. package/src/gdskills/bundled/agents/work-planner.md +62 -0
  23. package/src/gdskills/bundled/install-manifest.json +797 -0
  24. package/src/gdskills/bundled/rules/core/model-selection.mdc +51 -0
  25. package/src/gdskills/bundled/rules/core/skill-lifecycle.mdc +29 -1
  26. package/src/gdskills/bundled/rules/core/skills-storage-workflow.mdc +2 -2
  27. package/src/gdskills/bundled/skills/review/code-style-review/SKILL.md +1 -1
  28. package/src/gdskills/bundled/skills/review/review-jev-rules/SKILL.md +267 -0
  29. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +26 -0
  30. package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +75 -247
  31. package/src/gdskills/bundled/skills/review/review-orchestrator/output-contract.schema.json +19 -0
  32. package/src/gdskills/bundled/skills/review/review-orchestrator/reviewer-finding.schema.json +10 -0
  33. package/src/gdskills/bundled/skills/review/review-orchestrator/reviewer-input.schema.json +5 -0
  34. package/src/gdskills/bundled/skills/review/review-orchestrator/templates/pr-comment-backend.md +50 -0
  35. package/src/gdskills/bundled/skills/review/review-orchestrator/templates/pr-comment-frontend.md +52 -0
  36. package/src/gdskills/bundled/skills/review/review-orchestrator/templates/review-report.md +143 -0
  37. package/src/gdskills/bundled/stacks/angular/agent-refs.json +4 -0
  38. package/src/gdskills/bundled/stacks/angular/governance/eval.json +1751 -0
  39. package/src/gdskills/bundled/stacks/angular/governance/scout.json +32 -0
  40. package/src/gdskills/bundled/stacks/angular/pack.json +55 -0
  41. package/src/gdskills/bundled/stacks/angular/rules/coding-style.mdc +82 -0
  42. package/src/gdskills/bundled/stacks/angular/rules/patterns.mdc +84 -0
  43. package/src/gdskills/bundled/stacks/angular/rules/security.mdc +70 -0
  44. package/src/gdskills/bundled/stacks/angular/rules/testing.mdc +73 -0
  45. package/src/gdskills/bundled/stacks/angular/skills/angular-build-fix/SKILL.md +127 -0
  46. package/src/gdskills/bundled/stacks/angular/skills/angular-build-fix/evals.json +72 -0
  47. package/src/gdskills/bundled/stacks/angular/skills/angular-code-review/SKILL.md +98 -0
  48. package/src/gdskills/bundled/stacks/angular/skills/angular-code-review/evals.json +73 -0
  49. package/src/gdskills/bundled/stacks/angular/skills/angular-implementation/SKILL.md +112 -0
  50. package/src/gdskills/bundled/stacks/angular/skills/angular-implementation/evals.json +74 -0
  51. package/src/gdskills/bundled/stacks/angular/skills/angular-testing/SKILL.md +102 -0
  52. package/src/gdskills/bundled/stacks/angular/skills/angular-testing/evals.json +71 -0
  53. package/src/gdskills/bundled/stacks/go/agent-refs.json +3 -0
  54. package/src/gdskills/bundled/stacks/go/governance/eval.json +1745 -0
  55. package/src/gdskills/bundled/stacks/go/governance/scout.json +31 -0
  56. package/src/gdskills/bundled/stacks/go/pack.json +41 -0
  57. package/src/gdskills/bundled/stacks/go/rules/coding-style.mdc +85 -0
  58. package/src/gdskills/bundled/stacks/go/rules/patterns.mdc +65 -0
  59. package/src/gdskills/bundled/stacks/go/rules/security.mdc +73 -0
  60. package/src/gdskills/bundled/stacks/go/rules/testing.mdc +68 -0
  61. package/src/gdskills/bundled/stacks/go/skills/go-build-fix/SKILL.md +138 -0
  62. package/src/gdskills/bundled/stacks/go/skills/go-build-fix/evals.json +75 -0
  63. package/src/gdskills/bundled/stacks/go/skills/go-code-review/SKILL.md +121 -0
  64. package/src/gdskills/bundled/stacks/go/skills/go-code-review/evals.json +72 -0
  65. package/src/gdskills/bundled/stacks/go/skills/go-implementation/SKILL.md +122 -0
  66. package/src/gdskills/bundled/stacks/go/skills/go-implementation/evals.json +76 -0
  67. package/src/gdskills/bundled/stacks/go/skills/go-testing/SKILL.md +126 -0
  68. package/src/gdskills/bundled/stacks/go/skills/go-testing/evals.json +73 -0
  69. package/src/gdskills/bundled/stacks/mobx/agent-refs.json +4 -0
  70. package/src/gdskills/bundled/stacks/mobx/governance/eval.json +904 -0
  71. package/src/gdskills/bundled/stacks/mobx/governance/scout.json +18 -0
  72. package/src/gdskills/bundled/stacks/mobx/pack.json +28 -0
  73. package/src/gdskills/bundled/stacks/mobx/rules/coding-style.mdc +91 -0
  74. package/src/gdskills/bundled/stacks/mobx/rules/patterns.mdc +122 -0
  75. package/src/gdskills/bundled/stacks/mobx/rules/security.mdc +56 -0
  76. package/src/gdskills/bundled/stacks/mobx/rules/testing.mdc +63 -0
  77. package/src/gdskills/bundled/stacks/mobx/skills/mobx-observable-testing/SKILL.md +124 -0
  78. package/src/gdskills/bundled/stacks/mobx/skills/mobx-observable-testing/evals.json +73 -0
  79. package/src/gdskills/bundled/stacks/mobx/skills/mobx-store-implementation/SKILL.md +149 -0
  80. package/src/gdskills/bundled/stacks/mobx/skills/mobx-store-implementation/evals.json +74 -0
  81. package/src/gdskills/bundled/stacks/nestjs/agent-refs.json +4 -0
  82. package/src/gdskills/bundled/stacks/nestjs/governance/eval.json +1308 -0
  83. package/src/gdskills/bundled/stacks/nestjs/governance/scout.json +34 -0
  84. package/src/gdskills/bundled/stacks/nestjs/pack.json +53 -0
  85. package/src/gdskills/bundled/stacks/nestjs/rules/coding-style.mdc +70 -0
  86. package/src/gdskills/bundled/stacks/nestjs/rules/patterns.mdc +83 -0
  87. package/src/gdskills/bundled/stacks/nestjs/rules/security.mdc +73 -0
  88. package/src/gdskills/bundled/stacks/nestjs/rules/testing.mdc +69 -0
  89. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-build-fix/SKILL.md +157 -0
  90. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-build-fix/evals.json +70 -0
  91. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-implementation/SKILL.md +129 -0
  92. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-implementation/evals.json +71 -0
  93. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-testing/SKILL.md +143 -0
  94. package/src/gdskills/bundled/stacks/nestjs/skills/nestjs-testing/evals.json +69 -0
  95. package/src/gdskills/bundled/stacks/nextjs-nuxt/agent-refs.json +4 -0
  96. package/src/gdskills/bundled/stacks/nextjs-nuxt/governance/eval.json +2413 -0
  97. package/src/gdskills/bundled/stacks/nextjs-nuxt/governance/scout.json +42 -0
  98. package/src/gdskills/bundled/stacks/nextjs-nuxt/pack.json +42 -0
  99. package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/coding-style.mdc +69 -0
  100. package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/patterns.mdc +88 -0
  101. package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/security.mdc +72 -0
  102. package/src/gdskills/bundled/stacks/nextjs-nuxt/rules/testing.mdc +64 -0
  103. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-build-fix/SKILL.md +147 -0
  104. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-build-fix/evals.json +75 -0
  105. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-code-review/SKILL.md +118 -0
  106. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-code-review/evals.json +76 -0
  107. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-implementation/SKILL.md +135 -0
  108. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-implementation/evals.json +78 -0
  109. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-testing/SKILL.md +116 -0
  110. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-testing/evals.json +75 -0
  111. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-upgrade-migration/SKILL.md +134 -0
  112. package/src/gdskills/bundled/stacks/nextjs-nuxt/skills/nextjs-nuxt-upgrade-migration/evals.json +76 -0
  113. package/src/gdskills/bundled/stacks/python/agent-refs.json +3 -0
  114. package/src/gdskills/bundled/stacks/python/governance/eval.json +1758 -0
  115. package/src/gdskills/bundled/stacks/python/governance/scout.json +34 -0
  116. package/src/gdskills/bundled/stacks/python/pack.json +41 -0
  117. package/src/gdskills/bundled/stacks/python/rules/coding-style.mdc +63 -0
  118. package/src/gdskills/bundled/stacks/python/rules/patterns.mdc +88 -0
  119. package/src/gdskills/bundled/stacks/python/rules/security.mdc +84 -0
  120. package/src/gdskills/bundled/stacks/python/rules/testing.mdc +77 -0
  121. package/src/gdskills/bundled/stacks/python/skills/python-build-fix/SKILL.md +144 -0
  122. package/src/gdskills/bundled/stacks/python/skills/python-build-fix/evals.json +74 -0
  123. package/src/gdskills/bundled/stacks/python/skills/python-code-review/SKILL.md +155 -0
  124. package/src/gdskills/bundled/stacks/python/skills/python-code-review/evals.json +72 -0
  125. package/src/gdskills/bundled/stacks/python/skills/python-implementation/SKILL.md +143 -0
  126. package/src/gdskills/bundled/stacks/python/skills/python-implementation/evals.json +78 -0
  127. package/src/gdskills/bundled/stacks/python/skills/python-testing/SKILL.md +132 -0
  128. package/src/gdskills/bundled/stacks/python/skills/python-testing/evals.json +73 -0
  129. package/src/gdskills/bundled/stacks/react/agent-refs.json +4 -0
  130. package/src/gdskills/bundled/stacks/react/governance/eval.json +2188 -0
  131. package/src/gdskills/bundled/stacks/react/governance/scout.json +40 -0
  132. package/src/gdskills/bundled/stacks/react/pack.json +42 -0
  133. package/src/gdskills/bundled/stacks/react/rules/coding-style.mdc +58 -0
  134. package/src/gdskills/bundled/stacks/react/rules/patterns.mdc +79 -0
  135. package/src/gdskills/bundled/stacks/react/rules/security.mdc +70 -0
  136. package/src/gdskills/bundled/stacks/react/rules/testing.mdc +60 -0
  137. package/src/gdskills/bundled/stacks/react/skills/react-build-fix/SKILL.md +139 -0
  138. package/src/gdskills/bundled/stacks/react/skills/react-build-fix/evals.json +72 -0
  139. package/src/gdskills/bundled/stacks/react/skills/react-code-review/SKILL.md +148 -0
  140. package/src/gdskills/bundled/stacks/react/skills/react-code-review/evals.json +74 -0
  141. package/src/gdskills/bundled/stacks/react/skills/react-implementation/SKILL.md +140 -0
  142. package/src/gdskills/bundled/stacks/react/skills/react-implementation/evals.json +74 -0
  143. package/src/gdskills/bundled/stacks/react/skills/react-testing/SKILL.md +142 -0
  144. package/src/gdskills/bundled/stacks/react/skills/react-testing/evals.json +83 -0
  145. package/src/gdskills/bundled/stacks/react/skills/react-upgrade-migration/SKILL.md +155 -0
  146. package/src/gdskills/bundled/stacks/react/skills/react-upgrade-migration/evals.json +74 -0
  147. package/src/gdskills/bundled/stacks/ts-js-node/agent-refs.json +4 -0
  148. package/src/gdskills/bundled/stacks/ts-js-node/governance/eval.json +2155 -0
  149. package/src/gdskills/bundled/stacks/ts-js-node/governance/scout.json +40 -0
  150. package/src/gdskills/bundled/stacks/ts-js-node/pack.json +41 -0
  151. package/src/gdskills/bundled/stacks/ts-js-node/rules/coding-style.mdc +73 -0
  152. package/src/gdskills/bundled/stacks/ts-js-node/rules/patterns.mdc +61 -0
  153. package/src/gdskills/bundled/stacks/ts-js-node/rules/security.mdc +71 -0
  154. package/src/gdskills/bundled/stacks/ts-js-node/rules/testing.mdc +63 -0
  155. package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-build-fix/SKILL.md +137 -0
  156. package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-build-fix/evals.json +73 -0
  157. package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-code-review/SKILL.md +124 -0
  158. package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-code-review/evals.json +74 -0
  159. package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-esm-migration/SKILL.md +152 -0
  160. package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-esm-migration/evals.json +71 -0
  161. package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-implementation/SKILL.md +127 -0
  162. package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-implementation/evals.json +72 -0
  163. package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-testing/SKILL.md +134 -0
  164. package/src/gdskills/bundled/stacks/ts-js-node/skills/nodejs-testing/evals.json +70 -0
  165. package/src/gdskills/bundled/stacks/vue/agent-refs.json +4 -0
  166. package/src/gdskills/bundled/stacks/vue/governance/eval.json +2215 -0
  167. package/src/gdskills/bundled/stacks/vue/governance/scout.json +42 -0
  168. package/src/gdskills/bundled/stacks/vue/pack.json +42 -0
  169. package/src/gdskills/bundled/stacks/vue/rules/coding-style.mdc +73 -0
  170. package/src/gdskills/bundled/stacks/vue/rules/patterns.mdc +84 -0
  171. package/src/gdskills/bundled/stacks/vue/rules/security.mdc +60 -0
  172. package/src/gdskills/bundled/stacks/vue/rules/testing.mdc +69 -0
  173. package/src/gdskills/bundled/stacks/vue/skills/vue-build-fix/SKILL.md +137 -0
  174. package/src/gdskills/bundled/stacks/vue/skills/vue-build-fix/evals.json +72 -0
  175. package/src/gdskills/bundled/stacks/vue/skills/vue-code-review/SKILL.md +120 -0
  176. package/src/gdskills/bundled/stacks/vue/skills/vue-code-review/evals.json +71 -0
  177. package/src/gdskills/bundled/stacks/vue/skills/vue-implementation/SKILL.md +122 -0
  178. package/src/gdskills/bundled/stacks/vue/skills/vue-implementation/evals.json +72 -0
  179. package/src/gdskills/bundled/stacks/vue/skills/vue-testing/SKILL.md +115 -0
  180. package/src/gdskills/bundled/stacks/vue/skills/vue-testing/evals.json +72 -0
  181. package/src/gdskills/bundled/stacks/vue/skills/vue2-to-vue3-migration/SKILL.md +135 -0
  182. package/src/gdskills/bundled/stacks/vue/skills/vue2-to-vue3-migration/evals.json +71 -0
@@ -83,6 +83,57 @@ residue, and it is kept in the smallest shape that works:
83
83
  A model whose name is a codename carrying no size is unrankable by design. That
84
84
  is the honest outcome, not a gap to be patched with folklore.
85
85
 
86
+ **Version is deliberately not ranked here.** `rankModelId` never parses a
87
+ version number out of an id — two ids in the same size-word family but a
88
+ different generation (an older vs. a newer numbered release of the same
89
+ line) rank IDENTICALLY through this resolution, and a
90
+ `standard`/`light`/`deep` dispatch never distinguishes between them. That
91
+ stays true for `spawn_subagent` and `keryx review tier` after the decision
92
+ below; it is a DIFFERENT consumer's rule, not this one's.
93
+
94
+ The routing "derived default table" (`deriveDefaultTable`,
95
+ `src/harness/routing/derive-default-table.ts`, flow 327/PRD §6.3) is that
96
+ different consumer: it picks one model per routing category (`review`,
97
+ `subagents`, `quick`, …) from a whole connected catalogue, not a tier
98
+ relative to the session. The operator decided, 2026-09-25 (review of PR #718:
99
+ a session's newer release lost a tie to an older release of the SAME family
100
+ it should have beaten), that WITHIN the same family and vendor a newer
101
+ version ranks higher there — the family word (a size marker from this same
102
+ `MODEL_RANK_HINTS` list) already carries the size class, so version is free
103
+ to order generations within it. That table reuses
104
+ `rankModelId`/`MODEL_RANK_HINTS` UNMODIFIED for the size axis and adds its
105
+ own separate, conservative version parse on top of it — nothing above
106
+ changes for this module's own tier resolution.
107
+
108
+ That version parse (`parseModelVersion`/`familyKey`, round 2) reads BOTH a
109
+ dotted id (`vendor-opus-5.5`, `gemini-3.8-flash`) and a real
110
+ Anthropic-style HYPHENATED one, where a run of 2-3 adjacent short (1-2
111
+ digit) numeric tokens reads as a dotted version (`vendor-opus-4-8` → `4.8`).
112
+ It still refuses a date/snapshot-shaped token (6-8 bare digits) and more
113
+ than one numeric group in the same id — no version rather than a guess. A
114
+ `gpt-4o`-style id (a short number plus one trailing letter) parses the
115
+ number alone as the version, the letter kept only as an ignored variant
116
+ tag, so `gpt-4o` and `gpt-4.1` compare within the same family instead of
117
+ never comparing at all. Two versions of the same family are never decided
118
+ by alphabetical order — the id-string fallback only applies when neither
119
+ side has a parseable version.
120
+
121
+ **A parameter-size token is never a version (round 3, review of PR #718).**
122
+ `7b`, `32b`, `70b`, `1.5b`, `8x7b` — a total or MoE "N experts x M billion"
123
+ param count — always stay in `familyKey` and are never fed to
124
+ `parseModelVersion`. Before this fix, the letter-variant shape meant for
125
+ `gpt-4o` (`/^(\d{1,2})([a-z])$/`, any trailing lowercase letter) also
126
+ matched a size token, so `qwen2.5-coder-7b` and `qwen2.5-coder-32b` were
127
+ misread as the SAME family at "version" 7 vs. 32 — merging two genuinely
128
+ different-sized models and ranking the smaller one "newer". The letter
129
+ variant is now restricted to the one real vendor shape that needs it
130
+ (`4o`'s trailing `o`); every other trailing letter after a short digit run
131
+ is a size suffix, so `qwen2.5-coder-7b`/`qwen2.5-coder-32b` and
132
+ `llama-3.3-70b`/`llama-3.3-8b` key to different families and are never
133
+ version-compared against each other. A trailing `-latest`/`-preview` alias
134
+ word is also stripped from `familyKey` only (never from the id itself), so
135
+ `vendor-3-7-sonnet-latest` joins the same family as `vendor-sonnet-5`.
136
+
86
137
  ## Falling back
87
138
 
88
139
  Ranking is **refused**, and every tier keeps the session's provider and model,
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  description: "When and how agents verify and re-learn project-skills during implementation and review. Verify/learn live in the agent loop, not in git hooks."
3
3
  alwaysApply: false
4
+ version: 0.2.0
4
5
  ---
5
6
 
6
7
  # Skill Lifecycle — verify & learn in the loop
@@ -52,8 +53,35 @@ a different `review tier` invocation (e.g. `--scope narrow`, which resolves
52
53
  heavier than `light`).** The flagship's only job here is to review the
53
54
  proposal before apply.
54
55
 
56
+ ### Passive observation is not mutation
57
+
58
+ A hook registered under the shell hook runtime may append an
59
+ **observation event** to `.metaproject/data/learning/observations/` — the
60
+ shipped `keryx.learning-observer` hook does exactly this, on
61
+ `tool-start`/`tool-complete`/`tool-failed`/`user-prompt`/`session-start`/
62
+ `turn-stop`/`session-end`. This is explicitly permitted from a hook, unlike
63
+ every other write this rule governs, because it changes no skill, rule,
64
+ agent, or memory entry — it only accumulates evidence (redacted, digest-only)
65
+ that a human will later decide about.
66
+
67
+ `keryx learn extract`, which turns accumulated observations into a
68
+ `status: candidate` `learned-pattern` record under
69
+ `.metaproject/data/learning/candidates/`, is a separate, **command- or
70
+ schedule-triggered** step — never run inside a hook. It still produces no
71
+ `accepted` state and mutates no skill/rule/agent/memory file by itself.
72
+
73
+ Everything from `keryx learn accept` onward — Accept, project-skill Apply,
74
+ reviewer-profile Apply, Promote, Graduate apply — remains agent-loop,
75
+ human-triggered work under the rule above, with the same "dispatch as a
76
+ light-tier subagent, review before apply" discipline already specified for
77
+ `skills learn apply`. Each of accept/promote/graduate-apply additionally
78
+ refuses outright outside a real interactive terminal (no bypass flag).
79
+
55
80
  ## Never
56
81
 
57
- - Do not put `learn` (or any mutation) in a hook.
82
+ - Do not put `learn` (or any mutation) in a hook — **passive observation is
83
+ the single exception** (see above): a hook may append an observation event,
84
+ nothing more. `keryx learn extract` also never runs inside a hook, even
85
+ though it produces no `accepted` state.
58
86
  - Do not `apply` a proposal without reading it.
59
87
  - Do not trust a `stale` / `needs-review` skill as ground truth.
@@ -66,8 +66,8 @@ gets reconciled.
66
66
 
67
67
  `SKILL.md` frontmatter follows the published Agent Skills specification
68
68
  (agentskills.io — the format Zed adopted when it dropped its own Rules
69
- Library), with two deliberate additions. Flow 203 removed two accidental
70
- divergences that had crept in; keep both fixes intact when authoring or
69
+ Library), with two deliberate additions. Two earlier accidental divergences
70
+ have since been removed; keep both fixes intact when authoring or
71
71
  editing a skill:
72
72
 
73
73
  - **`version` lives in `metadata.version` only.** The spec has no top-level
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: code-style-review
3
- description: "Use when the legacy code style and architecture review profile (code-style-patterns.mdc) is asked for by name — naming, organization, patterns, and TypeScript usage on the current branch. NOT for: a general style review (review-style) or an architecture review (review-architecture)."
3
+ description: "Use when a caller wants the style/architecture checklist in code-style-patterns.mdc applied by its own name — checks naming, module and file layout, and TypeScript usage patterns against that specific ruleset, over this branch's diff. Narrower than a general style pass, and separate from the correctness-and-security AI review baseline. NOT for: a general style review (review-style) or an architecture review (review-architecture)."
4
4
  triggers:
5
5
  - "code-style-review"
6
6
  - "architecture style"
@@ -0,0 +1,267 @@
1
+ ---
2
+ name: review-jev-rules
3
+ model_tier: light
4
+ description: |
5
+ Use when: an ADDITIONAL, machine-scored pass is wanted over every changed hunk against
6
+ every applicable project rule clause — never replacing any other reviewer. Dispatched
7
+ by review-orchestrator in Wave B, via the CLI (`keryx review jev-rules`), when
8
+ `review.jev.rules: true` in .metaproject/tasks.config.json and a Jev/OpenRouter
9
+ credential is resolvable. It is not an LLM sub-agent: there is nothing to dispatch
10
+ through a platform-native agent mechanism, and no prose is generated by a model —
11
+ Jev only answers a `noul` violation probability per (hunk, rule clause) pair; keryx
12
+ writes every word of every finding.
13
+ NOT for: judgement calls a rule clause does not state (logic bugs, architecture,
14
+ security, performance — those stay with the reviewers that already cover them), and
15
+ not a substitute for reading the rules yourself. A finding here says a hunk likely
16
+ contradicts a clause of a rule this project already wrote down; it says nothing about
17
+ code that violates no documented rule.
18
+ triggers:
19
+ - "jev rules"
20
+ - "rule check"
21
+ - "review --jev-rules"
22
+ metadata:
23
+ author: "MrCipherSmith"
24
+ version: "1.0.0"
25
+ category: "review"
26
+ compatible_harnesses: "cursor,codex,zed,opencode,claude"
27
+ engine: "jev"
28
+ origin: "authored"
29
+ license: "MIT"
30
+ ---
31
+
32
+ # Review — Jev Rules (deterministic project-rule conformance)
33
+
34
+ An ADDITIONAL orchestrator reviewer, flow 330. It is a **keryx program**, not an
35
+ LLM sub-agent: `review-orchestrator` never dispatches a platform-native agent
36
+ for it, it runs `keryx review jev-rules` and reads the `--json` output. Every
37
+ finding is composed by keryx; Jev ("System One" on OpenRouter) supplies only a
38
+ `noul` probability per `(hunk, rule clause)` pair — it never writes prose, and
39
+ nothing it returns is quoted verbatim into a finding except the clause text
40
+ keryx already had.
41
+
42
+ ---
43
+
44
+ ## What it does
45
+
46
+ 1. Discovers rule sources deterministically (`src/review/jev-rules.ts`'s
47
+ module header names the exact scope): every file under
48
+ `.metaproject/rules/**` and `rules/**`, every project-skill or installed
49
+ gdskill whose name or `metadata.category` marks it a coding convention, and
50
+ anything named by `--rules <paths>`. **`--rules` bypass is scoped to ONE
51
+ FILE**: an entry naming a single document bypasses the category filter
52
+ below outright; an entry naming a DIRECTORY is walked and its files are
53
+ filtered exactly like auto-discovery — asking for a whole directory is not
54
+ the same as naming one document, and the whole point of the filter is lost
55
+ if it is.
56
+ 2. **Category filter, before clause extraction.** Each discovered source
57
+ (auto-discovered, or a file found by walking a `--rules` DIRECTORY — never
58
+ a `--rules` entry naming that ONE file explicitly) is classified
59
+ `code`/`process`/`docs` — explicit frontmatter first (`applies_to:
60
+ code|process|docs`, or `metadata.category`), else a documented
61
+ filename/title heuristic (`PROCESS_RULE_HEURISTIC_TERMS`: `commit`, `git`,
62
+ `tdd`, `workflow`, `definition-of-done`, `documentation`, `requirements`,
63
+ `plan`, `prompting`, `subagent`, `skill`, `jobs`, `orchestrat`,
64
+ `review-process`, `release`). A `process`/`docs` source never reaches
65
+ clause extraction at all — excluded, and reported with its reason under
66
+ `excludedSources`.
67
+ 3. Splits each remaining rule document into clauses with
68
+ `extractReferenceClauses` (the same deterministic splitter `review
69
+ conform` uses), drops authoring-template scaffolding
70
+ (`isPlaceholderClauseText`: an unfilled `[x] <criterion> — verified by
71
+ <test>` checklist line, text dominated by `<...>` placeholders, or a bare
72
+ code-fence line — never a real clause, dropped BEFORE tagging so it costs
73
+ nothing), then TAGS every remaining clause `state_kind:
74
+ "pr"|"report"|"hunk"` + `checkable` by REUSING `review conform`'s own
75
+ `applyClauseTags`/`buildClauseTagQuestions`/`clauseTagFromChoice` — an
76
+ explicit `[state:hunk]`/`[not-checkable: ...]` marker on the clause text
77
+ when the rule author wrote one, else one Jev `choice` call per doc's
78
+ untagged clauses (never per clause), cached by the doc's content hash at
79
+ `.metaproject/data/review-jev-rules/clause-tags.json` (a jev-rules-
80
+ specific cache file — `review conform` and `review-jev-rules` never race
81
+ on the same one). Only a clause tagged `state_kind: "hunk"` and
82
+ `checkable` is ever paired against a hunk; anything else is dropped and
83
+ reported under `droppedClauses`.
84
+ 4. Decides, per `(rule, changed file)` pair, whether an applicable clause
85
+ applies — by declared path globs (`metadata.paths`) and/or
86
+ `metadata.stack_requires`, failing toward inclusion exactly like `keryx
87
+ review stack`/`scope.ts` already do, THEN by file kind: a docs hunk
88
+ (`.md`/`.mdx`/`.txt`) pairs only with a `docs`-categorised source or a
89
+ clause explicitly tagged `[docs-applicable]`; a code source never pairs
90
+ with a docs hunk, and vice versa (`clauseFileKindApplicability`).
91
+ 5. Takes every changed hunk from `keryx review scope` (mechanical bulk
92
+ already dropped) and asks Jev, per applicable `(hunk, hunk-checkable
93
+ clause)` pair, **one question**: "does this hunk VIOLATE this clause?" —
94
+ batched under the vendor's 64k token budget, capped at `--max-calls`
95
+ (default 150). The budget is spread FAIRLY: hunks are ranked code, then
96
+ tests, then docs, and pairs are allocated round-robin across hunks rather
97
+ than draining the cap on the first hunk in diff order — every selection,
98
+ every drop, and per-hunk coverage (`selection.hunkCoverage`, how many
99
+ hunks were reached vs. never reached) is reported, never silent.
100
+ 6. Synthesizes findings deterministically: `problem` quotes the clause,
101
+ `impact` is the rule's own stated rationale (if it has a `## Rationale`/
102
+ `## Why` section) or a fixed template, `suggested_fix` names the clause to
103
+ bring the hunk in line with, `evidence` is the hunk location(s) + Jev's
104
+ probability. Severity is capped at `minor` unless the rule itself declares
105
+ a higher one (`[severity: major]` on the clause). One finding per
106
+ `(clause, file)`, deduped across every hunk of that file with a hunk list.
107
+
108
+ **Why steps 2-5's filters exist:** a live check of this repository's own
109
+ 41-doc `.metaproject/rules/**` corpus against a merged PR hand-labelled ~1/10
110
+ findings correct. Two causes, both fixed here: (a) process/agent-behaviour
111
+ rules (commit-message formatting, TDD workflow, an agent's own prompting
112
+ standard) paired against a code hunk they were never meant to describe,
113
+ because the corpus declares no `metadata.paths`/`stack_requires` and
114
+ applicability fails open to "applies everywhere" — the category and
115
+ file-kind filters (steps 2 and 4); (b) a re-measurement then found the
116
+ `--rules` bypass defeating the category filter for an entire directory, and
117
+ the whole `--max-calls` budget landing on a single docs hunk while no code
118
+ hunk was ever scored — the scoped bypass and fair round-robin budget (steps
119
+ 1 and 5). All four are cheap, deterministic, and applied BEFORE any
120
+ violation-scoring Jev call, so they cut cost as well as noise. See
121
+ `.metaproject/flows/330-*/journal.md` for the full before/after numbers.
122
+
123
+ ---
124
+
125
+ ## Input Contract
126
+
127
+ Not dispatched with a prompt — invoked as a CLI command:
128
+
129
+ ```text
130
+ keryx review jev-rules (--diff <ref> | --pr <n> | --scope <scope.json>)
131
+ [--rules <paths>] [--max-calls <n>] [--threshold <0..1>]
132
+ [--repo <owner/repo>] [--model <jev-1.13|jev-latest>]
133
+ [--fixtures <dir>] [--json]
134
+ ```
135
+
136
+ `review-orchestrator` passes `--scope <scope.json>` — the same `keryx review
137
+ scope --json` file every other reviewer's dispatch already reads — so this
138
+ reviewer checks exactly the same hunks, with the same mechanical-bulk drops,
139
+ as everyone else in the round.
140
+
141
+ ---
142
+
143
+ ## Opt-in and privacy
144
+
145
+ Refuses before any read or network call unless `.metaproject/tasks.config.json`
146
+ declares:
147
+
148
+ ```json
149
+ { "review": { "jev": { "rules": true } } }
150
+ ```
151
+
152
+ Every hunk, every rule-clause sent for violation scoring, and every clause
153
+ sent for tagging is redacted first through `src/security/service.ts` — the
154
+ same floor `review conform` and `review ci-triage` already apply. No cache or
155
+ output file stores a credential. Two caches, both gitignored and mode `0600`
156
+ under `.metaproject/data/review-jev-rules/`:
157
+
158
+ - `violation-cache.json`, keyed on the clause text and hunk location, so a
159
+ re-run over unchanged hunks and unchanged rules costs zero additional Jev
160
+ calls.
161
+ - `clause-tags.json`, keyed on the rule doc's own content hash — REUSING
162
+ `review conform`'s own clause-tagging cache module
163
+ (`src/review/conform-tag-cache.ts`) at this jev-rules-specific path, so a
164
+ re-run against a DIFFERENT diff but the SAME rule corpus costs zero
165
+ additional tagging calls even when the hunks (and so the violation cache)
166
+ miss.
167
+
168
+ ---
169
+
170
+ ## Output Contract
171
+
172
+ Emits a `REVIEW_RESULT`-shaped object matching
173
+ `.metaproject/skills/gdskills/review/review-orchestrator/reviewer-finding.schema.json`
174
+ — `status`, `reviewer: "review-jev-rules"`, `summary`, `findings`, `stats`,
175
+ plus `tokens` (with `taggingCalls`/`violationCalls` counted separately),
176
+ `selection` (`droppedClauses`'s count, and `hunkCoverage` — one
177
+ `{path, applicablePairs, selectedPairs}` per hunk with something to check,
178
+ plus `hunksWithPairs`/`hunksReached`/`hunksNeverReached`), `ruleSources`
179
+ (each with its `category`), `excludedSources` (category-filtered sources
180
+ with their reason), `droppedClauses` (tag-filtered `(ruleId, clauseId)`
181
+ pairs with their reason), and `droppedPlaceholderClauses`
182
+ (authoring-template clauses dropped before tagging, with their reason) —
183
+ under `--json`. `summary` names how many hunks the budget never reached in
184
+ plain text. Its findings merge into the consolidated array exactly like any
185
+ other reviewer's: same Quality Gate, same dedup, same Wave C verification.
186
+
187
+ ---
188
+
189
+ ## Orchestrator integration
190
+
191
+ - `keryx review reviewers --json` lists it under `bundled` with
192
+ `"engine": "jev"`.
193
+ - `review-orchestrator` dispatches it in **Wave B**, alongside the domain/
194
+ convention reviewers, when the opt-in is on and a Jev/OpenRouter credential
195
+ resolves (`resolveJevApiKeyResolution`). When either is false it is
196
+ **skipped with the stated reason**, recorded in `Skipped reviewers` —
197
+ never silently absent.
198
+ - It is **ADDITIONAL**: it never replaces `review-style`, the convention
199
+ reviewers, or any other pass. A rule clause this reviewer flags may also be
200
+ the exact concern a domain reviewer raises independently; the dedup pass
201
+ (by `dedupe_key`/`(file, quote, problem)`) merges the two rather than
202
+ double-counting.
203
+
204
+ ---
205
+
206
+ ### Shared laws (every reviewer)
207
+
208
+ 1. **A claim of runtime harm with no reproducible path is `info`.** If you cannot
209
+ name the input, call, or condition that reaches the code, you have an
210
+ observation, not a finding. Report it as `info` and say what would settle it.
211
+ 2. **Never flag the theoretical.** The path you describe must exist in the code
212
+ under review. Do not report a safe API because it could be misused, or a
213
+ pattern because it is often wrong elsewhere.
214
+ 3. **One finding per class, not one per occurrence.** When the same shape appears
215
+ at several sites, report it once and list every site. Ten findings that are one
216
+ finding hide the other nine problems.
217
+
218
+ This reviewer already satisfies all three by construction: `evidence` always
219
+ names the hunk and quotes it (never an unreproducible claim), a finding is
220
+ only synthesized when the hunk itself is checked against the clause (never a
221
+ theoretical pattern), and `synthesizeFindingsFromViolations` dedupes by
222
+ `(clause, file)` with every hunk listed, never one finding per hunk.
223
+
224
+ ---
225
+
226
+ ## Red Flags
227
+
228
+ | Rationalization | Why it is wrong |
229
+ |---|---|
230
+ | "Jev said 0.9, so this is definitely a violation" | A `noul` score is a probability, not a verdict — severity stays capped at `minor` unless the rule itself declares higher, precisely because Jev alone is not authoritative |
231
+ | "This rule clause doesn't really describe code, but the pair matched anyway" | Applicability defaults to "applies everywhere" when a rule declares no path/stack restriction — the clause-kind filter (step 3), the category filter (step 2), and the file-kind gate (step 4) catch most of this before scoring, but a `state_kind: "hunk"` clause on an off-topic rule can still pass through; that is a property of the rule corpus, not a defect in this reviewer |
232
+ | "No findings means the diff is clean" | `--max-calls` bounds how many pairs are checked; a capped run reports `selection.droppedPairs`/`hunksNeverReached` for exactly this reason — read the selection stats before treating silence as clean |
233
+ | "A rule wasn't checked and I don't know why" | Read `excludedSources` (category filter, at discovery), `droppedClauses` (tag filter, per clause), and `droppedPlaceholderClauses` (template scaffolding, per clause) before assuming a rule was silently skipped — each names the exact reason |
234
+ | "I passed `--rules` at a whole rule directory, so every file in it was checked" | Only a `--rules` entry naming ONE file bypasses the category filter; a directory's files are filtered exactly like auto-discovery — check `excludedSources` |
235
+ | "I'll skip the opt-in check since I trust this project" | The opt-in and credential gates exist because hunk and rule-clause text leaves the machine; skipping them is skipping consent, not a shortcut |
236
+
237
+ ---
238
+
239
+ ## Verification
240
+
241
+ Before trusting a run's findings:
242
+
243
+ 1. Read `selection.droppedPairs`/`selection.notApplicable`/
244
+ `selection.droppedClauses`/`selection.hunkCoverage` in the `--json`
245
+ output — a capped, narrowly-applicable, or tag-filtered run covered less
246
+ than "every changed hunk against every applicable clause", and
247
+ `hunksNeverReached` says exactly how many hunks the budget never got to.
248
+ 2. Read `excludedSources` — a process/meta rule doc excluded by the category
249
+ filter is reported by name and reason, never silently absent from
250
+ `ruleSources`.
251
+ 3. Spot-check a handful of findings against the actual hunk: does the quoted
252
+ line and the cited clause text support the claim? Jev's probability is not
253
+ evidence a human can inspect; the quote and the clause text are.
254
+ 4. Confirm every finding carries `reviewer: "review-jev-rules"` and a
255
+ `dedupe_key` — both are required for the orchestrator's Quality Gate and
256
+ Wave C verification to route it correctly.
257
+
258
+ ---
259
+
260
+ ## Scope Boundaries
261
+
262
+ | Concern | This skill | Use instead |
263
+ |---------|------------|-------------|
264
+ | A hunk contradicts a documented project rule clause | YES | — |
265
+ | Logic bugs, architecture, security, performance not named by any rule | NO | the matching domain reviewer |
266
+ | Judging whether a reported finding is real | NO | `review-verifier` |
267
+ | Writing or maintaining the rules themselves | NO | a human, or `rules/core/*.mdc` directly |
@@ -43,3 +43,29 @@ suppressing it would trade real coverage for tidiness. Record the drift in
43
43
  `review_context` and name it once in the report, so the next person knows the
44
44
  profile is due a re-read. Never file it as a finding against the code under
45
45
  review: it is a fact about the review, not about the diff.
46
+
47
+ ## CLI-engine reviewers — dispatched as a command, not a sub-agent
48
+
49
+ `review-jev-rules` (flow 330) is an ADDITIONAL reviewer, never replacing any
50
+ other, and its dispatch mechanism differs from every reviewer named in
51
+ SKILL.md's Routing Table: there is no platform-native agent to invoke,
52
+ because it is a deterministic **keryx program**. Run it with `keryx review
53
+ jev-rules --scope <scope.json> --json` — the SAME `scope.json` every other
54
+ Wave A/B reviewer's dispatch already reads, so it checks exactly the same
55
+ hunks. Read its `--json` output as a `REVIEW_RESULT` and merge its `findings`
56
+ into the consolidated array exactly like a sub-agent reviewer's: same
57
+ Sub-Agent Report Quality Gate, same dedup, same Wave C verification.
58
+
59
+ Gate it BEFORE running the command, not after: skip it — recorded in `Skipped
60
+ reviewers` with the reason, never silently absent — when `review.jev.rules`
61
+ is not `true` in `.metaproject/tasks.config.json`, or when no Jev/OpenRouter
62
+ credential is resolvable. Either gate failing means `keryx review jev-rules`
63
+ itself would refuse before any read or network call, so checking first saves
64
+ a doomed dispatch.
65
+
66
+ `keryx review reviewers --json` marks a CLI-engine reviewer with
67
+ `"engine": "jev"` on its `bundled` entry — the field's presence, not its
68
+ absence, is what distinguishes it from the default (an LLM sub-agent
69
+ dispatch). A future engine-backed reviewer follows the same pattern: gate on
70
+ its own opt-in and reachability, dispatch as a command, merge its `--json`
71
+ output the same way.