aireview 2.1.0 → 2.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +47 -0
- data/README.md +182 -5
- data/config/.aireview.yml.example +6 -0
- data/lib/aireview/cli.rb +11 -2
- data/lib/aireview/config.rb +26 -8
- data/lib/aireview/config_fallbacks.rb +43 -15
- data/lib/aireview/config_jev.rb +137 -0
- data/lib/aireview/config_layers.rb +1 -1
- data/lib/aireview/config_loader.rb +32 -6
- data/lib/aireview/context_builder.rb +13 -10
- data/lib/aireview/dry_run_prompts.rb +104 -0
- data/lib/aireview/dry_run_report.rb +40 -8
- data/lib/aireview/errors.rb +11 -0
- data/lib/aireview/jev_client.rb +129 -0
- data/lib/aireview/jev_critic.rb +387 -0
- data/lib/aireview/jev_shadow.rb +82 -0
- data/lib/aireview/jev_stage.rb +97 -0
- data/lib/aireview/llm_client.rb +52 -21
- data/lib/aireview/llm_router.rb +4 -4
- data/lib/aireview/model_candidate.rb +34 -5
- data/lib/aireview/model_checker.rb +46 -5
- data/lib/aireview/model_pool.rb +57 -28
- data/lib/aireview/prompts/jev_questions.yml +79 -0
- data/lib/aireview/review_marker.rb +7 -2
- data/lib/aireview/review_pipeline.rb +40 -55
- data/lib/aireview/review_renderer.rb +18 -2
- data/lib/aireview/stage_chains.rb +6 -3
- data/lib/aireview/version.rb +1 -1
- metadata +12 -4
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 7919de6ffa17aab9f5250ebf07dfec9ce83db9313ef01e6ba93a3b131da927ce
|
|
4
|
+
data.tar.gz: 348cda52953ff2d22ba56fb16b192876249efa30c5ba77f6f0736c1b96e5f569
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 372867e5a1547b4d595312498107dd8c33c6fe6721efb8f4c215850660bcf57706f5a864c9d9b7d46a9d2a1e29fdf18ab31e9ddb77c1859237ddbe5d979b2946
|
|
7
|
+
data.tar.gz: b3b079dbb2b5b92ffe39fc36c6de224d9bf2268e11b353129236e15b5c14689a9067012da053b8415bb53f64b2f52247ef252062831e9e73d789b943a611ed0a
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,52 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 2.2.1
|
|
4
|
+
|
|
5
|
+
- OpenAI reasoning models (`o1`, `o3`, `gpt-5`…) and search models
|
|
6
|
+
(`gpt-4o-search-preview`…) get no temperature: RubyLLM 2 sends it as given
|
|
7
|
+
(1.x replaced it with 1.0 or dropped it), and these models reject any
|
|
8
|
+
other value with a bad request, which failed the run without trying a
|
|
9
|
+
fallback model. The RubyLLM registry says which models take a
|
|
10
|
+
temperature; a model it does not know is judged by its name, as 1.x did.
|
|
11
|
+
`models check` is fixed the same way.
|
|
12
|
+
|
|
13
|
+
## 2.2.0
|
|
14
|
+
|
|
15
|
+
- Jev shadow mode (`llm.jev.shadow`, `LLM_JEV_SHADOW`, key `JEV_API_KEY`):
|
|
16
|
+
after the LLM Critique the same candidates go to Jev (TypeSafe), a fast
|
|
17
|
+
classifier, and its keep/reject/unverifiable decisions with the raw
|
|
18
|
+
probabilities are logged next to the Critique verdicts. The report and the
|
|
19
|
+
review key do not change; a Jev failure is a warning. The data is for
|
|
20
|
+
choosing between an LLM and Jev as the critic later.
|
|
21
|
+
- Critique engine (`llm.critique.engine`, `LLM_CRITIQUE_ENGINE`,
|
|
22
|
+
`--critique-engine`): `model` (the default, the LLM Critique as before) or
|
|
23
|
+
`jev`. With `jev` Jev decides keep/reject without refining the wording;
|
|
24
|
+
the candidates it cannot judge, and all of them when Jev fails, go to the
|
|
25
|
+
LLM Critique with `llm.jev.fallback: model` (the default), or are
|
|
26
|
+
rejected / fail the run with `fail`. Duplicates are dropped once over the
|
|
27
|
+
merged verdicts. The report says how Jev took part. The Jev thresholds
|
|
28
|
+
default to a provisional 0.5 until they are chosen from the shadow logs.
|
|
29
|
+
- The stages that need an LLM follow the engine: with Jev and
|
|
30
|
+
`fallback: fail` the critique model, its key, its share of the context
|
|
31
|
+
budget, the pool's critique policy and its `models check` probe are not
|
|
32
|
+
required. `models check` probes Jev as well (counted only for the engine).
|
|
33
|
+
- The review key includes the Jev version, thresholds, fallback and
|
|
34
|
+
question templates with `engine: jev`; with the default engine it does not
|
|
35
|
+
change, except for runs with `--no-critique`: the unused critique policy
|
|
36
|
+
of the pool (`critique.start`, `rank`, `allow_weaker`) no longer enters
|
|
37
|
+
their key and is no longer validated, so such a review is redone once.
|
|
38
|
+
- `--dry-run` prints the Jev request for the stub candidate.
|
|
39
|
+
- The gem now packages `lib/**/*.yml` (the Jev questions).
|
|
40
|
+
- A pool of models from several providers: OpenAI, Anthropic and OpenRouter
|
|
41
|
+
take their keys from `OPENAI_API_KEY(S)`, `ANTHROPIC_API_KEY(S)`,
|
|
42
|
+
`OPENROUTER_API_KEY(S)`, and a string like `anthropic/claude-opus-4.5` in
|
|
43
|
+
`LLM_MODELS` or a reserve names the provider (an OpenRouter model goes with
|
|
44
|
+
the prefix: `openrouter/qwen/qwen3-coder`).
|
|
45
|
+
- `api_base` for a model of the pool, a stage and a reserve: several servers
|
|
46
|
+
of your own in one pool. Such a model is named with its address, gets
|
|
47
|
+
`LLM_API_KEY` or no auth — never a provider's key — and needs no key to
|
|
48
|
+
start. Models without an address keep their names and review keys.
|
|
49
|
+
|
|
3
50
|
## 2.1.0
|
|
4
51
|
|
|
5
52
|
- An answer that is not valid JSON and that the provider cut off at the
|
data/README.md
CHANGED
|
@@ -79,11 +79,18 @@ REVIEW_LANGUAGE=ru
|
|
|
79
79
|
REVIEW_MODE=update
|
|
80
80
|
```
|
|
81
81
|
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
82
|
+
The providers are `gemini`, `ollama`, `openai`, `anthropic` and
|
|
83
|
+
`openrouter`, and any OpenAI-compatible server of your own (vLLM, llama.cpp,
|
|
84
|
+
LM Studio) as `openai` with an `api_base`. The provider and the model of each
|
|
85
|
+
stage are set through `LLM_GENERATE_PROVIDER`, `LLM_GENERATE_MODEL`,
|
|
86
|
+
`LLM_CRITIQUE_PROVIDER` and `LLM_CRITIQUE_MODEL`. `LLM_PROVIDER` stays the
|
|
87
|
+
shared default when a stage has no provider of its own. Gemini and Ollama
|
|
88
|
+
are the ones the image defaults and the verified models are about; for the
|
|
89
|
+
others run `aireview models check` before relying on a model.
|
|
90
|
+
|
|
91
|
+
Each provider takes its keys from its own variables:
|
|
92
|
+
`GEMINI_API_KEY(S)`, `OPENAI_API_KEY(S)`, `ANTHROPIC_API_KEY(S)`,
|
|
93
|
+
`OPENROUTER_API_KEY(S)`; `LLM_API_KEY` is the shared fallback.
|
|
87
94
|
|
|
88
95
|
`REVIEW_LANGUAGE` (or `review_language` in `.aireview.yml`) sets the language
|
|
89
96
|
of the review: both the LLM answers and the headings of the rendered report.
|
|
@@ -175,6 +182,58 @@ the end of the run (see "Fallback models and keys"). The separate smoke test
|
|
|
175
182
|
of every model with both schemas (`aireview models check`) runs on an image
|
|
176
183
|
release and on a schedule, not on MRs.
|
|
177
184
|
|
|
185
|
+
### A pool of models from several providers
|
|
186
|
+
|
|
187
|
+
The critic is best the strongest model you have, and it does not have to be
|
|
188
|
+
Gemini: the shared pool (see "Shared model pool") takes models of any
|
|
189
|
+
provider, ordered by strength as you see it — the order is the only measure,
|
|
190
|
+
names and release dates are not compared.
|
|
191
|
+
|
|
192
|
+
```bash
|
|
193
|
+
LLM_MODELS=anthropic/claude-opus-4.5,openai/gpt-5,gemini/gemini-3.8-flash,ollama/qwen2.5-coder:14b
|
|
194
|
+
LLM_GENERATE_START=gemini-3.8-flash
|
|
195
|
+
ANTHROPIC_API_KEY=xxx
|
|
196
|
+
OPENAI_API_KEY=xxx
|
|
197
|
+
GEMINI_API_KEY=xxx
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
In a string the provider goes in front of the model; an OpenRouter model has
|
|
201
|
+
a slash of its own, so it is written with the prefix:
|
|
202
|
+
`openrouter/qwen/qwen3-coder`. Servers of your own, one or several, are
|
|
203
|
+
listed in `.aireview.yml` with their address:
|
|
204
|
+
|
|
205
|
+
```yaml
|
|
206
|
+
llm:
|
|
207
|
+
models:
|
|
208
|
+
- provider: anthropic
|
|
209
|
+
model: claude-opus-4.5
|
|
210
|
+
- provider: openai # an OpenAI-compatible server of your own
|
|
211
|
+
model: qwen3-coder
|
|
212
|
+
api_base: http://gpu1:8000/v1
|
|
213
|
+
- provider: ollama
|
|
214
|
+
model: qwen2.5-coder:14b
|
|
215
|
+
api_base: http://gpu2:11434/v1
|
|
216
|
+
generate:
|
|
217
|
+
start: qwen2.5-coder:14b
|
|
218
|
+
```
|
|
219
|
+
|
|
220
|
+
`api_base` also works for a stage (`llm.generate.api_base`) and for a reserve
|
|
221
|
+
in `fallbacks`. A model with its own `api_base` is named with the address
|
|
222
|
+
(`openai/qwen3-coder@http://gpu1:8000/v1`), so two servers with the same model are
|
|
223
|
+
two models for quarantine and the review key, and a start may name either.
|
|
224
|
+
Such a server never gets a provider's key: it is called with `LLM_API_KEY`,
|
|
225
|
+
or without auth when that is not set, and needs no key to start. Without an
|
|
226
|
+
`api_base` a model goes to its provider's API, or to `LLM_API_BASE` for
|
|
227
|
+
Gemini, OpenAI and OpenRouter as before.
|
|
228
|
+
|
|
229
|
+
A stage's temperature goes only to the models that take one: OpenAI
|
|
230
|
+
reasoning models (`o1`, `o3`, `gpt-5`…) and search models
|
|
231
|
+
(`gpt-4o-search-preview`…) accept no other, so they get none and run with
|
|
232
|
+
their own default. The RubyLLM registry decides, and the name
|
|
233
|
+
for a model it does not know.
|
|
234
|
+
|
|
235
|
+
Instead of an LLM the critic can be Jev (see "Critique engine").
|
|
236
|
+
|
|
178
237
|
### Local Ollama
|
|
179
238
|
|
|
180
239
|
Install Ollama following the
|
|
@@ -522,6 +581,124 @@ Critique receives the result of the check in the candidate's `note` field
|
|
|
522
581
|
and decides keep/reject with it in mind. With `--no-critique` the check works
|
|
523
582
|
the same way, its marks just go straight to the report.
|
|
524
583
|
|
|
584
|
+
### Critique engine: an LLM or Jev
|
|
585
|
+
|
|
586
|
+
`llm.critique.engine` picks who checks the candidates of the first pass:
|
|
587
|
+
|
|
588
|
+
| | `model` (default) | `jev` |
|
|
589
|
+
|---|---|---|
|
|
590
|
+
| What checks | the LLM Critique chain or pool (`LLM_CRITIQUE_MODEL`, `critique.rank`) | [Jev](https://docs.typesafe.ai), a classifier by TypeSafe |
|
|
591
|
+
| Context | the whole prompt: MR, Jira, diff | the same, cut to Jev's limits (see below) |
|
|
592
|
+
| Wording of findings | refined by the critic | the first pass's own |
|
|
593
|
+
| Time and cost | tens of seconds, provider quota | under a second, per input token |
|
|
594
|
+
| Depends on Gemini | when the critique model is Gemini | no, with `fallback: fail` |
|
|
595
|
+
| Non-English reviews | fine | not declared by TypeSafe, check first |
|
|
596
|
+
| Code leaves for | the LLM provider | TypeSafe as well |
|
|
597
|
+
|
|
598
|
+
The default stays `model`. A stronger critic than the generator is set with
|
|
599
|
+
the existing settings (`LLM_CRITIQUE_MODEL`, or the order of the pool); Jev is
|
|
600
|
+
meant for when speed or independence from the LLM quota matters more than
|
|
601
|
+
refined wording, and only after its thresholds are chosen from the shadow
|
|
602
|
+
logs (see below).
|
|
603
|
+
|
|
604
|
+
```yaml
|
|
605
|
+
llm:
|
|
606
|
+
critique:
|
|
607
|
+
engine: jev # LLM_CRITIQUE_ENGINE, --critique-engine jev
|
|
608
|
+
jev:
|
|
609
|
+
model: jev-1.13.0 # LLM_JEV_MODEL; a pinned version, an alias is an error here
|
|
610
|
+
fallback: model # LLM_JEV_FALLBACK: model or fail
|
|
611
|
+
keep_above: 0.5 # LLM_JEV_KEEP_ABOVE; enough_context, version_claim, duplicate in YAML
|
|
612
|
+
```
|
|
613
|
+
|
|
614
|
+
With `engine: jev` Jev answers the questions of `prompts/jev_questions.yml`
|
|
615
|
+
about every candidate and decides:
|
|
616
|
+
|
|
617
|
+
- claims that a version does not exist — reject;
|
|
618
|
+
- not enough in the request to judge it, or too large for a Jev request —
|
|
619
|
+
unverifiable: with `fallback: model` such candidates go to the LLM
|
|
620
|
+
Critique (only them), with `fallback: fail` they are rejected;
|
|
621
|
+
- otherwise keep when `real_issue` reaches `keep_above`;
|
|
622
|
+
- duplicates are dropped once, over what Jev and the LLM kept together: the
|
|
623
|
+
more severe finding stays, on a tie the one the first pass listed first.
|
|
624
|
+
|
|
625
|
+
When Jev fails (network, 429/529 after one retry, an invalid answer), the
|
|
626
|
+
LLM Critique takes over with `fallback: model`, and the run fails with
|
|
627
|
+
`fallback: fail`. The report says how Jev took part: checked by Jev without
|
|
628
|
+
refining the wording, partly by the LLM, or by the LLM because Jev was
|
|
629
|
+
unavailable. `engine: jev` without `JEV_API_KEY` or with an alias is a
|
|
630
|
+
configuration error at start; `--no-critique` switches off every engine.
|
|
631
|
+
|
|
632
|
+
What an LLM is still needed for follows the engine: with `fallback: fail`
|
|
633
|
+
there is no LLM Critique at all, so its model and key are not required
|
|
634
|
+
(generate on Ollama with Jev needs no Gemini key), the context budget counts
|
|
635
|
+
generate only, the pool's critique policy is not used and `models check`
|
|
636
|
+
probes the generate stage only. `models check` also sends Jev a probe
|
|
637
|
+
request: with `engine: jev` it counts, in shadow mode it is shown but never
|
|
638
|
+
fails the check. `--dry-run` shows the Jev request for the stub candidate
|
|
639
|
+
(`=== JEV STATE ===`, `=== JEV QUESTIONS ===`).
|
|
640
|
+
|
|
641
|
+
The review key follows the engine too: with `engine: jev` it includes the Jev
|
|
642
|
+
version, every threshold, the fallback and the question templates; the LLM
|
|
643
|
+
Critique model counts only while Jev can fall back to it. With the default
|
|
644
|
+
engine the key is what it was before Jev, whatever `llm.jev` says. Without
|
|
645
|
+
an LLM Critique (Jev with `fallback: fail`, or `--no-critique`) the critique
|
|
646
|
+
settings — model, limit, start, rank, allow_weaker — are neither validated
|
|
647
|
+
nor part of the key, while the generate pool stays in it.
|
|
648
|
+
|
|
649
|
+
### Jev shadow (experiment)
|
|
650
|
+
|
|
651
|
+
Jev is a classifier, not a text model: it answers yes/no and choice
|
|
652
|
+
questions about a given state with probabilities. Before it decides anything
|
|
653
|
+
as the engine, it can run in shadow mode next to the default engine: after
|
|
654
|
+
the LLM Critique the same candidates go to Jev, and its decisions are logged
|
|
655
|
+
next to the Critique verdicts. The report and the review key do not change;
|
|
656
|
+
with `engine: jev` the shadow is ignored.
|
|
657
|
+
|
|
658
|
+
```yaml
|
|
659
|
+
llm:
|
|
660
|
+
jev:
|
|
661
|
+
shadow: true # LLM_JEV_SHADOW=true
|
|
662
|
+
model: jev-1.13.0 # LLM_JEV_MODEL; a pinned version, not jev-latest
|
|
663
|
+
timeout: 10 # seconds per request
|
|
664
|
+
keep_above: 0.5 # provisional thresholds, only for the log line
|
|
665
|
+
enough_context: 0.5
|
|
666
|
+
version_claim: 0.5
|
|
667
|
+
duplicate: 0.5
|
|
668
|
+
```
|
|
669
|
+
|
|
670
|
+
The key is `JEV_API_KEY`. For every candidate Jev is asked whether it is a
|
|
671
|
+
real, well-supported problem (the rules of `prompts/critique.txt`, asked in
|
|
672
|
+
`prompts/jev_questions.yml`), whether the state holds enough to judge it,
|
|
673
|
+
whether it claims that a version does not exist, which other candidate it
|
|
674
|
+
duplicates, and how severe it is. The log gets one line per candidate and a
|
|
675
|
+
summary:
|
|
676
|
+
|
|
677
|
+
```
|
|
678
|
+
Jev shadow C1: keep (real issue; real_issue=0.91 enough_context=0.88 version_claim=0.02 duplicate_of=none/0.97 severity=major/0.74), critique: keep
|
|
679
|
+
Jev shadow: agrees with critique on 2 of 2 decided candidate(s); keep 1, reject 1, unverifiable 1 (model=jev-1.13.0, requests=1, 0.4s)
|
|
680
|
+
```
|
|
681
|
+
|
|
682
|
+
The readable lines round the numbers; a third line, `Jev shadow data: {…}`,
|
|
683
|
+
carries every decision with the probabilities exactly as Jev returned them,
|
|
684
|
+
so thresholds can be chosen from the logs afterwards. The version that
|
|
685
|
+
answered is logged too, and a pinned version answering as another one is a
|
|
686
|
+
warning. A Jev failure (network, 429/529 after
|
|
687
|
+
one retry, an invalid answer, no key) is a warning and never affects the
|
|
688
|
+
review. Jev goes out through `LLM_HTTP_PROXY` like the LLM providers.
|
|
689
|
+
|
|
690
|
+
What leaves for TypeSafe: the MR title and description, the Jira section,
|
|
691
|
+
`review_instructions`, the diff shown to the models and the candidates, after
|
|
692
|
+
the same secret scrubbing as the LLM prompts. Jev limits a request to 32k
|
|
693
|
+
tokens for the state plus the longest question and 64k for the state plus all
|
|
694
|
+
questions; the state is cut to what the questions leave (the diff first, then
|
|
695
|
+
the MR and Jira sections, never the candidates and the hunks they point at),
|
|
696
|
+
and when the candidates do not fit together, each goes in a request of its
|
|
697
|
+
own. The size is estimated from characters; if Jev still rejects a request
|
|
698
|
+
as too large (a 422 that explicitly says it exceeds a limit), its candidates
|
|
699
|
+
are asked about one by one, and one that is too large alone stays
|
|
700
|
+
unverifiable. Any other 422 is not retried.
|
|
701
|
+
|
|
525
702
|
## Usage
|
|
526
703
|
|
|
527
704
|
```bash
|
|
@@ -36,3 +36,9 @@ llm:
|
|
|
36
36
|
model: gemini-3.8-flash
|
|
37
37
|
fallbacks:
|
|
38
38
|
- gemini-3.7-flash
|
|
39
|
+
# Jev as the critique engine (critique: {engine: jev}) or next to it, log only
|
|
40
|
+
# (see README, "Critique engine"); the key is JEV_API_KEY
|
|
41
|
+
# jev:
|
|
42
|
+
# shadow: true
|
|
43
|
+
# model: jev-1.13.0
|
|
44
|
+
# fallback: model
|
data/lib/aireview/cli.rb
CHANGED
|
@@ -99,9 +99,11 @@ module Aireview
|
|
|
99
99
|
critique_model: options[:critique_model],
|
|
100
100
|
generate_temperature: options[:generate_temperature],
|
|
101
101
|
critique_temperature: options[:critique_temperature],
|
|
102
|
-
|
|
102
|
+
critique_engine: options[:critique_engine],
|
|
103
|
+
no_fallbacks: options[:no_fallbacks] == true,
|
|
104
|
+
no_critique: options[:no_critique] == true
|
|
103
105
|
)
|
|
104
|
-
config.require_llm_configuration!
|
|
106
|
+
config.require_llm_configuration!(critique: !options[:no_critique])
|
|
105
107
|
config.warnings.each { |warning| @logger.warn(warning) }
|
|
106
108
|
config
|
|
107
109
|
end
|
|
@@ -335,6 +337,11 @@ module Aireview
|
|
|
335
337
|
options[:critique_temperature] = value
|
|
336
338
|
end
|
|
337
339
|
|
|
340
|
+
parser.on('--critique-engine ENGINE', Aireview::Config::CRITIQUE_ENGINES,
|
|
341
|
+
'Who checks the candidates: model (an LLM) or jev') do |value|
|
|
342
|
+
options[:critique_engine] = value
|
|
343
|
+
end
|
|
344
|
+
|
|
338
345
|
parser.on('--no-critique', 'Skip critique pass and render Generate candidates directly') do
|
|
339
346
|
options[:no_critique] = true
|
|
340
347
|
end
|
|
@@ -407,6 +414,8 @@ module Aireview
|
|
|
407
414
|
Override Generate pass temperature
|
|
408
415
|
--critique-temperature VALUE
|
|
409
416
|
Override Critique pass temperature
|
|
417
|
+
--critique-engine ENGINE
|
|
418
|
+
Who checks the candidates: model (an LLM, default) or jev
|
|
410
419
|
--config PATH Path to .aireview.yml
|
|
411
420
|
--review-mode MODE
|
|
412
421
|
How to treat an existing review: update (default) or once
|
data/lib/aireview/config.rb
CHANGED
|
@@ -6,6 +6,7 @@ require_relative 'errors'
|
|
|
6
6
|
require_relative 'utils'
|
|
7
7
|
require_relative 'config_limits'
|
|
8
8
|
require_relative 'config_fallbacks'
|
|
9
|
+
require_relative 'config_jev'
|
|
9
10
|
require_relative 'config_layers'
|
|
10
11
|
require_relative 'config_loader'
|
|
11
12
|
|
|
@@ -16,6 +17,7 @@ module Aireview
|
|
|
16
17
|
class Config
|
|
17
18
|
include ConfigLimits
|
|
18
19
|
include ConfigFallbacks
|
|
20
|
+
include ConfigJev
|
|
19
21
|
include ConfigLayers
|
|
20
22
|
|
|
21
23
|
DEFAULT_SECRET_FILES = [
|
|
@@ -76,20 +78,27 @@ module Aireview
|
|
|
76
78
|
|
|
77
79
|
# CLI overrides change only the primary model of a stage, the reserves
|
|
78
80
|
# from the config stay; no_fallbacks leaves one model and one key.
|
|
79
|
-
|
|
81
|
+
# no_critique — --no-critique: the run has no Critique, so its settings
|
|
82
|
+
# are neither routed nor validated (see llm_stages).
|
|
83
|
+
# Seven named flags of the CLI read better than a struct for its own sake.
|
|
84
|
+
def with_overrides( # rubocop:disable Metrics/ParameterLists
|
|
80
85
|
generate_model: nil,
|
|
81
86
|
critique_model: nil,
|
|
82
87
|
generate_temperature: nil,
|
|
83
88
|
critique_temperature: nil,
|
|
84
|
-
|
|
89
|
+
critique_engine: nil,
|
|
90
|
+
no_fallbacks: false,
|
|
91
|
+
no_critique: false
|
|
85
92
|
)
|
|
86
93
|
llm_config = {
|
|
87
94
|
'generate' => stage_overrides(model: generate_model, temperature: generate_temperature),
|
|
88
95
|
'critique' => stage_overrides(model: critique_model, temperature: critique_temperature)
|
|
96
|
+
.merge({'engine' => critique_engine}.compact)
|
|
89
97
|
}.reject { |_, overrides| overrides.empty? }
|
|
90
98
|
overrides = {}
|
|
91
99
|
overrides['llm'] = llm_config unless llm_config.empty?
|
|
92
100
|
overrides['fallbacks_disabled'] = true if no_fallbacks
|
|
101
|
+
overrides['critique_disabled'] = true if no_critique
|
|
93
102
|
return self if overrides.empty?
|
|
94
103
|
|
|
95
104
|
self.class.new(
|
|
@@ -137,8 +146,9 @@ module Aireview
|
|
|
137
146
|
routing.primary('generate').model
|
|
138
147
|
end
|
|
139
148
|
|
|
149
|
+
# nil when no LLM Critique can run (see llm_stages).
|
|
140
150
|
def critique_model
|
|
141
|
-
routing.primary('critique').model
|
|
151
|
+
routing.stage?('critique') ? routing.primary('critique').model : nil
|
|
142
152
|
end
|
|
143
153
|
|
|
144
154
|
def generate_provider
|
|
@@ -147,18 +157,26 @@ module Aireview
|
|
|
147
157
|
|
|
148
158
|
# Everything besides the prompt that affects the review result goes into
|
|
149
159
|
# the note key (see ReviewMarker): provider, model and temperature of the
|
|
150
|
-
# stages, the shared pool with its critique
|
|
151
|
-
# chains do not change the result.
|
|
160
|
+
# stages, Jev as the critique engine, the shared pool with its critique
|
|
161
|
+
# policy. Reserves of per-stage chains do not change the result.
|
|
152
162
|
def result_signature
|
|
153
163
|
{
|
|
154
|
-
'generate' =>
|
|
155
|
-
'critique' =>
|
|
164
|
+
'generate' => stage_signature('generate', generate_temperature),
|
|
165
|
+
'critique' => (routing.stage?('critique') ? stage_signature('critique', critique_temperature) : nil),
|
|
166
|
+
'critique_engine' => jev_signature,
|
|
156
167
|
'pool' => routing.signature
|
|
157
168
|
}
|
|
158
169
|
end
|
|
159
170
|
|
|
160
171
|
def critique_provider
|
|
161
|
-
routing.primary('critique').provider
|
|
172
|
+
routing.stage?('critique') ? routing.primary('critique').provider : nil
|
|
173
|
+
end
|
|
174
|
+
|
|
175
|
+
# The primary of a stage on a server of your own goes into the review key
|
|
176
|
+
# with its address; without one the signature is what it was.
|
|
177
|
+
def stage_signature(stage, temperature)
|
|
178
|
+
primary = routing.primary(stage)
|
|
179
|
+
[primary.provider, primary.model, temperature, *primary.api_base]
|
|
162
180
|
end
|
|
163
181
|
|
|
164
182
|
def generate_temperature
|
|
@@ -27,6 +27,19 @@ module Aireview
|
|
|
27
27
|
routing.chain(stage)
|
|
28
28
|
end
|
|
29
29
|
|
|
30
|
+
# The keys a model is called with. A server of your own (a model with
|
|
31
|
+
# api_base) never gets the provider's key — OPENAI_API_KEY must not leave
|
|
32
|
+
# for a self-hosted server — but LLM_API_KEY, or a placeholder for a
|
|
33
|
+
# server without auth, since the client wants some key.
|
|
34
|
+
SELF_HOSTED_PLACEHOLDER_KEY = 'no-key'
|
|
35
|
+
|
|
36
|
+
def candidate_api_keys(candidate)
|
|
37
|
+
return provider_api_keys(candidate.provider) unless candidate.api_base
|
|
38
|
+
return [nil] if KEYLESS_PROVIDERS.include?(candidate.provider.to_s)
|
|
39
|
+
|
|
40
|
+
[Aireview::Utils.presence(llm_api_key) || SELF_HOSTED_PLACEHOLDER_KEY]
|
|
41
|
+
end
|
|
42
|
+
|
|
30
43
|
# Keys in order of preference; a single nil for a keyless provider, so
|
|
31
44
|
# that walking the chain does not depend on the provider.
|
|
32
45
|
def provider_api_keys(provider)
|
|
@@ -47,9 +60,9 @@ module Aireview
|
|
|
47
60
|
end
|
|
48
61
|
|
|
49
62
|
# Only the number of keys per provider, for --dry-run; the values never
|
|
50
|
-
# leave.
|
|
63
|
+
# leave. Servers of your own are not counted: they take LLM_API_KEY.
|
|
51
64
|
def api_key_counts(stages)
|
|
52
|
-
providers = stages.flat_map { |stage| stage_chain(stage).map(&:provider) }.uniq
|
|
65
|
+
providers = stages.flat_map { |stage| stage_chain(stage).reject(&:api_base).map(&:provider) }.uniq
|
|
53
66
|
providers.reject { |provider| KEYLESS_PROVIDERS.include?(provider.to_s) }
|
|
54
67
|
.to_h { |provider| [provider, provider_api_keys(provider).size] }
|
|
55
68
|
end
|
|
@@ -67,39 +80,53 @@ module Aireview
|
|
|
67
80
|
'llm.overloaded_quarantine')
|
|
68
81
|
end
|
|
69
82
|
|
|
70
|
-
|
|
83
|
+
# critique — the run has a critique (false with --no-critique); only the
|
|
84
|
+
# stages that need an LLM in such a run are required (see llm_stages).
|
|
85
|
+
def require_models!(critique: true)
|
|
86
|
+
stages = llm_stages(critique: critique)
|
|
71
87
|
missing = []
|
|
72
88
|
missing << 'llm.generate.model (or LLM_GENERATE_MODEL)' if Aireview::Utils.blank?(generate_model)
|
|
73
|
-
|
|
89
|
+
if stages.include?('critique') && Aireview::Utils.blank?(critique_model)
|
|
90
|
+
missing << 'llm.critique.model (or LLM_CRITIQUE_MODEL)'
|
|
91
|
+
end
|
|
74
92
|
raise ConfigError, "LLM models are required: #{missing.join(', ')}" unless missing.empty?
|
|
75
93
|
end
|
|
76
94
|
|
|
77
95
|
# The plan is built whole (pool and critique policy validated) before the
|
|
78
|
-
# first request, not in Critique after a paid-for Generate.
|
|
79
|
-
|
|
80
|
-
|
|
96
|
+
# first request, not in Critique after a paid-for Generate. Keys are
|
|
97
|
+
# required for the stages that need an LLM: generate on Ollama with Jev
|
|
98
|
+
# as the critic and no fallback needs no Gemini key.
|
|
99
|
+
def require_llm_configuration!(critique: true)
|
|
100
|
+
require_models!(critique: critique)
|
|
101
|
+
require_jev!(critique: critique)
|
|
81
102
|
routing
|
|
82
103
|
|
|
83
|
-
missing_keys =
|
|
84
|
-
providers = stage_chain(stage).map(&:provider).uniq
|
|
85
|
-
providers.
|
|
104
|
+
missing_keys = llm_stages(critique: critique).flat_map do |stage|
|
|
105
|
+
providers = stage_chain(stage).reject(&:api_base).map(&:provider).uniq
|
|
106
|
+
providers.reject { |provider| provider_keys_present?(provider) }
|
|
107
|
+
.map { |provider| "#{stage}: API key is required for provider #{provider.inspect}" }
|
|
86
108
|
end
|
|
87
109
|
raise ConfigError, missing_keys.join(', ') unless missing_keys.empty?
|
|
88
110
|
end
|
|
89
111
|
|
|
90
112
|
private
|
|
91
113
|
|
|
114
|
+
# Only the stages that go to an LLM are routed (llm_stages): without an
|
|
115
|
+
# LLM Critique (Jev with fallback: fail, --no-critique) its model, limit,
|
|
116
|
+
# start, rank and allow_weaker are not used, so they must neither fail
|
|
117
|
+
# the run nor change the review key, nor push the pool out of it.
|
|
92
118
|
def build_routing
|
|
93
|
-
|
|
119
|
+
stages = llm_stages
|
|
120
|
+
settings = stages.to_h { |stage| [stage, stage_settings(stage)] }
|
|
94
121
|
models = Array(dig('llm', 'models'))
|
|
95
122
|
return StageChains.build(settings, only_primary: fallbacks_disabled?) if models.empty?
|
|
96
123
|
|
|
97
124
|
own = settings.select { |_, stage_settings| Aireview::Utils.present?(stage_settings[:model]) }
|
|
98
125
|
ModelPool.new(
|
|
99
|
-
items: models, provider: llm_provider,
|
|
100
|
-
limits:
|
|
101
|
-
starts:
|
|
102
|
-
inherited_starts:
|
|
126
|
+
items: models, provider: llm_provider, stages: stages,
|
|
127
|
+
limits: stages.to_h { |stage| [stage, max_prompt_chars(stage)] },
|
|
128
|
+
starts: stages.to_h { |stage| [stage, dig('llm', stage, 'start')] },
|
|
129
|
+
inherited_starts: stages.select { |stage| start_inherited?(stage) },
|
|
103
130
|
rank: dig('llm', 'critique', 'rank'), allow_weaker: dig('llm', 'critique', 'allow_weaker'),
|
|
104
131
|
own_chains: StageChains.build(own, only_primary: fallbacks_disabled?), only_primary: fallbacks_disabled?
|
|
105
132
|
)
|
|
@@ -109,6 +136,7 @@ module Aireview
|
|
|
109
136
|
{
|
|
110
137
|
provider: stage_setting(stage, 'provider') || llm_provider,
|
|
111
138
|
model: dig('llm', stage, 'model'),
|
|
139
|
+
api_base: dig('llm', stage, 'api_base'),
|
|
112
140
|
fallbacks: dig('llm', stage, 'fallbacks'),
|
|
113
141
|
max_prompt_chars: max_prompt_chars(stage)
|
|
114
142
|
}
|
|
@@ -0,0 +1,137 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
require_relative 'errors'
|
|
3
|
+
require_relative 'stages'
|
|
4
|
+
require_relative 'utils'
|
|
5
|
+
|
|
6
|
+
module Aireview
|
|
7
|
+
# The critique engine and the Jev (TypeSafe) settings. llm.critique.engine
|
|
8
|
+
# picks who checks the candidates: an LLM (model, the default) or Jev, a
|
|
9
|
+
# fast classifier that decides keep/reject but cannot refine a finding.
|
|
10
|
+
# llm.jev.fallback says what happens when Jev cannot decide: the LLM
|
|
11
|
+
# Critique takes over (model) or the run fails (fail). llm.jev.shadow runs
|
|
12
|
+
# Jev next to the LLM Critique for the log only.
|
|
13
|
+
module ConfigJev
|
|
14
|
+
CRITIQUE_ENGINES = %w[model jev].freeze
|
|
15
|
+
JEV_FALLBACKS = %w[model fail].freeze
|
|
16
|
+
# A pinned version, not an alias: thresholds are tuned against one
|
|
17
|
+
# version, and an alias moves when a release ships.
|
|
18
|
+
DEFAULT_JEV_MODEL = 'jev-1.13.0'
|
|
19
|
+
JEV_ALIASES = %w[jev-latest jev-preview].freeze
|
|
20
|
+
DEFAULT_JEV_TIMEOUT = 10
|
|
21
|
+
# Provisional values until real thresholds are chosen from the shadow
|
|
22
|
+
# logs; the logs carry the raw probabilities for that.
|
|
23
|
+
JEV_THRESHOLD_DEFAULTS = {
|
|
24
|
+
'keep_above' => 0.5,
|
|
25
|
+
'enough_context' => 0.5,
|
|
26
|
+
'version_claim' => 0.5,
|
|
27
|
+
'duplicate' => 0.5
|
|
28
|
+
}.freeze
|
|
29
|
+
|
|
30
|
+
def critique_engine
|
|
31
|
+
one_of!(dig('llm', 'critique', 'engine') || 'model', CRITIQUE_ENGINES, 'llm.critique.engine')
|
|
32
|
+
end
|
|
33
|
+
|
|
34
|
+
def jev_fallback
|
|
35
|
+
one_of!(dig('llm', 'jev', 'fallback') || 'model', JEV_FALLBACKS, 'llm.jev.fallback')
|
|
36
|
+
end
|
|
37
|
+
|
|
38
|
+
# Jev decides in this run: the engine is jev and the run has a critique
|
|
39
|
+
# at all (--no-critique switches off every engine, as the critique:
|
|
40
|
+
# argument or as a CLI override of the config).
|
|
41
|
+
def jev_critique?(critique: true)
|
|
42
|
+
critique && !critique_disabled? && critique_engine == 'jev'
|
|
43
|
+
end
|
|
44
|
+
|
|
45
|
+
# The stages that need an LLM in a run. Jev with fallback: model still
|
|
46
|
+
# needs a ready LLM Critique; with fallback: fail it needs none.
|
|
47
|
+
def llm_stages(critique: true)
|
|
48
|
+
return ['generate'] unless critique && !critique_disabled?
|
|
49
|
+
return ['generate'] if jev_critique? && jev_fallback == 'fail'
|
|
50
|
+
|
|
51
|
+
STAGES
|
|
52
|
+
end
|
|
53
|
+
|
|
54
|
+
def critique_disabled?
|
|
55
|
+
@data['critique_disabled'] == true
|
|
56
|
+
end
|
|
57
|
+
|
|
58
|
+
def jev_shadow?
|
|
59
|
+
value = dig('llm', 'jev', 'shadow')
|
|
60
|
+
return false if value.nil?
|
|
61
|
+
return value if [true, false].include?(value)
|
|
62
|
+
|
|
63
|
+
raise ConfigError, "llm.jev.shadow must be true or false, got #{value.inspect}"
|
|
64
|
+
end
|
|
65
|
+
|
|
66
|
+
def jev_model
|
|
67
|
+
Aireview::Utils.presence(dig('llm', 'jev', 'model')) || DEFAULT_JEV_MODEL
|
|
68
|
+
end
|
|
69
|
+
|
|
70
|
+
def jev_api_key
|
|
71
|
+
@data['jev_api_key']
|
|
72
|
+
end
|
|
73
|
+
|
|
74
|
+
def jev_timeout
|
|
75
|
+
positive_integer!(dig('llm', 'jev', 'timeout') || DEFAULT_JEV_TIMEOUT, 'llm.jev.timeout')
|
|
76
|
+
end
|
|
77
|
+
|
|
78
|
+
def jev_thresholds
|
|
79
|
+
JEV_THRESHOLD_DEFAULTS.to_h do |name, default|
|
|
80
|
+
value = dig('llm', 'jev', name)
|
|
81
|
+
[name.to_sym, value.nil? ? default : probability!(value, "llm.jev.#{name}")]
|
|
82
|
+
end
|
|
83
|
+
end
|
|
84
|
+
|
|
85
|
+
# Everything besides the questions that changes a Jev decision; nil when
|
|
86
|
+
# Jev does not decide. The fallback is part of it: it decides what
|
|
87
|
+
# happens to the candidates Jev could not judge.
|
|
88
|
+
def jev_signature
|
|
89
|
+
return nil unless critique_engine == 'jev'
|
|
90
|
+
|
|
91
|
+
['jev', jev_model, *jev_thresholds.values, jev_fallback]
|
|
92
|
+
end
|
|
93
|
+
|
|
94
|
+
# Jev as the critic needs a key and a pinned version: the thresholds are
|
|
95
|
+
# tuned against one version, and falling back to the LLM silently would
|
|
96
|
+
# hide a broken setup.
|
|
97
|
+
def require_jev!(critique: true)
|
|
98
|
+
return unless jev_critique?(critique: critique)
|
|
99
|
+
raise ConfigError, 'llm.critique.engine is jev, but JEV_API_KEY is not set' if Aireview::Utils.blank?(jev_api_key)
|
|
100
|
+
return unless JEV_ALIASES.include?(jev_model)
|
|
101
|
+
|
|
102
|
+
raise ConfigError, "llm.jev.model #{jev_model.inspect} is an alias; with llm.critique.engine: jev " \
|
|
103
|
+
'pin a version such as jev-1.13.0'
|
|
104
|
+
end
|
|
105
|
+
|
|
106
|
+
def jev_warnings
|
|
107
|
+
return [] unless jev_shadow?
|
|
108
|
+
return ['llm.jev.shadow is ignored: Jev already decides as the critique engine'] if critique_engine == 'jev'
|
|
109
|
+
|
|
110
|
+
warnings = []
|
|
111
|
+
if Aireview::Utils.blank?(jev_api_key)
|
|
112
|
+
warnings << 'llm.jev.shadow is on, but JEV_API_KEY is not set: Jev is skipped'
|
|
113
|
+
end
|
|
114
|
+
if JEV_ALIASES.include?(jev_model)
|
|
115
|
+
warnings << "llm.jev.model #{jev_model.inspect} is an alias: its answers change when TypeSafe ships " \
|
|
116
|
+
'a release; pin a version such as jev-1.13.0'
|
|
117
|
+
end
|
|
118
|
+
warnings
|
|
119
|
+
end
|
|
120
|
+
|
|
121
|
+
private
|
|
122
|
+
|
|
123
|
+
def one_of!(value, allowed, name)
|
|
124
|
+
value = value.to_s
|
|
125
|
+
return value if allowed.include?(value)
|
|
126
|
+
|
|
127
|
+
raise ConfigError, "#{name} must be one of #{allowed.join(', ')}, got #{value.inspect}"
|
|
128
|
+
end
|
|
129
|
+
|
|
130
|
+
def probability!(value, name)
|
|
131
|
+
number = Float(value, exception: false) if value.is_a?(Numeric) || value.is_a?(String)
|
|
132
|
+
return number if number&.between?(0, 1)
|
|
133
|
+
|
|
134
|
+
raise ConfigError, "#{name} must be a number from 0 to 1, got #{value.inspect}"
|
|
135
|
+
end
|
|
136
|
+
end
|
|
137
|
+
end
|
|
@@ -65,7 +65,7 @@ module Aireview
|
|
|
65
65
|
|
|
66
66
|
# Configuration and plan warnings; the CLI and --dry-run print them.
|
|
67
67
|
def warnings
|
|
68
|
-
|
|
68
|
+
llm_stages.flat_map { |stage| stage_provider_warnings(stage) } + routing.warnings + jev_warnings
|
|
69
69
|
end
|
|
70
70
|
|
|
71
71
|
# The paths of the file layers, for --dry-run.
|