aireview 2.0.0 → 2.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +73 -0
- data/README.md +206 -12
- data/config/.aireview.yml.example +6 -0
- data/lib/aireview/cli.rb +44 -9
- data/lib/aireview/config.rb +26 -8
- data/lib/aireview/config_fallbacks.rb +43 -15
- data/lib/aireview/config_jev.rb +137 -0
- data/lib/aireview/config_layers.rb +1 -1
- data/lib/aireview/config_loader.rb +32 -6
- data/lib/aireview/context_builder.rb +13 -10
- data/lib/aireview/dry_run_prompts.rb +104 -0
- data/lib/aireview/dry_run_report.rb +40 -8
- data/lib/aireview/errors.rb +11 -0
- data/lib/aireview/jev_client.rb +129 -0
- data/lib/aireview/jev_critic.rb +387 -0
- data/lib/aireview/jev_shadow.rb +82 -0
- data/lib/aireview/jev_stage.rb +97 -0
- data/lib/aireview/llm_client.rb +53 -16
- data/lib/aireview/llm_router.rb +4 -4
- data/lib/aireview/model_candidate.rb +34 -5
- data/lib/aireview/model_checker.rb +47 -6
- data/lib/aireview/model_pool.rb +57 -28
- data/lib/aireview/output_schemas.rb +3 -3
- data/lib/aireview/prompts/jev_questions.yml +79 -0
- data/lib/aireview/review_marker.rb +7 -2
- data/lib/aireview/review_pipeline.rb +56 -56
- data/lib/aireview/review_renderer.rb +18 -2
- data/lib/aireview/reviewer.rb +10 -1
- data/lib/aireview/stage_chains.rb +6 -3
- data/lib/aireview/version.rb +1 -1
- metadata +14 -6
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: c37a072f5cb03c8a4e95944a3c6fc5386d1abe6c340c63ebee9715db5d44ec6a
|
|
4
|
+
data.tar.gz: b2e2ef104f8e4310dafc0dd5041a899c7ff06af5d2098cff25e0e38f010f9847
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: '09524e49ed6896e32e6e73d3fcc5f5e27d19eaa255273017acf699e8e6992f44a8a950e34d96f446066e58b31a6ee625188891e913ac3a4c0db27ac2b39e95a8'
|
|
7
|
+
data.tar.gz: 4a193a705c06609c07ceb2833211d51d124e1299378368feb50fc08e0fac48a6ff09f553b47406b0bd367e83ebfe7b55fd66d018f4781551725aa27234e95113
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,78 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 2.2.0
|
|
4
|
+
|
|
5
|
+
- Jev shadow mode (`llm.jev.shadow`, `LLM_JEV_SHADOW`, key `JEV_API_KEY`):
|
|
6
|
+
after the LLM Critique the same candidates go to Jev (TypeSafe), a fast
|
|
7
|
+
classifier, and its keep/reject/unverifiable decisions with the raw
|
|
8
|
+
probabilities are logged next to the Critique verdicts. The report and the
|
|
9
|
+
review key do not change; a Jev failure is a warning. The data is for
|
|
10
|
+
choosing between an LLM and Jev as the critic later.
|
|
11
|
+
- Critique engine (`llm.critique.engine`, `LLM_CRITIQUE_ENGINE`,
|
|
12
|
+
`--critique-engine`): `model` (the default, the LLM Critique as before) or
|
|
13
|
+
`jev`. With `jev` Jev decides keep/reject without refining the wording;
|
|
14
|
+
the candidates it cannot judge, and all of them when Jev fails, go to the
|
|
15
|
+
LLM Critique with `llm.jev.fallback: model` (the default), or are
|
|
16
|
+
rejected / fail the run with `fail`. Duplicates are dropped once over the
|
|
17
|
+
merged verdicts. The report says how Jev took part. The Jev thresholds
|
|
18
|
+
default to a provisional 0.5 until they are chosen from the shadow logs.
|
|
19
|
+
- The stages that need an LLM follow the engine: with Jev and
|
|
20
|
+
`fallback: fail` the critique model, its key, its share of the context
|
|
21
|
+
budget, the pool's critique policy and its `models check` probe are not
|
|
22
|
+
required. `models check` probes Jev as well (counted only for the engine).
|
|
23
|
+
- The review key includes the Jev version, thresholds, fallback and
|
|
24
|
+
question templates with `engine: jev`; with the default engine it does not
|
|
25
|
+
change, except for runs with `--no-critique`: the unused critique policy
|
|
26
|
+
of the pool (`critique.start`, `rank`, `allow_weaker`) no longer enters
|
|
27
|
+
their key and is no longer validated, so such a review is redone once.
|
|
28
|
+
- `--dry-run` prints the Jev request for the stub candidate.
|
|
29
|
+
- The gem now packages `lib/**/*.yml` (the Jev questions).
|
|
30
|
+
- A pool of models from several providers: OpenAI, Anthropic and OpenRouter
|
|
31
|
+
take their keys from `OPENAI_API_KEY(S)`, `ANTHROPIC_API_KEY(S)`,
|
|
32
|
+
`OPENROUTER_API_KEY(S)`, and a string like `anthropic/claude-opus-4.5` in
|
|
33
|
+
`LLM_MODELS` or a reserve names the provider (an OpenRouter model goes with
|
|
34
|
+
the prefix: `openrouter/qwen/qwen3-coder`).
|
|
35
|
+
- `api_base` for a model of the pool, a stage and a reserve: several servers
|
|
36
|
+
of your own in one pool. Such a model is named with its address, gets
|
|
37
|
+
`LLM_API_KEY` or no auth — never a provider's key — and needs no key to
|
|
38
|
+
start. Models without an address keep their names and review keys.
|
|
39
|
+
|
|
40
|
+
## 2.1.0
|
|
41
|
+
|
|
42
|
+
- An answer that is not valid JSON and that the provider cut off at the
|
|
43
|
+
output limit (`max_tokens`) or blocked (`content_filter`) sends the stage
|
|
44
|
+
to the next model without a repair request: the same model would cut the
|
|
45
|
+
repair off too. The log names the reason. A valid answer is taken
|
|
46
|
+
whatever the finish reason.
|
|
47
|
+
- Every LLM request logs the token counts the provider reported
|
|
48
|
+
(`tokens: input=… output=… thinking=… cache_read=… cache_write=…`), one
|
|
49
|
+
line per attempt; cache counts only when not zero. They are not summed:
|
|
50
|
+
Gemini already counts thinking into output, and input leaves out cached
|
|
51
|
+
tokens, so the prompt size is input + cache_read + cache_write.
|
|
52
|
+
- RubyLLM 2.0 (was 1.16). Output schemas are built with Schematist, which
|
|
53
|
+
RubyLLM now ships instead of `ruby_llm-schema`; `json` goes back to 2.x,
|
|
54
|
+
RubyLLM 2.0 requires `json < 3`.
|
|
55
|
+
- A structured answer is read as with RubyLLM 1.16: JSON that breaks the
|
|
56
|
+
schema sends the stage to the next model at once, only a text that is not
|
|
57
|
+
JSON gets a repair request.
|
|
58
|
+
- OpenAI stays on Chat Completions: RubyLLM 2.0 would switch it to the
|
|
59
|
+
Responses API, which a compatible server behind `LLM_API_BASE` usually
|
|
60
|
+
lacks.
|
|
61
|
+
- The critique schema goes to Chat Completions providers (Ollama,
|
|
62
|
+
OpenRouter, OpenAI) with `strict: false`: its `refinement` is optional,
|
|
63
|
+
and OpenAI strict mode rejects optional properties. Generate stays strict;
|
|
64
|
+
Gemini gets no strict flag, as before.
|
|
65
|
+
|
|
66
|
+
## 2.0.2
|
|
67
|
+
|
|
68
|
+
- When the merge request changes while the review runs, the skip warning
|
|
69
|
+
names the fields that changed (`sha`, `diff_refs`, `target_branch`,
|
|
70
|
+
`title`, `description`) instead of printing only the new head.
|
|
71
|
+
- In `review_mode: once` the skip message names both the mode and whether
|
|
72
|
+
the review is up to date (`the review is up to date`, `review inputs
|
|
73
|
+
changed`, `review freshness is unknown`); a stale review gets a hint:
|
|
74
|
+
retry the job in CI, `--force` outside of it.
|
|
75
|
+
|
|
3
76
|
## 2.0.0
|
|
4
77
|
|
|
5
78
|
Breaking: the review key of a shared pool includes the pool order and the
|
data/README.md
CHANGED
|
@@ -79,11 +79,18 @@ REVIEW_LANGUAGE=ru
|
|
|
79
79
|
REVIEW_MODE=update
|
|
80
80
|
```
|
|
81
81
|
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
82
|
+
The providers are `gemini`, `ollama`, `openai`, `anthropic` and
|
|
83
|
+
`openrouter`, and any OpenAI-compatible server of your own (vLLM, llama.cpp,
|
|
84
|
+
LM Studio) as `openai` with an `api_base`. The provider and the model of each
|
|
85
|
+
stage are set through `LLM_GENERATE_PROVIDER`, `LLM_GENERATE_MODEL`,
|
|
86
|
+
`LLM_CRITIQUE_PROVIDER` and `LLM_CRITIQUE_MODEL`. `LLM_PROVIDER` stays the
|
|
87
|
+
shared default when a stage has no provider of its own. Gemini and Ollama
|
|
88
|
+
are the ones the image defaults and the verified models are about; for the
|
|
89
|
+
others run `aireview models check` before relying on a model.
|
|
90
|
+
|
|
91
|
+
Each provider takes its keys from its own variables:
|
|
92
|
+
`GEMINI_API_KEY(S)`, `OPENAI_API_KEY(S)`, `ANTHROPIC_API_KEY(S)`,
|
|
93
|
+
`OPENROUTER_API_KEY(S)`; `LLM_API_KEY` is the shared fallback.
|
|
87
94
|
|
|
88
95
|
`REVIEW_LANGUAGE` (or `review_language` in `.aireview.yml`) sets the language
|
|
89
96
|
of the review: both the LLM answers and the headings of the rendered report.
|
|
@@ -175,6 +182,52 @@ the end of the run (see "Fallback models and keys"). The separate smoke test
|
|
|
175
182
|
of every model with both schemas (`aireview models check`) runs on an image
|
|
176
183
|
release and on a schedule, not on MRs.
|
|
177
184
|
|
|
185
|
+
### A pool of models from several providers
|
|
186
|
+
|
|
187
|
+
The critic is best the strongest model you have, and it does not have to be
|
|
188
|
+
Gemini: the shared pool (see "Shared model pool") takes models of any
|
|
189
|
+
provider, ordered by strength as you see it — the order is the only measure,
|
|
190
|
+
names and release dates are not compared.
|
|
191
|
+
|
|
192
|
+
```bash
|
|
193
|
+
LLM_MODELS=anthropic/claude-opus-4.5,openai/gpt-5,gemini/gemini-3.8-flash,ollama/qwen2.5-coder:14b
|
|
194
|
+
LLM_GENERATE_START=gemini-3.8-flash
|
|
195
|
+
ANTHROPIC_API_KEY=xxx
|
|
196
|
+
OPENAI_API_KEY=xxx
|
|
197
|
+
GEMINI_API_KEY=xxx
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
In a string the provider goes in front of the model; an OpenRouter model has
|
|
201
|
+
a slash of its own, so it is written with the prefix:
|
|
202
|
+
`openrouter/qwen/qwen3-coder`. Servers of your own, one or several, are
|
|
203
|
+
listed in `.aireview.yml` with their address:
|
|
204
|
+
|
|
205
|
+
```yaml
|
|
206
|
+
llm:
|
|
207
|
+
models:
|
|
208
|
+
- provider: anthropic
|
|
209
|
+
model: claude-opus-4.5
|
|
210
|
+
- provider: openai # an OpenAI-compatible server of your own
|
|
211
|
+
model: qwen3-coder
|
|
212
|
+
api_base: http://gpu1:8000/v1
|
|
213
|
+
- provider: ollama
|
|
214
|
+
model: qwen2.5-coder:14b
|
|
215
|
+
api_base: http://gpu2:11434/v1
|
|
216
|
+
generate:
|
|
217
|
+
start: qwen2.5-coder:14b
|
|
218
|
+
```
|
|
219
|
+
|
|
220
|
+
`api_base` also works for a stage (`llm.generate.api_base`) and for a reserve
|
|
221
|
+
in `fallbacks`. A model with its own `api_base` is named with the address
|
|
222
|
+
(`openai/qwen3-coder@http://gpu1:8000/v1`), so two servers with the same model are
|
|
223
|
+
two models for quarantine and the review key, and a start may name either.
|
|
224
|
+
Such a server never gets a provider's key: it is called with `LLM_API_KEY`,
|
|
225
|
+
or without auth when that is not set, and needs no key to start. Without an
|
|
226
|
+
`api_base` a model goes to its provider's API, or to `LLM_API_BASE` for
|
|
227
|
+
Gemini, OpenAI and OpenRouter as before.
|
|
228
|
+
|
|
229
|
+
Instead of an LLM the critic can be Jev (see "Critique engine").
|
|
230
|
+
|
|
178
231
|
### Local Ollama
|
|
179
232
|
|
|
180
233
|
Install Ollama following the
|
|
@@ -248,7 +301,7 @@ A local Ollama needs no API key. The providers can be swapped by changing
|
|
|
248
301
|
`LLM_GENERATE_PROVIDER`, `LLM_CRITIQUE_PROVIDER` and the corresponding models.
|
|
249
302
|
To run both stages locally, set `ollama` in both provider variables. The
|
|
250
303
|
address with `/v1` matches the
|
|
251
|
-
[Ollama configuration in RubyLLM](https://rubyllm.com/configuration
|
|
304
|
+
[Ollama configuration in RubyLLM](https://rubyllm.com/configuration-providers/).
|
|
252
305
|
`LLM_TIMEOUT` sets the timeout of every LLM request in seconds; for a slow
|
|
253
306
|
local model it can be raised, but not without limit: a hung request holds
|
|
254
307
|
the job for exactly that long while a fallback model sits idle. An
|
|
@@ -522,6 +575,124 @@ Critique receives the result of the check in the candidate's `note` field
|
|
|
522
575
|
and decides keep/reject with it in mind. With `--no-critique` the check works
|
|
523
576
|
the same way, its marks just go straight to the report.
|
|
524
577
|
|
|
578
|
+
### Critique engine: an LLM or Jev
|
|
579
|
+
|
|
580
|
+
`llm.critique.engine` picks who checks the candidates of the first pass:
|
|
581
|
+
|
|
582
|
+
| | `model` (default) | `jev` |
|
|
583
|
+
|---|---|---|
|
|
584
|
+
| What checks | the LLM Critique chain or pool (`LLM_CRITIQUE_MODEL`, `critique.rank`) | [Jev](https://docs.typesafe.ai), a classifier by TypeSafe |
|
|
585
|
+
| Context | the whole prompt: MR, Jira, diff | the same, cut to Jev's limits (see below) |
|
|
586
|
+
| Wording of findings | refined by the critic | the first pass's own |
|
|
587
|
+
| Time and cost | tens of seconds, provider quota | under a second, per input token |
|
|
588
|
+
| Depends on Gemini | when the critique model is Gemini | no, with `fallback: fail` |
|
|
589
|
+
| Non-English reviews | fine | not declared by TypeSafe, check first |
|
|
590
|
+
| Code leaves for | the LLM provider | TypeSafe as well |
|
|
591
|
+
|
|
592
|
+
The default stays `model`. A stronger critic than the generator is set with
|
|
593
|
+
the existing settings (`LLM_CRITIQUE_MODEL`, or the order of the pool); Jev is
|
|
594
|
+
meant for when speed or independence from the LLM quota matters more than
|
|
595
|
+
refined wording, and only after its thresholds are chosen from the shadow
|
|
596
|
+
logs (see below).
|
|
597
|
+
|
|
598
|
+
```yaml
|
|
599
|
+
llm:
|
|
600
|
+
critique:
|
|
601
|
+
engine: jev # LLM_CRITIQUE_ENGINE, --critique-engine jev
|
|
602
|
+
jev:
|
|
603
|
+
model: jev-1.13.0 # LLM_JEV_MODEL; a pinned version, an alias is an error here
|
|
604
|
+
fallback: model # LLM_JEV_FALLBACK: model or fail
|
|
605
|
+
keep_above: 0.5 # LLM_JEV_KEEP_ABOVE; enough_context, version_claim, duplicate in YAML
|
|
606
|
+
```
|
|
607
|
+
|
|
608
|
+
With `engine: jev` Jev answers the questions of `prompts/jev_questions.yml`
|
|
609
|
+
about every candidate and decides:
|
|
610
|
+
|
|
611
|
+
- claims that a version does not exist — reject;
|
|
612
|
+
- not enough in the request to judge it, or too large for a Jev request —
|
|
613
|
+
unverifiable: with `fallback: model` such candidates go to the LLM
|
|
614
|
+
Critique (only them), with `fallback: fail` they are rejected;
|
|
615
|
+
- otherwise keep when `real_issue` reaches `keep_above`;
|
|
616
|
+
- duplicates are dropped once, over what Jev and the LLM kept together: the
|
|
617
|
+
more severe finding stays, on a tie the one the first pass listed first.
|
|
618
|
+
|
|
619
|
+
When Jev fails (network, 429/529 after one retry, an invalid answer), the
|
|
620
|
+
LLM Critique takes over with `fallback: model`, and the run fails with
|
|
621
|
+
`fallback: fail`. The report says how Jev took part: checked by Jev without
|
|
622
|
+
refining the wording, partly by the LLM, or by the LLM because Jev was
|
|
623
|
+
unavailable. `engine: jev` without `JEV_API_KEY` or with an alias is a
|
|
624
|
+
configuration error at start; `--no-critique` switches off every engine.
|
|
625
|
+
|
|
626
|
+
What an LLM is still needed for follows the engine: with `fallback: fail`
|
|
627
|
+
there is no LLM Critique at all, so its model and key are not required
|
|
628
|
+
(generate on Ollama with Jev needs no Gemini key), the context budget counts
|
|
629
|
+
generate only, the pool's critique policy is not used and `models check`
|
|
630
|
+
probes the generate stage only. `models check` also sends Jev a probe
|
|
631
|
+
request: with `engine: jev` it counts, in shadow mode it is shown but never
|
|
632
|
+
fails the check. `--dry-run` shows the Jev request for the stub candidate
|
|
633
|
+
(`=== JEV STATE ===`, `=== JEV QUESTIONS ===`).
|
|
634
|
+
|
|
635
|
+
The review key follows the engine too: with `engine: jev` it includes the Jev
|
|
636
|
+
version, every threshold, the fallback and the question templates; the LLM
|
|
637
|
+
Critique model counts only while Jev can fall back to it. With the default
|
|
638
|
+
engine the key is what it was before Jev, whatever `llm.jev` says. Without
|
|
639
|
+
an LLM Critique (Jev with `fallback: fail`, or `--no-critique`) the critique
|
|
640
|
+
settings — model, limit, start, rank, allow_weaker — are neither validated
|
|
641
|
+
nor part of the key, while the generate pool stays in it.
|
|
642
|
+
|
|
643
|
+
### Jev shadow (experiment)
|
|
644
|
+
|
|
645
|
+
Jev is a classifier, not a text model: it answers yes/no and choice
|
|
646
|
+
questions about a given state with probabilities. Before it decides anything
|
|
647
|
+
as the engine, it can run in shadow mode next to the default engine: after
|
|
648
|
+
the LLM Critique the same candidates go to Jev, and its decisions are logged
|
|
649
|
+
next to the Critique verdicts. The report and the review key do not change;
|
|
650
|
+
with `engine: jev` the shadow is ignored.
|
|
651
|
+
|
|
652
|
+
```yaml
|
|
653
|
+
llm:
|
|
654
|
+
jev:
|
|
655
|
+
shadow: true # LLM_JEV_SHADOW=true
|
|
656
|
+
model: jev-1.13.0 # LLM_JEV_MODEL; a pinned version, not jev-latest
|
|
657
|
+
timeout: 10 # seconds per request
|
|
658
|
+
keep_above: 0.5 # provisional thresholds, only for the log line
|
|
659
|
+
enough_context: 0.5
|
|
660
|
+
version_claim: 0.5
|
|
661
|
+
duplicate: 0.5
|
|
662
|
+
```
|
|
663
|
+
|
|
664
|
+
The key is `JEV_API_KEY`. For every candidate Jev is asked whether it is a
|
|
665
|
+
real, well-supported problem (the rules of `prompts/critique.txt`, asked in
|
|
666
|
+
`prompts/jev_questions.yml`), whether the state holds enough to judge it,
|
|
667
|
+
whether it claims that a version does not exist, which other candidate it
|
|
668
|
+
duplicates, and how severe it is. The log gets one line per candidate and a
|
|
669
|
+
summary:
|
|
670
|
+
|
|
671
|
+
```
|
|
672
|
+
Jev shadow C1: keep (real issue; real_issue=0.91 enough_context=0.88 version_claim=0.02 duplicate_of=none/0.97 severity=major/0.74), critique: keep
|
|
673
|
+
Jev shadow: agrees with critique on 2 of 2 decided candidate(s); keep 1, reject 1, unverifiable 1 (model=jev-1.13.0, requests=1, 0.4s)
|
|
674
|
+
```
|
|
675
|
+
|
|
676
|
+
The readable lines round the numbers; a third line, `Jev shadow data: {…}`,
|
|
677
|
+
carries every decision with the probabilities exactly as Jev returned them,
|
|
678
|
+
so thresholds can be chosen from the logs afterwards. The version that
|
|
679
|
+
answered is logged too, and a pinned version answering as another one is a
|
|
680
|
+
warning. A Jev failure (network, 429/529 after
|
|
681
|
+
one retry, an invalid answer, no key) is a warning and never affects the
|
|
682
|
+
review. Jev goes out through `LLM_HTTP_PROXY` like the LLM providers.
|
|
683
|
+
|
|
684
|
+
What leaves for TypeSafe: the MR title and description, the Jira section,
|
|
685
|
+
`review_instructions`, the diff shown to the models and the candidates, after
|
|
686
|
+
the same secret scrubbing as the LLM prompts. Jev limits a request to 32k
|
|
687
|
+
tokens for the state plus the longest question and 64k for the state plus all
|
|
688
|
+
questions; the state is cut to what the questions leave (the diff first, then
|
|
689
|
+
the MR and Jira sections, never the candidates and the hunks they point at),
|
|
690
|
+
and when the candidates do not fit together, each goes in a request of its
|
|
691
|
+
own. The size is estimated from characters; if Jev still rejects a request
|
|
692
|
+
as too large (a 422 that explicitly says it exceeds a limit), its candidates
|
|
693
|
+
are asked about one by one, and one that is too large alone stays
|
|
694
|
+
unverifiable. Any other 422 is not retried.
|
|
695
|
+
|
|
525
696
|
## Usage
|
|
526
697
|
|
|
527
698
|
```bash
|
|
@@ -698,12 +869,35 @@ The template is written for shell-executor runners (the image runs through
|
|
|
698
869
|
`docker run` with secrets passed by name) and carries `[skip review]`,
|
|
699
870
|
`resource_group`, `allow_failure` and `REVIEW_MODE=once`. Models come from
|
|
700
871
|
the image defaults, `.aireview.yml` is mounted into the container only when
|
|
701
|
-
the project has one. The job runs in the `.post` stage
|
|
702
|
-
pipeline, so a project does not declare
|
|
703
|
-
|
|
704
|
-
|
|
705
|
-
|
|
706
|
-
|
|
872
|
+
the project has one. The job runs in the `.post` stage: it exists in every
|
|
873
|
+
pipeline, so a project does not declare it.
|
|
874
|
+
|
|
875
|
+
GitLab does not start a pipeline in which, after `rules` and `only` are
|
|
876
|
+
applied, only `.pre` and `.post` jobs remain. That happens when a project has
|
|
877
|
+
no other merge request jobs (a build that runs on tags only does not count).
|
|
878
|
+
In that case add a `review` stage to the project's existing `stages`, without
|
|
879
|
+
replacing them, and move the job there:
|
|
880
|
+
|
|
881
|
+
```yaml
|
|
882
|
+
stages:
|
|
883
|
+
- review # added to the project's stages
|
|
884
|
+
- build
|
|
885
|
+
|
|
886
|
+
aireview:
|
|
887
|
+
stage: review
|
|
888
|
+
```
|
|
889
|
+
|
|
890
|
+
To keep the review from waiting for other stages to finish, add `needs: []`
|
|
891
|
+
to the `aireview` job.
|
|
892
|
+
|
|
893
|
+
The template sets `REVIEW_MODE` as a job variable, and an environment variable
|
|
894
|
+
beats `.aireview.yml`: `review_mode` in the project file has no effect once
|
|
895
|
+
the template is included. Change the mode with a `REVIEW_MODE` CI/CD variable
|
|
896
|
+
of the project or with `aireview: {variables: {REVIEW_MODE: update}}`.
|
|
897
|
+
|
|
898
|
+
Every variable the config reads (`Config.env_names`: models, providers,
|
|
899
|
+
reserves, temperatures, limits, `OLLAMA_API_BASE` and so on) is passed into
|
|
900
|
+
the container by name, so a project can override anything through its CI/CD
|
|
707
901
|
variables — for instance, swap an unavailable model with `LLM_CRITIQUE_MODEL`
|
|
708
902
|
without waiting for an image release.
|
|
709
903
|
|
|
@@ -36,3 +36,9 @@ llm:
|
|
|
36
36
|
model: gemini-3.8-flash
|
|
37
37
|
fallbacks:
|
|
38
38
|
- gemini-3.7-flash
|
|
39
|
+
# Jev as the critique engine (critique: {engine: jev}) or next to it, log only
|
|
40
|
+
# (see README, "Critique engine"); the key is JEV_API_KEY
|
|
41
|
+
# jev:
|
|
42
|
+
# shadow: true
|
|
43
|
+
# model: jev-1.13.0
|
|
44
|
+
# fallback: model
|
data/lib/aireview/cli.rb
CHANGED
|
@@ -99,9 +99,11 @@ module Aireview
|
|
|
99
99
|
critique_model: options[:critique_model],
|
|
100
100
|
generate_temperature: options[:generate_temperature],
|
|
101
101
|
critique_temperature: options[:critique_temperature],
|
|
102
|
-
|
|
102
|
+
critique_engine: options[:critique_engine],
|
|
103
|
+
no_fallbacks: options[:no_fallbacks] == true,
|
|
104
|
+
no_critique: options[:no_critique] == true
|
|
103
105
|
)
|
|
104
|
-
config.require_llm_configuration!
|
|
106
|
+
config.require_llm_configuration!(critique: !options[:no_critique])
|
|
105
107
|
config.warnings.each { |warning| @logger.warn(warning) }
|
|
106
108
|
config
|
|
107
109
|
end
|
|
@@ -212,18 +214,41 @@ module Aireview
|
|
|
212
214
|
up_to_date = existing[:key] == key
|
|
213
215
|
return false unless up_to_date || (mode == 'once' && !retried_ci_job?(gitlab_client))
|
|
214
216
|
|
|
215
|
-
|
|
216
|
-
@out.puts("Review skipped: #{reason}")
|
|
217
|
+
@out.puts("Review skipped: #{skip_reason(existing, up_to_date: up_to_date, mode: mode)}")
|
|
217
218
|
true
|
|
218
219
|
end
|
|
219
220
|
|
|
221
|
+
# In once mode the review is not repeated on new pushes, even when the MR
|
|
222
|
+
# has changed. So besides the mode the message says whether the review is
|
|
223
|
+
# up to date: if it is not, the Retry button of the GitLab job updates it.
|
|
224
|
+
def skip_reason(existing, up_to_date:, mode:)
|
|
225
|
+
return 'existing review is up to date' unless mode == 'once'
|
|
226
|
+
|
|
227
|
+
state = if up_to_date then 'the review is up to date'
|
|
228
|
+
elsif existing[:key].nil? then "review freshness is unknown: #{update_hint}"
|
|
229
|
+
else "review inputs changed: #{update_hint}"
|
|
230
|
+
end
|
|
231
|
+
"merge request already reviewed (review_mode=once), #{state}"
|
|
232
|
+
end
|
|
233
|
+
|
|
234
|
+
def update_hint
|
|
235
|
+
ci_job_context ? 'retry the job to update' : 'use --force to review again'
|
|
236
|
+
end
|
|
237
|
+
|
|
220
238
|
def retried_ci_job?(gitlab_client)
|
|
221
|
-
project_id, job_id =
|
|
222
|
-
return false
|
|
239
|
+
project_id, job_id = ci_job_context
|
|
240
|
+
return false unless project_id
|
|
223
241
|
|
|
224
242
|
gitlab_client.retried_job?(project_id, job_id)
|
|
225
243
|
end
|
|
226
244
|
|
|
245
|
+
def ci_job_context
|
|
246
|
+
project_id, job_id = @env.values_at('CI_PROJECT_ID', 'CI_JOB_ID')
|
|
247
|
+
return if Aireview::Utils.blank?(project_id) || Aireview::Utils.blank?(job_id)
|
|
248
|
+
|
|
249
|
+
[project_id, job_id]
|
|
250
|
+
end
|
|
251
|
+
|
|
227
252
|
def publish_review(review, context, publication)
|
|
228
253
|
return if merge_request_moved?(context)
|
|
229
254
|
|
|
@@ -245,10 +270,13 @@ module Aireview
|
|
|
245
270
|
context[:parser_result].project_id,
|
|
246
271
|
context[:parser_result].iid
|
|
247
272
|
)
|
|
248
|
-
|
|
273
|
+
before = ReviewMarker.state(context[:merge_request])
|
|
274
|
+
after = ReviewMarker.state(current)
|
|
275
|
+
changed = after.keys.reject { |field| after[field] == before[field] }
|
|
276
|
+
return false if changed.empty?
|
|
249
277
|
|
|
250
|
-
@logger.warn("Merge request
|
|
251
|
-
'
|
|
278
|
+
@logger.warn("Merge request changed while review was running (#{changed.join(', ')}); " \
|
|
279
|
+
'skipping publication')
|
|
252
280
|
true
|
|
253
281
|
end
|
|
254
282
|
|
|
@@ -309,6 +337,11 @@ module Aireview
|
|
|
309
337
|
options[:critique_temperature] = value
|
|
310
338
|
end
|
|
311
339
|
|
|
340
|
+
parser.on('--critique-engine ENGINE', Aireview::Config::CRITIQUE_ENGINES,
|
|
341
|
+
'Who checks the candidates: model (an LLM) or jev') do |value|
|
|
342
|
+
options[:critique_engine] = value
|
|
343
|
+
end
|
|
344
|
+
|
|
312
345
|
parser.on('--no-critique', 'Skip critique pass and render Generate candidates directly') do
|
|
313
346
|
options[:no_critique] = true
|
|
314
347
|
end
|
|
@@ -381,6 +414,8 @@ module Aireview
|
|
|
381
414
|
Override Generate pass temperature
|
|
382
415
|
--critique-temperature VALUE
|
|
383
416
|
Override Critique pass temperature
|
|
417
|
+
--critique-engine ENGINE
|
|
418
|
+
Who checks the candidates: model (an LLM, default) or jev
|
|
384
419
|
--config PATH Path to .aireview.yml
|
|
385
420
|
--review-mode MODE
|
|
386
421
|
How to treat an existing review: update (default) or once
|
data/lib/aireview/config.rb
CHANGED
|
@@ -6,6 +6,7 @@ require_relative 'errors'
|
|
|
6
6
|
require_relative 'utils'
|
|
7
7
|
require_relative 'config_limits'
|
|
8
8
|
require_relative 'config_fallbacks'
|
|
9
|
+
require_relative 'config_jev'
|
|
9
10
|
require_relative 'config_layers'
|
|
10
11
|
require_relative 'config_loader'
|
|
11
12
|
|
|
@@ -16,6 +17,7 @@ module Aireview
|
|
|
16
17
|
class Config
|
|
17
18
|
include ConfigLimits
|
|
18
19
|
include ConfigFallbacks
|
|
20
|
+
include ConfigJev
|
|
19
21
|
include ConfigLayers
|
|
20
22
|
|
|
21
23
|
DEFAULT_SECRET_FILES = [
|
|
@@ -76,20 +78,27 @@ module Aireview
|
|
|
76
78
|
|
|
77
79
|
# CLI overrides change only the primary model of a stage, the reserves
|
|
78
80
|
# from the config stay; no_fallbacks leaves one model and one key.
|
|
79
|
-
|
|
81
|
+
# no_critique — --no-critique: the run has no Critique, so its settings
|
|
82
|
+
# are neither routed nor validated (see llm_stages).
|
|
83
|
+
# Seven named flags of the CLI read better than a struct for its own sake.
|
|
84
|
+
def with_overrides( # rubocop:disable Metrics/ParameterLists
|
|
80
85
|
generate_model: nil,
|
|
81
86
|
critique_model: nil,
|
|
82
87
|
generate_temperature: nil,
|
|
83
88
|
critique_temperature: nil,
|
|
84
|
-
|
|
89
|
+
critique_engine: nil,
|
|
90
|
+
no_fallbacks: false,
|
|
91
|
+
no_critique: false
|
|
85
92
|
)
|
|
86
93
|
llm_config = {
|
|
87
94
|
'generate' => stage_overrides(model: generate_model, temperature: generate_temperature),
|
|
88
95
|
'critique' => stage_overrides(model: critique_model, temperature: critique_temperature)
|
|
96
|
+
.merge({'engine' => critique_engine}.compact)
|
|
89
97
|
}.reject { |_, overrides| overrides.empty? }
|
|
90
98
|
overrides = {}
|
|
91
99
|
overrides['llm'] = llm_config unless llm_config.empty?
|
|
92
100
|
overrides['fallbacks_disabled'] = true if no_fallbacks
|
|
101
|
+
overrides['critique_disabled'] = true if no_critique
|
|
93
102
|
return self if overrides.empty?
|
|
94
103
|
|
|
95
104
|
self.class.new(
|
|
@@ -137,8 +146,9 @@ module Aireview
|
|
|
137
146
|
routing.primary('generate').model
|
|
138
147
|
end
|
|
139
148
|
|
|
149
|
+
# nil when no LLM Critique can run (see llm_stages).
|
|
140
150
|
def critique_model
|
|
141
|
-
routing.primary('critique').model
|
|
151
|
+
routing.stage?('critique') ? routing.primary('critique').model : nil
|
|
142
152
|
end
|
|
143
153
|
|
|
144
154
|
def generate_provider
|
|
@@ -147,18 +157,26 @@ module Aireview
|
|
|
147
157
|
|
|
148
158
|
# Everything besides the prompt that affects the review result goes into
|
|
149
159
|
# the note key (see ReviewMarker): provider, model and temperature of the
|
|
150
|
-
# stages, the shared pool with its critique
|
|
151
|
-
# chains do not change the result.
|
|
160
|
+
# stages, Jev as the critique engine, the shared pool with its critique
|
|
161
|
+
# policy. Reserves of per-stage chains do not change the result.
|
|
152
162
|
def result_signature
|
|
153
163
|
{
|
|
154
|
-
'generate' =>
|
|
155
|
-
'critique' =>
|
|
164
|
+
'generate' => stage_signature('generate', generate_temperature),
|
|
165
|
+
'critique' => (routing.stage?('critique') ? stage_signature('critique', critique_temperature) : nil),
|
|
166
|
+
'critique_engine' => jev_signature,
|
|
156
167
|
'pool' => routing.signature
|
|
157
168
|
}
|
|
158
169
|
end
|
|
159
170
|
|
|
160
171
|
def critique_provider
|
|
161
|
-
routing.primary('critique').provider
|
|
172
|
+
routing.stage?('critique') ? routing.primary('critique').provider : nil
|
|
173
|
+
end
|
|
174
|
+
|
|
175
|
+
# The primary of a stage on a server of your own goes into the review key
|
|
176
|
+
# with its address; without one the signature is what it was.
|
|
177
|
+
def stage_signature(stage, temperature)
|
|
178
|
+
primary = routing.primary(stage)
|
|
179
|
+
[primary.provider, primary.model, temperature, *primary.api_base]
|
|
162
180
|
end
|
|
163
181
|
|
|
164
182
|
def generate_temperature
|
|
@@ -27,6 +27,19 @@ module Aireview
|
|
|
27
27
|
routing.chain(stage)
|
|
28
28
|
end
|
|
29
29
|
|
|
30
|
+
# The keys a model is called with. A server of your own (a model with
|
|
31
|
+
# api_base) never gets the provider's key — OPENAI_API_KEY must not leave
|
|
32
|
+
# for a self-hosted server — but LLM_API_KEY, or a placeholder for a
|
|
33
|
+
# server without auth, since the client wants some key.
|
|
34
|
+
SELF_HOSTED_PLACEHOLDER_KEY = 'no-key'
|
|
35
|
+
|
|
36
|
+
def candidate_api_keys(candidate)
|
|
37
|
+
return provider_api_keys(candidate.provider) unless candidate.api_base
|
|
38
|
+
return [nil] if KEYLESS_PROVIDERS.include?(candidate.provider.to_s)
|
|
39
|
+
|
|
40
|
+
[Aireview::Utils.presence(llm_api_key) || SELF_HOSTED_PLACEHOLDER_KEY]
|
|
41
|
+
end
|
|
42
|
+
|
|
30
43
|
# Keys in order of preference; a single nil for a keyless provider, so
|
|
31
44
|
# that walking the chain does not depend on the provider.
|
|
32
45
|
def provider_api_keys(provider)
|
|
@@ -47,9 +60,9 @@ module Aireview
|
|
|
47
60
|
end
|
|
48
61
|
|
|
49
62
|
# Only the number of keys per provider, for --dry-run; the values never
|
|
50
|
-
# leave.
|
|
63
|
+
# leave. Servers of your own are not counted: they take LLM_API_KEY.
|
|
51
64
|
def api_key_counts(stages)
|
|
52
|
-
providers = stages.flat_map { |stage| stage_chain(stage).map(&:provider) }.uniq
|
|
65
|
+
providers = stages.flat_map { |stage| stage_chain(stage).reject(&:api_base).map(&:provider) }.uniq
|
|
53
66
|
providers.reject { |provider| KEYLESS_PROVIDERS.include?(provider.to_s) }
|
|
54
67
|
.to_h { |provider| [provider, provider_api_keys(provider).size] }
|
|
55
68
|
end
|
|
@@ -67,39 +80,53 @@ module Aireview
|
|
|
67
80
|
'llm.overloaded_quarantine')
|
|
68
81
|
end
|
|
69
82
|
|
|
70
|
-
|
|
83
|
+
# critique — the run has a critique (false with --no-critique); only the
|
|
84
|
+
# stages that need an LLM in such a run are required (see llm_stages).
|
|
85
|
+
def require_models!(critique: true)
|
|
86
|
+
stages = llm_stages(critique: critique)
|
|
71
87
|
missing = []
|
|
72
88
|
missing << 'llm.generate.model (or LLM_GENERATE_MODEL)' if Aireview::Utils.blank?(generate_model)
|
|
73
|
-
|
|
89
|
+
if stages.include?('critique') && Aireview::Utils.blank?(critique_model)
|
|
90
|
+
missing << 'llm.critique.model (or LLM_CRITIQUE_MODEL)'
|
|
91
|
+
end
|
|
74
92
|
raise ConfigError, "LLM models are required: #{missing.join(', ')}" unless missing.empty?
|
|
75
93
|
end
|
|
76
94
|
|
|
77
95
|
# The plan is built whole (pool and critique policy validated) before the
|
|
78
|
-
# first request, not in Critique after a paid-for Generate.
|
|
79
|
-
|
|
80
|
-
|
|
96
|
+
# first request, not in Critique after a paid-for Generate. Keys are
|
|
97
|
+
# required for the stages that need an LLM: generate on Ollama with Jev
|
|
98
|
+
# as the critic and no fallback needs no Gemini key.
|
|
99
|
+
def require_llm_configuration!(critique: true)
|
|
100
|
+
require_models!(critique: critique)
|
|
101
|
+
require_jev!(critique: critique)
|
|
81
102
|
routing
|
|
82
103
|
|
|
83
|
-
missing_keys =
|
|
84
|
-
providers = stage_chain(stage).map(&:provider).uniq
|
|
85
|
-
providers.
|
|
104
|
+
missing_keys = llm_stages(critique: critique).flat_map do |stage|
|
|
105
|
+
providers = stage_chain(stage).reject(&:api_base).map(&:provider).uniq
|
|
106
|
+
providers.reject { |provider| provider_keys_present?(provider) }
|
|
107
|
+
.map { |provider| "#{stage}: API key is required for provider #{provider.inspect}" }
|
|
86
108
|
end
|
|
87
109
|
raise ConfigError, missing_keys.join(', ') unless missing_keys.empty?
|
|
88
110
|
end
|
|
89
111
|
|
|
90
112
|
private
|
|
91
113
|
|
|
114
|
+
# Only the stages that go to an LLM are routed (llm_stages): without an
|
|
115
|
+
# LLM Critique (Jev with fallback: fail, --no-critique) its model, limit,
|
|
116
|
+
# start, rank and allow_weaker are not used, so they must neither fail
|
|
117
|
+
# the run nor change the review key, nor push the pool out of it.
|
|
92
118
|
def build_routing
|
|
93
|
-
|
|
119
|
+
stages = llm_stages
|
|
120
|
+
settings = stages.to_h { |stage| [stage, stage_settings(stage)] }
|
|
94
121
|
models = Array(dig('llm', 'models'))
|
|
95
122
|
return StageChains.build(settings, only_primary: fallbacks_disabled?) if models.empty?
|
|
96
123
|
|
|
97
124
|
own = settings.select { |_, stage_settings| Aireview::Utils.present?(stage_settings[:model]) }
|
|
98
125
|
ModelPool.new(
|
|
99
|
-
items: models, provider: llm_provider,
|
|
100
|
-
limits:
|
|
101
|
-
starts:
|
|
102
|
-
inherited_starts:
|
|
126
|
+
items: models, provider: llm_provider, stages: stages,
|
|
127
|
+
limits: stages.to_h { |stage| [stage, max_prompt_chars(stage)] },
|
|
128
|
+
starts: stages.to_h { |stage| [stage, dig('llm', stage, 'start')] },
|
|
129
|
+
inherited_starts: stages.select { |stage| start_inherited?(stage) },
|
|
103
130
|
rank: dig('llm', 'critique', 'rank'), allow_weaker: dig('llm', 'critique', 'allow_weaker'),
|
|
104
131
|
own_chains: StageChains.build(own, only_primary: fallbacks_disabled?), only_primary: fallbacks_disabled?
|
|
105
132
|
)
|
|
@@ -109,6 +136,7 @@ module Aireview
|
|
|
109
136
|
{
|
|
110
137
|
provider: stage_setting(stage, 'provider') || llm_provider,
|
|
111
138
|
model: dig('llm', stage, 'model'),
|
|
139
|
+
api_base: dig('llm', stage, 'api_base'),
|
|
112
140
|
fallbacks: dig('llm', stage, 'fallbacks'),
|
|
113
141
|
max_prompt_chars: max_prompt_chars(stage)
|
|
114
142
|
}
|