aireview 2.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 3f326ae1b11f59f7c447e8f128ab94845ef1543b0b377753f053fd4f69e03bad
4
- data.tar.gz: c5908173ca8020a0563b6d4ee1fd4d8ce4ebe609fda7ded26d5b7eb40c84dda1
3
+ metadata.gz: 1e262ff91d3ae033f68f0a1a940af115dda8cf924e6715d7d494dfbe30e2e907
4
+ data.tar.gz: ea5bfca46774fe542338e67b198dc66e550b33404363646da055e8648c67f6b1
5
5
  SHA512:
6
- metadata.gz: 2de9b7dd7e6caaa440287339d03eed843dfdb73010ae8cf6a01dea4e9348e3b072fc4fb61a55bf30eb4ed324e857ff883e35e31b711bc542cb8d1495645041bc
7
- data.tar.gz: 5f0867f71415c019c03d99bd2215fc7f7ef7cadf98ab8b0db91e6eb98af54cfae3787b848c7611fc371cfa19cc94acd55e170047f094d8543f5c77f258e3313d
6
+ metadata.gz: 7fee8f2bbade71a38d72e66be714c748d38afc91643dde35329fbc64a357c88af6020bda5f1a89b9762922dc4ff4a413cdf777fab03b16dc9a09e073953ad400
7
+ data.tar.gz: 869d8f2e89afd98feab341028556c7a361ba46f9afb070cb728298a2b20aafc8df1aba5525f9ab9299c4b35169dcc3323cef72771764f6120ccfba0f34c34d69
data/CHANGELOG.md CHANGED
@@ -1,5 +1,41 @@
1
1
  # Changelog
2
2
 
3
+ ## 2.1.0
4
+
5
+ - An answer that is not valid JSON and that the provider cut off at the
6
+ output limit (`max_tokens`) or blocked (`content_filter`) sends the stage
7
+ to the next model without a repair request: the same model would cut the
8
+ repair off too. The log names the reason. A valid answer is taken
9
+ whatever the finish reason.
10
+ - Every LLM request logs the token counts the provider reported
11
+ (`tokens: input=… output=… thinking=… cache_read=… cache_write=…`), one
12
+ line per attempt; cache counts only when not zero. They are not summed:
13
+ Gemini already counts thinking into output, and input leaves out cached
14
+ tokens, so the prompt size is input + cache_read + cache_write.
15
+ - RubyLLM 2.0 (was 1.16). Output schemas are built with Schematist, which
16
+ RubyLLM now ships instead of `ruby_llm-schema`; `json` goes back to 2.x,
17
+ RubyLLM 2.0 requires `json < 3`.
18
+ - A structured answer is read as with RubyLLM 1.16: JSON that breaks the
19
+ schema sends the stage to the next model at once, only a text that is not
20
+ JSON gets a repair request.
21
+ - OpenAI stays on Chat Completions: RubyLLM 2.0 would switch it to the
22
+ Responses API, which a compatible server behind `LLM_API_BASE` usually
23
+ lacks.
24
+ - The critique schema goes to Chat Completions providers (Ollama,
25
+ OpenRouter, OpenAI) with `strict: false`: its `refinement` is optional,
26
+ and OpenAI strict mode rejects optional properties. Generate stays strict;
27
+ Gemini gets no strict flag, as before.
28
+
29
+ ## 2.0.2
30
+
31
+ - When the merge request changes while the review runs, the skip warning
32
+ names the fields that changed (`sha`, `diff_refs`, `target_branch`,
33
+ `title`, `description`) instead of printing only the new head.
34
+ - In `review_mode: once` the skip message names both the mode and whether
35
+ the review is up to date (`the review is up to date`, `review inputs
36
+ changed`, `review freshness is unknown`); a stale review gets a hint:
37
+ retry the job in CI, `--force` outside of it.
38
+
3
39
  ## 2.0.0
4
40
 
5
41
  Breaking: the review key of a shared pool includes the pool order and the
data/README.md CHANGED
@@ -248,7 +248,7 @@ A local Ollama needs no API key. The providers can be swapped by changing
248
248
  `LLM_GENERATE_PROVIDER`, `LLM_CRITIQUE_PROVIDER` and the corresponding models.
249
249
  To run both stages locally, set `ollama` in both provider variables. The
250
250
  address with `/v1` matches the
251
- [Ollama configuration in RubyLLM](https://rubyllm.com/configuration/#provider-configuration).
251
+ [Ollama configuration in RubyLLM](https://rubyllm.com/configuration-providers/).
252
252
  `LLM_TIMEOUT` sets the timeout of every LLM request in seconds; for a slow
253
253
  local model it can be raised, but not without limit: a hung request holds
254
254
  the job for exactly that long while a fallback model sits idle. An
@@ -698,12 +698,35 @@ The template is written for shell-executor runners (the image runs through
698
698
  `docker run` with secrets passed by name) and carries `[skip review]`,
699
699
  `resource_group`, `allow_failure` and `REVIEW_MODE=once`. Models come from
700
700
  the image defaults, `.aireview.yml` is mounted into the container only when
701
- the project has one. The job runs in the `.post` stage — it exists in every
702
- pipeline, so a project does not declare `stages`; to move the review to
703
- another stage, add `aireview: {stage: review}` to the project file. Every
704
- variable the config reads (`Config.env_names`: models, providers, reserves,
705
- temperatures, limits, `OLLAMA_API_BASE` and so on) is passed into the
706
- container by name, so a project can override anything through its CI/CD
701
+ the project has one. The job runs in the `.post` stage: it exists in every
702
+ pipeline, so a project does not declare it.
703
+
704
+ GitLab does not start a pipeline in which, after `rules` and `only` are
705
+ applied, only `.pre` and `.post` jobs remain. That happens when a project has
706
+ no other merge request jobs (a build that runs on tags only does not count).
707
+ In that case add a `review` stage to the project's existing `stages`, without
708
+ replacing them, and move the job there:
709
+
710
+ ```yaml
711
+ stages:
712
+ - review # added to the project's stages
713
+ - build
714
+
715
+ aireview:
716
+ stage: review
717
+ ```
718
+
719
+ To keep the review from waiting for other stages to finish, add `needs: []`
720
+ to the `aireview` job.
721
+
722
+ The template sets `REVIEW_MODE` as a job variable, and an environment variable
723
+ beats `.aireview.yml`: `review_mode` in the project file has no effect once
724
+ the template is included. Change the mode with a `REVIEW_MODE` CI/CD variable
725
+ of the project or with `aireview: {variables: {REVIEW_MODE: update}}`.
726
+
727
+ Every variable the config reads (`Config.env_names`: models, providers,
728
+ reserves, temperatures, limits, `OLLAMA_API_BASE` and so on) is passed into
729
+ the container by name, so a project can override anything through its CI/CD
707
730
  variables — for instance, swap an unavailable model with `LLM_CRITIQUE_MODEL`
708
731
  without waiting for an image release.
709
732
 
data/lib/aireview/cli.rb CHANGED
@@ -212,18 +212,41 @@ module Aireview
212
212
  up_to_date = existing[:key] == key
213
213
  return false unless up_to_date || (mode == 'once' && !retried_ci_job?(gitlab_client))
214
214
 
215
- reason = up_to_date ? 'existing review is up to date' : 'merge request already reviewed (review_mode=once)'
216
- @out.puts("Review skipped: #{reason}")
215
+ @out.puts("Review skipped: #{skip_reason(existing, up_to_date: up_to_date, mode: mode)}")
217
216
  true
218
217
  end
219
218
 
219
+ # In once mode the review is not repeated on new pushes, even when the MR
220
+ # has changed. So besides the mode the message says whether the review is
221
+ # up to date: if it is not, the Retry button of the GitLab job updates it.
222
+ def skip_reason(existing, up_to_date:, mode:)
223
+ return 'existing review is up to date' unless mode == 'once'
224
+
225
+ state = if up_to_date then 'the review is up to date'
226
+ elsif existing[:key].nil? then "review freshness is unknown: #{update_hint}"
227
+ else "review inputs changed: #{update_hint}"
228
+ end
229
+ "merge request already reviewed (review_mode=once), #{state}"
230
+ end
231
+
232
+ def update_hint
233
+ ci_job_context ? 'retry the job to update' : 'use --force to review again'
234
+ end
235
+
220
236
  def retried_ci_job?(gitlab_client)
221
- project_id, job_id = @env.values_at('CI_PROJECT_ID', 'CI_JOB_ID')
222
- return false if Aireview::Utils.blank?(project_id) || Aireview::Utils.blank?(job_id)
237
+ project_id, job_id = ci_job_context
238
+ return false unless project_id
223
239
 
224
240
  gitlab_client.retried_job?(project_id, job_id)
225
241
  end
226
242
 
243
+ def ci_job_context
244
+ project_id, job_id = @env.values_at('CI_PROJECT_ID', 'CI_JOB_ID')
245
+ return if Aireview::Utils.blank?(project_id) || Aireview::Utils.blank?(job_id)
246
+
247
+ [project_id, job_id]
248
+ end
249
+
227
250
  def publish_review(review, context, publication)
228
251
  return if merge_request_moved?(context)
229
252
 
@@ -245,10 +268,13 @@ module Aireview
245
268
  context[:parser_result].project_id,
246
269
  context[:parser_result].iid
247
270
  )
248
- return false if ReviewMarker.state(current) == ReviewMarker.state(context[:merge_request])
271
+ before = ReviewMarker.state(context[:merge_request])
272
+ after = ReviewMarker.state(current)
273
+ changed = after.keys.reject { |field| after[field] == before[field] }
274
+ return false if changed.empty?
249
275
 
250
- @logger.warn("Merge request moved to #{current['sha']} (#{current['target_branch']}) " \
251
- 'while review was running; skipping publication')
276
+ @logger.warn("Merge request changed while review was running (#{changed.join(', ')}); " \
277
+ 'skipping publication')
252
278
  true
253
279
  end
254
280
 
@@ -1,4 +1,5 @@
1
1
  # frozen_string_literal: true
2
+ require 'json'
2
3
  require 'timeout'
3
4
  require_relative 'errors'
4
5
  require_relative 'utils'
@@ -22,8 +23,8 @@ module Aireview
22
23
  @contexts = {}
23
24
  end
24
25
 
25
- # Returns the RubyLLM answer (content is text or a structure by the
26
- # schema). A request error is re-raised as is — LlmFailure classifies it.
26
+ # Returns the RubyLLM answer; read it with LlmClient.content. A request
27
+ # error is re-raised as is — LlmFailure classifies it.
27
28
  def request(prompt, candidate:, key:, timeout:, key_index: 0)
28
29
  load_ruby_llm
29
30
  stage = prompt.stage.to_s
@@ -36,15 +37,44 @@ module Aireview
36
37
  .with_schema(prompt.schema)
37
38
  chat.with_instructions(prompt.system)
38
39
  response = Timeout.timeout(timeout) { chat.ask(prompt.user) }
39
- @logger.info("LLM #{stage} request completed (model=#{model})")
40
+ @logger.info("LLM #{stage} request completed (model=#{model}#{token_counts(response)})")
40
41
  response
41
42
  rescue Timeout::Error
42
43
  @logger.warn("LLM #{stage} request timed out after #{timeout.round} seconds (model=#{model})")
43
44
  raise
44
45
  end
45
46
 
47
+ # The answer as RubyLLM 1.x gave it under a schema: the parsed JSON when
48
+ # the text is JSON, the text itself otherwise (the pipeline repairs it).
49
+ # RubyLLM 2 always returns the text, and a Hash that breaks the schema
50
+ # would go to a repair request instead of the next model. An empty
51
+ # answer stays an empty String (Message#parsed would turn it into nil);
52
+ # JSON null becomes nil, as in 1.x.
53
+ def self.content(response)
54
+ content = response.content
55
+ return content unless content.is_a?(String) && !content.empty?
56
+
57
+ response.parsed
58
+ rescue JSON::ParserError
59
+ content
60
+ end
61
+
46
62
  private
47
63
 
64
+ # Token counts as the provider reported them, one line per attempt. They
65
+ # are not summed: Gemini already counts thinking into output. Input is
66
+ # what was not read from or written to a cache; the prompt size is
67
+ # input + cache_read + cache_write, and a retry of the same prompt on
68
+ # Gemini is often served from its implicit cache.
69
+ def token_counts(response)
70
+ tokens = response.tokens
71
+ cached = {cache_read: tokens.cache_read, cache_write: tokens.cache_write}.reject { |_, count| count.to_i.zero? }
72
+ counts = {input: tokens.input, output: tokens.output, thinking: tokens.thinking}.compact.merge(cached)
73
+ return '' if counts.empty?
74
+
75
+ ", tokens: #{counts.map { |name, count| "#{name}=#{count}" }.join(' ')}"
76
+ end
77
+
48
78
  def load_ruby_llm
49
79
  require 'ruby_llm'
50
80
  rescue LoadError => e
@@ -103,8 +133,11 @@ module Aireview
103
133
  end
104
134
  end
105
135
 
136
+ # RubyLLM 2 sends OpenAI requests to the Responses API; a compatible
137
+ # server behind LLM_API_BASE usually has Chat Completions only.
106
138
  def configure_remote_provider(ruby_config, provider, api_key)
107
139
  ruby_config.public_send("#{provider}_api_key=", api_key)
140
+ ruby_config.openai_protocol = :chat_completions if provider == 'openai'
108
141
  return unless Aireview::Utils.present?(@config.llm_api_base)
109
142
 
110
143
  ruby_config.public_send("#{provider}_api_base=", @config.llm_api_base)
@@ -111,7 +111,7 @@ module Aireview
111
111
  schema: stage == 'critique' ? CritiqueOutputSchema : GenerateOutputSchema
112
112
  )
113
113
  key = @config.provider_api_keys(candidate.provider).first
114
- @client.request(request, candidate: candidate, key: key, timeout: @config.llm_timeout.to_f).content
114
+ LlmClient.content(@client.request(request, candidate: candidate, key: key, timeout: @config.llm_timeout.to_f))
115
115
  end
116
116
 
117
117
  # missing — the provider has no such model; unverified — the provider
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
- require 'ruby_llm/schema'
2
+ require 'schematist'
3
3
 
4
4
  module Aireview
5
5
  module OutputSchemaValues
@@ -18,7 +18,7 @@ module Aireview
18
18
  DECISIONS = %w[keep reject].freeze
19
19
  end
20
20
 
21
- class GenerateOutputSchema < RubyLLM::Schema
21
+ class GenerateOutputSchema < Schematist::Schema
22
22
  string :summary
23
23
  array :candidates, max_items: 3 do
24
24
  object do
@@ -38,7 +38,7 @@ module Aireview
38
38
  end
39
39
  end
40
40
 
41
- class CritiqueOutputSchema < RubyLLM::Schema
41
+ class CritiqueOutputSchema < Schematist::Schema
42
42
  array :verdicts do
43
43
  object do
44
44
  string :id
@@ -26,6 +26,11 @@ module Aireview
26
26
  DRY_RUN_CANDIDATES_JSON = '[{"id":"C1","file":"path/from/diff.rb","line":1,' \
27
27
  '"quoted_code":"...","problem":"...","why":"...","suggestion":"...",' \
28
28
  '"category":"bug","severity":"major"}]'
29
+ # Finish reasons after which an unparsable answer gets no repair.
30
+ CUT_OFF_REASONS = {
31
+ max_tokens: 'cut off at the output limit (max_tokens)',
32
+ content_filter: 'blocked by the provider (content_filter)'
33
+ }.freeze
29
34
 
30
35
  def initialize(config:, reviewer: nil, context_builder: nil, logger: Logger.new($stderr))
31
36
  @config = config
@@ -214,6 +219,7 @@ module Aireview
214
219
  def parse_string_with_repair(raw:, kind:, expected:, repair_stage:, critique_candidate_ids: nil)
215
220
  parse_expected_result(raw, expected, critique_candidate_ids: critique_candidate_ids)
216
221
  rescue JSON::ParserError, SchemaError => e
222
+ raise_if_cut_off(stage: repair_stage, kind: kind, error: e)
217
223
  @logger.warn("Invalid #{kind} JSON, requesting one repair: #{e.message}")
218
224
  repaired = repair_json(
219
225
  raw: raw,
@@ -225,10 +231,19 @@ module Aireview
225
231
  begin
226
232
  parse_expected_result(repaired, expected, critique_candidate_ids: critique_candidate_ids)
227
233
  rescue JSON::ParserError, SchemaError => second_error
234
+ raise_if_cut_off(stage: repair_stage, kind: "#{kind} repair", error: second_error)
228
235
  raise ParseError, "LLM returned invalid #{kind} JSON after repair: #{second_error.message}"
229
236
  end
230
237
  end
231
238
 
239
+ # An answer the provider cut off or blocked is not a JSON mistake: the
240
+ # same model would cut the repair off too, so the stage goes to the next
241
+ # model at once. A valid answer is taken whatever the reason.
242
+ def raise_if_cut_off(stage:, kind:, error:)
243
+ reason = CUT_OFF_REASONS[@reviewer.finish_reason(stage)]
244
+ raise ParseError, "LLM #{kind} was #{reason}: #{error.message}" if reason
245
+ end
246
+
232
247
  def parse_expected_result(raw, expected, critique_candidate_ids: nil)
233
248
  @parser.parse(raw, expected: expected, critique_candidate_ids: critique_candidate_ids)
234
249
  end
@@ -16,6 +16,7 @@ module Aireview
16
16
  @logger = logger
17
17
  @router = router || LlmRouter.new(config: config, logger: logger)
18
18
  @client = client || LlmClient.new(config: config, logger: logger)
19
+ @finish_reasons = {}
19
20
  end
20
21
 
21
22
  # pinned — a request only to the model that answered last in the stage
@@ -47,6 +48,12 @@ module Aireview
47
48
  @router.critique_weaker?
48
49
  end
49
50
 
51
+ # Why the last answer of the stage stopped (:stop, :max_tokens,
52
+ # :content_filter…), nil when the provider did not say or no answer came.
53
+ def finish_reason(stage)
54
+ @finish_reasons[stage.to_s]
55
+ end
56
+
50
57
  # An invalid result: the model is excluded for the stage, the next
51
58
  # request of the stage goes to another. Returns the excluded model or nil.
52
59
  def exclude_answered_model(stage:, reason:)
@@ -58,11 +65,13 @@ module Aireview
58
65
  # A pinned route giving up is not an API error for the pipeline but
59
66
  # "repair impossible": the same fate as an invalid result.
60
67
  def call_llm(prompt, pinned:)
68
+ @finish_reasons.delete(prompt.stage)
61
69
  response = @router.call(stage: prompt.stage, request_chars: prompt.chars, pinned: pinned) do |route, timeout|
62
70
  @client.request(prompt, candidate: route.candidate, key: route.key, key_index: route.key_index,
63
71
  timeout: timeout)
64
72
  end
65
- response.content
73
+ @finish_reasons[prompt.stage] = response.finish_reason
74
+ LlmClient.content(response)
66
75
  rescue RouteExhaustedError => e
67
76
  raise RepairImpossibleError, e.message
68
77
  end
@@ -1,4 +1,4 @@
1
1
  # frozen_string_literal: true
2
2
  module Aireview
3
- VERSION = '2.0.0'
3
+ VERSION = '2.1.0'
4
4
  end
metadata CHANGED
@@ -1,14 +1,14 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: aireview
3
3
  version: !ruby/object:Gem::Version
4
- version: 2.0.0
4
+ version: 2.1.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Denis Levenko
8
8
  autorequire:
9
9
  bindir: bin
10
10
  cert_chain: []
11
- date: 2026-09-19 00:00:00.000000000 Z
11
+ date: 2026-09-25 00:00:00.000000000 Z
12
12
  dependencies:
13
13
  - !ruby/object:Gem::Dependency
14
14
  name: dotenv
@@ -50,14 +50,14 @@ dependencies:
50
50
  requirements:
51
51
  - - '='
52
52
  - !ruby/object:Gem::Version
53
- version: 1.16.0
53
+ version: 2.0.0
54
54
  type: :runtime
55
55
  prerelease: false
56
56
  version_requirements: !ruby/object:Gem::Requirement
57
57
  requirements:
58
58
  - - '='
59
59
  - !ruby/object:Gem::Version
60
- version: 1.16.0
60
+ version: 2.0.0
61
61
  description: 'Reviews self-hosted GitLab merge requests with a two-pass LLM pipeline:
62
62
  the first pass finds candidate findings, the second one critiques them and drops
63
63
  the weak ones. Supports Gemini and local Ollama, optional Jira context and posting