aireview 0.1.1 → 0.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 100adbc2d1510ad1f6c0e9b152decbfa46e71f8db30dfde36a8ba977528ae215
4
- data.tar.gz: 071d392ef71d935a6b6a055e7e86fbbf06e8f4517251d460d8bbff8087cb8ecf
3
+ metadata.gz: 6f428bec4ae96692e65cd9f219856dff87e04e75dec45448dc9ef6806dd6694a
4
+ data.tar.gz: 515afe0ed7450df6cfcb2ce12040780a9c0f45bbbc45fef08f19c4601239908b
5
5
  SHA512:
6
- metadata.gz: 48d8222728f5bd0cdb27790c9971c316060c9efd89479e6be58ee6f41066e237e4e94cf2f864297b03a3012e1f225b92c4ece4b84ffeca3ad3ef66831a176e09
7
- data.tar.gz: e4cdac4c17edc362e318b548ac5d25198eaf9c916edd670424c7d21b0cdf4baab2f9f61527bb05547e737e30f192d243ee7dfcb80062a483b08bf410466bfcef
6
+ metadata.gz: 97cf66d34a0dee0d02823a3d37ad37b85aefeb4a0e2b997cd9f8c0780c3df9816d7c95f520709d3cc668d3910bf749f26553da2001b0878c0935bf117b6a486d
7
+ data.tar.gz: c9ede525f15c46b1ce3371c7a77394933599bc71c6900f79dcea5a6780a29621378e7cd6d5652480921cefe84e6513c4865bd81265523fb6436752f9b35a422d
data/CHANGELOG.md CHANGED
@@ -1,5 +1,30 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.2.1
4
+
5
+ - An overloaded LLM (503) gets a fifth attempt: the pauses are now about
6
+ 2, 5, 5 and 5 minutes. Rate limits and network errors keep their three
7
+ retries.
8
+
9
+ ## 0.2.0
10
+
11
+ - Context budget: `llm.max_prompt_chars` caps the request of each stage,
12
+ `context.max_*_chars` cap the MR description, the Jira description and
13
+ comments and the diff. The diff is cut by whole files and hunks, never in
14
+ the middle of a hunk, and both stages share the same context.
15
+ - Truncation is marked in the prompt and reported in the review: the result
16
+ line gets a `Partial review` suffix and a `Not reviewed` section lists the
17
+ files and sections that were left out. `--dry-run` and `--verbose` show the
18
+ sizes and the coverage.
19
+ - Renames, mode changes and empty new or deleted files are told apart from
20
+ diffs GitLab did not return (too large, binary, empty without a reason); the
21
+ latter are reported as not reviewed.
22
+ - Unparsable limits in the environment (`MAX_DIFF_CHARS=oops`) fail with a
23
+ `ConfigError` instead of silently falling back to the defaults.
24
+ - A run fails with a clear error when not even one hunk fits next to the
25
+ system prompt, or when the candidates push the Critique request over its
26
+ limit.
27
+
3
28
  ## 0.1.1
4
29
 
5
30
  - The prompts no longer ask the model to check whether dependency and image
data/CONTRIBUTORS.md ADDED
@@ -0,0 +1,8 @@
1
+ # Contributors
2
+
3
+ - [Denis Levenko](https://github.com/DenisDenis9331), author and maintainer.
4
+ - [Sergey Kondrashov](https://github.com/SergoHUH): Ollama support, Jira
5
+ self-hosted configuration, provider cleanup, prompt refactoring.
6
+
7
+ Contributions are welcome: open an issue or a pull request at
8
+ https://github.com/DenisDenis9331/aireview.
data/README.md CHANGED
@@ -163,8 +163,8 @@ address with `/v1` matches the
163
163
  [Ollama configuration in RubyLLM](https://rubyllm.com/configuration/#provider-configuration).
164
164
  `LLM_TIMEOUT` sets the timeout of every LLM request in seconds; for a slow
165
165
  local model it can be raised. It does not apply to an "overloaded" (503)
166
- answer from the provider: such a request gets up to four attempts, the
167
- original one and three retries with pauses of about 2, 5 and 5 minutes, and
166
+ answer from the provider: such a request gets up to five attempts, the
167
+ original one and four retries with pauses of about 2, 5, 5 and 5 minutes, and
168
168
  only the failed stage is repeated, not the whole run.
169
169
 
170
170
  Project rules live in `.aireview.yml`. In YAML the `generate.model` and
@@ -204,11 +204,18 @@ review_instructions: |
204
204
 
205
205
  ollama_api_base: http://localhost:11434/v1
206
206
 
207
+ context:
208
+ max_diff_chars: 120000
209
+ max_mr_description_chars: 8000
210
+ max_jira_description_chars: 8000
211
+ max_jira_comment_chars: 2000
212
+
207
213
  llm:
208
214
  provider: gemini
209
215
  temperature: 0
210
216
  timeout: 60
211
217
  http_proxy: http://127.0.0.1:8888
218
+ max_prompt_chars: 400000
212
219
  generate:
213
220
  provider: gemini
214
221
  model: gemini-3.7-flash
@@ -217,8 +224,46 @@ llm:
217
224
  provider: ollama
218
225
  model: qwen2.5-coder:7b
219
226
  temperature: 0
227
+ max_prompt_chars: 60000
220
228
  ```
221
229
 
230
+ ### Context budget
231
+
232
+ The request to each stage is capped by `llm.max_prompt_chars` (or
233
+ `LLM_MAX_PROMPT_CHARS`; per stage `llm.generate.max_prompt_chars` /
234
+ `LLM_GENERATE_MAX_PROMPT_CHARS` and the same for `critique`). The limits are in
235
+ characters, not tokens: there is no exact tokenizer for the providers locally,
236
+ and the Ollama window is set on the server where the client cannot see it. As
237
+ a rule of thumb one token is three to four characters, so for a local model
238
+ with `OLLAMA_CONTEXT_LENGTH=8192` set the stage limit to about 20 000
239
+ characters to leave room for the answer.
240
+
241
+ The MR and Jira context is assembled once per run and shared by both stages,
242
+ so it is sized for the tighter of the two: the Critique stage also has to fit
243
+ its system prompt and a reserve for the candidates. The sections are cut to
244
+ their own limits first, keeping the beginning: `context.max_diff_chars`,
245
+ `context.max_mr_description_chars`, `context.max_jira_description_chars` and
246
+ `context.max_jira_comment_chars` (`MAX_DIFF_CHARS`, `MAX_MR_DESCRIPTION_CHARS`,
247
+ `MAX_JIRA_DESCRIPTION_CHARS`, `MAX_JIRA_COMMENT_CHARS`). Only then is the diff
248
+ cut, and only by whole files and whole hunks: files in the order GitLab returns
249
+ them, a file that does not fit is shown hunk by hunk, everything after it is
250
+ left out, and a single hunk larger than the whole budget is skipped rather
251
+ than cut in the middle. Renames, mode changes and other files without text
252
+ changes are always listed; files whose diff GitLab did not return (too large,
253
+ binary) are listed as well and reported as not reviewed.
254
+
255
+ Everything that was cut is marked in the prompt, so the model knows that a
256
+ missing requirement or missing code may simply be outside the budget. The
257
+ review reports it too: the result line gets a `Partial review: ...` suffix
258
+ and a `Not reviewed` section lists the files and sections concerned. The
259
+ result itself (`ok` / `needs attention`) is still only about the findings.
260
+
261
+ When even one hunk cannot fit next to the system prompt, or the candidates
262
+ returned by Generate push the Critique request over its limit, the run stops
263
+ with an error instead of silently reviewing less. Raise the limits or extend
264
+ `ignore_paths`. `--dry-run` prints the sizes of every part and the coverage;
265
+ `--verbose` logs them during a real run.
266
+
222
267
  ## Usage
223
268
 
224
269
  ```bash
@@ -237,7 +282,7 @@ bundle _2.3.26_ exec bin/aireview review https://gitlab.company.com/team/project
237
282
  - `--critique-temperature VALUE` overrides the temperature for the Critique pass only.
238
283
  - `--config PATH` points at a specific `.aireview.yml`.
239
284
  - `--no-jira` turns off the Jira enrichment even when the MR carries an issue key.
240
- - `--dry-run` prints the LLM settings and the Generate prompt, plus the Critique prompt unless `--no-critique` is given.
285
+ - `--dry-run` prints the LLM settings, the context sizes and coverage, and the Generate prompt, plus the Critique prompt unless `--no-critique` is given.
241
286
  - `--no-critique` skips the second pass and renders the Generate candidates directly.
242
287
  - `--review-mode MODE` sets the behaviour when a review has already been published: `update` or `once`.
243
288
  - `--force` reviews again even when a review for this state of the MR is already published.
@@ -328,8 +373,8 @@ aireview:
328
373
  ```
329
374
 
330
375
  `timeout: 45m` is a chosen ceiling, not a guarantee that every retry fits in:
331
- when the provider is overloaded, one stage can wait up to ~14 minutes of
332
- pauses plus up to four requests of `LLM_TIMEOUT` each, and there are two
376
+ when the provider is overloaded, one stage can wait up to ~20 minutes of
377
+ pauses plus up to five requests of `LLM_TIMEOUT` each, and there are two
333
378
  stages.
334
379
 
335
380
  Set secrets such as `GITLAB_TOKEN`, `GEMINI_API_KEY` and the optional Jira
@@ -362,6 +407,7 @@ bundle _2.3.26_ exec rspec spec/secret_scrubber_spec.rb
362
407
 
363
408
  - The reviewer does not check whether the specified versions of dependencies and images exist: the model's knowledge of releases is outdated, and that is what CI is for. Syntax errors and contradictions with the MR/Jira requirements are checked as usual.
364
409
  - The CLI looks for `.aireview.yml` and `.env` walking up from the current working directory, so the project config can be kept in the repository root even when the tool is run from `aireview/`.
410
+ - How Ollama behaves when a request is still larger than its context window is up to the server, not to `aireview`: check the `ollama serve` log for truncation messages on your setup and size `max_prompt_chars` so it does not happen.
365
411
 
366
412
  ## Releasing
367
413
 
@@ -379,12 +425,18 @@ API key is stored anywhere. To cut a release:
379
425
  ```
380
426
 
381
427
  The workflow refuses to run when the tag does not match `Aireview::VERSION`,
382
- runs the test suite, builds the gem and pushes it.
428
+ runs the test suite, builds the gem, pushes it and then creates a GitHub
429
+ release for the tag with the matching `CHANGELOG.md` section as its notes and
430
+ the built `.gem` attached.
383
431
 
384
432
  ## Changelog
385
433
 
386
434
  See [CHANGELOG.md](CHANGELOG.md).
387
435
 
436
+ ## Contributors
437
+
438
+ See [CONTRIBUTORS.md](CONTRIBUTORS.md).
439
+
388
440
  ## License
389
441
 
390
442
  [MIT](LICENSE)
data/lib/aireview/cli.rb CHANGED
@@ -82,7 +82,7 @@ module Aireview
82
82
  parser_result: parser_result,
83
83
  gitlab_client: gitlab_client,
84
84
  merge_request: merge_request,
85
- changes_text: render_changes(changes, config),
85
+ changes: prepare_changes(changes, config),
86
86
  jira_issue: maybe_load_jira_issue(config, merge_request, options)
87
87
  }
88
88
  end
@@ -102,18 +102,18 @@ module Aireview
102
102
  [merge_request, changes]
103
103
  end
104
104
 
105
- def render_changes(changes, config)
105
+ # Дифф уходит дальше по файлам, а не одной строкой: бюджет контекста
106
+ # режет его по границам файлов и хунков.
107
+ def prepare_changes(changes, config)
106
108
  diff_fetcher = DiffFetcher.new(ignore_paths: config.ignore_paths, logger: @logger)
107
109
  filtered_changes = diff_fetcher.filter(changes)
108
110
  raise Error, 'No changes left after filtering ignore_paths' if filtered_changes.empty?
109
111
 
110
- scrubbed_changes = SecretScrubber.new(
112
+ SecretScrubber.new(
111
113
  secret_patterns: config.secret_patterns,
112
114
  secret_files: config.secret_files,
113
115
  logger: @logger
114
116
  ).scrub_changes(filtered_changes)
115
-
116
- diff_fetcher.render(scrubbed_changes)
117
117
  end
118
118
 
119
119
  def execute_review(config, context, options)
@@ -122,7 +122,7 @@ module Aireview
122
122
  if options[:dry_run]
123
123
  dry_run = pipeline.dry_run_prompts(
124
124
  merge_request: context[:merge_request],
125
- changes_text: context[:changes_text],
125
+ changes: context[:changes],
126
126
  jira_issue: context[:jira_issue],
127
127
  critique: !options[:no_critique]
128
128
  )
@@ -135,7 +135,7 @@ module Aireview
135
135
 
136
136
  review = pipeline.run(
137
137
  merge_request: context[:merge_request],
138
- changes_text: context[:changes_text],
138
+ changes: context[:changes],
139
139
  jira_issue: context[:jira_issue],
140
140
  critique: !options[:no_critique]
141
141
  )
@@ -157,7 +157,7 @@ module Aireview
157
157
  publisher = Publisher.new(gitlab_client: context[:gitlab_client], logger: @logger)
158
158
  prompts = pipeline.dry_run_prompts(
159
159
  merge_request: context[:merge_request],
160
- changes_text: context[:changes_text],
160
+ changes: context[:changes],
161
161
  jira_issue: context[:jira_issue],
162
162
  critique: !options[:no_critique]
163
163
  )
@@ -321,27 +321,7 @@ module Aireview
321
321
  end
322
322
 
323
323
  def render_dry_run(dry_run)
324
- @out.puts('=== LLM SETTINGS ===')
325
- @out.puts("Generate: #{dry_run[:generate_model]} temperature=#{dry_run[:generate_temperature]}")
326
- if dry_run[:critique_prompt]
327
- @out.puts("Critique: #{dry_run[:critique_model]} temperature=#{dry_run[:critique_temperature]}")
328
- else
329
- @out.puts('Critique: disabled')
330
- end
331
- @out.puts
332
- @out.puts('=== GENERATE SYSTEM PROMPT ===')
333
- @out.puts(dry_run.dig(:generate_prompt, :system_prompt))
334
- @out.puts
335
- @out.puts('=== GENERATE USER PROMPT ===')
336
- @out.puts(dry_run.dig(:generate_prompt, :user_prompt))
337
- return unless dry_run[:critique_prompt]
338
-
339
- @out.puts
340
- @out.puts('=== CRITIQUE SYSTEM PROMPT ===')
341
- @out.puts(dry_run.dig(:critique_prompt, :system_prompt))
342
- @out.puts
343
- @out.puts('=== CRITIQUE USER PROMPT ===')
344
- @out.puts(dry_run.dig(:critique_prompt, :user_prompt))
324
+ DryRunReport.new(@out).render(dry_run)
345
325
  end
346
326
 
347
327
  def help
@@ -4,9 +4,13 @@ require 'pathname'
4
4
  require 'yaml'
5
5
  require_relative 'errors'
6
6
  require_relative 'utils'
7
+ require_relative 'config_limits'
7
8
 
8
9
  module Aireview
9
10
  class Config
11
+ include ConfigLimits
12
+ extend ConfigLimits::ClassMethods
13
+
10
14
  DEFAULT_SECRET_FILES = [
11
15
  '.env',
12
16
  '.env.*',
@@ -21,7 +25,6 @@ module Aireview
21
25
  ].freeze
22
26
 
23
27
  REVIEW_MODES = %w[update once].freeze
24
-
25
28
  DEFAULTS = {
26
29
  'review_language' => 'en',
27
30
  'review_mode' => 'update',
@@ -35,8 +38,10 @@ module Aireview
35
38
  'llm' => {
36
39
  'provider' => 'gemini',
37
40
  'temperature' => 0,
38
- 'timeout' => 60
39
- }
41
+ 'timeout' => 60,
42
+ 'max_prompt_chars' => ConfigLimits::DEFAULT_MAX_PROMPT_CHARS
43
+ },
44
+ 'context' => ConfigLimits::CONTEXT_DEFAULTS
40
45
  }.freeze
41
46
 
42
47
  ENV_MAPPING = {
@@ -97,6 +102,7 @@ module Aireview
97
102
  def self.env_config(env)
98
103
  mapped_env_config(env)
99
104
  .merge('llm' => llm_env_config(env))
105
+ .merge(context_env_config(env))
100
106
  .merge(provider_key_env_config(env))
101
107
  .merge(generic_api_key_env_config(env))
102
108
  end
@@ -113,6 +119,7 @@ module Aireview
113
119
  'provider' => env['LLM_PROVIDER'],
114
120
  'temperature' => parse_float(env['LLM_TEMPERATURE']),
115
121
  'timeout' => parse_float(env['LLM_TIMEOUT']),
122
+ 'max_prompt_chars' => parse_integer(env['LLM_MAX_PROMPT_CHARS'], 'LLM_MAX_PROMPT_CHARS'),
116
123
  'generate' => llm_stage_env_config(env, 'GENERATE'),
117
124
  'critique' => llm_stage_env_config(env, 'CRITIQUE')
118
125
  }.compact.reject { |key, value| %w[generate critique].include?(key) && value.empty? }
@@ -122,7 +129,8 @@ module Aireview
122
129
  {
123
130
  'provider' => env["LLM_#{stage}_PROVIDER"],
124
131
  'model' => env["LLM_#{stage}_MODEL"],
125
- 'temperature' => parse_float(env["LLM_#{stage}_TEMPERATURE"])
132
+ 'temperature' => parse_float(env["LLM_#{stage}_TEMPERATURE"]),
133
+ 'max_prompt_chars' => parse_integer(env["LLM_#{stage}_MAX_PROMPT_CHARS"], "LLM_#{stage}_MAX_PROMPT_CHARS")
126
134
  }.compact
127
135
  end
128
136
 
@@ -0,0 +1,85 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Aireview
4
+ # Лимиты контекста в символах: точного токенизатора для провайдеров локально
5
+ # нет, а окно Ollama задаётся на сервере и клиенту не видно. Дефолты щедрые,
6
+ # под конкретную модель их задают в .aireview.yml.
7
+ module ConfigLimits
8
+ LLM_STAGES = %w[generate critique].freeze
9
+ DEFAULT_MAX_PROMPT_CHARS = 400_000
10
+ CONTEXT_DEFAULTS = {
11
+ 'max_diff_chars' => 120_000,
12
+ 'max_mr_description_chars' => 8_000,
13
+ 'max_jira_description_chars' => 8_000,
14
+ 'max_jira_comment_chars' => 2_000
15
+ }.freeze
16
+ CONTEXT_ENV = {
17
+ 'max_diff_chars' => 'MAX_DIFF_CHARS',
18
+ 'max_mr_description_chars' => 'MAX_MR_DESCRIPTION_CHARS',
19
+ 'max_jira_description_chars' => 'MAX_JIRA_DESCRIPTION_CHARS',
20
+ 'max_jira_comment_chars' => 'MAX_JIRA_COMMENT_CHARS'
21
+ }.freeze
22
+
23
+ module ClassMethods
24
+ def context_env_config(env)
25
+ context = CONTEXT_ENV.each_with_object({}) do |(key, env_key), config|
26
+ value = parse_integer(env[env_key], env_key)
27
+ config[key] = value unless value.nil?
28
+ end
29
+ context.empty? ? {} : {'context' => context}
30
+ end
31
+
32
+ # Лимит, который не разобрался, нельзя молча заменять дефолтом: запрос
33
+ # уйдёт в модель с окном, которого у неё нет.
34
+ def parse_integer(value, name)
35
+ return nil if Aireview::Utils.blank?(value)
36
+
37
+ Integer(value.to_s, 10)
38
+ rescue ArgumentError
39
+ raise ConfigError, "#{name} must be an integer, got #{value.inspect}"
40
+ end
41
+ end
42
+
43
+ # Лимит всего запроса стадии в символах: системный промпт плюс контекст
44
+ # (для критика ещё и кандидаты). Наследуется из llm как model/temperature.
45
+ def max_prompt_chars(stage)
46
+ stage = stage.to_s
47
+ raise ArgumentError, "unknown LLM stage #{stage.inspect}" unless LLM_STAGES.include?(stage)
48
+
49
+ positive_integer!(
50
+ dig('llm', stage, 'max_prompt_chars') || dig('llm', 'max_prompt_chars') || DEFAULT_MAX_PROMPT_CHARS,
51
+ "llm.#{stage}.max_prompt_chars"
52
+ )
53
+ end
54
+
55
+ def max_diff_chars
56
+ context_limit('max_diff_chars')
57
+ end
58
+
59
+ def max_mr_description_chars
60
+ context_limit('max_mr_description_chars')
61
+ end
62
+
63
+ def max_jira_description_chars
64
+ context_limit('max_jira_description_chars')
65
+ end
66
+
67
+ def max_jira_comment_chars
68
+ context_limit('max_jira_comment_chars')
69
+ end
70
+
71
+ private
72
+
73
+ def context_limit(key)
74
+ positive_integer!(dig('context', key) || CONTEXT_DEFAULTS.fetch(key), "context.#{key}")
75
+ end
76
+
77
+ def positive_integer!(value, name)
78
+ integer = Integer(value, exception: false) if value.is_a?(Integer) || value.is_a?(String)
79
+ integer = value.to_i if value.is_a?(Float) && value == value.floor
80
+ return integer if integer.is_a?(Integer) && integer.positive?
81
+
82
+ raise ConfigError, "#{name} must be a positive integer, got #{value.inspect}"
83
+ end
84
+ end
85
+ end
@@ -0,0 +1,197 @@
1
+ # frozen_string_literal: true
2
+ require_relative 'errors'
3
+
4
+ module Aireview
5
+ # Укладывает контекст ревью в бюджет символов и запоминает, что при этом не
6
+ # вошло. Секции MR и Jira режутся до своих лимитов с сохранением начала,
7
+ # дифф по целым файлам, затем по целым хункам; внутри хунка не режем.
8
+ module ContextBudget
9
+ # Пути, которые не вошли, перечисляются в конце диффа; список ограничен,
10
+ # чтобы сам не съел бюджет.
11
+ NOT_SHOWN_LIST_LIMIT = 20
12
+ TRAILER_RESERVE_CHARS = 400
13
+
14
+ Coverage = Struct.new(
15
+ :truncated_sections, :files_not_shown, :files_partial, :files_unavailable, :hunks_skipped,
16
+ keyword_init: true
17
+ ) do
18
+ def self.empty
19
+ new(truncated_sections: [], files_not_shown: [], files_partial: [], files_unavailable: [], hunks_skipped: [])
20
+ end
21
+
22
+ def complete?
23
+ to_h.values.all?(&:empty?)
24
+ end
25
+ end
26
+
27
+ Packed = Struct.new(:text, :shown_hunks, :total_hunks, keyword_init: true)
28
+
29
+ # Начало важнее конца: требования и критерии приёмки обычно там.
30
+ def self.truncate_section(text, limit:, label:, coverage:)
31
+ text = text.to_s
32
+ return text if text.length <= limit
33
+
34
+ coverage.truncated_sections << label
35
+ "#{text[0, limit]}\n[#{label} truncated: #{limit} of #{text.length} chars shown]"
36
+ end
37
+
38
+ def self.pack_entries(entries, budget:, coverage:)
39
+ Packer.new(entries, budget: budget, coverage: coverage).pack
40
+ end
41
+
42
+ # Файлы без хунков идут первыми: они дёшевы и всегда полезны для картины
43
+ # MR. Текстовые файлы идут в порядке GitLab, пока влезают; первый файл,
44
+ # который не влезает, показывается частично, всё после него не показывается.
45
+ # Хунк, который не влез бы даже в пустой бюджет, пропускается с пометкой,
46
+ # а не останавливает раскладку.
47
+ class Packer
48
+ def initialize(entries, budget:, coverage:)
49
+ @non_text, @text = entries.partition { |entry| !entry.text? }
50
+ @budget = budget
51
+ @coverage = coverage
52
+ @total_hunks = @text.sum { |entry| entry.hunks.size }
53
+ end
54
+
55
+ def pack
56
+ @non_text.each { |entry| @coverage.files_unavailable << entry.path if entry.unavailable? }
57
+ full = (@non_text + @text).map(&:render).join("\n")
58
+ return Packed.new(text: full, shown_hunks: @total_hunks, total_hunks: @total_hunks) if full.length <= @budget
59
+
60
+ pack_within_limit
61
+ end
62
+
63
+ private
64
+
65
+ # Что-то придётся опустить, значит нужен хвост со списком пропущенного.
66
+ # @used считает весь собранный текст, включая разделители между
67
+ # файлами: результат не должен выйти за бюджет ни на символ.
68
+ def pack_within_limit
69
+ @parts = @non_text.map(&:render)
70
+ @used = joined_length(@parts)
71
+ @limit = @budget - TRAILER_RESERVE_CHARS
72
+ raise_no_room!(:non_text) if @used > @limit
73
+
74
+ shown_hunks = pack_text_entries
75
+ raise_no_room!(:hunks) if shown_hunks.zero? && @total_hunks.positive?
76
+
77
+ @parts << not_shown_trailer unless @coverage.files_not_shown.empty?
78
+ Packed.new(text: @parts.join("\n"), shown_hunks: shown_hunks, total_hunks: @total_hunks)
79
+ end
80
+
81
+ def pack_text_entries
82
+ shown_hunks = 0
83
+ stopped = false
84
+ @text.each do |entry|
85
+ piece, shown, stopped = stopped ? ['', 0, true] : pack_entry(entry)
86
+ if shown.zero?
87
+ @coverage.files_not_shown << entry.path
88
+ next
89
+ end
90
+
91
+ @parts << piece
92
+ @used = joined_length(@parts)
93
+ shown_hunks += shown
94
+ end
95
+ shown_hunks
96
+ end
97
+
98
+ # Место под следующий кусок с учётом разделителя перед ним.
99
+ def remaining
100
+ @limit - @used - (@parts.empty? ? 0 : 1)
101
+ end
102
+
103
+ # Возвращает [текст, число показанных хунков, остановлена ли раскладка].
104
+ def pack_entry(entry)
105
+ full = entry.render
106
+ return [full, entry.hunks.size, false] if full.length <= remaining
107
+
108
+ body, shown, skipped, stopped = pack_hunks(entry)
109
+ return ['', 0, stopped] if shown.zero?
110
+
111
+ skipped.each { |hunk| @coverage.hunks_skipped << {path: entry.path, hunk: hunk} }
112
+ @coverage.files_partial << {path: entry.path, shown: shown, total: entry.hunks.size}
113
+ [entry.header + body + partial_marker(entry, shown), shown, stopped]
114
+ end
115
+
116
+ # Место сначала отдаётся хункам, которые можно показать, и только на
117
+ # остаток добавляются пометки о слишком больших: иначе пометки могли бы
118
+ # вытеснить единственный подходящий хунк. Факт пропуска в покрытие
119
+ # попадает независимо от того, есть ли для пометки место.
120
+ def pack_hunks(entry)
121
+ base = entry.header.length + partial_marker(entry, 0).length
122
+ shown = []
123
+ skipped = []
124
+ stopped = false
125
+ used = 0
126
+ entry.hunks.each_with_index do |hunk, index|
127
+ if base + hunk.length > @limit
128
+ skipped << index
129
+ next
130
+ end
131
+ if base + used + hunk.length > remaining
132
+ stopped = true
133
+ break
134
+ end
135
+
136
+ shown << index
137
+ used += hunk.length
138
+ end
139
+
140
+ body = render_hunks(entry, shown: shown, skipped: skipped, room: remaining - base - used)
141
+ [body, shown.size, skipped.map { |index| index + 1 }, stopped]
142
+ end
143
+
144
+ def render_hunks(entry, shown:, skipped:, room:)
145
+ marked = skipped.select do |index|
146
+ marker = skip_marker(entry, index)
147
+ next false if marker.length > room
148
+
149
+ room -= marker.length
150
+ true
151
+ end
152
+ entry.hunks.each_with_index.filter_map do |hunk, index|
153
+ next hunk if shown.include?(index)
154
+
155
+ skip_marker(entry, index) if marked.include?(index)
156
+ end.join
157
+ end
158
+
159
+ def skip_marker(entry, index)
160
+ "[hunk #{index + 1} of #{entry.hunks.size} skipped: larger than the context budget]\n"
161
+ end
162
+
163
+ def joined_length(parts)
164
+ parts.sum(&:length) + [parts.size - 1, 0].max
165
+ end
166
+
167
+ def partial_marker(entry, shown)
168
+ "[file #{entry.path}: #{shown} of #{entry.hunks.size} hunks shown]\n"
169
+ end
170
+
171
+ def not_shown_trailer
172
+ paths = @coverage.files_not_shown
173
+ listed = []
174
+ paths.first(NOT_SHOWN_LIST_LIMIT).each do |path|
175
+ break if listed.sum(&:length) + path.length > TRAILER_RESERVE_CHARS / 2
176
+
177
+ listed << path
178
+ end
179
+ rest = paths.size - listed.size
180
+ list = listed.join(', ')
181
+ list += " and #{rest} more" if rest.positive?
182
+ "[#{paths.size} file(s) not shown: #{list}]\n"
183
+ end
184
+
185
+ def raise_no_room!(reason)
186
+ detail = if reason == :non_text
187
+ 'the entries for files without text changes alone exceed it'
188
+ else
189
+ "not a single hunk fits after #{@used} chars of entries without text changes"
190
+ end
191
+ raise ContextBudgetError,
192
+ "Diff does not fit into the context budget of #{@budget} chars: #{detail}. " \
193
+ 'Raise llm.max_prompt_chars / context.max_diff_chars, or add paths to ignore_paths.'
194
+ end
195
+ end
196
+ end
197
+ end