minitest-impact 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 729696ea279a7dac126254382c8761f9e26ae5972e0047c1961b5c0b57e6037c
4
- data.tar.gz: d9abac4812e8a11862e87bc4c69c6b6b96f368c3b8b269b2f643c9547479557b
3
+ metadata.gz: db9b1457ea4098827caa398da16219f2ca2fa5034da0e2ccb9dfb2940c74c4f2
4
+ data.tar.gz: 50993034bce9425627b4f20784bb7154bff6a3dcd2ab55233ec72b3dda0c59db
5
5
  SHA512:
6
- metadata.gz: 5b5cb7ba14512aead7ede47003cd271ccbbe749f9c0880666b0cb372d2a6f69f5a28aed4ca96e6c998b7ee30e4b12df0e6a1e7be8b32896f9d541fe150fb23a6
7
- data.tar.gz: 0d83d45f8d06b7762b375e7428ea63b85f20189dce5bcf2d9d4bb08039995c165eda2855d323710d5b09809fc953c72c724734041ed222d02975939c6348ff3f
6
+ metadata.gz: 86d954e6cf0db47b51b0cacf6e672001097f039a2cc526a09b06424a3ff32a45b473780bdde073b1cfe4ceb90b3f610769b8b2a052d63c8669df752c4391c913
7
+ data.tar.gz: 6604fd3f172a91df47f5b76b4d22d9f582dee833c42cc7e2ace3c16eaba9b7d5a8a22a8e52c1ae7307571e3cc6fcb90a549e01349b66f8ecf91645f739acab20
data/CHANGELOG.md ADDED
@@ -0,0 +1,17 @@
1
+ # Changelog
2
+
3
+ ## 0.2.0 (2026-09-30)
4
+
5
+ - Record an app from outside its bundle: the recorder loads `json` only when it first writes, so an
6
+ app that pins another `json` still boots, and `SimpleCov.start` does nothing while recording.
7
+ - `run` switches SimpleCov off while it runs the selection.
8
+ - Record projects that live under a `tmp` directory (Linux's temporary directories, some CI
9
+ workspaces); before, every file was skipped.
10
+ - Trace files the application reads by path; Jev only ranks, it never drops an exact map hit.
11
+ - `run` exits 10 without running anything when the change needs the whole suite.
12
+ - `eval/baselines.rb` scores simple strategies on the same cases, and the README reports them for
13
+ Piou Piou and Fizzy.
14
+
15
+ ## 0.1.0 (2026-09-29)
16
+
17
+ - First release.
data/README.md CHANGED
@@ -16,7 +16,7 @@ Confidence: medium
16
16
  ...
17
17
  ```
18
18
 
19
- It works in two layers:
19
+ It works in three layers; the third is optional:
20
20
 
21
21
  1. **A coverage map, exact and free.** One recording run of your suite notes which project lines
22
22
  every test file executed (Ruby's `Coverage`, with `eval: true` so ERB views count too). Given a
@@ -28,8 +28,9 @@ It works in two layers:
28
28
  conventional test path, and tests that mention the new constant), locale keys (tests that use
29
29
  the key, views that render it), routes (their controllers), migrations and `schema.rb` (the
30
30
  models of the changed tables), fixtures, Stimulus controllers (the views that use them), files a
31
- test reads by path, and files that need the whole suite (`Gemfile.lock`, `test_helper.rb`, boot
32
- configuration).
31
+ test reads by path, files the application reads by path, folder or file name (prompts,
32
+ templates: the tests that ran the reading code), config keys the application reads by name, and
33
+ files that need the whole suite (`Gemfile.lock`, `test_helper.rb`, boot configuration).
33
34
  3. **Jev, optionally**, when the map and the rules are not enough: some file could not be traced,
34
35
  or the selection is too large to run in a loop. See [Jev](#jev) below.
35
36
 
@@ -63,10 +64,15 @@ nothing changes in your test helper. Each test process (Rails' forked parallel w
63
64
  appends one JSON line per test to its own file; when the command ends they are merged into
64
65
  `tmp/minitest-impact/map.json` (`--map PATH` to change it), stamped with the commit it describes.
65
66
 
66
- - **Turn SimpleCov off while recording** (`SimpleCov.start unless ENV["MINITEST_IMPACT_RECORD"]`,
67
- or your app's own switch). Ruby allows one coverage setup per process.
68
- - Recording is slower than a normal run (about 2 to 3 times on a Rails app), because every test
69
- reads and clears the coverage counters. Record on a quiet machine, or in CI, and refresh the map
67
+ - **SimpleCov is off while recording.** Ruby allows one coverage setup per process, so the
68
+ recorder makes `SimpleCov.start` do nothing for the run, with no change to your test helper.
69
+ Another tool that calls `Coverage.start` itself has to be turned off by you
70
+ (`ENV["MINITEST_IMPACT_RECORD"]` is set while recording).
71
+ - **It does not need to be in your Gemfile.** The recorder loads no gem before your app's bundle
72
+ does, so the CLI can run from a checkout of this repository, outside the app's bundle:
73
+ `ruby -I path/to/minitest-impact/lib path/to/minitest-impact/exe/minitest-impact record -- bin/rails test`.
74
+ - Recording is 2 to 3 times slower than a normal run on a Rails app, because every test reads
75
+ and clears the coverage counters. Record on a quiet machine, or in CI, and refresh the map
70
76
  when it drifts: a map a few hundred commits old still works, because methods are matched by name.
71
77
  - Record from a clean working tree; the map is stamped with `HEAD`.
72
78
 
@@ -81,38 +87,44 @@ $ minitest-impact run --since main # select, then bin/rails test the
81
87
  $ bin/rails test:impact SINCE=main # the same, as a Rake task (added by a Railtie)
82
88
  ```
83
89
 
84
- `--format paths` exits with status 10, and prints nothing, when the change needs the whole suite.
85
- `run` and `test:impact` run the whole suite themselves in that case. `--max N` keeps the N most
86
- likely files.
90
+ When the change needs the whole suite, `select --format paths` prints nothing and `run` runs
91
+ nothing, and both exit with status 10; `test:impact` runs the whole suite in that case. `run` and
92
+ `test:impact` switch SimpleCov off for the selected tests, since a minimum-coverage check on a
93
+ handful of tests always fails; with your own runner on `--format paths`, turn coverage off
94
+ yourself. `--max N` keeps the N most likely files.
87
95
 
88
96
  ### For coding agents
89
97
 
90
98
  Put this in the agent's instructions (`CLAUDE.md`, `AGENTS.md`):
91
99
 
92
- > While you work, run `bin/rails test:impact SINCE=main` instead of the whole suite. When it says
93
- > "Confidence: low", or when you are done, the full gate runs; you do not run it yourself.
100
+ > While you work, run `bin/rails test:impact SINCE=main` instead of the whole suite. Do not run
101
+ > the whole suite yourself: it runs after you finish.
94
102
 
95
- and run the full gate as code after the agent says it is done, feeding back only failures.
103
+ Then make that true in your harness: run the full suite as code once the agent says it is done,
104
+ and feed back only the failures. On "Confidence: low", `test:impact` already runs the whole suite.
96
105
 
97
106
  ## Jev
98
107
 
99
- [Jev](https://docs.typesafe.ai) is TypeSafe's System One model: a fast, cheap classifier that
100
- answers typed questions (yes/no, choice, score) about a state, with calibrated probabilities. It
101
- is used here the way TypeSafe's own guidance says to use it:
108
+ [Jev](https://docs.typesafe.ai) is TypeSafe's fast, cheap classifier: it answers typed questions
109
+ (yes/no, choice, score) about a state, with calibrated probabilities.
102
110
 
103
- - **Rules stay in code.** Jev never decides what a test file is, which files need the whole suite,
111
+ Measured on one app so far (see [the numbers](#with-jev-2026-09-30)): it puts the test written for
112
+ the change first far more often, and it never found a test the map and the rules had missed. It is
113
+ used as a ranker, not a filter.
114
+
115
+ - Rules stay in code. Jev never decides what a test file is, which files need the whole suite,
104
116
  or anything else a path can tell.
105
- - **One narrow question per judgment, all in one request.** The state is the change (paths, a
117
+ - One request per selection, one narrow question per judgment. The state is the change (paths, a
106
118
  trimmed diff, your `--intent`) and up to 48 candidate test files with their test names. The
107
119
  questions: one yes/no per candidate ("do these tests call, render or assert on something the
108
- change modifies?"), one choice of the candidate most directly written for the change, **with a
109
- "none" option**, and one yes/no for "does every test depend on this?".
110
- - **Fast search first, then re-rank**, as in TypeSafe's re-ranking cookbook: the candidates are
111
- the map's selection plus test files whose paths and test names share words with the change.
112
- - **Exact map hits are never dropped.** Jev can add tests, drop weak non-exact ones and reorder,
113
- but a test the map saw run the changed method stays.
114
- - **Thresholds live in one file** (`lib/minitest/impact/jev/questions.rb`) and **the model version
115
- is pinned** (`jev-1.13.0`), because a threshold tuned on one version does not carry over.
120
+ change modifies?"), one choice of the candidate most directly written for the change, with a
121
+ "none" option, and one yes/no for "does every test depend on this?".
122
+ - The candidates come from a fast search, and Jev only re-ranks them: the map's selection plus
123
+ test files whose paths and test names share words with the change.
124
+ - Jev never drops a test. It can add tests and reorder them. Letting it drop weak picks lost tests
125
+ the change needed (7 in 158 cases) and bought little.
126
+ - The thresholds live in one file (`lib/minitest/impact/jev/questions.rb`) and the model version
127
+ is pinned (`jev-1.13.0`), because a threshold tuned on one version does not carry over.
116
128
 
117
129
  Set `TYPESAFE_API_KEY` to turn it on (`TYPESAFE_BASE_URL` for another endpoint,
118
130
  `MINITEST_IMPACT_JEV_MODEL` to move the pin). Without a key, or with `--no-jev`, everything runs
@@ -153,35 +165,121 @@ Two kinds of labelled cases:
153
165
  Reported per case and on average: whether any expected test was selected (`caught`), whether all
154
166
  were (`all_caught`), recall, precision, and the share of the suite selected.
155
167
 
156
- ### First numbers: Piou Piou, map only (2026-09-28)
168
+ ### First numbers: one Rails app, map only (2026-09-28)
157
169
 
158
- The map was recorded at one commit: 396 test files (2,673 unit and 88 system tests), 580 KB of
159
- JSON (72 KB gzipped). Recording made the unit suite about 2 to 3 times slower. The evaluation made
160
- no Jev calls.
170
+ Measured on Piou Piou, a Rails 8.1 app, with the map recorded at one commit: 396 test files (2,673
171
+ unit and 88 system tests), 580 KB of JSON (72 KB gzipped). The evaluation made no Jev calls.
161
172
 
162
173
  | Cases | Caught (any expected test selected) | All caught | Recall | Precision | Share of suite selected | Share of suite time |
163
174
  |---|---:|---:|---:|---:|---:|---:|
164
175
  | 150 commits that changed code and tests together, method-level | 94.0% | 90.7% | 0.93 | 0.17 | 12.1% | 24.0% |
165
176
  | The same 150, file-level | 94.0% | 90.7% | 0.93 | 0.15 | 12.3% | 24.4% |
166
177
  | 17 failed CI runs on main | 70.6% | 64.7% | 0.68 | 0.07 | 38.5% | 42.2% |
167
- | The same, without 4 runs where only a flaky system test failed | 12 of 13 | 11 of 13 | 0.88 | 0.10 | 50.2% | 55.1% |
178
+ | The same, without 4 runs where only a flaky system test failed | 92.3% (12 of 13) | 84.6% (11 of 13) | 0.88 | 0.10 | 50.2% | 55.1% |
168
179
 
169
180
  How to read them:
170
181
 
171
- - Six of the 13 real CI breaks changed `Gemfile.lock`, `ci.yml`-tested config or `test_helper.rb`,
172
- so the rule selected the whole suite. That is correct but costly. On the other seven, the
173
- selection was 7.5% of the suite.
182
+ - Six of the 13 real CI breaks changed `Gemfile.lock`, `test_helper.rb` or boot configuration,
183
+ so the rules selected the whole suite. That is correct but costly, and it is why the share
184
+ selected rises to 50% once the flaky runs are left out. On the other seven, the selection was
185
+ 7.5% of the suite.
174
186
  - The one real break missed: a `config/piou.yml` change that failed
175
187
  `test/services/sandbox_container_test.rb`, which reads the setting through the app and never
176
188
  names the file.
177
189
  - Method-level tracing barely beats file-level on this history. Most changes land in small,
178
190
  focused files, where the two agree.
179
191
 
192
+ ### With Jev (2026-09-30)
193
+
194
+ The same app two days later, with a map recorded at one commit (371 test files), the rules for
195
+ files and config keys the application reads, and Jev as a ranker. "First" and "in the top 5" count
196
+ the cases where an expected test was ranked there; `--format json` lists `selected` in rank order.
197
+
198
+ | 150 commits that changed code and tests together | Caught | All caught | Recall | First | In the top 5 | Share of suite selected | Share of suite time |
199
+ |---|---:|---:|---:|---:|---:|---:|---:|
200
+ | Map and the older rules | 95.3% | 91.3% | 0.94 | 39% | 79% | 11.3% | 19.5% |
201
+ | Map and the older rules, Jev allowed to drop picks | 94.0% | 88.0% | 0.92 | 55% | 82% | 9.4% | 16.5% |
202
+ | Map and the current rules | 97.3% | 95.3% | 0.97 | 39% | 81% | 13.0% | 21.8% |
203
+ | Map, the current rules and Jev as a ranker | 97.3% | 95.3% | 0.97 | 61% | 85% | 13.0% | 21.9% |
204
+
205
+ How to read them:
206
+
207
+ - The rules for what the application reads recovered 6 of the 15 tests the older rules missed,
208
+ all of them prompts and templates read through a service. They also select 18% more files.
209
+ - Jev's gain is the order. An agent running `--max 5` gets the test written for its change in 85%
210
+ of cases, against 81% without it.
211
+ - Of the 9 tests still missed, 3 belong to a commit that added a config key and its tests, with no
212
+ application code reading it yet: nothing but those tests could have pointed at them. Two are
213
+ architecture tests that read the whole source tree. The rest follow a seeds change, an importmap
214
+ change, a one-word config key (`technical`, too common to search for), and one change spread
215
+ over an agent, two jobs and a model.
216
+ - Jev answered in about 4.7 seconds per selection, one request each.
217
+ - Only 8 failed CI runs fit the newer map, too few to report.
218
+
219
+ ### A second app: Fizzy (2026-09-30)
220
+
221
+ [Fizzy](https://github.com/basecamp/fizzy), Basecamp's open-source Rails app, with the map
222
+ recorded at `a703bf1de` (2026-09-29): 261 test files (1,703 unit tests; the system tests were not
223
+ recorded), 686 KB of JSON (62 KB gzipped), on SQLite in a Docker container. Recording took 44
224
+ seconds against 39 for a plain run. The evaluation made no Jev calls.
225
+
226
+ | Cases | Caught | All caught | Recall | Precision | Share of suite selected | Share of suite time |
227
+ |---|---:|---:|---:|---:|---:|---:|
228
+ | 200 commits that changed code and tests together (December 2025 to September 2026), method-level | 95.5% | 93.5% | 0.95 | 0.30 | 13.1% | 20.7% |
229
+
230
+ How to read them:
231
+
232
+ - 6 cases changed a file that needs the whole suite. 43 ended with "Confidence: low"; an agent
233
+ that runs the suite on those, as `test:impact` does, gets 95.5% all caught at 35% of suite time.
234
+ - 13 cases missed a test. 5 of them selected nothing: a new Action Text patch in `lib/rails_ext`,
235
+ a service worker view, a SQLite search adapter that no longer exists at the map's commit, a
236
+ partial changed with the SaaS lockfile, and a commit whose only code change was
237
+ `test/test_helper.rb`, which the evaluation hides from the selector along with the tests. A
238
+ `config/routes.rb` change selected its controllers but not `test/routes_test.rb`.
239
+ - No failed CI runs could be used. GitHub keeps Actions logs for 90 days; the 15 failed runs on
240
+ `main` whose logs were still there failed installing packages or gems, or on a flaky system
241
+ test in the SaaS bundle. None reported a failing unit test.
242
+
243
+ ### Against simpler strategies
244
+
245
+ `eval/baselines.rb` replays the same cases with two strategies a coding agent can follow with
246
+ no map: **conventional**, the changed tests plus the test named after each changed file
247
+ (`app/models/invoice.rb` to `test/models/invoice_test.rb`), and **mentions**, the tests that name
248
+ the constant a changed file defines. Both keep the gem's whole-suite rule, which needs no map.
249
+
250
+ ```console
251
+ $ ruby -Ilib eval/baselines.rb --repo ../app --map map.json --results eval.json [--cases cases.json]
252
+ ```
253
+
254
+ | Cases | Strategy | Caught | All caught | Recall | Precision | Share of suite selected | Share of suite time |
255
+ |---|---|---:|---:|---:|---:|---:|---:|
256
+ | Piou Piou, 150 commits | minitest-impact | 94.0% | 90.7% | 0.93 | 0.17 | 12.1% | 24.0% |
257
+ | | conventional | 78.0% | 50.7% | 0.66 | 0.53 | 2.8% | 4.0% |
258
+ | | mentions | 78.0% | 62.0% | 0.72 | 0.18 | 6.4% | 12.4% |
259
+ | | both | 81.3% | 66.0% | 0.76 | 0.21 | 6.4% | 12.4% |
260
+ | Piou Piou, 17 failed CI runs | minitest-impact | 70.6% | 64.7% | 0.68 | 0.07 | 38.5% | 42.2% |
261
+ | | conventional, mentions or both | 47.1% | 47.1% | 0.47 | 0.02 | 36.3% | 37.3% |
262
+ | Fizzy, 200 commits | minitest-impact | 95.5% | 93.5% | 0.95 | 0.30 | 13.1% | 20.7% |
263
+ | | conventional | 68.5% | 50.0% | 0.60 | 0.54 | 3.6% | 4.4% |
264
+ | | mentions | 68.5% | 54.0% | 0.62 | 0.40 | 5.1% | 6.3% |
265
+ | | both | 75.5% | 61.0% | 0.69 | 0.46 | 5.2% | 6.6% |
266
+
267
+ How to read them:
268
+
269
+ - On commits, the map found every test a change needed in 25 (Piou Piou) and 32 (Fizzy) more
270
+ cases in 100 than the best simple strategy, at twice its test time on Piou Piou and three times
271
+ on Fizzy.
272
+ - The simple strategies are more precise. When a change touches one model and its test, they
273
+ pick that test; the map also picks the controllers and jobs that ran the changed method.
274
+ - Of the 13 real CI breaks on Piou Piou (the 4 other runs failed on a flaky system test), the map
275
+ caught 12 and every simple strategy 8, at about 40% of suite time for all: six of them needed
276
+ the whole suite.
277
+
180
278
  ## Prior art
181
279
 
182
280
  Nothing did most of this for Minitest, offline, when this gem was written (September 2026):
183
281
 
184
- | Project | What it is | What we took |
282
+ | Project | What it is | What this gem took |
185
283
  |---|---|---|
186
284
  | [Crystalball](https://github.com/toptal/crystalball) (and GitLab's fork) | Coverage-map test selection for RSpec | The shape: record per-test coverage, predict from the diff; views, locales and schema as separate strategies |
187
285
  | [affected_tests](https://rubygems.org/gems/affected_tests), [test_impact](https://rubygems.org/gems/test_impact) | 2026 map-based selectors, RSpec only | `Coverage.result(clear: true)` per test; exit code for "run everything"; a staleness warning |
data/eval/baselines.rb ADDED
@@ -0,0 +1,87 @@
1
+ # frozen_string_literal: true
2
+
3
+ # Compares minitest-impact's selection with two strategies a coding agent could follow with no
4
+ # map at all, on the same labelled cases, so the gem has to earn its place:
5
+ #
6
+ # - conventional: the changed test files, plus the test named after each changed file
7
+ # (app/models/invoice.rb => test/models/invoice_test.rb).
8
+ # - mentions: tests whose source names the constant a changed file defines.
9
+ #
10
+ # Both apply the gem's own whole-suite rule (Gemfile.lock, test_helper.rb...), which needs no map.
11
+ #
12
+ # ruby -Ilib eval/baselines.rb --repo ../app --map map.json --results eval-co.json [--cases ci-cases.json]
13
+ #
14
+ # --results is the JSON `minitest-impact eval --format json` wrote for the same cases; its per-case
15
+ # ids say which commits to replay. --cases gives base/head for cases that are not single commits.
16
+
17
+ require "json"
18
+ require "optparse"
19
+ require "minitest/impact"
20
+
21
+ module Minitest
22
+ module Impact
23
+ module Baselines
24
+ module_function
25
+
26
+ def selections(repo, map, base:, head:, hide_tests:)
27
+ files = Diff.parse(repo.git("diff", "--no-color", "-M", "--unified=0", base, head, allow_failure: true).to_s)
28
+ files = files.reject { |file| file.path.start_with?(Rules::TEST_ROOT) } if hide_tests
29
+ known = map.test_files
30
+ return { whole: true } if files.any? { |file| Rules.whole_suite?(file.path) }
31
+
32
+ changed_tests = files.map(&:path).select { |path| Rules.test_file?(path) }
33
+ conventional = changed_tests + files.flat_map { |file| Rules.conventional_tests(file.path) }
34
+ mentions = files.filter_map { |file| Rules.constant_for(file.path) }.uniq
35
+ .flat_map { |constant| repo.grep(constant, paths: ["test"], rev: head) }
36
+ .select { |path| Rules.test_file?(path) }
37
+ { whole: false, conventional: (conventional & known).uniq, mentions: ((changed_tests + mentions) & known).uniq }
38
+ end
39
+
40
+ def score(expected, selected, map)
41
+ hit = expected & selected
42
+ total = map.tests.values.sum(&:seconds)
43
+ { caught: hit.any?, all_caught: (expected - hit).empty?, recall: hit.size.to_f / expected.size,
44
+ precision: selected.empty? ? 0.0 : hit.size.to_f / selected.size,
45
+ selected_share: selected.size.to_f / map.test_files.size,
46
+ seconds_share: total.zero? ? 0.0 : selected.sum { |test| map.seconds(test).to_f } / total }
47
+ end
48
+
49
+ def summarize(rows)
50
+ n = rows.size.to_f
51
+ %i[caught all_caught].to_h { |key| [key, (rows.count { |row| row[key] } / n).round(3)] }
52
+ .merge(%i[recall precision selected_share seconds_share].to_h { |key| [key, (rows.sum { |row| row[key] } / n).round(3)] })
53
+ end
54
+ end
55
+ end
56
+ end
57
+
58
+ options = {}
59
+ OptionParser.new do |opts|
60
+ opts.on("--repo PATH") { options[:repo] = it }
61
+ opts.on("--map PATH") { options[:map] = it }
62
+ opts.on("--results PATH") { options[:results] = it }
63
+ opts.on("--cases PATH") { options[:cases] = it }
64
+ end.parse!
65
+
66
+ repo = Minitest::Impact::Repo.new(options.fetch(:repo))
67
+ map = Minitest::Impact::Map.load(options.fetch(:map))
68
+ results = JSON.parse(File.read(options.fetch(:results))).fetch("cases")
69
+ cases = options[:cases] ? JSON.parse(File.read(options[:cases])).to_h { [it["id"], it] } : {}
70
+
71
+ rows = Hash.new { |hash, key| hash[key] = [] }
72
+ results.each do |result|
73
+ kase = cases[result["id"]]
74
+ head = kase ? kase["head"] : result["id"]
75
+ base = kase ? kase["base"] : "#{head}^"
76
+ expected = result["expected"]
77
+ picks = Minitest::Impact::Baselines.selections(repo, map, base: base, head: head, hide_tests: kase.nil?)
78
+ all = map.test_files
79
+ conventional = picks[:whole] ? all : picks[:conventional]
80
+ mentions = picks[:whole] ? all : picks[:mentions]
81
+ rows[:gem] << Minitest::Impact::Baselines.score(expected, result["selected"], map)
82
+ rows[:conventional] << Minitest::Impact::Baselines.score(expected, conventional, map)
83
+ rows[:mentions] << Minitest::Impact::Baselines.score(expected, mentions, map)
84
+ rows[:conventional_or_mentions] << Minitest::Impact::Baselines.score(expected, (conventional | mentions), map)
85
+ end
86
+
87
+ puts JSON.pretty_generate(rows.transform_values { Minitest::Impact::Baselines.summarize(it) }.merge(cases: results.size))
@@ -67,12 +67,10 @@ module Minitest
67
67
 
68
68
  repo = Repo.new
69
69
  Dir.mktmpdir("minitest-impact") do |dir|
70
- lib = File.expand_path("../..", __dir__)
71
- bootstrap = File.join(__dir__, "record_bootstrap.rb")
72
70
  env = {
73
71
  "MINITEST_IMPACT_RECORD" => dir,
74
72
  "MINITEST_IMPACT_ROOT" => repo.root,
75
- "RUBYOPT" => ["-I#{lib}", "-r#{bootstrap}", @env["RUBYOPT"]].compact.join(" ")
73
+ "RUBYOPT" => rubyopt("record_bootstrap.rb")
76
74
  }
77
75
  ok = system(env, *argv)
78
76
  map = Map.merge(File.join(dir, Recorder::PARTS), commit: repo.head)
@@ -137,7 +135,12 @@ module Minitest
137
135
  return 0 if result.tests.empty?
138
136
 
139
137
  runner = File.exist?("bin/rails") ? ["bin/rails", "test"] : ["ruby", "-Itest", "-e", "ARGV.each { |f| require File.expand_path(f) }"]
140
- system(*runner, *result.tests) ? 0 : 1
138
+ system({ "RUBYOPT" => rubyopt("partial_bootstrap.rb") }, *runner, *result.tests) ? 0 : 1
139
+ end
140
+
141
+ # Loads +bootstrap+ into every Ruby process the command starts, before the application.
142
+ def rubyopt(bootstrap)
143
+ ["-I#{File.expand_path("../..", __dir__)}", "-r#{File.join(__dir__, bootstrap)}", @env["RUBYOPT"]].compact.join(" ")
141
144
  end
142
145
 
143
146
  def print_text(result)
@@ -8,7 +8,9 @@ module Minitest
8
8
  # over test paths and test names builds a shortlist; Jev reads the change once and answers one
9
9
  # question per candidate, in one request (docs.typesafe.ai/cookbooks/rerank_typesafe).
10
10
  #
11
- # Exact map hits are never dropped: the map saw those tests run the changed code.
11
+ # Jev only adds and reorders; it never drops a pick. Measured on Piou Piou's history, dropping
12
+ # lost tests the change needed and bought little, while the new order put the test written
13
+ # for the change first far more often.
12
14
  class Ranker
13
15
  EXACT = 0.9
14
16
 
@@ -96,8 +98,6 @@ module Minitest
96
98
  next if noul < Questions::KEEP
97
99
 
98
100
  by_test[test] = Pick.new(test: test, score: noul * 0.8, reasons: ["Jev: exercises the change (#{noul.round(2)})"], seconds: map.seconds(test))
99
- elsif pick.score < EXACT && noul < Questions::KEEP
100
- by_test.delete(test)
101
101
  else
102
102
  pick.score = pick.score >= EXACT ? pick.score : (pick.score + noul) / 2
103
103
  pick.reasons += ["Jev: #{noul.round(2)}"]
@@ -0,0 +1,14 @@
1
+ # frozen_string_literal: true
2
+
3
+ # Required through RUBYOPT by `minitest-impact run`, before the application loads. Coverage measured
4
+ # on a few selected tests says nothing, and an app's minimum-coverage check would turn a green run
5
+ # red; so SimpleCov does nothing for the run, as while recording.
6
+ module_name = Module.instance_method(:name)
7
+ trace = TracePoint.new(:class) do |event|
8
+ next unless module_name.bind_call(event.self) == "SimpleCov"
9
+
10
+ require_relative "simplecov_off"
11
+ event.self.singleton_class.prepend(Minitest::Impact::SimpleCovOff)
12
+ trace.disable
13
+ end
14
+ trace.enable
@@ -2,18 +2,31 @@
2
2
 
3
3
  # Required through RUBYOPT by `minitest-impact record`, before the application or Bundler load, so
4
4
  # coverage sees every project file from its first line. Minitest 6 no longer loads plugins on its
5
- # own, so a TracePoint waits for Minitest::Test to be defined and wraps every test from there.
5
+ # own, so a TracePoint waits for Minitest::Test to be defined and wraps every test from there. The
6
+ # same TracePoint switches SimpleCov off as soon as it is defined: Ruby allows one coverage setup
7
+ # per process, and an app that starts SimpleCov unconditionally would otherwise fail to boot.
6
8
  require_relative "recorder"
7
9
 
8
10
  if (dir = ENV["MINITEST_IMPACT_RECORD"])
9
11
  Minitest::Impact::Recorder.start(dir: dir, root: ENV.fetch("MINITEST_IMPACT_ROOT", Dir.pwd))
10
12
 
13
+ hooked = []
14
+ module_name = Module.instance_method(:name)
11
15
  trace = TracePoint.new(:class) do |event|
12
- next unless event.self.name == "Minitest::Test"
16
+ name = module_name.bind_call(event.self)
17
+ next if hooked.include?(name)
13
18
 
14
- trace.disable
15
- require_relative "recording_hooks"
16
- event.self.prepend(Minitest::Impact::RecordingHooks)
19
+ case name
20
+ when "Minitest::Test"
21
+ require_relative "recording_hooks"
22
+ event.self.prepend(Minitest::Impact::RecordingHooks)
23
+ when "SimpleCov"
24
+ require_relative "simplecov_off"
25
+ event.self.singleton_class.prepend(Minitest::Impact::SimpleCovOff)
26
+ else next
27
+ end
28
+ hooked << name
29
+ trace.disable if hooked.size == 2
17
30
  end
18
31
  trace.enable
19
32
  end
@@ -1,12 +1,14 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  require "coverage"
4
- require "json"
5
4
 
6
5
  module Minitest
7
6
  module Impact
8
7
  # Records which project lines each test executes. Loaded through RUBYOPT before the
9
- # application boots (see `minitest-impact record`), so it depends on the standard library only.
8
+ # application boots (see `minitest-impact record`), so it depends on the standard library only,
9
+ # and loads no gem there: json, a default gem, would be activated at its newest installed version
10
+ # and the app's bundle would then refuse to boot on any other. It is required at the first write,
11
+ # after the bundle chose its version.
10
12
  #
11
13
  # Each test appends one JSON line to a file named after its process, so forked parallel
12
14
  # workers never share a file handle; `Map.merge` folds the parts into one map afterwards.
@@ -94,6 +96,7 @@ module Minitest
94
96
  end
95
97
 
96
98
  def write(record)
99
+ require "json"
97
100
  File.open(File.join(dir, PARTS, "#{Process.pid}.ndjson"), "a") do |file|
98
101
  file.puts(JSON.generate(record))
99
102
  end
@@ -25,6 +25,15 @@ module Minitest
25
25
  QUIET = [%r{\A(?:docs|tmp|log)/}, /\.md\z/, %r{\A\.github/}, /\ALICENSE/, /\.txt\z/, %r{\A\.claude/}, /\A\.gitignore\z/].freeze
26
26
 
27
27
  ROUTES = "config/routes.rb"
28
+ CONFIG_YAML = %r{\Aconfig/.+\.ya?ml\z}
29
+ CONFIG_KEY = /\A\s*([A-Za-z_][\w-]*):/
30
+
31
+ # Where application code lives, for files it reads by path and config keys it reads by name.
32
+ CODE_ROOTS = %w[app lib].freeze
33
+
34
+ # A folder or bare file name found in more code files than this is too common to say who reads
35
+ # the changed file.
36
+ MAX_READERS = 3
28
37
  LOCALE = %r{\Aconfig/locales/.+\.ya?ml\z}
29
38
  MIGRATION = %r{\Adb/(migrate/.+\.rb|schema\.rb|.+_schema\.rb)\z}
30
39
  VIEW = %r{\Aapp/views/(.+?)/(_?)([^/]+?)\.[^/]+\z}
@@ -53,6 +62,10 @@ module Minitest
53
62
  def quiet?(path) = QUIET.any? { |pattern| pattern.match?(path) }
54
63
  def ruby?(path) = path.end_with?(".rb")
55
64
 
65
+ # A key named like `sandbox_ready_poll_seconds` is found in code by name; `model` or
66
+ # `technical` would match unrelated code everywhere.
67
+ def distinctive_key?(key) = key.include?("_") && key.size >= 8
68
+
56
69
  # app/models/billing/invoice.rb => test/models/billing/invoice_test.rb
57
70
  # lib/tasks/x.rb => test/lib/tasks/x_test.rb, test/tasks/x_test.rb
58
71
  def conventional_tests(path)
@@ -215,6 +215,8 @@ module Minitest
215
215
  resolved(change, found || reads.any? ? :convention : :unresolved)
216
216
  elsif Rules.ruby?(change.path) && Rules.constant_for(change.path)
217
217
  resolve_by_name(change, reads.any?)
218
+ elsif resolve_read_by_code(change)
219
+ resolved(change, :convention, "the code reads it")
218
220
  elsif reads.any?
219
221
  resolved(change, :exact, "tests read it by path")
220
222
  elsif Rules.quiet?(change.path)
@@ -282,6 +284,46 @@ module Minitest
282
284
  resolved(change, found ? :convention : :unresolved, constant)
283
285
  end
284
286
 
287
+ # A file the map cannot see (a prompt, a template, a config file) that application code reads:
288
+ # by its path, by the folder it sits in, by its bare file name, or, for config, by the keys
289
+ # the change touched. The tests that ran that code are the ones that can notice.
290
+ def resolve_read_by_code(change)
291
+ path = change.path
292
+ found = false
293
+ readers = code_mentioning(path, change)
294
+ reason = "reads #{path}"
295
+ [[File.dirname(path), "reads files under #{File.dirname(path)}"], [File.basename(path), reason]].each do |needle, why|
296
+ break unless readers.empty?
297
+
298
+ candidates = code_mentioning(needle, change)
299
+ readers, reason = candidates, why if candidates.size <= Rules::MAX_READERS
300
+ end
301
+ readers.each { |file| found = true if map_tests(file, :via_file, "#{reason} through #{file}") }
302
+ found |= resolve_config_keys(change) if Rules::CONFIG_YAML.match?(path)
303
+ found
304
+ end
305
+
306
+ def resolve_config_keys(change)
307
+ keys = config_keys(@repo.read(change.path, @head), change.new_lines) |
308
+ config_keys(@repo.read(change.old_path || change.path, @base), change.old_lines)
309
+ found = false
310
+ keys.select { |key| Rules.distinctive_key?(key) }.each do |key|
311
+ code_mentioning(key, change).each do |file|
312
+ found = true if map_tests(file, :via_file, "reads #{key} from #{change.path} through #{file}")
313
+ end
314
+ end
315
+ found
316
+ end
317
+
318
+ def config_keys(text, lines)
319
+ rows = text.to_s.lines
320
+ lines.filter_map { |line| rows[line - 1]&.[](Rules::CONFIG_KEY, 1) }
321
+ end
322
+
323
+ def code_mentioning(needle, change)
324
+ @repo.grep(needle, paths: Rules::CODE_ROOTS, rev: @head) - [change.path]
325
+ end
326
+
285
327
  def resolve_view_like(view, reason)
286
328
  return map_tests(view, :view, reason || "renders #{view}") if @map.known_file?(view)
287
329
 
@@ -0,0 +1,12 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Minitest
4
+ module Impact
5
+ # Prepended to SimpleCov while recording and during `run`. `SimpleCov.start` does nothing: no
6
+ # second coverage setup beside the recorder's, no report, and no minimum-coverage exit status
7
+ # measured on counters the recorder clears or on a handful of selected tests.
8
+ module SimpleCovOff
9
+ def start(*) = nil
10
+ end
11
+ end
12
+ end
@@ -2,6 +2,6 @@
2
2
 
3
3
  module Minitest
4
4
  module Impact
5
- VERSION = "0.1.0"
5
+ VERSION = "0.2.0"
6
6
  end
7
7
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: minitest-impact
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.1.0
4
+ version: 0.2.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Bruno Costanzo
@@ -20,8 +20,10 @@ executables:
20
20
  extensions: []
21
21
  extra_rdoc_files: []
22
22
  files:
23
+ - CHANGELOG.md
23
24
  - LICENSE.txt
24
25
  - README.md
26
+ - eval/baselines.rb
25
27
  - eval/ci_failures.rb
26
28
  - exe/minitest-impact
27
29
  - lib/minitest/impact.rb
@@ -34,6 +36,7 @@ files:
34
36
  - lib/minitest/impact/locale_keys.rb
35
37
  - lib/minitest/impact/map.rb
36
38
  - lib/minitest/impact/methods.rb
39
+ - lib/minitest/impact/partial_bootstrap.rb
37
40
  - lib/minitest/impact/railtie.rb
38
41
  - lib/minitest/impact/rake_task.rb
39
42
  - lib/minitest/impact/record_bootstrap.rb
@@ -42,13 +45,14 @@ files:
42
45
  - lib/minitest/impact/repo.rb
43
46
  - lib/minitest/impact/rules.rb
44
47
  - lib/minitest/impact/selector.rb
48
+ - lib/minitest/impact/simplecov_off.rb
45
49
  - lib/minitest/impact/version.rb
46
50
  homepage: https://github.com/bruno-costanzo/minitest-impact
47
51
  licenses:
48
52
  - MIT
49
53
  metadata:
50
54
  source_code_uri: https://github.com/bruno-costanzo/minitest-impact
51
- changelog_uri: https://github.com/bruno-costanzo/minitest-impact/releases
55
+ changelog_uri: https://github.com/bruno-costanzo/minitest-impact/blob/main/CHANGELOG.md
52
56
  rubygems_mfa_required: 'true'
53
57
  rdoc_options: []
54
58
  require_paths: