minitest-impact 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +17 -0
- data/README.md +133 -35
- data/eval/baselines.rb +87 -0
- data/lib/minitest/impact/cli.rb +7 -4
- data/lib/minitest/impact/jev/ranker.rb +3 -3
- data/lib/minitest/impact/partial_bootstrap.rb +14 -0
- data/lib/minitest/impact/record_bootstrap.rb +18 -5
- data/lib/minitest/impact/recorder.rb +5 -2
- data/lib/minitest/impact/rules.rb +13 -0
- data/lib/minitest/impact/selector.rb +42 -0
- data/lib/minitest/impact/simplecov_off.rb +12 -0
- data/lib/minitest/impact/version.rb +1 -1
- metadata +6 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: db9b1457ea4098827caa398da16219f2ca2fa5034da0e2ccb9dfb2940c74c4f2
|
|
4
|
+
data.tar.gz: 50993034bce9425627b4f20784bb7154bff6a3dcd2ab55233ec72b3dda0c59db
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 86d954e6cf0db47b51b0cacf6e672001097f039a2cc526a09b06424a3ff32a45b473780bdde073b1cfe4ceb90b3f610769b8b2a052d63c8669df752c4391c913
|
|
7
|
+
data.tar.gz: 6604fd3f172a91df47f5b76b4d22d9f582dee833c42cc7e2ace3c16eaba9b7d5a8a22a8e52c1ae7307571e3cc6fcb90a549e01349b66f8ecf91645f739acab20
|
data/CHANGELOG.md
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.2.0 (2026-09-30)
|
|
4
|
+
|
|
5
|
+
- Record an app from outside its bundle: the recorder loads `json` only when it first writes, so an
|
|
6
|
+
app that pins another `json` still boots, and `SimpleCov.start` does nothing while recording.
|
|
7
|
+
- `run` switches SimpleCov off while it runs the selection.
|
|
8
|
+
- Record projects that live under a `tmp` directory (Linux's temporary directories, some CI
|
|
9
|
+
workspaces); before, every file was skipped.
|
|
10
|
+
- Trace files the application reads by path; Jev only ranks, it never drops an exact map hit.
|
|
11
|
+
- `run` exits 10 without running anything when the change needs the whole suite.
|
|
12
|
+
- `eval/baselines.rb` scores simple strategies on the same cases, and the README reports them for
|
|
13
|
+
Piou Piou and Fizzy.
|
|
14
|
+
|
|
15
|
+
## 0.1.0 (2026-09-29)
|
|
16
|
+
|
|
17
|
+
- First release.
|
data/README.md
CHANGED
|
@@ -16,7 +16,7 @@ Confidence: medium
|
|
|
16
16
|
...
|
|
17
17
|
```
|
|
18
18
|
|
|
19
|
-
It works in
|
|
19
|
+
It works in three layers; the third is optional:
|
|
20
20
|
|
|
21
21
|
1. **A coverage map, exact and free.** One recording run of your suite notes which project lines
|
|
22
22
|
every test file executed (Ruby's `Coverage`, with `eval: true` so ERB views count too). Given a
|
|
@@ -28,8 +28,9 @@ It works in two layers:
|
|
|
28
28
|
conventional test path, and tests that mention the new constant), locale keys (tests that use
|
|
29
29
|
the key, views that render it), routes (their controllers), migrations and `schema.rb` (the
|
|
30
30
|
models of the changed tables), fixtures, Stimulus controllers (the views that use them), files a
|
|
31
|
-
test reads by path,
|
|
32
|
-
|
|
31
|
+
test reads by path, files the application reads by path, folder or file name (prompts,
|
|
32
|
+
templates: the tests that ran the reading code), config keys the application reads by name, and
|
|
33
|
+
files that need the whole suite (`Gemfile.lock`, `test_helper.rb`, boot configuration).
|
|
33
34
|
3. **Jev, optionally**, when the map and the rules are not enough: some file could not be traced,
|
|
34
35
|
or the selection is too large to run in a loop. See [Jev](#jev) below.
|
|
35
36
|
|
|
@@ -63,10 +64,15 @@ nothing changes in your test helper. Each test process (Rails' forked parallel w
|
|
|
63
64
|
appends one JSON line per test to its own file; when the command ends they are merged into
|
|
64
65
|
`tmp/minitest-impact/map.json` (`--map PATH` to change it), stamped with the commit it describes.
|
|
65
66
|
|
|
66
|
-
- **
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
67
|
+
- **SimpleCov is off while recording.** Ruby allows one coverage setup per process, so the
|
|
68
|
+
recorder makes `SimpleCov.start` do nothing for the run, with no change to your test helper.
|
|
69
|
+
Another tool that calls `Coverage.start` itself has to be turned off by you
|
|
70
|
+
(`ENV["MINITEST_IMPACT_RECORD"]` is set while recording).
|
|
71
|
+
- **It does not need to be in your Gemfile.** The recorder loads no gem before your app's bundle
|
|
72
|
+
does, so the CLI can run from a checkout of this repository, outside the app's bundle:
|
|
73
|
+
`ruby -I path/to/minitest-impact/lib path/to/minitest-impact/exe/minitest-impact record -- bin/rails test`.
|
|
74
|
+
- Recording is 2 to 3 times slower than a normal run on a Rails app, because every test reads
|
|
75
|
+
and clears the coverage counters. Record on a quiet machine, or in CI, and refresh the map
|
|
70
76
|
when it drifts: a map a few hundred commits old still works, because methods are matched by name.
|
|
71
77
|
- Record from a clean working tree; the map is stamped with `HEAD`.
|
|
72
78
|
|
|
@@ -81,38 +87,44 @@ $ minitest-impact run --since main # select, then bin/rails test the
|
|
|
81
87
|
$ bin/rails test:impact SINCE=main # the same, as a Rake task (added by a Railtie)
|
|
82
88
|
```
|
|
83
89
|
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
90
|
+
When the change needs the whole suite, `select --format paths` prints nothing and `run` runs
|
|
91
|
+
nothing, and both exit with status 10; `test:impact` runs the whole suite in that case. `run` and
|
|
92
|
+
`test:impact` switch SimpleCov off for the selected tests, since a minimum-coverage check on a
|
|
93
|
+
handful of tests always fails; with your own runner on `--format paths`, turn coverage off
|
|
94
|
+
yourself. `--max N` keeps the N most likely files.
|
|
87
95
|
|
|
88
96
|
### For coding agents
|
|
89
97
|
|
|
90
98
|
Put this in the agent's instructions (`CLAUDE.md`, `AGENTS.md`):
|
|
91
99
|
|
|
92
|
-
> While you work, run `bin/rails test:impact SINCE=main` instead of the whole suite.
|
|
93
|
-
>
|
|
100
|
+
> While you work, run `bin/rails test:impact SINCE=main` instead of the whole suite. Do not run
|
|
101
|
+
> the whole suite yourself: it runs after you finish.
|
|
94
102
|
|
|
95
|
-
|
|
103
|
+
Then make that true in your harness: run the full suite as code once the agent says it is done,
|
|
104
|
+
and feed back only the failures. On "Confidence: low", `test:impact` already runs the whole suite.
|
|
96
105
|
|
|
97
106
|
## Jev
|
|
98
107
|
|
|
99
|
-
[Jev](https://docs.typesafe.ai) is TypeSafe's
|
|
100
|
-
|
|
101
|
-
is used here the way TypeSafe's own guidance says to use it:
|
|
108
|
+
[Jev](https://docs.typesafe.ai) is TypeSafe's fast, cheap classifier: it answers typed questions
|
|
109
|
+
(yes/no, choice, score) about a state, with calibrated probabilities.
|
|
102
110
|
|
|
103
|
-
|
|
111
|
+
Measured on one app so far (see [the numbers](#with-jev-2026-09-30)): it puts the test written for
|
|
112
|
+
the change first far more often, and it never found a test the map and the rules had missed. It is
|
|
113
|
+
used as a ranker, not a filter.
|
|
114
|
+
|
|
115
|
+
- Rules stay in code. Jev never decides what a test file is, which files need the whole suite,
|
|
104
116
|
or anything else a path can tell.
|
|
105
|
-
-
|
|
117
|
+
- One request per selection, one narrow question per judgment. The state is the change (paths, a
|
|
106
118
|
trimmed diff, your `--intent`) and up to 48 candidate test files with their test names. The
|
|
107
119
|
questions: one yes/no per candidate ("do these tests call, render or assert on something the
|
|
108
|
-
change modifies?"), one choice of the candidate most directly written for the change,
|
|
109
|
-
"none" option
|
|
110
|
-
-
|
|
111
|
-
|
|
112
|
-
-
|
|
113
|
-
|
|
114
|
-
-
|
|
115
|
-
is pinned
|
|
120
|
+
change modifies?"), one choice of the candidate most directly written for the change, with a
|
|
121
|
+
"none" option, and one yes/no for "does every test depend on this?".
|
|
122
|
+
- The candidates come from a fast search, and Jev only re-ranks them: the map's selection plus
|
|
123
|
+
test files whose paths and test names share words with the change.
|
|
124
|
+
- Jev never drops a test. It can add tests and reorder them. Letting it drop weak picks lost tests
|
|
125
|
+
the change needed (7 in 158 cases) and bought little.
|
|
126
|
+
- The thresholds live in one file (`lib/minitest/impact/jev/questions.rb`) and the model version
|
|
127
|
+
is pinned (`jev-1.13.0`), because a threshold tuned on one version does not carry over.
|
|
116
128
|
|
|
117
129
|
Set `TYPESAFE_API_KEY` to turn it on (`TYPESAFE_BASE_URL` for another endpoint,
|
|
118
130
|
`MINITEST_IMPACT_JEV_MODEL` to move the pin). Without a key, or with `--no-jev`, everything runs
|
|
@@ -153,35 +165,121 @@ Two kinds of labelled cases:
|
|
|
153
165
|
Reported per case and on average: whether any expected test was selected (`caught`), whether all
|
|
154
166
|
were (`all_caught`), recall, precision, and the share of the suite selected.
|
|
155
167
|
|
|
156
|
-
### First numbers:
|
|
168
|
+
### First numbers: one Rails app, map only (2026-09-28)
|
|
157
169
|
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
no Jev calls.
|
|
170
|
+
Measured on Piou Piou, a Rails 8.1 app, with the map recorded at one commit: 396 test files (2,673
|
|
171
|
+
unit and 88 system tests), 580 KB of JSON (72 KB gzipped). The evaluation made no Jev calls.
|
|
161
172
|
|
|
162
173
|
| Cases | Caught (any expected test selected) | All caught | Recall | Precision | Share of suite selected | Share of suite time |
|
|
163
174
|
|---|---:|---:|---:|---:|---:|---:|
|
|
164
175
|
| 150 commits that changed code and tests together, method-level | 94.0% | 90.7% | 0.93 | 0.17 | 12.1% | 24.0% |
|
|
165
176
|
| The same 150, file-level | 94.0% | 90.7% | 0.93 | 0.15 | 12.3% | 24.4% |
|
|
166
177
|
| 17 failed CI runs on main | 70.6% | 64.7% | 0.68 | 0.07 | 38.5% | 42.2% |
|
|
167
|
-
| The same, without 4 runs where only a flaky system test failed | 12 of 13 | 11 of 13 | 0.88 | 0.10 | 50.2% | 55.1% |
|
|
178
|
+
| The same, without 4 runs where only a flaky system test failed | 92.3% (12 of 13) | 84.6% (11 of 13) | 0.88 | 0.10 | 50.2% | 55.1% |
|
|
168
179
|
|
|
169
180
|
How to read them:
|
|
170
181
|
|
|
171
|
-
- Six of the 13 real CI breaks changed `Gemfile.lock`, `
|
|
172
|
-
so the
|
|
173
|
-
|
|
182
|
+
- Six of the 13 real CI breaks changed `Gemfile.lock`, `test_helper.rb` or boot configuration,
|
|
183
|
+
so the rules selected the whole suite. That is correct but costly, and it is why the share
|
|
184
|
+
selected rises to 50% once the flaky runs are left out. On the other seven, the selection was
|
|
185
|
+
7.5% of the suite.
|
|
174
186
|
- The one real break missed: a `config/piou.yml` change that failed
|
|
175
187
|
`test/services/sandbox_container_test.rb`, which reads the setting through the app and never
|
|
176
188
|
names the file.
|
|
177
189
|
- Method-level tracing barely beats file-level on this history. Most changes land in small,
|
|
178
190
|
focused files, where the two agree.
|
|
179
191
|
|
|
192
|
+
### With Jev (2026-09-30)
|
|
193
|
+
|
|
194
|
+
The same app two days later, with a map recorded at one commit (371 test files), the rules for
|
|
195
|
+
files and config keys the application reads, and Jev as a ranker. "First" and "in the top 5" count
|
|
196
|
+
the cases where an expected test was ranked there; `--format json` lists `selected` in rank order.
|
|
197
|
+
|
|
198
|
+
| 150 commits that changed code and tests together | Caught | All caught | Recall | First | In the top 5 | Share of suite selected | Share of suite time |
|
|
199
|
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
|
200
|
+
| Map and the older rules | 95.3% | 91.3% | 0.94 | 39% | 79% | 11.3% | 19.5% |
|
|
201
|
+
| Map and the older rules, Jev allowed to drop picks | 94.0% | 88.0% | 0.92 | 55% | 82% | 9.4% | 16.5% |
|
|
202
|
+
| Map and the current rules | 97.3% | 95.3% | 0.97 | 39% | 81% | 13.0% | 21.8% |
|
|
203
|
+
| Map, the current rules and Jev as a ranker | 97.3% | 95.3% | 0.97 | 61% | 85% | 13.0% | 21.9% |
|
|
204
|
+
|
|
205
|
+
How to read them:
|
|
206
|
+
|
|
207
|
+
- The rules for what the application reads recovered 6 of the 15 tests the older rules missed,
|
|
208
|
+
all of them prompts and templates read through a service. They also select 18% more files.
|
|
209
|
+
- Jev's gain is the order. An agent running `--max 5` gets the test written for its change in 85%
|
|
210
|
+
of cases, against 81% without it.
|
|
211
|
+
- Of the 9 tests still missed, 3 belong to a commit that added a config key and its tests, with no
|
|
212
|
+
application code reading it yet: nothing but those tests could have pointed at them. Two are
|
|
213
|
+
architecture tests that read the whole source tree. The rest follow a seeds change, an importmap
|
|
214
|
+
change, a one-word config key (`technical`, too common to search for), and one change spread
|
|
215
|
+
over an agent, two jobs and a model.
|
|
216
|
+
- Jev answered in about 4.7 seconds per selection, one request each.
|
|
217
|
+
- Only 8 failed CI runs fit the newer map, too few to report.
|
|
218
|
+
|
|
219
|
+
### A second app: Fizzy (2026-09-30)
|
|
220
|
+
|
|
221
|
+
[Fizzy](https://github.com/basecamp/fizzy), Basecamp's open-source Rails app, with the map
|
|
222
|
+
recorded at `a703bf1de` (2026-09-29): 261 test files (1,703 unit tests; the system tests were not
|
|
223
|
+
recorded), 686 KB of JSON (62 KB gzipped), on SQLite in a Docker container. Recording took 44
|
|
224
|
+
seconds against 39 for a plain run. The evaluation made no Jev calls.
|
|
225
|
+
|
|
226
|
+
| Cases | Caught | All caught | Recall | Precision | Share of suite selected | Share of suite time |
|
|
227
|
+
|---|---:|---:|---:|---:|---:|---:|
|
|
228
|
+
| 200 commits that changed code and tests together (December 2025 to September 2026), method-level | 95.5% | 93.5% | 0.95 | 0.30 | 13.1% | 20.7% |
|
|
229
|
+
|
|
230
|
+
How to read them:
|
|
231
|
+
|
|
232
|
+
- 6 cases changed a file that needs the whole suite. 43 ended with "Confidence: low"; an agent
|
|
233
|
+
that runs the suite on those, as `test:impact` does, gets 95.5% all caught at 35% of suite time.
|
|
234
|
+
- 13 cases missed a test. 5 of them selected nothing: a new Action Text patch in `lib/rails_ext`,
|
|
235
|
+
a service worker view, a SQLite search adapter that no longer exists at the map's commit, a
|
|
236
|
+
partial changed with the SaaS lockfile, and a commit whose only code change was
|
|
237
|
+
`test/test_helper.rb`, which the evaluation hides from the selector along with the tests. A
|
|
238
|
+
`config/routes.rb` change selected its controllers but not `test/routes_test.rb`.
|
|
239
|
+
- No failed CI runs could be used. GitHub keeps Actions logs for 90 days; the 15 failed runs on
|
|
240
|
+
`main` whose logs were still there failed installing packages or gems, or on a flaky system
|
|
241
|
+
test in the SaaS bundle. None reported a failing unit test.
|
|
242
|
+
|
|
243
|
+
### Against simpler strategies
|
|
244
|
+
|
|
245
|
+
`eval/baselines.rb` replays the same cases with two strategies a coding agent can follow with
|
|
246
|
+
no map: **conventional**, the changed tests plus the test named after each changed file
|
|
247
|
+
(`app/models/invoice.rb` to `test/models/invoice_test.rb`), and **mentions**, the tests that name
|
|
248
|
+
the constant a changed file defines. Both keep the gem's whole-suite rule, which needs no map.
|
|
249
|
+
|
|
250
|
+
```console
|
|
251
|
+
$ ruby -Ilib eval/baselines.rb --repo ../app --map map.json --results eval.json [--cases cases.json]
|
|
252
|
+
```
|
|
253
|
+
|
|
254
|
+
| Cases | Strategy | Caught | All caught | Recall | Precision | Share of suite selected | Share of suite time |
|
|
255
|
+
|---|---|---:|---:|---:|---:|---:|---:|
|
|
256
|
+
| Piou Piou, 150 commits | minitest-impact | 94.0% | 90.7% | 0.93 | 0.17 | 12.1% | 24.0% |
|
|
257
|
+
| | conventional | 78.0% | 50.7% | 0.66 | 0.53 | 2.8% | 4.0% |
|
|
258
|
+
| | mentions | 78.0% | 62.0% | 0.72 | 0.18 | 6.4% | 12.4% |
|
|
259
|
+
| | both | 81.3% | 66.0% | 0.76 | 0.21 | 6.4% | 12.4% |
|
|
260
|
+
| Piou Piou, 17 failed CI runs | minitest-impact | 70.6% | 64.7% | 0.68 | 0.07 | 38.5% | 42.2% |
|
|
261
|
+
| | conventional, mentions or both | 47.1% | 47.1% | 0.47 | 0.02 | 36.3% | 37.3% |
|
|
262
|
+
| Fizzy, 200 commits | minitest-impact | 95.5% | 93.5% | 0.95 | 0.30 | 13.1% | 20.7% |
|
|
263
|
+
| | conventional | 68.5% | 50.0% | 0.60 | 0.54 | 3.6% | 4.4% |
|
|
264
|
+
| | mentions | 68.5% | 54.0% | 0.62 | 0.40 | 5.1% | 6.3% |
|
|
265
|
+
| | both | 75.5% | 61.0% | 0.69 | 0.46 | 5.2% | 6.6% |
|
|
266
|
+
|
|
267
|
+
How to read them:
|
|
268
|
+
|
|
269
|
+
- On commits, the map found every test a change needed in 25 (Piou Piou) and 32 (Fizzy) more
|
|
270
|
+
cases in 100 than the best simple strategy, at twice its test time on Piou Piou and three times
|
|
271
|
+
on Fizzy.
|
|
272
|
+
- The simple strategies are more precise. When a change touches one model and its test, they
|
|
273
|
+
pick that test; the map also picks the controllers and jobs that ran the changed method.
|
|
274
|
+
- Of the 13 real CI breaks on Piou Piou (the 4 other runs failed on a flaky system test), the map
|
|
275
|
+
caught 12 and every simple strategy 8, at about 40% of suite time for all: six of them needed
|
|
276
|
+
the whole suite.
|
|
277
|
+
|
|
180
278
|
## Prior art
|
|
181
279
|
|
|
182
280
|
Nothing did most of this for Minitest, offline, when this gem was written (September 2026):
|
|
183
281
|
|
|
184
|
-
| Project | What it is | What
|
|
282
|
+
| Project | What it is | What this gem took |
|
|
185
283
|
|---|---|---|
|
|
186
284
|
| [Crystalball](https://github.com/toptal/crystalball) (and GitLab's fork) | Coverage-map test selection for RSpec | The shape: record per-test coverage, predict from the diff; views, locales and schema as separate strategies |
|
|
187
285
|
| [affected_tests](https://rubygems.org/gems/affected_tests), [test_impact](https://rubygems.org/gems/test_impact) | 2026 map-based selectors, RSpec only | `Coverage.result(clear: true)` per test; exit code for "run everything"; a staleness warning |
|
data/eval/baselines.rb
ADDED
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
# Compares minitest-impact's selection with two strategies a coding agent could follow with no
|
|
4
|
+
# map at all, on the same labelled cases, so the gem has to earn its place:
|
|
5
|
+
#
|
|
6
|
+
# - conventional: the changed test files, plus the test named after each changed file
|
|
7
|
+
# (app/models/invoice.rb => test/models/invoice_test.rb).
|
|
8
|
+
# - mentions: tests whose source names the constant a changed file defines.
|
|
9
|
+
#
|
|
10
|
+
# Both apply the gem's own whole-suite rule (Gemfile.lock, test_helper.rb...), which needs no map.
|
|
11
|
+
#
|
|
12
|
+
# ruby -Ilib eval/baselines.rb --repo ../app --map map.json --results eval-co.json [--cases ci-cases.json]
|
|
13
|
+
#
|
|
14
|
+
# --results is the JSON `minitest-impact eval --format json` wrote for the same cases; its per-case
|
|
15
|
+
# ids say which commits to replay. --cases gives base/head for cases that are not single commits.
|
|
16
|
+
|
|
17
|
+
require "json"
|
|
18
|
+
require "optparse"
|
|
19
|
+
require "minitest/impact"
|
|
20
|
+
|
|
21
|
+
module Minitest
|
|
22
|
+
module Impact
|
|
23
|
+
module Baselines
|
|
24
|
+
module_function
|
|
25
|
+
|
|
26
|
+
def selections(repo, map, base:, head:, hide_tests:)
|
|
27
|
+
files = Diff.parse(repo.git("diff", "--no-color", "-M", "--unified=0", base, head, allow_failure: true).to_s)
|
|
28
|
+
files = files.reject { |file| file.path.start_with?(Rules::TEST_ROOT) } if hide_tests
|
|
29
|
+
known = map.test_files
|
|
30
|
+
return { whole: true } if files.any? { |file| Rules.whole_suite?(file.path) }
|
|
31
|
+
|
|
32
|
+
changed_tests = files.map(&:path).select { |path| Rules.test_file?(path) }
|
|
33
|
+
conventional = changed_tests + files.flat_map { |file| Rules.conventional_tests(file.path) }
|
|
34
|
+
mentions = files.filter_map { |file| Rules.constant_for(file.path) }.uniq
|
|
35
|
+
.flat_map { |constant| repo.grep(constant, paths: ["test"], rev: head) }
|
|
36
|
+
.select { |path| Rules.test_file?(path) }
|
|
37
|
+
{ whole: false, conventional: (conventional & known).uniq, mentions: ((changed_tests + mentions) & known).uniq }
|
|
38
|
+
end
|
|
39
|
+
|
|
40
|
+
def score(expected, selected, map)
|
|
41
|
+
hit = expected & selected
|
|
42
|
+
total = map.tests.values.sum(&:seconds)
|
|
43
|
+
{ caught: hit.any?, all_caught: (expected - hit).empty?, recall: hit.size.to_f / expected.size,
|
|
44
|
+
precision: selected.empty? ? 0.0 : hit.size.to_f / selected.size,
|
|
45
|
+
selected_share: selected.size.to_f / map.test_files.size,
|
|
46
|
+
seconds_share: total.zero? ? 0.0 : selected.sum { |test| map.seconds(test).to_f } / total }
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
def summarize(rows)
|
|
50
|
+
n = rows.size.to_f
|
|
51
|
+
%i[caught all_caught].to_h { |key| [key, (rows.count { |row| row[key] } / n).round(3)] }
|
|
52
|
+
.merge(%i[recall precision selected_share seconds_share].to_h { |key| [key, (rows.sum { |row| row[key] } / n).round(3)] })
|
|
53
|
+
end
|
|
54
|
+
end
|
|
55
|
+
end
|
|
56
|
+
end
|
|
57
|
+
|
|
58
|
+
options = {}
|
|
59
|
+
OptionParser.new do |opts|
|
|
60
|
+
opts.on("--repo PATH") { options[:repo] = it }
|
|
61
|
+
opts.on("--map PATH") { options[:map] = it }
|
|
62
|
+
opts.on("--results PATH") { options[:results] = it }
|
|
63
|
+
opts.on("--cases PATH") { options[:cases] = it }
|
|
64
|
+
end.parse!
|
|
65
|
+
|
|
66
|
+
repo = Minitest::Impact::Repo.new(options.fetch(:repo))
|
|
67
|
+
map = Minitest::Impact::Map.load(options.fetch(:map))
|
|
68
|
+
results = JSON.parse(File.read(options.fetch(:results))).fetch("cases")
|
|
69
|
+
cases = options[:cases] ? JSON.parse(File.read(options[:cases])).to_h { [it["id"], it] } : {}
|
|
70
|
+
|
|
71
|
+
rows = Hash.new { |hash, key| hash[key] = [] }
|
|
72
|
+
results.each do |result|
|
|
73
|
+
kase = cases[result["id"]]
|
|
74
|
+
head = kase ? kase["head"] : result["id"]
|
|
75
|
+
base = kase ? kase["base"] : "#{head}^"
|
|
76
|
+
expected = result["expected"]
|
|
77
|
+
picks = Minitest::Impact::Baselines.selections(repo, map, base: base, head: head, hide_tests: kase.nil?)
|
|
78
|
+
all = map.test_files
|
|
79
|
+
conventional = picks[:whole] ? all : picks[:conventional]
|
|
80
|
+
mentions = picks[:whole] ? all : picks[:mentions]
|
|
81
|
+
rows[:gem] << Minitest::Impact::Baselines.score(expected, result["selected"], map)
|
|
82
|
+
rows[:conventional] << Minitest::Impact::Baselines.score(expected, conventional, map)
|
|
83
|
+
rows[:mentions] << Minitest::Impact::Baselines.score(expected, mentions, map)
|
|
84
|
+
rows[:conventional_or_mentions] << Minitest::Impact::Baselines.score(expected, (conventional | mentions), map)
|
|
85
|
+
end
|
|
86
|
+
|
|
87
|
+
puts JSON.pretty_generate(rows.transform_values { Minitest::Impact::Baselines.summarize(it) }.merge(cases: results.size))
|
data/lib/minitest/impact/cli.rb
CHANGED
|
@@ -67,12 +67,10 @@ module Minitest
|
|
|
67
67
|
|
|
68
68
|
repo = Repo.new
|
|
69
69
|
Dir.mktmpdir("minitest-impact") do |dir|
|
|
70
|
-
lib = File.expand_path("../..", __dir__)
|
|
71
|
-
bootstrap = File.join(__dir__, "record_bootstrap.rb")
|
|
72
70
|
env = {
|
|
73
71
|
"MINITEST_IMPACT_RECORD" => dir,
|
|
74
72
|
"MINITEST_IMPACT_ROOT" => repo.root,
|
|
75
|
-
"RUBYOPT" =>
|
|
73
|
+
"RUBYOPT" => rubyopt("record_bootstrap.rb")
|
|
76
74
|
}
|
|
77
75
|
ok = system(env, *argv)
|
|
78
76
|
map = Map.merge(File.join(dir, Recorder::PARTS), commit: repo.head)
|
|
@@ -137,7 +135,12 @@ module Minitest
|
|
|
137
135
|
return 0 if result.tests.empty?
|
|
138
136
|
|
|
139
137
|
runner = File.exist?("bin/rails") ? ["bin/rails", "test"] : ["ruby", "-Itest", "-e", "ARGV.each { |f| require File.expand_path(f) }"]
|
|
140
|
-
system(*runner, *result.tests) ? 0 : 1
|
|
138
|
+
system({ "RUBYOPT" => rubyopt("partial_bootstrap.rb") }, *runner, *result.tests) ? 0 : 1
|
|
139
|
+
end
|
|
140
|
+
|
|
141
|
+
# Loads +bootstrap+ into every Ruby process the command starts, before the application.
|
|
142
|
+
def rubyopt(bootstrap)
|
|
143
|
+
["-I#{File.expand_path("../..", __dir__)}", "-r#{File.join(__dir__, bootstrap)}", @env["RUBYOPT"]].compact.join(" ")
|
|
141
144
|
end
|
|
142
145
|
|
|
143
146
|
def print_text(result)
|
|
@@ -8,7 +8,9 @@ module Minitest
|
|
|
8
8
|
# over test paths and test names builds a shortlist; Jev reads the change once and answers one
|
|
9
9
|
# question per candidate, in one request (docs.typesafe.ai/cookbooks/rerank_typesafe).
|
|
10
10
|
#
|
|
11
|
-
#
|
|
11
|
+
# Jev only adds and reorders; it never drops a pick. Measured on Piou Piou's history, dropping
|
|
12
|
+
# lost tests the change needed and bought little, while the new order put the test written
|
|
13
|
+
# for the change first far more often.
|
|
12
14
|
class Ranker
|
|
13
15
|
EXACT = 0.9
|
|
14
16
|
|
|
@@ -96,8 +98,6 @@ module Minitest
|
|
|
96
98
|
next if noul < Questions::KEEP
|
|
97
99
|
|
|
98
100
|
by_test[test] = Pick.new(test: test, score: noul * 0.8, reasons: ["Jev: exercises the change (#{noul.round(2)})"], seconds: map.seconds(test))
|
|
99
|
-
elsif pick.score < EXACT && noul < Questions::KEEP
|
|
100
|
-
by_test.delete(test)
|
|
101
101
|
else
|
|
102
102
|
pick.score = pick.score >= EXACT ? pick.score : (pick.score + noul) / 2
|
|
103
103
|
pick.reasons += ["Jev: #{noul.round(2)}"]
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
# Required through RUBYOPT by `minitest-impact run`, before the application loads. Coverage measured
|
|
4
|
+
# on a few selected tests says nothing, and an app's minimum-coverage check would turn a green run
|
|
5
|
+
# red; so SimpleCov does nothing for the run, as while recording.
|
|
6
|
+
module_name = Module.instance_method(:name)
|
|
7
|
+
trace = TracePoint.new(:class) do |event|
|
|
8
|
+
next unless module_name.bind_call(event.self) == "SimpleCov"
|
|
9
|
+
|
|
10
|
+
require_relative "simplecov_off"
|
|
11
|
+
event.self.singleton_class.prepend(Minitest::Impact::SimpleCovOff)
|
|
12
|
+
trace.disable
|
|
13
|
+
end
|
|
14
|
+
trace.enable
|
|
@@ -2,18 +2,31 @@
|
|
|
2
2
|
|
|
3
3
|
# Required through RUBYOPT by `minitest-impact record`, before the application or Bundler load, so
|
|
4
4
|
# coverage sees every project file from its first line. Minitest 6 no longer loads plugins on its
|
|
5
|
-
# own, so a TracePoint waits for Minitest::Test to be defined and wraps every test from there.
|
|
5
|
+
# own, so a TracePoint waits for Minitest::Test to be defined and wraps every test from there. The
|
|
6
|
+
# same TracePoint switches SimpleCov off as soon as it is defined: Ruby allows one coverage setup
|
|
7
|
+
# per process, and an app that starts SimpleCov unconditionally would otherwise fail to boot.
|
|
6
8
|
require_relative "recorder"
|
|
7
9
|
|
|
8
10
|
if (dir = ENV["MINITEST_IMPACT_RECORD"])
|
|
9
11
|
Minitest::Impact::Recorder.start(dir: dir, root: ENV.fetch("MINITEST_IMPACT_ROOT", Dir.pwd))
|
|
10
12
|
|
|
13
|
+
hooked = []
|
|
14
|
+
module_name = Module.instance_method(:name)
|
|
11
15
|
trace = TracePoint.new(:class) do |event|
|
|
12
|
-
|
|
16
|
+
name = module_name.bind_call(event.self)
|
|
17
|
+
next if hooked.include?(name)
|
|
13
18
|
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
19
|
+
case name
|
|
20
|
+
when "Minitest::Test"
|
|
21
|
+
require_relative "recording_hooks"
|
|
22
|
+
event.self.prepend(Minitest::Impact::RecordingHooks)
|
|
23
|
+
when "SimpleCov"
|
|
24
|
+
require_relative "simplecov_off"
|
|
25
|
+
event.self.singleton_class.prepend(Minitest::Impact::SimpleCovOff)
|
|
26
|
+
else next
|
|
27
|
+
end
|
|
28
|
+
hooked << name
|
|
29
|
+
trace.disable if hooked.size == 2
|
|
17
30
|
end
|
|
18
31
|
trace.enable
|
|
19
32
|
end
|
|
@@ -1,12 +1,14 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
3
|
require "coverage"
|
|
4
|
-
require "json"
|
|
5
4
|
|
|
6
5
|
module Minitest
|
|
7
6
|
module Impact
|
|
8
7
|
# Records which project lines each test executes. Loaded through RUBYOPT before the
|
|
9
|
-
# application boots (see `minitest-impact record`), so it depends on the standard library only
|
|
8
|
+
# application boots (see `minitest-impact record`), so it depends on the standard library only,
|
|
9
|
+
# and loads no gem there: json, a default gem, would be activated at its newest installed version
|
|
10
|
+
# and the app's bundle would then refuse to boot on any other. It is required at the first write,
|
|
11
|
+
# after the bundle chose its version.
|
|
10
12
|
#
|
|
11
13
|
# Each test appends one JSON line to a file named after its process, so forked parallel
|
|
12
14
|
# workers never share a file handle; `Map.merge` folds the parts into one map afterwards.
|
|
@@ -94,6 +96,7 @@ module Minitest
|
|
|
94
96
|
end
|
|
95
97
|
|
|
96
98
|
def write(record)
|
|
99
|
+
require "json"
|
|
97
100
|
File.open(File.join(dir, PARTS, "#{Process.pid}.ndjson"), "a") do |file|
|
|
98
101
|
file.puts(JSON.generate(record))
|
|
99
102
|
end
|
|
@@ -25,6 +25,15 @@ module Minitest
|
|
|
25
25
|
QUIET = [%r{\A(?:docs|tmp|log)/}, /\.md\z/, %r{\A\.github/}, /\ALICENSE/, /\.txt\z/, %r{\A\.claude/}, /\A\.gitignore\z/].freeze
|
|
26
26
|
|
|
27
27
|
ROUTES = "config/routes.rb"
|
|
28
|
+
CONFIG_YAML = %r{\Aconfig/.+\.ya?ml\z}
|
|
29
|
+
CONFIG_KEY = /\A\s*([A-Za-z_][\w-]*):/
|
|
30
|
+
|
|
31
|
+
# Where application code lives, for files it reads by path and config keys it reads by name.
|
|
32
|
+
CODE_ROOTS = %w[app lib].freeze
|
|
33
|
+
|
|
34
|
+
# A folder or bare file name found in more code files than this is too common to say who reads
|
|
35
|
+
# the changed file.
|
|
36
|
+
MAX_READERS = 3
|
|
28
37
|
LOCALE = %r{\Aconfig/locales/.+\.ya?ml\z}
|
|
29
38
|
MIGRATION = %r{\Adb/(migrate/.+\.rb|schema\.rb|.+_schema\.rb)\z}
|
|
30
39
|
VIEW = %r{\Aapp/views/(.+?)/(_?)([^/]+?)\.[^/]+\z}
|
|
@@ -53,6 +62,10 @@ module Minitest
|
|
|
53
62
|
def quiet?(path) = QUIET.any? { |pattern| pattern.match?(path) }
|
|
54
63
|
def ruby?(path) = path.end_with?(".rb")
|
|
55
64
|
|
|
65
|
+
# A key named like `sandbox_ready_poll_seconds` is found in code by name; `model` or
|
|
66
|
+
# `technical` would match unrelated code everywhere.
|
|
67
|
+
def distinctive_key?(key) = key.include?("_") && key.size >= 8
|
|
68
|
+
|
|
56
69
|
# app/models/billing/invoice.rb => test/models/billing/invoice_test.rb
|
|
57
70
|
# lib/tasks/x.rb => test/lib/tasks/x_test.rb, test/tasks/x_test.rb
|
|
58
71
|
def conventional_tests(path)
|
|
@@ -215,6 +215,8 @@ module Minitest
|
|
|
215
215
|
resolved(change, found || reads.any? ? :convention : :unresolved)
|
|
216
216
|
elsif Rules.ruby?(change.path) && Rules.constant_for(change.path)
|
|
217
217
|
resolve_by_name(change, reads.any?)
|
|
218
|
+
elsif resolve_read_by_code(change)
|
|
219
|
+
resolved(change, :convention, "the code reads it")
|
|
218
220
|
elsif reads.any?
|
|
219
221
|
resolved(change, :exact, "tests read it by path")
|
|
220
222
|
elsif Rules.quiet?(change.path)
|
|
@@ -282,6 +284,46 @@ module Minitest
|
|
|
282
284
|
resolved(change, found ? :convention : :unresolved, constant)
|
|
283
285
|
end
|
|
284
286
|
|
|
287
|
+
# A file the map cannot see (a prompt, a template, a config file) that application code reads:
|
|
288
|
+
# by its path, by the folder it sits in, by its bare file name, or, for config, by the keys
|
|
289
|
+
# the change touched. The tests that ran that code are the ones that can notice.
|
|
290
|
+
def resolve_read_by_code(change)
|
|
291
|
+
path = change.path
|
|
292
|
+
found = false
|
|
293
|
+
readers = code_mentioning(path, change)
|
|
294
|
+
reason = "reads #{path}"
|
|
295
|
+
[[File.dirname(path), "reads files under #{File.dirname(path)}"], [File.basename(path), reason]].each do |needle, why|
|
|
296
|
+
break unless readers.empty?
|
|
297
|
+
|
|
298
|
+
candidates = code_mentioning(needle, change)
|
|
299
|
+
readers, reason = candidates, why if candidates.size <= Rules::MAX_READERS
|
|
300
|
+
end
|
|
301
|
+
readers.each { |file| found = true if map_tests(file, :via_file, "#{reason} through #{file}") }
|
|
302
|
+
found |= resolve_config_keys(change) if Rules::CONFIG_YAML.match?(path)
|
|
303
|
+
found
|
|
304
|
+
end
|
|
305
|
+
|
|
306
|
+
def resolve_config_keys(change)
|
|
307
|
+
keys = config_keys(@repo.read(change.path, @head), change.new_lines) |
|
|
308
|
+
config_keys(@repo.read(change.old_path || change.path, @base), change.old_lines)
|
|
309
|
+
found = false
|
|
310
|
+
keys.select { |key| Rules.distinctive_key?(key) }.each do |key|
|
|
311
|
+
code_mentioning(key, change).each do |file|
|
|
312
|
+
found = true if map_tests(file, :via_file, "reads #{key} from #{change.path} through #{file}")
|
|
313
|
+
end
|
|
314
|
+
end
|
|
315
|
+
found
|
|
316
|
+
end
|
|
317
|
+
|
|
318
|
+
def config_keys(text, lines)
|
|
319
|
+
rows = text.to_s.lines
|
|
320
|
+
lines.filter_map { |line| rows[line - 1]&.[](Rules::CONFIG_KEY, 1) }
|
|
321
|
+
end
|
|
322
|
+
|
|
323
|
+
def code_mentioning(needle, change)
|
|
324
|
+
@repo.grep(needle, paths: Rules::CODE_ROOTS, rev: @head) - [change.path]
|
|
325
|
+
end
|
|
326
|
+
|
|
285
327
|
def resolve_view_like(view, reason)
|
|
286
328
|
return map_tests(view, :view, reason || "renders #{view}") if @map.known_file?(view)
|
|
287
329
|
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Minitest
|
|
4
|
+
module Impact
|
|
5
|
+
# Prepended to SimpleCov while recording and during `run`. `SimpleCov.start` does nothing: no
|
|
6
|
+
# second coverage setup beside the recorder's, no report, and no minimum-coverage exit status
|
|
7
|
+
# measured on counters the recorder clears or on a handful of selected tests.
|
|
8
|
+
module SimpleCovOff
|
|
9
|
+
def start(*) = nil
|
|
10
|
+
end
|
|
11
|
+
end
|
|
12
|
+
end
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: minitest-impact
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.2.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Bruno Costanzo
|
|
@@ -20,8 +20,10 @@ executables:
|
|
|
20
20
|
extensions: []
|
|
21
21
|
extra_rdoc_files: []
|
|
22
22
|
files:
|
|
23
|
+
- CHANGELOG.md
|
|
23
24
|
- LICENSE.txt
|
|
24
25
|
- README.md
|
|
26
|
+
- eval/baselines.rb
|
|
25
27
|
- eval/ci_failures.rb
|
|
26
28
|
- exe/minitest-impact
|
|
27
29
|
- lib/minitest/impact.rb
|
|
@@ -34,6 +36,7 @@ files:
|
|
|
34
36
|
- lib/minitest/impact/locale_keys.rb
|
|
35
37
|
- lib/minitest/impact/map.rb
|
|
36
38
|
- lib/minitest/impact/methods.rb
|
|
39
|
+
- lib/minitest/impact/partial_bootstrap.rb
|
|
37
40
|
- lib/minitest/impact/railtie.rb
|
|
38
41
|
- lib/minitest/impact/rake_task.rb
|
|
39
42
|
- lib/minitest/impact/record_bootstrap.rb
|
|
@@ -42,13 +45,14 @@ files:
|
|
|
42
45
|
- lib/minitest/impact/repo.rb
|
|
43
46
|
- lib/minitest/impact/rules.rb
|
|
44
47
|
- lib/minitest/impact/selector.rb
|
|
48
|
+
- lib/minitest/impact/simplecov_off.rb
|
|
45
49
|
- lib/minitest/impact/version.rb
|
|
46
50
|
homepage: https://github.com/bruno-costanzo/minitest-impact
|
|
47
51
|
licenses:
|
|
48
52
|
- MIT
|
|
49
53
|
metadata:
|
|
50
54
|
source_code_uri: https://github.com/bruno-costanzo/minitest-impact
|
|
51
|
-
changelog_uri: https://github.com/bruno-costanzo/minitest-impact/
|
|
55
|
+
changelog_uri: https://github.com/bruno-costanzo/minitest-impact/blob/main/CHANGELOG.md
|
|
52
56
|
rubygems_mfa_required: 'true'
|
|
53
57
|
rdoc_options: []
|
|
54
58
|
require_paths:
|