rspec-mergify 0.3.0-x86_64-linux → 0.4.0-x86_64-linux
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +49 -0
- data/lib/mergify/rspec/3.1/mergify_ci.so +0 -0
- data/lib/mergify/rspec/3.2/mergify_ci.so +0 -0
- data/lib/mergify/rspec/3.3/mergify_ci.so +0 -0
- data/lib/mergify/rspec/3.4/mergify_ci.so +0 -0
- data/lib/mergify/rspec/4.0/mergify_ci.so +0 -0
- data/lib/mergify/rspec/ci_insights.rb +202 -3
- data/lib/mergify/rspec/configuration.rb +160 -65
- data/lib/mergify/rspec/flaky_detection.rb +4 -1
- data/lib/mergify/rspec/formatter.rb +59 -7
- data/lib/mergify/rspec/quarantine.rb +7 -1
- data/lib/mergify/rspec/session_verdict.rb +86 -0
- data/lib/mergify/rspec/test_selection.rb +299 -0
- data/lib/mergify/rspec/utils.rb +7 -0
- data/lib/mergify/rspec/version.rb +1 -1
- metadata +4 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 575c4e6cb1bf916b62d8a980400b66d759f5ab93a0336f4b291c334b0fc135e3
|
|
4
|
+
data.tar.gz: 75b9c3987f3ce8f6af6b7cae5475996894ff792c9d7731c0dfdcdfefc9986a23
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 82c9c02a5e7cdea0cec4f4c0c72aede62fd081e530029c180b5e54f3be3561f0ea95883e4490a731a93cd897bdc6e56227724e7b44f71fdf40c6f591ed023ccb
|
|
7
|
+
data.tar.gz: c248034b68730876dc1c7dbc4725aa37ab5fb014a05c74c8c5a74de775b65334a680cc8db8a0a1b7ebe7fa1419a389554efc25d9d1ac17c47ef55d57f9c60088
|
data/README.md
CHANGED
|
@@ -7,6 +7,7 @@ RSpec plugin for [Mergify Test Insights](https://docs.mergify.com/ci-insights/).
|
|
|
7
7
|
- **Test tracing** — Sends OpenTelemetry traces for every test to Mergify's API
|
|
8
8
|
- **Flaky test detection** — Intelligently reruns tests to detect flakiness with budget constraints
|
|
9
9
|
- **Test quarantine** — Quarantines failing tests so they don't block CI
|
|
10
|
+
- **Test selection** — Runs only the previously-failing examples when Mergify's merge queue reruns a job
|
|
10
11
|
|
|
11
12
|
## Installation
|
|
12
13
|
|
|
@@ -51,6 +52,54 @@ The plugin activates automatically when running in CI (detected via the `CI` env
|
|
|
51
52
|
| `RSPEC_MERGIFY_DEBUG` | Print spans to console | `false` |
|
|
52
53
|
| `MERGIFY_TRACEPARENT` | W3C distributed trace context | — |
|
|
53
54
|
| `MERGIFY_TEST_JOB_NAME` | Mergify test job name | — |
|
|
55
|
+
| `MERGIFY_TEST_SELECTION_ENABLE` | Opt this job into test selection (see below) | `false` |
|
|
56
|
+
|
|
57
|
+
### Test selection
|
|
58
|
+
|
|
59
|
+
When Mergify's merge queue reruns a job — a retry, or a step of a batch
|
|
60
|
+
bisection — only the examples that failed on the previous attempt are
|
|
61
|
+
informative. The gem asks Mergify whether the current run is such a rerun and,
|
|
62
|
+
if so, runs only those examples.
|
|
63
|
+
|
|
64
|
+
**This is off until you turn it on, per job.** Set
|
|
65
|
+
`MERGIFY_TEST_SELECTION_ENABLE=true` on the job you want reduced. Installing the
|
|
66
|
+
gem is not enough — a feature that decides not to run tests starts only where
|
|
67
|
+
you wrote that it should. Anything that is not a recognised yes (unset, empty,
|
|
68
|
+
`false`, unparsable) means no, and a job that has not opted in never asks.
|
|
69
|
+
|
|
70
|
+
Once opted in, it only ever removes work:
|
|
71
|
+
|
|
72
|
+
- the answer is applied **after** RSpec's own filters — file and line
|
|
73
|
+
arguments, `--tag`, `-e`, `--only-failures` all still narrow the run, never
|
|
74
|
+
widen it;
|
|
75
|
+
- examples are matched by their id (`./spec/models/user_spec.rb[1:2]`), the
|
|
76
|
+
identity RSpec itself reruns failures by;
|
|
77
|
+
- if Mergify says the previous attempt already ran every one of these examples
|
|
78
|
+
and they passed, the run executes none and exits green;
|
|
79
|
+
- if any example Mergify asks for is not collected here, the full suite runs;
|
|
80
|
+
- any error, timeout, or unrecognised answer runs the full suite;
|
|
81
|
+
- if Mergify refuses to pick a previous attempt (several runs of this job
|
|
82
|
+
report under one name), the run **fails** with Mergify's explanation: give
|
|
83
|
+
each run its own `MERGIFY_TEST_JOB_NAME`;
|
|
84
|
+
- at the end of the run the gem sends Mergify what it concluded (the counts
|
|
85
|
+
and the failing examples) — unless the run failed outside any example (a
|
|
86
|
+
failing `after(:context)` or suite hook), in which case a retry runs the
|
|
87
|
+
full suite.
|
|
88
|
+
|
|
89
|
+
Each parallel_tests worker, and each CI job handed a slice of the spec files,
|
|
90
|
+
is matched to the same slice of the previous attempt, by the examples it
|
|
91
|
+
collects. This needs the slices to be the same on both attempts, which is the
|
|
92
|
+
case for a split by file size or by a committed runtime log; Knapsack Pro's
|
|
93
|
+
Queue Mode hands files out dynamically and is not supported.
|
|
94
|
+
|
|
95
|
+
### Parallel runs
|
|
96
|
+
|
|
97
|
+
[parallel_tests](https://github.com/grosser/parallel_tests) and [turbo_tests](https://github.com/serpapi/turbo_tests) need no setup. Each worker is a separate RSpec process, and each reports to Mergify as its own session:
|
|
98
|
+
|
|
99
|
+
- flaky detection sizes each worker's rerun budget from the tests that worker runs;
|
|
100
|
+
- each worker prints its own Mergify report, so a report covers only that worker's tests.
|
|
101
|
+
|
|
102
|
+
Suites split across CI jobs (knapsack, `circleci tests split`, Buildkite `parallelism`) behave the same way, one session per job.
|
|
54
103
|
|
|
55
104
|
For detailed documentation, see the [official guide](https://docs.mergify.com/ci-insights/test-frameworks/rspec/).
|
|
56
105
|
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
|
@@ -5,18 +5,38 @@ require_relative 'trace'
|
|
|
5
5
|
require_relative 'utils'
|
|
6
6
|
require_relative 'native'
|
|
7
7
|
require_relative 'resources/rspec'
|
|
8
|
+
require_relative 'test_selection'
|
|
9
|
+
require_relative 'session_verdict'
|
|
8
10
|
|
|
9
11
|
module Mergify
|
|
10
12
|
module RSpec
|
|
11
13
|
# Central orchestrator for Mergify Test Insights: records the run's spans,
|
|
12
|
-
# uploads them, and coordinates flaky detection and
|
|
14
|
+
# uploads them, and coordinates flaky detection, quarantine and test
|
|
15
|
+
# selection.
|
|
16
|
+
# rubocop:disable-next Metrics/ClassLength
|
|
13
17
|
class CIInsights
|
|
18
|
+
# The resource attributes pytest-mergify reports, under the same keys.
|
|
19
|
+
# The fingerprint and count describe what this run collected, on every
|
|
20
|
+
# run that took one; the selection keys echo what Mergify answered, as
|
|
21
|
+
# sent, and are absent on a run it never answered.
|
|
22
|
+
TEST_COLLECTION_FINGERPRINT = 'test.collection.fingerprint'
|
|
23
|
+
TEST_COLLECTION_COUNT = 'test.collection.count'
|
|
24
|
+
TEST_SELECTION_ANSWER = 'test.selection.answer'
|
|
25
|
+
TEST_SELECTION_REASON = 'test.selection.reason'
|
|
26
|
+
TEST_SELECTION_KEPT_COUNT = 'test.selection.kept_count'
|
|
27
|
+
TEST_SELECTION_NOT_APPLIED_REASON = 'test.selection.not_applied_reason'
|
|
28
|
+
|
|
29
|
+
# The largest value the engine's counters take (a signed 32-bit int).
|
|
30
|
+
MAX_COUNT = (2**31) - 1
|
|
31
|
+
|
|
14
32
|
attr_reader :token, :repo_name, :api_url, :test_run_id,
|
|
15
33
|
:recorder,
|
|
16
34
|
:branch_name,
|
|
17
|
-
:flaky_detector, :flaky_detector_error_message, :quarantined_tests
|
|
35
|
+
:flaky_detector, :flaky_detector_error_message, :quarantined_tests,
|
|
36
|
+
:test_selection, :session_verdict, :session_verdict_result,
|
|
37
|
+
:captured_session_verdict, :session_verdict_withheld
|
|
18
38
|
|
|
19
|
-
# rubocop:disable-next Metrics/MethodLength
|
|
39
|
+
# rubocop:disable-next Metrics/MethodLength,Metrics/AbcSize
|
|
20
40
|
def initialize
|
|
21
41
|
@token = ENV.fetch('MERGIFY_TOKEN', nil)
|
|
22
42
|
@repo_name = Native.detect_repository_name
|
|
@@ -28,10 +48,110 @@ module Mergify
|
|
|
28
48
|
@flaky_detector = nil
|
|
29
49
|
@flaky_detector_error_message = nil
|
|
30
50
|
@quarantined_tests = nil
|
|
51
|
+
@test_selection = nil
|
|
52
|
+
@selection_echo = nil
|
|
53
|
+
@session_verdict = SessionVerdict.new
|
|
54
|
+
@session_verdict_result = SessionVerdict::Result.new(sent: false, truncated: false)
|
|
55
|
+
@captured_session_verdict = nil
|
|
56
|
+
@session_verdict_withheld = false
|
|
31
57
|
|
|
32
58
|
setup_tracing if Utils.in_ci?
|
|
33
59
|
end
|
|
34
60
|
|
|
61
|
+
# Take the identity of what this run is about to execute, and ask
|
|
62
|
+
# Mergify whether part of it is enough.
|
|
63
|
+
#
|
|
64
|
+
# Called once RSpec has loaded the spec files and applied every filter
|
|
65
|
+
# of its own -- file and line arguments, `--tag`, `-e`, `--only-failures`
|
|
66
|
+
# -- so the fingerprint describes this process's examples and nothing
|
|
67
|
+
# else. A parallel_tests worker, or a CI job handed a slice of the spec
|
|
68
|
+
# files, collects its own slice and so reports its own identity: the
|
|
69
|
+
# engine matches each against the previous attempt of the same slice.
|
|
70
|
+
#
|
|
71
|
+
# The fingerprint is reported on every run that can take one, whether
|
|
72
|
+
# or not this job opted into the selection, as pytest-mergify does.
|
|
73
|
+
def on_examples_collected(example_ids)
|
|
74
|
+
return unless @recorder && Native.available?
|
|
75
|
+
|
|
76
|
+
fingerprint = Native.test_collection_fingerprint(example_ids)
|
|
77
|
+
@recorder.resource_attributes[TEST_COLLECTION_FINGERPRINT] = fingerprint
|
|
78
|
+
@recorder.resource_attributes[TEST_COLLECTION_COUNT] = example_ids.size
|
|
79
|
+
# A process with nothing to run has nothing to reduce, and asking would
|
|
80
|
+
# cost the job its next retry: every parallel worker left empty by a
|
|
81
|
+
# filter reports the same empty-set fingerprint under the same job, and
|
|
82
|
+
# Mergify refuses to choose between sessions it cannot tell apart.
|
|
83
|
+
load_test_selection(fingerprint) unless example_ids.empty?
|
|
84
|
+
end
|
|
85
|
+
|
|
86
|
+
# Report what Mergify answered and what this run made of it, once the
|
|
87
|
+
# answer has met the collection. Nothing is echoed unless Mergify
|
|
88
|
+
# actually answered: a run that never asked, or whose question went
|
|
89
|
+
# unanswered, was offered nothing.
|
|
90
|
+
# rubocop:disable-next Metrics/MethodLength
|
|
91
|
+
def on_selection_resolved(kept_count)
|
|
92
|
+
return unless @recorder && @test_selection&.served?
|
|
93
|
+
|
|
94
|
+
@selection_echo = {
|
|
95
|
+
'answer' => @test_selection.selection,
|
|
96
|
+
'reason' => @test_selection.reason,
|
|
97
|
+
'kept_count' => kept_count
|
|
98
|
+
}
|
|
99
|
+
not_applied = @test_selection.not_applied_reason
|
|
100
|
+
@selection_echo['not_applied_reason'] = not_applied if not_applied
|
|
101
|
+
|
|
102
|
+
resource = @recorder.resource_attributes
|
|
103
|
+
resource[TEST_SELECTION_ANSWER] = @selection_echo['answer']
|
|
104
|
+
resource[TEST_SELECTION_REASON] = @selection_echo['reason']
|
|
105
|
+
resource[TEST_SELECTION_KEPT_COUNT] = kept_count
|
|
106
|
+
resource[TEST_SELECTION_NOT_APPLIED_REASON] = not_applied if not_applied
|
|
107
|
+
end
|
|
108
|
+
|
|
109
|
+
# Write what this session concluded to Mergify, before the spans go.
|
|
110
|
+
#
|
|
111
|
+
# Sent exactly when the selection was asked for -- including when that
|
|
112
|
+
# request failed: the verdict is what the NEXT rerun of this job needs.
|
|
113
|
+
# Withheld when the run failed outside any example (a failing
|
|
114
|
+
# `after(:context)` or suite hook): its examples then read as passed,
|
|
115
|
+
# and a verdict naming no failure would have the retry run nothing and
|
|
116
|
+
# turn green on the same breakage. Without a verdict, the retry runs the
|
|
117
|
+
# whole suite. Never raises.
|
|
118
|
+
# rubocop:disable-next Metrics/MethodLength,Metrics/AbcSize,Metrics/CyclomaticComplexity,Metrics/PerceivedComplexity
|
|
119
|
+
def send_session_verdict(failed_outside_examples:)
|
|
120
|
+
return if @test_selection.nil?
|
|
121
|
+
|
|
122
|
+
if failed_outside_examples
|
|
123
|
+
@session_verdict_withheld = true
|
|
124
|
+
return
|
|
125
|
+
end
|
|
126
|
+
|
|
127
|
+
body = session_verdict_body
|
|
128
|
+
return if body.nil?
|
|
129
|
+
|
|
130
|
+
if test_mode?
|
|
131
|
+
@captured_session_verdict = body
|
|
132
|
+
@session_verdict_result = SessionVerdict::Result.new(sent: true, truncated: false)
|
|
133
|
+
return
|
|
134
|
+
end
|
|
135
|
+
|
|
136
|
+
if debug_mode?
|
|
137
|
+
puts "MERGIFY SESSION VERDICT: #{body}"
|
|
138
|
+
@session_verdict_result = SessionVerdict::Result.new(sent: true, truncated: false)
|
|
139
|
+
return
|
|
140
|
+
end
|
|
141
|
+
|
|
142
|
+
receipt = api_client.send_session_verdict(body)
|
|
143
|
+
# Dormant: the feature is not enabled for this repository, which the
|
|
144
|
+
# selection block already says.
|
|
145
|
+
return if receipt.nil?
|
|
146
|
+
|
|
147
|
+
@session_verdict_result = SessionVerdict::Result.new(sent: true, truncated: receipt['truncated'])
|
|
148
|
+
rescue StandardError => e
|
|
149
|
+
# Broader than ApiError on purpose: a value the binding cannot marshal
|
|
150
|
+
# raises something else, and a verdict must never cost the run.
|
|
151
|
+
message = e.is_a?(Native::ApiError) ? e.message : "#{e.class}: #{e.message}"
|
|
152
|
+
@session_verdict_result = SessionVerdict::Result.new(sent: false, truncated: false, error: message)
|
|
153
|
+
end
|
|
154
|
+
|
|
35
155
|
# Send the run, once, at the end. Failing to report a run is worth
|
|
36
156
|
# saying out loud but never worth failing a suite that just passed,
|
|
37
157
|
# so this answers with a message instead of raising.
|
|
@@ -56,6 +176,85 @@ module Mergify
|
|
|
56
176
|
|
|
57
177
|
private
|
|
58
178
|
|
|
179
|
+
# Opt-in, per job, and asked for only where the run's identity is
|
|
180
|
+
# complete: the head branch and revision (a merge-queue draft branch on a
|
|
181
|
+
# rerun) and the job coordinates, the very values each uploaded example
|
|
182
|
+
# carries, so the server can match its records.
|
|
183
|
+
# rubocop:disable-next Metrics/MethodLength,Metrics/AbcSize,Metrics/CyclomaticComplexity
|
|
184
|
+
def load_test_selection(fingerprint)
|
|
185
|
+
return unless TestSelection.enabled?
|
|
186
|
+
# An empty token is what a fork or a bot pull request gets for a secret.
|
|
187
|
+
return if @token.to_s.empty? || @repo_name.to_s.empty?
|
|
188
|
+
|
|
189
|
+
resource = @recorder.resource_attributes
|
|
190
|
+
branch = resource['vcs.ref.head.name']
|
|
191
|
+
head_sha = resource['vcs.ref.head.revision']
|
|
192
|
+
pipeline_name = resource['cicd.pipeline.name']
|
|
193
|
+
job_name = job_name_from(resource)
|
|
194
|
+
return if [branch, head_sha, pipeline_name, job_name].any? { |value| value.to_s.empty? }
|
|
195
|
+
|
|
196
|
+
answer = api_client.fetch_test_selection(branch.to_s, head_sha.to_s, pipeline_name.to_s,
|
|
197
|
+
job_name.to_s, fingerprint)
|
|
198
|
+
@test_selection = answer.nil? ? TestSelection.unanswered : TestSelection.served(answer)
|
|
199
|
+
rescue Native::ApiError, Utils::InvalidRepositoryFullNameError => e
|
|
200
|
+
@test_selection = TestSelection.unanswered(init_error_msg: e.message)
|
|
201
|
+
end
|
|
202
|
+
|
|
203
|
+
# `mergify.test.job.name` is the operator-set override; the provider's
|
|
204
|
+
# own task name is the fallback. Same precedence as the other clients.
|
|
205
|
+
def job_name_from(resource)
|
|
206
|
+
override = resource['mergify.test.job.name']
|
|
207
|
+
override.to_s.empty? ? resource['cicd.pipeline.task.name'] : override
|
|
208
|
+
end
|
|
209
|
+
|
|
210
|
+
# The verdict body, keyed on exactly what the selection call was keyed
|
|
211
|
+
# on, so a verdict is found by what the asking run knows.
|
|
212
|
+
# rubocop:disable-next Metrics/MethodLength,Metrics/AbcSize
|
|
213
|
+
def session_verdict_body
|
|
214
|
+
resource = @recorder.resource_attributes
|
|
215
|
+
fingerprint = resource[TEST_COLLECTION_FINGERPRINT]
|
|
216
|
+
head_sha = resource['vcs.ref.head.revision']
|
|
217
|
+
pipeline_name = resource['cicd.pipeline.name']
|
|
218
|
+
job_name = job_name_from(resource)
|
|
219
|
+
return nil if [fingerprint, head_sha, pipeline_name, job_name].any? { |value| value.to_s.empty? }
|
|
220
|
+
|
|
221
|
+
body = {
|
|
222
|
+
'test_run_id' => @test_run_id,
|
|
223
|
+
'head_sha' => head_sha.to_s,
|
|
224
|
+
'pipeline_name' => pipeline_name.to_s,
|
|
225
|
+
'job_name' => job_name.to_s,
|
|
226
|
+
'collection_fingerprint' => fingerprint,
|
|
227
|
+
'collection_count' => resource[TEST_COLLECTION_COUNT],
|
|
228
|
+
'total_test_runtime_ms' => @session_verdict.total_test_runtime_ms,
|
|
229
|
+
'failing_tests' => @session_verdict.failing_tests,
|
|
230
|
+
'quarantined_failing_tests' => @session_verdict.quarantined_failing_tests
|
|
231
|
+
}.merge(@session_verdict.counts)
|
|
232
|
+
head_branch = resource['vcs.ref.head.name']
|
|
233
|
+
body['head_branch'] = head_branch.to_s unless head_branch.to_s.empty?
|
|
234
|
+
add_run_identity(body, resource)
|
|
235
|
+
body['selection'] = @selection_echo.dup if @selection_echo
|
|
236
|
+
body
|
|
237
|
+
end
|
|
238
|
+
|
|
239
|
+
# `run_attempt` only next to a `run_id`: the engine refuses an attempt of
|
|
240
|
+
# nothing. An attempt the engine's counter cannot hold is left out
|
|
241
|
+
# rather than refused.
|
|
242
|
+
def add_run_identity(body, resource)
|
|
243
|
+
run_id = resource['cicd.pipeline.run.id']
|
|
244
|
+
return if run_id.to_s.empty?
|
|
245
|
+
|
|
246
|
+
body['run_id'] = run_id.to_s
|
|
247
|
+
attempt = resource['cicd.pipeline.run.attempt']
|
|
248
|
+
body['run_attempt'] = attempt if attempt.is_a?(Integer) && attempt.between?(0, MAX_COUNT)
|
|
249
|
+
end
|
|
250
|
+
|
|
251
|
+
def api_client
|
|
252
|
+
@api_client ||= begin
|
|
253
|
+
owner, repo = Utils.split_full_repo_name(@repo_name)
|
|
254
|
+
Native::Client.new(@api_url, @token, owner, repo, Mergify::RSpec::VERSION)
|
|
255
|
+
end
|
|
256
|
+
end
|
|
257
|
+
|
|
59
258
|
# Recording is unconditional; whether the run is *uploaded* is what the
|
|
60
259
|
# token and repository decide. Debug and test runs keep their spans and
|
|
61
260
|
# send nothing, which is what they always did -- the difference is that
|
|
@@ -1,20 +1,33 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
3
|
require 'set'
|
|
4
|
+
require_relative 'test_selection'
|
|
4
5
|
|
|
5
6
|
module Mergify
|
|
6
7
|
module RSpec
|
|
7
|
-
# Registers RSpec hooks for quarantine and flaky detection, and
|
|
8
|
+
# Registers RSpec hooks for quarantine and flaky detection, and attaches the
|
|
8
9
|
# Mergify Test Insights formatter when running inside CI.
|
|
10
|
+
# rubocop:disable-next Metrics/ModuleLength
|
|
9
11
|
module Configuration
|
|
12
|
+
QUARANTINE_MESSAGE = 'Test is quarantined from Mergify Test Insights'
|
|
13
|
+
|
|
10
14
|
module_function
|
|
11
15
|
|
|
12
|
-
# rubocop:disable Metrics/MethodLength,Metrics/BlockLength,Metrics/AbcSize
|
|
16
|
+
# rubocop:disable-next Metrics/MethodLength,Metrics/BlockLength,Metrics/AbcSize
|
|
13
17
|
# rubocop:disable-next Metrics/CyclomaticComplexity,Metrics/PerceivedComplexity
|
|
14
18
|
def setup!
|
|
15
19
|
::RSpec.configure do |config|
|
|
16
|
-
#
|
|
17
|
-
|
|
20
|
+
# Registered before the formatter's hook so that it runs after it --
|
|
21
|
+
# each `prepend_before` goes to the front: the formatter has to be
|
|
22
|
+
# attached already for the run's report and verdict to see what the
|
|
23
|
+
# selection did.
|
|
24
|
+
config.prepend_before(:suite) { Configuration.select_examples } if Utils.in_ci?
|
|
25
|
+
|
|
26
|
+
# Attached to the reporter once the suite starts, rather than added with
|
|
27
|
+
# `add_formatter`: RSpec only sets up its default formatter when no
|
|
28
|
+
# other was added, so adding this one took the progress output and the
|
|
29
|
+
# summary line away from every suite that had not chosen a format.
|
|
30
|
+
config.prepend_before(:suite) { Configuration.attach_formatter(config) } if Utils.in_ci?
|
|
18
31
|
|
|
19
32
|
# Flaky detection: prepare session with all example IDs
|
|
20
33
|
config.before(:suite) do
|
|
@@ -37,81 +50,163 @@ module Mergify
|
|
|
37
50
|
|
|
38
51
|
# Flaky detection: rerun tests within budget
|
|
39
52
|
config.around(:each) do |example|
|
|
40
|
-
|
|
41
|
-
fd
|
|
42
|
-
|
|
43
|
-
example.run
|
|
44
|
-
|
|
45
|
-
# Feed metrics from the initial run so the detector can evaluate
|
|
46
|
-
if fd
|
|
47
|
-
run_time = example.execution_result.run_time || 0.0
|
|
48
|
-
status = example.execution_result.status
|
|
49
|
-
fd.fill_metrics_from_report(example.id, 'setup', 0.0, status)
|
|
50
|
-
fd.fill_metrics_from_report(example.id, 'call', run_time, status)
|
|
51
|
-
fd.fill_metrics_from_report(example.id, 'teardown', 0.0, status)
|
|
52
|
-
end
|
|
53
|
-
|
|
54
|
-
next unless fd&.rerunning_test?(example.id)
|
|
53
|
+
fd = Mergify::RSpec.ci_insights&.flaky_detector
|
|
54
|
+
fd ? Configuration.run_detecting_flakiness(example, fd) : example.run
|
|
55
|
+
end
|
|
55
56
|
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
57
|
+
# Quarantine: override failed quarantined test results. The failure is
|
|
58
|
+
# noted for flaky detection first, which would otherwise take a
|
|
59
|
+
# quarantined attempt that failed for one that passed.
|
|
60
|
+
config.after(:each) do |example|
|
|
61
|
+
next unless example.metadata[:mergify_quarantined] && example.exception
|
|
59
62
|
|
|
60
|
-
|
|
61
|
-
|
|
63
|
+
example.metadata[:mergify_quarantined_failure] = true
|
|
64
|
+
example.instance_variable_set(:@exception, nil)
|
|
65
|
+
example.execution_result.status = :pending
|
|
66
|
+
example.execution_result.pending_message = QUARANTINE_MESSAGE
|
|
67
|
+
end
|
|
68
|
+
end
|
|
69
|
+
end
|
|
62
70
|
|
|
63
|
-
|
|
64
|
-
|
|
71
|
+
# Fingerprint what this process is about to run, ask Mergify whether part
|
|
72
|
+
# of it is enough, and remove the rest.
|
|
73
|
+
#
|
|
74
|
+
# A `before(:suite)` hook, because that is the first moment RSpec has both
|
|
75
|
+
# loaded the spec files and applied its filters, and the last one before
|
|
76
|
+
# an example group starts. What each group runs is read from
|
|
77
|
+
# `RSpec.world.filtered_examples`, which RSpec memoized while counting the
|
|
78
|
+
# run; narrowing those lists is what deselects.
|
|
79
|
+
# rubocop:disable-next Metrics/MethodLength,Metrics/AbcSize,Metrics/CyclomaticComplexity,Metrics/PerceivedComplexity
|
|
80
|
+
def select_examples
|
|
81
|
+
ci = Mergify::RSpec.ci_insights
|
|
82
|
+
return unless ci
|
|
83
|
+
|
|
84
|
+
groups = ::RSpec.world.example_groups.flat_map(&:descendants)
|
|
85
|
+
ids = groups.flat_map { |group| ::RSpec.world.filtered_examples[group] }.map(&:id)
|
|
86
|
+
ci.on_examples_collected(ids)
|
|
87
|
+
|
|
88
|
+
begin
|
|
89
|
+
keep = ci.test_selection&.resolve(ids)
|
|
90
|
+
rescue TestSelectionRefused => e
|
|
91
|
+
ci.on_selection_resolved(0)
|
|
92
|
+
return refuse(e.message)
|
|
93
|
+
end
|
|
65
94
|
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
95
|
+
deselect(groups, keep) if keep
|
|
96
|
+
ci.on_selection_resolved(keep ? ids.count { |id| keep.include?(id) } : ids.size)
|
|
97
|
+
rescue StandardError => e
|
|
98
|
+
# This runs on every CI run, opted in or not, and RSpec would answer a
|
|
99
|
+
# raise here by running nothing: a fault of this gem must cost the run
|
|
100
|
+
# its reduction, never its tests.
|
|
101
|
+
::RSpec.configuration.error_stream.puts(
|
|
102
|
+
"Mergify Test Selection failed, so the full suite runs: #{e.class}: #{e.message}"
|
|
103
|
+
)
|
|
104
|
+
end
|
|
70
105
|
|
|
71
|
-
|
|
72
|
-
|
|
106
|
+
# Stop the run before any example, the way RSpec stops one whose suite
|
|
107
|
+
# hook failed -- nothing runs, the exit code is red, and the summary
|
|
108
|
+
# counts an error outside of examples -- but with the server's
|
|
109
|
+
# explanation alone. Raised from the hook, the same failure is printed
|
|
110
|
+
# under a "Failure/Error:" header pointing at a line of this gem, which
|
|
111
|
+
# only buries the explanation.
|
|
112
|
+
def refuse(message)
|
|
113
|
+
::RSpec.configuration.error_stream.puts("\n#{message}\n")
|
|
114
|
+
::RSpec.world.wants_to_quit = true
|
|
115
|
+
reporter = ::RSpec.configuration.reporter
|
|
116
|
+
count = reporter.instance_variable_get(:@non_example_exception_count)
|
|
117
|
+
reporter.instance_variable_set(:@non_example_exception_count, count + 1) if count.is_a?(Integer)
|
|
118
|
+
end
|
|
73
119
|
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
inner = memoized.instance_variable_get(:@memoized)
|
|
83
|
-
inner&.clear
|
|
84
|
-
end
|
|
85
|
-
end
|
|
120
|
+
# A group left with no example runs none of its `before(:context)`
|
|
121
|
+
# hooks: RSpec decides that from these same lists, once the group starts.
|
|
122
|
+
def deselect(groups, keep)
|
|
123
|
+
groups.each do |group|
|
|
124
|
+
::RSpec.world.filtered_examples[group] =
|
|
125
|
+
::RSpec.world.filtered_examples[group].select { |example| keep.include?(example.id) }
|
|
126
|
+
end
|
|
127
|
+
end
|
|
86
128
|
|
|
87
|
-
|
|
129
|
+
# The reporter has sent `start` by the time suite hooks run, so the
|
|
130
|
+
# formatter is handed the same notification directly.
|
|
131
|
+
def attach_formatter(config)
|
|
132
|
+
formatter = Formatter.new(config.output_stream)
|
|
133
|
+
formatter.start(::RSpec::Core::Notifications::StartNotification.new(::RSpec.world.example_count, 0))
|
|
134
|
+
config.reporter.register_listener(formatter, *Formatter::NOTIFICATIONS)
|
|
135
|
+
end
|
|
88
136
|
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
137
|
+
# `example` is the Procsy an around hook is handed, and the example has no
|
|
138
|
+
# result yet: RSpec records its status and run time in `Example#finish`,
|
|
139
|
+
# once every around hook has returned. So each attempt is timed here, and
|
|
140
|
+
# its outcome read off the exception it left behind.
|
|
141
|
+
# rubocop:disable-next Metrics/AbcSize
|
|
142
|
+
def run_detecting_flakiness(example, detector)
|
|
143
|
+
status, duration = run_attempt(example)
|
|
144
|
+
detector.fill_metrics_from_report(example.id, 'setup', 0.0, status)
|
|
145
|
+
detector.fill_metrics_from_report(example.id, 'call', duration, status)
|
|
146
|
+
detector.fill_metrics_from_report(example.id, 'teardown', 0.0, status)
|
|
147
|
+
return unless detector.rerunning_test?(example.id)
|
|
148
|
+
|
|
149
|
+
# Mark as flaky detection candidate (even if too slow to rerun)
|
|
150
|
+
example.metadata[:mergify_flaky_detection] = true
|
|
151
|
+
example.metadata[:mergify_new_test] = true if detector.mode == 'new'
|
|
152
|
+
|
|
153
|
+
detector.set_test_deadline(example.id)
|
|
154
|
+
return if detector.test_too_slow?(example.id)
|
|
155
|
+
|
|
156
|
+
rerun(example, detector, status)
|
|
157
|
+
end
|
|
93
158
|
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
159
|
+
# Which attempt decides the result is pytest-mergify's rule: in 'new' mode
|
|
160
|
+
# any failure fails the test, so a new flaky test cannot be merged; in
|
|
161
|
+
# 'unhealthy' mode the reruns only learn, and the first attempt stands.
|
|
162
|
+
# rubocop:disable-next Metrics/MethodLength,Metrics/AbcSize
|
|
163
|
+
def rerun(example, detector, initial_status)
|
|
164
|
+
initial_exception = example.exception
|
|
165
|
+
first_failure = initial_exception
|
|
166
|
+
outcomes = Set[initial_status]
|
|
167
|
+
rerun_count = 0
|
|
168
|
+
|
|
169
|
+
until example.metadata[:is_last_rerun]
|
|
170
|
+
example.metadata[:is_last_rerun] = detector.last_rerun_for_test?(example.id)
|
|
171
|
+
reset_for_rerun(example)
|
|
172
|
+
|
|
173
|
+
status, duration = run_attempt(example)
|
|
174
|
+
detector.fill_metrics_from_report(example.id, 'call', duration, status)
|
|
175
|
+
outcomes << status
|
|
176
|
+
first_failure ||= example.exception
|
|
177
|
+
rerun_count += 1
|
|
178
|
+
end
|
|
97
179
|
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
180
|
+
example.metadata[:mergify_flaky] = true if outcomes.include?(:passed) && outcomes.include?(:failed)
|
|
181
|
+
example.metadata[:mergify_rerun_count] = rerun_count
|
|
182
|
+
final_exception = detector.mode == 'new' ? first_failure : initial_exception
|
|
183
|
+
example.example.instance_variable_set(:@exception, final_exception)
|
|
184
|
+
end
|
|
103
185
|
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
186
|
+
# Runs what the around hook wraps -- the before and after hooks included
|
|
187
|
+
# -- and answers its outcome and duration in seconds.
|
|
188
|
+
def run_attempt(example)
|
|
189
|
+
started = Process.clock_gettime(Process::CLOCK_MONOTONIC)
|
|
190
|
+
example.run
|
|
191
|
+
duration = Process.clock_gettime(Process::CLOCK_MONOTONIC) - started
|
|
192
|
+
failed = example.exception || example.metadata.delete(:mergify_quarantined_failure)
|
|
193
|
+
[failed ? :failed : :passed, duration]
|
|
194
|
+
end
|
|
107
195
|
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
196
|
+
# The exception lives on the example itself, not on the Procsy: setting
|
|
197
|
+
# it there would leave RSpec to fold every attempt's failure into one.
|
|
198
|
+
def reset_for_rerun(example)
|
|
199
|
+
example.example.instance_variable_set(:@exception, nil)
|
|
200
|
+
return unless example.example_group_instance
|
|
201
|
+
|
|
202
|
+
memoized = example.example_group_instance.instance_variable_get(:@__memoized)
|
|
203
|
+
if memoized.respond_to?(:clear)
|
|
204
|
+
memoized.clear
|
|
205
|
+
elsif memoized
|
|
206
|
+
# RSpec >= 3.12 uses ThreadsafeMemoized which wraps an internal hash
|
|
207
|
+
memoized.instance_variable_get(:@memoized)&.clear
|
|
112
208
|
end
|
|
113
209
|
end
|
|
114
|
-
# rubocop:enable Metrics/MethodLength,Metrics/BlockLength,Metrics/AbcSize
|
|
115
210
|
end
|
|
116
211
|
end
|
|
117
212
|
end
|
|
@@ -146,11 +146,14 @@ module Mergify
|
|
|
146
146
|
(metrics.initial_duration * min_exec) > metrics.remaining_time
|
|
147
147
|
end
|
|
148
148
|
|
|
149
|
+
# Asked before the rerun it guards. `rerun_count` counts executions, the
|
|
150
|
+
# initial one included, so the rerun about to start is the last one when
|
|
151
|
+
# it is the one reaching the cap.
|
|
149
152
|
def last_rerun_for_test?(test_id)
|
|
150
153
|
return false unless @metrics.key?(test_id)
|
|
151
154
|
|
|
152
155
|
metrics = @metrics[test_id]
|
|
153
|
-
metrics.will_exceed_deadline? || metrics.rerun_count >= @context['max_test_execution_count']
|
|
156
|
+
metrics.will_exceed_deadline? || metrics.rerun_count + 1 >= @context['max_test_execution_count']
|
|
154
157
|
end
|
|
155
158
|
|
|
156
159
|
def test_metrics(test_id)
|
|
@@ -2,6 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
require 'rspec/core/formatters/base_formatter'
|
|
4
4
|
require_relative 'trace'
|
|
5
|
+
require_relative 'configuration'
|
|
5
6
|
|
|
6
7
|
require 'mergify/rspec/native'
|
|
7
8
|
|
|
@@ -12,12 +13,17 @@ module Mergify
|
|
|
12
13
|
# test execution.
|
|
13
14
|
# rubocop:disable-next Metrics/ClassLength
|
|
14
15
|
class Formatter < ::RSpec::Core::Formatters::BaseFormatter
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
16
|
+
# What it listens to once attached, `start` excepted: it is attached after
|
|
17
|
+
# that notification, and handed it directly.
|
|
18
|
+
NOTIFICATIONS = %i[example_started example_finished example_pending stop].freeze
|
|
19
|
+
|
|
20
|
+
# @mergifyio/vitest's sentence, for the same withheld verdict.
|
|
21
|
+
WITHHELD_VERDICT = TestSelection.wrap(
|
|
22
|
+
"Mergify wasn't sent this run's results: it failed outside any example. If this merge-queue " \
|
|
23
|
+
'batch is retried, this job will run its full test suite.'
|
|
24
|
+
)
|
|
25
|
+
|
|
26
|
+
::RSpec::Core::Formatters.register self, :start, *NOTIFICATIONS
|
|
21
27
|
|
|
22
28
|
def start(notification)
|
|
23
29
|
super
|
|
@@ -47,7 +53,7 @@ module Mergify
|
|
|
47
53
|
@example_spans[example.id] = span
|
|
48
54
|
end
|
|
49
55
|
|
|
50
|
-
# rubocop:disable-next Metrics/MethodLength
|
|
56
|
+
# rubocop:disable-next Metrics/MethodLength,Metrics/AbcSize
|
|
51
57
|
def example_finished(notification)
|
|
52
58
|
return unless @example_spans
|
|
53
59
|
|
|
@@ -56,6 +62,7 @@ module Mergify
|
|
|
56
62
|
return unless span
|
|
57
63
|
|
|
58
64
|
result = example.execution_result
|
|
65
|
+
@ci_insights.session_verdict.record(example.id, verdict_status(example), result.run_time)
|
|
59
66
|
status = result.status.to_s
|
|
60
67
|
span.set_attribute('test.case.result.status', status)
|
|
61
68
|
set_flaky_attributes(span, example)
|
|
@@ -78,17 +85,50 @@ module Mergify
|
|
|
78
85
|
return unless span
|
|
79
86
|
|
|
80
87
|
span.set_attribute('test.case.result.status', 'skipped')
|
|
88
|
+
@ci_insights.session_verdict.record(example.id, verdict_status(example), example.execution_result.run_time)
|
|
81
89
|
@ci_insights.recorder.record(span)
|
|
82
90
|
end
|
|
83
91
|
|
|
84
92
|
def stop(_notification)
|
|
85
93
|
finish_session_span
|
|
94
|
+
# The verdict first, on purpose: it is what the next merge-queue rerun
|
|
95
|
+
# of this job is answered from, and it must never wait behind the
|
|
96
|
+
# trace upload's timeout and retries.
|
|
97
|
+
send_session_verdict
|
|
86
98
|
print_report
|
|
87
99
|
flush_and_shutdown
|
|
88
100
|
end
|
|
89
101
|
|
|
90
102
|
private
|
|
91
103
|
|
|
104
|
+
# The status RSpec reported, which is the one that decided the exit
|
|
105
|
+
# code. A pending example is a skip, unless it is a failure the
|
|
106
|
+
# quarantine absorbed.
|
|
107
|
+
def verdict_status(example)
|
|
108
|
+
result = example.execution_result
|
|
109
|
+
case result.status
|
|
110
|
+
when :passed then 'passed'
|
|
111
|
+
when :failed then 'failed'
|
|
112
|
+
else
|
|
113
|
+
quarantined = example.metadata[:mergify_quarantined] &&
|
|
114
|
+
result.pending_message == Configuration::QUARANTINE_MESSAGE
|
|
115
|
+
quarantined ? 'quarantined_failed' : 'skipped'
|
|
116
|
+
end
|
|
117
|
+
end
|
|
118
|
+
|
|
119
|
+
def send_session_verdict
|
|
120
|
+
return unless @ci_insights&.recorder
|
|
121
|
+
|
|
122
|
+
@ci_insights.send_session_verdict(failed_outside_examples: failed_outside_examples?)
|
|
123
|
+
end
|
|
124
|
+
|
|
125
|
+
# RSpec flags a failure no example carries -- a failing `after(:context)`
|
|
126
|
+
# or suite hook. A load error flags it too, but stops the run before the
|
|
127
|
+
# formatter is ever attached.
|
|
128
|
+
def failed_outside_examples?
|
|
129
|
+
::RSpec.world.non_example_failure ? true : false
|
|
130
|
+
end
|
|
131
|
+
|
|
92
132
|
def build_example_attributes(example, quarantined)
|
|
93
133
|
{
|
|
94
134
|
'test.scope' => 'case',
|
|
@@ -149,6 +189,7 @@ module Mergify
|
|
|
149
189
|
print_configuration_warnings
|
|
150
190
|
print_flaky_report
|
|
151
191
|
print_quarantine_report
|
|
192
|
+
print_test_selection_report
|
|
152
193
|
output.puts "MERGIFY_TEST_RUN_ID=#{@ci_insights.test_run_id}"
|
|
153
194
|
output.puts '------------------'
|
|
154
195
|
end
|
|
@@ -185,6 +226,17 @@ module Mergify
|
|
|
185
226
|
output.puts report if report
|
|
186
227
|
end
|
|
187
228
|
|
|
229
|
+
# One block whatever happened, once the job asked for a selection.
|
|
230
|
+
def print_test_selection_report
|
|
231
|
+
selection = @ci_insights.test_selection
|
|
232
|
+
return unless selection
|
|
233
|
+
|
|
234
|
+
output.puts selection.report
|
|
235
|
+
output.puts WITHHELD_VERDICT if @ci_insights.session_verdict_withheld
|
|
236
|
+
verdict_report = @ci_insights.session_verdict_result.report
|
|
237
|
+
output.puts verdict_report if verdict_report
|
|
238
|
+
end
|
|
239
|
+
|
|
188
240
|
# One upload, at the end. There is nothing to shut down any more: the
|
|
189
241
|
# recorder holds spans in memory and the client owns the connection.
|
|
190
242
|
def flush_and_shutdown
|
|
@@ -48,13 +48,19 @@ module Mergify
|
|
|
48
48
|
lines << " Quarantined tests run (#{used.size}):"
|
|
49
49
|
used.each { |t| lines << " - #{t}" }
|
|
50
50
|
lines << ''
|
|
51
|
-
lines << "
|
|
51
|
+
lines << " #{unused_label} (#{unused.size}):"
|
|
52
52
|
unused.each { |t| lines << " - #{t}" }
|
|
53
53
|
lines.join("\n")
|
|
54
54
|
end
|
|
55
55
|
|
|
56
56
|
private
|
|
57
57
|
|
|
58
|
+
# A parallel worker runs a share of the suite, so a quarantined test it
|
|
59
|
+
# did not run was most likely run by another worker, not left unused.
|
|
60
|
+
def unused_label
|
|
61
|
+
Utils.parallel_worker? ? 'Quarantined tests not run by this worker' : 'Unused quarantined tests'
|
|
62
|
+
end
|
|
63
|
+
|
|
58
64
|
# A nil list means the repository has no quarantine subscription, which is
|
|
59
65
|
# not an error: the session simply quarantines nothing. Anything that went
|
|
60
66
|
# genuinely wrong is recorded and the suite carries on -- this plugin has
|
|
@@ -0,0 +1,86 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require_relative 'test_selection'
|
|
4
|
+
|
|
5
|
+
module Mergify
|
|
6
|
+
module RSpec
|
|
7
|
+
# What the session concluded about each example, folded as RSpec reports
|
|
8
|
+
# it, and written to Mergify when the session ends.
|
|
9
|
+
#
|
|
10
|
+
# Test Selection answers a merge-queue rerun from its predecessor's
|
|
11
|
+
# verdict: which examples failed, whether anything ran at all. The verdict
|
|
12
|
+
# is sent in one request before the trace upload, so the answer is there
|
|
13
|
+
# seconds after the session ends rather than whenever trace ingestion gets
|
|
14
|
+
# to it (INC-2434).
|
|
15
|
+
#
|
|
16
|
+
# One status per example, its FINAL one -- the status RSpec reported once
|
|
17
|
+
# flaky detection's reruns were over, which is what decided the exit code.
|
|
18
|
+
# A verdict that disagreed with the exit code would either replay examples
|
|
19
|
+
# that did not gate, or skip the one that did.
|
|
20
|
+
class SessionVerdict
|
|
21
|
+
# A failure the quarantine absorbed is `quarantined_failed`: it did not
|
|
22
|
+
# gate the job, so a rerun must not replay it, but it did run and fail,
|
|
23
|
+
# so it is not green either.
|
|
24
|
+
PRECEDENCE = { 'passed' => 0, 'skipped' => 1, 'quarantined_failed' => 2, 'failed' => 3 }.freeze
|
|
25
|
+
|
|
26
|
+
# How sending went, for the report. `sent` is false both when nothing
|
|
27
|
+
# had to be sent and when the request failed; `error` tells them apart.
|
|
28
|
+
Result = Struct.new(:sent, :truncated, :error, keyword_init: true) do
|
|
29
|
+
# Said only when the verdict did not reach Mergify whole. Wording
|
|
30
|
+
# validated by Alexandre on 2026-09-15 (MRGFY-9313), pytest-mergify's
|
|
31
|
+
# verbatim; a change here is a product decision.
|
|
32
|
+
def report
|
|
33
|
+
if error
|
|
34
|
+
"#{TestSelection.wrap("Mergify couldn't record this run's results. If this merge-queue batch is " \
|
|
35
|
+
'retried, this job will run its full test suite.')}\nError: #{error}\n"
|
|
36
|
+
elsif truncated
|
|
37
|
+
"#{TestSelection.wrap("Mergify recorded this run's counts but not its failing tests. If this " \
|
|
38
|
+
'merge-queue batch is retried, this job will run its full test suite.')}\n"
|
|
39
|
+
end
|
|
40
|
+
end
|
|
41
|
+
end
|
|
42
|
+
|
|
43
|
+
def initialize
|
|
44
|
+
@final = {}
|
|
45
|
+
@runtime_seconds = 0.0
|
|
46
|
+
end
|
|
47
|
+
|
|
48
|
+
def record(example_id, status, run_time)
|
|
49
|
+
@runtime_seconds += run_time.to_f
|
|
50
|
+
previous = @final[example_id]
|
|
51
|
+
@final[example_id] = status if previous.nil? || PRECEDENCE.fetch(status) > PRECEDENCE.fetch(previous)
|
|
52
|
+
end
|
|
53
|
+
|
|
54
|
+
def total_test_runtime_ms
|
|
55
|
+
(@runtime_seconds * 1000).to_i
|
|
56
|
+
end
|
|
57
|
+
|
|
58
|
+
# The engine's own definitions: `executed` counts every example that
|
|
59
|
+
# reached a status, a skipped one included, and `failed` includes the
|
|
60
|
+
# quarantined failures.
|
|
61
|
+
def counts
|
|
62
|
+
statuses = @final.values
|
|
63
|
+
{
|
|
64
|
+
'executed_count' => statuses.size,
|
|
65
|
+
'passed_count' => statuses.count('passed'),
|
|
66
|
+
'failed_count' => statuses.count { |status| %w[failed quarantined_failed].include?(status) },
|
|
67
|
+
'skipped_count' => statuses.count('skipped')
|
|
68
|
+
}
|
|
69
|
+
end
|
|
70
|
+
|
|
71
|
+
def failing_tests
|
|
72
|
+
ids_with('failed')
|
|
73
|
+
end
|
|
74
|
+
|
|
75
|
+
def quarantined_failing_tests
|
|
76
|
+
ids_with('quarantined_failed')
|
|
77
|
+
end
|
|
78
|
+
|
|
79
|
+
private
|
|
80
|
+
|
|
81
|
+
def ids_with(status)
|
|
82
|
+
@final.filter_map { |id, final| id if final == status }
|
|
83
|
+
end
|
|
84
|
+
end
|
|
85
|
+
end
|
|
86
|
+
end
|
|
@@ -0,0 +1,299 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require 'set'
|
|
4
|
+
require_relative 'utils'
|
|
5
|
+
|
|
6
|
+
module Mergify
|
|
7
|
+
module RSpec
|
|
8
|
+
# Raised by `TestSelection#resolve` when Mergify refuses to choose a
|
|
9
|
+
# selection for this run, carrying the explanation to show. The caller
|
|
10
|
+
# stops the run with it.
|
|
11
|
+
class TestSelectionRefused < StandardError; end
|
|
12
|
+
|
|
13
|
+
# Whether this run should execute only part of what it collected.
|
|
14
|
+
#
|
|
15
|
+
# A merge-queue rerun only needs to replay the examples that failed on the
|
|
16
|
+
# previous attempt of the same job. Mergify decides that server-side from
|
|
17
|
+
# the run's own identity AND from the fingerprint of what the run
|
|
18
|
+
# collected, which is why the request leaves only once RSpec has loaded
|
|
19
|
+
# and filtered its examples. The answer is one of:
|
|
20
|
+
#
|
|
21
|
+
# * `full` -- run everything;
|
|
22
|
+
# * `subset` -- run only `tests`;
|
|
23
|
+
# * `empty` -- run nothing: the previous attempt ran these examples and
|
|
24
|
+
# they passed, so the run exits green having executed none;
|
|
25
|
+
# * `refused` -- Mergify holds several candidate sessions for this job and
|
|
26
|
+
# will not guess between them, so the run FAILS with its explanation.
|
|
27
|
+
#
|
|
28
|
+
# Every error, and every answer outside that list, runs the whole suite:
|
|
29
|
+
# the feature can remove work, never correctness. pytest-mergify's
|
|
30
|
+
# `TestSelection`, answer for answer and sentence for sentence.
|
|
31
|
+
#
|
|
32
|
+
# When an answer cannot be honoured the run executes everything and says
|
|
33
|
+
# so in `not_applied_reason`, never by rewriting `selection` or `reason`:
|
|
34
|
+
# those two carry Mergify's word.
|
|
35
|
+
# rubocop:disable-next Metrics/ClassLength
|
|
36
|
+
class TestSelection
|
|
37
|
+
ENABLE_ENV = 'MERGIFY_TEST_SELECTION_ENABLE'
|
|
38
|
+
|
|
39
|
+
CLIENT_NAME = 'rspec-mergify'
|
|
40
|
+
DOCS_URL = 'https://docs.mergify.com/ci-insights/test-frameworks/rspec/'
|
|
41
|
+
|
|
42
|
+
# What a refusal says when the server sent no wording of its own. The
|
|
43
|
+
# copy belongs to the server; this is a fallback, not the message.
|
|
44
|
+
FALLBACK_REFUSAL_MESSAGE = <<~MESSAGE.chomp
|
|
45
|
+
Mergify Test Selection stopped this run.
|
|
46
|
+
|
|
47
|
+
Several runs of this job report to Mergify under the same name, and they run the same tests — so Mergify cannot tell which one this run repeats, and it will not guess which tests to skip.
|
|
48
|
+
|
|
49
|
+
If this job runs more than once (a build matrix, for example), give each run its own name with MERGIFY_TEST_JOB_NAME:
|
|
50
|
+
#{DOCS_URL}
|
|
51
|
+
|
|
52
|
+
If this job only runs once, this is unexpected — please contact Mergify support.
|
|
53
|
+
MESSAGE
|
|
54
|
+
|
|
55
|
+
NOT_APPLIED_SENTENCE =
|
|
56
|
+
"Mergify's answer didn't match the tests this run collected, so the full suite ran."
|
|
57
|
+
|
|
58
|
+
# One sentence per reason the full suite ran. pytest-mergify's
|
|
59
|
+
# `_FULL_RUN_SENTENCES` verbatim (validated by Alexandre on 2026-09-11,
|
|
60
|
+
# MRGFY-8978), naming this gem where they name a client. A change here is
|
|
61
|
+
# a product decision.
|
|
62
|
+
FULL_RUN_SENTENCES = {
|
|
63
|
+
'no_predecessor' => 'First attempt of this batch, so the full suite ran.',
|
|
64
|
+
'not_a_merge_queue_run' => "This job isn't part of a merge queue run, so the full suite ran.",
|
|
65
|
+
'stale_run' => 'The batch branch was updated while this job was running, so the full suite ran.',
|
|
66
|
+
'no_matching_test_session' =>
|
|
67
|
+
"The previous attempt didn't run this exact set of tests, so the full suite ran.",
|
|
68
|
+
'matched_test_session_ran_no_test' => 'The previous attempt executed no tests, so the full suite ran.',
|
|
69
|
+
'predecessor_unknown' => "Mergify couldn't tell which previous run to start from, so the full suite ran.",
|
|
70
|
+
'indeterminate_test_session' =>
|
|
71
|
+
"Mergify couldn't tell which previous run to start from, so the full suite ran.",
|
|
72
|
+
'matched_test_session_partially_processed' =>
|
|
73
|
+
"Mergify didn't have the complete results of the previous attempt, so the full suite ran.",
|
|
74
|
+
'matched_test_session_dropped_cases' =>
|
|
75
|
+
"Mergify didn't have the complete results of the previous attempt, so the full suite ran.",
|
|
76
|
+
'matched_test_session_declaration_unreadable' =>
|
|
77
|
+
"Mergify didn't have the complete results of the previous attempt, so the full suite ran.",
|
|
78
|
+
'matched_test_session_incomplete' =>
|
|
79
|
+
'The previous attempt stopped before running all of its tests, so the full suite ran.',
|
|
80
|
+
'matched_test_session_failures_truncated' =>
|
|
81
|
+
'The previous attempt had too many failures for Mergify to list, so the full suite ran.',
|
|
82
|
+
'no_collection_fingerprint' =>
|
|
83
|
+
"This version of #{CLIENT_NAME} doesn't report what it collected, so the full suite ran. " \
|
|
84
|
+
'Upgrade it to let Mergify reduce reruns.',
|
|
85
|
+
'feature_disabled' => "Test selection isn't enabled for this organization yet, so the full suite ran.",
|
|
86
|
+
'not_requested' => "Test selection isn't available for this repository, so the full suite ran.",
|
|
87
|
+
'unrecognised_selection' =>
|
|
88
|
+
"Mergify answered in a way this version of #{CLIENT_NAME} doesn't understand, so the full suite " \
|
|
89
|
+
'ran. Upgrade it to let Mergify reduce reruns.',
|
|
90
|
+
'subset_served_without_tests' => NOT_APPLIED_SENTENCE,
|
|
91
|
+
'subset_matched_no_collected_test' => NOT_APPLIED_SENTENCE,
|
|
92
|
+
'subset_partly_absent_from_collection' => NOT_APPLIED_SENTENCE
|
|
93
|
+
}.freeze
|
|
94
|
+
|
|
95
|
+
# A newer engine may serve a reason this gem predates. The block must
|
|
96
|
+
# still read as a full run, and must never show the raw identifier.
|
|
97
|
+
UNKNOWN_REASON_SENTENCE = 'Mergify served the full suite.'
|
|
98
|
+
|
|
99
|
+
HEADER = '✂️ Test selection'
|
|
100
|
+
|
|
101
|
+
# How many re-executed examples the subset block lists before counting
|
|
102
|
+
# the rest.
|
|
103
|
+
LISTED_TESTS_MAX = 10
|
|
104
|
+
|
|
105
|
+
# The width the wording was validated at.
|
|
106
|
+
WRAP_WIDTH = 80
|
|
107
|
+
|
|
108
|
+
KNOWN_ANSWERS = %w[full subset empty refused].freeze
|
|
109
|
+
|
|
110
|
+
attr_reader :selection, :reason, :tests, :message, :init_error_msg,
|
|
111
|
+
:not_applied_reason, :kept_count, :deselected_count, :kept_tests
|
|
112
|
+
|
|
113
|
+
class << self
|
|
114
|
+
# Whether this job asked for test selection. Opt-in and per job: a job
|
|
115
|
+
# that has not opted in asks nothing, which is what lets Mergify tell
|
|
116
|
+
# a repository that never opted in from one that did (MRGFY-9172).
|
|
117
|
+
# Trimmed, and anything unrecognised is off -- the direction that runs
|
|
118
|
+
# the whole suite.
|
|
119
|
+
def enabled?
|
|
120
|
+
Utils::TRUTHY_STRINGS.include?(ENV.fetch(ENABLE_ENV, '').strip.downcase)
|
|
121
|
+
end
|
|
122
|
+
|
|
123
|
+
# What Mergify answered, as the binding hands it over.
|
|
124
|
+
def served(answer)
|
|
125
|
+
new(selection: answer['selection'], reason: answer['reason'], tests: answer['tests'] || [],
|
|
126
|
+
message: answer['message'], served: true)
|
|
127
|
+
end
|
|
128
|
+
|
|
129
|
+
# A run Mergify did not answer: the repository has no such feature, or
|
|
130
|
+
# the request failed. Both run everything, and neither was offered
|
|
131
|
+
# anything, so neither is echoed as an answer.
|
|
132
|
+
def unanswered(init_error_msg: nil)
|
|
133
|
+
new(selection: 'full', reason: 'not_requested', tests: [], message: nil,
|
|
134
|
+
served: false, init_error_msg: init_error_msg)
|
|
135
|
+
end
|
|
136
|
+
end
|
|
137
|
+
|
|
138
|
+
# rubocop:disable-next Metrics/ParameterLists,Metrics/MethodLength
|
|
139
|
+
def initialize(selection:, reason:, tests:, message:, served:, init_error_msg: nil)
|
|
140
|
+
@selection = selection
|
|
141
|
+
@reason = reason
|
|
142
|
+
@tests = tests
|
|
143
|
+
@message = message
|
|
144
|
+
@served = served
|
|
145
|
+
@init_error_msg = init_error_msg
|
|
146
|
+
@not_applied_reason = nil
|
|
147
|
+
@kept_count = nil
|
|
148
|
+
@deselected_count = 0
|
|
149
|
+
@kept_tests = []
|
|
150
|
+
|
|
151
|
+
# A subset is only honoured with a non-empty list, and an answer this
|
|
152
|
+
# gem predates is never acted on: both run everything, and say so.
|
|
153
|
+
if @selection == 'subset'
|
|
154
|
+
@not_applied_reason = 'subset_served_without_tests' if @tests.empty?
|
|
155
|
+
else
|
|
156
|
+
@not_applied_reason = 'unrecognised_selection' unless KNOWN_ANSWERS.include?(@selection)
|
|
157
|
+
@tests = []
|
|
158
|
+
end
|
|
159
|
+
end
|
|
160
|
+
|
|
161
|
+
def served?
|
|
162
|
+
@served
|
|
163
|
+
end
|
|
164
|
+
|
|
165
|
+
def refused?
|
|
166
|
+
@not_applied_reason.nil? && @selection == 'refused'
|
|
167
|
+
end
|
|
168
|
+
|
|
169
|
+
# Decide what the answer leaves of this collection to run: the ids to
|
|
170
|
+
# keep, or nil for all of them. Raises `TestSelectionRefused` on a
|
|
171
|
+
# refusal, carrying the server's explanation.
|
|
172
|
+
#
|
|
173
|
+
# Matching is by exact example id, the identity this gem uploads. A
|
|
174
|
+
# subset is honoured all or not at all: one served id this run did not
|
|
175
|
+
# collect declines the whole answer and runs everything, rather than a
|
|
176
|
+
# reduced run over an arbitrary part of what was asked for.
|
|
177
|
+
# rubocop:disable-next Metrics/MethodLength,Metrics/AbcSize,Metrics/CyclomaticComplexity,Metrics/PerceivedComplexity
|
|
178
|
+
def resolve(ids)
|
|
179
|
+
return nil unless @not_applied_reason.nil?
|
|
180
|
+
|
|
181
|
+
raise TestSelectionRefused, @message.to_s.empty? ? FALLBACK_REFUSAL_MESSAGE : @message if refused?
|
|
182
|
+
|
|
183
|
+
if @selection == 'empty'
|
|
184
|
+
# A collection already empty is left alone: the run is then empty for
|
|
185
|
+
# a reason of its own, and announcing a skip over it would mislead.
|
|
186
|
+
return nil if ids.empty?
|
|
187
|
+
|
|
188
|
+
@deselected_count = ids.size
|
|
189
|
+
return Set.new
|
|
190
|
+
end
|
|
191
|
+
|
|
192
|
+
return nil unless @selection == 'subset'
|
|
193
|
+
|
|
194
|
+
subset = @tests.to_set
|
|
195
|
+
kept = ids.select { |id| subset.include?(id) }
|
|
196
|
+
matched = kept.to_set
|
|
197
|
+
# Identities, not counts, so a repeated id can never stand in for a
|
|
198
|
+
# missing one.
|
|
199
|
+
if matched != subset
|
|
200
|
+
@not_applied_reason =
|
|
201
|
+
matched.empty? ? 'subset_matched_no_collected_test' : 'subset_partly_absent_from_collection'
|
|
202
|
+
return nil
|
|
203
|
+
end
|
|
204
|
+
|
|
205
|
+
@kept_count = kept.size
|
|
206
|
+
@deselected_count = ids.size - kept.size
|
|
207
|
+
@kept_tests = kept
|
|
208
|
+
subset
|
|
209
|
+
end
|
|
210
|
+
|
|
211
|
+
# The block printed in the gem's "Mergify CI" report: what reduced this
|
|
212
|
+
# run, whether it was deliberate, and whether its green can be trusted.
|
|
213
|
+
# Prose, and no identifier from the API ever reaches it.
|
|
214
|
+
# rubocop:disable-next Metrics/MethodLength
|
|
215
|
+
def report
|
|
216
|
+
if @init_error_msg
|
|
217
|
+
# The error text on its own line: it is what support will ask for,
|
|
218
|
+
# and it usually carries a URL that wrapping would split.
|
|
219
|
+
return block("Mergify couldn't be asked whether this run could be reduced, so the full suite ran.") +
|
|
220
|
+
"Error: #{@init_error_msg}\n"
|
|
221
|
+
end
|
|
222
|
+
|
|
223
|
+
if refused?
|
|
224
|
+
# The server's explanation was printed when the run stopped; the
|
|
225
|
+
# block only says where to look.
|
|
226
|
+
return block('Mergify stopped this run before any test ran; its explanation is in the error above.')
|
|
227
|
+
end
|
|
228
|
+
|
|
229
|
+
return block(FULL_RUN_SENTENCES.fetch(@not_applied_reason, UNKNOWN_REASON_SENTENCE)) if @not_applied_reason
|
|
230
|
+
return empty_block if @selection == 'empty'
|
|
231
|
+
return subset_block if @selection == 'subset' && @kept_count
|
|
232
|
+
|
|
233
|
+
block(FULL_RUN_SENTENCES.fetch(@reason, UNKNOWN_REASON_SENTENCE))
|
|
234
|
+
end
|
|
235
|
+
|
|
236
|
+
private
|
|
237
|
+
|
|
238
|
+
def empty_block
|
|
239
|
+
skipped = @deselected_count
|
|
240
|
+
# Nothing was collected: the block is the title alone, rather than a
|
|
241
|
+
# paragraph about skipping "all 0 tests".
|
|
242
|
+
return "#{HEADER}\n" if skipped.zero?
|
|
243
|
+
|
|
244
|
+
passed, them =
|
|
245
|
+
skipped == 1 ? ['its only test passed back then', 'it'] : ["all #{skipped} tests passed back then", 'them']
|
|
246
|
+
block("The code under test hasn't changed since the previous attempt of this job, and #{passed}. " \
|
|
247
|
+
"Mergify skipped #{them}: the job is green, and no test was executed.")
|
|
248
|
+
end
|
|
249
|
+
|
|
250
|
+
# rubocop:disable-next Metrics/MethodLength
|
|
251
|
+
def subset_block
|
|
252
|
+
failed = @kept_count
|
|
253
|
+
skipped = @deselected_count
|
|
254
|
+
sentence =
|
|
255
|
+
if skipped.zero?
|
|
256
|
+
which, them =
|
|
257
|
+
failed == 1 ? ['its only test failed', 'it'] : ["all #{failed} of its tests failed", 'all of them']
|
|
258
|
+
"The code under test hasn't changed since the previous attempt of this job, where #{which}. " \
|
|
259
|
+
"Mergify re-executed #{them}:"
|
|
260
|
+
else
|
|
261
|
+
those = failed == 1 ? 'that one' : "those #{failed}"
|
|
262
|
+
"The code under test hasn't changed since the previous attempt of this job, where #{failed} of its " \
|
|
263
|
+
"#{count_tests(failed + skipped)} failed. Mergify re-executed only #{those} and skipped the " \
|
|
264
|
+
"#{skipped} that had already passed:"
|
|
265
|
+
end
|
|
266
|
+
|
|
267
|
+
lines = @kept_tests.first(LISTED_TESTS_MAX).map { |id| " #{id}" }
|
|
268
|
+
remaining = @kept_tests.size - LISTED_TESTS_MAX
|
|
269
|
+
lines << " … and #{remaining} more" if remaining.positive?
|
|
270
|
+
"#{block(sentence)}\n#{lines.join("\n")}\n"
|
|
271
|
+
end
|
|
272
|
+
|
|
273
|
+
def count_tests(count)
|
|
274
|
+
count == 1 ? '1 test' : "#{count} tests"
|
|
275
|
+
end
|
|
276
|
+
|
|
277
|
+
def block(text)
|
|
278
|
+
"#{HEADER}\n\n#{TestSelection.wrap(text)}\n"
|
|
279
|
+
end
|
|
280
|
+
|
|
281
|
+
class << self
|
|
282
|
+
# Greedy word wrap, the way Python's `textwrap.fill` does it for this
|
|
283
|
+
# prose: words are never split, and a word longer than the width sits
|
|
284
|
+
# on a line of its own.
|
|
285
|
+
def wrap(text, width = WRAP_WIDTH)
|
|
286
|
+
lines = []
|
|
287
|
+
text.split.each do |word|
|
|
288
|
+
if lines.empty? || lines.last.length + 1 + word.length > width
|
|
289
|
+
lines << word.dup
|
|
290
|
+
else
|
|
291
|
+
lines.last << ' ' << word
|
|
292
|
+
end
|
|
293
|
+
end
|
|
294
|
+
lines.join("\n")
|
|
295
|
+
end
|
|
296
|
+
end
|
|
297
|
+
end
|
|
298
|
+
end
|
|
299
|
+
end
|
data/lib/mergify/rspec/utils.rb
CHANGED
|
@@ -45,6 +45,13 @@ module Mergify
|
|
|
45
45
|
env_truthy?('CI') || env_truthy?('RSPEC_MERGIFY_ENABLE')
|
|
46
46
|
end
|
|
47
47
|
|
|
48
|
+
# Returns true inside one of the processes parallel_tests or turbo_tests
|
|
49
|
+
# splits a suite across. Both set TEST_ENV_NUMBER, to an empty string for
|
|
50
|
+
# parallel_tests' first worker, so it is the key that counts.
|
|
51
|
+
def parallel_worker?
|
|
52
|
+
ENV.key?('TEST_ENV_NUMBER')
|
|
53
|
+
end
|
|
54
|
+
|
|
48
55
|
# Split "owner/repo" into [owner, repo].
|
|
49
56
|
# Raises InvalidRepositoryFullNameError when the format is wrong.
|
|
50
57
|
def split_full_repo_name(full_repo_name)
|
metadata
CHANGED
|
@@ -1,14 +1,14 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: rspec-mergify
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.4.0
|
|
5
5
|
platform: x86_64-linux
|
|
6
6
|
authors:
|
|
7
7
|
- Mergify
|
|
8
8
|
autorequire:
|
|
9
9
|
bindir: bin
|
|
10
10
|
cert_chain: []
|
|
11
|
-
date: 2026-
|
|
11
|
+
date: 2026-10-01 00:00:00.000000000 Z
|
|
12
12
|
dependencies:
|
|
13
13
|
- !ruby/object:Gem::Dependency
|
|
14
14
|
name: rspec-core
|
|
@@ -47,6 +47,8 @@ files:
|
|
|
47
47
|
- lib/mergify/rspec/native.rb
|
|
48
48
|
- lib/mergify/rspec/quarantine.rb
|
|
49
49
|
- lib/mergify/rspec/resources/rspec.rb
|
|
50
|
+
- lib/mergify/rspec/session_verdict.rb
|
|
51
|
+
- lib/mergify/rspec/test_selection.rb
|
|
50
52
|
- lib/mergify/rspec/trace.rb
|
|
51
53
|
- lib/mergify/rspec/utils.rb
|
|
52
54
|
- lib/mergify/rspec/version.rb
|