scrap4rb 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml ADDED
@@ -0,0 +1,7 @@
1
+ ---
2
+ SHA256:
3
+ metadata.gz: 9534e00ea290711ab140dcca603641ba180313ace9259a32c7b031cffaca788a
4
+ data.tar.gz: cd326a737dd566c651b5c25ef133415afe38d667756fb9138ae20cfbbb951709
5
+ SHA512:
6
+ metadata.gz: a372446e785fa9b0d5b915577df54a1526e77821357c7971095ab3b859b72778c23e85c9a8b984934a208a4d2a31c1c3d4f5df5f2a1904544925f21b8e8e7ed7
7
+ data.tar.gz: 8af1af6f60679abb63c4e8ae1454176be74292096d6f45a78bcdbcfaa00ebe8e0013929488a408d00cb6f334888b0adc54d8f86102b474e131501f3467617a7f
data/CHANGELOG.md ADDED
@@ -0,0 +1,23 @@
1
+ # Changelog
2
+
3
+ All notable changes are documented here. The format is based on
4
+ [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project
5
+ adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
6
+
7
+ ## [Unreleased]
8
+
9
+ ## [0.1.0] - 2026-10-02
10
+
11
+ ### Added
12
+ - Initial release. A Ruby port of SCRAP by Robert C. Martin for Minitest
13
+ test files: structure errors, a SCRAP score and smells per test, fuzzy
14
+ duplication with extraction pressure, coverage-matrix detection, and per
15
+ file a refactor pressure, a remediation mode and an AI actionability class.
16
+ - `scrap4rb` executable with `--all`, `--verbose`, `--json`,
17
+ `--write-baseline` and `--compare`.
18
+ - Support for `test "..." do`, `def test_*`, `define_method(:"test_...")` and minitest/spec
19
+ `describe`/`it` blocks. A `define_method` test inside a loop over a table
20
+ (`CASES.each`) is one table-driven test that spans the loop.
21
+
22
+ [Unreleased]: https://github.com/roberthopman/scrap4rb/compare/v0.1.0...HEAD
23
+ [0.1.0]: https://github.com/roberthopman/scrap4rb/releases/tag/v0.1.0
data/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Robert Hopman
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
data/README.md ADDED
@@ -0,0 +1,161 @@
1
+ # scrap4rb
2
+
3
+ `scrap4rb` scores Minitest test code for structural quality. It is a Ruby port of [SCRAP](https://github.com/unclebob/scrap) by Robert C. Martin (Uncle Bob). SCRAP does for test code what CRAP does for production code: it finds the tests that are too large, too weak, too logic-heavy or too mock-heavy, and the duplicated scaffolding that is worth a helper.
4
+
5
+ For each test file, scrap4rb answers three questions for a programmer or an AI assistant:
6
+
7
+ - Is this file poorly structured enough to refactor?
8
+ - Where is the worst structure: which tests, which lines?
9
+ - How to refactor: split the file, fix tests in place, make a table, or leave it alone?
10
+
11
+ The output is advice, not a verdict. Read the file before you act on it.
12
+
13
+ ## Installation
14
+
15
+ ```bash
16
+ gem install scrap4rb
17
+ ```
18
+
19
+ Or add it to a Gemfile:
20
+
21
+ ```ruby
22
+ gem "scrap4rb", group: :development
23
+ ```
24
+
25
+ From a checkout you can run it without installing:
26
+
27
+ ```bash
28
+ ruby -Ilib exe/scrap4rb path/to/project/test
29
+ ```
30
+
31
+ ## Usage
32
+
33
+ Run it from the project root. The default path is `test`.
34
+
35
+ ```bash
36
+ scrap4rb # files that need attention, plus totals
37
+ scrap4rb --all # every file, STABLE included
38
+ scrap4rb --verbose test/models # per-test metrics
39
+ scrap4rb --json > scrap.json # the full report as data
40
+ scrap4rb test/models/order_test.rb # one file
41
+ ```
42
+
43
+ Before and after a refactor:
44
+
45
+ ```bash
46
+ scrap4rb --write-baseline tmp/scrap.json
47
+ # refactor the tests
48
+ scrap4rb --compare tmp/scrap.json
49
+ ```
50
+
51
+ The comparison gives each changed file a verdict: `improved`, `worse`, `mixed` or `unchanged`. When the verdict is `worse`, revert the refactor or simplify the helpers.
52
+
53
+ ## Output
54
+
55
+ ```
56
+ test/models/smelly_test.rb
57
+ score 48.46 HIGH mode LOCAL ai AUTO_REFACTOR tests=6 avg=12.78 max=18.26
58
+ HIGH Strengthen assertions in weak tests before doing structural cleanup.
59
+ LOW Be skeptical of helper extraction that only hides setup; ...
60
+ 18.3 L59 it uses a long local helper low-assertion-density,helper-hidden-complexity
61
+ 17.3 L35 it handles a large payload low-assertion-density,literal-heavy-setup
62
+ 11.0 L8 it saves the record no-assertions
63
+ EXTRACT L4-17 net=31.74 (I=3 shared=22 variable=0): hours above 168 ... | negative hours ...
64
+ ```
65
+
66
+ ```
67
+ Field Meaning
68
+ ----- -------
69
+ score Refactor pressure of the file
70
+ level STABLE, LOW, MEDIUM, HIGH or CRITICAL
71
+ mode STABLE: leave it. LOCAL: fix tests in place. SPLIT: split the file first
72
+ ai LEAVE_ALONE, AUTO_TABLE_DRIVE, AUTO_REFACTOR, MANUAL_SPLIT or REVIEW_FIRST
73
+ HIGH/... Recommendations, ranked by confidence, at most four
74
+ tests The SCRAP score of each test, worst first, with its smells
75
+ EXTRACT A group of tests where a shared helper pays for itself
76
+ ```
77
+
78
+ ## What it checks
79
+
80
+ Structure errors:
81
+
82
+ - a `test` inside a `test`
83
+ - a `setup`, `teardown`, `describe`, `context` or `class` inside a test
84
+ - a second `def test_x` with the same name in one class (the first one never runs)
85
+ - parse errors
86
+
87
+ Smells per test, with the SCRAP penalty:
88
+
89
+ ```
90
+ Smell Fires when Penalty
91
+ ----- ---------- -------
92
+ no-assertions the test has no assertion 10
93
+ low-assertion-density one assertion in more than 10 lines 6
94
+ multiple-phases assert, act, assert again 5
95
+ high-mocking more than 3 stubs or mocks 4
96
+ large-example more than 20 lines 4
97
+ helper-hidden-complexity more than 8 lines hidden in local helpers 4
98
+ temp-resource-work temporary files, threads or shell calls 3
99
+ literal-heavy-setup a string over 5 lines, a hash or array over 10 3
100
+ ```
101
+
102
+ A short test with one or two assertions and few subjects is an "API contract" test. SCRAP excuses it from the low-assertion-density and large-example smells.
103
+
104
+ Duplication is structural and fuzzy. Names and literal values become placeholders, so two tests that differ only in their numbers have the same shape. scrap4rb separates three kinds:
105
+
106
+ - **Harmful duplication**: repeated setup, arrange or assertion code. It becomes an EXTRACT recommendation only when a helper pays for itself after its own cost.
107
+ - **Coverage matrix**: many small, similar tests. The advice is "make it one table-driven test", not "this is bad".
108
+ - **Subject repetition**: many tests of the same API. This is normal and costs almost nothing.
109
+
110
+ ## How SCRAP maps to Minitest
111
+
112
+ ```
113
+ SCRAP (speclj) scrap4rb (Minitest)
114
+ -------------- -------------------
115
+ describe, context a test class, or a minitest/spec describe block
116
+ it test "..." do, def test_*, it "..." do, define_method(:"test_...")
117
+ before, with setup blocks and def setup, inherited by every test in the class
118
+ let, binding travel_to, freeze_time, within, using_session,
119
+ perform_enqueued_jobs, with_*, and stub with a block
120
+ with-redefs stub, stub_request, any_instance, expects, stubs, Minitest::Mock.new
121
+ should* assert*, refute*, must_*, wont_*
122
+ doseq, for, every? each, each_with_index, map, times, all?, ...
123
+ if, when, cond, and if, unless, case, while, until, &&, ||, rescue
124
+ ```
125
+
126
+ The policy numbers (weights, thresholds, levels) are SCRAP's, unchanged.
127
+
128
+ ## Deviations from SCRAP
129
+
130
+ - A test whose whole body is one context block, such as `travel_to(...) do ... end`, is unwrapped before the phase count. In Ruby that block is test context, not arrange code.
131
+ - A call to an assertion helper defined in `test/test_helper.rb`, `test/application_system_test_case.rb`, `test/test_helpers/` or `test/support/` counts as one assertion. SCRAP only reads the spec file itself. Without this rule, every test that uses a shared assertion helper would be "no-assertions".
132
+ - A loop over a literal hash (`{ 5 => 0, 6 => 10 }.each`) counts as table-driven, the same as a loop over an array of arrays.
133
+
134
+ ## Notes for Rails suites
135
+
136
+ Two SCRAP rules can disagree with common Rails practice. Decide for your own suite.
137
+
138
+ - **multiple-phases**: a system test that asserts after each click (to prove that the page changed before the next step) has several phases, and SCRAP gives it 5 points.
139
+ - **Low-assertion files**: a file is not STABLE when more than 35% of its tests have one assertion or fewer. A suite with "one behavior, one assertion" tests will have many LOCAL files for this reason only.
140
+
141
+ ## Development
142
+
143
+ ```bash
144
+ bundle install
145
+ rake test # unit tests and golden snapshots
146
+ UPDATE_GOLDEN=1 rake test # rewrite the snapshots after an intended change
147
+ ```
148
+
149
+ The golden tests run the executable against the fixture suite in `test/fixtures/sample`. A golden diff means that the output changed. Accept it only when the change was intended.
150
+
151
+ ## Credits
152
+
153
+ SCRAP is by Robert C. Martin (Uncle Bob): [github.com/unclebob/scrap](https://github.com/unclebob/scrap). The algorithm, the metrics, the smells and the policy numbers in scrap4rb come from SCRAP. scrap4rb is an independent Ruby implementation. It contains no SCRAP source code.
154
+
155
+ scrap4rb 0.1.0 follows SCRAP at commit [`f793b28`](https://github.com/unclebob/scrap/tree/f793b28906d3c60e30b128e1ec08df6269a72b13) (2026-03-17). When SCRAP changes its rules or numbers, scrap4rb can differ until it follows again.
156
+
157
+ Related tools by the same author: [crap4clj](https://github.com/unclebob/crap4clj), [crap4java](https://github.com/unclebob/crap4java), [crap4go](https://github.com/unclebob/crap4go) and [clj-mutate](https://github.com/unclebob/clj-mutate).
158
+
159
+ ## License
160
+
161
+ scrap4rb is released under the MIT License. See [LICENSE](LICENSE). The license covers the code in this repository. It does not cover SCRAP itself.
data/exe/scrap4rb ADDED
@@ -0,0 +1,4 @@
1
+ #!/usr/bin/env ruby
2
+ require "scrap4rb"
3
+
4
+ Scrap4rb::CLI.run(ARGV)
@@ -0,0 +1,130 @@
1
+ require "json"
2
+ require "digest"
3
+ require "fileutils"
4
+ require "optparse"
5
+ require_relative "example"
6
+ require_relative "structure"
7
+ require_relative "summary"
8
+ require_relative "judgment"
9
+
10
+ module Scrap4rb
11
+ module CLI
12
+ extend self
13
+
14
+ LABEL = { 3 => "HIGH", 2 => "MEDIUM", 1 => "LOW" }.freeze
15
+
16
+ def run(argv)
17
+ opts = parse(argv)
18
+ files = (argv.empty? ? ["test"] : argv).flat_map { File.directory?(_1) ? Dir["#{_1}/**/*_test.rb"] : [_1] }.sort
19
+ shared = Example.shared_assert_helpers(Dir.pwd)
20
+ reports = files.map { analyze_file(_1, shared) }
21
+ write_baseline(opts[:write], reports) if opts[:write]
22
+ return puts(JSON.pretty_generate(json_safe(reports))) if opts[:json]
23
+
24
+ reports.each { print_report(_1, verbose: opts[:verbose]) if opts[:all] || _1[:mode] != "STABLE" || _1[:errors].any? }
25
+ print_comparison(opts[:compare], reports) if opts[:compare]
26
+ print_totals(files, reports)
27
+ end
28
+
29
+ def parse(argv)
30
+ opts = {}
31
+ OptionParser.new do |o|
32
+ o.banner = "Usage: scrap4rb [options] [paths...] (default path: test)"
33
+ o.on("--all", "Show STABLE files too") { opts[:all] = true }
34
+ o.on("--verbose", "Show per-test metrics") { opts[:verbose] = true }
35
+ o.on("--json", "Print the full report as JSON") { opts[:json] = true }
36
+ o.on("--write-baseline FILE", "Write a baseline document to FILE") { opts[:write] = _1 }
37
+ o.on("--compare FILE", "Compare with a baseline document") { opts[:compare] = _1 }
38
+ o.on("-v", "--version", "Show the version") { puts VERSION; exit }
39
+ o.on("-h", "--help", "Show this help") { puts o; exit }
40
+ end.parse!(argv)
41
+ opts
42
+ end
43
+
44
+ def analyze_file(path, shared_asserts)
45
+ src = File.read(path)
46
+ result = Prism.parse(src)
47
+ exs = Example.from_tree(result.value, shared_assert_helpers: shared_asserts)
48
+ summary = Summary.summarize(exs)
49
+ blocks = exs.group_by { _1[:describe_path] }.reject { |p, _| p.empty? }.map { |p, b| { path: p, summary: Summary.summarize(b) } }
50
+ mode = Judgment.remediation_mode(summary, blocks)
51
+ { path:, content_hash: Digest::SHA256.hexdigest(src)[0, 16], errors: Structure.errors(result), examples: exs, summary:,
52
+ file_score: Judgment.pressure_score(summary), level: Judgment.level(summary), mode:,
53
+ ai: Judgment.actionability(summary, mode), actions: Judgment.actions(summary, mode) }
54
+ end
55
+
56
+ def print_report(r, verbose:)
57
+ s = r[:summary]
58
+ puts "\n#{r[:path]}"
59
+ r[:errors].each { puts " #{_1}" }
60
+ puts " score #{r[:file_score]} #{r[:level]} mode #{r[:mode]} ai #{r[:ai]} tests=#{s[:example_count]} avg=#{s[:avg_scrap]} max=#{s[:max_scrap]}"
61
+ r[:actions].each { |c, t| puts " #{LABEL[c].ljust(6)} #{t}" }
62
+ worst = r[:examples].sort_by { -_1[:scrap] }
63
+ worst = worst.first(5) unless verbose
64
+ worst.each { puts example_line(_1, verbose) }
65
+ s[:recommended_extractions].each do |x|
66
+ puts " EXTRACT L#{x[:line_start]}-#{x[:line_end]} net=#{x[:net_benefit]} (I=#{x[:instances]} shared=#{x[:shared_forms]} variable=#{x[:variable_points]}): " +
67
+ x[:examples].map { _1[:name].to_s[0, 40] }.join(" | ")
68
+ end
69
+ end
70
+
71
+ def example_line(e, verbose)
72
+ extra = if verbose
73
+ " [a=#{e[:assertions]} b=#{e[:branches]} d=#{e[:setup_depth]} stub=#{e[:with_redefs]} h=#{e[:helper_calls]} " \
74
+ "hh=#{e[:helper_hidden_lines]} ph=#{e[:phases]}#{" table" if e[:table_driven]}#{" api" if e[:api_contract]}]"
75
+ else
76
+ ""
77
+ end
78
+ format(" %5.1f L%-4d %-58s %s%s", e[:scrap], e[:line], e[:name].to_s[0, 58], e[:smells].join(","), extra)
79
+ end
80
+
81
+ def write_baseline(file, reports)
82
+ FileUtils.mkdir_p(File.dirname(file))
83
+ File.write(file, JSON.pretty_generate(version: 1, reports: reports.map { json_safe(_1.slice(:path, :content_hash, :summary)) }))
84
+ end
85
+
86
+ def compare(baseline, reports)
87
+ before = baseline["reports"].to_h { [_1["path"], _1["summary"].transform_keys(&:to_sym)] }
88
+ reports.filter_map do |r|
89
+ b = before[r[:path]] or next
90
+ a = r[:summary]
91
+ cmp = { file_score_delta: (Judgment.pressure_score(a) - Judgment.pressure_score(b)).round(2),
92
+ max_scrap_delta: (a[:max_scrap] - b[:max_scrap]).round(2),
93
+ extraction_pressure_delta: (a[:effective_duplication_score] - b[:effective_duplication_score]).round(2),
94
+ case_matrix_delta: a[:case_matrix_repetition] - b[:case_matrix_repetition],
95
+ helper_hidden_delta: a[:helper_hidden_example_count] - b[:helper_hidden_example_count] }
96
+ [r[:path], cmp.merge(verdict: Judgment.verdict(cmp))]
97
+ end
98
+ end
99
+
100
+ def print_comparison(file, reports)
101
+ puts "\nComparison with #{file}:"
102
+ compare(JSON.parse(File.read(file)), reports).each do |path, c|
103
+ next if c[:verdict] == "unchanged"
104
+ puts " #{c[:verdict].ljust(9)} #{path} #{c.except(:verdict).map { "#{_1}=#{_2}" }.join(" ")}"
105
+ puts " the refactor made it worse: revert it, or simplify the helpers" if c[:verdict] == "worse"
106
+ end
107
+ end
108
+
109
+ def print_totals(files, reports)
110
+ exs = reports.flat_map { _1[:examples] }
111
+ tally = ->(values) { values.tally.sort.map { "#{_1}=#{_2}" }.join(" ") }
112
+ puts "\n#{files.size} files, #{exs.size} tests, #{reports.sum { _1[:errors].size }} structure errors"
113
+ puts "modes: " + tally.(reports.map { _1[:mode] })
114
+ puts "ai: " + tally.(reports.map { _1[:ai] })
115
+ puts "levels: " + tally.(reports.map { _1[:level] })
116
+ puts "smells: " + exs.flat_map { _1[:smells] }.tally.sort_by { -_2 }.map { "#{_1}=#{_2}" }.join(" ")
117
+ puts "extractions recommended: #{reports.sum { _1[:summary][:recommended_extraction_count] }}, " \
118
+ "table-drive candidates: #{reports.sum { _1[:summary][:coverage_matrix_candidates] }}"
119
+ end
120
+
121
+ def json_safe(value)
122
+ case value
123
+ when Set then value.to_a
124
+ when Hash then value.to_h { [_1, json_safe(_2)] }
125
+ when Array then value.map { json_safe(_1) }
126
+ else value
127
+ end
128
+ end
129
+ end
130
+ end
@@ -0,0 +1,195 @@
1
+ require_relative "syntax"
2
+ require_relative "policy"
3
+
4
+ module Scrap4rb
5
+ module Example
6
+ include Syntax
7
+ extend self
8
+
9
+ SHARED_HELPER_GLOBS = %w[test/test_helper.rb test/application_system_test_case.rb test/test_helpers/**/*.rb test/support/**/*.rb].freeze
10
+
11
+ EMPTY = { assertions: 0, branches: 0, table_branches: 0, with_redefs: 0, helper_calls: 0, helper_hidden_lines: 0,
12
+ temp_resources: 0, large_literals: 0, max_setup_depth: 0, subject_symbols: Set.new.freeze, table_driven: false }.freeze
13
+
14
+ Context = Struct.new(:helpers, :shared_assert_helpers, :cache)
15
+
16
+ def from_source(src, shared_assert_helpers:)
17
+ from_tree(Prism.parse(src).value, shared_assert_helpers:)
18
+ end
19
+
20
+ def from_tree(tree, shared_assert_helpers:)
21
+ collect(statements(tree.statements), Context.new(helper_defs(tree), shared_assert_helpers, {}))
22
+ end
23
+
24
+ def shared_assert_helpers(root)
25
+ SHARED_HELPER_GLOBS.flat_map { Dir[File.join(root, _1)] }.flat_map do |file|
26
+ descendants(Prism.parse_file(file).value).select do |d|
27
+ d.is_a?(Prism::DefNode) && descendants(d).any? { assertion?(_1) }
28
+ end.map(&:name)
29
+ end.to_set
30
+ end
31
+
32
+ def saturating(complexity)
33
+ c = Policy::COMPLEXITY
34
+ return c[:floor] if complexity <= 1
35
+ c[:floor] + (c[:cap] - c[:floor]) * (1.0 - Math.exp(-c[:rise_rate] * (complexity - 1)))
36
+ end
37
+
38
+ def combine(*metrics)
39
+ metrics.reduce(EMPTY) do |acc, m|
40
+ {
41
+ assertions: acc[:assertions] + m.fetch(:assertions, 0),
42
+ branches: acc[:branches] + m.fetch(:branches, 0),
43
+ table_branches: acc[:table_branches] + m.fetch(:table_branches, 0),
44
+ with_redefs: acc[:with_redefs] + m.fetch(:with_redefs, 0),
45
+ helper_calls: acc[:helper_calls] + m.fetch(:helper_calls, 0),
46
+ helper_hidden_lines: acc[:helper_hidden_lines] + m.fetch(:helper_hidden_lines, 0),
47
+ temp_resources: acc[:temp_resources] + m.fetch(:temp_resources, 0),
48
+ large_literals: acc[:large_literals] + m.fetch(:large_literals, 0),
49
+ max_setup_depth: [acc[:max_setup_depth], m.fetch(:max_setup_depth, 0)].max,
50
+ subject_symbols: acc[:subject_symbols] | m.fetch(:subject_symbols, Set.new),
51
+ table_driven: acc[:table_driven] || m.fetch(:table_driven, false)
52
+ }
53
+ end
54
+ end
55
+
56
+ def analyze(node, ctx, depth, stack)
57
+ helper = bare_call?(node) && ctx.helpers.key?(node.name)
58
+ shared_assert = bare_call?(node) && !helper && ctx.shared_assert_helpers.include?(node.name)
59
+ next_depth = context_head?(node) ? depth + 1 : depth
60
+ children = kids(node).map { analyze(_1, ctx, next_depth, stack) }
61
+ expanded = helper && !stack.include?(node.name) ? expand(ctx, node.name, stack | [node.name], depth) : nil
62
+ local = {
63
+ assertions: assertion?(node) || shared_assert ? 1 : 0,
64
+ branches: BRANCH_NODES.any? { node.is_a?(_1) } ? 1 : 0,
65
+ table_branches: call?(node) && TABLE_CALLS.include?(node.name) && node.block ? 1 : 0,
66
+ with_redefs: stub_call?(node) ? 1 : 0,
67
+ helper_calls: helper ? 1 : 0,
68
+ temp_resources: temp_resource?(node) ? 1 : 0,
69
+ large_literals: large_literal?(node) ? 1 : 0,
70
+ max_setup_depth: depth,
71
+ subject_symbols: subject_symbols(node, helper),
72
+ table_driven: table_driven_form?(node)
73
+ }
74
+ combine(local, *children, *[expanded].compact)
75
+ end
76
+
77
+ def subject_symbols(node, helper)
78
+ return Set.new unless call?(node) && !helper && !control_call?(node)
79
+ recv = node.receiver
80
+ constant = recv.is_a?(Prism::ConstantReadNode) || recv.is_a?(Prism::ConstantPathNode)
81
+ Set[constant ? "#{recv.slice}.#{node.name}" : node.name.to_s]
82
+ end
83
+
84
+ def expand(ctx, name, stack, depth)
85
+ ctx.cache[[name, depth]] ||= begin
86
+ body = ctx.helpers[name]
87
+ hidden = body.sum { [1, lines_of(_1)].max }
88
+ combine({ helper_hidden_lines: hidden }, *body.map { analyze(_1, ctx, depth, stack) })
89
+ end
90
+ end
91
+
92
+ def top_forms(body)
93
+ forms = body
94
+ forms = block_body(forms.first) while forms.size == 1 && context_head?(forms.first)
95
+ forms
96
+ end
97
+
98
+ def phase(form, ctx)
99
+ return :setup if context_head?(form)
100
+ analyze(form, ctx, 0, Set.new)[:assertions].positive? ? :assert : :action
101
+ end
102
+
103
+ def assertion_clusters(forms, ctx) = forms.map { phase(_1, ctx) }.chunk_while { _1 == _2 }.count { _1.first == :assert }
104
+
105
+ def api_contract?(m, line_count, phases)
106
+ m[:assertions] <= 2 && m[:branches] + m[:table_branches] <= 4 && m[:table_branches] <= 1 &&
107
+ m[:with_redefs] <= 0 && m[:helper_calls] <= 1 && m[:helper_hidden_lines] <= 0 && m[:temp_resources] <= 0 &&
108
+ m[:large_literals] <= 0 && m[:subject_symbols].size <= 4 && phases <= 1 && line_count <= 18
109
+ end
110
+
111
+ def smells(m, line_count, phases, table_driven, api_contract)
112
+ [
113
+ [m[:assertions].zero?, "no-assertions", 10],
114
+ [m[:assertions] == 1 && line_count > 10 && !table_driven && !api_contract, "low-assertion-density", 6],
115
+ [phases > 1, "multiple-phases", 5],
116
+ [m[:with_redefs] > 3, "high-mocking", 4],
117
+ [line_count > 20 && !api_contract, "large-example", 4],
118
+ [m[:temp_resources].positive?, "temp-resource-work", 3],
119
+ [m[:large_literals].positive?, "literal-heavy-setup", 3],
120
+ [m[:helper_hidden_lines] > 8, "helper-hidden-complexity", 4]
121
+ ].select(&:first).map { { label: _2, penalty: _3 } }
122
+ end
123
+
124
+ def score(name:, node:, body:, ctx:, inherited_setup:, describe_path:, table_driven: false)
125
+ m = combine(*body.map { analyze(_1, ctx, inherited_setup.size, Set.new) }, { table_driven: })
126
+ raw_lines = lines_of(node)
127
+ line_count = raw_lines + m[:helper_hidden_lines]
128
+ forms = top_forms(body)
129
+ phases = assertion_clusters(forms, ctx)
130
+ api = api_contract?(m, line_count, phases)
131
+ setup_depth = api ? [0, m[:max_setup_depth] - 2].max : m[:max_setup_depth]
132
+ branch_penalty = m[:table_driven] ? m[:table_branches] : m[:branches] + m[:table_branches]
133
+ branch_penalty = [0, branch_penalty - 2].max if api
134
+ found = smells(m, line_count, phases, m[:table_driven], api)
135
+ complexity_score = saturating(1 + branch_penalty + setup_depth + m[:helper_calls])
136
+
137
+ setup = inherited_setup + forms.select { context_head?(_1) }
138
+ asserts = forms.select { phase(_1, ctx) == :assert }
139
+ arrange = forms.reject { phase(_1, ctx) == :assert }
140
+ setup_signatures = form_signatures(setup)
141
+ literal_signatures = body.flat_map { large_literal_signatures(_1) }
142
+ {
143
+ name:, describe_path:, line: node.location.start_line, end_line: node.location.end_line,
144
+ line_count:, raw_line_count: raw_lines, assertions: m[:assertions],
145
+ branches: m[:branches] + m[:table_branches], setup_depth: m[:max_setup_depth], with_redefs: m[:with_redefs],
146
+ helper_calls: m[:helper_calls], helper_hidden_lines: m[:helper_hidden_lines], temp_resources: m[:temp_resources],
147
+ table_driven: m[:table_driven], api_contract: api, phases:,
148
+ complexity: 1 + m[:branches] + m[:max_setup_depth] + m[:helper_calls] + m[:helper_hidden_lines] / 8,
149
+ complexity_score: complexity_score.round(2),
150
+ subject_symbols: m[:subject_symbols],
151
+ assert_signatures: form_signatures(asserts), assert_features: form_features(asserts),
152
+ setup_signatures:, setup_features: form_features(setup),
153
+ arrange_signatures: form_signatures(arrange), arrange_features: form_features(arrange),
154
+ literal_signatures:,
155
+ fixture_features: setup_signatures.empty? ? Set.new : shape_features(setup_signatures.map { :string }.unshift(:vec)),
156
+ scrap: (complexity_score + found.sum { _1[:penalty] }).round(2),
157
+ smells: found.map { _1[:label] },
158
+ literal_features: literal_signatures.to_set
159
+ }
160
+ end
161
+
162
+ def setup_forms(scope_body)
163
+ scope_body.flat_map do |n|
164
+ if bare_call?(n) && SETUP_CALLS.include?(n.name) && n.block.is_a?(Prism::BlockNode) then block_body(n)
165
+ elsif n.is_a?(Prism::DefNode) && n.name == :setup then statements(n.body).reject { bare_call?(_1) && _1.name == :super }
166
+ else []
167
+ end
168
+ end
169
+ end
170
+
171
+ def test_body(node) = node.is_a?(Prism::DefNode) ? statements(node.body) : block_body(node)
172
+
173
+ def collect(forms, ctx, path = [], inherited = [], out = [])
174
+ setup = inherited + setup_forms(forms)
175
+ forms.each do |f|
176
+ if (name, body = scope_children(f))
177
+ collect(body, ctx, path + [name], setup, out)
178
+ elsif test_node?(f)
179
+ out << score(name: test_name(f), node: f, body: test_body(f), ctx:, inherited_setup: setup, describe_path: path)
180
+ elsif (inner = dynamic_test_in_loop(f))
181
+ out << score(name: dynamic_test_name(inner), node: f, body: block_body(inner), ctx:, inherited_setup: setup,
182
+ describe_path: path, table_driven: true)
183
+ elsif dynamic_test?(f)
184
+ out << score(name: dynamic_test_name(f), node: f, body: block_body(f), ctx:, inherited_setup: setup, describe_path: path)
185
+ end
186
+ end
187
+ out
188
+ end
189
+
190
+ def helper_defs(tree)
191
+ descendants(tree).select { _1.is_a?(Prism::DefNode) && !_1.name.start_with?("test_") && !%i[setup teardown].include?(_1.name) }
192
+ .to_h { [_1.name, statements(_1.body)] }
193
+ end
194
+ end
195
+ end
@@ -0,0 +1,111 @@
1
+ require_relative "policy"
2
+
3
+ module Scrap4rb
4
+ module Judgment
5
+ extend self
6
+
7
+ def ratio(n, d) = d.positive? ? n.fdiv(d) : 0.0
8
+
9
+ def stable?(s)
10
+ n = s[:example_count]
11
+ zero = ratio(s[:zero_assertion_examples], n)
12
+ low = ratio(s[:low_assertion_examples], n)
13
+ small = n <= 2 && s[:max_scrap] <= 10 && s[:effective_duplication_score] <= 1 && s[:helper_hidden_example_count].zero? && zero <= 0
14
+ general = n.positive? && s[:max_scrap] <= 12 && s[:effective_duplication_score] <= 3 && zero <= 0 && low <= 0.35
15
+ small || general
16
+ end
17
+
18
+ def pressure_score(s)
19
+ n = s[:example_count]
20
+ w = Policy::PRESSURE[:weights]
21
+ base = w[:avg_scrap] * s[:avg_scrap] + w[:max_scrap] * s[:max_scrap] +
22
+ w[:effective_duplication_score] * s[:effective_duplication_score] +
23
+ w[:low_assertion_ratio] * ratio(s[:low_assertion_examples], n) +
24
+ w[:branching_ratio] * ratio(s[:branching_examples], n) +
25
+ w[:with_redefs_ratio] * ratio(s[:with_redefs_examples], n) +
26
+ w[:helper_hidden_ratio] * ratio(s[:helper_hidden_example_count], n)
27
+ factor = Policy::PRESSURE[:size_factors].find { |up_to, _| up_to.nil? || n <= up_to }.last
28
+ [0, factor * base - Policy::PRESSURE[:matrix_credit] * s[:case_matrix_repetition]].max.round(2)
29
+ end
30
+
31
+ def level(s)
32
+ score = pressure_score(s)
33
+ l = Policy::PRESSURE[:levels]
34
+ if stable?(s) then "STABLE"
35
+ elsif score >= l[:critical] then "CRITICAL"
36
+ elsif score >= l[:high] then "HIGH"
37
+ elsif score >= l[:medium] then "MEDIUM"
38
+ else "LOW"
39
+ end
40
+ end
41
+
42
+ def remediation_mode(s, blocks)
43
+ sp = Policy::PRESSURE[:split]
44
+ hot_blocks = blocks.count { %w[HIGH CRITICAL].include?(level(_1[:summary])) }
45
+ split_pressure = s[:avg_scrap] >= sp[:avg_scrap] || s[:effective_duplication_score] >= sp[:effective_duplication_score] ||
46
+ s[:subject_repetition_score] >= sp[:subject_repetition_score] || s[:helper_hidden_example_count].positive?
47
+ if stable?(s) then "STABLE"
48
+ elsif s[:example_count] >= sp[:example_count] && (hot_blocks >= sp[:high_pressure_blocks] || s[:max_scrap] >= sp[:max_scrap]) && split_pressure
49
+ "SPLIT"
50
+ else "LOCAL"
51
+ end
52
+ end
53
+
54
+ def ratios(s)
55
+ n = s[:example_count]
56
+ { low: ratio(s[:low_assertion_examples], n), zero: ratio(s[:zero_assertion_examples], n),
57
+ branching: ratio(s[:branching_examples], n), mocking: ratio(s[:with_redefs_examples], n) }
58
+ end
59
+
60
+ def actionability(s, mode)
61
+ if mode == "STABLE" then "LEAVE_ALONE"
62
+ elsif matrix_heavy?(s) then "AUTO_TABLE_DRIVE"
63
+ elsif local_safe?(s, mode) then "AUTO_REFACTOR"
64
+ elsif mode == "SPLIT" then "MANUAL_SPLIT"
65
+ else "REVIEW_FIRST"
66
+ end
67
+ end
68
+
69
+ def matrix_heavy?(s)
70
+ r = ratios(s)
71
+ m = Policy::ACTIONABILITY[:matrix]
72
+ s[:coverage_matrix_candidates].positive? &&
73
+ s[:case_matrix_repetition] >= [m[:min_case_matrix_repetition], (s[:effective_duplication_score] / m[:effective_duplication_divisor]).floor].max &&
74
+ s[:max_scrap] <= m[:max_scrap] && r[:branching] <= m[:max_branching_ratio] && r[:mocking] < m[:max_mocking_ratio]
75
+ end
76
+
77
+ def local_safe?(s, mode)
78
+ r = ratios(s)
79
+ l = Policy::ACTIONABILITY[:local]
80
+ mode == "LOCAL" &&
81
+ (s[:effective_duplication_score].positive? || r[:zero].positive? || r[:low] > l[:low_assertion_ratio] || s[:max_scrap] > l[:max_scrap]) &&
82
+ r[:branching] <= l[:max_branching_ratio] && r[:mocking] < l[:max_mocking_ratio]
83
+ end
84
+
85
+ def actions(s, mode)
86
+ return [[3, "No refactor recommended; the file is structurally stable enough to leave alone."]] if mode == "STABLE"
87
+ r = ratios(s)
88
+ l = Policy::ACTIONABILITY[:local]
89
+ [
90
+ [mode == "SPLIT", 3, "Split this test file by responsibility before attempting local cleanup; the structural pressure is spread across multiple hotspots."],
91
+ [s[:coverage_matrix_candidates].positive?, 3, "Convert repeated low-complexity tests into table-driven checks; treat this as coverage-matrix repetition, not harmful duplication."],
92
+ [r[:zero].positive? || r[:low] > l[:low_assertion_ratio], 3, "Strengthen assertions in weak tests before doing structural cleanup."],
93
+ [mode == "LOCAL" && s[:max_scrap] > l[:max_scrap], 3, "Split oversized tests into narrower tests."],
94
+ [s[:effective_duplication_score].positive?, 2, "Extract shared setup or repeated assertion scaffolding only where harmful duplication is dominating."],
95
+ [r[:mocking] > l[:max_mocking_ratio], 2, "Reduce mocking and move coverage toward higher-level behaviors."],
96
+ [r[:branching] > l[:max_branching_ratio], 2, "Remove logic from tests or keep variation in explicit data tables rather than control flow."],
97
+ [s[:helper_hidden_example_count].positive?, 1, "Be skeptical of helper extraction that only hides setup; helper-hidden complexity should still count as complexity."],
98
+ [mode == "LOCAL" && s[:avg_scrap] > l[:avg_scrap], 1, "Consider splitting this file or block by responsibility."]
99
+ ].select(&:first).map { _1.drop(1) }.sort_by { [-_1[0], _1[1]] }.first(Policy::ACTIONABILITY[:max_actions])
100
+ end
101
+
102
+ def verdict(c)
103
+ if c[:file_score_delta] <= -5 && c[:extraction_pressure_delta] <= 0 && c[:max_scrap_delta] <= 0 then "improved"
104
+ elsif c[:helper_hidden_delta].positive? && c[:extraction_pressure_delta] >= 0 && c[:case_matrix_delta] <= 0 then "worse"
105
+ elsif c[:extraction_pressure_delta].positive? || c[:max_scrap_delta].positive? || c[:file_score_delta] >= 5 then "worse"
106
+ elsif c[:file_score_delta].zero? then "unchanged"
107
+ else "mixed"
108
+ end
109
+ end
110
+ end
111
+ end
@@ -0,0 +1,23 @@
1
+ module Scrap4rb
2
+ module Policy
3
+ DUPLICATION = { threshold: 0.5, matrix_max_scrap: 18, matrix_max_lines: 12, matrix_max_assertions: 1,
4
+ matrix_max_branches: 0, matrix_max_setup_depth: 2, matrix_max_with_redefs: 0,
5
+ matrix_max_temp_resources: 0, matrix_max_helper_hidden_lines: 0, matrix_max_subject_symbols: 2 }.freeze
6
+ COMPLEXITY = { cap: 25.0, rise_rate: 0.18, floor: 1.0 }.freeze
7
+ PRESSURE = {
8
+ size_factors: [[1, 0.25], [2, 0.40], [4, 0.65], [nil, 1.0]],
9
+ weights: { avg_scrap: 1.2, max_scrap: 0.6, effective_duplication_score: 0.8, low_assertion_ratio: 20,
10
+ branching_ratio: 15, with_redefs_ratio: 15, helper_hidden_ratio: 12 },
11
+ matrix_credit: 1.5,
12
+ levels: { critical: 55, high: 35, medium: 18 },
13
+ split: { avg_scrap: 10, effective_duplication_score: 20, subject_repetition_score: 12, example_count: 12,
14
+ high_pressure_blocks: 2, max_scrap: 35 }
15
+ }.freeze
16
+ ACTIONABILITY = {
17
+ matrix: { min_case_matrix_repetition: 2, effective_duplication_divisor: 3, max_scrap: 12,
18
+ max_branching_ratio: 0.15, max_mocking_ratio: 0.2 },
19
+ local: { low_assertion_ratio: 0.4, max_branching_ratio: 0.3, max_mocking_ratio: 0.35, max_scrap: 20, avg_scrap: 12 },
20
+ max_actions: 4
21
+ }.freeze
22
+ end
23
+ end
@@ -0,0 +1,42 @@
1
+ require_relative "syntax"
2
+
3
+ module Scrap4rb
4
+ module Structure
5
+ include Syntax
6
+ extend self
7
+
8
+ INSIDE_TEST_ERRORS = %i[test it specify setup teardown before after describe context].to_set
9
+
10
+ def errors(result)
11
+ parse_errors(result) + nesting_errors(result.value) + duplicate_name_errors(result.value)
12
+ end
13
+
14
+ def parse_errors(result) = result.errors.map { "ERROR line #{_1.location.start_line}: parse error: #{_1.message}" }
15
+
16
+ def nesting_errors(node, parent = nil, out = [])
17
+ if parent && misplaced?(node)
18
+ label = node.is_a?(Prism::ClassNode) ? "class" : node.name
19
+ out << "ERROR line #{node.location.start_line}: (#{label}) inside (#{parent[:form]}) at line #{parent[:line]}"
20
+ end
21
+ inner = test_node?(node) ? { form: node.is_a?(Prism::DefNode) ? "def #{node.name}" : node.name, line: node.location.start_line } : parent
22
+ kids(node).each { nesting_errors(_1, inner, out) }
23
+ out
24
+ end
25
+
26
+ def misplaced?(node)
27
+ test_node?(node) || (bare_call?(node) && INSIDE_TEST_ERRORS.include?(node.name) && node.block) ||
28
+ node.is_a?(Prism::ClassNode) || (node.is_a?(Prism::DefNode) && node.name == :setup)
29
+ end
30
+
31
+ def duplicate_name_errors(tree)
32
+ scopes = [tree] + descendants(tree).select { _1.is_a?(Prism::ClassNode) || (call?(_1) && BLOCK_CALLS.include?(_1.name)) }
33
+ scopes.flat_map do |scope|
34
+ body = scope.is_a?(Prism::ProgramNode) ? scope.statements.body : (scope_children(scope)&.last || [])
35
+ body.select { test_node?(_1) }.group_by { method_name_of(_1) }.filter_map do |name, nodes|
36
+ next if nodes.size < 2
37
+ "ERROR line #{nodes.last.location.start_line}: #{name} is defined again; line #{nodes.first.location.start_line} never runs"
38
+ end
39
+ end
40
+ end
41
+ end
42
+ end
@@ -0,0 +1,131 @@
1
+ require "set"
2
+ require_relative "policy"
3
+
4
+ module Scrap4rb
5
+ module Summary
6
+ extend self
7
+
8
+ def jaccard(a, b)
9
+ union = a | b
10
+ union.empty? ? 0.0 : (a & b).size.fdiv(union.size)
11
+ end
12
+
13
+ def duplication_cost(shared, instances, variable)
14
+ return 0.0 if shared <= 3 || variable > 4
15
+ (shared - 3) * (instances - 1)**1.5 / (variable + 1)
16
+ end
17
+
18
+ def summarize(exs)
19
+ total = exs.size
20
+ recs = extractions(exs)
21
+ matrix = coverage_matrix_count(exs)
22
+ dup = %i[setup assert fixture literal arrange].to_h { [_1, similar_count(exs, :"#{_1}_features")] }
23
+ pressure = recs.sum { _1[:net_benefit] }.round(2)
24
+ {
25
+ example_count: total,
26
+ avg_scrap: total.positive? ? (exs.sum { _1[:scrap] } / total).round(2) : 0.0,
27
+ max_scrap: exs.map { _1[:scrap] }.max || 0,
28
+ branching_examples: exs.count { _1[:branches].positive? && !_1[:table_driven] },
29
+ low_assertion_examples: exs.count { _1[:assertions] <= 1 },
30
+ with_redefs_examples: exs.count { _1[:with_redefs].positive? },
31
+ zero_assertion_examples: exs.count { _1[:assertions].zero? },
32
+ helper_hidden_example_count: exs.count { _1[:helper_hidden_lines].positive? },
33
+ table_driven_examples: exs.count { _1[:table_driven] },
34
+ setup_duplication_score: dup[:setup], assertion_duplication_score: dup[:assert],
35
+ fixture_duplication_score: dup[:fixture], literal_duplication_score: dup[:literal],
36
+ arrange_duplication_score: dup[:arrange],
37
+ subject_repetition_score: similar_count(exs, :subject_symbols),
38
+ duplication_score: dup.values.sum,
39
+ harmful_duplication_score: dup.values_at(:setup, :assert, :fixture, :arrange).sum,
40
+ avg_setup_similarity: average_similarity(exs, :setup_features).round(3),
41
+ avg_assert_similarity: average_similarity(exs, :assert_features).round(3),
42
+ avg_arrange_similarity: average_similarity(exs, :arrange_features).round(3),
43
+ avg_subject_similarity: average_similarity(exs, :subject_symbols).round(3),
44
+ coverage_matrix_candidates: matrix, case_matrix_repetition: matrix,
45
+ recommended_extraction_count: recs.size,
46
+ extraction_pressure_score: pressure,
47
+ effective_duplication_score: pressure,
48
+ recommended_extractions: recs
49
+ }
50
+ end
51
+
52
+ def extractions(exs)
53
+ components(exs).reject { |c| c.all? { extraction_matrix?(_1) } }.map { extraction(_1) }
54
+ .select { _1[:net_benefit].positive? }.sort_by { [_1[:line_start], -_1[:net_benefit]] }
55
+ end
56
+
57
+ def similar_to_any?(ex, exs, key) = ex[key].any? && exs.any? { !_1.equal?(ex) && jaccard(ex[key], _1[key]) >= Policy::DUPLICATION[:threshold] }
58
+ def similar_count(exs, key) = exs.count { similar_to_any?(_1, exs, key) }
59
+
60
+ def average_similarity(exs, key)
61
+ pairs = exs.combination(2).map { jaccard(_1[key], _2[key]) }
62
+ pairs.empty? ? 0.0 : pairs.sum / pairs.size
63
+ end
64
+
65
+ def harmful_features(ex) = ex[:setup_features] | ex[:assert_features] | ex[:fixture_features] | ex[:arrange_features]
66
+
67
+ def matrix_base?(ex)
68
+ d = Policy::DUPLICATION
69
+ ex[:scrap] <= d[:matrix_max_scrap] && ex[:line_count] <= d[:matrix_max_lines] && ex[:assertions] <= d[:matrix_max_assertions] &&
70
+ ex[:branches] <= d[:matrix_max_branches] && ex[:setup_depth] <= d[:matrix_max_setup_depth] &&
71
+ ex[:with_redefs] <= d[:matrix_max_with_redefs] && ex[:temp_resources] <= d[:matrix_max_temp_resources] &&
72
+ ex[:helper_hidden_lines] <= d[:matrix_max_helper_hidden_lines]
73
+ end
74
+
75
+ def few_subjects?(ex) = ex[:subject_symbols].size <= Policy::DUPLICATION[:matrix_max_subject_symbols]
76
+
77
+ def extraction_matrix?(ex) = matrix_base?(ex) && (ex[:table_driven] || (few_subjects?(ex) && harmful_features(ex).any?))
78
+
79
+ def summary_matrix?(ex)
80
+ matrix_base?(ex) && (ex[:table_driven] || (few_subjects?(ex) && (ex[:assert_features].any? || ex[:arrange_features].any?)))
81
+ end
82
+
83
+ def coverage_matrix_count(exs)
84
+ exs.count do |ex|
85
+ summary_matrix?(ex) && (similar_to_any?(ex, exs, :setup_features) || similar_to_any?(ex, exs, :arrange_features) ||
86
+ (similar_to_any?(ex, exs, :assert_features) && similar_to_any?(ex, exs, :subject_symbols)))
87
+ end
88
+ end
89
+
90
+ def components(exs)
91
+ features = exs.map { harmful_features(_1) }
92
+ adjacency = Hash.new { |h, k| h[k] = Set.new }
93
+ exs.each_index.to_a.combination(2).each do |i, j|
94
+ next unless jaccard(features[i], features[j]) >= Policy::DUPLICATION[:threshold]
95
+ adjacency[i] << j
96
+ adjacency[j] << i
97
+ end
98
+ seen = Set.new
99
+ exs.each_index.filter_map do |start|
100
+ next if seen.include?(start)
101
+ component = reachable(start, adjacency)
102
+ seen.merge(component)
103
+ component.sort.map { exs[_1] } if component.size > 1
104
+ end
105
+ end
106
+
107
+ def reachable(start, adjacency)
108
+ component = Set.new
109
+ stack = [start]
110
+ until stack.empty?
111
+ i = stack.pop
112
+ next if component.include?(i)
113
+ component << i
114
+ stack.concat(adjacency[i].to_a)
115
+ end
116
+ component
117
+ end
118
+
119
+ def extraction(cluster)
120
+ sets = cluster.map { harmful_features(_1) }
121
+ shared = sets.reduce(:&).size
122
+ variable = sets.reduce(:|).size - shared
123
+ before = duplication_cost(shared, cluster.size, variable)
124
+ helper_cost = shared + variable
125
+ { examples: cluster.map { _1.slice(:name, :line, :end_line) }, line_start: cluster.map { _1[:line] }.min,
126
+ line_end: cluster.map { _1[:end_line] }.max, instances: cluster.size, shared_forms: shared,
127
+ variable_points: variable, duplication_before: before.round(2), helper_cost:,
128
+ net_benefit: [0.0, before - helper_cost].max.round(2) }
129
+ end
130
+ end
131
+ end
@@ -0,0 +1,168 @@
1
+ require "prism"
2
+ require "set"
3
+
4
+ module Scrap4rb
5
+ module Syntax
6
+ ASSERTION = /\A(assert|refute|must_|wont_)/
7
+ CONTEXT_CALLS = %i[travel_to travel freeze_time within within_frame using_session perform_enqueued_jobs].to_set
8
+ STUB_CALLS = %i[stub stub_request any_instance expects stubs stub_const].to_set
9
+ TABLE_CALLS = %i[each each_with_index each_pair each_with_object each_slice map flat_map times all? any? none?].to_set
10
+ BRANCH_NODES = [Prism::IfNode, Prism::UnlessNode, Prism::CaseNode, Prism::CaseMatchNode, Prism::WhileNode,
11
+ Prism::UntilNode, Prism::AndNode, Prism::OrNode, Prism::RescueNode, Prism::RescueModifierNode].freeze
12
+ TEST_CALLS = %i[test it specify].to_set
13
+ BLOCK_CALLS = %i[describe context].to_set
14
+ SETUP_CALLS = %i[setup before].to_set
15
+
16
+ extend self
17
+
18
+ def kids(node) = node.compact_child_nodes
19
+ def descendants(node) = kids(node).flat_map { [_1, *descendants(_1)] }
20
+ def call?(node) = node.is_a?(Prism::CallNode)
21
+ def bare_call?(node) = call?(node) && node.receiver.nil?
22
+ def lines_of(node) = node.location.end_line - node.location.start_line + 1
23
+
24
+ def statements(body)
25
+ case body
26
+ when nil then []
27
+ when Prism::StatementsNode then body.body
28
+ when Prism::BeginNode then body.statements&.body || []
29
+ else [body]
30
+ end
31
+ end
32
+
33
+ def block_body(call) = call.block.is_a?(Prism::BlockNode) ? statements(call.block.body) : []
34
+
35
+ def context_head?(node)
36
+ return false unless call?(node) && node.block.is_a?(Prism::BlockNode)
37
+ CONTEXT_CALLS.include?(node.name) || node.name.start_with?("with_") || node.name == :stub
38
+ end
39
+
40
+ def stub_call?(node)
41
+ call?(node) && (STUB_CALLS.include?(node.name) ||
42
+ (node.name == :new && node.receiver&.slice == "Minitest::Mock"))
43
+ end
44
+
45
+ def assertion?(node) = call?(node) && ASSERTION.match?(node.name)
46
+
47
+ def table_literal?(node)
48
+ case node
49
+ when Prism::ArrayNode then node.elements.size >= 2 && node.elements.all? { _1.is_a?(Prism::ArrayNode) || _1.is_a?(Prism::HashNode) }
50
+ when Prism::HashNode then node.elements.size >= 2
51
+ else false
52
+ end
53
+ end
54
+
55
+ def table_driven_form?(node)
56
+ return false unless call?(node)
57
+ (TABLE_CALLS.include?(node.name) && node.block && table_literal?(node.receiver)) ||
58
+ (assertion?(node) && (node.arguments&.arguments || []).any? { table_literal?(_1) })
59
+ end
60
+
61
+ def temp_resource?(node)
62
+ return true if node.is_a?(Prism::XStringNode) || node.is_a?(Prism::InterpolatedXStringNode)
63
+ return false unless call?(node)
64
+ recv = node.receiver&.slice
65
+ %i[mktmpdir tmpdir system spawn fork].include?(node.name) ||
66
+ recv == "Tempfile" || recv == "Thread" ||
67
+ (recv == "File" && %i[write binwrite].include?(node.name)) ||
68
+ (recv == "FileUtils")
69
+ end
70
+
71
+ def large_literal?(node)
72
+ case node
73
+ when Prism::StringNode, Prism::InterpolatedStringNode then string_line_count(node) > 5
74
+ when Prism::HashNode, Prism::KeywordHashNode, Prism::ArrayNode then node.elements.size > 10
75
+ else false
76
+ end
77
+ end
78
+
79
+ def string_line_count(node)
80
+ if node.opening_loc&.slice&.start_with?("<<") && node.closing_loc
81
+ node.closing_loc.start_line - node.opening_loc.start_line - 1
82
+ elsif node.is_a?(Prism::StringNode) then node.unescaped.lines.size
83
+ else lines_of(node)
84
+ end
85
+ end
86
+
87
+ def control_call?(node)
88
+ assertion?(node) || context_head?(node) || stub_call?(node) || TABLE_CALLS.include?(node.name) ||
89
+ TEST_CALLS.include?(node.name) || SETUP_CALLS.include?(node.name) || !node.name.match?(/\A[a-z_]/i)
90
+ end
91
+
92
+ def test_node?(node)
93
+ (node.is_a?(Prism::DefNode) && node.name.start_with?("test_")) ||
94
+ (bare_call?(node) && TEST_CALLS.include?(node.name) && node.block.is_a?(Prism::BlockNode))
95
+ end
96
+
97
+ def dynamic_test?(node)
98
+ bare_call?(node) && node.name == :define_method && node.block.is_a?(Prism::BlockNode) &&
99
+ dynamic_test_name(node).to_s.start_with?("test_")
100
+ end
101
+
102
+ def dynamic_test_name(node)
103
+ arg = node.arguments&.arguments&.first
104
+ case arg
105
+ when Prism::SymbolNode, Prism::StringNode then arg.unescaped
106
+ when Prism::InterpolatedSymbolNode, Prism::InterpolatedStringNode then arg.parts.map(&:slice).join
107
+ end
108
+ end
109
+
110
+ def dynamic_test_in_loop(node)
111
+ return unless call?(node) && TABLE_CALLS.include?(node.name) && node.block.is_a?(Prism::BlockNode)
112
+ block_body(node).find { dynamic_test?(_1) }
113
+ end
114
+
115
+ def test_name(node)
116
+ if node.is_a?(Prism::DefNode) then node.name.to_s
117
+ else node.arguments&.arguments&.first.then { _1.respond_to?(:unescaped) ? _1.unescaped : _1&.slice.to_s }
118
+ end
119
+ end
120
+
121
+ def method_name_of(node) = node.is_a?(Prism::DefNode) ? node.name.to_s : "test_#{test_name(node).to_s.gsub(/\s+/, "_")}"
122
+
123
+ def scope_children(node)
124
+ case node
125
+ when Prism::ClassNode, Prism::ModuleNode then [node.constant_path.slice, statements(node.body)]
126
+ when Prism::CallNode
127
+ [test_name(node) || node.name.to_s, block_body(node)] if bare_call?(node) && BLOCK_CALLS.include?(node.name) && node.block
128
+ end
129
+ end
130
+
131
+ def shape(node)
132
+ case node
133
+ when nil then nil
134
+ when Prism::StringNode, Prism::InterpolatedStringNode, Prism::XStringNode, Prism::InterpolatedXStringNode then :string
135
+ when Prism::IntegerNode, Prism::FloatNode, Prism::RationalNode, Prism::ImaginaryNode then :number
136
+ when Prism::TrueNode, Prism::FalseNode then :boolean
137
+ when Prism::NilNode then :nil
138
+ when Prism::SymbolNode then :"#{node.unescaped}"
139
+ when Prism::LocalVariableReadNode, Prism::InstanceVariableReadNode, Prism::ConstantReadNode,
140
+ Prism::ConstantPathNode, Prism::SelfNode, Prism::ItLocalVariableReadNode then :sym
141
+ when Prism::HashNode, Prism::KeywordHashNode then [:hash, *node.elements.map { shape(_1) }.sort_by(&:inspect)]
142
+ when Prism::AssocNode then [shape(node.key), shape(node.value)]
143
+ when Prism::ArrayNode then [:vec, *node.elements.map { shape(_1) }]
144
+ when Prism::CallNode then [:sym, *[node.receiver, node.arguments, node.block].compact.map { shape(_1) }]
145
+ when Prism::ArgumentsNode then node.arguments.map { shape(_1) }
146
+ when Prism::StatementsNode then node.body.map { shape(_1) }
147
+ when Prism::LocalVariableWriteNode, Prism::InstanceVariableWriteNode then [:assign, :sym, shape(node.value)]
148
+ else [node.type, *kids(node).map { shape(_1) }]
149
+ end
150
+ end
151
+
152
+ def signature(value) = value.inspect
153
+
154
+ def shape_features(value, acc = Set.new)
155
+ acc << signature(value)
156
+ value.each { shape_features(_1, acc) } if value.is_a?(Array)
157
+ acc
158
+ end
159
+
160
+ def form_features(forms) = forms.each_with_object(Set.new) { |f, acc| shape_features(shape(f), acc) }
161
+ def form_signatures(forms) = forms.map { signature(shape(_1)) }
162
+
163
+ def large_literal_signatures(node)
164
+ own = large_literal?(node) ? [signature(shape(node))] : []
165
+ own + kids(node).flat_map { large_literal_signatures(_1) }
166
+ end
167
+ end
168
+ end
@@ -0,0 +1,3 @@
1
+ module Scrap4rb
2
+ VERSION = "0.1.0"
3
+ end
data/lib/scrap4rb.rb ADDED
@@ -0,0 +1,8 @@
1
+ require_relative "scrap4rb/version"
2
+ require_relative "scrap4rb/policy"
3
+ require_relative "scrap4rb/syntax"
4
+ require_relative "scrap4rb/example"
5
+ require_relative "scrap4rb/structure"
6
+ require_relative "scrap4rb/summary"
7
+ require_relative "scrap4rb/judgment"
8
+ require_relative "scrap4rb/cli"
metadata ADDED
@@ -0,0 +1,109 @@
1
+ --- !ruby/object:Gem::Specification
2
+ name: scrap4rb
3
+ version: !ruby/object:Gem::Version
4
+ version: 0.1.0
5
+ platform: ruby
6
+ authors:
7
+ - Robert Hopman
8
+ bindir: exe
9
+ cert_chain: []
10
+ date: 1980-01-02 00:00:00.000000000 Z
11
+ dependencies:
12
+ - !ruby/object:Gem::Dependency
13
+ name: prism
14
+ requirement: !ruby/object:Gem::Requirement
15
+ requirements:
16
+ - - ">="
17
+ - !ruby/object:Gem::Version
18
+ version: '1.0'
19
+ - - "<"
20
+ - !ruby/object:Gem::Version
21
+ version: '2'
22
+ type: :runtime
23
+ prerelease: false
24
+ version_requirements: !ruby/object:Gem::Requirement
25
+ requirements:
26
+ - - ">="
27
+ - !ruby/object:Gem::Version
28
+ version: '1.0'
29
+ - - "<"
30
+ - !ruby/object:Gem::Version
31
+ version: '2'
32
+ - !ruby/object:Gem::Dependency
33
+ name: minitest
34
+ requirement: !ruby/object:Gem::Requirement
35
+ requirements:
36
+ - - "~>"
37
+ - !ruby/object:Gem::Version
38
+ version: '5.0'
39
+ type: :development
40
+ prerelease: false
41
+ version_requirements: !ruby/object:Gem::Requirement
42
+ requirements:
43
+ - - "~>"
44
+ - !ruby/object:Gem::Version
45
+ version: '5.0'
46
+ - !ruby/object:Gem::Dependency
47
+ name: rake
48
+ requirement: !ruby/object:Gem::Requirement
49
+ requirements:
50
+ - - "~>"
51
+ - !ruby/object:Gem::Version
52
+ version: '13.0'
53
+ type: :development
54
+ prerelease: false
55
+ version_requirements: !ruby/object:Gem::Requirement
56
+ requirements:
57
+ - - "~>"
58
+ - !ruby/object:Gem::Version
59
+ version: '13.0'
60
+ description: scrap4rb reads Minitest test files and scores each test for size, branching,
61
+ mocking, weak assertions and duplicated scaffolding. Per file it reports a refactor
62
+ pressure, a remediation mode and an actionability class for an AI assistant. A Ruby
63
+ port of the SCRAP algorithm by Robert C. Martin.
64
+ email:
65
+ - hopman.r@gmail.com
66
+ executables:
67
+ - scrap4rb
68
+ extensions: []
69
+ extra_rdoc_files: []
70
+ files:
71
+ - CHANGELOG.md
72
+ - LICENSE
73
+ - README.md
74
+ - exe/scrap4rb
75
+ - lib/scrap4rb.rb
76
+ - lib/scrap4rb/cli.rb
77
+ - lib/scrap4rb/example.rb
78
+ - lib/scrap4rb/judgment.rb
79
+ - lib/scrap4rb/policy.rb
80
+ - lib/scrap4rb/structure.rb
81
+ - lib/scrap4rb/summary.rb
82
+ - lib/scrap4rb/syntax.rb
83
+ - lib/scrap4rb/version.rb
84
+ homepage: https://github.com/roberthopman/scrap4rb
85
+ licenses:
86
+ - MIT
87
+ metadata:
88
+ source_code_uri: https://github.com/roberthopman/scrap4rb
89
+ changelog_uri: https://github.com/roberthopman/scrap4rb/blob/main/CHANGELOG.md
90
+ allowed_push_host: https://rubygems.org
91
+ rubygems_mfa_required: 'true'
92
+ rdoc_options: []
93
+ require_paths:
94
+ - lib
95
+ required_ruby_version: !ruby/object:Gem::Requirement
96
+ requirements:
97
+ - - ">="
98
+ - !ruby/object:Gem::Version
99
+ version: '3.2'
100
+ required_rubygems_version: !ruby/object:Gem::Requirement
101
+ requirements:
102
+ - - ">="
103
+ - !ruby/object:Gem::Version
104
+ version: '0'
105
+ requirements: []
106
+ rubygems_version: 3.6.9
107
+ specification_version: 4
108
+ summary: 'SCRAP for Minitest: structural quality scores for Ruby test code.'
109
+ test_files: []