sloplint 0.8.1 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 7b924df142be1939d6da4e27d2e925b2de378cbe835e5c9ea41e3c3b0b046633
4
- data.tar.gz: 128a3a13dd8d4dc81610266ede34b52e73293b61af560ff7fa26d9dbd9708dd5
3
+ metadata.gz: 9031ddb2a2a80637f1958c3ec24014261a23b017c25cc3032e72e85f9acf0313
4
+ data.tar.gz: 53389164a5fdf74367d985bc1d3eff0d661c47cb70d22469ab89e55f56e4ac22
5
5
  SHA512:
6
- metadata.gz: ffc7954ff2c5ccab7e431e007be51bddbb0a42cf27fe54a1c8629ec90c9c0c47c9a7ceba91bc03116f38303241efcf7b4263a84376aedae67a95dae2877d3db8
7
- data.tar.gz: f0c893e80441a5e738a5a037897a064b9cf201fb1bca6f5ca5e1b160eabbf2d5c7d870800a0d8ac36263b00a8dbc0ca331673da4f302a3e2e4a3f41346248e3e
6
+ metadata.gz: 2029db5c55bdaed221ff62c629eaf9989bf05044274f340295ecb5122ba5cd063c773566c9c9bf8e201050348284103f6910aed13b6b24745b538ed030738612
7
+ data.tar.gz: 46130960f17c5fd70b7a8752065fc85520b8eeac94498588b3d0b798daa3bdb8683ab5084e258e561a9b502227793d9746c0fc7e2d7ab2f143a0b5f7364c2fc7
data/CHANGELOG.md CHANGED
@@ -3,7 +3,110 @@
3
3
  All notable changes to this project are documented here. Format loosely
4
4
  follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
5
5
 
6
- ## [Unreleased]
6
+ ## sloplint-judge
7
+
8
+ ### [0.1.0] - 2026-09-21
9
+
10
+ - New sentence rule `unnamed-authority` (`info`, `medium`): a claim handed to
11
+ experts, studies, research, critics or many, an authority the reader could
12
+ not find and that speaks for nobody. Officials, a spokesperson and the
13
+ other sources news quotes by convention pass, and so does an abstract's
14
+ prior work. The reading behind the regex `vague-attribution`.
15
+ - New sentence rule `stated-stakes` (`info`, `low`, off by default): a sentence
16
+ that says something is crucial, vital or key and gives no fact, number or
17
+ consequence, in it or in the sentence after it. Off by default because the
18
+ model rarely answers it above low confidence; `--select stated-stakes` or
19
+ `--strict` runs it.
20
+ - A rule at `low` confidence (`same-weight`, `matched-shape`) now reports its
21
+ notes when `--select` names it; before, they ran and printed nothing
22
+ without `--strict`. The rule's `low` caps the note's confidence; what a
23
+ default run drops is an answer the model itself gave at low confidence.
24
+ - `check --judge` always writes the `{"notes", "judge"}` object, with 0
25
+ requests when no judge rule survived selection. `compare` accepts
26
+ `--drift` after the two files and refuses a third. `status` refuses a key
27
+ with a control character the way `check --judge` does.
28
+ - Graveyard: `owned-claim` was retried as `no-actor` and stays out, and
29
+ `buried-verbs`, the paragraph reading of the nominalization shift, joins
30
+ it; neither crossed the flag on any side. The entries in `docs/JUDGE.md`
31
+ say so.
32
+ - New paragraph rule `promotional` (`warning`, `medium`): a paragraph in
33
+ which every judgment is favorable, none comes with a measure and no
34
+ drawback appears. The paragraph-level reading behind `puffery-words`,
35
+ which catches the register after the watch words have aged out.
36
+ - New paragraph rule `self-narration` (`info`, `medium`): a paragraph whose
37
+ sentences signpost the document, what comes first, what a section covers,
38
+ what the reader should take away, instead of saying something about the
39
+ subject. `throat-clearing` sees the first sentence; this is the paragraph.
40
+ - `script/calibrate run` also prints how many units each side would have
41
+ flagged at the confidence `check` reports. A rare flag barely moves the
42
+ rank statistic, so for such a rule the count is the number to quote.
43
+ - New paragraph rule `same-weight` (`info`, `low`, off by default): a paragraph
44
+ that states its inferences and opinions as flatly as its measurements, with
45
+ no probably, no we think, and no reason given. Separates model from human
46
+ text in news and abstracts; off by default because a design document argues
47
+ in flat sentences on purpose. `--select same-weight` or `--strict` runs it.
48
+ - New sentence rule `trailing-gloss` (`info`, `medium`): a sentence that ends
49
+ on a comma and an -ing clause that interprets the fact before it
50
+ ("highlighting the value of", "underscoring the importance of") rather than
51
+ adding a fact or a consequence. The reading the regex
52
+ `trailing-significance-participle` could not do.
53
+ - Graveyard: `redundancy` was retried as `restatement` with the fairness
54
+ reading it was owed, reversed in news against every model side, and stays
55
+ out. The entry in `docs/JUDGE.md` says why.
56
+ - `wrap-up` also flags a last sentence that admits problems and then
57
+ promises a bright future ("Despite these challenges, the future looks
58
+ promising"). Its flagged level says so, its message now reads "Paragraph
59
+ ends on a summary, a moral or a hope", and fixtures pin both sides: the
60
+ hope that names nothing is flagged, the same frame closing on a dated
61
+ fact is not. Wikipedia's field guide documents the formula across
62
+ unrelated topics, and the GPT family ends documents on it.
63
+
64
+ - First release of the judge gem: the rule catalog in `docs/JUDGE.md`, the
65
+ Jev backend, `check`, `compare`, `rules`, `explain`, and `script/calibrate`.
66
+ Requires sloplint 0.9; see "Phase two" in `docs/JUDGE.md` for what is next.
67
+ - `check` reports what the judge spent: JSON output is `{"notes", "judge"}`
68
+ with the backend, request count, token counts and cost in dollars under
69
+ `judge`, and the same line goes to stderr. Jev returns no price, so the
70
+ cost is computed from TypeSafe's public price: $42 per billion input
71
+ tokens, output tokens free.
72
+ - The key can live in the OS keychain instead of the environment.
73
+ `sloplint-judge key set` stores it once (the keychain tool prompts, so the
74
+ key is never on a command line), the Jev adapter reads it after
75
+ `TYPESAFE_API_KEY`, and `sloplint-judge status` says whether a run could
76
+ happen and where the key is, without reading it. `sloplint-judge key
77
+ unset` removes the item. The check skill probes with `status`.
78
+ - `SYSTEMONE_URL` must point at a `typesafe.ai` host, not only be `https`:
79
+ the key and the document go there, so an injected URL on the sanctioned
80
+ command is refused with exit 2.
81
+
82
+ ## sloplint
83
+
84
+ ## [0.9.0] - 2026-09-21
85
+
86
+ ### Added
87
+
88
+ - `sloplint check --judge` runs the rules of sloplint-judge alongside the
89
+ regex catalog and merges the notes in document order. The judge is a second
90
+ gem built from this repository (`sloplint-judge.gemspec`, `exe/sloplint-judge`)
91
+ whose rules are questions put to a System One model. sloplint loads it by
92
+ name only and exits 2 with an install hint when it is missing. A backend
93
+ failure under `--judge` is exit 3 and writes no notes. See `docs/JUDGE.md`.
94
+ - `Sloplint::Split`, a paragraph and sentence splitter that keeps offsets into
95
+ the source, and `Engine.context_window`, the note's context drawn for any
96
+ span rather than only a regex match. Both are used by the judge.
97
+ - `script/calibrate`, which measures a judge backend against RAID and a
98
+ current-model side generated from RAID's own prompts, and reports the
99
+ pass lines from `docs/JUDGE.md`.
100
+
101
+ ### Fixed
102
+
103
+ - `--markdown` opens and closes a fenced code block only at the start of a
104
+ line, at any indent, so a fence under `10. ` or a nested bullet still
105
+ pairs. A fence quoted inside a sentence (```` ``` ````) used to open a
106
+ block there, and every fence after it paired wrong for the rest of the file.
107
+ - The splitter also drops indented code blocks (four spaces or a tab), lines
108
+ that are one HTML tag, and YAML front matter under `--markdown`, so none of
109
+ them reach the judge as prose.
7
110
 
8
111
  ## [0.8.1] - 2026-09-20
9
112
 
data/README.md CHANGED
@@ -4,6 +4,8 @@ A dependency-free CLI that scans prose for the tells of AI-generated **slop** an
4
4
 
5
5
  The primary reader is an agent (Claude Code and friends) that runs sloplint, reads the JSON, and rewrites what it flags. Humans are the secondary reader, and everything is built to keep the false-positive rate low enough that a flag is worth trusting.
6
6
 
7
+ For the tells a regex cannot see, there is a second gem, [sloplint-judge](#sloplint-judge), whose rules are questions put to a model and whose notes come back in the same JSON. It is optional, it needs an API key, and it is the only part of sloplint that sends your text anywhere.
8
+
7
9
  ## What it catches, and what it doesn't
8
10
 
9
11
  A pattern earns a place in the catalog only if it shows up constantly in AI writing and rarely in careful human writing. Passive voice, weak adverbs, wordiness, clichés a person reaches for too: those belong in `proselint` or `write-good`, not here. sloplint is not a general prose linter and never tries to be. It hunts the specific fingerprints of a language model, so an agent can act on a flag instead of second-guessing it.
@@ -118,9 +120,13 @@ Flags: No fluff, no filler, no jargon.
118
120
  Does not: No parking on Sundays.
119
121
  ```
120
122
 
123
+ ### `--judge`
124
+
125
+ `sloplint check --judge` adds the rules of [sloplint-judge](#sloplint-judge) to the run, questions put to a model rather than regexes, and merges the notes into the same array in document order. It needs the sloplint-judge gem and an API key; the section below covers both.
126
+
121
127
  ## The note
122
128
 
123
- One match is one note. JSON output is an array of these, or an object keyed by path when more than one file is scanned. The schema is the contract:
129
+ One match is one note. JSON output is an array of these, or an object keyed by path when more than one file is scanned. Under `--judge` that array or object sits under a `notes` key next to a `judge` key with the backend name, request count and token counts (see [sloplint-judge](#sloplint-judge)). The schema is the contract:
124
130
 
125
131
  ```json
126
132
  {
@@ -144,16 +150,95 @@ One match is one note. JSON output is an array of these, or an object keyed by p
144
150
 
145
151
  ## Exit codes
146
152
 
147
- Three codes carry the contract. A crash exits nonzero on its own.
153
+ Four codes carry the contract. A crash exits nonzero on its own.
148
154
 
149
155
  | code | meaning |
150
156
  |------|---------|
151
157
  | 0 | ran, no notes |
152
158
  | 1 | ran, notes found |
153
159
  | 2 | bad arguments or usage error |
160
+ | 3 | `--judge` only: the model could not be reached, no notes written |
154
161
 
155
162
  An unknown id or category in `--select`/`--ignore` is a usage error (exit 2, naming the id) rather than a silent no-op, so a typo can't masquerade as a clean scan. Input that is empty or only whitespace is exit 2 for the same reason: a pipe that delivered nothing must not read as a clean draft. Only when every source is empty — one empty file among several named ones is taken as deliberate.
156
163
 
164
+ ## sloplint-judge
165
+
166
+ A second gem in this repository, for the tells a regex cannot see. Its rules are questions put to a System One model (Jev, from TypeSafe) about one paragraph or one sentence at a time: does this paragraph end on a summary, a moral or a hope, does this sentence tell the stated reader anything they did not know, does it name anything a reader could check. The answers come back as sloplint notes, same fields, same JSON, same exit codes, so anything that already reads sloplint's output reads the judge's without change.
167
+
168
+ One thing is different from the rest of sloplint: the judge sends your text to an API. sloplint on its own never leaves the machine. Every paragraph the judge examines goes to `api.typesafe.ai` over HTTPS, and nothing goes anywhere until you set a key, so the plain `sloplint check` stays offline whether or not the judge is installed.
169
+
170
+ ### Install
171
+
172
+ ```bash
173
+ gem install sloplint sloplint-judge
174
+ sloplint-judge key set # stores your TypeSafe API key in the OS keychain; it prompts for it
175
+ ```
176
+
177
+ `key set` hands your terminal to the keychain tool (`security` on macOS, `secret-tool` from libsecret on Linux), which asks for the key with echo off, so the key is never on a command line, in shell history or in a dotfile. Setting `TYPESAFE_API_KEY` in the environment works too and takes precedence. One thing the keychain does not do: it keeps the key out of the agent's environment, not out of your account, since any process running as you can read the item back. `sloplint-judge status` says whether a key was found and where, without printing it, and `sloplint-judge key unset` removes the item.
178
+
179
+ Requires Ruby 3.3+ and sloplint 0.9 or later. The judge is not a plugin of its own: the Claude Code plugin at the root of this repository already carries it, and the `/sloplint:check` skill asks before it runs the judge. A key in the environment makes the judge possible; it does not make it run. The skill puts the question once per conversation, says what leaves the machine and what it costs, and stays offline unless the answer is yes or the request already asked for the judge.
180
+
181
+ ### Run
182
+
183
+ The one command to know, and the one an agent should use:
184
+
185
+ ```bash
186
+ sloplint check --judge --markdown -o json draft.md
187
+ ```
188
+
189
+ That runs both catalogs and merges the notes in document order. Because the judge spent money, the JSON says how much: the notes sit under `notes` and a `judge` object carries the backend, the number of requests, the token counts the backend reported and the cost in dollars. The same figures go to stderr in one line for the human formats.
190
+
191
+ ```json
192
+ {
193
+ "notes": [ ... ],
194
+ "judge": { "backend": "jev-latest", "requests": 9, "input_tokens": 14200, "output_tokens": 610, "cost_usd": 0.000596 }
195
+ }
196
+ ```
197
+
198
+ Without the sloplint-judge gem it exits 2 and says to install it. Without a key it exits 2 and says which variable to set. If the model cannot be reached, or answers in a shape the judge does not understand, it exits 3 and writes no notes at all, the regex ones included, so a partial run can never pass as a clean one.
199
+
200
+ The gem also puts a `sloplint-judge` executable on your path for the judge on its own:
201
+
202
+ ```
203
+ sloplint-judge [-o full|json] [--register TEXT] [--backend NAME] [command] [args]
204
+
205
+ check scan paths (or stdin) with the judge's rules only [default]
206
+ compare A B which of two passages a plain-prose editor keeps (--drift for rewrites)
207
+ rules list the judge's rule catalog (add --json)
208
+ explain ID print one rule's question, levels, rationale and fixtures
209
+ status say whether a run could happen here, and where the key is, without reading it
210
+ key set store the backend's key in the OS keychain (the keychain tool prompts for it)
211
+ key unset remove it from the OS keychain
212
+ version print the sloplint-judge version
213
+ ```
214
+
215
+ `check` takes `--markdown`, `--select`, `--ignore` and `--strict` with the same meanings as sloplint's. `--strict` runs the three rules that are off by default, runs the sentence rules on every sentence, rather than only in the paragraphs a paragraph rule flagged or skipped as too short, and keeps the notes the model was not confident about. `--register TEXT` says who the reader is; the default is an engineer on the team reading a design document, and every question is asked on that reader's behalf, so a rule such as `no-news` flags a sentence that reader already knows rather than one anybody would.
216
+
217
+ ### The rules
218
+
219
+ Fourteen rules in two categories. `sloplint-judge rules` lists them and `sloplint-judge explain ID` prints the question the model is asked, the answer that flags, and the fixtures. The bar is a little different from the regex catalog's: a judge rule ships when a reader shown the flagged unit agrees it should go, whoever wrote it, and how sharply it separates model prose from human prose sets its severity. So `throat-clearing` is `info`, not gone: human abstracts open by announcing the paper, and it is dead weight either way.
220
+
221
+ - **paragraph** (6): `particulars`, a paragraph that names nothing a reader could check; `wrap-up`, a paragraph that ends on a summary, a moral or a hope; `throat-clearing`, a paragraph that opens by announcing its topic; `self-narration`, a paragraph that signposts the document instead of saying something; `promotional`, a paragraph that praises its subject and measures nothing; `same-weight`, a paragraph that states its guesses and opinions as flatly as its measurements. `same-weight` is off by default; name it in `--select` or pass `--strict`.
222
+ - **sentence** (8): `stock-figure`, a stock figure of speech; `no-news`, a sentence that explains what the stated reader already knows; `names-nothing`, a sentence with no specific noun in it; `ends-on-verdict`, a sentence that ends by grading the fact it just stated; `trailing-gloss`, a sentence that ends on an -ing clause drawing its own moral; `unnamed-authority`, a claim handed to experts, studies or many; `stated-stakes`, a sentence that says something matters and not why; `matched-shape`, a pair or triple built to a rhythm rather than to the content. `stated-stakes` and `matched-shape` are off by default; name them in `--select` or pass `--strict`.
223
+
224
+ Each note's `confidence` is the lower of the rule's own ceiling and how sure the model was of that answer. A note the model was unsure about is dropped unless you pass `--strict`, the same way sloplint drops its low-confidence rules.
225
+
226
+ ### Cost and configuration
227
+
228
+ One request per paragraph carries the paragraph questions, and one request per examined sentence carries the sentence questions, about eight in parallel. A 2,000-word document runs in a few seconds. Every run reports requests, tokens and cost, in the JSON under `judge` and on stderr. Jev returns token counts and no price, so the dollar figure is computed from TypeSafe's public price: $42 per billion input tokens, and output tokens are free. At that rate a 2,000-word document costs well under a cent.
229
+
230
+ Configuration is from the environment, plus the OS keychain for the key:
231
+
232
+ | variable | default | meaning |
233
+ |---|---|---|
234
+ | `TYPESAFE_API_KEY` | none, required | bearer key sent with every request; from the environment, else the keychain item `key set` wrote |
235
+ | `SYSTEMONE_MODEL` | `jev-latest` | model name |
236
+ | `SYSTEMONE_URL` | `https://api.typesafe.ai/v1/systemone` | endpoint; must be `https` on a `typesafe.ai` host |
237
+ | `SLOPLINT_JUDGE_BACKEND` | `jev` | which adapter to use |
238
+ | `SLOPLINT_JUDGE_CONCURRENCY` | `8` | parallel requests |
239
+
240
+ The design, the calibration that decides which rules ship, and how to add a backend are in [docs/JUDGE.md](docs/JUDGE.md).
241
+
157
242
  ## The rule catalog
158
243
 
159
244
  82 rules across nine categories, each named for the rhetorical move the construct makes. `sloplint rules` prints them; `sloplint rules --json` gives an agent the enumerable form.
data/docs/SPEC.md CHANGED
@@ -96,7 +96,8 @@ dependencies at runtime.
96
96
 
97
97
  Rationale: rules are regexes; the whole thing is a scanner plus an output
98
98
  formatter. A dependency-free `gem install sloplint` is the robust, boring
99
- choice. The name is free on RubyGems (taken on PyPI, so this also sidesteps the
99
+ choice. The one exception is `check --judge`, which loads the separate
100
+ sloplint-judge gem and calls a model over the network; see `docs/JUDGE.md`. The name is free on RubyGems (taken on PyPI, so this also sidesteps the
100
101
  collision). Ships a `sloplint` executable from `exe/` (claude.ai rejects a
101
102
  plugin with a top-level `bin/`, and sloplint ships as a Claude Code plugin
102
103
  too).
@@ -124,7 +125,7 @@ Rakefile # rake spec
124
125
  ```
125
126
 
126
127
  RSpec is a **development** dependency (in the gemspec's `add_development_
127
- dependency`), so the runtime stays dependency-free.
128
+ dependency`), so the runtime stays dependency-free unless `--judge` is passed.
128
129
 
129
130
  ## CLI surface
130
131
 
@@ -160,13 +161,14 @@ Deliberately left out of v1 (add when a real need shows up, not before):
160
161
 
161
162
  ## Exit codes
162
163
 
163
- Three codes carry the contract. A crash just exits nonzero on its own.
164
+ Four codes carry the contract. A crash just exits nonzero on its own.
164
165
 
165
166
  | code | meaning |
166
167
  |------|---------|
167
168
  | 0 | ran, **no notes** |
168
169
  | 1 | ran, **notes found** |
169
170
  | 2 | bad arguments / usage error |
171
+ | 3 | `--judge` only: the model backend could not be reached, no notes written |
170
172
 
171
173
  Empty or whitespace-only input is exit 2, like a mistyped rule id: a scan of
172
174
  nothing must not report as a clean scan. The text is tested before
@@ -176,7 +178,9 @@ block still exits 0. Only when every source is empty.
176
178
  ## Note (the diagnostic object)
177
179
 
178
180
  One match = one Note. JSON output is an array of these (or an object keyed by
179
- path when multiple files are scanned).
181
+ path when multiple files are scanned). Under `--judge` the array or object sits
182
+ under `"notes"` beside a `"judge"` object with the backend name, request count
183
+ and token counts; see `docs/JUDGE.md` "Note".
180
184
 
181
185
  ```json
182
186
  {
@@ -249,7 +253,9 @@ RULES = [
249
253
  ```
250
254
 
251
255
  Adding a rule = appending one entry + one bad and one ok fixture. That's the
252
- whole extension story. No new files, no plugin system (YAGNI).
256
+ whole extension story. The only thing that plugs in is sloplint-judge, and it
257
+ plugs in by name: `check --judge` does a lazy `require "sloplint/judge"` and
258
+ exits 2 with an install hint when it is missing. See `docs/JUDGE.md`.
253
259
 
254
260
  ### Why Ruby literals, not JSON/YAML
255
261
 
@@ -476,6 +482,21 @@ RSpec (dev dependency), run via `rake spec`.
476
482
  - LSP server mode, editor plugins, autofix/rewrite. sloplint *flags*; the agent
477
483
  rewrites. Autofix is a separate tool if ever.
478
484
  - Non-English. Languages other than English are a v2 conversation.
485
+ - Grammar-frequency tells from the corpus-linguistics literature. Reinhart et
486
+ al. (PNAS 2025) measure GPT-4o at 2.1 times the human rate for
487
+ nominalizations and 2.6 times for "that" relative clauses on a subject noun
488
+ ("a framework that enables real-time analysis"). Both shifts are real in RAID
489
+ and in Claude Sonnet 5 text generated from RAID's prompts, and neither is a
490
+ shape a regex can flag. A single that-relative is ordinary English ("a move
491
+ that is likely to put a dent in its accounts", human BBC news) and human
492
+ news carries one every 650 words; two in one sentence separate no better,
493
+ and flagging either pushes a writer toward the nominalization instead.
494
+ Nominalization is a density, not a sentence shape: five or more suffix nouns
495
+ in one sentence fires on human abstracts at the 2023 model rate, and 50k
496
+ words of README prose gave 22 hits, every one a bullet list or feature
497
+ table. Both belong to a judge that reads a paragraph, not to this catalog.
498
+ Sentence-initial clausal subjects ("That the cache was stale is not in
499
+ dispute.") were also tried and are near zero on both sides everywhere.
479
500
  - ML/embedding-based detection. This is a regex linter on purpose — fast,
480
501
  explainable, zero-dependency. Statistical detection is a different product.
481
502
  - Scraped corpora. Platform terms prohibit automated collection, and finding a
data/exe/sloplint CHANGED
@@ -9,5 +9,6 @@ if (RUBY_VERSION.split(".").map(&:to_i).first(2) <=> [3, 3]) < 0
9
9
  "macOS ships Ruby 2.6 at /usr/bin/ruby. Install a current one with `brew install ruby`."
10
10
  end
11
11
 
12
- require_relative "../lib/sloplint/cli"
12
+ $LOAD_PATH.unshift(File.expand_path("../lib", __dir__))
13
+ require "sloplint/cli"
13
14
  exit Sloplint::CLI.run(ARGV)
data/lib/sloplint/cli.rb CHANGED
@@ -56,6 +56,10 @@ module Sloplint
56
56
  strict = false
57
57
  select = nil
58
58
  ignore = nil
59
+ judge = false
60
+ register = nil
61
+ backend = nil
62
+ help = false
59
63
  p = OptionParser.new do |o|
60
64
  o.banner = "usage: sloplint check [options] [paths...] (\"-\" or no paths = stdin)"
61
65
  o.on("-o", "--output-format FORMAT", %w[full json],
@@ -63,21 +67,143 @@ module Sloplint
63
67
  o.on("--markdown", "Skip fenced/inline code spans, HTML comments, and URLs before scanning.") { markdown = true }
64
68
  o.on("--select IDS", "Only run these rules (comma-separated rule ids or category names).") { |v| select = v.split(",").map(&:strip) }
65
69
  o.on("--ignore IDS", "Skip these rules (comma-separated rule ids or category names).") { |v| ignore = v.split(",").map(&:strip) }
66
- o.on("--strict", "Run every rule, including the ones that are off by default.") { strict = true }
70
+ o.on("--strict", "Run every rule, including the ones that are off by default;",
71
+ "with --judge, also ask the sentence rules about every sentence, about three times the requests.") { strict = true }
72
+ o.on("--judge", "Also run sloplint-judge's rules, which ask a model (needs the gem and a key).") { judge = true }
73
+ o.on("--register TEXT", "With --judge: who the reader is.") { |v| register = v }
74
+ o.on("--backend NAME", "With --judge: which model adapter to use.") { |v| backend = v }
75
+ # OptionParser answers -h itself when nobody else does, and it answers
76
+ # it on the real stdout and ends the process. This command prints to
77
+ # the out it was given and returns, as every other one does.
78
+ o.on("-h", "--help", "Show this help.") { out.puts(o.help); help = true }
67
79
  end
68
- p.order!(argv)
80
+ # permute!, so a flag written after the path is a flag: `sloplint check
81
+ # draft.md --judge` reads the way anyone would write it.
82
+ p.permute!(argv)
83
+ return 0 if help
69
84
 
70
- unknown = unknown_rule_refs(select) + unknown_rule_refs(ignore)
85
+ # The judge is a separate gem that depends on this one. The bare require
86
+ # goes through the load path: exe/sloplint puts this checkout's lib/ at
87
+ # the front, so from the plugin tree the judge files beside this one win,
88
+ # and an installed sloplint-judge gem is found otherwise. Only a missing
89
+ # judge is the install hint; any other LoadError is a real one.
90
+ if judge
91
+ begin
92
+ require "sloplint/judge"
93
+ # A Gem::LoadError is an installed sloplint-judge that RubyGems will
94
+ # not activate beside this sloplint: a conflict, or no version that
95
+ # fits. It carries no path, and to the reader it is the same thing as
96
+ # no judge at all, which is how `--help` already reports it.
97
+ rescue LoadError => e
98
+ raise unless e.path == "sloplint/judge" || e.is_a?(Gem::LoadError)
99
+
100
+ err.puts("sloplint: --judge needs the sloplint-judge gem: gem install sloplint-judge")
101
+ # An installed judge that RubyGems will not activate is a version
102
+ # conflict, and the hint above tells the reader to install what
103
+ # they already have. RubyGems says which versions fell out, so
104
+ # print that too rather than throwing it away.
105
+ err.puts("sloplint: #{e.message}") if e.is_a?(Gem::LoadError)
106
+ return 2
107
+ end
108
+ end
109
+ # Accepted and then dropped, these two read as a judge run that was
110
+ # never asked for: `sloplint check --register "a lawyer" brief.md` runs
111
+ # the regex rules and says nothing about the reader it was given.
112
+ unless judge
113
+ { "--register" => register, "--backend" => backend }.each do |flag, value|
114
+ next if value.nil?
115
+
116
+ err.puts("sloplint: #{flag} needs --judge: without it no model is asked anything.")
117
+ return 2
118
+ end
119
+ end
120
+ catalog = judge ? RULES + Judge::RULES : RULES
121
+
122
+ unknown = unknown_rule_refs(select, catalog) + unknown_rule_refs(ignore, catalog)
71
123
  unless unknown.empty?
72
124
  err.puts("sloplint: unknown rule or category: #{unknown.join(", ")}")
73
- err.puts("run `sloplint rules` to list them.")
125
+ # Under --judge the selection is read against both catalogs, and
126
+ # `sloplint rules` lists only one of them: a judge rule id typed with
127
+ # a letter wrong is not in the list the hint would send you to.
128
+ lists = judge ? "`sloplint rules` and `sloplint-judge rules`" : "`sloplint rules`"
129
+ err.puts("run #{lists} to list them.")
74
130
  return 2
75
131
  end
76
132
 
77
- rules = select_rules(select, ignore, strict)
133
+ rules = select_rules(select, ignore, strict, catalog)
78
134
  paths = argv.empty? ? ["-"] : argv
79
135
  by_path = paths.reject { |x| x == "-" }.size > 1
80
136
 
137
+ # Only the read is invalid input. An ArgumentError out of a scan or a
138
+ # judge run is a bug in this code, and it raises like one instead of
139
+ # sending the reader to look for bad bytes in a file that has none.
140
+ sources = begin
141
+ read_sources(paths, err:, stdin:)
142
+ rescue ArgumentError, Encoding::CompatibilityError => e
143
+ err.puts("sloplint: invalid input: #{e.message}")
144
+ return 2
145
+ end
146
+ return 2 unless sources
147
+
148
+ regex_rules, judge_rules = rules.partition { |r| r.is_a?(Rule) }
149
+
150
+ judge_usage = nil
151
+ # Under --judge the output is the {"notes", "judge"} object whether or
152
+ # not a judge rule survived selection: a caller that asked for the judge
153
+ # reads the wrapper, and an empty selection is 0 requests, not a
154
+ # different output shape. 0 requests also means no key: `--judge
155
+ # --select em-dash` must run on a machine that has none.
156
+ #
157
+ # Before the regex scan, not after it: a missing key or a bad
158
+ # SYSTEMONE_URL is known without asking anything, and finding it out
159
+ # after every file has been scanned throws that work away to print the
160
+ # same message.
161
+ if judge
162
+ # With nothing to ask, no key is read: the backend is named and never
163
+ # built, and `sloplint-judge check` reports an empty selection through
164
+ # the same helper, so the two commands print one line. Either way an
165
+ # unknown --backend or a key that is not there is the same usage
166
+ # error, so one rescue answers for both.
167
+ begin
168
+ if judge_rules.empty?
169
+ judge_usage = Judge::Engine.nothing_asked(backend)
170
+ else
171
+ judge_backend = Judge::Backend.load(backend)
172
+ end
173
+ rescue ArgumentError => e
174
+ err.puts("sloplint: --judge: #{e.message}")
175
+ return 2
176
+ end
177
+ end
178
+
179
+ all_notes = sources.flat_map do |label, text|
180
+ Engine.scan(text, rules: regex_rules, markdown:, path: label)
181
+ end
182
+
183
+ if judge_usage
184
+ err.puts("sloplint: judge #{Judge::Engine.usage_line(judge_usage)}")
185
+ elsif judge_backend
186
+ begin
187
+ judged = Judge::Engine.scan_sources(sources, name: "sloplint: judge", err:, rules: judge_rules, backend: judge_backend,
188
+ markdown:, register: register || Judge::Engine::DEFAULT_REGISTER, strict:)
189
+ # Exit 3 withholds the regex notes too: a caller that asked for both
190
+ # and got one would read it as a clean judge run.
191
+ rescue Judge::BackendError => e
192
+ err.puts("sloplint: judge backend failure: #{e.message}")
193
+ return 3
194
+ end
195
+ order = sources.each_with_index.to_h { |(label, _), i| [label, i] }
196
+ all_notes = (all_notes + judged.notes).sort_by.with_index { |n, i| [order[n.path], n.line, n.column, i] }
197
+ judge_usage = judged.usage
198
+ end
199
+
200
+ emit(all_notes, opts[:format], out:, by_path:, judge: judge_usage)
201
+ all_notes.empty? ? 0 : 1
202
+ end
203
+
204
+ # Read every path (or stdin for "-") as UTF-8. Returns [[label, text], ...]
205
+ # or nil after writing the error, so the caller exits 2.
206
+ def read_sources(paths, err:, stdin: $stdin, name: "sloplint")
81
207
  sources = []
82
208
  paths.each do |path|
83
209
  # Read as UTF-8 whatever the locale says. A sandbox with no LANG set
@@ -90,11 +216,20 @@ module Sloplint
90
216
  stdin.read.force_encoding(Encoding::UTF_8)
91
217
  else
92
218
  unless File.file?(path)
93
- err.puts("sloplint: no such file: #{path}")
94
- return 2
219
+ err.puts("#{name}: no such file: #{path}")
220
+ return nil
95
221
  end
96
222
  File.read(path, encoding: Encoding::UTF_8)
97
223
  end
224
+ # Prose that is not valid UTF-8 is an invalid input, said here rather
225
+ # than left to whatever reads the text first: a scan raises on the
226
+ # first regex, and the judge sends the text to a backend, where it
227
+ # would be a JSON error after the request was paid for.
228
+ unless text.valid_encoding?
229
+ err.puts("#{name}: invalid input: #{path == "-" ? "stdin" : path} is not valid UTF-8")
230
+ return nil
231
+ end
232
+
98
233
  sources << [path == "-" ? "-" : path, text]
99
234
  end
100
235
 
@@ -104,29 +239,19 @@ module Sloplint
104
239
  # mistyped rule id sets, and it exits 2 for the same reason.
105
240
  if sources.all? { |_, text| text.strip.empty? }
106
241
  names = sources.map { |label, _| label == "-" ? "stdin" : label }
107
- err.puts("sloplint: empty input: nothing to check in #{names.join(", ")}")
108
- return 2
109
- end
110
-
111
- all_notes = sources.flat_map do |label, text|
112
- Engine.scan(text, rules:, markdown:, path: label)
242
+ err.puts("#{name}: empty input: nothing to check in #{names.join(", ")}")
243
+ return nil
113
244
  end
245
+ sources
246
+ end
114
247
 
115
- case opts[:format]
116
- when "json"
117
- out.puts(Output.format_json(all_notes, by_path:))
248
+ def emit(notes, format, out:, by_path:, judge: nil)
249
+ if format == "json"
250
+ out.puts(Output.format_json(notes, by_path:, judge:))
118
251
  else
119
- text = Output.format_human(all_notes)
252
+ text = Output.format_human(notes)
120
253
  out.puts(text) unless text.empty?
121
254
  end
122
-
123
- all_notes.empty? ? 0 : 1
124
- # Invalid UTF-8 reaches this two ways: String#strip in the empty check
125
- # raises Encoding::CompatibilityError, the engine's regexes raise
126
- # ArgumentError. Both are the same thing to the reader.
127
- rescue ArgumentError, Encoding::CompatibilityError => e
128
- err.puts("sloplint: invalid input: #{e.message}")
129
- 2
130
255
  end
131
256
 
132
257
  # ── rules ───────────────────────────────────────────────────────────────
@@ -137,19 +262,26 @@ module Sloplint
137
262
  o.on("--json", "Emit the catalog as JSON for machine enumeration.") { as_json = true }
138
263
  end.order!(argv)
139
264
 
140
- if as_json
141
- payload = RULES.map do |r|
265
+ render_rules(RULES, json: as_json, out:)
266
+ 0
267
+ end
268
+
269
+ # The catalog as a table or as JSON. Shared with sloplint-judge, whose
270
+ # rules also carry a unit.
271
+ def render_rules(catalog, json:, out:)
272
+ if json
273
+ payload = catalog.map do |r|
142
274
  { id: r.id, category: r.category, severity: r.severity, confidence: r.confidence,
143
275
  message: r.message, rationale: r.rationale, suggestion: r.suggestion }
276
+ .merge(r.respond_to?(:unit) ? { unit: r.unit } : {})
144
277
  end
145
278
  out.puts(JSON.pretty_generate(payload))
146
279
  else
147
- RULES.each do |r|
280
+ catalog.each do |r|
148
281
  off = r.confidence == "low" ? " [off by default]" : ""
149
282
  out.puts("#{r.id.ljust(24)} #{r.category.ljust(18)} #{r.severity.ljust(8)} #{r.confidence.ljust(7)} #{r.message}#{off}")
150
283
  end
151
284
  end
152
- 0
153
285
  end
154
286
 
155
287
  # ── explain ID ────────────────────────────────────────────────────────
@@ -190,10 +322,10 @@ module Sloplint
190
322
  # ── helpers ─────────────────────────────────────────────────────────────
191
323
  # Ids/categories in refs that match no rule in the catalog. nil (no --select
192
324
  # or --ignore given) passes through as no unknowns.
193
- def unknown_rule_refs(refs)
325
+ def unknown_rule_refs(refs, catalog = RULES)
194
326
  return [] unless refs
195
327
 
196
- known = RULES.flat_map { |r| [r.id, r.category] }.uniq
328
+ known = catalog.flat_map { |r| [r.id, r.category] }.uniq
197
329
  refs - known
198
330
  end
199
331
 
@@ -201,15 +333,15 @@ module Sloplint
201
333
  # excludes low-confidence rules unless they are explicitly selected. A
202
334
  # category ref selects only that category's non-low rules unless --strict
203
335
  # is set; naming a rule by its own id still selects it whatever its
204
- # confidence.
205
- def select_rules(select, ignore, strict = false)
336
+ # confidence. catalog is RULES, or RULES plus the judge's under --judge.
337
+ def select_rules(select, ignore, strict = false, catalog = RULES)
206
338
  runs_by_default = ->(r) { r.confidence != "low" }
207
339
  rules = if select
208
- RULES.select { |r| select.include?(r.id) || (select.include?(r.category) && (runs_by_default.call(r) || strict)) }
340
+ catalog.select { |r| select.include?(r.id) || (select.include?(r.category) && (runs_by_default.call(r) || strict)) }
209
341
  elsif strict
210
- RULES
342
+ catalog
211
343
  else
212
- RULES.select(&runs_by_default)
344
+ catalog.select(&runs_by_default)
213
345
  end
214
346
  if ignore
215
347
  rules = rules.reject { |r| ignore.include?(r.id) || ignore.include?(r.category) }
@@ -217,9 +349,37 @@ module Sloplint
217
349
  rules
218
350
  end
219
351
 
220
- def global_parser(opts, out:)
221
- OptionParser.new do |o|
222
- o.banner = <<~BANNER
352
+ # The judge's part of the agent recipe, which depends on whether the gem
353
+ # is here. Looked up on the load path without loading it: help must not
354
+ # pull a model adapter in. The text says the one thing an agent must know
355
+ # before the flag: the judge sends the text to an API, so ask the person.
356
+ def judge_recipe
357
+ # Installed as a gem, the judge is not on the load path until RubyGems
358
+ # activates it, so the load path alone reports a missing gem. A gem on
359
+ # disk that asks for a different sloplint cannot be activated beside
360
+ # this one, and advertising it would send the agent into a conflict.
361
+ installed = $LOAD_PATH.resolve_feature_path("sloplint/judge") ||
362
+ Gem::Specification.find_all_by_name("sloplint-judge").any? { |spec|
363
+ spec.dependencies.find { |d| d.name == "sloplint" }&.match?("sloplint", Sloplint::VERSION)
364
+ }
365
+ unless installed
366
+ return "# sloplint-judge (not installed here) adds model-backed rules for what a regex cannot see: gem install sloplint-judge\n"
367
+ end
368
+
369
+ <<~RECIPE
370
+ # With the judge (sloplint-judge is installed here). It sends the text to api.typesafe.ai,
371
+ # so an agent asks the person before running it. Check first, ask, then run:
372
+ sloplint-judge status # exit 0 = a key is set up (reads no key), 2 = not
373
+ sloplint check --judge --markdown -o json FILE # output becomes {"notes": [...], "judge": {requests, tokens, cost_usd}}
374
+ # exit 3 = the model could not be reached; nothing was checked, not even the regex rules
375
+ RECIPE
376
+ end
377
+
378
+ # The banner names whether the judge gem is here, which costs a scan of
379
+ # the installed gems. Only --help reads it, so it is built when it is
380
+ # read and never on the way to a scan.
381
+ def global_banner
382
+ <<~BANNER
223
383
  sloplint — flag the rhetorical tics and puffery that mark AI-generated prose.
224
384
 
225
385
  # Recommended for agents:
@@ -228,17 +388,23 @@ module Sloplint
228
388
  # each note: {path,line,column,severity,confidence,rule,category,message,excerpt,context,rationale,suggestion,count}
229
389
  # (count is present only for the rules that tally items)
230
390
 
391
+ #{judge_recipe}
231
392
  usage: sloplint [-o full|json] [command] [args]
232
393
 
233
394
  commands:
234
395
  check scan paths (or stdin) for AI-slop tells and report notes [default]
235
396
  a first argument that is not a command name is taken as a path
397
+ --judge adds sloplint-judge's model-backed rules when that gem is installed
236
398
  rules list the rule catalog (add --json for the machine-readable form)
237
399
  explain ID print one rule's message, rationale, and a bad/ok example
238
400
  version print the sloplint version
239
401
 
240
402
  global options:
241
- BANNER
403
+ BANNER
404
+ end
405
+
406
+ def global_parser(opts, out:)
407
+ parser = OptionParser.new do |o|
242
408
  o.on("-o", "--output-format FORMAT", %w[full json],
243
409
  "Output format: 'full' (human-readable text) or 'json' (default: full).") do |v|
244
410
  opts[:format] = v
@@ -254,6 +420,8 @@ module Sloplint
254
420
  o.separator ""
255
421
  o.separator "See `sloplint explain <id>` for any rule, or docs/SPEC.md for the JSON contract."
256
422
  end
423
+ parser.define_singleton_method(:banner) { CLI.global_banner }
424
+ parser
257
425
  end
258
426
  end
259
427
  end
@@ -59,14 +59,21 @@ module Sloplint
59
59
  # over blanked text shows code and URLs as a run of spaces. Offsets here are
60
60
  # character offsets (MatchData#begin), matching the char-based line_starts_for.
61
61
  def context_for(source, match)
62
- return "[#{match[0].gsub(/\s+/, " ").strip}]" if match[0].length >= CONTEXT_CHARS
62
+ context_window(source, match.begin(0), match.end(0))
63
+ end
64
+
65
+ # The same window for any span [b, e) of source, so a tool that locates a
66
+ # sentence rather than a regex match (the judge) draws its context the
67
+ # way sloplint does.
68
+ def context_window(source, b, e)
69
+ span = source[b...e]
70
+ return "[#{span.gsub(/\s+/, " ").strip}]" if span.length >= CONTEXT_CHARS
63
71
 
64
- b, e = match.begin(0), match.end(0)
65
72
  pre = source[[b - CONTEXT_CHARS, 0].max...b]
66
73
  post = source[e, CONTEXT_CHARS].to_s
67
74
  pre = "…#{pre.sub(/\A\S*\s+/, "")}" if b > CONTEXT_CHARS
68
75
  post = "#{post.sub(/\s+\S*\z/, "")}…" if e + CONTEXT_CHARS < source.length
69
- "#{pre}[#{match[0]}]#{post}".gsub(/\s+/, " ").strip
76
+ "#{pre}[#{span}]#{post}".gsub(/\s+/, " ").strip
70
77
  end
71
78
 
72
79
  # 1-indexed line and column for a char offset into text. Binary-searches a
@@ -101,8 +108,28 @@ module Sloplint
101
108
  # with one alternation, so whichever construct opens first is the one that
102
109
  # gets consumed: a `<!--` quoted inside backticks is inline code, and a
103
110
  # backtick inside a comment is part of the comment.
111
+ # A URL ends before the punctuation that ends the sentence it sits in:
112
+ # "See https://example.com. Then do X." is two sentences, and swallowing
113
+ # the first period would make it one. A URL that really ends in one of
114
+ # these, a Wikipedia link closing on a bracket, loses that character to
115
+ # the sentence instead; the text is only being blanked, so what it costs
116
+ # is one visible character, not a broken link.
117
+ # A fence opens and closes at the start of a line, with three or more
118
+ # backticks; the opener may carry an info string and the closer nothing
119
+ # but whitespace. The line start is the shape that matters: a fence
120
+ # quoted inside a code span, which is how a document explains fences,
121
+ # opened a block in the middle of a sentence and every fence after it in
122
+ # the file paired with the wrong one.
123
+ #
124
+ # Any indent is allowed, because CommonMark measures a fence's indent
125
+ # from its container and a fence under "10. " or a nested bullet stands
126
+ # further in than three spaces. What that costs is an indented code block
127
+ # whose own content has a line of backticks in it, which is rare, and the
128
+ # only thing it costs there is more blanking.
129
+ MARKDOWN_NOISE = /(?<block>^[ \t]*`{3,}[^\n]*\n.*?^[ \t]*`{3,}[ \t]*$|<!--.*?-->)|(?<inline>`[^`\n]*`|https?:\/\/\S*[^\s.,;:!?)\]])/m
130
+
104
131
  def blank_markdown(text)
105
- text.gsub(/```.*?```|<!--.*?-->|`[^`\n]*`|https?:\/\/\S+/m) { |s| s.gsub(/[^\n]/, " ") }
132
+ text.gsub(MARKDOWN_NOISE) { |s| s.gsub(/[^\n]/, " ") }
106
133
  end
107
134
  end
108
135
  end
@@ -20,13 +20,17 @@ module Sloplint
20
20
  end
21
21
 
22
22
  # JSON: an array of notes, or an object keyed by path when >1 file was scanned.
23
- def format_json(notes, by_path: false)
24
- if by_path
25
- grouped = notes.group_by(&:path).transform_values { |ns| ns.map { |n| note_hash(n) } }
26
- JSON.pretty_generate(grouped)
23
+ # judge: when the judge ran, its backend name and usage; the notes then
24
+ # sit under "notes" and the usage under "judge", so a caller who asked for
25
+ # the model's opinion also gets what it cost.
26
+ def format_json(notes, by_path: false, judge: nil)
27
+ payload = if by_path
28
+ notes.group_by(&:path).transform_values { |ns| ns.map { |n| note_hash(n) } }
27
29
  else
28
- JSON.pretty_generate(notes.map { |n| note_hash(n) })
30
+ notes.map { |n| note_hash(n) }
29
31
  end
32
+ payload = { "notes" => payload, "judge" => judge } if judge
33
+ JSON.pretty_generate(payload)
30
34
  end
31
35
 
32
36
  def note_hash(note)
@@ -0,0 +1,348 @@
1
+ # frozen_string_literal: true
2
+
3
+ require_relative "rules"
4
+ require_relative "engine"
5
+
6
+ module Sloplint
7
+ # Splits a document into paragraphs and sentences, keeping each one's offset
8
+ # and length in the original text so a note can land on the file as written.
9
+ #
10
+ # The regex engine does not use this. It exists for the judge and for any
11
+ # other tool that asks questions about a paragraph or a sentence rather than
12
+ # matching a pattern across the whole text. See docs/JUDGE.md "Splitting".
13
+ module Split
14
+ # text is the span as the file has it, which is what a note quotes.
15
+ # asked is the same span with what --markdown skips taken out, which is
16
+ # what the model is shown. They are the same string without --markdown.
17
+ Sentence = Data.define(:offset, :length, :text, :asked)
18
+ Paragraph = Data.define(:offset, :length, :text, :asked, :sentences)
19
+
20
+ # Furniture: a line that is not prose, whatever a regex would think.
21
+ # Headings, list items, table rows, block quotes, horizontal rules,
22
+ # reference-style link definitions and a lone image or link line.
23
+ FURNITURE = /\A[ \t]*(?:\#{1,6}[ \t]|[-*+][ \t]|\||>|[-*_]{3,}[ \t]*\z|\[[^\]]+\]:[ \t]|!\[)/
24
+
25
+ # An ordered-list item, which is furniture too but only sometimes: see
26
+ # ordered_item?.
27
+ ORDERED = /\A[ \t]*(\d+)[.)][ \t]/
28
+
29
+ # Abbreviations whose trailing period does not end a sentence. Case
30
+ # matters: "No." is an abbreviation, "no." is the end of a sentence.
31
+ ABBREV = /\b(?:e\.g|i\.e|vs|etc|cf|Fig|No|Dr|Mr|Mrs|Ms|St|Inc|Ltd|Sec|Ch|Vol)\.\z/
32
+
33
+ # The other period that does not end a sentence, matched on its shape
34
+ # rather than on a list of words: letters written one at a time with a
35
+ # period after each, which is "U.S.", "U.K.", "e.g.", "i.e." and the rest
36
+ # of them. Two pairs at least, so the last letter of a word is not one:
37
+ # "bigger than AT&T." and "a free upgrade from 512K." still end their
38
+ # sentences. The run is bounded, so a piece is read in one pass.
39
+ #
40
+ # One period of this shape is left alone: a single capital standing as a
41
+ # word. It is an initial in "J. K. Rowling wrote the book.", which is
42
+ # split into three sentences here, and it is a whole sentence in "Service
43
+ # A. Service B handles the rest.", which a design document writes far more
44
+ # often than it names anyone by initial.
45
+ #
46
+ # What the run costs instead: a sentence that really ends in one of these
47
+ # joins the one after it, so "It was in the U.S. It cost $2." is read as
48
+ # one sentence.
49
+ INITIALS = /(?<![[:alpha:]])(?:[[:alpha:]]\.){2,6}\z/
50
+
51
+ # A number and a period with nothing else between two boundaries is an
52
+ # enumerator, not a sentence: the "2." of a list item that CommonMark
53
+ # keeps inside the paragraph above it, because a list may interrupt a
54
+ # paragraph only when it starts at 1. The whole piece has to be the
55
+ # number, so a year still closes the sentence it sits in -- "It was
56
+ # rewritten in\n2021. Then it shipped." is two sentences.
57
+ ENUMERATOR = /\A\d+\.\z/
58
+
59
+ # The judge's view of Markdown differs from the regex engine's in one
60
+ # way. Fenced code and comments become spaces, as there, but an inline
61
+ # code span or a URL becomes a run of this character, same length: it is
62
+ # a word in its sentence, so "`x` runs fast." starts at the backtick and
63
+ # "Done. `x` runs." is two sentences. Blanked to spaces, the first lost
64
+ # its head and the second merged. It is a private-use codepoint, which
65
+ # ordinary prose cannot contain, so a line of real text is never taken
66
+ # for a blanked one -- a line reading "XXX" used to be.
67
+ INLINE = "\uE000"
68
+
69
+ # A sentence ends at .!? (with an optional closing quote or bracket)
70
+ # followed by whitespace and a capital, digit, opening quote/bracket, or
71
+ # a code span, which starts a sentence as any other word does.
72
+ BOUNDARY = /(?<=[.!?]|[.!?]["'”’)\]])\s+(?=["'“‘(\[A-Z0-9#{INLINE}])/
73
+
74
+ # The split keeps its separators, so one of the blocks it hands back is
75
+ # the break itself. Both of these are written out here rather than inside
76
+ # the walk, where the interpolation would build the same pattern again for
77
+ # every block of every document.
78
+ BREAK_ONLY = /\A#{PARA_BREAK}\z/
79
+ KEEP_BREAK = /(#{PARA_BREAK})/
80
+ KEEP_BOUNDARY = /(#{BOUNDARY})/
81
+
82
+ module_function
83
+
84
+ # text: the source. markdown: blank code, HTML comments and URLs before
85
+ # splitting, and blank furniture lines so they neither count as prose nor
86
+ # take the prose around them with them. Splitting happens on the blanked
87
+ # copy; the texts returned are cut from the original at the same offsets,
88
+ # so an excerpt is always a string that is in the file. Blanking is
89
+ # character for character, so the offsets agree.
90
+ def paragraphs(text, markdown: false)
91
+ scan, asked = markdown ? blank_furniture(blank(text), blank_blocks(text)) : [text, text]
92
+ # A document with nothing to leave out is asked about as written, and
93
+ # then every text is cut once instead of twice.
94
+ asked = text if asked == text
95
+ written = Cursor.new(text)
96
+ shown = asked.equal?(text) ? nil : Cursor.new(asked)
97
+ out = []
98
+ at = 0
99
+ # Split keeping the separators, and walk the offsets arithmetically:
100
+ # String#index with a start position rescans from 0 on non-ASCII text.
101
+ scan.split(KEEP_BREAK).each do |block|
102
+ start = at
103
+ at += block.length
104
+ next if block.match?(BREAK_ONLY)
105
+
106
+ # Both ends trimmed by the same measure. String#strip takes a leading
107
+ # NUL off as well as whitespace and /\A\s*/ does not, so a block that
108
+ # starts with one used to put every offset in the paragraph one
109
+ # character early and cut its last character off.
110
+ from_lead = block.lstrip
111
+ lead = block.length - from_lead.length
112
+ body = from_lead.rstrip
113
+ next if body.empty?
114
+
115
+ offset = start + lead
116
+ # The paragraph is cut out of the original once, and its sentences are
117
+ # cut out of that: cutting each of them out of the whole document by
118
+ # character offset is what made the walk take the square of the
119
+ # document's size.
120
+ para = written.cut(offset, body.length)
121
+ para_asked = shown ? shown.cut(offset, body.length) : para
122
+ sents = sentences(body, base: offset, text: para, asked: para_asked)
123
+ # A one-sentence block ending in a colon is the lead-in to whatever
124
+ # follows (a code block, a list), not a paragraph.
125
+ next if markdown && sents.size == 1 && body.end_with?(":")
126
+
127
+ unless sents.empty?
128
+ flat = squash(para)
129
+ out << Paragraph.new(offset:, length: body.length, text: flat, sentences: sents,
130
+ asked: para.equal?(para_asked) ? flat : squash(para_asked))
131
+ end
132
+ end
133
+ out
134
+ end
135
+
136
+ # Sentences of one paragraph. base is where the paragraph starts in the
137
+ # document, which each sentence's offset is counted from. text is the
138
+ # paragraph as the file has it and asked is the copy the model is shown,
139
+ # both cut at the paragraph's own offsets; a sentence is cut out of them
140
+ # and its whitespace collapsed, so hard-wrapped lines read as one. A
141
+ # caller that hands over a bare paragraph and neither copy gets the body
142
+ # itself as both.
143
+ def sentences(body, base: 0, text: nil, asked: nil)
144
+ text ||= body
145
+ asked ||= text
146
+ # The same cursors as a paragraph is cut with, for the same reason: a
147
+ # document whose sentences all sit in one paragraph is one long string
148
+ # to cut them out of.
149
+ written = Cursor.new(text)
150
+ shown = asked.equal?(text) ? nil : Cursor.new(asked)
151
+ cut = lambda do |at, length|
152
+ flat = squash(written.cut(at, length))
153
+ Sentence.new(offset: base + at, length:, text: flat,
154
+ asked: shown ? squash(shown.cut(at, length)) : flat)
155
+ end
156
+ parts = []
157
+ buf_start = nil
158
+ pos = 0
159
+ body.split(KEEP_BOUNDARY).each do |piece|
160
+ at = pos
161
+ pos = at + piece.length
162
+ next if piece.match?(/\A\s+\z/)
163
+
164
+ buf_start ||= at
165
+ # The abbreviation ends where the piece does, and a piece ends at a
166
+ # .!? that a boundary follows, so it is always inside this piece: the
167
+ # whole buffer never has to be cut out of the body to see it.
168
+ next if piece.match?(ABBREV) || piece.match?(INITIALS) || piece.match?(ENUMERATOR)
169
+
170
+ parts << cut.call(buf_start, pos - buf_start)
171
+ buf_start = nil
172
+ end
173
+ parts << cut.call(buf_start, pos - buf_start) if buf_start
174
+ parts
175
+ end
176
+
177
+ def squash(s) = s.gsub(/\s+/, " ").strip
178
+
179
+ def blank(text)
180
+ text.gsub(Engine::MARKDOWN_NOISE) do |s|
181
+ Regexp.last_match[:inline] ? INLINE * s.length : s.gsub(/[^\n]/, " ")
182
+ end
183
+ end
184
+
185
+ # The copy the model is shown: fenced code and HTML comments go, because
186
+ # --markdown says they do, and a code span or URL stays as written,
187
+ # because it is a word of its sentence and the model has to read it.
188
+ def blank_blocks(text)
189
+ text.gsub(Engine::MARKDOWN_NOISE) do |s|
190
+ Regexp.last_match[:inline] ? s : s.gsub(/[^\n]/, " ")
191
+ end
192
+ end
193
+
194
+ # Blank furniture lines to same-length spaces, and the indented lines that
195
+ # continue a blanked one: a wrapped bullet is one bullet, and its second
196
+ # line is not a sentence of its own. A line that is only code spans or
197
+ # URLs (a command on a line of its own, a bare link) is furniture too.
198
+ LONE_INLINE = /\A[ \t]*#{INLINE}[ \t#{INLINE}]*\z/
199
+ # The same line with sentence punctuation after it. A URL stops before
200
+ # that punctuation now, so the full stop on a bare link line is written
201
+ # there rather than part of the link.
202
+ LONE_INLINE_ENDED = /\A[ \t]*#{INLINE}[ \t#{INLINE}]*[.,;:!?)\]]+\z/
203
+
204
+ # A bullet. FURNITURE matches one too, among everything else it matches;
205
+ # this is here because a bullet and an ordered item are the only two
206
+ # kinds of furniture a continuation line can belong to.
207
+ BULLET = /\A[ \t]*[-*+][ \t]/
208
+ # The second line of a list item, indented under the first.
209
+ CONTINUED = /\A(?:[ ]{2,}|\t)\S/
210
+ # An indented code block: four spaces or a tab. Two spaces are still
211
+ # prose, so an indented paragraph under a heading is kept, which is what
212
+ # CommonMark says as well -- a code block starts at four.
213
+ CODE_INDENT = /\A(?:[ ]{4}|\t)/
214
+ # A line that is one HTML tag, opening or closing, matched on that shape
215
+ # rather than on a list of element names: <details>, <div>, <br/> and
216
+ # whatever else a document drops into Markdown all look the same from
217
+ # here. A sentence that names a tag in running text does not look like
218
+ # this: "<p> is the tag for a paragraph." has words after the ">".
219
+ HTML_LINE = %r{\A[ \t]{0,3}</?[A-Za-z][^\n]*>[ \t]*\z}
220
+ # The --- that opens YAML front matter. FURNITURE reads it as a
221
+ # horizontal rule wherever it appears; front matter is the block it
222
+ # opens, and only on the first line of the file.
223
+ FRONT = /\A---[ \t]*\z/
224
+
225
+ # How many lines of YAML front matter the document opens with: the ---
226
+ # on the first line, everything to the next --- line, and that line. Zero
227
+ # for a document that does not open with one, so a --- between two
228
+ # paragraphs stays the horizontal rule it is.
229
+ def front_matter(lines)
230
+ return 0 unless lines.first&.chomp&.match?(FRONT)
231
+
232
+ close = lines.drop(1).index { |l| l.chomp.match?(FRONT) } or return 0
233
+
234
+ close + 2
235
+ end
236
+
237
+ # Both copies at once, line by line: the furniture is decided on the
238
+ # splitter's copy, where a code span is a placeholder, and the same lines
239
+ # are blanked in the copy the model is shown. Every line keeps its own
240
+ # length and its line ending, so an offset means the same thing in the
241
+ # original and in both copies. The line ending is taken off before the
242
+ # tests run: on a CRLF file it would otherwise sit between the line and
243
+ # the \z that a horizontal rule or a bare link ends at, and neither would
244
+ # be recognised as furniture.
245
+ def blank_furniture(scan, shown)
246
+ lines = scan.each_line.to_a
247
+ front = front_matter(lines)
248
+ item = false
249
+ prose = false
250
+ code = false
251
+ opens = true
252
+ pairs = lines.zip(shown.each_line.to_a).each_with_index.map do |(whole, also), i|
253
+ l = whole.chomp
254
+ ending = whole[l.length..]
255
+ blank = l.strip.empty?
256
+ # Only a list item runs on to the next line. An indented line under a
257
+ # heading, a table row, a horizontal rule or a link definition is an
258
+ # indented paragraph, and chaining from those dropped the prose along
259
+ # with the furniture above it.
260
+ listed = l.match?(BULLET) || ordered_item?(l, prose)
261
+ indented = l.match?(CODE_INDENT)
262
+ # An indented code block opens where a paragraph cannot be running
263
+ # already -- the first line, or after a blank or furniture line --
264
+ # and where the indent is not a list item's second line. It then runs
265
+ # for as long as the indent holds.
266
+ code = (code && (indented || blank)) || (indented && !item && opens)
267
+ dropped = i < front || code || listed || l.match?(FURNITURE) || l.match?(HTML_LINE) ||
268
+ lone_inline?(l, prose) || (item && l.match?(CONTINUED))
269
+ item = listed || (item && l.match?(CONTINUED))
270
+ prose = !dropped && !blank
271
+ opens = blank || dropped
272
+ dropped ? ["#{" " * l.length}#{ending}"] * 2 : [whole, also]
273
+ end
274
+ [pairs.map(&:first).join, pairs.map(&:last).join]
275
+ end
276
+
277
+ # A line that is only code spans or URLs is furniture: a command on a
278
+ # line of its own, a bare link. With sentence punctuation after it, only
279
+ # where no sentence can have been running already -- after a line of
280
+ # prose it is the wrapped tail of that sentence, and "Run it with\n`make
281
+ # test`." is one sentence that would otherwise lose its object.
282
+ def lone_inline?(line, after_prose)
283
+ return true if line.match?(LONE_INLINE)
284
+
285
+ !after_prose && line.match?(LONE_INLINE_ENDED)
286
+ end
287
+
288
+ # CommonMark lets an ordered list interrupt a paragraph only when it
289
+ # starts at 1. So a hard-wrapped prose line that begins with a year --
290
+ # "The library was released in\n2019. It was rewritten in\n2021." -- is
291
+ # prose, not two list items, and the sentences on it stay in view.
292
+ def ordered_item?(line, after_prose)
293
+ m = ORDERED.match(line) or return false
294
+ !after_prose || m[1] == "1"
295
+ end
296
+
297
+ # Cutting a span out of a string by character offset walks the string from
298
+ # its start, because a character is not a fixed number of bytes, so cutting
299
+ # every paragraph of a document out of it one at a time takes the square of
300
+ # the document's size. The spans are asked for in the order they appear, so
301
+ # a cursor holds the byte offset of the last character cut at and answers
302
+ # each one in the length of the span itself: byteslice takes a byte offset,
303
+ # and a byte offset costs nothing to reach.
304
+ class Cursor
305
+ def initialize(text)
306
+ @text = text
307
+ @char = 0
308
+ @byte = 0
309
+ end
310
+
311
+ # The span of length characters at character offset at. A cursor only
312
+ # goes forward: asked to go back it would answer with the wrong span,
313
+ # and a note would quote a string that is not in the file.
314
+ def cut(at, length)
315
+ raise ArgumentError, "a cursor at #{@char} was asked for #{at}" if at < @char
316
+
317
+ advance(at - @char)
318
+ advance(length)
319
+ end
320
+
321
+ private
322
+
323
+ def advance(count)
324
+ return "" unless count.positive?
325
+
326
+ s = window(count)
327
+ @char += s.length
328
+ @byte += s.bytesize
329
+ s
330
+ end
331
+
332
+ # A character is at most four bytes in UTF-8, so four bytes per character
333
+ # is a window wide enough to hold the ones asked for. Any other encoding
334
+ # widens it until it is, or until the string ends. A character split by
335
+ # the far edge of the window is past the count asked for, so it is never
336
+ # one of the ones returned.
337
+ def window(count)
338
+ width = count * 4
339
+ loop do
340
+ s = @text.byteslice(@byte, width)
341
+ return s[0, count] if s.length >= count || s.bytesize < width
342
+
343
+ width *= 2
344
+ end
345
+ end
346
+ end
347
+ end
348
+ end
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Sloplint
4
- VERSION = "0.8.1"
4
+ VERSION = "0.9.0"
5
5
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: sloplint
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.8.1
4
+ version: 0.9.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Benjamin Jackson
@@ -58,6 +58,7 @@ files:
58
58
  - lib/sloplint/engine.rb
59
59
  - lib/sloplint/output.rb
60
60
  - lib/sloplint/rules.rb
61
+ - lib/sloplint/split.rb
61
62
  - lib/sloplint/version.rb
62
63
  homepage: https://github.com/benjaminjackson/sloplint
63
64
  licenses: