sloplint 0.8.0 → 0.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +123 -1
- data/README.md +90 -5
- data/docs/SPEC.md +31 -6
- data/exe/sloplint +2 -1
- data/lib/sloplint/cli.rb +208 -40
- data/lib/sloplint/engine.rb +31 -4
- data/lib/sloplint/output.rb +9 -5
- data/lib/sloplint/rules.rb +127 -0
- data/lib/sloplint/split.rb +348 -0
- data/lib/sloplint/version.rb +1 -1
- metadata +2 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 9031ddb2a2a80637f1958c3ec24014261a23b017c25cc3032e72e85f9acf0313
|
|
4
|
+
data.tar.gz: 53389164a5fdf74367d985bc1d3eff0d661c47cb70d22469ab89e55f56e4ac22
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 2029db5c55bdaed221ff62c629eaf9989bf05044274f340295ecb5122ba5cd063c773566c9c9bf8e201050348284103f6910aed13b6b24745b538ed030738612
|
|
7
|
+
data.tar.gz: 46130960f17c5fd70b7a8752065fc85520b8eeac94498588b3d0b798daa3bdb8683ab5084e258e561a9b502227793d9746c0fc7e2d7ab2f143a0b5f7364c2fc7
|
data/CHANGELOG.md
CHANGED
|
@@ -3,7 +3,129 @@
|
|
|
3
3
|
All notable changes to this project are documented here. Format loosely
|
|
4
4
|
follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
|
5
5
|
|
|
6
|
-
##
|
|
6
|
+
## sloplint-judge
|
|
7
|
+
|
|
8
|
+
### [0.1.0] - 2026-09-21
|
|
9
|
+
|
|
10
|
+
- New sentence rule `unnamed-authority` (`info`, `medium`): a claim handed to
|
|
11
|
+
experts, studies, research, critics or many, an authority the reader could
|
|
12
|
+
not find and that speaks for nobody. Officials, a spokesperson and the
|
|
13
|
+
other sources news quotes by convention pass, and so does an abstract's
|
|
14
|
+
prior work. The reading behind the regex `vague-attribution`.
|
|
15
|
+
- New sentence rule `stated-stakes` (`info`, `low`, off by default): a sentence
|
|
16
|
+
that says something is crucial, vital or key and gives no fact, number or
|
|
17
|
+
consequence, in it or in the sentence after it. Off by default because the
|
|
18
|
+
model rarely answers it above low confidence; `--select stated-stakes` or
|
|
19
|
+
`--strict` runs it.
|
|
20
|
+
- A rule at `low` confidence (`same-weight`, `matched-shape`) now reports its
|
|
21
|
+
notes when `--select` names it; before, they ran and printed nothing
|
|
22
|
+
without `--strict`. The rule's `low` caps the note's confidence; what a
|
|
23
|
+
default run drops is an answer the model itself gave at low confidence.
|
|
24
|
+
- `check --judge` always writes the `{"notes", "judge"}` object, with 0
|
|
25
|
+
requests when no judge rule survived selection. `compare` accepts
|
|
26
|
+
`--drift` after the two files and refuses a third. `status` refuses a key
|
|
27
|
+
with a control character the way `check --judge` does.
|
|
28
|
+
- Graveyard: `owned-claim` was retried as `no-actor` and stays out, and
|
|
29
|
+
`buried-verbs`, the paragraph reading of the nominalization shift, joins
|
|
30
|
+
it; neither crossed the flag on any side. The entries in `docs/JUDGE.md`
|
|
31
|
+
say so.
|
|
32
|
+
- New paragraph rule `promotional` (`warning`, `medium`): a paragraph in
|
|
33
|
+
which every judgment is favorable, none comes with a measure and no
|
|
34
|
+
drawback appears. The paragraph-level reading behind `puffery-words`,
|
|
35
|
+
which catches the register after the watch words have aged out.
|
|
36
|
+
- New paragraph rule `self-narration` (`info`, `medium`): a paragraph whose
|
|
37
|
+
sentences signpost the document, what comes first, what a section covers,
|
|
38
|
+
what the reader should take away, instead of saying something about the
|
|
39
|
+
subject. `throat-clearing` sees the first sentence; this is the paragraph.
|
|
40
|
+
- `script/calibrate run` also prints how many units each side would have
|
|
41
|
+
flagged at the confidence `check` reports. A rare flag barely moves the
|
|
42
|
+
rank statistic, so for such a rule the count is the number to quote.
|
|
43
|
+
- New paragraph rule `same-weight` (`info`, `low`, off by default): a paragraph
|
|
44
|
+
that states its inferences and opinions as flatly as its measurements, with
|
|
45
|
+
no probably, no we think, and no reason given. Separates model from human
|
|
46
|
+
text in news and abstracts; off by default because a design document argues
|
|
47
|
+
in flat sentences on purpose. `--select same-weight` or `--strict` runs it.
|
|
48
|
+
- New sentence rule `trailing-gloss` (`info`, `medium`): a sentence that ends
|
|
49
|
+
on a comma and an -ing clause that interprets the fact before it
|
|
50
|
+
("highlighting the value of", "underscoring the importance of") rather than
|
|
51
|
+
adding a fact or a consequence. The reading the regex
|
|
52
|
+
`trailing-significance-participle` could not do.
|
|
53
|
+
- Graveyard: `redundancy` was retried as `restatement` with the fairness
|
|
54
|
+
reading it was owed, reversed in news against every model side, and stays
|
|
55
|
+
out. The entry in `docs/JUDGE.md` says why.
|
|
56
|
+
- `wrap-up` also flags a last sentence that admits problems and then
|
|
57
|
+
promises a bright future ("Despite these challenges, the future looks
|
|
58
|
+
promising"). Its flagged level says so, its message now reads "Paragraph
|
|
59
|
+
ends on a summary, a moral or a hope", and fixtures pin both sides: the
|
|
60
|
+
hope that names nothing is flagged, the same frame closing on a dated
|
|
61
|
+
fact is not. Wikipedia's field guide documents the formula across
|
|
62
|
+
unrelated topics, and the GPT family ends documents on it.
|
|
63
|
+
|
|
64
|
+
- First release of the judge gem: the rule catalog in `docs/JUDGE.md`, the
|
|
65
|
+
Jev backend, `check`, `compare`, `rules`, `explain`, and `script/calibrate`.
|
|
66
|
+
Requires sloplint 0.9; see "Phase two" in `docs/JUDGE.md` for what is next.
|
|
67
|
+
- `check` reports what the judge spent: JSON output is `{"notes", "judge"}`
|
|
68
|
+
with the backend, request count, token counts and cost in dollars under
|
|
69
|
+
`judge`, and the same line goes to stderr. Jev returns no price, so the
|
|
70
|
+
cost is computed from TypeSafe's public price: $42 per billion input
|
|
71
|
+
tokens, output tokens free.
|
|
72
|
+
- The key can live in the OS keychain instead of the environment.
|
|
73
|
+
`sloplint-judge key set` stores it once (the keychain tool prompts, so the
|
|
74
|
+
key is never on a command line), the Jev adapter reads it after
|
|
75
|
+
`TYPESAFE_API_KEY`, and `sloplint-judge status` says whether a run could
|
|
76
|
+
happen and where the key is, without reading it. `sloplint-judge key
|
|
77
|
+
unset` removes the item. The check skill probes with `status`.
|
|
78
|
+
- `SYSTEMONE_URL` must point at a `typesafe.ai` host, not only be `https`:
|
|
79
|
+
the key and the document go there, so an injected URL on the sanctioned
|
|
80
|
+
command is refused with exit 2.
|
|
81
|
+
|
|
82
|
+
## sloplint
|
|
83
|
+
|
|
84
|
+
## [0.9.0] - 2026-09-21
|
|
85
|
+
|
|
86
|
+
### Added
|
|
87
|
+
|
|
88
|
+
- `sloplint check --judge` runs the rules of sloplint-judge alongside the
|
|
89
|
+
regex catalog and merges the notes in document order. The judge is a second
|
|
90
|
+
gem built from this repository (`sloplint-judge.gemspec`, `exe/sloplint-judge`)
|
|
91
|
+
whose rules are questions put to a System One model. sloplint loads it by
|
|
92
|
+
name only and exits 2 with an install hint when it is missing. A backend
|
|
93
|
+
failure under `--judge` is exit 3 and writes no notes. See `docs/JUDGE.md`.
|
|
94
|
+
- `Sloplint::Split`, a paragraph and sentence splitter that keeps offsets into
|
|
95
|
+
the source, and `Engine.context_window`, the note's context drawn for any
|
|
96
|
+
span rather than only a regex match. Both are used by the judge.
|
|
97
|
+
- `script/calibrate`, which measures a judge backend against RAID and a
|
|
98
|
+
current-model side generated from RAID's own prompts, and reports the
|
|
99
|
+
pass lines from `docs/JUDGE.md`.
|
|
100
|
+
|
|
101
|
+
### Fixed
|
|
102
|
+
|
|
103
|
+
- `--markdown` opens and closes a fenced code block only at the start of a
|
|
104
|
+
line, at any indent, so a fence under `10. ` or a nested bullet still
|
|
105
|
+
pairs. A fence quoted inside a sentence (```` ``` ````) used to open a
|
|
106
|
+
block there, and every fence after it paired wrong for the rest of the file.
|
|
107
|
+
- The splitter also drops indented code blocks (four spaces or a tab), lines
|
|
108
|
+
that are one HTML tag, and YAML front matter under `--markdown`, so none of
|
|
109
|
+
them reach the judge as prose.
|
|
110
|
+
|
|
111
|
+
## [0.8.1] - 2026-09-20
|
|
112
|
+
|
|
113
|
+
### Added
|
|
114
|
+
|
|
115
|
+
- `cataphoric-teaser` (self-rating, warning, high confidence): the
|
|
116
|
+
forward-pointing tease that sells a claim as rare knowledge before making
|
|
117
|
+
it. "Here's what nobody tells you about hiring.", "The part most people get
|
|
118
|
+
wrong is the rollback." Three closed frames (a forward demonstrative, a
|
|
119
|
+
noun slot, a scarcity subject and a verb of telling); the bare-subject
|
|
120
|
+
form must reach a copula or colon, and "that's what nobody understands"
|
|
121
|
+
points backward and is left alone.
|
|
122
|
+
- `punch-sentence` (cadence, info, medium confidence): the verbless beat of
|
|
123
|
+
one to three words wedged between two long sentences in the middle of a
|
|
124
|
+
paragraph, "Not anymore." and "Simple.", which `mic-drop-closer` and
|
|
125
|
+
`bare-auxiliary-closer` only see at a paragraph's end. Two closed shapes:
|
|
126
|
+
a negator with up to two words after it, or one capitalised word of five
|
|
127
|
+
letters or more. Titles, initials and case citations before the stop are
|
|
128
|
+
guarded, and short clauses with a subject ("He agrees.") are not matched.
|
|
7
129
|
|
|
8
130
|
## [0.8.0] - 2026-09-15
|
|
9
131
|
|
data/README.md
CHANGED
|
@@ -4,6 +4,8 @@ A dependency-free CLI that scans prose for the tells of AI-generated **slop** an
|
|
|
4
4
|
|
|
5
5
|
The primary reader is an agent (Claude Code and friends) that runs sloplint, reads the JSON, and rewrites what it flags. Humans are the secondary reader, and everything is built to keep the false-positive rate low enough that a flag is worth trusting.
|
|
6
6
|
|
|
7
|
+
For the tells a regex cannot see, there is a second gem, [sloplint-judge](#sloplint-judge), whose rules are questions put to a model and whose notes come back in the same JSON. It is optional, it needs an API key, and it is the only part of sloplint that sends your text anywhere.
|
|
8
|
+
|
|
7
9
|
## What it catches, and what it doesn't
|
|
8
10
|
|
|
9
11
|
A pattern earns a place in the catalog only if it shows up constantly in AI writing and rarely in careful human writing. Passive voice, weak adverbs, wordiness, clichés a person reaches for too: those belong in `proselint` or `write-good`, not here. sloplint is not a general prose linter and never tries to be. It hunts the specific fingerprints of a language model, so an agent can act on a flag instead of second-guessing it.
|
|
@@ -118,9 +120,13 @@ Flags: No fluff, no filler, no jargon.
|
|
|
118
120
|
Does not: No parking on Sundays.
|
|
119
121
|
```
|
|
120
122
|
|
|
123
|
+
### `--judge`
|
|
124
|
+
|
|
125
|
+
`sloplint check --judge` adds the rules of [sloplint-judge](#sloplint-judge) to the run, questions put to a model rather than regexes, and merges the notes into the same array in document order. It needs the sloplint-judge gem and an API key; the section below covers both.
|
|
126
|
+
|
|
121
127
|
## The note
|
|
122
128
|
|
|
123
|
-
One match is one note. JSON output is an array of these, or an object keyed by path when more than one file is scanned. The schema is the contract:
|
|
129
|
+
One match is one note. JSON output is an array of these, or an object keyed by path when more than one file is scanned. Under `--judge` that array or object sits under a `notes` key next to a `judge` key with the backend name, request count and token counts (see [sloplint-judge](#sloplint-judge)). The schema is the contract:
|
|
124
130
|
|
|
125
131
|
```json
|
|
126
132
|
{
|
|
@@ -144,23 +150,102 @@ One match is one note. JSON output is an array of these, or an object keyed by p
|
|
|
144
150
|
|
|
145
151
|
## Exit codes
|
|
146
152
|
|
|
147
|
-
|
|
153
|
+
Four codes carry the contract. A crash exits nonzero on its own.
|
|
148
154
|
|
|
149
155
|
| code | meaning |
|
|
150
156
|
|------|---------|
|
|
151
157
|
| 0 | ran, no notes |
|
|
152
158
|
| 1 | ran, notes found |
|
|
153
159
|
| 2 | bad arguments or usage error |
|
|
160
|
+
| 3 | `--judge` only: the model could not be reached, no notes written |
|
|
154
161
|
|
|
155
162
|
An unknown id or category in `--select`/`--ignore` is a usage error (exit 2, naming the id) rather than a silent no-op, so a typo can't masquerade as a clean scan. Input that is empty or only whitespace is exit 2 for the same reason: a pipe that delivered nothing must not read as a clean draft. Only when every source is empty — one empty file among several named ones is taken as deliberate.
|
|
156
163
|
|
|
164
|
+
## sloplint-judge
|
|
165
|
+
|
|
166
|
+
A second gem in this repository, for the tells a regex cannot see. Its rules are questions put to a System One model (Jev, from TypeSafe) about one paragraph or one sentence at a time: does this paragraph end on a summary, a moral or a hope, does this sentence tell the stated reader anything they did not know, does it name anything a reader could check. The answers come back as sloplint notes, same fields, same JSON, same exit codes, so anything that already reads sloplint's output reads the judge's without change.
|
|
167
|
+
|
|
168
|
+
One thing is different from the rest of sloplint: the judge sends your text to an API. sloplint on its own never leaves the machine. Every paragraph the judge examines goes to `api.typesafe.ai` over HTTPS, and nothing goes anywhere until you set a key, so the plain `sloplint check` stays offline whether or not the judge is installed.
|
|
169
|
+
|
|
170
|
+
### Install
|
|
171
|
+
|
|
172
|
+
```bash
|
|
173
|
+
gem install sloplint sloplint-judge
|
|
174
|
+
sloplint-judge key set # stores your TypeSafe API key in the OS keychain; it prompts for it
|
|
175
|
+
```
|
|
176
|
+
|
|
177
|
+
`key set` hands your terminal to the keychain tool (`security` on macOS, `secret-tool` from libsecret on Linux), which asks for the key with echo off, so the key is never on a command line, in shell history or in a dotfile. Setting `TYPESAFE_API_KEY` in the environment works too and takes precedence. One thing the keychain does not do: it keeps the key out of the agent's environment, not out of your account, since any process running as you can read the item back. `sloplint-judge status` says whether a key was found and where, without printing it, and `sloplint-judge key unset` removes the item.
|
|
178
|
+
|
|
179
|
+
Requires Ruby 3.3+ and sloplint 0.9 or later. The judge is not a plugin of its own: the Claude Code plugin at the root of this repository already carries it, and the `/sloplint:check` skill asks before it runs the judge. A key in the environment makes the judge possible; it does not make it run. The skill puts the question once per conversation, says what leaves the machine and what it costs, and stays offline unless the answer is yes or the request already asked for the judge.
|
|
180
|
+
|
|
181
|
+
### Run
|
|
182
|
+
|
|
183
|
+
The one command to know, and the one an agent should use:
|
|
184
|
+
|
|
185
|
+
```bash
|
|
186
|
+
sloplint check --judge --markdown -o json draft.md
|
|
187
|
+
```
|
|
188
|
+
|
|
189
|
+
That runs both catalogs and merges the notes in document order. Because the judge spent money, the JSON says how much: the notes sit under `notes` and a `judge` object carries the backend, the number of requests, the token counts the backend reported and the cost in dollars. The same figures go to stderr in one line for the human formats.
|
|
190
|
+
|
|
191
|
+
```json
|
|
192
|
+
{
|
|
193
|
+
"notes": [ ... ],
|
|
194
|
+
"judge": { "backend": "jev-latest", "requests": 9, "input_tokens": 14200, "output_tokens": 610, "cost_usd": 0.000596 }
|
|
195
|
+
}
|
|
196
|
+
```
|
|
197
|
+
|
|
198
|
+
Without the sloplint-judge gem it exits 2 and says to install it. Without a key it exits 2 and says which variable to set. If the model cannot be reached, or answers in a shape the judge does not understand, it exits 3 and writes no notes at all, the regex ones included, so a partial run can never pass as a clean one.
|
|
199
|
+
|
|
200
|
+
The gem also puts a `sloplint-judge` executable on your path for the judge on its own:
|
|
201
|
+
|
|
202
|
+
```
|
|
203
|
+
sloplint-judge [-o full|json] [--register TEXT] [--backend NAME] [command] [args]
|
|
204
|
+
|
|
205
|
+
check scan paths (or stdin) with the judge's rules only [default]
|
|
206
|
+
compare A B which of two passages a plain-prose editor keeps (--drift for rewrites)
|
|
207
|
+
rules list the judge's rule catalog (add --json)
|
|
208
|
+
explain ID print one rule's question, levels, rationale and fixtures
|
|
209
|
+
status say whether a run could happen here, and where the key is, without reading it
|
|
210
|
+
key set store the backend's key in the OS keychain (the keychain tool prompts for it)
|
|
211
|
+
key unset remove it from the OS keychain
|
|
212
|
+
version print the sloplint-judge version
|
|
213
|
+
```
|
|
214
|
+
|
|
215
|
+
`check` takes `--markdown`, `--select`, `--ignore` and `--strict` with the same meanings as sloplint's. `--strict` runs the three rules that are off by default, runs the sentence rules on every sentence, rather than only in the paragraphs a paragraph rule flagged or skipped as too short, and keeps the notes the model was not confident about. `--register TEXT` says who the reader is; the default is an engineer on the team reading a design document, and every question is asked on that reader's behalf, so a rule such as `no-news` flags a sentence that reader already knows rather than one anybody would.
|
|
216
|
+
|
|
217
|
+
### The rules
|
|
218
|
+
|
|
219
|
+
Fourteen rules in two categories. `sloplint-judge rules` lists them and `sloplint-judge explain ID` prints the question the model is asked, the answer that flags, and the fixtures. The bar is a little different from the regex catalog's: a judge rule ships when a reader shown the flagged unit agrees it should go, whoever wrote it, and how sharply it separates model prose from human prose sets its severity. So `throat-clearing` is `info`, not gone: human abstracts open by announcing the paper, and it is dead weight either way.
|
|
220
|
+
|
|
221
|
+
- **paragraph** (6): `particulars`, a paragraph that names nothing a reader could check; `wrap-up`, a paragraph that ends on a summary, a moral or a hope; `throat-clearing`, a paragraph that opens by announcing its topic; `self-narration`, a paragraph that signposts the document instead of saying something; `promotional`, a paragraph that praises its subject and measures nothing; `same-weight`, a paragraph that states its guesses and opinions as flatly as its measurements. `same-weight` is off by default; name it in `--select` or pass `--strict`.
|
|
222
|
+
- **sentence** (8): `stock-figure`, a stock figure of speech; `no-news`, a sentence that explains what the stated reader already knows; `names-nothing`, a sentence with no specific noun in it; `ends-on-verdict`, a sentence that ends by grading the fact it just stated; `trailing-gloss`, a sentence that ends on an -ing clause drawing its own moral; `unnamed-authority`, a claim handed to experts, studies or many; `stated-stakes`, a sentence that says something matters and not why; `matched-shape`, a pair or triple built to a rhythm rather than to the content. `stated-stakes` and `matched-shape` are off by default; name them in `--select` or pass `--strict`.
|
|
223
|
+
|
|
224
|
+
Each note's `confidence` is the lower of the rule's own ceiling and how sure the model was of that answer. A note the model was unsure about is dropped unless you pass `--strict`, the same way sloplint drops its low-confidence rules.
|
|
225
|
+
|
|
226
|
+
### Cost and configuration
|
|
227
|
+
|
|
228
|
+
One request per paragraph carries the paragraph questions, and one request per examined sentence carries the sentence questions, about eight in parallel. A 2,000-word document runs in a few seconds. Every run reports requests, tokens and cost, in the JSON under `judge` and on stderr. Jev returns token counts and no price, so the dollar figure is computed from TypeSafe's public price: $42 per billion input tokens, and output tokens are free. At that rate a 2,000-word document costs well under a cent.
|
|
229
|
+
|
|
230
|
+
Configuration is from the environment, plus the OS keychain for the key:
|
|
231
|
+
|
|
232
|
+
| variable | default | meaning |
|
|
233
|
+
|---|---|---|
|
|
234
|
+
| `TYPESAFE_API_KEY` | none, required | bearer key sent with every request; from the environment, else the keychain item `key set` wrote |
|
|
235
|
+
| `SYSTEMONE_MODEL` | `jev-latest` | model name |
|
|
236
|
+
| `SYSTEMONE_URL` | `https://api.typesafe.ai/v1/systemone` | endpoint; must be `https` on a `typesafe.ai` host |
|
|
237
|
+
| `SLOPLINT_JUDGE_BACKEND` | `jev` | which adapter to use |
|
|
238
|
+
| `SLOPLINT_JUDGE_CONCURRENCY` | `8` | parallel requests |
|
|
239
|
+
|
|
240
|
+
The design, the calibration that decides which rules ship, and how to add a backend are in [docs/JUDGE.md](docs/JUDGE.md).
|
|
241
|
+
|
|
157
242
|
## The rule catalog
|
|
158
243
|
|
|
159
|
-
|
|
244
|
+
82 rules across nine categories, each named for the rhetorical move the construct makes. `sloplint rules` prints them; `sloplint rules --json` gives an agent the enumerable form.
|
|
160
245
|
|
|
161
|
-
- **self-rating** (
|
|
246
|
+
- **self-rating** (16) the writer grades their own prose or claim: `clean-x`, `clean-count`, `cleanest-x`, `cleanly`, `honest-x`, `most-honest-x`, `honestly` (the honesty family, built the same way as the four `clean` rules), `worth-naming`, `worth-saying-plainly`, `earns-its-place`, `does-a-lot-of-work`, `exact-exactly`, `genuinely` (off by default), `the-punchline-is`, `announced-takeaway`, `cataphoric-teaser` ("Here's what nobody tells you", "the part most people get wrong").
|
|
162
247
|
- **closer** (12) closes by restating or announcing the point: `thats-the-whole`, `is-the-whole-x` (the same closer on any subject: "Consistency is the real test.", at `medium` confidence), `is-the-entire`, `the-entire-is`, `thats-how-x`, `thats-the-tension`, `right-up-until`, `and-nothing-else` (the trailing "…, and nothing else"), `nothing-else-frag`, `bare-equative` ("The lesson is the handoff.", at `medium` confidence), `trailing-restatement` (the "…, which means …" tail that says the sentence again, off by default), `and-what-it-should` (the elliptical tail: "…, and what it should.").
|
|
163
|
-
- **cadence** (
|
|
248
|
+
- **cadence** (18) rhythm: repetition, parallelism, and the long-then-short kicker: `no-x-no-y`, `no-x-no-y-frag`, `did-not-x-did-not-y`, `one-x-one-y` ("one reviewer, one queue, one deadline"), `from-x-to-y-chain` ("from guessing to measuring, from hoping to knowing"), `same-determiner-chain` (any other repeated determiner, at `medium` confidence), `real-x-real-y`, `epistrophe` (off by default), `phrase-echo` (off by default), `is-is` (doubled copula), `the-x-is-the-x` ("the problem with A is the problem with B"), `rule-of-three` (off by default), `everyone-nobody` (the comma-spliced antithesis: "Everyone wants the dashboard, nobody maintains it."), `short-run` (three sentences of thirty characters or fewer in a row, at `medium` confidence), `mic-drop-closer` (the short quantifier-led sentence that ends a paragraph after a long one, at `medium` confidence), `bare-auxiliary-closer` (the same shape, but the closer's verb is elided down to a bare auxiliary: "The agent did.", at `medium` confidence), `np-fragment-and` (the verbless "A named owner and a quarterly review.", at `medium` confidence), `punch-sentence` (the verbless beat of three words or fewer between two long sentences, "Not anymore.", at `medium` confidence).
|
|
164
249
|
- **puffery** (8) inflates the subject: `puffery-words` (vibrant, nestled, groundbreaking, in the heart of), `rich-tapestry`, `vital-role`, `stands-serves-as`, `underscores-highlights`, `impact-noun-vague` ("a significant impact", "make an impact"), `trailing-significance-participle` (the "…, showcasing its importance" clause), `abstract-lives-in` ("the value sits in the follow-up", at `medium` confidence).
|
|
165
250
|
- **false-correction** (8) corrects a reading nobody offered: `not-just-x-but-y`, `not-x-but-y` (the bare corrective), `not-by-x-but-by-y` ("not by luck, but by design"), `isnt-x-its-y` (the same corrective split across two clauses: "It isn't the tool. It's the habit."), `question-isnt` (the corrective frame in interrogative dress), `less-about-more-about`, `actually-not-x`, `dont-verb-it`.
|
|
166
251
|
- **false-concession** (7) performs balance or candour and gives nothing up: `two-things-true`, `none-of-this-is-to-say`, `is-real-and-not`, `not-nothing`, `vague-attribution` ("some critics argue", "it is widely regarded"), `if-im-being-honest` (the candor preamble, from slopwash.com's "false intimacy"), `and-thats-fine`.
|
data/docs/SPEC.md
CHANGED
|
@@ -47,6 +47,9 @@ so the source never needs to enter the repository.
|
|
|
47
47
|
- **A reference corpus of human prose must be public domain.** Copyrighted text
|
|
48
48
|
can't be redistributed, so a corpus built from it can't live in the repo, and
|
|
49
49
|
neither can the false-positive check that depends on it.
|
|
50
|
+
- **The probe corpus is fetched, not committed.** `script/probe-raid fetch` pulls
|
|
51
|
+
RAID (Dugan et al. 2024, MIT) into `.corpus/`, which is ignored. The script
|
|
52
|
+
and the numbers it prints may be quoted in a commit message; the texts may not.
|
|
50
53
|
|
|
51
54
|
This repository is MIT-licensed, so anything committed is redistributable by
|
|
52
55
|
anyone. Writing the fixtures ourselves also means nobody's prose gets held up as
|
|
@@ -93,7 +96,8 @@ dependencies at runtime.
|
|
|
93
96
|
|
|
94
97
|
Rationale: rules are regexes; the whole thing is a scanner plus an output
|
|
95
98
|
formatter. A dependency-free `gem install sloplint` is the robust, boring
|
|
96
|
-
choice. The
|
|
99
|
+
choice. The one exception is `check --judge`, which loads the separate
|
|
100
|
+
sloplint-judge gem and calls a model over the network; see `docs/JUDGE.md`. The name is free on RubyGems (taken on PyPI, so this also sidesteps the
|
|
97
101
|
collision). Ships a `sloplint` executable from `exe/` (claude.ai rejects a
|
|
98
102
|
plugin with a top-level `bin/`, and sloplint ships as a Claude Code plugin
|
|
99
103
|
too).
|
|
@@ -121,7 +125,7 @@ Rakefile # rake spec
|
|
|
121
125
|
```
|
|
122
126
|
|
|
123
127
|
RSpec is a **development** dependency (in the gemspec's `add_development_
|
|
124
|
-
dependency`), so the runtime stays dependency-free.
|
|
128
|
+
dependency`), so the runtime stays dependency-free unless `--judge` is passed.
|
|
125
129
|
|
|
126
130
|
## CLI surface
|
|
127
131
|
|
|
@@ -157,13 +161,14 @@ Deliberately left out of v1 (add when a real need shows up, not before):
|
|
|
157
161
|
|
|
158
162
|
## Exit codes
|
|
159
163
|
|
|
160
|
-
|
|
164
|
+
Four codes carry the contract. A crash just exits nonzero on its own.
|
|
161
165
|
|
|
162
166
|
| code | meaning |
|
|
163
167
|
|------|---------|
|
|
164
168
|
| 0 | ran, **no notes** |
|
|
165
169
|
| 1 | ran, **notes found** |
|
|
166
170
|
| 2 | bad arguments / usage error |
|
|
171
|
+
| 3 | `--judge` only: the model backend could not be reached, no notes written |
|
|
167
172
|
|
|
168
173
|
Empty or whitespace-only input is exit 2, like a mistyped rule id: a scan of
|
|
169
174
|
nothing must not report as a clean scan. The text is tested before
|
|
@@ -173,7 +178,9 @@ block still exits 0. Only when every source is empty.
|
|
|
173
178
|
## Note (the diagnostic object)
|
|
174
179
|
|
|
175
180
|
One match = one Note. JSON output is an array of these (or an object keyed by
|
|
176
|
-
path when multiple files are scanned).
|
|
181
|
+
path when multiple files are scanned). Under `--judge` the array or object sits
|
|
182
|
+
under `"notes"` beside a `"judge"` object with the backend name, request count
|
|
183
|
+
and token counts; see `docs/JUDGE.md` "Note".
|
|
177
184
|
|
|
178
185
|
```json
|
|
179
186
|
{
|
|
@@ -246,7 +253,9 @@ RULES = [
|
|
|
246
253
|
```
|
|
247
254
|
|
|
248
255
|
Adding a rule = appending one entry + one bad and one ok fixture. That's the
|
|
249
|
-
whole extension story.
|
|
256
|
+
whole extension story. The only thing that plugs in is sloplint-judge, and it
|
|
257
|
+
plugs in by name: `check --judge` does a lazy `require "sloplint/judge"` and
|
|
258
|
+
exits 2 with an install hint when it is missing. See `docs/JUDGE.md`.
|
|
250
259
|
|
|
251
260
|
### Why Ruby literals, not JSON/YAML
|
|
252
261
|
|
|
@@ -313,6 +322,7 @@ The writer grades their own prose or claim.
|
|
|
313
322
|
- `genuinely` — any "genuinely"; `low` confidence, off by default; no narrowing holds
|
|
314
323
|
- `the-punchline-is` — "the punchline is/:/?", "the honest answer/version is"; "short version" left out; ordinary writing
|
|
315
324
|
- `announced-takeaway` — colon-led label: "The pattern/lesson/takeaway…:"; sentence-initial
|
|
325
|
+
- `cataphoric-teaser` — the forward-pointing tease, sentence-initial, in three closed frames: an optional forward demonstrative ("here's", "here is", "this is", with one lead-in word allowed: "So here's"), a noun slot (what / the part / the thing / the bit / the secret), a scarcity subject (nobody, no one, most X, few people, hardly anyone, almost nobody, people), up to two adverbs, and a verb of telling, getting wrong, talking about, admitting or noticing: "Here's what nobody tells you about hiring.", "The part most people get wrong is the rollback." Without the demonstrative the sentence must reach a copula or a colon ("What nobody tells you is…"), since "What nobody tells you gets forgotten." is a report, and "they" is allowed only after a demonstrative, since a bare "What they don't realize is" refers to the people in the story. "That's what nobody understands" points backward and is left alone.
|
|
316
326
|
|
|
317
327
|
### closer
|
|
318
328
|
|
|
@@ -351,7 +361,7 @@ Rhythm: repetition, parallelism, and the long-then-short kicker.
|
|
|
351
361
|
- `short-run` — three consecutive sentences of thirty characters or fewer, each letter-led and closing on a full stop, no quotation marks or digits, no lone-letter labels (either case, though a possessive is not one) or abbreviations, starting at a real sentence boundary (never on a wrap continuation), crossing a hard wrap but not a paragraph break; `medium` confidence. A list marker may open a run but never sit inside one, so consecutive bullets are a list — the marker set covers glyphs, the literal "o" plain-text documents use as a bullet, and numbered and lettered items. Two whole-run exclusions keep document furniture out, each a property of the run rather than of any sentence in it: three single-word sentences in a row is a citation line ("Natl. Inst. Stand. Technol."), and three lowercase-led sentences in a row is transcribed speech ("we all get gas. we go to divert to Albany."). One single-word sentence is the archetypal kicker and stays, and a run that reaches a capital anywhere is prose, so identifier-initial writing ("npm was slow. git blame helped. We moved on.") is untouched. One is a question; a draft that repeats it is the tell.
|
|
352
362
|
- `mic-drop-closer` — a sentence of 60+ characters, then a paragraph-final closer of two to eight words opening on a **quantifier** (Nothing, Most, None, Everything, Everyone, Nobody, Then, Neither, Both): "Nothing here needs a new login."; `medium` confidence. The bare demonstratives (That, This, It) were in the list and are out: procedural writing ends a step with one as a matter of course ("This completes the roughing operations."). A blank line or the end of the text must follow the closer; both sentences may be hard-wrapped; whitespace runs in the long sentence are capped so blanked Markdown cannot make it, and the long-sentence prefix is an atomic group so an unpunctuated stretch cannot send it into catastrophic backtracking. The note points at the closer. One means nothing; a draft where it repeats is the tell.
|
|
353
363
|
- `bare-auxiliary-closer` — the same long-sentence-then-short-closer shape as `mic-drop-closer`, but the tell sits in the verb rather than the subject: the closer's verb phrase is elided down to a bare auxiliary with no object, "The agent did."; `medium` confidence. No subject list is needed — a one-to-three word subject runs straight into a bare `did/does/do/was/were/is/are/had/has/can/could/would/will/should/might/must` (contracted forms included) and a period, and the closer must be the last thing in the paragraph, same as `mic-drop-closer`, so a closing quotation mark after the period excludes quoted dialogue the same way `short-run` excludes it. A negative lookahead drops a closer that still holds "what", "that", "which", "who", "why" or "how", since those introduce a subordinate clause supplying its own complement rather than an elided one ("Nobody knew who did." asks who did it). Reuses `mic-drop-closer`'s long-sentence prefix rather than a second copy of it. One means nothing; a draft where it repeats is the tell.
|
|
354
|
-
- `np-fragment-and` — a whole sentence made of two noun phrases and an "and", opening on A/An/One at a sentence start or after a list marker ("A named owner and a quarterly review."). One to three words a side, no auxiliary or modal anywhere (contractions included); a lexical verb is invisible, so "A car and a truck collided." flags, which is why it ships at `medium` confidence.
|
|
364
|
+
- `punch-sentence` — the same long-then-short shape in the middle of a paragraph, where the closer rules refuse it: after the shared sixty-character prefix, a verbless beat of one to three words closing on a full stop, then one or two spaces and a sentence of forty or more characters on the same paragraph ("…three nights running that week. Not anymore. The job now checks…"); `info`, `medium` confidence. The beat is one of two closed shapes: a negator (Not, No, Never, Nothing, None) followed by up to two lowercase words, or a single capitalised word of five letters or more ("Simple."). The stop before the beat may not close a capitalised word of one to four letters or a lone lowercase letter, so a title, an initial or a case citation ("Mrs. Cadaver.", "John F. Kennedy.", "Swift v. Tyson.") never yields a one-word sentence. Short clauses with a subject ("He agrees.", "Dawes refuses.") are how plot summaries and news briefs move and are not matched; a one-word imperative in procedural prose ("Drain.") is, at about one per 25,000 words of recipes. A run of short sentences is `short-run`'s, and a paragraph-final beat is `mic-drop-closer`'s or `bare-auxiliary-closer`'s.- `np-fragment-and` — a whole sentence made of two noun phrases and an "and", opening on A/An/One at a sentence start or after a list marker ("A named owner and a quarterly review."). One to three words a side, no auxiliary or modal anywhere (contractions included); a lexical verb is invisible, so "A car and a truck collided." flags, which is why it ships at `medium` confidence.
|
|
355
365
|
|
|
356
366
|
### puffery
|
|
357
367
|
|
|
@@ -472,6 +482,21 @@ RSpec (dev dependency), run via `rake spec`.
|
|
|
472
482
|
- LSP server mode, editor plugins, autofix/rewrite. sloplint *flags*; the agent
|
|
473
483
|
rewrites. Autofix is a separate tool if ever.
|
|
474
484
|
- Non-English. Languages other than English are a v2 conversation.
|
|
485
|
+
- Grammar-frequency tells from the corpus-linguistics literature. Reinhart et
|
|
486
|
+
al. (PNAS 2025) measure GPT-4o at 2.1 times the human rate for
|
|
487
|
+
nominalizations and 2.6 times for "that" relative clauses on a subject noun
|
|
488
|
+
("a framework that enables real-time analysis"). Both shifts are real in RAID
|
|
489
|
+
and in Claude Sonnet 5 text generated from RAID's prompts, and neither is a
|
|
490
|
+
shape a regex can flag. A single that-relative is ordinary English ("a move
|
|
491
|
+
that is likely to put a dent in its accounts", human BBC news) and human
|
|
492
|
+
news carries one every 650 words; two in one sentence separate no better,
|
|
493
|
+
and flagging either pushes a writer toward the nominalization instead.
|
|
494
|
+
Nominalization is a density, not a sentence shape: five or more suffix nouns
|
|
495
|
+
in one sentence fires on human abstracts at the 2023 model rate, and 50k
|
|
496
|
+
words of README prose gave 22 hits, every one a bullet list or feature
|
|
497
|
+
table. Both belong to a judge that reads a paragraph, not to this catalog.
|
|
498
|
+
Sentence-initial clausal subjects ("That the cache was stale is not in
|
|
499
|
+
dispute.") were also tried and are near zero on both sides everywhere.
|
|
475
500
|
- ML/embedding-based detection. This is a regex linter on purpose — fast,
|
|
476
501
|
explainable, zero-dependency. Statistical detection is a different product.
|
|
477
502
|
- Scraped corpora. Platform terms prohibit automated collection, and finding a
|
data/exe/sloplint
CHANGED
|
@@ -9,5 +9,6 @@ if (RUBY_VERSION.split(".").map(&:to_i).first(2) <=> [3, 3]) < 0
|
|
|
9
9
|
"macOS ships Ruby 2.6 at /usr/bin/ruby. Install a current one with `brew install ruby`."
|
|
10
10
|
end
|
|
11
11
|
|
|
12
|
-
|
|
12
|
+
$LOAD_PATH.unshift(File.expand_path("../lib", __dir__))
|
|
13
|
+
require "sloplint/cli"
|
|
13
14
|
exit Sloplint::CLI.run(ARGV)
|