simple_english 0.1.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 7c05cc78b762e2b44f4c841ef5a78537e91d9bcd408ed8739fcc00f3538f14eb
4
- data.tar.gz: 9abac46821a9260b3ec7312a1c258464e08a943262eaccab893383104479dcee
3
+ metadata.gz: 5c1efe5da86f44ce1618aea7d2c49f3e5dd0aec897a8f103ace5276be73d9ef4
4
+ data.tar.gz: cd8597b07d037a0631d16ea11377144fb9b93d6f6e604a316027d5c43a72d50d
5
5
  SHA512:
6
- metadata.gz: 6c4ac5c79f9a62c848ae37977dfb48ec2b815b27e5a38af7ab448a489206b14b8b26dec4f560b5115dd5be3c3688bfed254ebecbfd5e57c83c84d4a545367780
7
- data.tar.gz: 699067a5a056a54186fe40d62ee17dd358492409dd86a76fc17cd2e1be8d110ee1147daa3ecb5f4d8b995cfd98f4108e96adbf6ed96603a43566ee820fd48729
6
+ metadata.gz: 838420d06071034d7e6d48c074fc0d5893b6443e6f1dd387247e86239f74b620ff7e4adecd69e4b37a3d58e98ebb193bbe53a113dc96759a075cb08fdd8247c7
7
+ data.tar.gz: 7fcc415dcf6cc5a01ba01807118ca6e3ea8ff7f7d4573960f5b31fe7971c3612559e15d5f7ee0c122363cdf2b620b8ede79ed7b0e0e193dd188a0446483045fc
data/README.md CHANGED
@@ -1,3 +1,7 @@
1
+ [![CI](https://github.com/TonyCTHsu/simple-english/actions/workflows/ci.yml/badge.svg)](https://github.com/TonyCTHsu/simple-english/actions/workflows/ci.yml) <!-- se: ignore=SE_PARAGRAPH_TOO_LONG -->
2
+ [![Gem Version](https://badge.fury.io/rb/simple_english.svg)](https://badge.fury.io/rb/simple_english)
3
+ [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://github.com/TonyCTHsu/simple-english/blob/master/LICENSE)
4
+
1
5
  # simple_english
2
6
 
3
7
  > Write for human readers, not for reviewers or another AI.
@@ -7,16 +11,18 @@ the slop it leaves behind. Every finding says what to write
7
11
  instead.
8
12
 
9
13
  ```console
10
- $ printf 'You should leverage this tool in order to make sure that your docs are readable.' > note.md
14
+ $ printf "The config was written by setup — don't edit it; the daemon caches rules, making the first lint slow." > note.md
11
15
  $ se note.md
12
- note.md:1: [SE_MODAL_RESTRICTED] Use can, will, or must. State the requirement exactly.
13
- note.md:1: [SE_SLOP_IN_ORDER_TO] Write "to".
14
- note.md:1: [SE_SLOP_LEVERAGE] Write "use".
16
+ note.md:1:12-26: [SE_ACTIVE_VOICE] "was written by" - Use the active voice. Say who does the action.
17
+ note.md:1:73-81: [SE_ING_AFTER_COMMA] ", making" - Start a new sentence instead of the -ing phrase.
18
+ note.md:1:37-40: [SE_NO_CONTRACTIONS] "n't" - Write the words in full. No contractions.
19
+ note.md:1:33-34: [SE_NO_EMDASH] "—" - Write two sentences, or use a comma.
20
+ note.md:1:48-49: [SE_NO_SEMICOLON] ";" - Write two sentences, or name the relation.
15
21
  ```
16
22
 
17
23
  Markdown prose plus code comments in Python, Ruby, JavaScript,
18
- TypeScript, Go, Rust, Java, C#, Kotlin, bash, and YAML. Output as plain
19
- text, JSON, or SARIF.
24
+ TypeScript, Go, Rust, Java, C#, C++, Kotlin, bash, and YAML. Output as
25
+ plain text, JSON, or SARIF.
20
26
 
21
27
  ## The rules
22
28
 
@@ -27,7 +33,7 @@ text, JSON, or SARIF.
27
33
  - **Contractions:** write every word in full.
28
34
  - **Sentence shape:** condition before command, no `-ing` phrase after a comma.
29
35
  - **Word choice:** about 50 substitution rules, from `leverage` to `in conclusion`. `make sure that` keeps its "that".
30
- - **Code comments:** same pattern rules, with line and column.
36
+ - **Code comments:** same pattern rules, with line and column range.
31
37
  - **Counts (Markdown only):** 20 words per sentence in list items, 25 in paragraphs, six sentences per paragraph at most.
32
38
 
33
39
  The full list, with a wrong and a right example for each rule:
@@ -89,7 +95,11 @@ git diff --name-only --diff-filter=ACM main | xargs se
89
95
 
90
96
  ### Outputs
91
97
 
92
- Findings print as `file:line: [RULE_ID] message`. Code-comment findings also carry a column in `--format json` and `--format sarif`:
98
+ Pattern findings print an exclusive column range as
99
+ `file:line:start-end: [RULE_ID] message`. A range that crosses lines ends
100
+ with `end-line:end-column`. Columns count UTF-16 code units. Counting findings
101
+ identify only the paragraph's first line. JSON and SARIF output expose the same
102
+ source ranges:
93
103
 
94
104
  ```bash
95
105
  se --format json docs/
@@ -176,7 +186,7 @@ Any tool or language can lint through its HTTP API:
176
186
 
177
187
  ```bash
178
188
  curl -d "text=Don't do this." http://localhost:8181/lint
179
- # [{"line":1,"column":null,"rule":"SE_NO_CONTRACTIONS","message":"..."}]
189
+ # [{"line":1,"column":3,"end_line":1,"end_column":6,"rule":"SE_NO_CONTRACTIONS","message":"..."}]
180
190
  ```
181
191
 
182
192
  The full wire format: [docs/DAEMON.md](docs/DAEMON.md).
@@ -122,6 +122,7 @@ module SimpleEnglish
122
122
  when "json"
123
123
  puts JSON.pretty_generate(results.map do |path, finding|
124
124
  {"path" => path, "line" => finding.line, "column" => finding.column,
125
+ "end_line" => finding.end_line, "end_column" => finding.end_column,
125
126
  "rule" => finding.rule, "message" => finding.message}
126
127
  end)
127
128
  when "sarif"
@@ -129,10 +130,17 @@ module SimpleEnglish
129
130
  "version" => "2.1.0",
130
131
  "$schema" => "https://json.schemastore.org/sarif-2.1.0.json",
131
132
  "runs" => [{
133
+ "columnKind" => "utf16CodeUnits",
132
134
  "tool" => {"driver" => {"name" => "se"}},
133
135
  "results" => results.map do |path, finding|
134
136
  region = {"startLine" => finding.line}
135
- region["startColumn"] = finding.column if finding.column
137
+ if finding.column
138
+ region["startColumn"] = finding.column
139
+ if finding.end_line && finding.end_column
140
+ region["endLine"] = finding.end_line
141
+ region["endColumn"] = finding.end_column
142
+ end
143
+ end
136
144
  {"ruleId" => finding.rule, "level" => "error",
137
145
  "message" => {"text" => finding.message},
138
146
  "locations" => [{"physicalLocation" => {
@@ -144,7 +152,16 @@ module SimpleEnglish
144
152
  )
145
153
  else
146
154
  results.each do |path, finding|
147
- puts "#{path}:#{finding.line}: [#{finding.rule}] #{finding.message}"
155
+ location = "#{path}:#{finding.line}"
156
+ if finding.column
157
+ location += ":#{finding.column}"
158
+ if finding.end_line && finding.end_column
159
+ finish = (finding.end_line == finding.line) ? finding.end_column :
160
+ "#{finding.end_line}:#{finding.end_column}"
161
+ location += "-#{finish}"
162
+ end
163
+ end
164
+ puts "#{location}: [#{finding.rule}] #{finding.message}"
148
165
  end
149
166
  end
150
167
  end
@@ -14,19 +14,23 @@ module SimpleEnglish
14
14
  module Client
15
15
  module_function
16
16
 
17
- # LanguageTool reports offsets in Java UTF-16 code units.
18
- # Astral characters (emoji) count as two. Count them so a
19
- # newline comparison cannot drift past a boundary. A match
20
- # starting at the newline itself belongs to the next line.
21
- def offset_to_line(text, offset)
17
+ # LanguageTool reports offsets and lengths in Java UTF-16 code units.
18
+ # Return a 1-based line and UTF-16 column for its 0-based offset.
19
+ def offset_to_position(text, offset)
22
20
  line = 1
21
+ column = 1
23
22
  units = 0
24
23
  text.each_char do |char|
25
- line += 1 if char == "\n" && units <= offset
26
- return line if units >= offset
24
+ return [line, column] if units >= offset
27
25
  units += (char.ord > 0xFFFF) ? 2 : 1
26
+ if char == "\n"
27
+ line += 1
28
+ column = 1
29
+ else
30
+ column += (char.ord > 0xFFFF) ? 2 : 1
31
+ end
28
32
  end
29
- line
33
+ [line, column]
30
34
  end
31
35
 
32
36
  DEFAULT_PORT = 8181
@@ -62,17 +66,45 @@ module SimpleEnglish
62
66
  def parse_matches(matches, payload)
63
67
  payload = to_payload(payload)
64
68
  matches.map do |match|
65
- line, column = payload.locate(match.fetch("offset"))
69
+ offset = match.fetch("offset")
70
+ line, column = payload.locate(offset)
71
+ end_line, end_column = payload.locate(offset + match.fetch("length"))
66
72
  Finding.new(line: line, column: column,
67
- rule: match.fetch("rule").fetch("id"), message: match.fetch("message"))
73
+ end_line: end_line, end_column: end_column,
74
+ rule: match.fetch("rule").fetch("id"),
75
+ message: with_context(match.fetch("message"), match))
68
76
  end
69
77
  end
70
78
 
79
+ # Prefix the offending text, so a finding says what to change,
80
+ # not only how. Context positions use Java UTF-16 code units.
81
+ def with_context(message, match)
82
+ context = match["context"] or return message
83
+ matched = utf16_slice(context.fetch("text", ""),
84
+ context.fetch("offset", 0).to_i, context.fetch("length", 0).to_i)
85
+ matched.empty? ? message : "\"#{matched}\" - #{message}"
86
+ end
87
+
88
+ def utf16_slice(text, offset, length)
89
+ first = utf16_index(text, offset)
90
+ last = utf16_index(text, offset + length)
91
+ text[first...last]
92
+ end
93
+
94
+ def utf16_index(text, offset)
95
+ units = 0
96
+ text.each_char.with_index do |char, index|
97
+ return index if units >= offset
98
+ units += (char.ord > 0xFFFF) ? 2 : 1
99
+ end
100
+ text.length
101
+ end
102
+
71
103
  def to_payload(payload)
72
104
  payload.is_a?(String) ? PlainText.new(payload) : payload
73
105
  end
74
106
 
75
- private_class_method :to_payload
107
+ private_class_method :to_payload, :with_context, :utf16_slice, :utf16_index
76
108
 
77
109
  # Full lint via the se daemon. Raw Markdown in, or code
78
110
  # source with a language for the comment pipeline. nil when
@@ -87,6 +119,7 @@ module SimpleEnglish
87
119
  return nil unless body.is_a?(Array)
88
120
  body.map do |hash|
89
121
  Finding.new(line: hash.fetch("line"), column: hash["column"],
122
+ end_line: hash["end_line"], end_column: hash["end_column"],
90
123
  rule: hash.fetch("rule"), message: hash.fetch("message"))
91
124
  end
92
125
  rescue SystemCallError, SocketError, Timeout::Error,
@@ -109,15 +142,15 @@ module SimpleEnglish
109
142
  # caller owns the exit status. Diagnostics go to stderr here, where the cause
110
143
  # is known.
111
144
  def ensure_up(install: SimpleEnglish::Install.from_env)
112
- unless File.exist?(install.server_jar)
113
- warn install.setup_error
114
- return false
115
- end
116
145
  return true if up?
117
146
  if ENV["SE_SERVER_URL"]
118
147
  warn "error: SE_SERVER_URL is set but #{url} does not answer."
119
148
  return false
120
149
  end
150
+ unless File.exist?(install.server_jar)
151
+ warn install.setup_error
152
+ return false
153
+ end
121
154
  warn "se: daemon not running; starting it (first lint takes ~15s)..."
122
155
  bin = File.expand_path("../../bin/se", __dir__)
123
156
  Process.spawn(RbConfig.ruby, bin, "serve", out: File::NULL, err: File::NULL)
@@ -20,7 +20,12 @@ module SimpleEnglish
20
20
  ".java" => "java",
21
21
  ".sh" => "bash",
22
22
  ".kt" => "kotlin",
23
- ".cs" => "csharp"
23
+ ".cs" => "csharp",
24
+ ".cpp" => "cpp",
25
+ ".cc" => "cpp",
26
+ ".cxx" => "cpp",
27
+ ".hpp" => "cpp",
28
+ ".h" => "cpp"
24
29
  }.freeze
25
30
 
26
31
  module_function
@@ -1,8 +1,9 @@
1
1
  # frozen_string_literal: true
2
2
 
3
- # One lint result. The line and column are 1-based. The column is
4
- # nil when the rule has no position within the line.
3
+ # One lint result. Positions are 1-based. The end position is exclusive.
4
+ # Counting findings have no columns or end positions. Findings from an older
5
+ # daemon can have a start column without an end position.
5
6
 
6
7
  module SimpleEnglish
7
- Finding = Struct.new(:line, :column, :rule, :message)
8
+ Finding = Struct.new(:line, :column, :end_line, :end_column, :rule, :message)
8
9
  end
@@ -9,7 +9,10 @@ require_relative "paragraph"
9
9
 
10
10
  module SimpleEnglish
11
11
  module Markdown
12
- SENTENCE_END = /[.!?]/
12
+ # A terminator ends a sentence only before whitespace or at the
13
+ # end of the text, so periods inside URLs and file names do not
14
+ # split sentences.
15
+ SENTENCE_END = /(?<=[.!?])(?:\s+|$)/
13
16
  # Non-prose blocks: their lines never reach the counting rules or
14
17
  # LanguageTool.
15
18
  BLANKED_BLOCKS = %w[fenced_code_block indented_code_block thematic_break].freeze
@@ -18,16 +21,37 @@ module SimpleEnglish
18
21
  module_function
19
22
 
20
23
  # Blank out code and other non-prose blocks and neutralize inline
21
- # code and heading markers. The line count stays identical to the
22
- # source.
24
+ # code and heading markers. Line count and UTF-16 widths stay identical
25
+ # to the source so LanguageTool offsets remain source positions.
23
26
  def strip(text)
24
- lines = text.lines.map(&:chomp)
25
- blank_rows(text).each { |row| lines[row] = "" }
26
- lines.map { |line| strip_line(line) }.join("\n")
27
+ blank = blank_rows(text).to_h { |row| [row, true] }
28
+ text.lines.each_with_index.map do |line, row|
29
+ ending = if line.end_with?("\r\n")
30
+ "\r\n"
31
+ elsif line.end_with?("\n")
32
+ "\n"
33
+ else
34
+ ""
35
+ end
36
+ body = line.delete_suffix(ending)
37
+ (blank[row] ? spaces_for(body) : strip_line(body)) + ending
38
+ end.join
27
39
  end
28
40
 
29
41
  def strip_line(line)
30
- line.gsub(/`[^`]*`/, "X").sub(/\A\#{1,6} /, "")
42
+ # Width-preserving replacements keep LanguageTool positions aligned
43
+ # with the source.
44
+ line
45
+ .gsub(/`[^`]*`/) { |code| "X" * utf16_length(code) }
46
+ .sub(/\A\#{1,6} /) { |marker| " " * marker.length }
47
+ end
48
+
49
+ def spaces_for(text)
50
+ " " * utf16_length(text)
51
+ end
52
+
53
+ def utf16_length(text)
54
+ text.each_char.sum { |char| (char.ord > 0xFFFF) ? 2 : 1 }
31
55
  end
32
56
 
33
57
  # A vertical list is not one paragraph: each list item is its own.
@@ -91,7 +115,7 @@ module SimpleEnglish
91
115
  TreeSitterLanguagePack.get_parser("markdown").parse(text).root_node
92
116
  end
93
117
 
94
- private_class_method :strip_line, :node_rows, :blank_rows, :each_node,
95
- :inside_list_item?, :parse
118
+ private_class_method :strip_line, :spaces_for, :utf16_length, :node_rows,
119
+ :blank_rows, :each_node, :inside_list_item?, :parse
96
120
  end
97
121
  end
@@ -10,7 +10,7 @@ module SimpleEnglish
10
10
  end
11
11
 
12
12
  def locate(utf16_offset)
13
- [Client.offset_to_line(text, utf16_offset), nil]
13
+ Client.offset_to_position(text, utf16_offset)
14
14
  end
15
15
  end
16
16
  end
@@ -59,10 +59,12 @@ module SimpleEnglish
59
59
  rules_dir = Dir.mktmpdir("se-rules")
60
60
  stage_rules(rules_dir)
61
61
  # The inner JVM's stderr goes to a file so failure messages can quote
62
- # its first line. The dev log has a path. $stderr does not, so fall
63
- # back to a temp file that Ruby unlinks when the process exits.
62
+ # its first line. Only an explicitly opened dev log (a File) is reused
63
+ # for that. $stderr reports path "<STDERR>", so it creates a file
64
+ # by that name in the CWD. Anything else falls back to a temp file
65
+ # that Ruby unlinks when the process exits.
64
66
  inner_log_path =
65
- if log.respond_to?(:path)
67
+ if log.is_a?(File)
66
68
  log.path
67
69
  else
68
70
  inner_log = Tempfile.new("se-inner")
@@ -0,0 +1,5 @@
1
+ # frozen_string_literal: true
2
+
3
+ module SimpleEnglish
4
+ VERSION = "0.3.0"
5
+ end
@@ -4,6 +4,7 @@
4
4
  # Pattern rules run on LanguageTool. Counting rules run here.
5
5
  # This file is the composition root. The pieces live in lib/simple_english/.
6
6
 
7
+ require_relative "simple_english/version"
7
8
  require_relative "simple_english/finding"
8
9
  require_relative "simple_english/markdown"
9
10
  require_relative "simple_english/counts"
@@ -19,8 +20,6 @@ require_relative "simple_english/client"
19
20
  require_relative "simple_english/server"
20
21
 
21
22
  module SimpleEnglish
22
- VERSION = "0.1.0"
23
-
24
23
  module_function
25
24
 
26
25
  # Returns findings, or nil when the daemon is unreachable. The
@@ -56,10 +56,14 @@
56
56
  </rule>
57
57
 
58
58
  <rule id="SE_ING_AFTER_COMMA" name="No -ing after a comma">
59
- <pattern><token>,</token><token postag="VBG"/></pattern>
59
+ <pattern><token>,</token><token postag="VBG"><exception scope="next" regexp="yes">[,)]|and|or</exception></token></pattern>
60
60
  <message>Start a new sentence instead of the -ing phrase.</message>
61
61
  <example type="incorrect">The tool runs<marker>, making</marker> it easy.</example>
62
62
  <example type="correct">The tool runs. It is easy.</example>
63
+ <example type="correct">The walkthrough (accounts, VPC, secrets, troubleshooting) is in the guide.</example>
64
+ <example type="correct">The guide covers billing, troubleshooting, and support for each environment.</example>
65
+ <example type="correct">The guide covers billing, troubleshooting and support for each environment.</example>
66
+ <example type="correct">The guide covers billing, troubleshooting or support for each environment.</example>
63
67
  </rule>
64
68
 
65
69
  <rule id="SE_KEEP_THAT" name="make sure that">
metadata CHANGED
@@ -1,14 +1,14 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: simple_english
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.1.0
4
+ version: 0.3.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - TonyCTHsu
8
8
  autorequire:
9
9
  bindir: bin
10
10
  cert_chain: []
11
- date: 2026-09-28 00:00:00.000000000 Z
11
+ date: 2026-09-29 00:00:00.000000000 Z
12
12
  dependencies:
13
13
  - !ruby/object:Gem::Dependency
14
14
  name: tree_sitter_language_pack
@@ -72,14 +72,28 @@ dependencies:
72
72
  requirements:
73
73
  - - "~>"
74
74
  - !ruby/object:Gem::Version
75
- version: '5.0'
75
+ version: '6.0'
76
76
  type: :development
77
77
  prerelease: false
78
78
  version_requirements: !ruby/object:Gem::Requirement
79
79
  requirements:
80
80
  - - "~>"
81
81
  - !ruby/object:Gem::Version
82
- version: '5.0'
82
+ version: '6.0'
83
+ - !ruby/object:Gem::Dependency
84
+ name: minitest-mock
85
+ requirement: !ruby/object:Gem::Requirement
86
+ requirements:
87
+ - - "~>"
88
+ - !ruby/object:Gem::Version
89
+ version: '5.27'
90
+ type: :development
91
+ prerelease: false
92
+ version_requirements: !ruby/object:Gem::Requirement
93
+ requirements:
94
+ - - "~>"
95
+ - !ruby/object:Gem::Version
96
+ version: '5.27'
83
97
  - !ruby/object:Gem::Dependency
84
98
  name: json
85
99
  requirement: !ruby/object:Gem::Requirement
@@ -140,11 +154,15 @@ files:
140
154
  - lib/simple_english/server.rb
141
155
  - lib/simple_english/span.rb
142
156
  - lib/simple_english/suppressions.rb
157
+ - lib/simple_english/version.rb
143
158
  - rules/simple-english.xml
144
159
  homepage: https://github.com/TonyCTHsu/simple-english
145
160
  licenses:
146
161
  - MIT
147
- metadata: {}
162
+ metadata:
163
+ homepage_uri: https://github.com/TonyCTHsu/simple-english
164
+ source_code_uri: https://github.com/TonyCTHsu/simple-english
165
+ changelog_uri: https://github.com/TonyCTHsu/simple-english/blob/v0.3.0/CHANGELOG.md
148
166
  post_install_message:
149
167
  rdoc_options: []
150
168
  require_paths: