json-repair 0.12.0 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: aef10e86ea82fb56d9666ad7470317cf751f00730ef51910094ce8b2ee876a53
4
- data.tar.gz: 682f0cdacc02896687e6c39e534f92f0beda52110679da8b7b7f43721aa6c2a4
3
+ metadata.gz: 3e0898f751c91f7c3a026ccb107c319b983212296090031ba6b7a80713de18eb
4
+ data.tar.gz: 21028aa56a685b77ea42506333753deb1a903108b860843a877181b6af66b1a2
5
5
  SHA512:
6
- metadata.gz: 572225e5c09ac6ab7795d21d179e9ac07bc2c3bffedb5f1df48afa5a33ab8a73923af90429730532a822f3a78004181f3a9d2d7a2bddf17b9b649819b169ec65
7
- data.tar.gz: a62311d2002b538b81132f6efb8525e6c6801baaa5a8f9eac020abbc2726e5ed783ee2719aeed9dc77301492f869f79deb807b9eee74f3a56e0a37d4712b7279
6
+ metadata.gz: 05752fda1b0ff27c2dcde8b217099460bdecb57efa4a3ab73ce534f27dfecf51673f68683f3009c8d159079bfb7fee3f0892f2c86e3f15d0e86a58d4c497c8f0
7
+ data.tar.gz: 46a8a689d2a969c3c7ca1395dbd4a5555e191e304c0cfffd3bbaec44d22066b2a8527efc7323267c62c21738364c2dddf917d20af3388ae87cf993d872d014a0
data/CHANGELOG.md CHANGED
@@ -1,5 +1,42 @@
1
1
  # Changes
2
2
 
3
+ ### 2026-06-17 (0.14.0)
4
+
5
+ * Repair a leading `+` on numbers, like in JSON5: `+1.23` → `1.23`,
6
+ `[+1]` → `[1]`, `{"a": +5}` → `{"a":5}`, `+.5` → `0.5`. The `+` is
7
+ consumed and dropped (JSON has no explicit plus), so a leading-zero
8
+ number is quoted exactly like its unsigned form (`+05` → `"05"`),
9
+ while a `+` inside an exponent is left untouched (`2e+5` →
10
+ `200000.0`). The `+` must be followed by a digit or a leading dot, so
11
+ a bare `+` (and `+abc`, `++1`, `+e5`) still raises — and requiring a
12
+ digit/dot keeps the 0.12.0 empty-mantissa exponent guard intact.
13
+ Divergence from upstream
14
+ [jsonrepair](https://github.com/josdejong/jsonrepair) (v3.14.0 leaves
15
+ `+1.23` unrepaired, raising "Unexpected end of json string"),
16
+ commented at the site. Mirrors the leading-dot number repair (0.9.0).
17
+
18
+ ### 2026-06-12 (0.13.0)
19
+
20
+ * Repair `#` hash line comments, like in Python, YAML, or Hjson:
21
+ `{"a": 1 # comment\n}` → `{"a":1}`, `{ # note\n "a": 1}` →
22
+ `{"a":1}`, `# lead\n{"a": 1}` → `{"a":1}`. Divergence from upstream
23
+ [jsonrepair](https://github.com/josdejong/jsonrepair) (v3.14.0
24
+ raises on all of these), commented at the site. Recognition is
25
+ context-aware so unquoted values starting with `#` keep repairing
26
+ into strings — `{"color": #ff0000}` → `{"color":"#ff0000"}`,
27
+ `{#tag: 1}` → `{"#tag":1}`, `#standalone` → `"#standalone"` — where
28
+ Python's `json_repair` silently loses them (`{"color": ""}`).
29
+ Where a value or key is expected, a `#` token reaching a structural
30
+ delimiter (`,` `}` `]` `:`) before any whitespace, or running to
31
+ end-of-input without a newline, stays a value; anything else is a
32
+ comment stripped to the end of the line. The tradeoff: a `#` token
33
+ followed by whitespace or a newline at a value position now reads
34
+ as a comment — `{"a": #b c}` → `{"a":null}` and `{"a": #tag\n}` →
35
+ `{"a":null}`, where 0.12.0 kept them as strings (Python drops them
36
+ too), and a comment-only document now raises like `// only a
37
+ comment` always has. Pinned in the spec suite as conscious
38
+ decisions.
39
+
3
40
  ### 2026-06-12 (0.12.0)
4
41
 
5
42
  * Repair the three known input families that raised `Internal error:
data/CLAUDE.md CHANGED
@@ -24,7 +24,7 @@ A few repair heuristics deliberately go **beyond** upstream (leading-dot numbers
24
24
 
25
25
  ### Entry point
26
26
 
27
- `JSON.repair(str)` in `lib/json/repair.rb` first tries stdlib `JSON.parse` (fast path; opt out with `skip_json_loads: true`), and falls back to `JSON::Repairer.new(str).repair` when that raises. Either way the result is re-serialized with `JSON.generate`, so **output is canonical** — whitespace collapsed, numbers normalized, duplicate keys last-write-wins — and both paths agree on it. A `REPAIR_REQUIRED_PATTERN` regex routes inputs containing comments or invalid escapes straight to the Repairer even though the bundled `json` gem would accept them. `return_objects: true` returns the parsed Ruby value instead of a string; `JSON.repair_file(path)` / `JSON.repair_io(io)` are convenience wrappers forwarding both options.
27
+ `JSON.repair(str)` in `lib/json/repair.rb` first tries stdlib `JSON.parse` (fast path; opt out with `skip_json_loads: true`), and falls back to `JSON::Repairer.new(str).repair` when that raises. Either way the result is re-serialized with `JSON.generate`, so **output is canonical** — whitespace collapsed, numbers normalized, duplicate keys last-write-wins — and both paths agree on it. There is no pre-filter regex: the fast path simply runs `JSON.parse` and falls back to the Repairer on any `JSON::ParserError`. Comments are handled either way — the bundled `json` gem currently accepts `//` and `/*` comments (an implementation detail, not part of the `JSON.parse` contract) and `JSON.generate` drops them (`{} // c` → `{}`); were a future `json` release to reject them, they would simply raise and route to the Repairer, which strips comments too. Invalid escapes, NDJSON, Markdown fences, and trailing prose already make `JSON.parse` raise and so route to the Repairer. `return_objects: true` returns the parsed Ruby value instead of a string; `JSON.repair_file(path)` / `JSON.repair_io(io)` are convenience wrappers forwarding both options.
28
28
 
29
29
  `JSON::JSONRepairError` is the only error raised for unrecoverable inputs; it exposes the failure `#position`. If the Repairer ever emits a string stdlib cannot parse (a Repairer bug), the `JSON::ParserError` is wrapped in `JSONRepairError` rather than leaked.
30
30
 
@@ -2,6 +2,6 @@
2
2
 
3
3
  module JSON
4
4
  module Repair
5
- VERSION = '0.12.0'
5
+ VERSION = '0.14.0'
6
6
  end
7
7
  end
data/lib/json/repairer.rb CHANGED
@@ -71,7 +71,7 @@ module JSON
71
71
  # repair redundant end quotes
72
72
  while [CLOSING_BRACE, CLOSING_BRACKET].include?(@json[@index])
73
73
  @index += 1
74
- parse_whitespace_and_skip_comments
74
+ parse_whitespace_and_skip_comments(value_expected: false)
75
75
  end
76
76
 
77
77
  if @index >= @json.length
@@ -93,17 +93,17 @@ module JSON
93
93
  parse_keywords ||
94
94
  parse_unquoted_string(false) ||
95
95
  parse_regex
96
- parse_whitespace_and_skip_comments
96
+ parse_whitespace_and_skip_comments(value_expected: false)
97
97
 
98
98
  process
99
99
  end
100
100
 
101
- def parse_whitespace_and_skip_comments(skip_newline: true)
101
+ def parse_whitespace_and_skip_comments(skip_newline: true, value_expected: true)
102
102
  start = @index
103
103
 
104
104
  changed = parse_whitespace(skip_newline: skip_newline)
105
105
  loop do
106
- changed = parse_comment
106
+ changed = parse_comment(value_expected: value_expected)
107
107
  changed = parse_whitespace(skip_newline: skip_newline) if changed
108
108
  break unless changed
109
109
  end
@@ -131,7 +131,7 @@ module JSON
131
131
  false
132
132
  end
133
133
 
134
- def parse_comment
134
+ def parse_comment(value_expected: true)
135
135
  if @json[@index] == '/' && @json[@index + 1] == '*'
136
136
  # Block comment
137
137
  @index += 2
@@ -143,11 +143,42 @@ module JSON
143
143
  @index += 2
144
144
  @index += 1 until @json[@index].nil? || @json[@index] == "\n"
145
145
  true
146
+ elsif @json[@index] == '#' && hash_comment?(value_expected)
147
+ # Hash line comment, like in Python, YAML, or Hjson (divergence
148
+ # from upstream, which raises on `#` as of v3.14.0)
149
+ @index += 1
150
+ @index += 1 until @json[@index].nil? || @json[@index] == "\n"
151
+ true
146
152
  else
147
153
  false
148
154
  end
149
155
  end
150
156
 
157
+ # Decide whether the `#` at @index starts a line comment or an
158
+ # unquoted value like {"color": #ff0000}, {#tag: 1}, or a root
159
+ # #hashtag (which Python's json_repair eats as comments, losing
160
+ # data). Where no value or key is expected an unquoted token would
161
+ # be junk anyway, so `#` is always a comment, exactly like `//`.
162
+ # Where one is expected, scan the rest of the line: a structural
163
+ # delimiter (`,` `}` `]` `:`) before any whitespace means the token
164
+ # reads as a value in context; whitespace (including the newline
165
+ # itself) first means comment prose; reaching EOF without a newline
166
+ # keeps the token a value, so truncated input like {"a": #tag is
167
+ # repaired, not dropped. Divergence from upstream, as above.
168
+ def hash_comment?(value_expected)
169
+ return true unless value_expected
170
+
171
+ i = @index + 1
172
+ while (char = @json[i])
173
+ return true if whitespace_or_special?(char)
174
+ return false if [COMMA, COLON, CLOSING_BRACE, CLOSING_BRACKET].include?(char)
175
+
176
+ i += 1
177
+ end
178
+
179
+ false
180
+ end
181
+
151
182
  # Find and skip over a Markdown fenced code block:
152
183
  # ``` ... ```
153
184
  # or
@@ -268,7 +299,7 @@ module JSON
268
299
  break
269
300
  end
270
301
 
271
- parse_whitespace_and_skip_comments
302
+ parse_whitespace_and_skip_comments(value_expected: false)
272
303
  processed_colon = parse_character(COLON)
273
304
  truncated_text = @index >= @json.length
274
305
  unless processed_colon
@@ -491,7 +522,7 @@ module JSON
491
522
  @index += 1
492
523
  @output << str
493
524
 
494
- parse_whitespace_and_skip_comments(skip_newline: false)
525
+ parse_whitespace_and_skip_comments(skip_newline: false, value_expected: false)
495
526
 
496
527
  if stop_at_delimiter ||
497
528
  @index >= @json.length ||
@@ -710,6 +741,24 @@ module JSON
710
741
  # Parse a number like 2.4 or 2.4e6
711
742
  def parse_number
712
743
  start = @index
744
+
745
+ # Divergence from upstream: accept and discard a leading "+". JSON5
746
+ # permits an explicit plus; JSON does not, so "+1.23" -> "1.23" and
747
+ # "{"a": +5}" -> {"a":5}. The "+" must be followed by a digit or a
748
+ # leading dot (mirroring the "-" branch below); otherwise this is not
749
+ # a number and we backtrack to the "+" so a bare "+" still raises.
750
+ # `start` stays on the "+", so the reset paths restore @index cleanly
751
+ # and the @index > start exponent guard below still implies a digit
752
+ # was consumed; the two emission sites drop the leading "+" before
753
+ # quoting. Upstream leaves "+1.23" unrepaired.
754
+ if @json[@index] == PLUS
755
+ @index += 1
756
+ unless digit?(@json[@index]) || @json[@index] == DOT
757
+ @index = start
758
+ return false
759
+ end
760
+ end
761
+
713
762
  if @json[@index] == '-'
714
763
  @index += 1
715
764
  if at_end_of_number?
@@ -771,7 +820,9 @@ module JSON
771
820
 
772
821
  if @index > start
773
822
  # repair a number with leading zeros like "00789"
774
- num = @json[start...@index]
823
+ # drop a leading "+" first (see the PLUS branch above), so "+05"
824
+ # quotes like "05" and "+1.23" emits as "1.23"
825
+ num = @json[start...@index].delete_prefix(PLUS)
775
826
  # the optional sign quotes "-05" like "05" (divergence from
776
827
  # upstream, whose unsigned check lets "-05" through unrepaired)
777
828
  has_invalid_leading_zero = num.match?(/^-?0\d/)
@@ -846,11 +897,11 @@ module JSON
846
897
  def parse_concatenated_string
847
898
  processed = false
848
899
 
849
- parse_whitespace_and_skip_comments
900
+ parse_whitespace_and_skip_comments(value_expected: false)
850
901
  while @json[@index] == PLUS
851
902
  processed = true
852
903
  @index += 1
853
- parse_whitespace_and_skip_comments
904
+ parse_whitespace_and_skip_comments(value_expected: false)
854
905
 
855
906
  # repair: remove the end quote of the first string
856
907
  @output = strip_last_occurrence(@output, '"', strip_remaining_text: true)
@@ -877,7 +928,8 @@ module JSON
877
928
  # repair numbers cut off at the end
878
929
  # this will only be called when we end after a '.', '-', or 'e' and does not
879
930
  # change the number more than it needs to make it valid JSON
880
- num = "#{@json[start...@index]}0"
931
+ # delete_prefix drops a leading "+" (see the PLUS branch in parse_number)
932
+ num = "#{@json[start...@index]}0".delete_prefix(PLUS)
881
933
  # quote a padded token that has an invalid leading zero, like "05e" ->
882
934
  # "05e0", applying the same rule as the end of parse_number (divergence
883
935
  # from upstream, which emits the invalid number raw)
@@ -37,7 +37,9 @@ module JSON
37
37
 
38
38
  def parse_whitespace: (?skip_newline: bool) -> bool
39
39
 
40
- def parse_comment: () -> bool
40
+ def parse_comment: (?value_expected: bool) -> bool
41
+
42
+ def hash_comment?: (bool value_expected) -> bool
41
43
 
42
44
  # Find and skip over a Markdown fenced code block
43
45
  def parse_markdown_code_block: (::Array[::String] blocks) -> bool
@@ -87,7 +89,7 @@ module JSON
87
89
 
88
90
  def parse_character: (::String char) -> bool
89
91
 
90
- def parse_whitespace_and_skip_comments: (?skip_newline: bool) -> bool
92
+ def parse_whitespace_and_skip_comments: (?skip_newline: bool, ?value_expected: bool) -> bool
91
93
 
92
94
  # Parse a number like 2.4 or 2.4e6
93
95
  def parse_number: () -> bool
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: json-repair
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.12.0
4
+ version: 0.14.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Aleksandr Zykov