json-repair 0.11.3 → 0.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +53 -0
- data/lib/json/repair/version.rb +1 -1
- data/lib/json/repairer.rb +69 -16
- data/sig/json/repairer.rbs +4 -2
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: cf5d06053ce264da6b60de4beef4a5718fd3482c81d295deea6dbf90fc63c098
|
|
4
|
+
data.tar.gz: 9d343f549cf34414e0de45618cb19bd4a9b58b12e53c2f52da0380fc73b4d95a
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 269f7de75ec6cadde857d2f7c10186b1d0c63a19aa3e863744957c95b98db8d0cef790deef8814d532ca62c0405d97cfc5f441c37be9f1c4e6c7f5cfe07eaa67
|
|
7
|
+
data.tar.gz: 9cbc099462defe41b8bf96f590833d61de90f79c8e5b0f2d8a7d0ae04ee72dcc7e512c6f3886f67cb0e700dfb1046abd143b2adbde560c502d25700035e69834
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,58 @@
|
|
|
1
1
|
# Changes
|
|
2
2
|
|
|
3
|
+
### 2026-06-12 (0.13.0)
|
|
4
|
+
|
|
5
|
+
* Repair `#` hash line comments, like in Python, YAML, or Hjson:
|
|
6
|
+
`{"a": 1 # comment\n}` → `{"a":1}`, `{ # note\n "a": 1}` →
|
|
7
|
+
`{"a":1}`, `# lead\n{"a": 1}` → `{"a":1}`. Divergence from upstream
|
|
8
|
+
[jsonrepair](https://github.com/josdejong/jsonrepair) (v3.14.0
|
|
9
|
+
raises on all of these), commented at the site. Recognition is
|
|
10
|
+
context-aware so unquoted values starting with `#` keep repairing
|
|
11
|
+
into strings — `{"color": #ff0000}` → `{"color":"#ff0000"}`,
|
|
12
|
+
`{#tag: 1}` → `{"#tag":1}`, `#standalone` → `"#standalone"` — where
|
|
13
|
+
Python's `json_repair` silently loses them (`{"color": ""}`).
|
|
14
|
+
Where a value or key is expected, a `#` token reaching a structural
|
|
15
|
+
delimiter (`,` `}` `]` `:`) before any whitespace, or running to
|
|
16
|
+
end-of-input without a newline, stays a value; anything else is a
|
|
17
|
+
comment stripped to the end of the line. The tradeoff: a `#` token
|
|
18
|
+
followed by whitespace or a newline at a value position now reads
|
|
19
|
+
as a comment — `{"a": #b c}` → `{"a":null}` and `{"a": #tag\n}` →
|
|
20
|
+
`{"a":null}`, where 0.12.0 kept them as strings (Python drops them
|
|
21
|
+
too), and a comment-only document now raises like `// only a
|
|
22
|
+
comment` always has. Pinned in the spec suite as conscious
|
|
23
|
+
decisions.
|
|
24
|
+
|
|
25
|
+
### 2026-06-12 (0.12.0)
|
|
26
|
+
|
|
27
|
+
* Repair the three known input families that raised `Internal error:
|
|
28
|
+
repaired output is not valid JSON` — cases where upstream
|
|
29
|
+
[jsonrepair](https://github.com/josdejong/jsonrepair) (v3.14.0, still
|
|
30
|
+
its latest release) emits invalid JSON and this gem's canonical
|
|
31
|
+
re-serialize guard caught it but blamed the Repairer. All three are
|
|
32
|
+
deliberate divergences from upstream, commented at each site:
|
|
33
|
+
* A stray `e`/`E` with no mantissa is now an unquoted string instead
|
|
34
|
+
of an empty-mantissa exponent: `[e]` → `["e"]`, `[e5]` → `["e5"]`,
|
|
35
|
+
`[truee]` → `[true,"e"]`, `{"k": e}` → `{"k":"e"}` (upstream emits
|
|
36
|
+
`e0` / raw `e5`). Numbers truncated at a real exponent (`[2e]` →
|
|
37
|
+
`[2.0]`) are unchanged.
|
|
38
|
+
* Negative leading-zero numbers are quoted like positive ones:
|
|
39
|
+
`{"n": -05}` → `{"n":"-05"}`, matching the existing `{"n": 05}` →
|
|
40
|
+
`{"n":"05"}` (upstream emits `-05` unrepaired). The same rule now
|
|
41
|
+
also covers the truncated-number repair, which bypassed it:
|
|
42
|
+
`[05e]` → `["05e0"]`, `00.` → `"00.0"` (upstream emits `05e0` /
|
|
43
|
+
`00.0` unrepaired). Valid `-0` / `-0.5` / `0e` / `0.` are
|
|
44
|
+
unchanged.
|
|
45
|
+
* The trailing-comma repair no longer strips a comma belonging to the
|
|
46
|
+
enclosing container when an inner object/array fails on its first
|
|
47
|
+
key or value: `[{{]` → `[{},{}]`, `[1,[}]` → `[1,[]]`,
|
|
48
|
+
`{"a": 1, "b": [}` → `{"a":1,"b":[]}` (upstream emits `[{}{}]`,
|
|
49
|
+
`[1[]]`, `{"a": 1 "b": []}`).
|
|
50
|
+
Validated by differential testing against upstream over a 270-input
|
|
51
|
+
grid of these shapes in every container context: the only behavior
|
|
52
|
+
changes vs 0.11.3 are the 123 previously-`Internal error` inputs now
|
|
53
|
+
repairing (or, for `e+` shapes where upstream emits invalid `e+0`,
|
|
54
|
+
raising a clean position-bearing error). Benchmarks flat.
|
|
55
|
+
|
|
3
56
|
### 2026-06-12 (0.11.3)
|
|
4
57
|
|
|
5
58
|
* Fix infinite recursion (`SystemStackError`) on a quoted string
|
data/lib/json/repair/version.rb
CHANGED
data/lib/json/repairer.rb
CHANGED
|
@@ -71,7 +71,7 @@ module JSON
|
|
|
71
71
|
# repair redundant end quotes
|
|
72
72
|
while [CLOSING_BRACE, CLOSING_BRACKET].include?(@json[@index])
|
|
73
73
|
@index += 1
|
|
74
|
-
parse_whitespace_and_skip_comments
|
|
74
|
+
parse_whitespace_and_skip_comments(value_expected: false)
|
|
75
75
|
end
|
|
76
76
|
|
|
77
77
|
if @index >= @json.length
|
|
@@ -93,17 +93,17 @@ module JSON
|
|
|
93
93
|
parse_keywords ||
|
|
94
94
|
parse_unquoted_string(false) ||
|
|
95
95
|
parse_regex
|
|
96
|
-
parse_whitespace_and_skip_comments
|
|
96
|
+
parse_whitespace_and_skip_comments(value_expected: false)
|
|
97
97
|
|
|
98
98
|
process
|
|
99
99
|
end
|
|
100
100
|
|
|
101
|
-
def parse_whitespace_and_skip_comments(skip_newline: true)
|
|
101
|
+
def parse_whitespace_and_skip_comments(skip_newline: true, value_expected: true)
|
|
102
102
|
start = @index
|
|
103
103
|
|
|
104
104
|
changed = parse_whitespace(skip_newline: skip_newline)
|
|
105
105
|
loop do
|
|
106
|
-
changed = parse_comment
|
|
106
|
+
changed = parse_comment(value_expected: value_expected)
|
|
107
107
|
changed = parse_whitespace(skip_newline: skip_newline) if changed
|
|
108
108
|
break unless changed
|
|
109
109
|
end
|
|
@@ -131,7 +131,7 @@ module JSON
|
|
|
131
131
|
false
|
|
132
132
|
end
|
|
133
133
|
|
|
134
|
-
def parse_comment
|
|
134
|
+
def parse_comment(value_expected: true)
|
|
135
135
|
if @json[@index] == '/' && @json[@index + 1] == '*'
|
|
136
136
|
# Block comment
|
|
137
137
|
@index += 2
|
|
@@ -143,11 +143,42 @@ module JSON
|
|
|
143
143
|
@index += 2
|
|
144
144
|
@index += 1 until @json[@index].nil? || @json[@index] == "\n"
|
|
145
145
|
true
|
|
146
|
+
elsif @json[@index] == '#' && hash_comment?(value_expected)
|
|
147
|
+
# Hash line comment, like in Python, YAML, or Hjson (divergence
|
|
148
|
+
# from upstream, which raises on `#` as of v3.14.0)
|
|
149
|
+
@index += 1
|
|
150
|
+
@index += 1 until @json[@index].nil? || @json[@index] == "\n"
|
|
151
|
+
true
|
|
146
152
|
else
|
|
147
153
|
false
|
|
148
154
|
end
|
|
149
155
|
end
|
|
150
156
|
|
|
157
|
+
# Decide whether the `#` at @index starts a line comment or an
|
|
158
|
+
# unquoted value like {"color": #ff0000}, {#tag: 1}, or a root
|
|
159
|
+
# #hashtag (which Python's json_repair eats as comments, losing
|
|
160
|
+
# data). Where no value or key is expected an unquoted token would
|
|
161
|
+
# be junk anyway, so `#` is always a comment, exactly like `//`.
|
|
162
|
+
# Where one is expected, scan the rest of the line: a structural
|
|
163
|
+
# delimiter (`,` `}` `]` `:`) before any whitespace means the token
|
|
164
|
+
# reads as a value in context; whitespace (including the newline
|
|
165
|
+
# itself) first means comment prose; reaching EOF without a newline
|
|
166
|
+
# keeps the token a value, so truncated input like {"a": #tag is
|
|
167
|
+
# repaired, not dropped. Divergence from upstream, as above.
|
|
168
|
+
def hash_comment?(value_expected)
|
|
169
|
+
return true unless value_expected
|
|
170
|
+
|
|
171
|
+
i = @index + 1
|
|
172
|
+
while (char = @json[i])
|
|
173
|
+
return true if whitespace_or_special?(char)
|
|
174
|
+
return false if [COMMA, COLON, CLOSING_BRACE, CLOSING_BRACKET].include?(char)
|
|
175
|
+
|
|
176
|
+
i += 1
|
|
177
|
+
end
|
|
178
|
+
|
|
179
|
+
false
|
|
180
|
+
end
|
|
181
|
+
|
|
151
182
|
# Find and skip over a Markdown fenced code block:
|
|
152
183
|
# ``` ... ```
|
|
153
184
|
# or
|
|
@@ -237,6 +268,7 @@ module JSON
|
|
|
237
268
|
|
|
238
269
|
initial = true
|
|
239
270
|
while @index < @json.length && @json[@index] != CLOSING_BRACE
|
|
271
|
+
first_pair = initial
|
|
240
272
|
if initial
|
|
241
273
|
initial = false
|
|
242
274
|
else
|
|
@@ -255,15 +287,19 @@ module JSON
|
|
|
255
287
|
if @json[@index] == CLOSING_BRACE || @json[@index] == OPENING_BRACE ||
|
|
256
288
|
@json[@index] == CLOSING_BRACKET || @json[@index] == OPENING_BRACKET ||
|
|
257
289
|
@json[@index].nil?
|
|
258
|
-
# repair trailing comma
|
|
259
|
-
|
|
290
|
+
# repair trailing comma — but only the one this object's own loop
|
|
291
|
+
# emitted or inserted; on the first pair the buffer's last
|
|
292
|
+
# comma belongs to the enclosing container, like in [{{] or
|
|
293
|
+
# {"a": 1, "b": {] (divergence from upstream, which strips
|
|
294
|
+
# the parent's comma and emits invalid JSON like [{}{}])
|
|
295
|
+
@output = strip_last_occurrence(@output, ',') unless first_pair
|
|
260
296
|
else
|
|
261
297
|
throw_object_key_expected
|
|
262
298
|
end
|
|
263
299
|
break
|
|
264
300
|
end
|
|
265
301
|
|
|
266
|
-
parse_whitespace_and_skip_comments
|
|
302
|
+
parse_whitespace_and_skip_comments(value_expected: false)
|
|
267
303
|
processed_colon = parse_character(COLON)
|
|
268
304
|
truncated_text = @index >= @json.length
|
|
269
305
|
unless processed_colon
|
|
@@ -486,7 +522,7 @@ module JSON
|
|
|
486
522
|
@index += 1
|
|
487
523
|
@output << str
|
|
488
524
|
|
|
489
|
-
parse_whitespace_and_skip_comments(skip_newline: false)
|
|
525
|
+
parse_whitespace_and_skip_comments(skip_newline: false, value_expected: false)
|
|
490
526
|
|
|
491
527
|
if stop_at_delimiter ||
|
|
492
528
|
@index >= @json.length ||
|
|
@@ -738,7 +774,13 @@ module JSON
|
|
|
738
774
|
@index += 1 while digit?(@json[@index])
|
|
739
775
|
end
|
|
740
776
|
|
|
741
|
-
|
|
777
|
+
# Divergence from upstream: only enter the exponent branch when a
|
|
778
|
+
# mantissa was consumed — at this point @index > start implies at
|
|
779
|
+
# least one digit (the '-' and '.' paths reset otherwise). Upstream
|
|
780
|
+
# accepts a bare "e"/"E" here and emits invalid JSON like `e0` or
|
|
781
|
+
# raw `e5`; declining lets the token fall through to
|
|
782
|
+
# parse_unquoted_string, matching how "-e5" already becomes "-e5".
|
|
783
|
+
if @index > start && @json[@index] && @json[@index].downcase == 'e'
|
|
742
784
|
@index += 1
|
|
743
785
|
@index += 1 if ['-', '+'].include?(@json[@index])
|
|
744
786
|
if at_end_of_number?
|
|
@@ -761,7 +803,9 @@ module JSON
|
|
|
761
803
|
if @index > start
|
|
762
804
|
# repair a number with leading zeros like "00789"
|
|
763
805
|
num = @json[start...@index]
|
|
764
|
-
|
|
806
|
+
# the optional sign quotes "-05" like "05" (divergence from
|
|
807
|
+
# upstream, whose unsigned check lets "-05" through unrepaired)
|
|
808
|
+
has_invalid_leading_zero = num.match?(/^-?0\d/)
|
|
765
809
|
|
|
766
810
|
@output << (has_invalid_leading_zero ? "\"#{num}\"" : repair_leading_dot_number(num))
|
|
767
811
|
return true
|
|
@@ -786,6 +830,7 @@ module JSON
|
|
|
786
830
|
|
|
787
831
|
initial = true
|
|
788
832
|
while @index < @json.length && @json[@index] != CLOSING_BRACKET
|
|
833
|
+
first_item = initial
|
|
789
834
|
if initial
|
|
790
835
|
initial = false
|
|
791
836
|
else
|
|
@@ -799,8 +844,12 @@ module JSON
|
|
|
799
844
|
processed_value = parse_value
|
|
800
845
|
next if processed_value
|
|
801
846
|
|
|
802
|
-
# repair trailing comma
|
|
803
|
-
|
|
847
|
+
# repair trailing comma — but only the one this array's own loop
|
|
848
|
+
# emitted or inserted; on the first item the buffer's last
|
|
849
|
+
# comma belongs to the enclosing container, like in [1,[}] or
|
|
850
|
+
# {"a": 1, "b": [} (divergence from upstream, which strips
|
|
851
|
+
# the parent's comma and emits invalid JSON like [1[]])
|
|
852
|
+
@output = strip_last_occurrence(@output, ',') unless first_item
|
|
804
853
|
break
|
|
805
854
|
end
|
|
806
855
|
|
|
@@ -828,11 +877,11 @@ module JSON
|
|
|
828
877
|
def parse_concatenated_string
|
|
829
878
|
processed = false
|
|
830
879
|
|
|
831
|
-
parse_whitespace_and_skip_comments
|
|
880
|
+
parse_whitespace_and_skip_comments(value_expected: false)
|
|
832
881
|
while @json[@index] == PLUS
|
|
833
882
|
processed = true
|
|
834
883
|
@index += 1
|
|
835
|
-
parse_whitespace_and_skip_comments
|
|
884
|
+
parse_whitespace_and_skip_comments(value_expected: false)
|
|
836
885
|
|
|
837
886
|
# repair: remove the end quote of the first string
|
|
838
887
|
@output = strip_last_occurrence(@output, '"', strip_remaining_text: true)
|
|
@@ -859,7 +908,11 @@ module JSON
|
|
|
859
908
|
# repair numbers cut off at the end
|
|
860
909
|
# this will only be called when we end after a '.', '-', or 'e' and does not
|
|
861
910
|
# change the number more than it needs to make it valid JSON
|
|
862
|
-
|
|
911
|
+
num = "#{@json[start...@index]}0"
|
|
912
|
+
# quote a padded token that has an invalid leading zero, like "05e" ->
|
|
913
|
+
# "05e0", applying the same rule as the end of parse_number (divergence
|
|
914
|
+
# from upstream, which emits the invalid number raw)
|
|
915
|
+
@output << (num.match?(/^-?0\d/) ? "\"#{num}\"" : repair_leading_dot_number(num))
|
|
863
916
|
end
|
|
864
917
|
|
|
865
918
|
# Repair a number missing its digit before the decimal point, like ".5"
|
data/sig/json/repairer.rbs
CHANGED
|
@@ -37,7 +37,9 @@ module JSON
|
|
|
37
37
|
|
|
38
38
|
def parse_whitespace: (?skip_newline: bool) -> bool
|
|
39
39
|
|
|
40
|
-
def parse_comment: () -> bool
|
|
40
|
+
def parse_comment: (?value_expected: bool) -> bool
|
|
41
|
+
|
|
42
|
+
def hash_comment?: (bool value_expected) -> bool
|
|
41
43
|
|
|
42
44
|
# Find and skip over a Markdown fenced code block
|
|
43
45
|
def parse_markdown_code_block: (::Array[::String] blocks) -> bool
|
|
@@ -87,7 +89,7 @@ module JSON
|
|
|
87
89
|
|
|
88
90
|
def parse_character: (::String char) -> bool
|
|
89
91
|
|
|
90
|
-
def parse_whitespace_and_skip_comments: (?skip_newline: bool) -> bool
|
|
92
|
+
def parse_whitespace_and_skip_comments: (?skip_newline: bool, ?value_expected: bool) -> bool
|
|
91
93
|
|
|
92
94
|
# Parse a number like 2.4 or 2.4e6
|
|
93
95
|
def parse_number: () -> bool
|