json-repair 0.12.0 → 0.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +37 -0
- data/CLAUDE.md +1 -1
- data/lib/json/repair/version.rb +1 -1
- data/lib/json/repairer.rb +63 -11
- data/sig/json/repairer.rbs +4 -2
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 3e0898f751c91f7c3a026ccb107c319b983212296090031ba6b7a80713de18eb
|
|
4
|
+
data.tar.gz: 21028aa56a685b77ea42506333753deb1a903108b860843a877181b6af66b1a2
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 05752fda1b0ff27c2dcde8b217099460bdecb57efa4a3ab73ce534f27dfecf51673f68683f3009c8d159079bfb7fee3f0892f2c86e3f15d0e86a58d4c497c8f0
|
|
7
|
+
data.tar.gz: 46a8a689d2a969c3c7ca1395dbd4a5555e191e304c0cfffd3bbaec44d22066b2a8527efc7323267c62c21738364c2dddf917d20af3388ae87cf993d872d014a0
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,42 @@
|
|
|
1
1
|
# Changes
|
|
2
2
|
|
|
3
|
+
### 2026-06-17 (0.14.0)
|
|
4
|
+
|
|
5
|
+
* Repair a leading `+` on numbers, like in JSON5: `+1.23` → `1.23`,
|
|
6
|
+
`[+1]` → `[1]`, `{"a": +5}` → `{"a":5}`, `+.5` → `0.5`. The `+` is
|
|
7
|
+
consumed and dropped (JSON has no explicit plus), so a leading-zero
|
|
8
|
+
number is quoted exactly like its unsigned form (`+05` → `"05"`),
|
|
9
|
+
while a `+` inside an exponent is left untouched (`2e+5` →
|
|
10
|
+
`200000.0`). The `+` must be followed by a digit or a leading dot, so
|
|
11
|
+
a bare `+` (and `+abc`, `++1`, `+e5`) still raises — and requiring a
|
|
12
|
+
digit/dot keeps the 0.12.0 empty-mantissa exponent guard intact.
|
|
13
|
+
Divergence from upstream
|
|
14
|
+
[jsonrepair](https://github.com/josdejong/jsonrepair) (v3.14.0 leaves
|
|
15
|
+
`+1.23` unrepaired, raising "Unexpected end of json string"),
|
|
16
|
+
commented at the site. Mirrors the leading-dot number repair (0.9.0).
|
|
17
|
+
|
|
18
|
+
### 2026-06-12 (0.13.0)
|
|
19
|
+
|
|
20
|
+
* Repair `#` hash line comments, like in Python, YAML, or Hjson:
|
|
21
|
+
`{"a": 1 # comment\n}` → `{"a":1}`, `{ # note\n "a": 1}` →
|
|
22
|
+
`{"a":1}`, `# lead\n{"a": 1}` → `{"a":1}`. Divergence from upstream
|
|
23
|
+
[jsonrepair](https://github.com/josdejong/jsonrepair) (v3.14.0
|
|
24
|
+
raises on all of these), commented at the site. Recognition is
|
|
25
|
+
context-aware so unquoted values starting with `#` keep repairing
|
|
26
|
+
into strings — `{"color": #ff0000}` → `{"color":"#ff0000"}`,
|
|
27
|
+
`{#tag: 1}` → `{"#tag":1}`, `#standalone` → `"#standalone"` — where
|
|
28
|
+
Python's `json_repair` silently loses them (`{"color": ""}`).
|
|
29
|
+
Where a value or key is expected, a `#` token reaching a structural
|
|
30
|
+
delimiter (`,` `}` `]` `:`) before any whitespace, or running to
|
|
31
|
+
end-of-input without a newline, stays a value; anything else is a
|
|
32
|
+
comment stripped to the end of the line. The tradeoff: a `#` token
|
|
33
|
+
followed by whitespace or a newline at a value position now reads
|
|
34
|
+
as a comment — `{"a": #b c}` → `{"a":null}` and `{"a": #tag\n}` →
|
|
35
|
+
`{"a":null}`, where 0.12.0 kept them as strings (Python drops them
|
|
36
|
+
too), and a comment-only document now raises like `// only a
|
|
37
|
+
comment` always has. Pinned in the spec suite as conscious
|
|
38
|
+
decisions.
|
|
39
|
+
|
|
3
40
|
### 2026-06-12 (0.12.0)
|
|
4
41
|
|
|
5
42
|
* Repair the three known input families that raised `Internal error:
|
data/CLAUDE.md
CHANGED
|
@@ -24,7 +24,7 @@ A few repair heuristics deliberately go **beyond** upstream (leading-dot numbers
|
|
|
24
24
|
|
|
25
25
|
### Entry point
|
|
26
26
|
|
|
27
|
-
`JSON.repair(str)` in `lib/json/repair.rb` first tries stdlib `JSON.parse` (fast path; opt out with `skip_json_loads: true`), and falls back to `JSON::Repairer.new(str).repair` when that raises. Either way the result is re-serialized with `JSON.generate`, so **output is canonical** — whitespace collapsed, numbers normalized, duplicate keys last-write-wins — and both paths agree on it.
|
|
27
|
+
`JSON.repair(str)` in `lib/json/repair.rb` first tries stdlib `JSON.parse` (fast path; opt out with `skip_json_loads: true`), and falls back to `JSON::Repairer.new(str).repair` when that raises. Either way the result is re-serialized with `JSON.generate`, so **output is canonical** — whitespace collapsed, numbers normalized, duplicate keys last-write-wins — and both paths agree on it. There is no pre-filter regex: the fast path simply runs `JSON.parse` and falls back to the Repairer on any `JSON::ParserError`. Comments are handled either way — the bundled `json` gem currently accepts `//` and `/*` comments (an implementation detail, not part of the `JSON.parse` contract) and `JSON.generate` drops them (`{} // c` → `{}`); were a future `json` release to reject them, they would simply raise and route to the Repairer, which strips comments too. Invalid escapes, NDJSON, Markdown fences, and trailing prose already make `JSON.parse` raise and so route to the Repairer. `return_objects: true` returns the parsed Ruby value instead of a string; `JSON.repair_file(path)` / `JSON.repair_io(io)` are convenience wrappers forwarding both options.
|
|
28
28
|
|
|
29
29
|
`JSON::JSONRepairError` is the only error raised for unrecoverable inputs; it exposes the failure `#position`. If the Repairer ever emits a string stdlib cannot parse (a Repairer bug), the `JSON::ParserError` is wrapped in `JSONRepairError` rather than leaked.
|
|
30
30
|
|
data/lib/json/repair/version.rb
CHANGED
data/lib/json/repairer.rb
CHANGED
|
@@ -71,7 +71,7 @@ module JSON
|
|
|
71
71
|
# repair redundant end quotes
|
|
72
72
|
while [CLOSING_BRACE, CLOSING_BRACKET].include?(@json[@index])
|
|
73
73
|
@index += 1
|
|
74
|
-
parse_whitespace_and_skip_comments
|
|
74
|
+
parse_whitespace_and_skip_comments(value_expected: false)
|
|
75
75
|
end
|
|
76
76
|
|
|
77
77
|
if @index >= @json.length
|
|
@@ -93,17 +93,17 @@ module JSON
|
|
|
93
93
|
parse_keywords ||
|
|
94
94
|
parse_unquoted_string(false) ||
|
|
95
95
|
parse_regex
|
|
96
|
-
parse_whitespace_and_skip_comments
|
|
96
|
+
parse_whitespace_and_skip_comments(value_expected: false)
|
|
97
97
|
|
|
98
98
|
process
|
|
99
99
|
end
|
|
100
100
|
|
|
101
|
-
def parse_whitespace_and_skip_comments(skip_newline: true)
|
|
101
|
+
def parse_whitespace_and_skip_comments(skip_newline: true, value_expected: true)
|
|
102
102
|
start = @index
|
|
103
103
|
|
|
104
104
|
changed = parse_whitespace(skip_newline: skip_newline)
|
|
105
105
|
loop do
|
|
106
|
-
changed = parse_comment
|
|
106
|
+
changed = parse_comment(value_expected: value_expected)
|
|
107
107
|
changed = parse_whitespace(skip_newline: skip_newline) if changed
|
|
108
108
|
break unless changed
|
|
109
109
|
end
|
|
@@ -131,7 +131,7 @@ module JSON
|
|
|
131
131
|
false
|
|
132
132
|
end
|
|
133
133
|
|
|
134
|
-
def parse_comment
|
|
134
|
+
def parse_comment(value_expected: true)
|
|
135
135
|
if @json[@index] == '/' && @json[@index + 1] == '*'
|
|
136
136
|
# Block comment
|
|
137
137
|
@index += 2
|
|
@@ -143,11 +143,42 @@ module JSON
|
|
|
143
143
|
@index += 2
|
|
144
144
|
@index += 1 until @json[@index].nil? || @json[@index] == "\n"
|
|
145
145
|
true
|
|
146
|
+
elsif @json[@index] == '#' && hash_comment?(value_expected)
|
|
147
|
+
# Hash line comment, like in Python, YAML, or Hjson (divergence
|
|
148
|
+
# from upstream, which raises on `#` as of v3.14.0)
|
|
149
|
+
@index += 1
|
|
150
|
+
@index += 1 until @json[@index].nil? || @json[@index] == "\n"
|
|
151
|
+
true
|
|
146
152
|
else
|
|
147
153
|
false
|
|
148
154
|
end
|
|
149
155
|
end
|
|
150
156
|
|
|
157
|
+
# Decide whether the `#` at @index starts a line comment or an
|
|
158
|
+
# unquoted value like {"color": #ff0000}, {#tag: 1}, or a root
|
|
159
|
+
# #hashtag (which Python's json_repair eats as comments, losing
|
|
160
|
+
# data). Where no value or key is expected an unquoted token would
|
|
161
|
+
# be junk anyway, so `#` is always a comment, exactly like `//`.
|
|
162
|
+
# Where one is expected, scan the rest of the line: a structural
|
|
163
|
+
# delimiter (`,` `}` `]` `:`) before any whitespace means the token
|
|
164
|
+
# reads as a value in context; whitespace (including the newline
|
|
165
|
+
# itself) first means comment prose; reaching EOF without a newline
|
|
166
|
+
# keeps the token a value, so truncated input like {"a": #tag is
|
|
167
|
+
# repaired, not dropped. Divergence from upstream, as above.
|
|
168
|
+
def hash_comment?(value_expected)
|
|
169
|
+
return true unless value_expected
|
|
170
|
+
|
|
171
|
+
i = @index + 1
|
|
172
|
+
while (char = @json[i])
|
|
173
|
+
return true if whitespace_or_special?(char)
|
|
174
|
+
return false if [COMMA, COLON, CLOSING_BRACE, CLOSING_BRACKET].include?(char)
|
|
175
|
+
|
|
176
|
+
i += 1
|
|
177
|
+
end
|
|
178
|
+
|
|
179
|
+
false
|
|
180
|
+
end
|
|
181
|
+
|
|
151
182
|
# Find and skip over a Markdown fenced code block:
|
|
152
183
|
# ``` ... ```
|
|
153
184
|
# or
|
|
@@ -268,7 +299,7 @@ module JSON
|
|
|
268
299
|
break
|
|
269
300
|
end
|
|
270
301
|
|
|
271
|
-
parse_whitespace_and_skip_comments
|
|
302
|
+
parse_whitespace_and_skip_comments(value_expected: false)
|
|
272
303
|
processed_colon = parse_character(COLON)
|
|
273
304
|
truncated_text = @index >= @json.length
|
|
274
305
|
unless processed_colon
|
|
@@ -491,7 +522,7 @@ module JSON
|
|
|
491
522
|
@index += 1
|
|
492
523
|
@output << str
|
|
493
524
|
|
|
494
|
-
parse_whitespace_and_skip_comments(skip_newline: false)
|
|
525
|
+
parse_whitespace_and_skip_comments(skip_newline: false, value_expected: false)
|
|
495
526
|
|
|
496
527
|
if stop_at_delimiter ||
|
|
497
528
|
@index >= @json.length ||
|
|
@@ -710,6 +741,24 @@ module JSON
|
|
|
710
741
|
# Parse a number like 2.4 or 2.4e6
|
|
711
742
|
def parse_number
|
|
712
743
|
start = @index
|
|
744
|
+
|
|
745
|
+
# Divergence from upstream: accept and discard a leading "+". JSON5
|
|
746
|
+
# permits an explicit plus; JSON does not, so "+1.23" -> "1.23" and
|
|
747
|
+
# "{"a": +5}" -> {"a":5}. The "+" must be followed by a digit or a
|
|
748
|
+
# leading dot (mirroring the "-" branch below); otherwise this is not
|
|
749
|
+
# a number and we backtrack to the "+" so a bare "+" still raises.
|
|
750
|
+
# `start` stays on the "+", so the reset paths restore @index cleanly
|
|
751
|
+
# and the @index > start exponent guard below still implies a digit
|
|
752
|
+
# was consumed; the two emission sites drop the leading "+" before
|
|
753
|
+
# quoting. Upstream leaves "+1.23" unrepaired.
|
|
754
|
+
if @json[@index] == PLUS
|
|
755
|
+
@index += 1
|
|
756
|
+
unless digit?(@json[@index]) || @json[@index] == DOT
|
|
757
|
+
@index = start
|
|
758
|
+
return false
|
|
759
|
+
end
|
|
760
|
+
end
|
|
761
|
+
|
|
713
762
|
if @json[@index] == '-'
|
|
714
763
|
@index += 1
|
|
715
764
|
if at_end_of_number?
|
|
@@ -771,7 +820,9 @@ module JSON
|
|
|
771
820
|
|
|
772
821
|
if @index > start
|
|
773
822
|
# repair a number with leading zeros like "00789"
|
|
774
|
-
|
|
823
|
+
# drop a leading "+" first (see the PLUS branch above), so "+05"
|
|
824
|
+
# quotes like "05" and "+1.23" emits as "1.23"
|
|
825
|
+
num = @json[start...@index].delete_prefix(PLUS)
|
|
775
826
|
# the optional sign quotes "-05" like "05" (divergence from
|
|
776
827
|
# upstream, whose unsigned check lets "-05" through unrepaired)
|
|
777
828
|
has_invalid_leading_zero = num.match?(/^-?0\d/)
|
|
@@ -846,11 +897,11 @@ module JSON
|
|
|
846
897
|
def parse_concatenated_string
|
|
847
898
|
processed = false
|
|
848
899
|
|
|
849
|
-
parse_whitespace_and_skip_comments
|
|
900
|
+
parse_whitespace_and_skip_comments(value_expected: false)
|
|
850
901
|
while @json[@index] == PLUS
|
|
851
902
|
processed = true
|
|
852
903
|
@index += 1
|
|
853
|
-
parse_whitespace_and_skip_comments
|
|
904
|
+
parse_whitespace_and_skip_comments(value_expected: false)
|
|
854
905
|
|
|
855
906
|
# repair: remove the end quote of the first string
|
|
856
907
|
@output = strip_last_occurrence(@output, '"', strip_remaining_text: true)
|
|
@@ -877,7 +928,8 @@ module JSON
|
|
|
877
928
|
# repair numbers cut off at the end
|
|
878
929
|
# this will only be called when we end after a '.', '-', or 'e' and does not
|
|
879
930
|
# change the number more than it needs to make it valid JSON
|
|
880
|
-
|
|
931
|
+
# delete_prefix drops a leading "+" (see the PLUS branch in parse_number)
|
|
932
|
+
num = "#{@json[start...@index]}0".delete_prefix(PLUS)
|
|
881
933
|
# quote a padded token that has an invalid leading zero, like "05e" ->
|
|
882
934
|
# "05e0", applying the same rule as the end of parse_number (divergence
|
|
883
935
|
# from upstream, which emits the invalid number raw)
|
data/sig/json/repairer.rbs
CHANGED
|
@@ -37,7 +37,9 @@ module JSON
|
|
|
37
37
|
|
|
38
38
|
def parse_whitespace: (?skip_newline: bool) -> bool
|
|
39
39
|
|
|
40
|
-
def parse_comment: () -> bool
|
|
40
|
+
def parse_comment: (?value_expected: bool) -> bool
|
|
41
|
+
|
|
42
|
+
def hash_comment?: (bool value_expected) -> bool
|
|
41
43
|
|
|
42
44
|
# Find and skip over a Markdown fenced code block
|
|
43
45
|
def parse_markdown_code_block: (::Array[::String] blocks) -> bool
|
|
@@ -87,7 +89,7 @@ module JSON
|
|
|
87
89
|
|
|
88
90
|
def parse_character: (::String char) -> bool
|
|
89
91
|
|
|
90
|
-
def parse_whitespace_and_skip_comments: (?skip_newline: bool) -> bool
|
|
92
|
+
def parse_whitespace_and_skip_comments: (?skip_newline: bool, ?value_expected: bool) -> bool
|
|
91
93
|
|
|
92
94
|
# Parse a number like 2.4 or 2.4e6
|
|
93
95
|
def parse_number: () -> bool
|