ata-validator 1.21.0 → 1.22.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +24 -0
- package/README.md +4 -4
- package/index.js +22 -10
- package/lib/enrich-error.js +83 -26
- package/lib/formats.js +158 -30
- package/lib/js-compiler.js +12 -1
- package/lib/levenshtein.js +19 -7
- package/lib/pointer.js +46 -0
- package/lib/retry-message.js +26 -2
- package/lib/suggestions.js +27 -24
- package/lib/version.js +1 -1
- package/package.json +9 -9
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,30 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to ata-validator are documented here. The format follows [Keep a Changelog](https://keepachangelog.com/), and this project adheres to semantic versioning.
|
|
4
4
|
|
|
5
|
+
## 1.22.0 - 2026-09-14
|
|
6
|
+
|
|
7
|
+
### Fixed
|
|
8
|
+
|
|
9
|
+
- A schema validated once with `isValidObject()` lost its generated error function for the rest of the process. The verdict-only fast path compiles that one function and seeds the shared compile cache with the other two set to null, meaning "not built yet"; the full compile read those nulls as "the compiler declined this schema" and never tried again, so errors fell back to the interpreted engine or the addon. Every server hits that order, because most documents are valid. Cache entries now record whether they are complete. Verdicts and error content were never affected, and `tests/test_compile_cache_order.js` runs both orders in separate processes and holds the engine and the error list equal.
|
|
10
|
+
- Measured on a five-field schema with an enum and a format, after validating one valid document: `validate(bad).errors` 3.89 µs to 0.66 µs, the Standard Schema path 3.52 µs to 0.07 µs. With no addon installed at all (`ATA_NO_NATIVE=1`), where the fallback was the interpreter rather than the addon, 1.05 µs to 0.66 µs and 0.42 µs to 0.07 µs.
|
|
11
|
+
- `engine()` reported `closure` for a schema that compiles, once a valid document had been validated first. It now reports `codegen` in either order.
|
|
12
|
+
|
|
13
|
+
### Changed
|
|
14
|
+
|
|
15
|
+
- The official test suite submodule moved 60 commits forward, to 2026-09-13. All three dialects stay at zero regressions and the case counts rose: Draft 2020-12 1299 to 1301, draft 7 927 to 929, the v1 dialect 1133 to 1135, identical with code generation blocked. The buffer path agrees with `validate()` on all 3365 of them. Every figure in the README, the docs and the contributor notes was remeasured and updated together, and the floors in `tests/test_no_eval.js` were raised so the new cases cannot be lost silently.
|
|
16
|
+
|
|
17
|
+
- `uri` is one walk over the string instead of three, and reads its character classes out of tables instead of comparison chains. It was the most expensive thing an ordinary document carried: on a schema with five formats over a 1.6 KB payload, format assertion was 94 percent of the validation cost and `uri` was most of that, while the structural check of the whole nested document was 43 ns. The authority's landmarks, the last `@`, the colons after it and whether a bracket came before it, are now recorded during the character scan rather than by cutting the authority out with `slice` and asking the copy for `lastIndexOf` and two regular expressions.
|
|
18
|
+
- `format: uri` on one string, through a compiled validator: 92 to 42 ns. The same product schema's document: 697 to 423 ns. End to end, where `JSON.parse` is the other 75 percent: 2.98 to 2.69 µs.
|
|
19
|
+
- Measured, and not taken: nine `===` tests per character cost more than the loop they guarded, a lookup table is 2.3x that chain, and a three-regex formulation of the whole rule landed at 65 ns against the table walk's 41.
|
|
20
|
+
- The code generator hoists the walk once per compiled function and calls it, rather than inlining it at each call site where the tables would have to be rebuilt. This is the shape `email` already had. `tests/test_uri_helper_parity.js` holds the function, the generated text and a compiled validator to the same answer over 35,677 strings built from the walk's own boundaries.
|
|
21
|
+
|
|
22
|
+
- Enriching a rejection no longer reads more of the document than the diagnostic it produces. `received` describes the offending value in at most 60 characters, and it used to find out whether a value fitted by serialising it whole, so a `required` error on a large payload cost a `JSON.stringify` of that payload, once per error. Sizing is now bounded by those 60 characters and stops as soon as they are spent. Enriching one error on a document holding a thousand rows: 61.5 µs to 160 ns, and flat in the size of the document rather than linear. On the schema-benchmarks product schema, where `validate()` reports 16 errors, reading `.errors` went from 9.84 to 5.10 µs (48 percent less) on the same machine, separate processes. Verdict paths, `isValidObject()` and the Standard Schema bridge are unchanged.
|
|
23
|
+
- An object too large to show reports its shape rather than its serialised size: `[object, 5 keys]` where it used to say `[object, ~0.1KB]`. This is the only change to error output; every value that printed its body before prints the same body now, including nested objects, arrays and values with their own `toJSON`.
|
|
24
|
+
- The error's value is resolved from its pointer once and handed to both the `received` repr and the suggestion sources. `lib/suggestions.js` had a second pointer walk that split the string and mapped over the segments; both paths now use `resolvePointer` in `lib/pointer.js`.
|
|
25
|
+
- A `required` typo hint is offered on containers of up to 64 keys. Past that the nearest key stops being evidence of a typo, which is the cap `suggestEnumTypo` already applied to its candidate list, and the scan is a distance computation per key.
|
|
26
|
+
- `levenshtein` reads characters by code and swaps its two rows through a temporary. The destructured swap it used built an array per row. About 9 percent faster on the key pairs the suggestion sources compare.
|
|
27
|
+
- `tests/test_enrich_received.js` holds enrichment to within 4x across a hundredfold increase in document volume. Before this change the same measurement was 122x.
|
|
28
|
+
|
|
5
29
|
## 1.17.2 - 2026-09-13
|
|
6
30
|
|
|
7
31
|
### Changed
|
package/README.md
CHANGED
|
@@ -20,11 +20,11 @@ The `ata-validator` package itself is pure JavaScript. The native accelerator (s
|
|
|
20
20
|
npm install ata-validator --omit=optional
|
|
21
21
|
```
|
|
22
22
|
|
|
23
|
-
or set `ATA_NO_NATIVE=1` at runtime. Typical schemas compile to specialized JS; shapes the compiler cannot represent (some `$dynamicRef`, cyclic `$ref`, unusual keyword interactions) fall back to an interpreted engine, so every schema validates in every environment. The pure-JS setup scores the same on the official suite as the native one,
|
|
23
|
+
or set `ATA_NO_NATIVE=1` at runtime. Typical schemas compile to specialized JS; shapes the compiler cannot represent (some `$dynamicRef`, cyclic `$ref`, unusual keyword interactions) fall back to an interpreted engine, so every schema validates in every environment. The pure-JS setup scores the same on the official suite as the native one, 1301 of 1301 Draft 2020-12 cases. Only the buffer and parallel APIs (`isValid` on raw buffers, `countValid`, `batchIsValid`, `validateAndParse`) need the native engine and say so with a clear error.
|
|
24
24
|
|
|
25
|
-
Those four now agree with `validate()` on every case of the official suite,
|
|
25
|
+
Those four now agree with `validate()` on every case of the official suite, 3365 across three dialects. The native walker behind them does not handle every shape (`contains`, `unevaluatedProperties`, `patternProperties`, tuple `items`, cross-document `$ref`, a few formats), so for schemas using one of those the buffer APIs parse the bytes and answer through `validate()`; the list is in `lib/buffer-gate.js`. Typical request schemas stay on the zero-copy path. `npm test` holds the disagreement count at zero.
|
|
26
26
|
|
|
27
|
-
Where `new Function` is refused altogether, on Cloudflare Workers, Deno Deploy or under a strict Content-Security-Policy, ata drops to the interpreted engine and scores the same
|
|
27
|
+
Where `new Function` is refused altogether, on Cloudflare Workers, Deno Deploy or under a strict Content-Security-Policy, ata drops to the interpreted engine and scores the same 1301 of 1301 with code generation blocked. No flags, and on Workers no `nodejs_compat` either. See [docs/edge-runtimes.md](docs/edge-runtimes.md).
|
|
28
28
|
|
|
29
29
|
Node has a switch for exactly that environment, so this takes thirty seconds to check for yourself, on ata or on whatever you use today:
|
|
30
30
|
|
|
@@ -560,7 +560,7 @@ Both are implemented in the interpreted engine, so a v1 schema that uses `$dynam
|
|
|
560
560
|
|
|
561
561
|
### Known limitations
|
|
562
562
|
|
|
563
|
-
Running the whole Draft 2020-12 suite with nothing excluded, `format` and `default` under specification semantics (`assertFormat: false`, `useDefaults: false`), gives
|
|
563
|
+
Running the whole Draft 2020-12 suite with nothing excluded, `format` and `default` under specification semantics (`assertFormat: false`, `useDefaults: false`), gives 1301 of 1301 cases. Draft 7 gives 929 of 929. The v1 dialect gives 1135 of 1135. `npm run test:suite` reproduces all three.
|
|
564
564
|
|
|
565
565
|
Areas that remain deliberate scope decisions for 1.x:
|
|
566
566
|
|
package/index.js
CHANGED
|
@@ -458,13 +458,17 @@ Object.defineProperty(RichRejection.prototype, 'errors', {
|
|
|
458
458
|
const enrich = this._enrich;
|
|
459
459
|
let raw = this._result.errors || [];
|
|
460
460
|
if (raw.length > 1) raw = sortErrorsBySchemaOrder(this._root, raw);
|
|
461
|
-
|
|
462
|
-
|
|
461
|
+
// One options object for the whole list, not one per error.
|
|
462
|
+
const opts = enrich && raw.length
|
|
463
|
+
? {
|
|
463
464
|
data: this._data,
|
|
464
465
|
positions: this._positions,
|
|
465
466
|
schemaPositions: self._schemaPositions,
|
|
466
467
|
schemaFile: self._source ? self._source.path : undefined,
|
|
467
|
-
}
|
|
468
|
+
}
|
|
469
|
+
: null;
|
|
470
|
+
const cached = opts
|
|
471
|
+
? raw.map((e) => enrich(e, opts))
|
|
468
472
|
// The v0.14 shape is a fixed key set; the ordering key the
|
|
469
473
|
// generated code carries is dropped from it here.
|
|
470
474
|
: raw.map(stripOrdinal);
|
|
@@ -1005,7 +1009,12 @@ class Validator {
|
|
|
1005
1009
|
// there.
|
|
1006
1010
|
if (this._v1Dynamic || this._usesKeywords || !codegenAvailable()) {
|
|
1007
1011
|
jsFn = null; jsCombinedFn = null; jsErrFn = null;
|
|
1008
|
-
|
|
1012
|
+
// `full` separates an entry that holds every compiled function from one
|
|
1013
|
+
// the verdict-only fast path seeded, where `combined` and `errFn` are null
|
|
1014
|
+
// because nothing has tried to build them yet. Both halves of that
|
|
1015
|
+
// distinction are null, and reading the second as the first costs this
|
|
1016
|
+
// schema its generated error function for the life of the process.
|
|
1017
|
+
} else if (cached && cached.full && !_forceNapi) {
|
|
1009
1018
|
jsFn = cached.jsFn;
|
|
1010
1019
|
jsCombinedFn = cached.combined;
|
|
1011
1020
|
jsErrFn = cached.errFn;
|
|
@@ -1019,13 +1028,13 @@ class Validator {
|
|
|
1019
1028
|
_isCodegen = !!_cgFn;
|
|
1020
1029
|
this._engine = _cgFn ? 'codegen' : jsFn ? 'closure' : null;
|
|
1021
1030
|
if (!uf) {
|
|
1022
|
-
_compileCache.set(mapKey, { jsFn, combined: jsCombinedFn, errFn: jsErrFn, isCodegen: _isCodegen });
|
|
1031
|
+
_compileCache.set(mapKey, { jsFn, combined: jsCombinedFn, errFn: jsErrFn, isCodegen: _isCodegen, full: true });
|
|
1023
1032
|
}
|
|
1024
1033
|
} else {
|
|
1025
1034
|
jsFn = null; jsCombinedFn = null; jsErrFn = null;
|
|
1026
1035
|
}
|
|
1027
1036
|
this._jsFn = jsFn;
|
|
1028
|
-
if (this._engine === undefined) this._engine = cached ? (cached.isCodegen ? 'codegen' : jsFn ? 'closure' : null) : null;
|
|
1037
|
+
if (this._engine === undefined) this._engine = (cached && cached.full) ? (cached.isCodegen ? 'codegen' : jsFn ? 'closure' : null) : null;
|
|
1029
1038
|
|
|
1030
1039
|
// Data mutators -- try codegen first (12x faster), fallback to closure arrays.
|
|
1031
1040
|
// Follow cross-refs so coercion/defaults/removeAdditional see the referenced
|
|
@@ -1549,12 +1558,13 @@ class Validator {
|
|
|
1549
1558
|
// Declaration order, as validate() applies it; the text path
|
|
1550
1559
|
// used to enrich in emission order.
|
|
1551
1560
|
const ordered = result.errors.length > 1 ? sortErrorsBySchemaOrder(this._schemaObj, result.errors) : result.errors;
|
|
1552
|
-
const
|
|
1561
|
+
const enrichOpts = {
|
|
1553
1562
|
data: parsedData,
|
|
1554
1563
|
positions,
|
|
1555
1564
|
schemaPositions: this._schemaPositions,
|
|
1556
1565
|
schemaFile: this._source ? this._source.path : undefined,
|
|
1557
|
-
}
|
|
1566
|
+
};
|
|
1567
|
+
const enriched = ordered.map((e) => enrich(e, enrichOpts));
|
|
1558
1568
|
if (enriched.length > 1) attachRelated(enriched);
|
|
1559
1569
|
attachDiagnosticSource(enriched, {
|
|
1560
1570
|
data: parsedData,
|
|
@@ -1776,9 +1786,11 @@ class Validator {
|
|
|
1776
1786
|
this._jsFn = jsFn;
|
|
1777
1787
|
if (jsFn) {
|
|
1778
1788
|
this.isValidObject = jsFn;
|
|
1779
|
-
//
|
|
1789
|
+
// A partial entry: the verdict function is real, the other two are not
|
|
1790
|
+
// built yet rather than declined. `full: false` says so, so the next
|
|
1791
|
+
// caller that needs errors compiles them instead of inheriting nulls.
|
|
1780
1792
|
if (!uf) {
|
|
1781
|
-
if (!cached) _compileCache.set(mapKey, { jsFn, combined: null, errFn: null });
|
|
1793
|
+
if (!cached) _compileCache.set(mapKey, { jsFn, combined: null, errFn: null, full: false });
|
|
1782
1794
|
else cached.jsFn = jsFn;
|
|
1783
1795
|
}
|
|
1784
1796
|
}
|
package/lib/enrich-error.js
CHANGED
|
@@ -2,9 +2,70 @@
|
|
|
2
2
|
|
|
3
3
|
const { CODES, codeFor, fromNative } = require('./error-codes');
|
|
4
4
|
const { suggestFor } = require('./suggestions');
|
|
5
|
+
const { resolvePointer } = require('./pointer');
|
|
6
|
+
|
|
7
|
+
// A pointer that walked out of the document: no value to report, not even the
|
|
8
|
+
// word 'undefined', which is what an absent key reports.
|
|
9
|
+
const MISSING = Symbol('ata.received.missing');
|
|
5
10
|
|
|
6
11
|
const DOC_BASE = 'https://ata-validator.com/e/';
|
|
7
12
|
|
|
13
|
+
// The number of characters JSON.stringify would emit for `v`, or -1 as soon as
|
|
14
|
+
// that passes `budget` or the value holds something a plain walk cannot size.
|
|
15
|
+
// The point is the bound: `reprValue` only ever shows a body of 60 characters,
|
|
16
|
+
// so sizing one never needs to read further than that, and a `required` error
|
|
17
|
+
// on a large document used to serialise the whole container to find out it was
|
|
18
|
+
// too big. Keys are counted unescaped, which makes the result a lower bound,
|
|
19
|
+
// and the caller rechecks the real length on the string it ends up building.
|
|
20
|
+
function jsonSizeWithin (v, budget) {
|
|
21
|
+
if (budget < 0) return -1;
|
|
22
|
+
if (v === null) return 4;
|
|
23
|
+
const t = typeof v;
|
|
24
|
+
if (t === 'number') return Number.isFinite(v) ? String(v).length : 4;
|
|
25
|
+
if (t === 'boolean') return v ? 4 : 5;
|
|
26
|
+
if (t === 'string') return v.length + 2 > budget ? -1 : v.length + 2;
|
|
27
|
+
if (t !== 'object') return -1;
|
|
28
|
+
// A value with its own toJSON serialises to something only that method can
|
|
29
|
+
// say. Sizing it exactly would mean running it, and a Date's toJSON alone
|
|
30
|
+
// costs more than serialising the object around it, so this returns the
|
|
31
|
+
// smallest thing it could produce. Undershooting is safe: the caller only
|
|
32
|
+
// uses the estimate to decide whether to try, and rechecks the real length.
|
|
33
|
+
if (typeof v.toJSON === 'function') return 2;
|
|
34
|
+
if (Array.isArray(v)) {
|
|
35
|
+
let n = 2;
|
|
36
|
+
for (let i = 0; i < v.length; i++) {
|
|
37
|
+
n += i === 0 ? 0 : 1;
|
|
38
|
+
if (n > budget) return -1;
|
|
39
|
+
const c = jsonSizeWithin(v[i], budget - n);
|
|
40
|
+
if (c < 0) return -1;
|
|
41
|
+
n += c;
|
|
42
|
+
if (n > budget) return -1;
|
|
43
|
+
}
|
|
44
|
+
return n;
|
|
45
|
+
}
|
|
46
|
+
const proto = Object.getPrototypeOf(v);
|
|
47
|
+
if (proto !== Object.prototype && proto !== null) return -1;
|
|
48
|
+
let n = 2;
|
|
49
|
+
let first = true;
|
|
50
|
+
// for-in rather than Object.keys: the prototype is checked above, so there
|
|
51
|
+
// is nothing inherited to enumerate, and this walk runs per error.
|
|
52
|
+
for (const k in v) {
|
|
53
|
+
const val = v[k];
|
|
54
|
+
// Keys JSON drops cost nothing, but they also cost nothing to skip.
|
|
55
|
+
if (val === undefined || typeof val === 'function' || typeof val === 'symbol') continue;
|
|
56
|
+
// Charge the key before reading the value, so a budget already spent
|
|
57
|
+
// returns without descending into it.
|
|
58
|
+
n += k.length + 3 + (first ? 0 : 1);
|
|
59
|
+
if (n > budget) return -1;
|
|
60
|
+
first = false;
|
|
61
|
+
const c = jsonSizeWithin(val, budget - n);
|
|
62
|
+
if (c < 0) return -1;
|
|
63
|
+
n += c;
|
|
64
|
+
if (n > budget) return -1;
|
|
65
|
+
}
|
|
66
|
+
return n;
|
|
67
|
+
}
|
|
68
|
+
|
|
8
69
|
function reprValue (v) {
|
|
9
70
|
if (v === undefined) return 'undefined';
|
|
10
71
|
if (v === null) return 'null';
|
|
@@ -16,13 +77,21 @@ function reprValue (v) {
|
|
|
16
77
|
if (t === 'number' || t === 'boolean') return String(v);
|
|
17
78
|
if (Array.isArray(v)) return `[array, ${v.length} items]`;
|
|
18
79
|
if (t === 'object') {
|
|
80
|
+
if (jsonSizeWithin(v, 60) >= 0) {
|
|
81
|
+
try {
|
|
82
|
+
const s = JSON.stringify(v);
|
|
83
|
+
if (s !== undefined && s.length <= 60) return s;
|
|
84
|
+
} catch {
|
|
85
|
+
return '[object, unserializable]';
|
|
86
|
+
}
|
|
87
|
+
}
|
|
88
|
+
let n;
|
|
19
89
|
try {
|
|
20
|
-
|
|
21
|
-
if (s.length <= 60) return s;
|
|
22
|
-
return `[object, ~${(s.length / 1024).toFixed(1)}KB]`;
|
|
90
|
+
n = Object.keys(v).length;
|
|
23
91
|
} catch {
|
|
24
92
|
return '[object, unserializable]';
|
|
25
93
|
}
|
|
94
|
+
return `[object, ${n} ${n === 1 ? 'key' : 'keys'}]`;
|
|
26
95
|
}
|
|
27
96
|
return `[${t}]`;
|
|
28
97
|
}
|
|
@@ -45,28 +114,14 @@ function expectedFor (err) {
|
|
|
45
114
|
}
|
|
46
115
|
}
|
|
47
116
|
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
// pass on the segments that carry no `~`.
|
|
117
|
+
// The value the error's pointer resolves to, or MISSING when the pointer walks
|
|
118
|
+
// out of the document. Resolved once per error and handed to both the
|
|
119
|
+
// `received` repr and the suggestion sources, which used to walk it each.
|
|
120
|
+
function resolveAt (err, data) {
|
|
121
|
+
if (!data && data !== 0 && data !== false) return MISSING;
|
|
54
122
|
const p = err.instancePath || err.path || '';
|
|
55
|
-
if (!p) return
|
|
56
|
-
|
|
57
|
-
let cur = data;
|
|
58
|
-
let i = p.charCodeAt(0) === 47 ? 1 : 0; // 47 is '/'
|
|
59
|
-
for (;;) {
|
|
60
|
-
let j = p.indexOf('/', i);
|
|
61
|
-
if (j === -1) j = len;
|
|
62
|
-
let seg = p.slice(i, j);
|
|
63
|
-
if (seg.indexOf('~') !== -1) seg = seg.replace(/~1/g, '/').replace(/~0/g, '~');
|
|
64
|
-
if (cur == null) return undefined;
|
|
65
|
-
cur = cur[seg];
|
|
66
|
-
if (j === len) break;
|
|
67
|
-
i = j + 1;
|
|
68
|
-
}
|
|
69
|
-
return reprValue(cur);
|
|
123
|
+
if (!p) return data;
|
|
124
|
+
return resolvePointer(data, p, MISSING);
|
|
70
125
|
}
|
|
71
126
|
|
|
72
127
|
/**
|
|
@@ -146,13 +201,15 @@ function enrich (rawErr, opts) {
|
|
|
146
201
|
const meta = CODES[code];
|
|
147
202
|
const path = rawErr.instancePath != null ? rawErr.instancePath : (rawErr.path || '');
|
|
148
203
|
|
|
204
|
+
const at = data !== undefined ? resolveAt(rawErr, data) : MISSING;
|
|
205
|
+
|
|
149
206
|
const out = {
|
|
150
207
|
code,
|
|
151
208
|
message: rawErr.message || (meta && meta.headline) || 'validation failed',
|
|
152
209
|
keyword,
|
|
153
210
|
path,
|
|
154
211
|
expected: expectedFor(rawErr),
|
|
155
|
-
received:
|
|
212
|
+
received: at === MISSING ? undefined : reprValue(at),
|
|
156
213
|
schemaPath: rawErr.schemaPath,
|
|
157
214
|
docUrl: DOC_BASE + code,
|
|
158
215
|
// Back-compat aliases (additive, present in both rich and legacy paths)
|
|
@@ -187,7 +244,7 @@ function enrich (rawErr, opts) {
|
|
|
187
244
|
// Suggestion attachment runs last so it can read `received`, `params`, and
|
|
188
245
|
// `keyword` from the enriched shape. `data` is the full input object so the
|
|
189
246
|
// required-typo source can scan sibling keys.
|
|
190
|
-
const sugg = suggestFor(out,
|
|
247
|
+
const sugg = suggestFor(out, data, at === MISSING ? undefined : at);
|
|
191
248
|
if (sugg) out.suggestion = sugg;
|
|
192
249
|
|
|
193
250
|
const detail = detailFor(rawErr, out);
|
package/lib/formats.js
CHANGED
|
@@ -224,24 +224,102 @@ function hostname (s) {
|
|
|
224
224
|
return labelLength !== 0 && previous !== 45;
|
|
225
225
|
}
|
|
226
226
|
|
|
227
|
-
//
|
|
228
|
-
//
|
|
229
|
-
//
|
|
227
|
+
// Character classes as tables rather than comparison chains. Nine `===` tests
|
|
228
|
+
// per character more than doubled the cost of the scan they guarded, measured
|
|
229
|
+
// against the same loop doing a single range check; an indexed byte read costs
|
|
230
|
+
// the same whatever the class holds.
|
|
231
|
+
const URI_CHAR = new Uint8Array(128);
|
|
232
|
+
for (let i = 33; i < 127; i++) URI_CHAR[i] = 1;
|
|
233
|
+
// " < > \ ^ ` { | } are printable and still not allowed in a URI.
|
|
234
|
+
for (const c of [34, 60, 62, 92, 94, 96, 123, 124, 125]) URI_CHAR[c] = 0;
|
|
235
|
+
const HEX_CHAR = new Uint8Array(128);
|
|
236
|
+
for (let i = 48; i < 58; i++) HEX_CHAR[i] = 1;
|
|
237
|
+
for (let i = 97; i < 103; i++) HEX_CHAR[i] = 1;
|
|
238
|
+
for (let i = 65; i < 71; i++) HEX_CHAR[i] = 1;
|
|
239
|
+
// Scheme characters: letters, digits, "+", "-", ".".
|
|
240
|
+
const SCHEME_CHAR = new Uint8Array(128);
|
|
241
|
+
for (let i = 48; i < 58; i++) SCHEME_CHAR[i] = 1;
|
|
242
|
+
for (let i = 97; i < 123; i++) SCHEME_CHAR[i] = 1;
|
|
243
|
+
for (let i = 65; i < 91; i++) SCHEME_CHAR[i] = 1;
|
|
244
|
+
SCHEME_CHAR[43] = 1; SCHEME_CHAR[45] = 1; SCHEME_CHAR[46] = 1;
|
|
245
|
+
|
|
246
|
+
// A scheme followed by a colon, an optional authority, and no character a URI
|
|
247
|
+
// cannot hold anywhere after the colon.
|
|
248
|
+
//
|
|
249
|
+
// One walk answers all of it. The authority's landmarks, the last "@", the
|
|
250
|
+
// colons after it and whether a bracket appeared before it, are recorded while
|
|
251
|
+
// the characters are being checked, so the authority is never cut out of the
|
|
252
|
+
// string or scanned again. Reading it out with `slice` and asking the copy for
|
|
253
|
+
// `lastIndexOf` and two regular expressions allocated a string per URI, and a
|
|
254
|
+
// document carrying a handful of URLs paid that per field.
|
|
230
255
|
function uri (s) {
|
|
231
256
|
const n = s.length;
|
|
232
257
|
if (n === 0) return false;
|
|
233
|
-
|
|
234
|
-
|
|
258
|
+
let c = s.charCodeAt(0);
|
|
259
|
+
// A scheme opens with a letter, never a digit or a sign.
|
|
260
|
+
if (c > 127 || SCHEME_CHAR[c] === 0 || (c >= 48 && c <= 57) || c === 43 || c === 45 || c === 46) return false;
|
|
235
261
|
let colon = -1;
|
|
236
262
|
for (let i = 1; i < n; i++) {
|
|
237
|
-
|
|
263
|
+
c = s.charCodeAt(i);
|
|
238
264
|
if (c === 58) { colon = i; break; }
|
|
239
|
-
|
|
240
|
-
if (!(isDigit(c) || (c >= 97 && c <= 122) || (c >= 65 && c <= 90) ||
|
|
241
|
-
c === 43 || c === 45 || c === 46)) return false;
|
|
265
|
+
if (c > 127 || SCHEME_CHAR[c] === 0) return false;
|
|
242
266
|
}
|
|
243
267
|
if (colon === -1) return false;
|
|
244
|
-
|
|
268
|
+
|
|
269
|
+
let i = colon + 1;
|
|
270
|
+
let authEnd = n;
|
|
271
|
+
if (s.charCodeAt(i) === 47 && s.charCodeAt(i + 1) === 47) {
|
|
272
|
+
const authStart = i + 2;
|
|
273
|
+
let at = -1, firstColon = -1, lastColon = -1, sawBracket = 0, bracketBeforeAt = 0;
|
|
274
|
+
let hostStart = authStart;
|
|
275
|
+
authEnd = -1;
|
|
276
|
+
for (i = authStart; i < n; i++) {
|
|
277
|
+
c = s.charCodeAt(i);
|
|
278
|
+
if (c > 127 || URI_CHAR[c] === 0) return false;
|
|
279
|
+
if (c === 47 || c === 63 || c === 35) { authEnd = i; break; }
|
|
280
|
+
if (c === 37) {
|
|
281
|
+
const h1 = s.charCodeAt(i + 1), h2 = s.charCodeAt(i + 2);
|
|
282
|
+
// A "%" at the end of the string reads as NaN, which no comparison
|
|
283
|
+
// admits; `<= 127` is written that way so NaN takes the reject.
|
|
284
|
+
if (!(h1 <= 127) || !(h2 <= 127) || HEX_CHAR[h1] === 0 || HEX_CHAR[h2] === 0) return false;
|
|
285
|
+
i += 2;
|
|
286
|
+
continue;
|
|
287
|
+
}
|
|
288
|
+
// The last "@" ends userinfo, so everything recorded before it belongs
|
|
289
|
+
// to a part that has its own rules and is dropped here.
|
|
290
|
+
if (c === 64) { bracketBeforeAt = sawBracket; at = i; firstColon = -1; lastColon = -1; hostStart = i + 1; }
|
|
291
|
+
else if (c === 58) { if (firstColon === -1) firstColon = i; lastColon = i; }
|
|
292
|
+
else if (c === 91 || c === 93) { sawBracket = 1; }
|
|
293
|
+
}
|
|
294
|
+
if (authEnd === -1) authEnd = n;
|
|
295
|
+
// Brackets belong to the host, so one before the "@" is not userinfo.
|
|
296
|
+
if (at !== -1 && bracketBeforeAt) return false;
|
|
297
|
+
if (s.charCodeAt(hostStart) === 91) { // '['
|
|
298
|
+
let close = -1;
|
|
299
|
+
for (let j = hostStart + 1; j < authEnd; j++) { if (s.charCodeAt(j) === 93) { close = j; break; } }
|
|
300
|
+
if (close === -1) return false;
|
|
301
|
+
if (close + 1 !== authEnd) {
|
|
302
|
+
if (s.charCodeAt(close + 1) !== 58) return false;
|
|
303
|
+
if (!allDigits(s, close + 2, authEnd)) return false;
|
|
304
|
+
}
|
|
305
|
+
} else if (lastColon !== -1) {
|
|
306
|
+
// A host holding more than one colon is an IPv6 address, and those must
|
|
307
|
+
// be bracketed. Without that rule the last group reads as a port number.
|
|
308
|
+
if (firstColon !== lastColon) return false;
|
|
309
|
+
if (!allDigits(s, lastColon + 1, authEnd)) return false;
|
|
310
|
+
}
|
|
311
|
+
i = authEnd;
|
|
312
|
+
}
|
|
313
|
+
for (; i < n; i++) {
|
|
314
|
+
c = s.charCodeAt(i);
|
|
315
|
+
if (c > 127 || URI_CHAR[c] === 0) return false;
|
|
316
|
+
if (c === 37) {
|
|
317
|
+
const h1 = s.charCodeAt(i + 1), h2 = s.charCodeAt(i + 2);
|
|
318
|
+
if (!(h1 <= 127) || !(h2 <= 127) || HEX_CHAR[h1] === 0 || HEX_CHAR[h2] === 0) return false;
|
|
319
|
+
i += 2;
|
|
320
|
+
}
|
|
321
|
+
}
|
|
322
|
+
return true;
|
|
245
323
|
}
|
|
246
324
|
|
|
247
325
|
// The characters a URI cannot hold: C0 controls, space, DEL, and the rest of
|
|
@@ -287,34 +365,68 @@ const uriCharsSource = (v, from) =>
|
|
|
287
365
|
'if(!((_h1>=48&&_h1<=57)||(_h1>=97&&_h1<=102)||(_h1>=65&&_h1<=70))||' +
|
|
288
366
|
'!((_h2>=48&&_h2<=57)||(_h2>=97&&_h2<=102)||(_h2>=65&&_h2<=70)))return false;_ri+=2}}';
|
|
289
367
|
|
|
368
|
+
function allDigits (s, from, to) {
|
|
369
|
+
for (let i = from; i < to; i++) {
|
|
370
|
+
const c = s.charCodeAt(i);
|
|
371
|
+
if (c < 48 || c > 57) return false;
|
|
372
|
+
}
|
|
373
|
+
return true;
|
|
374
|
+
}
|
|
375
|
+
|
|
290
376
|
// The authority sits between "//" and the next "/", "?" or "#". A bracketed
|
|
291
377
|
// host must close, and a port is digits.
|
|
378
|
+
//
|
|
379
|
+
// Everything here reads the original string through index bounds. Cutting the
|
|
380
|
+
// authority out with `slice` and asking it for `lastIndexOf` and two regular
|
|
381
|
+
// expressions allocated a string for every URI validated, and a document
|
|
382
|
+
// carrying a handful of URLs paid that per field; the authority of an ordinary
|
|
383
|
+
// URL is a dozen characters, so walking it twice costs less than copying it
|
|
384
|
+
// once.
|
|
292
385
|
function uriAuthority (s, start) {
|
|
293
386
|
if (s.charCodeAt(start) !== 47 || s.charCodeAt(start + 1) !== 47) return true;
|
|
294
|
-
|
|
295
|
-
|
|
387
|
+
const n = s.length;
|
|
388
|
+
const authStart = start + 2;
|
|
389
|
+
let authEnd = n;
|
|
390
|
+
for (let i = authStart; i < n; i++) {
|
|
296
391
|
const c = s.charCodeAt(i);
|
|
297
|
-
if (c === 47 || c === 63 || c === 35) {
|
|
298
|
-
}
|
|
299
|
-
|
|
300
|
-
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
392
|
+
if (c === 47 || c === 63 || c === 35) { authEnd = i; break; }
|
|
393
|
+
}
|
|
394
|
+
// The last "@" splits userinfo from the host. Brackets belong to the host,
|
|
395
|
+
// so one before the "@" is not userinfo.
|
|
396
|
+
let at = -1;
|
|
397
|
+
for (let i = authEnd - 1; i >= authStart; i--) {
|
|
398
|
+
if (s.charCodeAt(i) === 64) { at = i; break; }
|
|
399
|
+
}
|
|
400
|
+
if (at !== -1) {
|
|
401
|
+
for (let i = authStart; i < at; i++) {
|
|
402
|
+
const c = s.charCodeAt(i);
|
|
403
|
+
if (c === 91 || c === 93) return false;
|
|
404
|
+
}
|
|
405
|
+
}
|
|
406
|
+
const hpStart = at === -1 ? authStart : at + 1;
|
|
407
|
+
if (s.charCodeAt(hpStart) === 91) { // '['
|
|
408
|
+
let close = -1;
|
|
409
|
+
for (let i = hpStart + 1; i < authEnd; i++) {
|
|
410
|
+
if (s.charCodeAt(i) === 93) { close = i; break; }
|
|
411
|
+
}
|
|
306
412
|
if (close === -1) return false;
|
|
307
|
-
|
|
308
|
-
if (
|
|
309
|
-
|
|
310
|
-
|
|
413
|
+
if (close + 1 === authEnd) return true;
|
|
414
|
+
if (s.charCodeAt(close + 1) !== 58) return false;
|
|
415
|
+
return allDigits(s, close + 2, authEnd);
|
|
416
|
+
}
|
|
417
|
+
let firstColon = -1;
|
|
418
|
+
let lastColon = -1;
|
|
419
|
+
for (let i = hpStart; i < authEnd; i++) {
|
|
420
|
+
if (s.charCodeAt(i) === 58) {
|
|
421
|
+
if (firstColon === -1) firstColon = i;
|
|
422
|
+
lastColon = i;
|
|
423
|
+
}
|
|
311
424
|
}
|
|
312
|
-
|
|
313
|
-
if (colon === -1) return true;
|
|
425
|
+
if (lastColon === -1) return true;
|
|
314
426
|
// A host holding more than one colon is an IPv6 address, and those must be
|
|
315
427
|
// bracketed. Without that rule the last group reads as a port number.
|
|
316
|
-
if (
|
|
317
|
-
return
|
|
428
|
+
if (firstColon !== lastColon) return false;
|
|
429
|
+
return allDigits(s, lastColon + 1, authEnd);
|
|
318
430
|
}
|
|
319
431
|
const uriAuthoritySource = (v, start) =>
|
|
320
432
|
`if(${v}.charCodeAt(${start})===47&&${v}.charCodeAt(${start}+1)===47){` +
|
|
@@ -622,6 +734,22 @@ function timeSource (v, isStr) {
|
|
|
622
734
|
'if(_u!==1439)return false}');
|
|
623
735
|
}
|
|
624
736
|
|
|
737
|
+
// The same walk as `uri` above, as source, for the code generator to hoist
|
|
738
|
+
// once per compiled function and call. Two copies of one algorithm is what
|
|
739
|
+
// this file has always done for every format, so they sit together here and
|
|
740
|
+
// tests/test_uri_helper_parity.js holds them to the same answer over a corpus
|
|
741
|
+
// built from the walk's own boundaries.
|
|
742
|
+
const URI_HELPER_TABLES = 'const _uct=new Uint8Array(128);for(let _i=33;_i<127;_i++)_uct[_i]=1;' +
|
|
743
|
+
'_uct[34]=_uct[60]=_uct[62]=_uct[92]=_uct[94]=_uct[96]=_uct[123]=_uct[124]=_uct[125]=0;' +
|
|
744
|
+
'const _uch=new Uint8Array(128);for(let _i=48;_i<58;_i++)_uch[_i]=1;' +
|
|
745
|
+
'for(let _i=97;_i<103;_i++)_uch[_i]=1;for(let _i=65;_i<71;_i++)_uch[_i]=1;' +
|
|
746
|
+
'const _ucs=new Uint8Array(128);for(let _i=48;_i<58;_i++)_ucs[_i]=1;' +
|
|
747
|
+
'for(let _i=97;_i<123;_i++)_ucs[_i]=1;for(let _i=65;_i<91;_i++)_ucs[_i]=1;' +
|
|
748
|
+
'_ucs[43]=1;_ucs[45]=1;_ucs[46]=1';
|
|
749
|
+
const URI_HELPER_BODY = 'const _n=_s.length;if(_n===0)return false;let _c=_s.charCodeAt(0);if(_c>127||_ucs[_c]===0||(_c>=48&&_c<=57)||_c===43||_c===45||_c===46)return false;let _co=-1;for(let _i=1;_i<_n;_i++){_c=_s.charCodeAt(_i);if(_c===58){_co=_i;break}if(_c>127||_ucs[_c]===0)return false}if(_co===-1)return false;let _i=_co+1;let _ae=_n;if(_s.charCodeAt(_i)===47&&_s.charCodeAt(_i+1)===47){const _as=_i+2;let _at=-1,_fc=-1,_lc=-1,_br=0,_ba=0,_hs=_as;_ae=-1;for(_i=_as;_i<_n;_i++){_c=_s.charCodeAt(_i);if(_c>127||_uct[_c]===0)return false;if(_c===47||_c===63||_c===35){_ae=_i;break}if(_c===37){const _h1=_s.charCodeAt(_i+1),_h2=_s.charCodeAt(_i+2);if(!(_h1<=127)||!(_h2<=127)||_uch[_h1]===0||_uch[_h2]===0)return false;_i+=2;continue}if(_c===64){_ba=_br;_at=_i;_fc=-1;_lc=-1;_hs=_i+1}else if(_c===58){if(_fc===-1)_fc=_i;_lc=_i}else if(_c===91||_c===93){_br=1}}if(_ae===-1)_ae=_n;if(_at!==-1&&_ba)return false;if(_s.charCodeAt(_hs)===91){let _cl=-1;for(let _j=_hs+1;_j<_ae;_j++){if(_s.charCodeAt(_j)===93){_cl=_j;break}}if(_cl===-1)return false;if(_cl+1!==_ae){if(_s.charCodeAt(_cl+1)!==58)return false;for(let _j=_cl+2;_j<_ae;_j++){const _d=_s.charCodeAt(_j);if(_d<48||_d>57)return false}}}else if(_lc!==-1){if(_fc!==_lc)return false;for(let _j=_lc+1;_j<_ae;_j++){const _d=_s.charCodeAt(_j);if(_d<48||_d>57)return false}}_i=_ae}for(;_i<_n;_i++){_c=_s.charCodeAt(_i);if(_c>127||_uct[_c]===0)return false;if(_c===37){const _h1=_s.charCodeAt(_i+1),_h2=_s.charCodeAt(_i+2);if(!(_h1<=127)||!(_h2<=127)||_uch[_h1]===0||_uch[_h2]===0)return false;_i+=2}}return true;';
|
|
750
|
+
const uriHelperSource = (name) =>
|
|
751
|
+
URI_HELPER_TABLES + ';function ' + name + '(_s){' + URI_HELPER_BODY + '}';
|
|
752
|
+
|
|
625
753
|
function uriSource (v, isStr) {
|
|
626
754
|
return guard(v, isStr, `const _n=${v}.length;if(_n===0)return false;` +
|
|
627
755
|
`const _f=${v}.charCodeAt(0);if(!((_f>=97&&_f<=122)||(_f>=65&&_f<=90)))return false;` +
|
|
@@ -814,4 +942,4 @@ function durationSource (v, isStr) {
|
|
|
814
942
|
return isStr ? inner : `if(typeof ${v}==='string'&&${inner.slice(3)}`;
|
|
815
943
|
}
|
|
816
944
|
|
|
817
|
-
module.exports = { date, ipv4, dateTime, ipv6, hostname, uri, uriReference, uuid, time, noReserved, jsonPointer, relativeJsonPointer, uriTemplate, email, duration, uriChars, uriAuthority, uriCharsSource, uriAuthoritySource, iri, iriReference, idnEmail, iriSource, iriReferenceSource, idnEmailSource, dateSource, ipv4Source, dateTimeSource, ipv6Source, hostnameSource, uriSource, uuidSource, timeSource, noReservedSource, jsonPointerSource, relativeJsonPointerSource, uriTemplateSource, emailSource, durationSource };
|
|
945
|
+
module.exports = { uriHelperSource, date, ipv4, dateTime, ipv6, hostname, uri, uriReference, uuid, time, noReserved, jsonPointer, relativeJsonPointer, uriTemplate, email, duration, uriChars, uriAuthority, uriCharsSource, uriAuthoritySource, iri, iriReference, idnEmail, iriSource, iriReferenceSource, idnEmailSource, dateSource, ipv4Source, dateTimeSource, ipv6Source, hostnameSource, uriSource, uuidSource, timeSource, noReservedSource, jsonPointerSource, relativeJsonPointerSource, uriTemplateSource, emailSource, durationSource };
|
package/lib/js-compiler.js
CHANGED
|
@@ -3037,6 +3037,13 @@ function genCode(schema, v, lines, ctx, knownType) {
|
|
|
3037
3037
|
// the check was inlined at each of them.
|
|
3038
3038
|
const EMAIL_HELPER = `function _em(_s){${_formats.emailSource('_s', true)}return true}`
|
|
3039
3039
|
|
|
3040
|
+
// `uri` is the most expensive format an ordinary document carries: a schema
|
|
3041
|
+
// with a handful of URL fields spent more time here than on every structural
|
|
3042
|
+
// check put together. The walk reads its character classes out of hoisted
|
|
3043
|
+
// tables, so it is declared once per compiled function rather than inlined at
|
|
3044
|
+
// each call site, where the tables would have to be rebuilt.
|
|
3045
|
+
const URI_HELPER = _formats.uriHelperSource('_uri')
|
|
3046
|
+
|
|
3040
3047
|
const FORMAT_CODEGEN = {
|
|
3041
3048
|
email: (v, isStr, ctx) => {
|
|
3042
3049
|
if (!ctx) return _formats.emailSource(v, isStr)
|
|
@@ -3058,7 +3065,11 @@ const FORMAT_CODEGEN = {
|
|
|
3058
3065
|
'date-time': _formats.dateTimeSource,
|
|
3059
3066
|
time: _formats.timeSource,
|
|
3060
3067
|
duration: _formats.durationSource,
|
|
3061
|
-
uri:
|
|
3068
|
+
uri: (v, isStr, ctx) => {
|
|
3069
|
+
if (!ctx) return _formats.uriSource(v, isStr)
|
|
3070
|
+
hoistOnce(ctx, '_uriHoisted', URI_HELPER)
|
|
3071
|
+
return isStr ? `if(!_uri(${v}))return false` : `if(typeof ${v}==='string'&&!_uri(${v}))return false`
|
|
3072
|
+
},
|
|
3062
3073
|
'uri-reference': (v, isStr) => isStr
|
|
3063
3074
|
? `{${_formats.uriCharsSource(v, '0')}}`
|
|
3064
3075
|
: `if(typeof ${v}==='string'){${_formats.uriCharsSource(v, '0')}}`,
|
package/lib/levenshtein.js
CHANGED
|
@@ -21,19 +21,31 @@ function levenshtein (a, b, maxDistance) {
|
|
|
21
21
|
}
|
|
22
22
|
let prev = scratchA;
|
|
23
23
|
let curr = scratchB;
|
|
24
|
-
|
|
24
|
+
const bn = b.length;
|
|
25
|
+
for (let j = 0; j <= bn; j++) prev[j] = j;
|
|
25
26
|
for (let i = 1; i <= a.length; i++) {
|
|
26
27
|
curr[0] = i;
|
|
27
28
|
let rowMin = i;
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
29
|
+
// charCodeAt rather than indexing: reading a character by index allocates
|
|
30
|
+
// a one-character string, and this is the innermost loop.
|
|
31
|
+
const ca = a.charCodeAt(i - 1);
|
|
32
|
+
for (let j = 1; j <= bn; j++) {
|
|
33
|
+
const cost = ca === b.charCodeAt(j - 1) ? 0 : 1;
|
|
34
|
+
let m = prev[j - 1] + cost;
|
|
35
|
+
const del = curr[j - 1] + 1;
|
|
36
|
+
if (del < m) m = del;
|
|
37
|
+
const ins = prev[j] + 1;
|
|
38
|
+
if (ins < m) m = ins;
|
|
39
|
+
curr[j] = m;
|
|
40
|
+
if (m < rowMin) rowMin = m;
|
|
32
41
|
}
|
|
33
42
|
if (rowMin > max) return Infinity;
|
|
34
|
-
|
|
43
|
+
// A destructured swap builds an array per row; a temporary does not.
|
|
44
|
+
const t = prev;
|
|
45
|
+
prev = curr;
|
|
46
|
+
curr = t;
|
|
35
47
|
}
|
|
36
|
-
return prev[
|
|
48
|
+
return prev[bn];
|
|
37
49
|
}
|
|
38
50
|
|
|
39
51
|
module.exports = { levenshtein };
|
package/lib/pointer.js
ADDED
|
@@ -0,0 +1,46 @@
|
|
|
1
|
+
'use strict';
|
|
2
|
+
|
|
3
|
+
/**
|
|
4
|
+
* Resolve a JSON pointer against a document, for the diagnostic paths that run
|
|
5
|
+
* once per error on a rejected payload.
|
|
6
|
+
*
|
|
7
|
+
* Segments are read straight out of the pointer string: no leading-slash
|
|
8
|
+
* regex, no parts array, and no unescape pass on the segments that carry no
|
|
9
|
+
* `~`. A pointer that leaves the document returns rather than throwing,
|
|
10
|
+
* because a diagnostic must not fail where validation succeeded.
|
|
11
|
+
*
|
|
12
|
+
* `missing` separates the two ways a pointer yields nothing: a key that is
|
|
13
|
+
* absent from a container that exists resolves to undefined, while a pointer
|
|
14
|
+
* that walks through a null or a primitive returns `missing`. Callers that
|
|
15
|
+
* format the value need that apart, since the first is a value worth printing
|
|
16
|
+
* and the second is a path that was never in the document.
|
|
17
|
+
*
|
|
18
|
+
* @param {*} data the document the pointer is read against
|
|
19
|
+
* @param {string} pointer an RFC 6901 pointer, '' for the document itself
|
|
20
|
+
* @param {*} [missing] returned when the walk leaves the document
|
|
21
|
+
* @returns {*} the value at the pointer, `missing` if the walk broke
|
|
22
|
+
*/
|
|
23
|
+
function resolvePointer (data, pointer, missing) {
|
|
24
|
+
if (!pointer) return data;
|
|
25
|
+
const len = pointer.length;
|
|
26
|
+
let cur = data;
|
|
27
|
+
let i = pointer.charCodeAt(0) === 47 ? 1 : 0; // 47 is '/'
|
|
28
|
+
for (;;) {
|
|
29
|
+
let j = pointer.indexOf('/', i);
|
|
30
|
+
if (j === -1) j = len;
|
|
31
|
+
let seg = pointer.slice(i, j);
|
|
32
|
+
if (seg.indexOf('~') !== -1) seg = seg.replace(/~1/g, '/').replace(/~0/g, '~');
|
|
33
|
+
if (cur == null) return missing;
|
|
34
|
+
cur = cur[seg];
|
|
35
|
+
if (j === len) break;
|
|
36
|
+
i = j + 1;
|
|
37
|
+
}
|
|
38
|
+
return cur;
|
|
39
|
+
}
|
|
40
|
+
|
|
41
|
+
// Passed where a caller has no resolved value to offer, which is not the same
|
|
42
|
+
// as resolving to undefined: the first means walk it yourself, the second
|
|
43
|
+
// means the document really holds nothing there.
|
|
44
|
+
const UNRESOLVED = Symbol('ata.pointer.unresolved');
|
|
45
|
+
|
|
46
|
+
module.exports = { resolvePointer, UNRESOLVED };
|
package/lib/retry-message.js
CHANGED
|
@@ -22,15 +22,39 @@
|
|
|
22
22
|
// `received`. The trap is that joining `message`, which is the obvious thing to
|
|
23
23
|
// do, throws both away. So this is one call that does not.
|
|
24
24
|
|
|
25
|
+
// `multipleOf: 0.01` is how a schema says money, and a model acts on "rounded
|
|
26
|
+
// to 2 decimal places" where it does not reliably act on "a multiple of 0.01".
|
|
27
|
+
// describeSchema found that by measurement; the same wording belongs here, or
|
|
28
|
+
// the two halves of the loop describe the same rule differently.
|
|
29
|
+
function phrase (error) {
|
|
30
|
+
const p = error.params || {};
|
|
31
|
+
if (error.keyword === 'multipleOf' && typeof p.multipleOf === 'number') {
|
|
32
|
+
const m = p.multipleOf;
|
|
33
|
+
if (m > 0 && m < 1) {
|
|
34
|
+
const places = Math.round(Math.log10(1 / m));
|
|
35
|
+
if (Math.abs(Math.pow(10, -places) - m) < Number.EPSILON * 8) {
|
|
36
|
+
return `must be rounded to ${places} decimal place${places === 1 ? '' : 's'}`;
|
|
37
|
+
}
|
|
38
|
+
}
|
|
39
|
+
}
|
|
40
|
+
return null;
|
|
41
|
+
}
|
|
42
|
+
|
|
25
43
|
function line (error) {
|
|
26
44
|
const where = error.instancePath || error.path || '';
|
|
27
|
-
const body =
|
|
45
|
+
const body = phrase(error) ||
|
|
46
|
+
(typeof error.detail === 'string' && error.detail ? error.detail : error.message);
|
|
28
47
|
// `detail` usually quotes the offending value already; saying it twice reads
|
|
29
48
|
// like a stutter in a string that is going into a prompt. A container is
|
|
30
49
|
// summarised as `[object, ~0.1KB]` rather than shown, which tells a model
|
|
31
50
|
// nothing it cannot see in its own output, so that is left out too.
|
|
32
51
|
const raw = error.received === undefined || error.received === null ? '' : String(error.received);
|
|
33
|
-
|
|
52
|
+
// A container tells a model nothing its own output does not already show,
|
|
53
|
+
// and it is the most expensive thing you can put in a prompt. That covers
|
|
54
|
+
// both the `[object, ~0.1KB]` summary and a whole serialized object.
|
|
55
|
+
const head = raw.charAt(0);
|
|
56
|
+
const useless = (head === '[' && /^\[(object|array)\b/.test(raw)) || head === '{' ||
|
|
57
|
+
(head === '[' && raw.charAt(raw.length - 1) === ']');
|
|
34
58
|
const received = useless ? '' : raw;
|
|
35
59
|
const got = received && body && !body.includes(received) ? `, got ${received}` : '';
|
|
36
60
|
return `${where || '/'}: ${body}${got}`;
|
package/lib/suggestions.js
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
'use strict';
|
|
2
2
|
|
|
3
3
|
const { levenshtein } = require('./levenshtein');
|
|
4
|
+
const { resolvePointer: walk, UNRESOLVED } = require('./pointer');
|
|
4
5
|
|
|
5
6
|
// Hand-coded format hints. Keep <=60 chars per text.
|
|
6
7
|
const FORMAT_HINTS = {
|
|
@@ -46,8 +47,15 @@ function suggestEnumTypo (received, enumValues) {
|
|
|
46
47
|
return null;
|
|
47
48
|
}
|
|
48
49
|
|
|
50
|
+
// The widest container a typo hint is offered on. `suggestEnumTypo` caps its
|
|
51
|
+
// candidate list for the same two reasons: past a certain width the nearest
|
|
52
|
+
// key stops being evidence of a typo, and the scan is a distance computation
|
|
53
|
+
// per key on a value the caller has already rejected.
|
|
54
|
+
const MAX_TYPO_CANDIDATES = 64;
|
|
55
|
+
|
|
49
56
|
function suggestRequiredTypo (missing, presentKeys) {
|
|
50
57
|
if (!missing || !Array.isArray(presentKeys)) return null;
|
|
58
|
+
if (presentKeys.length > MAX_TYPO_CANDIDATES) return null;
|
|
51
59
|
for (const k of presentKeys) {
|
|
52
60
|
if (typeof k !== 'string') continue;
|
|
53
61
|
const d = levenshtein(missing, k, 2);
|
|
@@ -58,22 +66,14 @@ function suggestRequiredTypo (missing, presentKeys) {
|
|
|
58
66
|
return null;
|
|
59
67
|
}
|
|
60
68
|
|
|
61
|
-
function suggestFormat (format,
|
|
69
|
+
function suggestFormat (format, raw) {
|
|
62
70
|
const fn = FORMAT_HINTS[format];
|
|
63
71
|
if (!fn) return null;
|
|
64
|
-
// received is repr-truncated string like '"foo"'; unwrap for testing
|
|
65
|
-
let raw = received;
|
|
66
|
-
if (typeof raw === 'string' && raw.startsWith('"') && raw.endsWith('"')) {
|
|
67
|
-
try { raw = JSON.parse(raw); } catch {}
|
|
68
|
-
}
|
|
69
72
|
const text = fn(raw);
|
|
70
73
|
return text ? { text, kind: 'format' } : null;
|
|
71
74
|
}
|
|
72
75
|
|
|
73
|
-
function suggestCoercion (expectedType,
|
|
74
|
-
if (typeof received !== 'string' || !received.startsWith('"') || !received.endsWith('"')) return null;
|
|
75
|
-
let raw;
|
|
76
|
-
try { raw = JSON.parse(received); } catch { return null; }
|
|
76
|
+
function suggestCoercion (expectedType, raw) {
|
|
77
77
|
if (typeof raw !== 'string') return null;
|
|
78
78
|
if (expectedType === 'integer' && /^-?\d+$/.test(raw)) {
|
|
79
79
|
return { text: 'value would coerce; enable `coerceTypes` or pass an integer', kind: 'coercion' };
|
|
@@ -91,26 +91,37 @@ function suggestCoercion (expectedType, received) {
|
|
|
91
91
|
* Apply suggestion sources in priority order. Returns the first hit, or null.
|
|
92
92
|
* @param err Enriched ValidationError (with `received`, `params`, `keyword`)
|
|
93
93
|
* @param data The full input data (for required-typo)
|
|
94
|
+
* @param at The value at `err.path`, already resolved by the caller, or
|
|
95
|
+
* UNRESOLVED when the caller has none to offer
|
|
94
96
|
*/
|
|
95
|
-
function suggestFor (err, data) {
|
|
97
|
+
function suggestFor (err, data, at = UNRESOLVED) {
|
|
98
|
+
// The offending value itself. The caller resolves it once per error and
|
|
99
|
+
// passes it here rather than having each source walk the document again.
|
|
100
|
+
// A caller that passes only an enriched error falls back to un-quoting the
|
|
101
|
+
// repr, which is what every source used to read.
|
|
102
|
+
const raw = at !== UNRESOLVED ? at : parseReceived(err.received);
|
|
96
103
|
if (err.keyword === 'enum') {
|
|
97
|
-
return suggestEnumTypo(
|
|
104
|
+
return suggestEnumTypo(raw, err.params && err.params.allowedValues);
|
|
98
105
|
}
|
|
99
106
|
if (err.keyword === 'required') {
|
|
100
107
|
const missing = err.params && err.params.missingProperty;
|
|
101
|
-
|
|
108
|
+
const path = err.path || '';
|
|
109
|
+
let parentPath = path;
|
|
102
110
|
if (parentPath.endsWith('/' + missing)) parentPath = parentPath.slice(0, -missing.length - 1);
|
|
103
|
-
|
|
111
|
+
// A `required` error usually names the container it failed on, so the
|
|
112
|
+
// value the caller already resolved is the parent. Only the spelling that
|
|
113
|
+
// points at the missing key itself needs a second walk.
|
|
114
|
+
const parent = (parentPath === path && at !== UNRESOLVED) ? at : walk(data, parentPath);
|
|
104
115
|
if (parent && typeof parent === 'object') {
|
|
105
116
|
return suggestRequiredTypo(missing, Object.keys(parent));
|
|
106
117
|
}
|
|
107
118
|
return null;
|
|
108
119
|
}
|
|
109
120
|
if (err.keyword === 'format') {
|
|
110
|
-
return suggestFormat(err.params && err.params.format,
|
|
121
|
+
return suggestFormat(err.params && err.params.format, raw);
|
|
111
122
|
}
|
|
112
123
|
if (err.keyword === 'type') {
|
|
113
|
-
return suggestCoercion(err.params && err.params.type,
|
|
124
|
+
return suggestCoercion(err.params && err.params.type, raw);
|
|
114
125
|
}
|
|
115
126
|
return null;
|
|
116
127
|
}
|
|
@@ -121,12 +132,4 @@ function parseReceived (r) {
|
|
|
121
132
|
return r;
|
|
122
133
|
}
|
|
123
134
|
|
|
124
|
-
function walk (data, pointer) {
|
|
125
|
-
if (!pointer) return data;
|
|
126
|
-
const parts = pointer.replace(/^\//, '').split('/').map(s => s.replace(/~1/g, '/').replace(/~0/g, '~'));
|
|
127
|
-
let cur = data;
|
|
128
|
-
for (const p of parts) { if (cur == null) return undefined; cur = cur[p]; }
|
|
129
|
-
return cur;
|
|
130
|
-
}
|
|
131
|
-
|
|
132
135
|
module.exports = { suggestFor, suggestEnumTypo, suggestRequiredTypo, suggestFormat, suggestCoercion };
|
package/lib/version.js
CHANGED
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "ata-validator",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.22.0",
|
|
4
4
|
"description": "JSON Schema validation with first-class TypeScript and zero runtime cost. AOT compile to per-schema ESM modules with zero validator dependency. Generic Validator<T> for TypeBox/Zod/Valibot composition. Optional runtime API. Standard Schema V1 compatible.",
|
|
5
5
|
"main": "index.js",
|
|
6
6
|
"module": "index.mjs",
|
|
@@ -51,7 +51,7 @@
|
|
|
51
51
|
"release:check": "node scripts/regen-safe-regex-source.js && node tests/test_pack_purity.js && node scripts/check-doc-coverage.js && node tests/test_error_codes_lock.js && node tests/test_safe_regex_source_sync.js && node tests/test_version_sync.js",
|
|
52
52
|
"build": "cmake-js build --target ata",
|
|
53
53
|
"rebuild": "cmake-js rebuild --target ata",
|
|
54
|
-
"test": "node test.js && node tests/test_removed_aot_methods.js && node tests/test_no_native.js && node tests/test_no_eval.js && node tests/test_property_dependencies.js && node tests/test_v1_dialect.js && node tests/test_buffer_path_parity.js && node tests/test_buffer_gate.js && node tests/test_buffer_reject_cost.js && node tests/test_nan_verdict.js && node tests/test_remove_additional_nested.js && node tests/test_exclusive_bounds.js && node tests/test_engine_differential.js && node tests/test_error_shape_differential.js && node tests/test_output_format.js && node tests/test_retry_message.js && node tests/test_describe_schema.js && node tests/test_aot_parse.js && node tests/test_draft7_semantics.js && node tests/test_metaschema_ref.js && node tests/test_pure_js_unsupported.js && node tests/test_native_load_order.js && node tests/test_pack_purity.js && node tests/test_make_native_package.js && node tests/test_browser_nofs.js && node tests/test_browser_imports_guard.js && node tests/test_version_sync.js && node tests/test_native_loaded.js && node tests/test_safe_regex_source_sync.js && node tests/test_t_builder.js && node tests/test_async_refine.js && node tests/test_safe_regex.js && node tests/test_safe_regex_integration.js && node tests/test_aot_build.js && node tests/test_aot_differential.js && node tests/test_aot_cli_build.js && node tests/test_aot_cli_smoke.js && node tests/test_bundle_standalone.js && node tests/test_standalone_anyof.js && node tests/test_standalone_formats.js && node tests/test_aot_format_mode.js && node tests/test_aot_additional_props_errors.js && node tests/test_id_anchor_refs.js && node tests/test_engine_routing.js && node tests/test_engine_diagnostic.js && node tests/test_format_engine_parity.js && node tests/test_formats_single_pass.js && node tests/test_format_single_error.js && node tests/test_defs_pointer_alias.js && node tests/test_cross_doc_root_ref.js && node tests/test_codegen_entrypoint_agreement.js && node tests/test_hybrid_agreement.js && node tests/test_codegen_edge_shapes.js && node tests/test_ajv_errors.js && node tests/test_custom_keywords.js && node tests/test_aot_external_checks.js && node tests/test_ajv_parity.js && node tests/test_user_format_error_path.js && node tests/test_unevaluated_error_path.js && node tests/test_pattern_properties_errors.js && node tests/test_no_input_mutation.js && node tests/test_typed_validator_runner.js && node tests/test_define_schema.js && node tests/test_error_codes_lock.js && node tests/test_error_code_lookup.js && node tests/test_error_order.js && node tests/test_error_order_ordinal.js && node tests/test_rejection_shape.js && node tests/test_lazy_normalization.js && node tests/test_value_equality.js && node tests/test_vocabulary.js && node tests/test_schema_scan.js && node tests/test_lazy_errors.js && node tests/test_lazy_instance.js && node tests/test_cyclic_input.js && node tests/test_native_error_codes.js && node tests/test_verdict_preprocess.js && node tests/test_plan_compiler.js && node tests/test_nullable.js && node tests/test_validate_and_parse.js && node tests/test_validate_data.js && node tests/test_enrich_error.js && node tests/test_enrich_received.js && node tests/test_rich_errors_optout.js && node tests/test_error_messages.js && node tests/test_source_positions.js && node tests/fuzz_positions.js && node tests/test_data_positions.js && node tests/test_render_shared.js && node tests/test_renderers.js && node tests/test_additive_fields.js && node tests/test_diagnostic_source.js && node tests/test_diagnose.js && node tests/test_correlate.js && node tests/test_diagnostics_score.js && node tests/test_runtime_error_dx.js && node tests/test_aot_error_dx.js && node tests/test_abort_early.js && node tests/test_branch_collapse.js && node tests/test_suggestions.js && node tests/test_cli_validate.js && node tests/test_cli_version.js && node benchmark/bench_aot_size.mjs",
|
|
54
|
+
"test": "node test.js && node tests/test_removed_aot_methods.js && node tests/test_no_native.js && node tests/test_no_eval.js && node tests/test_property_dependencies.js && node tests/test_v1_dialect.js && node tests/test_buffer_path_parity.js && node tests/test_buffer_gate.js && node tests/test_buffer_reject_cost.js && node tests/test_nan_verdict.js && node tests/test_remove_additional_nested.js && node tests/test_exclusive_bounds.js && node tests/test_engine_differential.js && node tests/test_error_shape_differential.js && node tests/test_output_format.js && node tests/test_retry_message.js && node tests/test_describe_schema.js && node tests/test_aot_parse.js && node tests/test_draft7_semantics.js && node tests/test_metaschema_ref.js && node tests/test_pure_js_unsupported.js && node tests/test_native_load_order.js && node tests/test_pack_purity.js && node tests/test_make_native_package.js && node tests/test_browser_nofs.js && node tests/test_browser_imports_guard.js && node tests/test_version_sync.js && node tests/test_native_loaded.js && node tests/test_safe_regex_source_sync.js && node tests/test_t_builder.js && node tests/test_async_refine.js && node tests/test_safe_regex.js && node tests/test_safe_regex_integration.js && node tests/test_aot_build.js && node tests/test_aot_differential.js && node tests/test_aot_cli_build.js && node tests/test_aot_cli_smoke.js && node tests/test_bundle_standalone.js && node tests/test_standalone_anyof.js && node tests/test_standalone_formats.js && node tests/test_aot_format_mode.js && node tests/test_aot_additional_props_errors.js && node tests/test_id_anchor_refs.js && node tests/test_engine_routing.js && node tests/test_compile_cache_order.js && node tests/test_engine_diagnostic.js && node tests/test_format_engine_parity.js && node tests/test_uri_helper_parity.js && node tests/test_formats_single_pass.js && node tests/test_format_single_error.js && node tests/test_defs_pointer_alias.js && node tests/test_cross_doc_root_ref.js && node tests/test_codegen_entrypoint_agreement.js && node tests/test_hybrid_agreement.js && node tests/test_codegen_edge_shapes.js && node tests/test_ajv_errors.js && node tests/test_custom_keywords.js && node tests/test_aot_external_checks.js && node tests/test_ajv_parity.js && node tests/test_user_format_error_path.js && node tests/test_unevaluated_error_path.js && node tests/test_pattern_properties_errors.js && node tests/test_no_input_mutation.js && node tests/test_typed_validator_runner.js && node tests/test_define_schema.js && node tests/test_error_codes_lock.js && node tests/test_error_code_lookup.js && node tests/test_error_order.js && node tests/test_error_order_ordinal.js && node tests/test_rejection_shape.js && node tests/test_lazy_normalization.js && node tests/test_value_equality.js && node tests/test_vocabulary.js && node tests/test_schema_scan.js && node tests/test_lazy_errors.js && node tests/test_lazy_instance.js && node tests/test_cyclic_input.js && node tests/test_native_error_codes.js && node tests/test_verdict_preprocess.js && node tests/test_plan_compiler.js && node tests/test_nullable.js && node tests/test_validate_and_parse.js && node tests/test_validate_data.js && node tests/test_enrich_error.js && node tests/test_enrich_received.js && node tests/test_rich_errors_optout.js && node tests/test_error_messages.js && node tests/test_source_positions.js && node tests/fuzz_positions.js && node tests/test_data_positions.js && node tests/test_render_shared.js && node tests/test_renderers.js && node tests/test_additive_fields.js && node tests/test_diagnostic_source.js && node tests/test_diagnose.js && node tests/test_correlate.js && node tests/test_diagnostics_score.js && node tests/test_runtime_error_dx.js && node tests/test_aot_error_dx.js && node tests/test_abort_early.js && node tests/test_branch_collapse.js && node tests/test_suggestions.js && node tests/test_cli_validate.js && node tests/test_cli_version.js && node benchmark/bench_aot_size.mjs",
|
|
55
55
|
"bench:size": "node benchmark/bench_aot_size.mjs",
|
|
56
56
|
"test:suite": "node tests/run_suite.js && node tests/run_suite.js draft7 && node tests/run_suite.js v1",
|
|
57
57
|
"test:compat": "node tests/test_compat.js",
|
|
@@ -122,13 +122,13 @@
|
|
|
122
122
|
"LICENSE"
|
|
123
123
|
],
|
|
124
124
|
"optionalDependencies": {
|
|
125
|
-
"@ata-validator/native-darwin-arm64": "1.
|
|
126
|
-
"@ata-validator/native-darwin-x64": "1.
|
|
127
|
-
"@ata-validator/native-linux-arm64-gnu": "1.
|
|
128
|
-
"@ata-validator/native-linux-arm64-musl": "1.
|
|
129
|
-
"@ata-validator/native-linux-x64-gnu": "1.
|
|
130
|
-
"@ata-validator/native-linux-x64-musl": "1.
|
|
131
|
-
"@ata-validator/native-win32-x64": "1.
|
|
125
|
+
"@ata-validator/native-darwin-arm64": "1.22.0",
|
|
126
|
+
"@ata-validator/native-darwin-x64": "1.22.0",
|
|
127
|
+
"@ata-validator/native-linux-arm64-gnu": "1.22.0",
|
|
128
|
+
"@ata-validator/native-linux-arm64-musl": "1.22.0",
|
|
129
|
+
"@ata-validator/native-linux-x64-gnu": "1.22.0",
|
|
130
|
+
"@ata-validator/native-linux-x64-musl": "1.22.0",
|
|
131
|
+
"@ata-validator/native-win32-x64": "1.22.0"
|
|
132
132
|
},
|
|
133
133
|
"peerDependencies": {
|
|
134
134
|
"yaml": "^2.0.0"
|