ata-validator 1.7.4 → 1.8.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,32 @@
2
2
 
3
3
  All notable changes to ata-validator are documented here. The format follows [Keep a Changelog](https://keepachangelog.com/), and this project adheres to semantic versioning.
4
4
 
5
+ ## 1.8.1 - 2026-08-27
6
+
7
+ ### Changed
8
+
9
+ - Constructing a validator no longer clones and serializes every schema to find out whether it needed normalizing. It did that on the root and on each registered schema: serialize, deep clone, normalize the clone, serialize again, compare. For a schema with no `nullable` field and no draft-07 keyword the answer is always no, and the work was thrown away. Every question the normalizers ask is whether some key appears anywhere in the tree, so one walk answers all of them, and the clone now happens only for schemas that need it. No registry and ten fields goes from 18.5 to 4.6 µs; fifty registered schemas of fifty fields, which is a small server's worth, from 2570 to 260 µs.
10
+
11
+ The walk descends through every object-valued key rather than the subschema keywords the normalizers recurse through, so it is a superset of what they visit: it can report work where there is none, costing one clone, and cannot miss work there is. Across the 2344 schemas in the official suite it reports work on 335 where normalization changes 334. `tests/test_schema_scan.js` asserts that direction over the whole suite, and asserts the check is load-bearing by breaking the scan deliberately and confirming the broken one is caught.
12
+
13
+ The answer is remembered against the schema object, the way whole compiled validators already are, and so is the schema map built from a registry. Both were being redone once per validator, so a server building one validator per route over a shared registry paid for that registry once per route. Fifty routes over twenty shared schemas boots in 0.046 ms rather than 20.66 ms. Validators share the map, and anything that writes to one, which is `addSchema()` and the meta-schema registration during compilation, takes a private copy first.
14
+
15
+ Nothing about what any schema validates to changes. The comparison that decided the old answer is still there behind the walk, so a schema the walk sends down the slow path gets exactly the result it got before.
16
+
17
+ ## 1.8.0 - 2026-08-27
18
+
19
+ ### Added
20
+
21
+ - `$vocabulary` is honoured when a schema names a custom meta-schema in `$schema`. A dialect is a set of vocabularies and a vocabulary is a set of keywords, so a keyword whose vocabulary the meta-schema does not declare is not part of that dialect at all: it is an unknown keyword, and an unknown keyword is ignored. A schema written against a meta-schema which declares only the core and applicator vocabularies now has its `minimum`, `type` and the rest of the validation keywords ignored, while `properties` and the other applicators still apply. The keywords belonging to each vocabulary are read from the vendored meta-schemas rather than from a list kept by hand, so they come from the specification. Removing the keyword before compilation means every engine agrees without any of them knowing what a vocabulary is.
22
+
23
+ This closes the last case ata missed. Draft 2020-12 goes from 1298 to **1299 of 1299**, with draft 7 at 927 of 927 and the v1 dialect at 1133 of 1133, the same with code generation blocked. `tests/run_suite.js` has no known failures left.
24
+
25
+ Two things it does not do. It does not refuse a schema whose meta-schema requires, with `true`, a vocabulary ata does not recognise; the specification says an implementation must refuse there, and turning what has always been silently accepted into a hard failure is a separate decision, so such a schema is evaluated as before with every keyword applied. And a separate document reached through `$ref` keeps its own keywords rather than inheriting the referring dialect, so a document which should follow one needs to say so with its own `$schema`.
26
+
27
+ ### Fixed
28
+
29
+ - `bundleStandalone` and `bundleCompact` built the error-reporting function from the schema the caller passed rather than the one the validator compiled. Anything which prepares a schema therefore reached the boolean path and not the error path: with `assertFormat: false` the bundle reported a `format` error the caller had turned off, and only once some other keyword failed first, since the error function runs only after the boolean one says no. `toStandalone` already read the compiled schema; these two now do as well.
30
+
5
31
  ## 1.7.4 - 2026-08-26
6
32
 
7
33
  ### Fixed
@@ -10,7 +36,7 @@ All notable changes to ata-validator are documented here. The format follows [Ke
10
36
 
11
37
  ### Changed
12
38
 
13
- - Rejecting a document through the buffer APIs no longer costs more than accepting one. The On-Demand plan answers first, and a `false` from it used to be ambiguous: it could mean the document failed a constraint, or that the plan could not decide. The caller had to assume the second, so it re-created the padded view, parsed the whole document again into a DOM, tried the generated plan, and then walked the tree. A rejected document was therefore parsed twice and walked up to three times. The plan now reports whether it stopped because the document failed a constraint or because it could not be read, and only the second falls through. simdjson's `INCORRECT_TYPE` is a constraint failure rather than a read failure, which is the distinction that makes this work: a property holding a string where the schema wants an integer surfaces as a read error from `get<int64>`, not as a type mismatch. On a 62 KB array of a thousand objects with one bad element: 316 µs to 11 µs when the bad element is first, 315 µs to 40 µs when it is last, and rejecting is now never dearer than accepting. `tests/test_buffer_path_parity.js` holds the buffer path to the same answers as `validate()` across all 3359 suite cases, and still reports zero disagreements.
39
+ - Rejecting a document through the buffer APIs no longer costs more than accepting one. The On-Demand plan answers first, and a `false` from it used to be ambiguous: it could mean the document failed a constraint, or that the plan could not decide. The caller had to assume the second, so it re-created the padded view, parsed the whole document again into a DOM, tried the generated plan, and then walked the tree. A rejected document was therefore parsed twice and walked up to three times. The plan now reports whether it stopped because the document failed a constraint or because it could not be read, and only the second falls through. simdjson's `INCORRECT_TYPE` is a constraint failure rather than a read failure, which is the distinction that makes this work: a property holding a string where the schema wants an integer surfaces as a read error from `get<int64>`, not as a type mismatch. On a 62 KB array of a thousand objects with one bad element: 316 µs to 11 µs when the bad element is first, 315 µs to 40 µs when it is last. On that shape rejecting is no longer dearer than accepting, and when the bad element is early it is several times cheaper. It is still dearer on a small document, where there is no bulk of parsing for an early exit to save: on a two-field object, 129 ns to accept and 278 ns to reject. `tests/test_buffer_path_parity.js` holds the buffer path to the same answers as `validate()` across all 3359 suite cases, and still reports zero disagreements.
14
40
 
15
41
  ## 1.7.3 - 2026-08-26
16
42
 
package/README.md CHANGED
@@ -14,17 +14,29 @@ npm install --save-dev ata-validator
14
14
  npx ata build 'schemas/*.json' --out-dir src/generated
15
15
  ```
16
16
 
17
- The `ata-validator` package itself is pure JavaScript. The native accelerator (simdjson parsing, parallel NDJSON, buffer APIs) ships as per-platform optional packages that npm installs automatically where they fit, the same pattern Vite uses for esbuild. For a guaranteed zero-binary install:
17
+ The `ata-validator` package itself is pure JavaScript. The native accelerator (simdjson parsing, parallel NDJSON, buffer APIs) ships as per-platform optional packages that npm installs automatically where they fit, the same pattern Vite uses for esbuild. Seven targets are built: macOS on arm64 and x64, Linux on x64 and arm64 against both glibc and musl, and Windows on x64. A platform without a prebuild still installs and validates, on the pure-JS engine. For a guaranteed zero-binary install:
18
18
 
19
19
  ```bash
20
20
  npm install ata-validator --omit=optional
21
21
  ```
22
22
 
23
- or set `ATA_NO_NATIVE=1` at runtime. Typical schemas compile to specialized JS; shapes the compiler cannot represent (some `$dynamicRef`, cyclic `$ref`, unusual keyword interactions) fall back to an interpreted engine, so every schema validates in every environment. The pure-JS setup scores the same on the official suite as the native one, 1298 of 1299 Draft 2020-12 cases; the one case both miss needs `$vocabulary`, which ata does not implement. Only the buffer and parallel APIs (`isValid` on raw buffers, `countValid`, `batchIsValid`, `validateAndParse`) need the native engine and say so with a clear error.
23
+ or set `ATA_NO_NATIVE=1` at runtime. Typical schemas compile to specialized JS; shapes the compiler cannot represent (some `$dynamicRef`, cyclic `$ref`, unusual keyword interactions) fall back to an interpreted engine, so every schema validates in every environment. The pure-JS setup scores the same on the official suite as the native one, 1299 of 1299 Draft 2020-12 cases. Only the buffer and parallel APIs (`isValid` on raw buffers, `countValid`, `batchIsValid`, `validateAndParse`) need the native engine and say so with a clear error.
24
24
 
25
25
  Those four now agree with `validate()` on every case of the official suite, 3359 across three dialects. The native walker behind them does not handle every shape (`contains`, `unevaluatedProperties`, `patternProperties`, tuple `items`, cross-document `$ref`, a few formats), so for schemas using one of those the buffer APIs parse the bytes and answer through `validate()`; the list is in `lib/buffer-gate.js`. Typical request schemas stay on the zero-copy path. `npm test` holds the disagreement count at zero.
26
26
 
27
- Where `new Function` is refused altogether, on Cloudflare Workers, Deno Deploy or under a strict Content-Security-Policy, ata drops to the interpreted engine and scores the same 1298 of 1299 with code generation blocked. No flags, and on Workers no `nodejs_compat` either. See [docs/edge-runtimes.md](docs/edge-runtimes.md).
27
+ Where `new Function` is refused altogether, on Cloudflare Workers, Deno Deploy or under a strict Content-Security-Policy, ata drops to the interpreted engine and scores the same 1299 of 1299 with code generation blocked. No flags, and on Workers no `nodejs_compat` either. See [docs/edge-runtimes.md](docs/edge-runtimes.md).
28
+
29
+ Node has a switch for exactly that environment, so this takes thirty seconds to check for yourself, on ata or on whatever you use today:
30
+
31
+ ```bash
32
+ node --disallow-code-generation-from-strings -e "
33
+ const { Validator } = require('ata-validator')
34
+ const v = new Validator({ type: 'object', properties: { id: { type: 'integer' } }, required: ['id'] })
35
+ console.log(v.validate({ id: 1 }).valid, v.validate({ id: 'x' }).valid)
36
+ "
37
+ ```
38
+
39
+ A validator that reaches its speed by generating source and calling `new Function` cannot run under that flag at all, which is the same reason it cannot run on a Worker. That is a deliberate trade rather than an oversight, and it is worth knowing which side of it your validator is on before you deploy to an edge runtime.
28
40
 
29
41
  In your code:
30
42
 
@@ -428,7 +440,7 @@ const result = v['~standard'].validate(data);
428
440
 
429
441
  ### Fastify Plugin
430
442
 
431
- Measured against Fastify's own schema and validation test files with ata swapped in as the default validator: 181 of 187 tests pass, and the remaining six assert the default validator's own extension API rather than validation behavior. Validation errors follow the schema's keyword declaration order, so error-message plugins and `transformErrors` hooks written against the default validator see the same order.
443
+ Measured against Fastify's own schema and validation test files with ata swapped in as the default validator: 178 of 184 tests pass, and the remaining six assert the default validator's own extension API rather than validation behavior. Validation errors follow the schema's keyword declaration order, so error-message plugins and `transformErrors` hooks written against the default validator see the same order.
432
444
 
433
445
  ```bash
434
446
  npm install fastify-ata
@@ -507,12 +519,12 @@ Both are implemented in the interpreted engine, so a v1 schema that uses `$dynam
507
519
 
508
520
  ### Known limitations
509
521
 
510
- Running the whole Draft 2020-12 suite with nothing excluded, `format` and `default` under specification semantics (`assertFormat: false`, `useDefaults: false`), gives 1298 of 1299 cases, 99.9%. Draft 7 gives 927 of 927. The v1 dialect gives 1133 of 1133. The one 2020-12 case that fails uses a custom meta-schema that drops the validation vocabulary through `$vocabulary`, which ata does not implement, so the validation keywords still apply. `npm run test:suite` reproduces all three and names it.
522
+ Running the whole Draft 2020-12 suite with nothing excluded, `format` and `default` under specification semantics (`assertFormat: false`, `useDefaults: false`), gives 1299 of 1299 cases. Draft 7 gives 927 of 927. The v1 dialect gives 1133 of 1133. `npm run test:suite` reproduces all three.
511
523
 
512
524
  Areas that remain deliberate scope decisions for 1.x:
513
525
 
514
526
  - **Remote `$ref` over the network** is not fetched. Register cross-schema refs explicitly with `schemas: [...]` or `addSchema()`.
515
- - **`$vocabulary`** is not implemented; custom vocabularies are ignored rather than enforced. The keyword was extracted from the specification for the stable release as incomplete, so it is not part of v1.
527
+ - **`$vocabulary`** is honoured for the document which names the meta-schema: a keyword whose vocabulary that meta-schema does not declare is not part of the dialect, so it is treated as unknown and does not apply. Two things it does not do. It does not refuse a schema whose meta-schema requires a vocabulary ata does not recognise, which the specification says an implementation must do; that schema is evaluated with every keyword applied, as before. And a separate document reached through `$ref` keeps its own keywords rather than inheriting the referring dialect, so register it with its own `$schema` if it should follow one.
516
528
  - **`contentEncoding` / `contentMediaType` / `contentSchema`** are annotation-only, as the spec permits, and are not validated.
517
529
  - **`minLength`/`maxLength`** count UTF-16 code units, not grapheme clusters.
518
530
  - **Infinite-loop detection** relies on a recursion depth guard that cuts off circular `$ref` chains.
package/index.js CHANGED
@@ -10,6 +10,8 @@ const {
10
10
  compileToJSCombined,
11
11
  } = require("./lib/js-compiler");
12
12
  const { normalizeDraft7, normalizeNullable, stripFormatAssertions } = require("./lib/draft7");
13
+ const { enabledKeywords, stripDisabledKeywords } = require("./lib/vocabularies");
14
+ const { needsNormalization } = require("./lib/schema-scan");
13
15
  const { isV1Dialect } = require("./lib/dialect");
14
16
  const { classify } = require("./lib/shape-classifier");
15
17
  const { buildTier0Plan, tier0Validate } = require("./lib/tier0");
@@ -459,11 +461,20 @@ function _normalizeCallerSchema(s, inheritDraft7) {
459
461
  const needsDraft7 = declares
460
462
  ? (s.$schema === 'http://json-schema.org/draft-07/schema#' || s.$schema === 'http://json-schema.org/draft-07/schema')
461
463
  : !!inheritDraft7
464
+ // One walk answers whether there is anything to do. Almost always there is
465
+ // not, and then the serialize, clone, normalize, serialize, compare below is
466
+ // work spent to find that out. The walk over-reports rather than under, so a
467
+ // schema it clears is one no normalizer would have touched;
468
+ // `tests/test_schema_scan.js` holds that direction against the whole suite.
469
+ if (!needsNormalization(s, needsDraft7)) return s
470
+
462
471
  const str = JSON.stringify(s)
463
472
  const copy = _deepCloneWithSymbols(s)
464
473
  if (needsDraft7) normalizeDraft7(copy, true)
465
474
  normalizeNullable(copy)
466
475
  // Return original when normalization produced no change, copy otherwise.
476
+ // Kept even though the walk has already said there is work, so that a walk
477
+ // which over-reports still returns exactly what it returned before.
467
478
  // Change-detection uses JSON content only; symbols do not affect it.
468
479
  return JSON.stringify(copy) === str ? s : copy
469
480
  }
@@ -487,8 +498,33 @@ function declaredId(original, normalized) {
487
498
 
488
499
  // `inheritDraft7` is true when the root schema is draft-07: a retrieved
489
500
  // document that declares no dialect is read under the root's draft.
501
+ // The map is derived entirely from what the caller passed, so the same
502
+ // `schemas` gives the same map. A server building one validator per route over
503
+ // a shared registry rebuilt it once per route, normalizing and re-reading the
504
+ // `$id` of every registered schema each time. Keyed by the registry object,
505
+ // and by the draft it is read under, since that changes what normalization
506
+ // does to a document which declares no dialect of its own.
507
+ //
508
+ // Validators share the returned map, so anything that mutates one calls
509
+ // `_ownSchemaMap()` first. There are two such places: registering the vendored
510
+ // meta-schemas during compilation, and `addSchema()`.
511
+ const _schemaMapCache = new WeakMap()
512
+
490
513
  function buildSchemaMap(schemas, inheritDraft7) {
491
514
  if (!schemas) return null
515
+ const byDraft = _schemaMapCache.get(schemas)
516
+ if (byDraft) {
517
+ const hit = byDraft[inheritDraft7 ? 1 : 0]
518
+ if (hit) return hit
519
+ }
520
+ const map = _buildSchemaMap(schemas, inheritDraft7)
521
+ const slot = byDraft || [null, null]
522
+ slot[inheritDraft7 ? 1 : 0] = map
523
+ if (!byDraft) _schemaMapCache.set(schemas, slot)
524
+ return map
525
+ }
526
+
527
+ function _buildSchemaMap(schemas, inheritDraft7) {
492
528
  const map = new Map()
493
529
  if (Array.isArray(schemas)) {
494
530
  for (const s of schemas) {
@@ -511,6 +547,29 @@ function buildSchemaMap(schemas, inheritDraft7) {
511
547
  return map
512
548
  }
513
549
 
550
+ // A schema which names a custom meta-schema in `$schema` is written against
551
+ // whatever dialect that meta-schema declares. A keyword from a vocabulary the
552
+ // dialect does not have is not part of the dialect, so it is an unknown
553
+ // keyword and does not apply. Removing it here means every engine sees the
554
+ // same schema and none of them needs to know about vocabularies.
555
+ //
556
+ // Only the root is consulted. A subschema naming its own `$schema` is its own
557
+ // resource under its own dialect, and the walk stops there rather than
558
+ // applying this dialect's answer to it.
559
+ function _applyVocabularies(schemaObj, original, schemaMap) {
560
+ if (!schemaObj || typeof schemaObj !== 'object') return schemaObj
561
+ const declared = schemaObj.$schema
562
+ if (typeof declared !== 'string') return schemaObj
563
+ const enabled = enabledKeywords(schemaMap.get(declared))
564
+ if (!enabled) return schemaObj
565
+ // `original` is the caller's own object when it reached here unchanged, and
566
+ // that one is never mutated.
567
+ const copy = schemaObj === original
568
+ ? _deepCloneWithSymbols(schemaObj)
569
+ : schemaObj
570
+ return stripDisabledKeywords(copy, enabled)
571
+ }
572
+
514
573
  // Compile-cache key for a root schema plus its external schemas. Must include
515
574
  // the external schema CONTENT, not just their $ids: two validators can share a
516
575
  // root schema string and the same $id while pointing that $id at different
@@ -621,6 +680,15 @@ class Validator {
621
680
  );
622
681
  }
623
682
 
683
+ // Built here rather than below because `$vocabulary` is resolved against
684
+ // it, and that resolution waits until compilation so a meta-schema
685
+ // registered by addSchema() still counts.
686
+ const shared = buildSchemaMap(options.schemas, rootIsDraft7);
687
+ const schemaMap = shared || new Map();
688
+ this._schemaMapShared = shared !== null;
689
+ this._schemaIsCallers = schemaObj === schema;
690
+ this._vocabulariesApplied = false;
691
+
624
692
  this._schemaStr = null; // lazy: computed on first use
625
693
  this._schemaObj = schemaObj;
626
694
  this._options = options;
@@ -634,7 +702,7 @@ class Validator {
634
702
  this._applyDefaults = null;
635
703
 
636
704
  // Schema map for cross-schema $ref resolution
637
- this._schemaMap = buildSchemaMap(options.schemas, rootIsDraft7) || new Map();
705
+ this._schemaMap = schemaMap;
638
706
 
639
707
  // User-supplied format checkers: { formatName: (value) => boolean }.
640
708
  // Looked up at runtime when a schema references a format the built-in
@@ -758,8 +826,27 @@ class Validator {
758
826
  }
759
827
  }
760
828
 
829
+ // `$vocabulary` says which keywords the dialect has, and answering needs the
830
+ // meta-schema, which addSchema() may only have registered just now. Run once,
831
+ // before anything reads the schema, and before `_schemaStr` is computed from
832
+ // it. After this addSchema() is refused, so the answer cannot go stale.
833
+ _ensureVocabularies() {
834
+ if (this._vocabulariesApplied) return;
835
+ this._vocabulariesApplied = true;
836
+ const stripped = _applyVocabularies(
837
+ this._schemaObj,
838
+ this._schemaIsCallers ? this._schemaObj : null,
839
+ this._schemaMap,
840
+ );
841
+ if (stripped !== this._schemaObj) {
842
+ this._schemaObj = stripped;
843
+ this._schemaStr = null;
844
+ }
845
+ }
846
+
761
847
  _ensureCompiled() {
762
848
  if (this._initialized) return;
849
+ this._ensureVocabularies();
763
850
  this._initialized = true;
764
851
 
765
852
  const schemaObj = this._schemaObj;
@@ -774,6 +861,7 @@ class Validator {
774
861
  // reference pay for the lookup.
775
862
  if (this._schemaStr.includes('json-schema.org/draft')) {
776
863
  const { METASCHEMAS } = require('./lib/metaschemas');
864
+ this._ownSchemaMap();
777
865
  for (const [id, meta] of METASCHEMAS) {
778
866
  const bare = id.replace(/#$/, '');
779
867
  for (const key of [id, bare, bare + '#', bare.replace(/^https:/, 'http:'), bare.replace(/^http:/, 'https:')]) {
@@ -1482,11 +1570,21 @@ class Validator {
1482
1570
  const rootIsDraft7 = !!(root && typeof root === 'object' && typeof root.$schema === 'string' &&
1483
1571
  (root.$schema === 'http://json-schema.org/draft-07/schema#' || root.$schema === 'http://json-schema.org/draft-07/schema'))
1484
1572
  const normalized = _normalizeCallerSchema(schema, rootIsDraft7)
1573
+ this._ownSchemaMap()
1485
1574
  this._schemaMap.set(normalized.$id, normalized)
1486
1575
  }
1487
1576
 
1577
+ // buildSchemaMap hands the same map to every validator built from the same
1578
+ // registry. Take a private copy before writing to it.
1579
+ _ownSchemaMap() {
1580
+ if (!this._schemaMapShared) return
1581
+ this._schemaMap = new Map(this._schemaMap)
1582
+ this._schemaMapShared = false
1583
+ }
1584
+
1488
1585
  _ensureCodegen() {
1489
1586
  if (this._jsFn) return;
1587
+ this._ensureVocabularies();
1490
1588
  if (typeof process !== 'undefined' && process.env && process.env.ATA_FORCE_NAPI) return;
1491
1589
  if (!this._schemaStr) this._schemaStr = JSON.stringify(this._schemaObj);
1492
1590
  const sm = this._schemaMap.size > 0 ? this._schemaMap : null;
package/lib/aot.js CHANGED
@@ -285,8 +285,13 @@ function bundleStandalone(Validator, schemas, opts) {
285
285
  v._ensureCompiled();
286
286
  const jsFn = v._jsFn;
287
287
  if (!jsFn || !jsFn._hybridSource) return 'null';
288
+ // The schema the validator compiled, not the one the caller passed. They
289
+ // differ whenever anything prepared the schema: draft-07 normalization,
290
+ // `format` removed under assertFormat: false, keywords outside the
291
+ // dialect's vocabularies. Reading the original here would leave the
292
+ // boolean path and the error path disagreeing about the same document.
288
293
  const jsErrFn = compileToJSCodegenWithErrors(
289
- typeof schema === 'string' ? JSON.parse(schema) : schema,
294
+ v._schemaObj,
290
295
  v._schemaMap,
291
296
  v._userFormats,
292
297
  );
@@ -311,7 +316,7 @@ function bundleStandalone(Validator, schemas, opts) {
311
316
  }
312
317
  if (opts && opts.verbose) {
313
318
  // Embed the schema and a small resolver so errors carry parentSchema.
314
- const schemaLit = JSON.stringify(typeof schema === 'string' ? JSON.parse(schema) : schema);
319
+ const schemaLit = JSON.stringify(v._schemaObj);
315
320
  return `(function(R){${preamble}var _S=${schemaLit};function _PS(p){if(!p||p[0]!=='#')return undefined;var s=p.slice(1);if(!s)return _S;var ps=s.split('/').filter(Boolean).map(function(x){return x.replace(/~1/g,'/').replace(/~0/g,'~')});var t=_S;for(var i=0;i<ps.length-1;i++){if(t==null||typeof t!=='object')return undefined;t=t[ps[i]]}return t}var E=function(d){var _all=true;${errBody}};var _v=function(d){${jsFn._hybridSource}};return function(d){var r=_v(d);if(r&&r.valid===false&&r.errors){var es=[];for(var i=0;i<r.errors.length;i++){var e=r.errors[i];es.push(Object.assign({},e,{parentSchema:_PS(e.schemaPath)}))}return{valid:false,errors:es}}return r}})(R)`;
316
321
  }
317
322
  return `(function(R){${preamble}var E=function(d){var _all=true;${errBody}};return function(d){${jsFn._hybridSource}}})(R)`;
@@ -343,8 +348,13 @@ function bundleCompact(Validator, schemas, opts) {
343
348
  v._ensureCompiled();
344
349
  const jsFn = v._jsFn;
345
350
  if (!jsFn || !jsFn._hybridSource) return null;
351
+ // The schema the validator compiled, not the one the caller passed. They
352
+ // differ whenever anything prepared the schema: draft-07 normalization,
353
+ // `format` removed under assertFormat: false, keywords outside the
354
+ // dialect's vocabularies. Reading the original here would leave the
355
+ // boolean path and the error path disagreeing about the same document.
346
356
  const jsErrFn = compileToJSCodegenWithErrors(
347
- typeof schema === 'string' ? JSON.parse(schema) : schema,
357
+ v._schemaObj,
348
358
  v._schemaMap,
349
359
  v._userFormats,
350
360
  );
@@ -0,0 +1,115 @@
1
+ 'use strict';
2
+
3
+ // One walk of a schema, answering in a single integer the questions that are
4
+ // currently answered by doing work and looking at the result.
5
+ //
6
+ // `_normalizeCallerSchema` decides whether a schema needs normalizing by
7
+ // serializing it, cloning it, normalizing the clone and serializing again to
8
+ // compare. Two serializations and a deep copy, on every schema, to find out
9
+ // that a modern schema needs nothing. The questions the normalizers actually
10
+ // ask are all of the form "does this key appear anywhere in the tree", and one
11
+ // traversal answers all of them at once. Daniel Lemire's rule from the C++
12
+ // side, a layer up: do not do the work to find out whether the work is needed.
13
+ //
14
+ // The walk deliberately descends through **every** object-valued key, not the
15
+ // list of subschema keywords the normalizers recurse through. That makes it a
16
+ // superset of what they visit, so it can report work where there is none, but
17
+ // never miss work there is. A false positive costs one clone. A false negative
18
+ // would hand an un-normalized schema to the engines, which is the silent
19
+ // acceptance this codebase treats as its worst failure. `tests/test_schema_scan.js`
20
+ // pins the direction of that error against the whole suite.
21
+
22
+ // One bit per question. Kept small and explicit; a schema which trips none of
23
+ // them scans to 0 and can skip normalization entirely.
24
+ const NULLABLE = 1 << 0; // `nullable`, which normalizeNullable folds into `type`
25
+ const REF_SIBLINGS = 1 << 1; // a draft-07 `$ref` carrying keywords which are ignored
26
+ const ANCHOR_ID = 1 << 2; // a fragment-only `$id`, which is an anchor in draft-07
27
+ const DEFINITIONS = 1 << 3; // `definitions`, renamed to `$defs`
28
+ const DEPENDENCIES = 1 << 4; // `dependencies`, split into two keywords
29
+ const TUPLE_ITEMS = 1 << 5; // array-valued `items`, split into `prefixItems`
30
+
31
+ // Everything only draft-07 normalization acts on. A schema of another dialect
32
+ // can carry these without them meaning anything, so they are read together
33
+ // with whether the document is draft-07 at all.
34
+ const DRAFT7_WORK = REF_SIBLINGS | ANCHOR_ID | DEFINITIONS | DEPENDENCIES | TUPLE_ITEMS;
35
+
36
+ // Copied from the normalizer rather than shared, so that a change there which
37
+ // forgets this file shows up as a differential test failure rather than as a
38
+ // silently wider scan. The test asserts the two agree.
39
+ const REF_SIBLINGS_KEPT = new Set([
40
+ '$ref', '$defs', 'definitions', '$schema', '$comment',
41
+ 'title', 'description', 'examples', 'default', 'readOnly', 'writeOnly',
42
+ ]);
43
+ const ANCHOR_ID_RE = /^#[A-Za-z][A-Za-z0-9_.:-]*$/;
44
+
45
+ // The answer depends only on the object, so it is remembered against it. A
46
+ // server building one validator per route over a shared registry hands the
47
+ // same registry objects to every one of them: fifty routes over twenty shared
48
+ // schemas scanned those twenty a thousand times. Identity caching is what
49
+ // `_identityCache` already does for whole compiled validators.
50
+ //
51
+ // This assumes a schema is not mutated after being handed to a Validator,
52
+ // which is already true of everything else here: the compile cache, the
53
+ // identity cache and the schema map would all be stale too.
54
+ const CACHE = new WeakMap();
55
+
56
+ function scan(schema) {
57
+ if (typeof schema !== 'object' || schema === null) return 0;
58
+ const hit = CACHE.get(schema);
59
+ if (hit !== undefined) return hit;
60
+
61
+ let bits = 0;
62
+ const seen = new Set();
63
+
64
+ const walk = (node) => {
65
+ if (typeof node !== 'object' || node === null) return;
66
+ if (Array.isArray(node)) {
67
+ for (let i = 0; i < node.length; i++) walk(node[i]);
68
+ return;
69
+ }
70
+ if (seen.has(node)) return;
71
+ seen.add(node);
72
+
73
+ if ('nullable' in node) bits |= NULLABLE;
74
+ if (node.definitions !== undefined && node.$defs === undefined) bits |= DEFINITIONS;
75
+ if (node.dependencies !== undefined) bits |= DEPENDENCIES;
76
+ if (Array.isArray(node.items)) bits |= TUPLE_ITEMS;
77
+ if (typeof node.$id === 'string' && ANCHOR_ID_RE.test(node.$id)) bits |= ANCHOR_ID;
78
+
79
+ // `Object.keys` rather than `for...in`: the prototype chain has nothing
80
+ // to contribute here, and walking it is the slower path in V8.
81
+ const keys = Object.keys(node);
82
+ if (typeof node.$ref === 'string') {
83
+ for (let i = 0; i < keys.length; i++) {
84
+ if (!REF_SIBLINGS_KEPT.has(keys[i])) { bits |= REF_SIBLINGS; break; }
85
+ }
86
+ }
87
+ // Every key, not just the subschema keywords, so this cannot miss a place
88
+ // the normalizers would reach.
89
+ for (let i = 0; i < keys.length; i++) walk(node[keys[i]]);
90
+ };
91
+
92
+ walk(schema);
93
+ CACHE.set(schema, bits);
94
+ return bits;
95
+ }
96
+
97
+ // Would normalization change this schema? `isDraft7` says whether the draft-07
98
+ // rules apply at all, which the caller already knows.
99
+ function needsNormalization(schema, isDraft7) {
100
+ const bits = scan(schema);
101
+ if (bits & NULLABLE) return true;
102
+ return isDraft7 ? (bits & DRAFT7_WORK) !== 0 : false;
103
+ }
104
+
105
+ module.exports = {
106
+ scan,
107
+ needsNormalization,
108
+ NULLABLE,
109
+ REF_SIBLINGS,
110
+ ANCHOR_ID,
111
+ DEFINITIONS,
112
+ DEPENDENCIES,
113
+ TUPLE_ITEMS,
114
+ DRAFT7_WORK,
115
+ };
package/lib/version.js CHANGED
@@ -7,4 +7,4 @@
7
7
  //
8
8
  // Kept in lockstep with package.json by `tests/test_version_sync.js`.
9
9
 
10
- module.exports = '1.7.4';
10
+ module.exports = '1.8.1';
@@ -0,0 +1,199 @@
1
+ 'use strict';
2
+
3
+ // `$vocabulary`, as far as evaluation is concerned.
4
+ //
5
+ // A dialect is a set of vocabularies, and a vocabulary is a set of keywords.
6
+ // When a schema names a custom meta-schema in `$schema`, that meta-schema's
7
+ // `$vocabulary` says which vocabularies the dialect has. A keyword whose
8
+ // vocabulary is not there is not part of the dialect at all, so it is an
9
+ // unknown keyword, and an unknown keyword is ignored. Deleting it is exactly
10
+ // that reading, and it means every engine gets this for free rather than each
11
+ // of them growing a notion of vocabularies.
12
+ //
13
+ // The keyword lists are not written out here. Each official vocabulary has a
14
+ // meta-schema whose `properties` are precisely that vocabulary's keywords,
15
+ // and those documents are already vendored in `metaschemas.js`, so the lists
16
+ // come from the specification rather than from someone keeping a copy in
17
+ // step with it.
18
+ //
19
+ // What this deliberately does not do: refuse a schema whose meta-schema
20
+ // requires a vocabulary ata does not know. The specification says an
21
+ // implementation must refuse there, and ata does not, because until now it
22
+ // ignored `$vocabulary` entirely and turning that into a hard failure would
23
+ // break schemas which work today. Such a schema is left exactly as it was,
24
+ // evaluated with every keyword applied. See `docs/` for the limitation.
25
+
26
+ const { METASCHEMAS } = require('./metaschemas')
27
+
28
+ const CORE_VOCABULARY = 'https://json-schema.org/draft/2020-12/vocab/core'
29
+
30
+ // vocabulary URI -> Set of the keywords it defines, built once on first use.
31
+ let KEYWORDS_BY_VOCABULARY = null
32
+
33
+ function keywordsByVocabulary() {
34
+ if (KEYWORDS_BY_VOCABULARY) return KEYWORDS_BY_VOCABULARY
35
+ const byVocabulary = new Map()
36
+ for (const name of [
37
+ 'core',
38
+ 'applicator',
39
+ 'unevaluated',
40
+ 'validation',
41
+ 'meta-data',
42
+ 'format-annotation',
43
+ 'content',
44
+ ]) {
45
+ const meta = METASCHEMAS.get(
46
+ `https://json-schema.org/draft/2020-12/meta/${name}`,
47
+ )
48
+ if (!meta || !meta.properties) continue
49
+ byVocabulary.set(
50
+ `https://json-schema.org/draft/2020-12/vocab/${name}`,
51
+ new Set(Object.keys(meta.properties)),
52
+ )
53
+ }
54
+ // The one vocabulary with no document to read. `format-assertion` is an
55
+ // alternative to `format-annotation` rather than part of the standard
56
+ // dialect, so its meta-schema is not among the vendored ones. It defines
57
+ // the single keyword `format`, which is what the other one defines, and a
58
+ // vocabulary's keyword set does not change once it is published. Without
59
+ // this a meta-schema asking for format assertion, which is the standard
60
+ // way to ask, would be a vocabulary ata does not know and no filtering
61
+ // would happen at all.
62
+ byVocabulary.set(
63
+ 'https://json-schema.org/draft/2020-12/vocab/format-assertion',
64
+ new Set(['format']),
65
+ )
66
+
67
+ KEYWORDS_BY_VOCABULARY = byVocabulary
68
+ return byVocabulary
69
+ }
70
+
71
+ // Every keyword ata can attribute to a vocabulary. A keyword outside this set
72
+ // belongs to no vocabulary ata knows, so `$vocabulary` says nothing about it
73
+ // and it is left alone.
74
+ function allVocabularyKeywords() {
75
+ const all = new Set()
76
+ for (const keywords of keywordsByVocabulary().values()) {
77
+ for (const keyword of keywords) all.add(keyword)
78
+ }
79
+ return all
80
+ }
81
+
82
+ // Given a meta-schema, the keywords a schema written against it may use, or
83
+ // `null` when the question does not arise or cannot be answered:
84
+ //
85
+ // - the document declares no `$vocabulary`, so it does not describe a
86
+ // dialect in these terms and every keyword stands
87
+ // - it requires a vocabulary ata does not know, so ata cannot say which
88
+ // keywords the dialect has and does not guess
89
+ function enabledKeywords(metaschema) {
90
+ if (!metaschema || typeof metaschema !== 'object') return null
91
+ const vocabularies = metaschema.$vocabulary
92
+ if (!vocabularies || typeof vocabularies !== 'object') return null
93
+
94
+ const known = keywordsByVocabulary()
95
+ const enabled = new Set()
96
+ for (const [uri, required] of Object.entries(vocabularies)) {
97
+ const keywords = known.get(uri)
98
+ if (keywords) {
99
+ for (const keyword of keywords) enabled.add(keyword)
100
+ } else if (required === true) {
101
+ return null // a required vocabulary ata cannot account for
102
+ }
103
+ }
104
+
105
+ // Core is what makes a document a schema at all: `$ref`, `$defs`, `$id`.
106
+ // A meta-schema which omits it is not describing something ata could
107
+ // evaluate, so treat core as present rather than stripping the plumbing.
108
+ for (const keyword of known.get(CORE_VOCABULARY) || []) enabled.add(keyword)
109
+
110
+ // The usual case is a meta-schema which has every vocabulary, and then
111
+ // there is nothing to answer. Saying so here keeps the caller from copying
112
+ // a schema it would not have changed.
113
+ for (const keyword of allVocabularyKeywords()) {
114
+ if (!enabled.has(keyword)) return enabled
115
+ }
116
+ return null
117
+ }
118
+
119
+ // Delete every keyword which belongs to a vocabulary the dialect does not
120
+ // have. Mutates, so callers pass a copy.
121
+ function stripDisabledKeywords(schema, enabled) {
122
+ const removable = allVocabularyKeywords()
123
+ const disabled = new Set()
124
+ for (const keyword of removable) {
125
+ if (!enabled.has(keyword)) disabled.add(keyword)
126
+ }
127
+ if (disabled.size === 0) return schema
128
+ _strip(schema, disabled, new Set(), true)
129
+ return schema
130
+ }
131
+
132
+ function _strip(schema, disabled, seen, isRoot) {
133
+ if (typeof schema !== 'object' || schema === null) return
134
+ if (Array.isArray(schema)) {
135
+ for (const each of schema) _strip(each, disabled, seen, false)
136
+ return
137
+ }
138
+ if (seen.has(schema)) return
139
+ seen.add(schema)
140
+
141
+ // A subschema which names its own `$schema` is its own resource under its
142
+ // own dialect, so this dialect's vocabularies say nothing about it. The
143
+ // root names one by definition, which is how this dialect was chosen.
144
+ if (!isRoot && typeof schema.$schema === 'string') return
145
+
146
+ for (const keyword of disabled) {
147
+ if (keyword in schema) delete schema[keyword]
148
+ }
149
+
150
+ // Walk what is left. A disabled applicator has already been deleted, so
151
+ // this only descends through subschemas which still apply.
152
+ const objSubs = [
153
+ 'properties',
154
+ 'patternProperties',
155
+ '$defs',
156
+ 'definitions',
157
+ 'dependentSchemas',
158
+ ]
159
+ for (const key of objSubs) {
160
+ const box = schema[key]
161
+ if (box && typeof box === 'object' && !Array.isArray(box)) {
162
+ for (const each of Object.values(box)) _strip(each, disabled, seen, false)
163
+ }
164
+ }
165
+ const arrSubs = ['allOf', 'anyOf', 'oneOf', 'prefixItems']
166
+ for (const key of arrSubs) {
167
+ if (Array.isArray(schema[key])) {
168
+ for (const each of schema[key]) _strip(each, disabled, seen, false)
169
+ }
170
+ }
171
+ const singleSubs = [
172
+ 'items',
173
+ 'additionalItems',
174
+ 'contains',
175
+ 'not',
176
+ 'if',
177
+ 'then',
178
+ 'else',
179
+ 'additionalProperties',
180
+ 'propertyNames',
181
+ 'unevaluatedItems',
182
+ 'unevaluatedProperties',
183
+ 'contentSchema',
184
+ ]
185
+ for (const key of singleSubs) {
186
+ const sub = schema[key]
187
+ if (Array.isArray(sub)) {
188
+ for (const each of sub) _strip(each, disabled, seen, false)
189
+ } else {
190
+ _strip(sub, disabled, seen, false)
191
+ }
192
+ }
193
+ }
194
+
195
+ module.exports = {
196
+ enabledKeywords,
197
+ stripDisabledKeywords,
198
+ keywordsByVocabulary,
199
+ }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ata-validator",
3
- "version": "1.7.4",
3
+ "version": "1.8.1",
4
4
  "description": "JSON Schema validation with first-class TypeScript and zero runtime cost. AOT compile to per-schema ESM modules with zero validator dependency. Generic Validator<T> for TypeBox/Zod/Valibot composition. Optional runtime API. Standard Schema V1 compatible.",
5
5
  "main": "index.js",
6
6
  "module": "index.mjs",
@@ -44,7 +44,7 @@
44
44
  "release:check": "node scripts/regen-safe-regex-source.js && node tests/test_pack_purity.js && node scripts/check-doc-coverage.js && node tests/test_error_codes_lock.js && node tests/test_safe_regex_source_sync.js && node tests/test_version_sync.js",
45
45
  "build": "cmake-js build --target ata",
46
46
  "rebuild": "cmake-js rebuild --target ata",
47
- "test": "node test.js && node tests/test_removed_aot_methods.js && node tests/test_no_native.js && node tests/test_no_eval.js && node tests/test_property_dependencies.js && node tests/test_v1_dialect.js && node tests/test_buffer_path_parity.js && node tests/test_buffer_gate.js && node tests/test_draft7_semantics.js && node tests/test_metaschema_ref.js && node tests/test_pure_js_unsupported.js && node tests/test_native_load_order.js && node tests/test_pack_purity.js && node tests/test_make_native_package.js && node tests/test_browser_nofs.js && node tests/test_browser_imports_guard.js && node tests/test_version_sync.js && node tests/test_native_loaded.js && node tests/test_safe_regex_source_sync.js && node tests/test_t_builder.js && node tests/test_async_refine.js && node tests/test_safe_regex.js && node tests/test_safe_regex_integration.js && node tests/test_aot_build.js && node tests/test_aot_differential.js && node tests/test_aot_cli_build.js && node tests/test_aot_cli_smoke.js && node tests/test_bundle_standalone.js && node tests/test_standalone_anyof.js && node tests/test_standalone_formats.js && node tests/test_aot_format_mode.js && node tests/test_aot_additional_props_errors.js && node tests/test_id_anchor_refs.js && node tests/test_engine_routing.js && node tests/test_engine_diagnostic.js && node tests/test_format_engine_parity.js && node tests/test_defs_pointer_alias.js && node tests/test_cross_doc_root_ref.js && node tests/test_codegen_entrypoint_agreement.js && node tests/test_pattern_properties_errors.js && node tests/test_no_input_mutation.js && node tests/test_typed_validator_runner.js && node tests/test_define_schema.js && node tests/test_error_codes_lock.js && node tests/test_error_code_lookup.js && node tests/test_error_order.js && node tests/test_value_equality.js && node tests/test_lazy_errors.js && node tests/test_plan_compiler.js && node tests/test_nullable.js && node tests/test_validate_and_parse.js && node tests/test_validate_data.js && node tests/test_enrich_error.js && node tests/test_enrich_received.js && node tests/test_rich_errors_optout.js && node tests/test_error_messages.js && node tests/test_source_positions.js && node tests/fuzz_positions.js && node tests/test_data_positions.js && node tests/test_render_shared.js && node tests/test_renderers.js && node tests/test_runtime_error_dx.js && node tests/test_aot_error_dx.js && node tests/test_abort_early.js && node tests/test_branch_collapse.js && node tests/test_suggestions.js && node tests/test_cli_validate.js && node tests/test_cli_version.js && node benchmark/bench_aot_size.mjs",
47
+ "test": "node test.js && node tests/test_removed_aot_methods.js && node tests/test_no_native.js && node tests/test_no_eval.js && node tests/test_property_dependencies.js && node tests/test_v1_dialect.js && node tests/test_buffer_path_parity.js && node tests/test_buffer_gate.js && node tests/test_draft7_semantics.js && node tests/test_metaschema_ref.js && node tests/test_pure_js_unsupported.js && node tests/test_native_load_order.js && node tests/test_pack_purity.js && node tests/test_make_native_package.js && node tests/test_browser_nofs.js && node tests/test_browser_imports_guard.js && node tests/test_version_sync.js && node tests/test_native_loaded.js && node tests/test_safe_regex_source_sync.js && node tests/test_t_builder.js && node tests/test_async_refine.js && node tests/test_safe_regex.js && node tests/test_safe_regex_integration.js && node tests/test_aot_build.js && node tests/test_aot_differential.js && node tests/test_aot_cli_build.js && node tests/test_aot_cli_smoke.js && node tests/test_bundle_standalone.js && node tests/test_standalone_anyof.js && node tests/test_standalone_formats.js && node tests/test_aot_format_mode.js && node tests/test_aot_additional_props_errors.js && node tests/test_id_anchor_refs.js && node tests/test_engine_routing.js && node tests/test_engine_diagnostic.js && node tests/test_format_engine_parity.js && node tests/test_defs_pointer_alias.js && node tests/test_cross_doc_root_ref.js && node tests/test_codegen_entrypoint_agreement.js && node tests/test_pattern_properties_errors.js && node tests/test_no_input_mutation.js && node tests/test_typed_validator_runner.js && node tests/test_define_schema.js && node tests/test_error_codes_lock.js && node tests/test_error_code_lookup.js && node tests/test_error_order.js && node tests/test_value_equality.js && node tests/test_vocabulary.js && node tests/test_schema_scan.js && node tests/test_lazy_errors.js && node tests/test_plan_compiler.js && node tests/test_nullable.js && node tests/test_validate_and_parse.js && node tests/test_validate_data.js && node tests/test_enrich_error.js && node tests/test_enrich_received.js && node tests/test_rich_errors_optout.js && node tests/test_error_messages.js && node tests/test_source_positions.js && node tests/fuzz_positions.js && node tests/test_data_positions.js && node tests/test_render_shared.js && node tests/test_renderers.js && node tests/test_runtime_error_dx.js && node tests/test_aot_error_dx.js && node tests/test_abort_early.js && node tests/test_branch_collapse.js && node tests/test_suggestions.js && node tests/test_cli_validate.js && node tests/test_cli_version.js && node benchmark/bench_aot_size.mjs",
48
48
  "bench:size": "node benchmark/bench_aot_size.mjs",
49
49
  "test:suite": "node tests/run_suite.js && node tests/run_suite.js draft7 && node tests/run_suite.js v1",
50
50
  "test:compat": "node tests/test_compat.js",
@@ -112,13 +112,13 @@
112
112
  "LICENSE"
113
113
  ],
114
114
  "optionalDependencies": {
115
- "@ata-validator/native-darwin-arm64": "1.7.4",
116
- "@ata-validator/native-darwin-x64": "1.7.4",
117
- "@ata-validator/native-linux-x64-gnu": "1.7.4",
118
- "@ata-validator/native-linux-arm64-gnu": "1.7.4",
119
- "@ata-validator/native-linux-x64-musl": "1.7.4",
120
- "@ata-validator/native-linux-arm64-musl": "1.7.4",
121
- "@ata-validator/native-win32-x64": "1.7.4"
115
+ "@ata-validator/native-darwin-arm64": "1.8.1",
116
+ "@ata-validator/native-darwin-x64": "1.8.1",
117
+ "@ata-validator/native-linux-x64-gnu": "1.8.1",
118
+ "@ata-validator/native-linux-arm64-gnu": "1.8.1",
119
+ "@ata-validator/native-linux-x64-musl": "1.8.1",
120
+ "@ata-validator/native-linux-arm64-musl": "1.8.1",
121
+ "@ata-validator/native-win32-x64": "1.8.1"
122
122
  },
123
123
  "peerDependencies": {
124
124
  "yaml": "^2.0.0"