ochre-sdk 1.0.74 → 1.0.75

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -145,13 +145,13 @@ const queries: Query = {
145
145
  };
146
146
  ```
147
147
 
148
- Every `<string>` node in that layer holds a single OCR word, so `includes` matches each search term as its own word anywhere in the layer, in any order, with `*` and `?` wildcards supported. `exact` instead matches the terms as a run of adjacent whole words, so `"THE COLLEGE"` matches a page carrying that phrase but not one where the two words merely appear apart. An item matches when the OCR layer of the item itself or of any of its child Resources matches.
148
+ Every word node in that layer holds a single OCR word in its `CONTENT` attribute, so `includes` matches each search term as its own word anywhere in the layer, in any order, with `*` and `?` wildcards supported. `exact` instead matches the terms as a run of adjacent whole words, so `"THE COLLEGE"` matches a page carrying that phrase but not one where the two words merely appear apart. An item matches when the OCR layer of the item itself or of any of its child Resources matches.
149
149
 
150
150
  Set item projections do not carry the OCR layer, so an `ocr` leaf is resolved by an extra index-only search over the Resource documents whose matching UUIDs are then joined back onto the Set items. It still composes with `and`, `or`, and `isNegated` like any other leaf, and repeating the same OCR search inside one tree only costs one search.
151
151
 
152
152
  ## OCR Data
153
153
 
154
- A Resource may carry an `<ocr>` layer holding the positioned output of an OCR run. The node hierarchy inside that layer is irregular and is not parsed, but any `<string>` node within it, at any depth, is read as one positioned OCR string.
154
+ A Resource may carry an `<ocr>` layer holding the positioned output of an OCR run. The node hierarchy inside that layer is irregular and is not parsed, but any word node within it, at any depth, is read as one positioned OCR string. A word node is any element named `string` in any casing and any namespace, and its text comes from the `CONTENT` attribute rather than from the element's text content.
155
155
 
156
156
  ```ts
157
157
  import { fetchItemOcrData } from "ochre-sdk";
@@ -165,7 +165,7 @@ for (const ocrString of result.ocrStrings ?? []) {
165
165
  }
166
166
  ```
167
167
 
168
- `x` and `y` come from `HPOS` and `VPOS` and give the top-left corner of the box, `width` and `height` its size, and `vertices` its full bounding polygon, which is not always rectangular. Each geometry field is null when the source attribute is absent or unparseable. `resourceUuid` names the Resource that owns the OCR layer, which differs from the requested item when the OCR lives on a child Resource.
168
+ `x` and `y` come from `HPOS` and `VPOS` and give the top-left corner of the box, `width` and `height` its size, and `vertices` comes from `VERTICES` and is its full bounding polygon, which is not always rectangular. Each geometry field is null when the source attribute is absent or unparseable, and `vertices` is then empty. `resourceUuid` names the Resource that owns the OCR layer, which differs from the requested item when the OCR lives on a child Resource.
169
169
 
170
170
  Matching defaults to case-insensitive `includes` and runs against each string's `CONTENT`. Because a `<string>` holds a single OCR word, a multi-word search value is split on whitespace and a string is returned when it matches any one term. Requesting an item that does not exist is an error; an item with no OCR layer, or no matches, returns an empty array.
171
171
 
@@ -4,12 +4,15 @@ import { OcrString } from "../types/index.mjs";
4
4
  * Fetches and parses the OCR strings of an OCHRE item that match a search value
5
5
  *
6
6
  * Resources may carry an `<ocr>` layer whose internal hierarchy is irregular
7
- * and therefore not parsed. Only its `<string>` nodes are returned, wherever
8
- * they occur in that subtree, in document order. Matching runs per string, and
9
- * each `<string>` holds a single OCR word, so a multi-word search value is
10
- * split on whitespace and a string is returned when it matches any one term.
11
- * Nested child Resources are searched too, with `resourceUuid` naming the
12
- * Resource each match belongs to.
7
+ * and therefore not parsed. Only its word nodes are returned, wherever they
8
+ * occur in that subtree, in document order. A word node is any element whose
9
+ * name is `string` in any casing and any namespace, and matching reads its
10
+ * `CONTENT` attribute rather than its text content.
11
+ *
12
+ * Each word node holds a single OCR word, so a multi-word search value is split
13
+ * on whitespace and a node is returned when it matches any one term. Nested
14
+ * child Resources are searched too, with `resourceUuid` naming the Resource
15
+ * each match belongs to.
13
16
  *
14
17
  * @param uuid - The UUID of the OCHRE item to read the OCR layer of
15
18
  * @param value - The search value to match against each OCR string's content
@@ -52,9 +52,14 @@ function parseOcrStringVertices(value) {
52
52
  * Build an XQuery string to fetch matching OCR strings from the OCHRE API
53
53
  *
54
54
  * The `<ocr>` layer is marked supplemental, so it is deliberately read without
55
- * the supplemental stripping the other fetchers apply. Only the `<string>`
56
- * nodes are projected, at any depth, because OCHRE does not guarantee the shape
57
- * of the surrounding hierarchy.
55
+ * the supplemental stripping the other fetchers apply. Word nodes are projected
56
+ * at any depth, because OCHRE does not guarantee the shape of the surrounding
57
+ * hierarchy.
58
+ *
59
+ * Both the container and the word nodes are matched on a case-folded
60
+ * `local-name()` rather than a name test, because OCHRE varies the casing of
61
+ * these elements and may serve them in a namespace. A plain `//ocr//string`
62
+ * name test silently matches nothing in either of those cases.
58
63
  *
59
64
  * The matches are wrapped in an `<ocrStrings>` element rather than returned
60
65
  * directly under `<ochre>`: the API collapses an `<ochre>` element that has no
@@ -78,10 +83,10 @@ declare variable $terms := (${termValues.join(", ")});
78
83
 
79
84
  let $ochre := doc(${stringLiteral(uuid)})/ochre
80
85
  let $ocrStrings :=
81
- for $string in $ochre//ocr//string[@CONTENT]
86
+ for $string in $ochre//*[lower-case(local-name(.)) = "ocr"]//*[lower-case(local-name(.)) = "string"][@CONTENT]
82
87
  where (some $term in $terms satisfies ${matchExpression})
83
88
  return <ocrString
84
- resourceUuid="{string($string/ancestor::resource[1]/@uuid)}"
89
+ resourceUuid="{string($string/ancestor::*[local-name(.) = "resource"][1]/@uuid)}"
85
90
  content="{string($string/@CONTENT)}"
86
91
  x="{string($string/@HPOS)}"
87
92
  y="{string($string/@VPOS)}"
@@ -95,12 +100,15 @@ return <ochre><ocrStrings found="{exists($ochre)}">{$ocrStrings}</ocrStrings></o
95
100
  * Fetches and parses the OCR strings of an OCHRE item that match a search value
96
101
  *
97
102
  * Resources may carry an `<ocr>` layer whose internal hierarchy is irregular
98
- * and therefore not parsed. Only its `<string>` nodes are returned, wherever
99
- * they occur in that subtree, in document order. Matching runs per string, and
100
- * each `<string>` holds a single OCR word, so a multi-word search value is
101
- * split on whitespace and a string is returned when it matches any one term.
102
- * Nested child Resources are searched too, with `resourceUuid` naming the
103
- * Resource each match belongs to.
103
+ * and therefore not parsed. Only its word nodes are returned, wherever they
104
+ * occur in that subtree, in document order. A word node is any element whose
105
+ * name is `string` in any casing and any namespace, and matching reads its
106
+ * `CONTENT` attribute rather than its text content.
107
+ *
108
+ * Each word node holds a single OCR word, so a multi-word search value is split
109
+ * on whitespace and a node is returned when it matches any one term. Nested
110
+ * child Resources are searched too, with `resourceUuid` naming the Resource
111
+ * each match belongs to.
104
112
  *
105
113
  * @param uuid - The UUID of the OCHRE item to read the OCR layer of
106
114
  * @param value - The search value to match against each OCR string's content
package/dist/query.mjs CHANGED
@@ -472,16 +472,17 @@ function buildItemStringQueryExpression(parameters) {
472
472
  language
473
473
  })]);
474
474
  }
475
- function tokenizeOcrPhraseValue(value) {
475
+ const OCR_STRING_QNAMES = `(xs:QName("String"), fn:QName("http://www.loc.gov/standards/alto/ns-v2#", "string"))`;
476
+ function tokenizeOcrExactValue(value) {
476
477
  const terms = [];
477
478
  for (const term of value.split(/\s+/u)) if (term !== "") terms.push(term);
478
479
  return terms;
479
480
  }
480
481
  /**
481
482
  * Word queries against the OCR layer cannot carry a stemming option: the OCHRE
482
- * database has unstemmed word searches turned off, and asking an element word
483
- * query for `unstemmed` fails with `XDMP-WORDSEARCH`. Omitting the option
484
- * altogether resolves the term against the database default instead.
483
+ * database has unstemmed word searches turned off, and asking a word query for
484
+ * `unstemmed` fails with `XDMP-WORDSEARCH`. Omitting the option altogether
485
+ * resolves the term against the database default instead.
485
486
  */
486
487
  function buildOcrWordQueryExpression(parameters) {
487
488
  const { value, isCaseSensitive } = parameters;
@@ -492,30 +493,39 @@ function buildOcrWordQueryExpression(parameters) {
492
493
  "whitespace-insensitive"
493
494
  ];
494
495
  if (hasWildcardCharacters(value)) options.push("wildcarded");
495
- return `cts:element-word-query(xs:QName("string"), ${stringLiteral(value)}, (${options.map((option) => stringLiteral(option)).join(", ")}))`;
496
+ return `cts:element-attribute-word-query(${OCR_STRING_QNAMES}, xs:QName("CONTENT"), ${stringLiteral(value)}, (${options.map((option) => stringLiteral(option)).join(", ")}))`;
497
+ }
498
+ function buildOcrValueQueryExpression(parameters) {
499
+ const { value, isCaseSensitive } = parameters;
500
+ return `cts:element-attribute-value-query(${OCR_STRING_QNAMES}, xs:QName("CONTENT"), ${stringLiteral(value)}, ${buildWordQueryOptionsExpression({
501
+ matchMode: "exact",
502
+ isCaseSensitive
503
+ })})`;
496
504
  }
497
505
  /**
498
506
  * Compile an OCR text search into a query over the `<ocr>` layer of a Resource
499
507
  * document
500
508
  *
501
- * Every `<string>` node in that layer holds a single OCR word, so `includes`
502
- * matches each search term as its own word anywhere in the layer, and `exact`
503
- * matches the terms as a run of adjacent whole string values. A phrase cannot
504
- * be a word query here: word positions do not carry across the `<string>`
505
- * boundaries, which makes `cts:near-query` the only phrase mechanism, and its
506
- * distance is a total span rather than a pairwise gap.
509
+ * Each word node in that layer holds a single OCR word in its `CONTENT`
510
+ * attribute, so `includes` matches every search term as a word inside that
511
+ * attribute anywhere in the layer, and `exact` requires every term to equal a
512
+ * whole `CONTENT` value.
513
+ *
514
+ * The conjunction is only an index narrowing for `exact`. Attribute values
515
+ * carry no word positions, so `cts:near-query` over them silently degenerates
516
+ * into a conjunction and cannot express a phrase at all. Word order and
517
+ * adjacency are instead enforced by {@link registerOcrPhraseHelper} over the
518
+ * documents this narrowing returns.
507
519
  */
508
520
  function buildOcrQueryExpression(query) {
509
521
  const { value, matchMode, isCaseSensitive } = query;
510
522
  if (matchMode === "exact") {
511
- const terms = tokenizeOcrPhraseValue(value);
523
+ const terms = tokenizeOcrExactValue(value);
512
524
  if (terms.length === 0) return "cts:false-query()";
513
- const termQueryExpressions = Array.from(terms, (term) => buildCtsElementValueQueryExpression({
514
- elementName: "string",
525
+ return buildNestedElementQuery(["ocr"], buildAndCtsQueryExpressionInternal(Array.from(terms, (term) => buildOcrValueQueryExpression({
515
526
  value: term,
516
527
  isCaseSensitive
517
- }));
518
- return buildNestedElementQuery(["ocr"], termQueryExpressions.length === 1 ? termQueryExpressions[0] ?? "cts:false-query()" : `cts:near-query((${termQueryExpressions.join(", ")}), ${termQueryExpressions.length - 1}, ("ordered"))`);
528
+ }))));
519
529
  }
520
530
  const terms = tokenizeIncludesSearchValue({
521
531
  value,
@@ -528,6 +538,31 @@ function buildOcrQueryExpression(query) {
528
538
  }))));
529
539
  }
530
540
  /**
541
+ * Declare the filter that holds an `exact` multi-term search to a run of
542
+ * adjacent OCR words, which no CTS query over the layer can express
543
+ */
544
+ function registerOcrPhraseHelper(context) {
545
+ const helperName = "local:ocrHasPhrase";
546
+ if (context.helperNamesByKey.has(helperName)) return helperName;
547
+ context.helperNamesByKey.set(helperName, helperName);
548
+ context.helperDeclarations.push(`declare function ${helperName}($resource as node(), $terms as xs:string*, $isCaseSensitive as xs:boolean) as xs:boolean {
549
+ let $contents :=
550
+ for $word in $resource//*[lower-case(local-name(.)) = "ocr"]//*[lower-case(local-name(.)) = "string"][@CONTENT]
551
+ return if ($isCaseSensitive) then string($word/@CONTENT) else lower-case(string($word/@CONTENT))
552
+ let $needles :=
553
+ for $term in $terms
554
+ return if ($isCaseSensitive) then $term else lower-case($term)
555
+ let $length := count($needles)
556
+ return
557
+ some $start in (1 to (count($contents) - $length + 1))
558
+ satisfies (
559
+ every $offset in (1 to $length)
560
+ satisfies $contents[$start + $offset - 1] = $needles[$offset]
561
+ )
562
+ };`);
563
+ return helperName;
564
+ }
565
+ /**
531
566
  * Bind the UUIDs of the Resource documents whose OCR layer matches a query,
532
567
  * reusing the binding when the same search is requested more than once
533
568
  */
@@ -541,10 +576,15 @@ function registerOcrBinding(context, query) {
541
576
  if (existingName != null) return existingName;
542
577
  const name = `$ocrItemUuids${context.ocrBindings.length + 1}`;
543
578
  const queryExpression = buildOcrQueryExpression(query);
579
+ const phraseTerms = query.matchMode === "exact" ? tokenizeOcrExactValue(query.value) : [];
544
580
  context.ocrBindingNamesByKey.set(key, name);
581
+ const searchExpression = `cts:search(/ochre/resource, ${queryExpression})`;
582
+ const phraseHelperName = phraseTerms.length > 1 ? registerOcrPhraseHelper(context) : null;
545
583
  context.ocrBindings.push({
546
584
  name,
547
- expression: queryExpression === "cts:false-query()" ? "()" : `cts:search(/ochre/resource, ${queryExpression})/@uuid/string()`
585
+ expression: queryExpression === "cts:false-query()" ? "()" : phraseHelperName == null ? `${searchExpression}/@uuid/string()` : `for $ocrResource in ${searchExpression}
586
+ where ${phraseHelperName}($ocrResource, (${phraseTerms.map((term) => stringLiteral(term)).join(", ")}), ${query.isCaseSensitive ? "true()" : "false()"})
587
+ return string($ocrResource/@uuid)`
548
588
  });
549
589
  return name;
550
590
  }
@@ -266,11 +266,15 @@ type ImageMap = {
266
266
  * Positioned OCR string in OCHRE
267
267
  *
268
268
  * OCHRE gives no guarantee about the node hierarchy inside a Resource's
269
- * `<ocr>` layer, so only `<string>` nodes are parsed, at whatever depth they
270
- * occur. `x` and `y` come from `HPOS` and `VPOS` and locate the top-left
271
- * corner of the string's box, while `vertices` is its full bounding polygon,
272
- * which is not necessarily rectangular. Every geometry attribute is optional
273
- * in the source, so each one is null when absent or unparseable.
269
+ * `<ocr>` layer, so only its word nodes are parsed, at whatever depth they
270
+ * occur. A word node is any element named `string` in any casing and any
271
+ * namespace, and its text is read from the `CONTENT` attribute.
272
+ *
273
+ * `x` and `y` come from `HPOS` and `VPOS` and locate the top-left corner of
274
+ * the word's box, while `vertices` comes from `VERTICES` and is its full
275
+ * bounding polygon, which is not necessarily rectangular. Every geometry
276
+ * attribute is optional in the source, so each one is null when absent or
277
+ * unparseable, and `vertices` is then empty.
274
278
  *
275
279
  * `resourceUuid` is the Resource that owns the OCR layer, which differs from
276
280
  * the requested item when the OCR lives on a child Resource.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ochre-sdk",
3
- "version": "1.0.74",
3
+ "version": "1.0.75",
4
4
  "type": "module",
5
5
  "license": "MIT",
6
6
  "description": "Node.js library for working with OCHRE (Online Cultural and Historical Research Environment) data",