cry-search 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (64) hide show
  1. package/CLAUDE.md +285 -0
  2. package/LICENSE.md +67 -0
  3. package/README.md +283 -0
  4. package/UNIVERSE.md +818 -0
  5. package/dist/common/SearchUniverse.d.ts +210 -0
  6. package/dist/common/SearchUniverse.d.ts.map +1 -0
  7. package/dist/common/findInArray.d.ts +37 -0
  8. package/dist/common/findInArray.d.ts.map +1 -0
  9. package/dist/common/findInArrayReturnDataAndMeta.d.ts +52 -0
  10. package/dist/common/findInArrayReturnDataAndMeta.d.ts.map +1 -0
  11. package/dist/common/findInLinkedArrays.d.ts +66 -0
  12. package/dist/common/findInLinkedArrays.d.ts.map +1 -0
  13. package/dist/common/numeric/createSearchArrayMetadata.d.ts +65 -0
  14. package/dist/common/numeric/createSearchArrayMetadata.d.ts.map +1 -0
  15. package/dist/common/numeric/createSearchLinkedMetadata.d.ts +46 -0
  16. package/dist/common/numeric/createSearchLinkedMetadata.d.ts.map +1 -0
  17. package/dist/common/numeric/findInLinkedArrays.d.ts +32 -0
  18. package/dist/common/numeric/findInLinkedArrays.d.ts.map +1 -0
  19. package/dist/common/numeric/index.d.ts +9 -0
  20. package/dist/common/numeric/index.d.ts.map +1 -0
  21. package/dist/common/numeric/searchInData.d.ts +31 -0
  22. package/dist/common/numeric/searchInData.d.ts.map +1 -0
  23. package/dist/common/numeric/updateSearchMetadata.d.ts +60 -0
  24. package/dist/common/numeric/updateSearchMetadata.d.ts.map +1 -0
  25. package/dist/common/string/createSearchArrayMetadata.d.ts +117 -0
  26. package/dist/common/string/createSearchArrayMetadata.d.ts.map +1 -0
  27. package/dist/common/string/createSearchLinkedMetadata.d.ts +36 -0
  28. package/dist/common/string/createSearchLinkedMetadata.d.ts.map +1 -0
  29. package/dist/common/string/index.d.ts +10 -0
  30. package/dist/common/string/index.d.ts.map +1 -0
  31. package/dist/common/string/updateSearchMetadata.d.ts +103 -0
  32. package/dist/common/string/updateSearchMetadata.d.ts.map +1 -0
  33. package/dist/common/syncSearchArrayMetadata.d.ts +64 -0
  34. package/dist/common/syncSearchArrayMetadata.d.ts.map +1 -0
  35. package/dist/common/updateSearchLinkedMetadata.d.ts +129 -0
  36. package/dist/common/updateSearchLinkedMetadata.d.ts.map +1 -0
  37. package/dist/index.cjs +1827 -0
  38. package/dist/index.d.cts +32 -0
  39. package/dist/index.d.ts +32 -0
  40. package/dist/index.d.ts.map +1 -0
  41. package/dist/index.js +1795 -0
  42. package/dist/types/index.d.ts +300 -0
  43. package/dist/types/index.d.ts.map +1 -0
  44. package/dist/types.d.ts +6 -0
  45. package/dist/types.d.ts.map +1 -0
  46. package/dist/utils/matchToken.d.ts +65 -0
  47. package/dist/utils/matchToken.d.ts.map +1 -0
  48. package/dist/utils/matchTokenNumeric.d.ts +94 -0
  49. package/dist/utils/matchTokenNumeric.d.ts.map +1 -0
  50. package/dist/utils/normalizeDates.d.ts +39 -0
  51. package/dist/utils/normalizeDates.d.ts.map +1 -0
  52. package/dist/utils/prepareStringForSearch.d.ts +28 -0
  53. package/dist/utils/prepareStringForSearch.d.ts.map +1 -0
  54. package/dist/utils/prepareStringForSearchNumeric.d.ts +42 -0
  55. package/dist/utils/prepareStringForSearchNumeric.d.ts.map +1 -0
  56. package/dist/utils/preprocessString.d.ts +26 -0
  57. package/dist/utils/preprocessString.d.ts.map +1 -0
  58. package/dist/utils/sanitiseString.d.ts +26 -0
  59. package/dist/utils/sanitiseString.d.ts.map +1 -0
  60. package/dist/utils/tokenRegistry.d.ts +70 -0
  61. package/dist/utils/tokenRegistry.d.ts.map +1 -0
  62. package/dist/utils/tokenize.d.ts +46 -0
  63. package/dist/utils/tokenize.d.ts.map +1 -0
  64. package/package.json +67 -0
package/CLAUDE.md ADDED
@@ -0,0 +1,285 @@
1
+ # cry-search library
2
+
3
+ cry-search is a library for fast searching in large datasets.
4
+
5
+ ## Tech stack
6
+ 1. Bun bundler for development, testing, and bundling
7
+ 2. ESM jscript in dist for distribution
8
+
9
+ ## Architecture
10
+
11
+ The library has two implementations:
12
+ 1. **Numeric (default)** - Memory-optimized using Uint32Array for token storage
13
+ - Location: `src/common/numeric/`
14
+ - ~45% less memory than string implementation
15
+ - Global token registry with array-based lookups
16
+ - Tokens stored as numeric IDs in Uint32Array
17
+ 2. **String (legacy)** - Original string-based implementation
18
+ - Location: `src/common/string/`
19
+ - Uses string arrays for token storage
20
+ - Available with `*String` suffix (e.g., `createSearchArrayMetadataString`)
21
+
22
+ ### Token Registry (Numeric Implementation)
23
+
24
+ The numeric implementation uses a global token registry (`src/utils/tokenRegistry.ts`):
25
+ - `tokens: string[]` - Global array of all tokens (preallocated 10k)
26
+ - Token ID = index in array
27
+ - During metadata build: temporary `Map<string, number>` for fast lookups, cleared after use
28
+ - During search: `findTokenId()` linear search (no Map in memory)
29
+ - Match modes encoded in high 3 bits of Uint32 token ID
30
+
31
+ ## Requirements
32
+ 1. Metadata
33
+ 1. searching uses metadata for object arrays to be searched
34
+ 2. metadata is prepared in advance
35
+ 3. the purpose of having metadata is to make searches as fast as possible
36
+ 2. Search requirements
37
+ 1. input preparation
38
+ 1. date normalization (first step)
39
+ 1. ISO 8601: "2025-12-27T11:35:25.814Z" or "2025-12-27" → "20251227"
40
+ 2. ISO year-month: "2025-12" → "202512"
41
+ 3. d.m.yyyy: "1.3.2025" or "01.03.2025" → "20250301"
42
+ 4. d.m.yy: "27.12.25" → "20251227" (2-digit year assumes 2000s)
43
+ 5. m.yyyy: "3.2025" or "03.2025" → "202503"
44
+ 6. month name + year: "mar 2025", "march 2025", "marec 2025" → "202503"
45
+ 7. month names supported in English and Slovenian (full and abbreviated)
46
+ 2. preprocessing
47
+ 1. before splitting the string, we find possible numbers and replace decimal commas with decimal periods and drop leading zeros
48
+ 1. eg "abc 12,3x" -> "abc 12.3x"
49
+ 2. eg "abc ,3x" -> "abc .3x"
50
+ 3. eg "abc 0,3x" -> "abc .3x"
51
+ 4. eg "abc 0.3x" -> "abc .3x"
52
+ 5. eg "007" -> "7"
53
+ 6. multiple periods stay as-is: "1.2.3" -> "1.2.3"
54
+ 2. we remove spaces in pattern digits - digits
55
+ 1. eg. "12 - 25" => "12-25"
56
+ 3. we insert spaces between starting letters and digits (only at letter-to-digit boundaries)
57
+ 1. eg "SI123" -> "SI 123"
58
+ 2. eg "123SI" -> "123 SI"
59
+ 3. eg "A1B2" -> "A1B2"
60
+ 3. sanitation
61
+ 1. case-insensitive
62
+ 2. diacritics-insensitive (č,ć,c -> same c)
63
+ 3. we only keep letters (a-z), digits (0-9), periods (.), minuses (-), spaces (" ") in search
64
+ 4. other characters are replaced with space (" ")
65
+ 2. tokenization
66
+ 1. we split strings on spaces
67
+ 2. then we trim all substrings
68
+ 3. matching
69
+ 1. we match words in any order
70
+ 2. we match numbers from start or from end
71
+ 1. eg. 123 matches 12345 and 45123 but not 7812345
72
+ 4. nested data extraction
73
+ 1. nested objects and arrays are automatically extracted recursively
74
+ 2. Date objects are converted to ISO strings before tokenization
75
+ 3. example: { customer: { firstName: 'John' }, items: [{ opis: 'A' }] } extracts "john", "a"
76
+ 5. dot notation in searchInFields and fieldMatchModes
77
+ 1. use dot notation to target specific nested fields: 'customer.firstName', 'items.opis'
78
+ 2. arrays of objects: 'items.opis' extracts from all array elements
79
+ 3. deeply nested: 'data.level1.level2.value'
80
+ 4. fieldMatchModes supports dot notation: { 'items.barcode': 'startEnd' }
81
+ 6. building metadata
82
+ 1. function prepareStringForSearch -> tokens[]
83
+ 2. function createSearchArrayMetadata (numeric, default)
84
+ 1. all elements in array must have unique _id:string field
85
+ 2. prepares input as metadata for each array item
86
+ 3. stores tokens as Uint32Array of token IDs
87
+ 4. creates temporary tokensMap, clears after use
88
+ 5. supports per-field match modes (encoded in token ID high bits)
89
+ 7. searching
90
+ 1. searchInData / searchInDataReturnIds
91
+ 1. prepares query using prepareStringForSearch
92
+ 2. matches query tokens against numeric metadata
93
+ 3. returns matching IDs
94
+ 2. searchInDataReturnObjects
95
+ 1. calls searchInDataReturnIds and looks up objects
96
+ 8. linked arrays searching
97
+ 1. purpose: search across two related arrays (e.g., stranke + pacienti) where secondary items belong to primary items via foreign key
98
+ 2. building linked metadata
99
+ 1. function createSearchLinkedMetadata
100
+ 1. takes primary items + their metadata
101
+ 2. takes secondary items + their metadata
102
+ 3. takes foreignKeyGetter function to extract primary ID from secondary item
103
+ 4. builds reverse index: primary._id -> secondary._id[]
104
+ 5. no data duplication - stores references and indexes only
105
+ 3. searching linked arrays
106
+ 1. function findInLinkedArrays
107
+ 1. for each primary item, collects all tokens from primary + all its secondaries
108
+ 2. matches query tokens against combined token pool
109
+ 3. determines matchedIn: 'primary' | 'secondary' | 'both'
110
+ 4. returns only secondaries that contributed to the match (not all secondaries)
111
+ 5. tokens that match in primary are excluded from secondary matching to avoid foreign key redundancy
112
+ 2. function findInLinkedArraysSimple
113
+ 1. calls findInLinkedArrays and returns just the results array
114
+ 4. example use case
115
+ 1. search "krajnik angie" matches stranka "Krajnik" with pacient "Angie"
116
+ 2. returns: { primary: Stranka, secondaries: [matching Pacient], matchedIn: 'both' }
117
+
118
+ ## Types
119
+
120
+ 1. Id: string
121
+ 2. NumericToken: number (index into global tokens array)
122
+ 3. NumericTokenSortedList: Uint32Array (sorted numeric token IDs)
123
+ 4. NumericSearchMetadata: Map<Id, NumericTokenSortedList> (default SearchMetadata)
124
+ 5. StringSearchMetadata: Map<Id, string[]> (legacy, string-based)
125
+ 6. LinkedSearchMetadata: { primaryMeta, linkedMeta, primaryToLinked: Map<Id, Id[]> }
126
+ 7. SearchableObject: { _id: Id, _deleted?: Date, _blocked?: Date }
127
+ 8. SearchableData<T>: Map<Id, T> where T extends SearchableObject
128
+ 9. SearchableObjectSpec<T extends SearchableObject, C = any>:
129
+ 1. searchInFields?: (keyof T | string)[] - supports dot notation like 'customer.firstName'
130
+ 2. dontSearchInFields?: (keyof T)[]
131
+ 3. searchContext?: C
132
+ 4. extractSearchableStringFn?: (obj: SearchableObject) => string | undefined
133
+ 5. enrichSearchObjectFn?: (obj: SearchableObject, searchContext: C) => void
134
+ 6. fieldMatchModes?: Partial<Record<keyof T | string, MatchMode>> - per-field match modes
135
+ 10. MatchMode: 'start' | 'end' | 'startEnd' | 'anywhere' | 'whole'
136
+ 1. Field token prefixes: '<' (start), '>' (end), '+' (startEnd), '*' (anywhere), '!' (whole)
137
+ 2. Numeric: encoded in high 3 bits of Uint32 token ID
138
+ 11. SearchableLinkedObjectsSpec: primary SearchableObjectSpec, linked SearchableObjectSpec, foreignKeyGetter
139
+
140
+ ## Search Query Prefixes
141
+
142
+ Override match behavior per token at search time. Priority: query prefix > field mode > global SearchOpts.
143
+
144
+ ### User-friendly prefixes
145
+ | Prefix | Mode | Description | Example |
146
+ |--------|------|-------------|---------|
147
+ | `word..` | start | Match tokens starting with word | `APL..` matches "APL-IP15P" |
148
+ | `..word` | end | Match tokens ending with word | `..Pro` matches "MacBookPro" |
149
+ | `=word` | whole | Exact match only | `=jan` matches "jan" not "jana" |
150
+ | `?word` | anywhere | Match word anywhere in token | `?book` matches "MacBook" |
151
+ | `--word` | negation | Exclude results containing word | `apple --iphone` excludes iPhones |
152
+ | `~word` | negation | Same as `--word` | `apple ~iphone` |
153
+
154
+ ### Programmatic prefixes
155
+ | Prefix | Mode |
156
+ |--------|------|
157
+ | `<word` | start |
158
+ | `>word` | end |
159
+ | `+word` | startEnd |
160
+ | `*word` | anywhere |
161
+ | `!word` | whole |
162
+
163
+ ### Examples
164
+ 1. `"jana useni"` → finds "jana usenik" (default startEnd matching)
165
+ 2. `"jana =useni"` → no match (whole match required, "useni" ≠ "usenik")
166
+ 3. `"jana --macka"` → finds "jana" but excludes any result containing "macka"
167
+ 4. `"APL.."` → finds all tokens starting with "APL" (APL-IP15P, APL-MBP16)
168
+ 5. `"..nik"` → finds tokens ending with "nik" (usenik, krajnik)
169
+
170
+
171
+ ## Main Functions (Default = Numeric)
172
+
173
+ 1. prepareStringForSearch(input: string): string[]
174
+ 1. pipeline: normalizeDates → preprocessString → sanitiseString → tokenize → sort
175
+ 2. createSearchArrayMetadata(data: SearchableData, spec?): NumericSearchMetadata
176
+ 1. creates temporary tokensMap for fast string→number lookups
177
+ 2. stores tokens as Uint32Array per item
178
+ 3. clears tokensMap after use (only global tokens[] remains)
179
+ 4. for _deleted or _blocked nothing is added
180
+ 3. prepareObjectForSearch(obj, spec?, tokensMap): NumericTokenSortedList | undefined
181
+ 1. returns undefined if _deleted or _blocked
182
+ 2. if custom extractSearchableStringFn provided, uses simple processing
183
+ 3. otherwise processes fields individually to support per-field match modes
184
+ 4. recursively extracts strings from nested objects/arrays
185
+ 4. searchInData / searchInDataReturnIds(query: string, metadata): Id[]
186
+ 5. searchInDataReturnObjects<T>(query, metadata, data): Map<Id, T>
187
+ 6. searchInDataWithLimit(query, metadata, limit): Id[]
188
+ 7. updateSearchMetadata(metadata, objectId, obj, spec)
189
+ 1. updates or deletes metadata for this object
190
+ 2. taking into account _blocked and _deleted
191
+ 8. updateSearchMetadataBatch(metadata, batch, spec)
192
+ 1. creates temporary tokensMap, clears after use
193
+ 2. updates or deletes metadata for batch of objects
194
+
195
+ ## Token Registry Functions
196
+
197
+ 1. tokens: string[] - global array (read-only)
198
+ 2. getTokenCount(): number
199
+ 3. getTokenById(id: number): string | undefined
200
+ 4. findTokenId(token: string): number (-1 if not found)
201
+ 5. registerToken(token: string): number
202
+ 6. getOrCreateTokenId(token: string, tokensMap): number
203
+ 7. createTokensMap(): Map<string, number>
204
+ 8. clearTokensMap(tokensMap): void
205
+ 9. resetTokenRegistry(): void (for testing only)
206
+
207
+ ## String Implementation (Legacy)
208
+
209
+ Available with `*String` suffix:
210
+ - createSearchArrayMetadataString
211
+ - createSearchLinkedMetadataString
212
+ - updateSearchMetadataString
213
+ - updateSearchMetadataBatchString
214
+ - findInArray (uses string metadata)
215
+ - findInArrayReturnDataAndMeta
216
+ - findInLinkedArrays
217
+
218
+ ## File Structure
219
+
220
+ ```
221
+ src/
222
+ ├── common/
223
+ │ ├── numeric/ # Default numeric implementation
224
+ │ │ ├── index.ts
225
+ │ │ ├── createSearchArrayMetadata.ts
226
+ │ │ ├── createSearchLinkedMetadata.ts
227
+ │ │ ├── updateSearchMetadata.ts
228
+ │ │ └── searchInData.ts
229
+ │ ├── string/ # Legacy string implementation
230
+ │ │ ├── index.ts
231
+ │ │ ├── createSearchArrayMetadata.ts
232
+ │ │ ├── createSearchLinkedMetadata.ts
233
+ │ │ └── updateSearchMetadata.ts
234
+ │ ├── findInArray.ts # String-based search
235
+ │ ├── findInLinkedArrays.ts # String-based linked search
236
+ │ ├── syncSearchArrayMetadata.ts
237
+ │ ├── updateSearchLinkedMetadata.ts
238
+ │ └── SearchUniverse.ts # Multi-collection manager (uses numeric)
239
+ ├── utils/
240
+ │ ├── tokenRegistry.ts # Global tokens array
241
+ │ ├── prepareStringForSearch.ts
242
+ │ ├── prepareStringForSearchNumeric.ts
243
+ │ ├── matchToken.ts # String-based matching
244
+ │ ├── matchTokenNumeric.ts # Numeric matching
245
+ │ └── ...
246
+ ├── types/
247
+ │ └── index.ts
248
+ └── index.ts # Main exports (numeric as default)
249
+ ```
250
+
251
+ ## Coding rules
252
+
253
+ 1. each function is in separate file
254
+ 1. preprocessString
255
+ 2. sanitiseString
256
+ 3. tokenize
257
+ 4. matchToken / matchTokenNumeric
258
+ 5. searchInData
259
+
260
+ ## Testing
261
+
262
+ 1. Large real-life datasets are in ./test-data (arikli, stranke, pacienti).
263
+ 2. tests are in ./test
264
+ 3. tests per implementation are in ./test/(IMPLEMENTATION)
265
+ 4. We have interactive.ts that
266
+ 1. loads all three datasets and creates SearchMetadata for each
267
+ 2. measures memory consumption for each loaded dataset and for created SearchMetadata
268
+ 3. lets user enter search string in a loop and presents
269
+ 1. 3 results from each dataset (where found)
270
+ 1. no fieldnames, just values
271
+ 2. prints time spent searching in each dataset
272
+ 4. Supports command lines for changing data
273
+ 1. 621771cadf380d819d0b0970: Angie -> Ančka
274
+ 1. find object with _id===621771cadf380d819d0b0970
275
+ 2. find key with value "Angie"
276
+ 3. change value to Ančka
277
+ 2. 621771cadf380d819d0b0970: x siamska -> siamska
278
+ 3. 621771cadf380d819d0b0970: _blocked -> true/false
279
+ 1. sets _blocked to new Date() or deletes the field
280
+ 4. 621771cadf380d819d0b0970: _deleted -> true/false
281
+ 1. like blocked
282
+ 5. benchmark-numeric-vs-string.ts compares implementations
283
+ 1. Memory: numeric ~45% less
284
+ 2. Build time: ~same
285
+ 3. Search time: varies by match mode
package/LICENSE.md ADDED
@@ -0,0 +1,67 @@
1
+ # Proprietary License
2
+
3
+ Copyright (c) 2025. All Rights Reserved.
4
+
5
+ ## Terms and Conditions
6
+
7
+ ### 1. Definitions
8
+
9
+ - "Software" refers to the cry-search library, including all source code, documentation, and associated files.
10
+ - "Licensor" refers to the copyright holder(s) of this Software.
11
+ - "You" refers to any individual or entity that accesses, uses, or attempts to use this Software.
12
+
13
+ ### 2. Grant of License
14
+
15
+ Subject to the terms of this License, the Licensor grants You a limited, non-exclusive, non-transferable, revocable license to use this Software **solely for personal, non-commercial, evaluation purposes**.
16
+
17
+ ### 3. Restrictions
18
+
19
+ You are expressly **PROHIBITED** from:
20
+
21
+ 1. **Commercial Use**: Using the Software, in whole or in part, for any commercial purpose, including but not limited to:
22
+ - Incorporating the Software into commercial products or services
23
+ - Using the Software to provide services to third parties for compensation
24
+ - Using the Software in any revenue-generating activity
25
+ - Using the Software in a business or organizational context
26
+
27
+ 2. **Modification**: Modifying, adapting, translating, reverse engineering, decompiling, disassembling, or creating derivative works based on the Software.
28
+
29
+ 3. **Distribution**: Distributing, sublicensing, leasing, renting, lending, selling, reselling, or otherwise transferring the Software or any rights therein to any third party.
30
+
31
+ 4. **Sharing**: Forwarding, sharing, publishing, uploading, or making the Software available to any third party through any means, including but not limited to:
32
+ - Public or private repositories
33
+ - File sharing services
34
+ - Email or messaging platforms
35
+ - Any form of electronic or physical distribution
36
+
37
+ 5. **Copying**: Copying or reproducing the Software except as strictly necessary for personal evaluation use on a single machine.
38
+
39
+ 6. **Removal of Notices**: Removing, altering, or obscuring any copyright, trademark, or other proprietary notices contained in the Software.
40
+
41
+ ### 4. Intellectual Property
42
+
43
+ The Software and all copies thereof are proprietary to the Licensor and title thereto remains exclusively with the Licensor. All rights in the Software not specifically granted in this License are reserved to the Licensor.
44
+
45
+ ### 5. No Warranty
46
+
47
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, AND NONINFRINGEMENT. IN NO EVENT SHALL THE LICENSOR BE LIABLE FOR ANY CLAIM, DAMAGES, OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT, OR OTHERWISE, ARISING FROM, OUT OF, OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
48
+
49
+ ### 6. Termination
50
+
51
+ This License is effective until terminated. Your rights under this License will terminate automatically without notice if You fail to comply with any of its terms. Upon termination, You must immediately cease all use of the Software and destroy all copies in Your possession or control.
52
+
53
+ ### 7. Governing Law
54
+
55
+ This License shall be governed by and construed in accordance with applicable laws, without regard to conflicts of law principles.
56
+
57
+ ### 8. Entire Agreement
58
+
59
+ This License constitutes the entire agreement between You and the Licensor concerning the Software and supersedes all prior or contemporaneous oral or written communications, proposals, and representations with respect to the Software.
60
+
61
+ ### 9. Contact
62
+
63
+ For licensing inquiries, including requests for commercial licenses, please contact the copyright holder.
64
+
65
+ ---
66
+
67
+ **BY USING THIS SOFTWARE, YOU ACKNOWLEDGE THAT YOU HAVE READ THIS LICENSE, UNDERSTAND IT, AND AGREE TO BE BOUND BY ITS TERMS AND CONDITIONS.**
package/README.md ADDED
@@ -0,0 +1,283 @@
1
+ # cry-search
2
+
3
+ A fast, memory-efficient search library for large datasets with support for tokenized matching, linked collections, and per-field match modes.
4
+
5
+ ## Features
6
+
7
+ - **Fast search** - Pre-built metadata enables instant queries on large datasets
8
+ - **Memory efficient** - Numeric token storage uses ~45% less memory than string-based approaches
9
+ - **Updatable** - Add, update, or remove items without rebuilding metadata
10
+ - **Flexible matching** - Per-field match modes and query prefixes for prefix, suffix, anywhere, exact, and negation
11
+ - **Search prefixed** - use `--word` to exclude it, `...word` for suffix, and more
12
+ - **Automatic normalization** - Handles diacritics, case, dates, and number formatting
13
+ - **[Universal search](./UNIVERSE.md)** - Search and update across linked collections (e.g., customers with their pets)
14
+
15
+ ## Table of Contents
16
+
17
+ - [cry-search](#cry-search)
18
+ - [Features](#features)
19
+ - [Table of Contents](#table-of-contents)
20
+ - [Installation](#installation)
21
+ - [Examples](#examples)
22
+ - [Single Table Search](#single-table-search)
23
+ - [SearchUniverse with Linked Collections](#searchuniverse-with-linked-collections)
24
+ - [Updating Existing Data](#updating-existing-data)
25
+ - [Architecture](#architecture)
26
+ - [Text Processing Pipeline](#text-processing-pipeline)
27
+ - [Metadata Building](#metadata-building)
28
+ - [Searching](#searching)
29
+ - [Memory Optimization](#memory-optimization)
30
+ - [Query Prefixes](#query-prefixes)
31
+ - [Legacy String Implementation](#legacy-string-implementation)
32
+ - [Specification](#specification)
33
+ - [License](#license)
34
+
35
+ ## Installation
36
+
37
+ ```bash
38
+ npm install cry-search
39
+ ```
40
+
41
+ ## Examples
42
+
43
+ ### Single Table Search
44
+
45
+ ```typescript
46
+ import {
47
+ createSearchArrayMetadata,
48
+ searchInDataReturnObjects,
49
+ arrayToSearchableData,
50
+ type SearchableObject,
51
+ } from 'cry-search';
52
+
53
+ interface Product extends SearchableObject {
54
+ _id: string;
55
+ name: string;
56
+ description: string;
57
+ sku: string;
58
+ }
59
+
60
+ const products: Product[] = [
61
+ { _id: '1', name: 'iPhone 15 Pro', description: 'Latest Apple smartphone', sku: 'APL-IP15P' },
62
+ { _id: '2', name: 'Samsung Galaxy S24', description: 'Android flagship phone', sku: 'SAM-GS24' },
63
+ { _id: '3', name: 'MacBook Pro 16"', description: 'Apple laptop for professionals', sku: 'APL-MBP16' },
64
+ ];
65
+
66
+ // Convert array to searchable data (Map<id, item>)
67
+ const data = arrayToSearchableData(products);
68
+
69
+ // Build search metadata (do this once, reuse for all searches)
70
+ const metadata = createSearchArrayMetadata(data, {
71
+ searchInFields: ['name', 'description', 'sku'],
72
+ });
73
+
74
+ // Search
75
+ const results = searchInDataReturnObjects('apple', metadata, data);
76
+ // Returns: Map with products 1 and 3
77
+
78
+ const results2 = searchInDataReturnObjects('samsung galaxy', metadata, data);
79
+ // Returns: Map with product 2
80
+
81
+ // Search is diacritics and case insensitive
82
+ const results3 = searchInDataReturnObjects('IPHONE', metadata, data);
83
+ // Returns: Map with product 1
84
+ ```
85
+
86
+ ### SearchUniverse with Linked Collections
87
+
88
+ For related data (like customers and their pets), use `SearchUniverse` to search across collections:
89
+
90
+ ```typescript
91
+ import { SearchUniverse, type SearchableObject } from 'cry-search';
92
+
93
+ interface Stranka extends SearchableObject {
94
+ _id: string;
95
+ name: string;
96
+ address: string;
97
+ }
98
+
99
+ interface Pacient extends SearchableObject {
100
+ _id: string;
101
+ stranka_id: string; // Foreign key to Stranka
102
+ name: string;
103
+ species: string;
104
+ }
105
+
106
+ // Create universe and register collections
107
+ const universe = new SearchUniverse();
108
+
109
+ universe.registerCollection<Stranka>('stranke', {
110
+ spec: { searchInFields: ['name', 'address'] },
111
+ });
112
+
113
+ universe.registerCollection<Pacient>('pacienti', {
114
+ spec: { dontSearchInFields: ['stranka_id'] },
115
+ linkedTo: {
116
+ collectionName: 'stranke',
117
+ foreignKeyGetter: (p) => p.stranka_id,
118
+ },
119
+ });
120
+
121
+ // Load data
122
+ universe.loadCollection('stranke', [
123
+ { _id: 's1', name: 'Krajnik', address: 'Ljubljana' },
124
+ { _id: 's2', name: 'Novak', address: 'Maribor' },
125
+ ]);
126
+
127
+ universe.loadCollection('pacienti', [
128
+ { _id: 'p1', stranka_id: 's1', name: 'Angie', species: 'cat' },
129
+ { _id: 'p2', stranka_id: 's1', name: 'Rex', species: 'dog' },
130
+ { _id: 'p3', stranka_id: 's2', name: 'Bella', species: 'cat' },
131
+ ]);
132
+
133
+ // Search across all collections
134
+ const results = universe.searchUniversally('krajnik angie');
135
+ // Returns:
136
+ // {
137
+ // stranke: [{ _id: 's1', name: 'Krajnik', ... }],
138
+ // pacienti: [{
139
+ // primary: { _id: 's1', name: 'Krajnik', ... },
140
+ // linked: [{ _id: 'p1', name: 'Angie', ... }],
141
+ // matchedIn: 'both'
142
+ // }]
143
+ // }
144
+
145
+ // Search only in linked items
146
+ const results2 = universe.searchUniversally('rex');
147
+ // Returns pacienti results with matchedIn: 'linked'
148
+
149
+ // Search only in primary - returns all linked items
150
+ const results3 = universe.searchUniversally('novak');
151
+ // Returns Novak stranka with Bella in linked array
152
+ ```
153
+
154
+ ### Updating Existing Data
155
+
156
+ When data changes, update the metadata to keep search in sync:
157
+
158
+ ```typescript
159
+ import {
160
+ createSearchArrayMetadata,
161
+ updateSearchMetadata,
162
+ updateSearchMetadataBatch,
163
+ arrayToSearchableData,
164
+ } from 'cry-search';
165
+
166
+ // Initial setup
167
+ const data = arrayToSearchableData(products);
168
+ const metadata = createSearchArrayMetadata(data);
169
+
170
+ // Update a single item
171
+ const updatedProduct = { ...products[0], name: 'iPhone 16 Pro' };
172
+ data.set(updatedProduct._id, updatedProduct);
173
+ updateSearchMetadata(metadata, updatedProduct._id, updatedProduct);
174
+
175
+ // Update multiple items at once
176
+ const updates = [
177
+ { _id: '1', name: 'iPhone 16 Pro Max', description: 'Newest Apple phone', sku: 'APL-IP16PM' },
178
+ { _id: '4', name: 'iPad Pro', description: 'Apple tablet', sku: 'APL-IPAD' },
179
+ ];
180
+ for (const item of updates) {
181
+ data.set(item._id, item);
182
+ }
183
+ updateSearchMetadataBatch(metadata, updates);
184
+
185
+ // Mark item as deleted (removes from search but keeps in data)
186
+ const deletedProduct = { ...products[1], _deleted: new Date() };
187
+ data.set(deletedProduct._id, deletedProduct);
188
+ updateSearchMetadata(metadata, deletedProduct._id, deletedProduct);
189
+
190
+ // With SearchUniverse - handles linked indexes automatically
191
+ universe.updateSearchMetadata('pacienti', 'p1', {
192
+ _id: 'p1',
193
+ stranka_id: 's2', // Changed owner from s1 to s2
194
+ name: 'Angie',
195
+ species: 'cat',
196
+ });
197
+ ```
198
+
199
+ ## Architecture
200
+
201
+ cry-search uses a two-phase approach: **build** metadata once, then **search** instantly.
202
+
203
+ ### Text Processing Pipeline
204
+
205
+ Both metadata building and search queries go through the same pipeline:
206
+
207
+ ```
208
+ Input: "Čokolada 27.12.2025 SI123"
209
+ ↓
210
+ 1. Normalize dates → "Čokolada 20251227 SI123"
211
+ 2. Preprocess → "Čokolada 20251227 SI 123" (split letter-digit)
212
+ 3. Sanitize → "cokolada 20251227 si 123" (lowercase, remove diacritics)
213
+ 4. Tokenize → ["cokolada", "20251227", "si", "123"]
214
+ 5. Sort → ["123", "20251227", "cokolada", "si"]
215
+ 6. Numeric IDs → Uint32Array [42, 891, 156, 7] (indices into global registry)
216
+ ```
217
+
218
+ ### Metadata Building
219
+
220
+ Tokens are stored as numeric IDs in a global registry. Each item's tokens are stored in a sorted `Uint32Array` for memory efficiency and fast binary search.
221
+
222
+ ### Searching
223
+
224
+ Query tokens are matched against metadata using binary search. All query tokens must match for an item to be returned. Match modes (start, end, startEnd, anywhere, whole) control how tokens are compared.
225
+
226
+ ### Memory Optimization
227
+
228
+ The numeric implementation stores tokens as `Uint32Array` indices instead of string arrays, reducing memory usage by ~45% compared to the string-based implementation.
229
+
230
+ ## Query Prefixes
231
+
232
+ Override match behavior per token at search time using prefixes:
233
+
234
+ | Prefix | Mode | Description |
235
+ |--------|------|-------------|
236
+ | `word..` | start | Match tokens starting with "word" |
237
+ | `..word` | end | Match tokens ending with "word" |
238
+ | `=word` | whole | Exact match only |
239
+ | `?word` | anywhere | Match "word" anywhere in token |
240
+ | `--word` or `~word` | negation | Exclude results containing "word" |
241
+
242
+ Alternative prefixes (for programmatic use): `<word` (start), `>word` (end), `+word` (startEnd), `*word` (anywhere), `!word` (whole)
243
+
244
+ ```typescript
245
+ // Find "jana" but only exact match on "usenik"
246
+ searchInDataReturnObjects('jana =usenik', metadata, data);
247
+ // Matches "Jana Usenik" but not "Jana Useniker"
248
+
249
+ // Find "apple" but exclude results containing "iphone"
250
+ searchInDataReturnObjects('apple --iphone', metadata, data);
251
+ // Matches MacBook Pro but not iPhone
252
+
253
+ // Find products with SKU starting with "APL"
254
+ searchInDataReturnObjects('APL..', metadata, data);
255
+ // Matches APL-IP15P, APL-MBP16, etc.
256
+ ```
257
+
258
+ Priority: query prefix > field match mode > global SearchOpts default.
259
+
260
+ ## Legacy String Implementation
261
+
262
+ A string-based implementation is available for backwards compatibility. Import with the `String` suffix:
263
+
264
+ ```typescript
265
+ import {
266
+ createSearchArrayMetadataString,
267
+ updateSearchMetadataString,
268
+ findInArray, // String-based search function
269
+ } from 'cry-search';
270
+ ```
271
+
272
+ ## Specification
273
+
274
+ See [CLAUDE.md](./CLAUDE.md) for the complete technical specification including:
275
+ - Input processing pipeline (date normalization, preprocessing, sanitization)
276
+ - Tokenization rules
277
+ - Matching behavior
278
+ - Linked search semantics
279
+ - Type definitions
280
+
281
+ ## License
282
+
283
+ See [LICENSE.md](./LICENSE.md) for license terms.