cry-search 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CLAUDE.md +285 -0
- package/LICENSE.md +67 -0
- package/README.md +283 -0
- package/UNIVERSE.md +818 -0
- package/dist/common/SearchUniverse.d.ts +210 -0
- package/dist/common/SearchUniverse.d.ts.map +1 -0
- package/dist/common/findInArray.d.ts +37 -0
- package/dist/common/findInArray.d.ts.map +1 -0
- package/dist/common/findInArrayReturnDataAndMeta.d.ts +52 -0
- package/dist/common/findInArrayReturnDataAndMeta.d.ts.map +1 -0
- package/dist/common/findInLinkedArrays.d.ts +66 -0
- package/dist/common/findInLinkedArrays.d.ts.map +1 -0
- package/dist/common/numeric/createSearchArrayMetadata.d.ts +65 -0
- package/dist/common/numeric/createSearchArrayMetadata.d.ts.map +1 -0
- package/dist/common/numeric/createSearchLinkedMetadata.d.ts +46 -0
- package/dist/common/numeric/createSearchLinkedMetadata.d.ts.map +1 -0
- package/dist/common/numeric/findInLinkedArrays.d.ts +32 -0
- package/dist/common/numeric/findInLinkedArrays.d.ts.map +1 -0
- package/dist/common/numeric/index.d.ts +9 -0
- package/dist/common/numeric/index.d.ts.map +1 -0
- package/dist/common/numeric/searchInData.d.ts +31 -0
- package/dist/common/numeric/searchInData.d.ts.map +1 -0
- package/dist/common/numeric/updateSearchMetadata.d.ts +60 -0
- package/dist/common/numeric/updateSearchMetadata.d.ts.map +1 -0
- package/dist/common/string/createSearchArrayMetadata.d.ts +117 -0
- package/dist/common/string/createSearchArrayMetadata.d.ts.map +1 -0
- package/dist/common/string/createSearchLinkedMetadata.d.ts +36 -0
- package/dist/common/string/createSearchLinkedMetadata.d.ts.map +1 -0
- package/dist/common/string/index.d.ts +10 -0
- package/dist/common/string/index.d.ts.map +1 -0
- package/dist/common/string/updateSearchMetadata.d.ts +103 -0
- package/dist/common/string/updateSearchMetadata.d.ts.map +1 -0
- package/dist/common/syncSearchArrayMetadata.d.ts +64 -0
- package/dist/common/syncSearchArrayMetadata.d.ts.map +1 -0
- package/dist/common/updateSearchLinkedMetadata.d.ts +129 -0
- package/dist/common/updateSearchLinkedMetadata.d.ts.map +1 -0
- package/dist/index.cjs +1827 -0
- package/dist/index.d.cts +32 -0
- package/dist/index.d.ts +32 -0
- package/dist/index.d.ts.map +1 -0
- package/dist/index.js +1795 -0
- package/dist/types/index.d.ts +300 -0
- package/dist/types/index.d.ts.map +1 -0
- package/dist/types.d.ts +6 -0
- package/dist/types.d.ts.map +1 -0
- package/dist/utils/matchToken.d.ts +65 -0
- package/dist/utils/matchToken.d.ts.map +1 -0
- package/dist/utils/matchTokenNumeric.d.ts +94 -0
- package/dist/utils/matchTokenNumeric.d.ts.map +1 -0
- package/dist/utils/normalizeDates.d.ts +39 -0
- package/dist/utils/normalizeDates.d.ts.map +1 -0
- package/dist/utils/prepareStringForSearch.d.ts +28 -0
- package/dist/utils/prepareStringForSearch.d.ts.map +1 -0
- package/dist/utils/prepareStringForSearchNumeric.d.ts +42 -0
- package/dist/utils/prepareStringForSearchNumeric.d.ts.map +1 -0
- package/dist/utils/preprocessString.d.ts +26 -0
- package/dist/utils/preprocessString.d.ts.map +1 -0
- package/dist/utils/sanitiseString.d.ts +26 -0
- package/dist/utils/sanitiseString.d.ts.map +1 -0
- package/dist/utils/tokenRegistry.d.ts +70 -0
- package/dist/utils/tokenRegistry.d.ts.map +1 -0
- package/dist/utils/tokenize.d.ts +46 -0
- package/dist/utils/tokenize.d.ts.map +1 -0
- package/package.json +67 -0
package/CLAUDE.md
ADDED
|
@@ -0,0 +1,285 @@
|
|
|
1
|
+
# cry-search library
|
|
2
|
+
|
|
3
|
+
cry-search is a library for fast searching in large datasets.
|
|
4
|
+
|
|
5
|
+
## Tech stack
|
|
6
|
+
1. Bun bundler for development, testing, and bundling
|
|
7
|
+
2. ESM jscript in dist for distribution
|
|
8
|
+
|
|
9
|
+
## Architecture
|
|
10
|
+
|
|
11
|
+
The library has two implementations:
|
|
12
|
+
1. **Numeric (default)** - Memory-optimized using Uint32Array for token storage
|
|
13
|
+
- Location: `src/common/numeric/`
|
|
14
|
+
- ~45% less memory than string implementation
|
|
15
|
+
- Global token registry with array-based lookups
|
|
16
|
+
- Tokens stored as numeric IDs in Uint32Array
|
|
17
|
+
2. **String (legacy)** - Original string-based implementation
|
|
18
|
+
- Location: `src/common/string/`
|
|
19
|
+
- Uses string arrays for token storage
|
|
20
|
+
- Available with `*String` suffix (e.g., `createSearchArrayMetadataString`)
|
|
21
|
+
|
|
22
|
+
### Token Registry (Numeric Implementation)
|
|
23
|
+
|
|
24
|
+
The numeric implementation uses a global token registry (`src/utils/tokenRegistry.ts`):
|
|
25
|
+
- `tokens: string[]` - Global array of all tokens (preallocated 10k)
|
|
26
|
+
- Token ID = index in array
|
|
27
|
+
- During metadata build: temporary `Map<string, number>` for fast lookups, cleared after use
|
|
28
|
+
- During search: `findTokenId()` linear search (no Map in memory)
|
|
29
|
+
- Match modes encoded in high 3 bits of Uint32 token ID
|
|
30
|
+
|
|
31
|
+
## Requirements
|
|
32
|
+
1. Metadata
|
|
33
|
+
1. searching uses metadata for object arrays to be searched
|
|
34
|
+
2. metadata is prepared in advance
|
|
35
|
+
3. the purpose of having metadata is to make searches as fast as possible
|
|
36
|
+
2. Search requirements
|
|
37
|
+
1. input preparation
|
|
38
|
+
1. date normalization (first step)
|
|
39
|
+
1. ISO 8601: "2025-12-27T11:35:25.814Z" or "2025-12-27" → "20251227"
|
|
40
|
+
2. ISO year-month: "2025-12" → "202512"
|
|
41
|
+
3. d.m.yyyy: "1.3.2025" or "01.03.2025" → "20250301"
|
|
42
|
+
4. d.m.yy: "27.12.25" → "20251227" (2-digit year assumes 2000s)
|
|
43
|
+
5. m.yyyy: "3.2025" or "03.2025" → "202503"
|
|
44
|
+
6. month name + year: "mar 2025", "march 2025", "marec 2025" → "202503"
|
|
45
|
+
7. month names supported in English and Slovenian (full and abbreviated)
|
|
46
|
+
2. preprocessing
|
|
47
|
+
1. before splitting the string, we find possible numbers and replace decimal commas with decimal periods and drop leading zeros
|
|
48
|
+
1. eg "abc 12,3x" -> "abc 12.3x"
|
|
49
|
+
2. eg "abc ,3x" -> "abc .3x"
|
|
50
|
+
3. eg "abc 0,3x" -> "abc .3x"
|
|
51
|
+
4. eg "abc 0.3x" -> "abc .3x"
|
|
52
|
+
5. eg "007" -> "7"
|
|
53
|
+
6. multiple periods stay as-is: "1.2.3" -> "1.2.3"
|
|
54
|
+
2. we remove spaces in pattern digits - digits
|
|
55
|
+
1. eg. "12 - 25" => "12-25"
|
|
56
|
+
3. we insert spaces between starting letters and digits (only at letter-to-digit boundaries)
|
|
57
|
+
1. eg "SI123" -> "SI 123"
|
|
58
|
+
2. eg "123SI" -> "123 SI"
|
|
59
|
+
3. eg "A1B2" -> "A1B2"
|
|
60
|
+
3. sanitation
|
|
61
|
+
1. case-insensitive
|
|
62
|
+
2. diacritics-insensitive (č,ć,c -> same c)
|
|
63
|
+
3. we only keep letters (a-z), digits (0-9), periods (.), minuses (-), spaces (" ") in search
|
|
64
|
+
4. other characters are replaced with space (" ")
|
|
65
|
+
2. tokenization
|
|
66
|
+
1. we split strings on spaces
|
|
67
|
+
2. then we trim all substrings
|
|
68
|
+
3. matching
|
|
69
|
+
1. we match words in any order
|
|
70
|
+
2. we match numbers from start or from end
|
|
71
|
+
1. eg. 123 matches 12345 and 45123 but not 7812345
|
|
72
|
+
4. nested data extraction
|
|
73
|
+
1. nested objects and arrays are automatically extracted recursively
|
|
74
|
+
2. Date objects are converted to ISO strings before tokenization
|
|
75
|
+
3. example: { customer: { firstName: 'John' }, items: [{ opis: 'A' }] } extracts "john", "a"
|
|
76
|
+
5. dot notation in searchInFields and fieldMatchModes
|
|
77
|
+
1. use dot notation to target specific nested fields: 'customer.firstName', 'items.opis'
|
|
78
|
+
2. arrays of objects: 'items.opis' extracts from all array elements
|
|
79
|
+
3. deeply nested: 'data.level1.level2.value'
|
|
80
|
+
4. fieldMatchModes supports dot notation: { 'items.barcode': 'startEnd' }
|
|
81
|
+
6. building metadata
|
|
82
|
+
1. function prepareStringForSearch -> tokens[]
|
|
83
|
+
2. function createSearchArrayMetadata (numeric, default)
|
|
84
|
+
1. all elements in array must have unique _id:string field
|
|
85
|
+
2. prepares input as metadata for each array item
|
|
86
|
+
3. stores tokens as Uint32Array of token IDs
|
|
87
|
+
4. creates temporary tokensMap, clears after use
|
|
88
|
+
5. supports per-field match modes (encoded in token ID high bits)
|
|
89
|
+
7. searching
|
|
90
|
+
1. searchInData / searchInDataReturnIds
|
|
91
|
+
1. prepares query using prepareStringForSearch
|
|
92
|
+
2. matches query tokens against numeric metadata
|
|
93
|
+
3. returns matching IDs
|
|
94
|
+
2. searchInDataReturnObjects
|
|
95
|
+
1. calls searchInDataReturnIds and looks up objects
|
|
96
|
+
8. linked arrays searching
|
|
97
|
+
1. purpose: search across two related arrays (e.g., stranke + pacienti) where secondary items belong to primary items via foreign key
|
|
98
|
+
2. building linked metadata
|
|
99
|
+
1. function createSearchLinkedMetadata
|
|
100
|
+
1. takes primary items + their metadata
|
|
101
|
+
2. takes secondary items + their metadata
|
|
102
|
+
3. takes foreignKeyGetter function to extract primary ID from secondary item
|
|
103
|
+
4. builds reverse index: primary._id -> secondary._id[]
|
|
104
|
+
5. no data duplication - stores references and indexes only
|
|
105
|
+
3. searching linked arrays
|
|
106
|
+
1. function findInLinkedArrays
|
|
107
|
+
1. for each primary item, collects all tokens from primary + all its secondaries
|
|
108
|
+
2. matches query tokens against combined token pool
|
|
109
|
+
3. determines matchedIn: 'primary' | 'secondary' | 'both'
|
|
110
|
+
4. returns only secondaries that contributed to the match (not all secondaries)
|
|
111
|
+
5. tokens that match in primary are excluded from secondary matching to avoid foreign key redundancy
|
|
112
|
+
2. function findInLinkedArraysSimple
|
|
113
|
+
1. calls findInLinkedArrays and returns just the results array
|
|
114
|
+
4. example use case
|
|
115
|
+
1. search "krajnik angie" matches stranka "Krajnik" with pacient "Angie"
|
|
116
|
+
2. returns: { primary: Stranka, secondaries: [matching Pacient], matchedIn: 'both' }
|
|
117
|
+
|
|
118
|
+
## Types
|
|
119
|
+
|
|
120
|
+
1. Id: string
|
|
121
|
+
2. NumericToken: number (index into global tokens array)
|
|
122
|
+
3. NumericTokenSortedList: Uint32Array (sorted numeric token IDs)
|
|
123
|
+
4. NumericSearchMetadata: Map<Id, NumericTokenSortedList> (default SearchMetadata)
|
|
124
|
+
5. StringSearchMetadata: Map<Id, string[]> (legacy, string-based)
|
|
125
|
+
6. LinkedSearchMetadata: { primaryMeta, linkedMeta, primaryToLinked: Map<Id, Id[]> }
|
|
126
|
+
7. SearchableObject: { _id: Id, _deleted?: Date, _blocked?: Date }
|
|
127
|
+
8. SearchableData<T>: Map<Id, T> where T extends SearchableObject
|
|
128
|
+
9. SearchableObjectSpec<T extends SearchableObject, C = any>:
|
|
129
|
+
1. searchInFields?: (keyof T | string)[] - supports dot notation like 'customer.firstName'
|
|
130
|
+
2. dontSearchInFields?: (keyof T)[]
|
|
131
|
+
3. searchContext?: C
|
|
132
|
+
4. extractSearchableStringFn?: (obj: SearchableObject) => string | undefined
|
|
133
|
+
5. enrichSearchObjectFn?: (obj: SearchableObject, searchContext: C) => void
|
|
134
|
+
6. fieldMatchModes?: Partial<Record<keyof T | string, MatchMode>> - per-field match modes
|
|
135
|
+
10. MatchMode: 'start' | 'end' | 'startEnd' | 'anywhere' | 'whole'
|
|
136
|
+
1. Field token prefixes: '<' (start), '>' (end), '+' (startEnd), '*' (anywhere), '!' (whole)
|
|
137
|
+
2. Numeric: encoded in high 3 bits of Uint32 token ID
|
|
138
|
+
11. SearchableLinkedObjectsSpec: primary SearchableObjectSpec, linked SearchableObjectSpec, foreignKeyGetter
|
|
139
|
+
|
|
140
|
+
## Search Query Prefixes
|
|
141
|
+
|
|
142
|
+
Override match behavior per token at search time. Priority: query prefix > field mode > global SearchOpts.
|
|
143
|
+
|
|
144
|
+
### User-friendly prefixes
|
|
145
|
+
| Prefix | Mode | Description | Example |
|
|
146
|
+
|--------|------|-------------|---------|
|
|
147
|
+
| `word..` | start | Match tokens starting with word | `APL..` matches "APL-IP15P" |
|
|
148
|
+
| `..word` | end | Match tokens ending with word | `..Pro` matches "MacBookPro" |
|
|
149
|
+
| `=word` | whole | Exact match only | `=jan` matches "jan" not "jana" |
|
|
150
|
+
| `?word` | anywhere | Match word anywhere in token | `?book` matches "MacBook" |
|
|
151
|
+
| `--word` | negation | Exclude results containing word | `apple --iphone` excludes iPhones |
|
|
152
|
+
| `~word` | negation | Same as `--word` | `apple ~iphone` |
|
|
153
|
+
|
|
154
|
+
### Programmatic prefixes
|
|
155
|
+
| Prefix | Mode |
|
|
156
|
+
|--------|------|
|
|
157
|
+
| `<word` | start |
|
|
158
|
+
| `>word` | end |
|
|
159
|
+
| `+word` | startEnd |
|
|
160
|
+
| `*word` | anywhere |
|
|
161
|
+
| `!word` | whole |
|
|
162
|
+
|
|
163
|
+
### Examples
|
|
164
|
+
1. `"jana useni"` → finds "jana usenik" (default startEnd matching)
|
|
165
|
+
2. `"jana =useni"` → no match (whole match required, "useni" ≠ "usenik")
|
|
166
|
+
3. `"jana --macka"` → finds "jana" but excludes any result containing "macka"
|
|
167
|
+
4. `"APL.."` → finds all tokens starting with "APL" (APL-IP15P, APL-MBP16)
|
|
168
|
+
5. `"..nik"` → finds tokens ending with "nik" (usenik, krajnik)
|
|
169
|
+
|
|
170
|
+
|
|
171
|
+
## Main Functions (Default = Numeric)
|
|
172
|
+
|
|
173
|
+
1. prepareStringForSearch(input: string): string[]
|
|
174
|
+
1. pipeline: normalizeDates → preprocessString → sanitiseString → tokenize → sort
|
|
175
|
+
2. createSearchArrayMetadata(data: SearchableData, spec?): NumericSearchMetadata
|
|
176
|
+
1. creates temporary tokensMap for fast string→number lookups
|
|
177
|
+
2. stores tokens as Uint32Array per item
|
|
178
|
+
3. clears tokensMap after use (only global tokens[] remains)
|
|
179
|
+
4. for _deleted or _blocked nothing is added
|
|
180
|
+
3. prepareObjectForSearch(obj, spec?, tokensMap): NumericTokenSortedList | undefined
|
|
181
|
+
1. returns undefined if _deleted or _blocked
|
|
182
|
+
2. if custom extractSearchableStringFn provided, uses simple processing
|
|
183
|
+
3. otherwise processes fields individually to support per-field match modes
|
|
184
|
+
4. recursively extracts strings from nested objects/arrays
|
|
185
|
+
4. searchInData / searchInDataReturnIds(query: string, metadata): Id[]
|
|
186
|
+
5. searchInDataReturnObjects<T>(query, metadata, data): Map<Id, T>
|
|
187
|
+
6. searchInDataWithLimit(query, metadata, limit): Id[]
|
|
188
|
+
7. updateSearchMetadata(metadata, objectId, obj, spec)
|
|
189
|
+
1. updates or deletes metadata for this object
|
|
190
|
+
2. taking into account _blocked and _deleted
|
|
191
|
+
8. updateSearchMetadataBatch(metadata, batch, spec)
|
|
192
|
+
1. creates temporary tokensMap, clears after use
|
|
193
|
+
2. updates or deletes metadata for batch of objects
|
|
194
|
+
|
|
195
|
+
## Token Registry Functions
|
|
196
|
+
|
|
197
|
+
1. tokens: string[] - global array (read-only)
|
|
198
|
+
2. getTokenCount(): number
|
|
199
|
+
3. getTokenById(id: number): string | undefined
|
|
200
|
+
4. findTokenId(token: string): number (-1 if not found)
|
|
201
|
+
5. registerToken(token: string): number
|
|
202
|
+
6. getOrCreateTokenId(token: string, tokensMap): number
|
|
203
|
+
7. createTokensMap(): Map<string, number>
|
|
204
|
+
8. clearTokensMap(tokensMap): void
|
|
205
|
+
9. resetTokenRegistry(): void (for testing only)
|
|
206
|
+
|
|
207
|
+
## String Implementation (Legacy)
|
|
208
|
+
|
|
209
|
+
Available with `*String` suffix:
|
|
210
|
+
- createSearchArrayMetadataString
|
|
211
|
+
- createSearchLinkedMetadataString
|
|
212
|
+
- updateSearchMetadataString
|
|
213
|
+
- updateSearchMetadataBatchString
|
|
214
|
+
- findInArray (uses string metadata)
|
|
215
|
+
- findInArrayReturnDataAndMeta
|
|
216
|
+
- findInLinkedArrays
|
|
217
|
+
|
|
218
|
+
## File Structure
|
|
219
|
+
|
|
220
|
+
```
|
|
221
|
+
src/
|
|
222
|
+
├── common/
|
|
223
|
+
│ ├── numeric/ # Default numeric implementation
|
|
224
|
+
│ │ ├── index.ts
|
|
225
|
+
│ │ ├── createSearchArrayMetadata.ts
|
|
226
|
+
│ │ ├── createSearchLinkedMetadata.ts
|
|
227
|
+
│ │ ├── updateSearchMetadata.ts
|
|
228
|
+
│ │ └── searchInData.ts
|
|
229
|
+
│ ├── string/ # Legacy string implementation
|
|
230
|
+
│ │ ├── index.ts
|
|
231
|
+
│ │ ├── createSearchArrayMetadata.ts
|
|
232
|
+
│ │ ├── createSearchLinkedMetadata.ts
|
|
233
|
+
│ │ └── updateSearchMetadata.ts
|
|
234
|
+
│ ├── findInArray.ts # String-based search
|
|
235
|
+
│ ├── findInLinkedArrays.ts # String-based linked search
|
|
236
|
+
│ ├── syncSearchArrayMetadata.ts
|
|
237
|
+
│ ├── updateSearchLinkedMetadata.ts
|
|
238
|
+
│ └── SearchUniverse.ts # Multi-collection manager (uses numeric)
|
|
239
|
+
├── utils/
|
|
240
|
+
│ ├── tokenRegistry.ts # Global tokens array
|
|
241
|
+
│ ├── prepareStringForSearch.ts
|
|
242
|
+
│ ├── prepareStringForSearchNumeric.ts
|
|
243
|
+
│ ├── matchToken.ts # String-based matching
|
|
244
|
+
│ ├── matchTokenNumeric.ts # Numeric matching
|
|
245
|
+
│ └── ...
|
|
246
|
+
├── types/
|
|
247
|
+
│ └── index.ts
|
|
248
|
+
└── index.ts # Main exports (numeric as default)
|
|
249
|
+
```
|
|
250
|
+
|
|
251
|
+
## Coding rules
|
|
252
|
+
|
|
253
|
+
1. each function is in separate file
|
|
254
|
+
1. preprocessString
|
|
255
|
+
2. sanitiseString
|
|
256
|
+
3. tokenize
|
|
257
|
+
4. matchToken / matchTokenNumeric
|
|
258
|
+
5. searchInData
|
|
259
|
+
|
|
260
|
+
## Testing
|
|
261
|
+
|
|
262
|
+
1. Large real-life datasets are in ./test-data (arikli, stranke, pacienti).
|
|
263
|
+
2. tests are in ./test
|
|
264
|
+
3. tests per implementation are in ./test/(IMPLEMENTATION)
|
|
265
|
+
4. We have interactive.ts that
|
|
266
|
+
1. loads all three datasets and creates SearchMetadata for each
|
|
267
|
+
2. measures memory consumption for each loaded dataset and for created SearchMetadata
|
|
268
|
+
3. lets user enter search string in a loop and presents
|
|
269
|
+
1. 3 results from each dataset (where found)
|
|
270
|
+
1. no fieldnames, just values
|
|
271
|
+
2. prints time spent searching in each dataset
|
|
272
|
+
4. Supports command lines for changing data
|
|
273
|
+
1. 621771cadf380d819d0b0970: Angie -> Ančka
|
|
274
|
+
1. find object with _id===621771cadf380d819d0b0970
|
|
275
|
+
2. find key with value "Angie"
|
|
276
|
+
3. change value to Ančka
|
|
277
|
+
2. 621771cadf380d819d0b0970: x siamska -> siamska
|
|
278
|
+
3. 621771cadf380d819d0b0970: _blocked -> true/false
|
|
279
|
+
1. sets _blocked to new Date() or deletes the field
|
|
280
|
+
4. 621771cadf380d819d0b0970: _deleted -> true/false
|
|
281
|
+
1. like blocked
|
|
282
|
+
5. benchmark-numeric-vs-string.ts compares implementations
|
|
283
|
+
1. Memory: numeric ~45% less
|
|
284
|
+
2. Build time: ~same
|
|
285
|
+
3. Search time: varies by match mode
|
package/LICENSE.md
ADDED
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
# Proprietary License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2025. All Rights Reserved.
|
|
4
|
+
|
|
5
|
+
## Terms and Conditions
|
|
6
|
+
|
|
7
|
+
### 1. Definitions
|
|
8
|
+
|
|
9
|
+
- "Software" refers to the cry-search library, including all source code, documentation, and associated files.
|
|
10
|
+
- "Licensor" refers to the copyright holder(s) of this Software.
|
|
11
|
+
- "You" refers to any individual or entity that accesses, uses, or attempts to use this Software.
|
|
12
|
+
|
|
13
|
+
### 2. Grant of License
|
|
14
|
+
|
|
15
|
+
Subject to the terms of this License, the Licensor grants You a limited, non-exclusive, non-transferable, revocable license to use this Software **solely for personal, non-commercial, evaluation purposes**.
|
|
16
|
+
|
|
17
|
+
### 3. Restrictions
|
|
18
|
+
|
|
19
|
+
You are expressly **PROHIBITED** from:
|
|
20
|
+
|
|
21
|
+
1. **Commercial Use**: Using the Software, in whole or in part, for any commercial purpose, including but not limited to:
|
|
22
|
+
- Incorporating the Software into commercial products or services
|
|
23
|
+
- Using the Software to provide services to third parties for compensation
|
|
24
|
+
- Using the Software in any revenue-generating activity
|
|
25
|
+
- Using the Software in a business or organizational context
|
|
26
|
+
|
|
27
|
+
2. **Modification**: Modifying, adapting, translating, reverse engineering, decompiling, disassembling, or creating derivative works based on the Software.
|
|
28
|
+
|
|
29
|
+
3. **Distribution**: Distributing, sublicensing, leasing, renting, lending, selling, reselling, or otherwise transferring the Software or any rights therein to any third party.
|
|
30
|
+
|
|
31
|
+
4. **Sharing**: Forwarding, sharing, publishing, uploading, or making the Software available to any third party through any means, including but not limited to:
|
|
32
|
+
- Public or private repositories
|
|
33
|
+
- File sharing services
|
|
34
|
+
- Email or messaging platforms
|
|
35
|
+
- Any form of electronic or physical distribution
|
|
36
|
+
|
|
37
|
+
5. **Copying**: Copying or reproducing the Software except as strictly necessary for personal evaluation use on a single machine.
|
|
38
|
+
|
|
39
|
+
6. **Removal of Notices**: Removing, altering, or obscuring any copyright, trademark, or other proprietary notices contained in the Software.
|
|
40
|
+
|
|
41
|
+
### 4. Intellectual Property
|
|
42
|
+
|
|
43
|
+
The Software and all copies thereof are proprietary to the Licensor and title thereto remains exclusively with the Licensor. All rights in the Software not specifically granted in this License are reserved to the Licensor.
|
|
44
|
+
|
|
45
|
+
### 5. No Warranty
|
|
46
|
+
|
|
47
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, AND NONINFRINGEMENT. IN NO EVENT SHALL THE LICENSOR BE LIABLE FOR ANY CLAIM, DAMAGES, OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT, OR OTHERWISE, ARISING FROM, OUT OF, OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
|
48
|
+
|
|
49
|
+
### 6. Termination
|
|
50
|
+
|
|
51
|
+
This License is effective until terminated. Your rights under this License will terminate automatically without notice if You fail to comply with any of its terms. Upon termination, You must immediately cease all use of the Software and destroy all copies in Your possession or control.
|
|
52
|
+
|
|
53
|
+
### 7. Governing Law
|
|
54
|
+
|
|
55
|
+
This License shall be governed by and construed in accordance with applicable laws, without regard to conflicts of law principles.
|
|
56
|
+
|
|
57
|
+
### 8. Entire Agreement
|
|
58
|
+
|
|
59
|
+
This License constitutes the entire agreement between You and the Licensor concerning the Software and supersedes all prior or contemporaneous oral or written communications, proposals, and representations with respect to the Software.
|
|
60
|
+
|
|
61
|
+
### 9. Contact
|
|
62
|
+
|
|
63
|
+
For licensing inquiries, including requests for commercial licenses, please contact the copyright holder.
|
|
64
|
+
|
|
65
|
+
---
|
|
66
|
+
|
|
67
|
+
**BY USING THIS SOFTWARE, YOU ACKNOWLEDGE THAT YOU HAVE READ THIS LICENSE, UNDERSTAND IT, AND AGREE TO BE BOUND BY ITS TERMS AND CONDITIONS.**
|
package/README.md
ADDED
|
@@ -0,0 +1,283 @@
|
|
|
1
|
+
# cry-search
|
|
2
|
+
|
|
3
|
+
A fast, memory-efficient search library for large datasets with support for tokenized matching, linked collections, and per-field match modes.
|
|
4
|
+
|
|
5
|
+
## Features
|
|
6
|
+
|
|
7
|
+
- **Fast search** - Pre-built metadata enables instant queries on large datasets
|
|
8
|
+
- **Memory efficient** - Numeric token storage uses ~45% less memory than string-based approaches
|
|
9
|
+
- **Updatable** - Add, update, or remove items without rebuilding metadata
|
|
10
|
+
- **Flexible matching** - Per-field match modes and query prefixes for prefix, suffix, anywhere, exact, and negation
|
|
11
|
+
- **Search prefixed** - use `--word` to exclude it, `...word` for suffix, and more
|
|
12
|
+
- **Automatic normalization** - Handles diacritics, case, dates, and number formatting
|
|
13
|
+
- **[Universal search](./UNIVERSE.md)** - Search and update across linked collections (e.g., customers with their pets)
|
|
14
|
+
|
|
15
|
+
## Table of Contents
|
|
16
|
+
|
|
17
|
+
- [cry-search](#cry-search)
|
|
18
|
+
- [Features](#features)
|
|
19
|
+
- [Table of Contents](#table-of-contents)
|
|
20
|
+
- [Installation](#installation)
|
|
21
|
+
- [Examples](#examples)
|
|
22
|
+
- [Single Table Search](#single-table-search)
|
|
23
|
+
- [SearchUniverse with Linked Collections](#searchuniverse-with-linked-collections)
|
|
24
|
+
- [Updating Existing Data](#updating-existing-data)
|
|
25
|
+
- [Architecture](#architecture)
|
|
26
|
+
- [Text Processing Pipeline](#text-processing-pipeline)
|
|
27
|
+
- [Metadata Building](#metadata-building)
|
|
28
|
+
- [Searching](#searching)
|
|
29
|
+
- [Memory Optimization](#memory-optimization)
|
|
30
|
+
- [Query Prefixes](#query-prefixes)
|
|
31
|
+
- [Legacy String Implementation](#legacy-string-implementation)
|
|
32
|
+
- [Specification](#specification)
|
|
33
|
+
- [License](#license)
|
|
34
|
+
|
|
35
|
+
## Installation
|
|
36
|
+
|
|
37
|
+
```bash
|
|
38
|
+
npm install cry-search
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
## Examples
|
|
42
|
+
|
|
43
|
+
### Single Table Search
|
|
44
|
+
|
|
45
|
+
```typescript
|
|
46
|
+
import {
|
|
47
|
+
createSearchArrayMetadata,
|
|
48
|
+
searchInDataReturnObjects,
|
|
49
|
+
arrayToSearchableData,
|
|
50
|
+
type SearchableObject,
|
|
51
|
+
} from 'cry-search';
|
|
52
|
+
|
|
53
|
+
interface Product extends SearchableObject {
|
|
54
|
+
_id: string;
|
|
55
|
+
name: string;
|
|
56
|
+
description: string;
|
|
57
|
+
sku: string;
|
|
58
|
+
}
|
|
59
|
+
|
|
60
|
+
const products: Product[] = [
|
|
61
|
+
{ _id: '1', name: 'iPhone 15 Pro', description: 'Latest Apple smartphone', sku: 'APL-IP15P' },
|
|
62
|
+
{ _id: '2', name: 'Samsung Galaxy S24', description: 'Android flagship phone', sku: 'SAM-GS24' },
|
|
63
|
+
{ _id: '3', name: 'MacBook Pro 16"', description: 'Apple laptop for professionals', sku: 'APL-MBP16' },
|
|
64
|
+
];
|
|
65
|
+
|
|
66
|
+
// Convert array to searchable data (Map<id, item>)
|
|
67
|
+
const data = arrayToSearchableData(products);
|
|
68
|
+
|
|
69
|
+
// Build search metadata (do this once, reuse for all searches)
|
|
70
|
+
const metadata = createSearchArrayMetadata(data, {
|
|
71
|
+
searchInFields: ['name', 'description', 'sku'],
|
|
72
|
+
});
|
|
73
|
+
|
|
74
|
+
// Search
|
|
75
|
+
const results = searchInDataReturnObjects('apple', metadata, data);
|
|
76
|
+
// Returns: Map with products 1 and 3
|
|
77
|
+
|
|
78
|
+
const results2 = searchInDataReturnObjects('samsung galaxy', metadata, data);
|
|
79
|
+
// Returns: Map with product 2
|
|
80
|
+
|
|
81
|
+
// Search is diacritics and case insensitive
|
|
82
|
+
const results3 = searchInDataReturnObjects('IPHONE', metadata, data);
|
|
83
|
+
// Returns: Map with product 1
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
### SearchUniverse with Linked Collections
|
|
87
|
+
|
|
88
|
+
For related data (like customers and their pets), use `SearchUniverse` to search across collections:
|
|
89
|
+
|
|
90
|
+
```typescript
|
|
91
|
+
import { SearchUniverse, type SearchableObject } from 'cry-search';
|
|
92
|
+
|
|
93
|
+
interface Stranka extends SearchableObject {
|
|
94
|
+
_id: string;
|
|
95
|
+
name: string;
|
|
96
|
+
address: string;
|
|
97
|
+
}
|
|
98
|
+
|
|
99
|
+
interface Pacient extends SearchableObject {
|
|
100
|
+
_id: string;
|
|
101
|
+
stranka_id: string; // Foreign key to Stranka
|
|
102
|
+
name: string;
|
|
103
|
+
species: string;
|
|
104
|
+
}
|
|
105
|
+
|
|
106
|
+
// Create universe and register collections
|
|
107
|
+
const universe = new SearchUniverse();
|
|
108
|
+
|
|
109
|
+
universe.registerCollection<Stranka>('stranke', {
|
|
110
|
+
spec: { searchInFields: ['name', 'address'] },
|
|
111
|
+
});
|
|
112
|
+
|
|
113
|
+
universe.registerCollection<Pacient>('pacienti', {
|
|
114
|
+
spec: { dontSearchInFields: ['stranka_id'] },
|
|
115
|
+
linkedTo: {
|
|
116
|
+
collectionName: 'stranke',
|
|
117
|
+
foreignKeyGetter: (p) => p.stranka_id,
|
|
118
|
+
},
|
|
119
|
+
});
|
|
120
|
+
|
|
121
|
+
// Load data
|
|
122
|
+
universe.loadCollection('stranke', [
|
|
123
|
+
{ _id: 's1', name: 'Krajnik', address: 'Ljubljana' },
|
|
124
|
+
{ _id: 's2', name: 'Novak', address: 'Maribor' },
|
|
125
|
+
]);
|
|
126
|
+
|
|
127
|
+
universe.loadCollection('pacienti', [
|
|
128
|
+
{ _id: 'p1', stranka_id: 's1', name: 'Angie', species: 'cat' },
|
|
129
|
+
{ _id: 'p2', stranka_id: 's1', name: 'Rex', species: 'dog' },
|
|
130
|
+
{ _id: 'p3', stranka_id: 's2', name: 'Bella', species: 'cat' },
|
|
131
|
+
]);
|
|
132
|
+
|
|
133
|
+
// Search across all collections
|
|
134
|
+
const results = universe.searchUniversally('krajnik angie');
|
|
135
|
+
// Returns:
|
|
136
|
+
// {
|
|
137
|
+
// stranke: [{ _id: 's1', name: 'Krajnik', ... }],
|
|
138
|
+
// pacienti: [{
|
|
139
|
+
// primary: { _id: 's1', name: 'Krajnik', ... },
|
|
140
|
+
// linked: [{ _id: 'p1', name: 'Angie', ... }],
|
|
141
|
+
// matchedIn: 'both'
|
|
142
|
+
// }]
|
|
143
|
+
// }
|
|
144
|
+
|
|
145
|
+
// Search only in linked items
|
|
146
|
+
const results2 = universe.searchUniversally('rex');
|
|
147
|
+
// Returns pacienti results with matchedIn: 'linked'
|
|
148
|
+
|
|
149
|
+
// Search only in primary - returns all linked items
|
|
150
|
+
const results3 = universe.searchUniversally('novak');
|
|
151
|
+
// Returns Novak stranka with Bella in linked array
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
### Updating Existing Data
|
|
155
|
+
|
|
156
|
+
When data changes, update the metadata to keep search in sync:
|
|
157
|
+
|
|
158
|
+
```typescript
|
|
159
|
+
import {
|
|
160
|
+
createSearchArrayMetadata,
|
|
161
|
+
updateSearchMetadata,
|
|
162
|
+
updateSearchMetadataBatch,
|
|
163
|
+
arrayToSearchableData,
|
|
164
|
+
} from 'cry-search';
|
|
165
|
+
|
|
166
|
+
// Initial setup
|
|
167
|
+
const data = arrayToSearchableData(products);
|
|
168
|
+
const metadata = createSearchArrayMetadata(data);
|
|
169
|
+
|
|
170
|
+
// Update a single item
|
|
171
|
+
const updatedProduct = { ...products[0], name: 'iPhone 16 Pro' };
|
|
172
|
+
data.set(updatedProduct._id, updatedProduct);
|
|
173
|
+
updateSearchMetadata(metadata, updatedProduct._id, updatedProduct);
|
|
174
|
+
|
|
175
|
+
// Update multiple items at once
|
|
176
|
+
const updates = [
|
|
177
|
+
{ _id: '1', name: 'iPhone 16 Pro Max', description: 'Newest Apple phone', sku: 'APL-IP16PM' },
|
|
178
|
+
{ _id: '4', name: 'iPad Pro', description: 'Apple tablet', sku: 'APL-IPAD' },
|
|
179
|
+
];
|
|
180
|
+
for (const item of updates) {
|
|
181
|
+
data.set(item._id, item);
|
|
182
|
+
}
|
|
183
|
+
updateSearchMetadataBatch(metadata, updates);
|
|
184
|
+
|
|
185
|
+
// Mark item as deleted (removes from search but keeps in data)
|
|
186
|
+
const deletedProduct = { ...products[1], _deleted: new Date() };
|
|
187
|
+
data.set(deletedProduct._id, deletedProduct);
|
|
188
|
+
updateSearchMetadata(metadata, deletedProduct._id, deletedProduct);
|
|
189
|
+
|
|
190
|
+
// With SearchUniverse - handles linked indexes automatically
|
|
191
|
+
universe.updateSearchMetadata('pacienti', 'p1', {
|
|
192
|
+
_id: 'p1',
|
|
193
|
+
stranka_id: 's2', // Changed owner from s1 to s2
|
|
194
|
+
name: 'Angie',
|
|
195
|
+
species: 'cat',
|
|
196
|
+
});
|
|
197
|
+
```
|
|
198
|
+
|
|
199
|
+
## Architecture
|
|
200
|
+
|
|
201
|
+
cry-search uses a two-phase approach: **build** metadata once, then **search** instantly.
|
|
202
|
+
|
|
203
|
+
### Text Processing Pipeline
|
|
204
|
+
|
|
205
|
+
Both metadata building and search queries go through the same pipeline:
|
|
206
|
+
|
|
207
|
+
```
|
|
208
|
+
Input: "Čokolada 27.12.2025 SI123"
|
|
209
|
+
↓
|
|
210
|
+
1. Normalize dates → "Čokolada 20251227 SI123"
|
|
211
|
+
2. Preprocess → "Čokolada 20251227 SI 123" (split letter-digit)
|
|
212
|
+
3. Sanitize → "cokolada 20251227 si 123" (lowercase, remove diacritics)
|
|
213
|
+
4. Tokenize → ["cokolada", "20251227", "si", "123"]
|
|
214
|
+
5. Sort → ["123", "20251227", "cokolada", "si"]
|
|
215
|
+
6. Numeric IDs → Uint32Array [42, 891, 156, 7] (indices into global registry)
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
### Metadata Building
|
|
219
|
+
|
|
220
|
+
Tokens are stored as numeric IDs in a global registry. Each item's tokens are stored in a sorted `Uint32Array` for memory efficiency and fast binary search.
|
|
221
|
+
|
|
222
|
+
### Searching
|
|
223
|
+
|
|
224
|
+
Query tokens are matched against metadata using binary search. All query tokens must match for an item to be returned. Match modes (start, end, startEnd, anywhere, whole) control how tokens are compared.
|
|
225
|
+
|
|
226
|
+
### Memory Optimization
|
|
227
|
+
|
|
228
|
+
The numeric implementation stores tokens as `Uint32Array` indices instead of string arrays, reducing memory usage by ~45% compared to the string-based implementation.
|
|
229
|
+
|
|
230
|
+
## Query Prefixes
|
|
231
|
+
|
|
232
|
+
Override match behavior per token at search time using prefixes:
|
|
233
|
+
|
|
234
|
+
| Prefix | Mode | Description |
|
|
235
|
+
|--------|------|-------------|
|
|
236
|
+
| `word..` | start | Match tokens starting with "word" |
|
|
237
|
+
| `..word` | end | Match tokens ending with "word" |
|
|
238
|
+
| `=word` | whole | Exact match only |
|
|
239
|
+
| `?word` | anywhere | Match "word" anywhere in token |
|
|
240
|
+
| `--word` or `~word` | negation | Exclude results containing "word" |
|
|
241
|
+
|
|
242
|
+
Alternative prefixes (for programmatic use): `<word` (start), `>word` (end), `+word` (startEnd), `*word` (anywhere), `!word` (whole)
|
|
243
|
+
|
|
244
|
+
```typescript
|
|
245
|
+
// Find "jana" but only exact match on "usenik"
|
|
246
|
+
searchInDataReturnObjects('jana =usenik', metadata, data);
|
|
247
|
+
// Matches "Jana Usenik" but not "Jana Useniker"
|
|
248
|
+
|
|
249
|
+
// Find "apple" but exclude results containing "iphone"
|
|
250
|
+
searchInDataReturnObjects('apple --iphone', metadata, data);
|
|
251
|
+
// Matches MacBook Pro but not iPhone
|
|
252
|
+
|
|
253
|
+
// Find products with SKU starting with "APL"
|
|
254
|
+
searchInDataReturnObjects('APL..', metadata, data);
|
|
255
|
+
// Matches APL-IP15P, APL-MBP16, etc.
|
|
256
|
+
```
|
|
257
|
+
|
|
258
|
+
Priority: query prefix > field match mode > global SearchOpts default.
|
|
259
|
+
|
|
260
|
+
## Legacy String Implementation
|
|
261
|
+
|
|
262
|
+
A string-based implementation is available for backwards compatibility. Import with the `String` suffix:
|
|
263
|
+
|
|
264
|
+
```typescript
|
|
265
|
+
import {
|
|
266
|
+
createSearchArrayMetadataString,
|
|
267
|
+
updateSearchMetadataString,
|
|
268
|
+
findInArray, // String-based search function
|
|
269
|
+
} from 'cry-search';
|
|
270
|
+
```
|
|
271
|
+
|
|
272
|
+
## Specification
|
|
273
|
+
|
|
274
|
+
See [CLAUDE.md](./CLAUDE.md) for the complete technical specification including:
|
|
275
|
+
- Input processing pipeline (date normalization, preprocessing, sanitization)
|
|
276
|
+
- Tokenization rules
|
|
277
|
+
- Matching behavior
|
|
278
|
+
- Linked search semantics
|
|
279
|
+
- Type definitions
|
|
280
|
+
|
|
281
|
+
## License
|
|
282
|
+
|
|
283
|
+
See [LICENSE.md](./LICENSE.md) for license terms.
|