@equationalapplications/core-llm-wiki 5.5.0 → 5.5.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +133 -0
- package/dist/{chunk-QVY5DFJV.mjs → chunk-JQRA6BYQ.mjs} +79 -12
- package/dist/chunk-JQRA6BYQ.mjs.map +1 -0
- package/dist/index.js +77 -10
- package/dist/index.js.map +1 -1
- package/dist/index.mjs +2 -2
- package/dist/testing.js +77 -10
- package/dist/testing.js.map +1 -1
- package/dist/testing.mjs +1 -1
- package/package.json +2 -2
- package/dist/chunk-QVY5DFJV.mjs.map +0 -1
package/README.md
CHANGED
|
@@ -52,6 +52,45 @@ Platform-agnostic TypeScript engine for hybrid LLM memory. Features episodic fac
|
|
|
52
52
|
- **Interoperability:** Supports [Open Knowledge Format (OKF) v0.1](https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/okf) import and export via the [llm-wiki OKF profile (llm-wiki/1)](https://github.com/equationalapplications/expo-llm-wiki/blob/main/docs/okf-profile.md).
|
|
53
53
|
- **Per-entity seeded ontology** — Optional Strict, Emergent, or Off modes govern LLM graph extraction; seed taxonomies per entity and persist typed facts with inline edges.
|
|
54
54
|
|
|
55
|
+
## GraphRAG & Multi-Modal Retrieval
|
|
56
|
+
|
|
57
|
+
`@equationalapplications/core-llm-wiki` exposes three complementary retrieval modes, each addressing a different shape of query:
|
|
58
|
+
|
|
59
|
+
| Mode | API | Best for |
|
|
60
|
+
|---|---|---|
|
|
61
|
+
| **Semantic** (vector cosine) | `wiki.read(entityId, query)` with `embed` configured | Open-ended natural-language questions; "what do I know about X" |
|
|
62
|
+
| **Keyword** (MiniSearch) | `wiki.read(entityId, query)` with `embed` absent or offline | Exact terms, identifiers, names; offline fallback |
|
|
63
|
+
| **GraphRAG** (recursive CTE) | `wiki.traverseGraph(entityId, options)` + `formatGraphContext(result)` | Structural questions; "what connects to X", "everything two hops from this fact", "summarise the people, places, and projects linked to Alice" |
|
|
64
|
+
|
|
65
|
+
The GraphRAG path is structurally distinct: it doesn't rank by relevance to a query string, it walks `llm_wiki_edges` from a known anchor fact. The result is dense and connected — subgraphs, not loose top-K hits.
|
|
66
|
+
|
|
67
|
+
### Graph traversal APIs
|
|
68
|
+
|
|
69
|
+
```typescript
|
|
70
|
+
import { WikiMemory, formatGraphContext } from '@equationalapplications/core-llm-wiki';
|
|
71
|
+
|
|
72
|
+
const graph = await wikiMemory.traverseGraph('user-123', {
|
|
73
|
+
sourceId: '<anchor-fact-id>',
|
|
74
|
+
maxDepth: 2,
|
|
75
|
+
direction: 'both', // 'inbound' | 'outbound' | 'both'
|
|
76
|
+
edgeTypes: ['reports_to'], // optional filter
|
|
77
|
+
excludeSourceTypes: ['immutable_document'],
|
|
78
|
+
minTraversalConfidence: 'inferred',
|
|
79
|
+
maxTraversalNodes: 20,
|
|
80
|
+
});
|
|
81
|
+
|
|
82
|
+
const promptContext = formatGraphContext(graph);
|
|
83
|
+
// → dense text block ready for prompt injection
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
`traverseGraph` runs as a single recursive CTE in SQLite (see the root README's ["The SQL: how traversal works in one query"](../../README.md#the-sql-how-traversal-works-in-one-query) for the query shape). No external graph database.
|
|
87
|
+
|
|
88
|
+
### Deterministic graph seeding (no LLM)
|
|
89
|
+
|
|
90
|
+
For programmatic pipelines — importing pre-classified data, building a GraphRAG corpus from a CSV, or seed-loading from a JSON file — use `upsertGraph()`. It writes nodes and edges directly under the same `(sourceRef, sourceHash)` ownership semantics as `ingestDocument()`, but skips the LLM extraction step. See [Direct Graph Write](#direct-graph-write) for the canonical signature, including the required `SQLiteAdapter` argument and transactional semantics.
|
|
91
|
+
|
|
92
|
+
This is the GraphRAG seed path: load a corpus, walk it.
|
|
93
|
+
|
|
55
94
|
## Installation
|
|
56
95
|
|
|
57
96
|
```bash
|
|
@@ -827,6 +866,100 @@ configureRandomSource(getRandomValues);
|
|
|
827
866
|
|
|
828
867
|
`@equationalapplications/expo-llm-wiki` does this automatically on import (main entry and `/factory` subpath). If you use `@equationalapplications/core-llm-wiki` directly on React Native without the expo package, you must call `configureRandomSource()` yourself or polyfill `globalThis.crypto.getRandomValues`.
|
|
829
868
|
|
|
869
|
+
## Entity Enumeration
|
|
870
|
+
|
|
871
|
+
List all entities that have stored data in the wiki:
|
|
872
|
+
|
|
873
|
+
```typescript
|
|
874
|
+
const entityIds = await wikiMemory.listEntityIds();
|
|
875
|
+
// Returns all entity_ids with at least one row (including soft-deleted-only entities)
|
|
876
|
+
// Optional prefix filter: await wikiMemory.listEntityIds({ prefix: 'tier_' });
|
|
877
|
+
```
|
|
878
|
+
|
|
879
|
+
Use this for maintenance scheduling, multi-entity operations, or discovering which namespaces exist. Includes entities with only soft-deleted rows so `runPrune()` can reclaim orphaned storage.
|
|
880
|
+
|
|
881
|
+
## Source Reference Enumeration
|
|
882
|
+
|
|
883
|
+
List all documents currently stored for an entity:
|
|
884
|
+
|
|
885
|
+
```typescript
|
|
886
|
+
const sourceRefs = await wikiMemory.listSourceRefs('user-123');
|
|
887
|
+
// One row per live sourceRef (soft-deleted rows are excluded):
|
|
888
|
+
// Array<{ sourceRef: string; sourceHash: string | null; factCount: number; lastIngestedAt: number }>
|
|
889
|
+
// factCount — number of live facts under that sourceRef
|
|
890
|
+
// lastIngestedAt — Unix timestamp in ms from the most recently updated live entry
|
|
891
|
+
```
|
|
892
|
+
|
|
893
|
+
Use this to audit stored documents, validate external sync state, or preview the blast radius before `forget()` operations.
|
|
894
|
+
|
|
895
|
+
## Direct Graph Write
|
|
896
|
+
|
|
897
|
+
Write structured graph data directly without LLM extraction — useful for programmatic fact ingestion, parsers, and deterministic pipelines:
|
|
898
|
+
|
|
899
|
+
```typescript
|
|
900
|
+
const { nodesWritten, edgesWritten, superseded } = await wikiMemory.upsertGraph('entity-123', {
|
|
901
|
+
sourceRef: 'codebase_main.ts',
|
|
902
|
+
sourceHash: sha256(sourceCode),
|
|
903
|
+
nodes: [
|
|
904
|
+
{ id: 'fn_processData', type: 'function', title: 'processData', body: 'Processes user data' },
|
|
905
|
+
{ id: 'class_UserService', type: 'class', title: 'UserService', body: 'User management service' },
|
|
906
|
+
],
|
|
907
|
+
edges: [
|
|
908
|
+
{ type: 'calls', sourceId: 'fn_processData', targetId: 'class_UserService' },
|
|
909
|
+
],
|
|
910
|
+
}, adapter); // SQLiteAdapter from your platform driver — writes join the caller's transaction
|
|
911
|
+
```
|
|
912
|
+
|
|
913
|
+
`upsertGraph` is "the tail of `ingestDocument` with the middle (LLM extraction) step removed" — it accepts caller-supplied nodes (`{ id, type, title, body? }`) and edges (`{ type, sourceId, targetId, id? }`) and writes them under the same `(sourceRef, sourceHash)` semantics. If a *different* live `sourceRef` already holds the same `sourceHash`, it throws `WikiSourceRefHashCollision`; re-writing the identical `(sourceRef, sourceHash)` is a no-op returning zero counts. The adapter parameter is required so writes participate in the caller's transaction.
|
|
914
|
+
|
|
915
|
+
## Duplicate Hash Detection
|
|
916
|
+
|
|
917
|
+
Control behavior when a different live `sourceRef` already holds the same `sourceHash`. The option is passed as a third argument to `ingestDocument`, not inside the params object:
|
|
918
|
+
|
|
919
|
+
```typescript
|
|
920
|
+
await wikiMemory.ingestDocument(
|
|
921
|
+
'entity-123',
|
|
922
|
+
{
|
|
923
|
+
sourceRef: 'doc.md',
|
|
924
|
+
sourceHash: sha256(content),
|
|
925
|
+
documentChunk: content,
|
|
926
|
+
},
|
|
927
|
+
{ onDuplicateHash: 'ingest' } // 'ingest' (default) | 'skip' | 'throw'
|
|
928
|
+
);
|
|
929
|
+
```
|
|
930
|
+
|
|
931
|
+
- `'ingest'` (default): No duplicate pre-check; extraction proceeds as before this option existed. If a different live `sourceRef` is found holding the same hash at commit time (a concurrent-writer race caught by the source-ref unique index), the call still throws `WikiDuplicateHashError`.
|
|
932
|
+
- `'skip'`: Pre-check before any LLM call; if a different live `sourceRef` already holds the hash, return a zero-chunk result without writing.
|
|
933
|
+
- `'throw'`: Pre-check before any LLM call; throw `WikiDuplicateHashError` (carries the canonical `sourceRef`).
|
|
934
|
+
- The guard only considers **live** references — soft-deleted refs do not trigger it in any mode.
|
|
935
|
+
|
|
936
|
+
## Batch Change Detection
|
|
937
|
+
|
|
938
|
+
Check multiple documents for changes in one call:
|
|
939
|
+
|
|
940
|
+
```typescript
|
|
941
|
+
const batch = [
|
|
942
|
+
{ sourceRef: 'doc1.md', sourceHash: sha256(content1) },
|
|
943
|
+
{ sourceRef: 'doc2.md', sourceHash: sha256(content2) },
|
|
944
|
+
{ sourceRef: 'doc3.md', sourceHash: sha256(content3) },
|
|
945
|
+
];
|
|
946
|
+
const changes = await wikiMemory.hasChanged('entity-123', batch);
|
|
947
|
+
// changes: Array<{ sourceRef: string; changed: boolean; duplicateOf?: string }>
|
|
948
|
+
// duplicateOf — when present, the canonical stored different sourceRef holding
|
|
949
|
+
// the same hash (DB-normalized spelling; sourceRef echoes the raw caller value).
|
|
950
|
+
// Per-document change detection; internally batched across queries
|
|
951
|
+
```
|
|
952
|
+
|
|
953
|
+
## Dry-Run Deletion
|
|
954
|
+
|
|
955
|
+
Preview deletion impact without writing:
|
|
956
|
+
|
|
957
|
+
```typescript
|
|
958
|
+
const preview = await wikiMemory.forget('entity-123', { sourceRef: 'doc.md' }, { dryRun: true });
|
|
959
|
+
// preview: { deleted: { entries: number; tasks: number } }
|
|
960
|
+
// No database writes performed; safe for blast-radius validation
|
|
961
|
+
```
|
|
962
|
+
|
|
830
963
|
## Chunking Utilities
|
|
831
964
|
|
|
832
965
|
`ingestDocument()` splits a document into chunks before extraction. That same chunking is exported as a pure function, so a consumer can reproduce ingest-time chunk boundaries exactly — useful for recovering the passage a fact was extracted from, by re-chunking the source and ranking chunks against the fact's stored embedding.
|
|
@@ -1220,7 +1220,12 @@ function extractParsePosition(err) {
|
|
|
1220
1220
|
return match ? Number(match[1]) : null;
|
|
1221
1221
|
}
|
|
1222
1222
|
function safeErrorToString(e) {
|
|
1223
|
-
|
|
1223
|
+
let isErrorLike = false;
|
|
1224
|
+
try {
|
|
1225
|
+
isErrorLike = e instanceof Error;
|
|
1226
|
+
} catch {
|
|
1227
|
+
}
|
|
1228
|
+
if (isErrorLike) {
|
|
1224
1229
|
const msg = readErrorField(e, "message");
|
|
1225
1230
|
if (typeof msg === "string" && msg.length > 0) return msg;
|
|
1226
1231
|
const name = readErrorField(e, "name");
|
|
@@ -1245,16 +1250,54 @@ function readErrorField(e, key) {
|
|
|
1245
1250
|
}
|
|
1246
1251
|
}
|
|
1247
1252
|
function sanitizeRankerError(err, sanitizeRankerErrors) {
|
|
1253
|
+
let isErrorLike = false;
|
|
1254
|
+
try {
|
|
1255
|
+
isErrorLike = err instanceof Error;
|
|
1256
|
+
} catch {
|
|
1257
|
+
}
|
|
1248
1258
|
if (sanitizeRankerErrors === false) {
|
|
1249
|
-
return
|
|
1259
|
+
return isErrorLike ? err : new Error(safeErrorToString(err));
|
|
1260
|
+
}
|
|
1261
|
+
let errLike = null;
|
|
1262
|
+
if (isErrorLike) errLike = err;
|
|
1263
|
+
let typeName;
|
|
1264
|
+
try {
|
|
1265
|
+
if (errLike) {
|
|
1266
|
+
const rawName = errLike.constructor?.name;
|
|
1267
|
+
typeName = typeof rawName === "string" ? rawName : "Error";
|
|
1268
|
+
} else {
|
|
1269
|
+
typeName = typeof err;
|
|
1270
|
+
}
|
|
1271
|
+
} catch {
|
|
1272
|
+
typeName = isErrorLike ? "Error" : typeof err;
|
|
1273
|
+
}
|
|
1274
|
+
let innerCause;
|
|
1275
|
+
if (errLike) {
|
|
1276
|
+
let cause;
|
|
1277
|
+
try {
|
|
1278
|
+
cause = errLike.cause;
|
|
1279
|
+
} catch {
|
|
1280
|
+
cause = void 0;
|
|
1281
|
+
}
|
|
1282
|
+
if (cause !== void 0) {
|
|
1283
|
+
let causeName;
|
|
1284
|
+
try {
|
|
1285
|
+
const rawCauseName = cause?.constructor?.name;
|
|
1286
|
+
causeName = typeof rawCauseName === "string" ? rawCauseName : typeof cause;
|
|
1287
|
+
} catch {
|
|
1288
|
+
causeName = typeof cause;
|
|
1289
|
+
}
|
|
1290
|
+
innerCause = new Error(`Caused by: ${causeName}`);
|
|
1291
|
+
}
|
|
1250
1292
|
}
|
|
1251
|
-
const typeName = err instanceof Error ? err.constructor?.name ?? "Error" : typeof err;
|
|
1252
|
-
const innerCause = err instanceof Error && err.cause !== void 0 ? new Error(`Caused by: ${err.cause?.constructor?.name ?? typeof err.cause}`) : void 0;
|
|
1253
1293
|
const sanitized = new Error(
|
|
1254
1294
|
`VectorRanker ${typeName} (message scrubbed for security)`,
|
|
1255
1295
|
innerCause ? { cause: innerCause } : void 0
|
|
1256
1296
|
);
|
|
1257
|
-
|
|
1297
|
+
try {
|
|
1298
|
+
sanitized.name = typeName;
|
|
1299
|
+
} catch {
|
|
1300
|
+
}
|
|
1258
1301
|
return sanitized;
|
|
1259
1302
|
}
|
|
1260
1303
|
function safeSlice(value, start, end) {
|
|
@@ -2198,7 +2241,12 @@ var TRUNCATION_PATTERNS = [
|
|
|
2198
2241
|
];
|
|
2199
2242
|
var EXCEEDS_LIMIT_PATTERN = /exceed[a-z]*[^.]{0,40}\b(model|context)?[ _-]?limit/i;
|
|
2200
2243
|
function isTruncationError(err) {
|
|
2201
|
-
|
|
2244
|
+
let message;
|
|
2245
|
+
try {
|
|
2246
|
+
message = err instanceof Error ? err.message : String(err ?? "");
|
|
2247
|
+
} catch {
|
|
2248
|
+
return false;
|
|
2249
|
+
}
|
|
2202
2250
|
if (EXCEEDS_LIMIT_PATTERN.test(message)) return false;
|
|
2203
2251
|
return TRUNCATION_PATTERNS.some((pattern) => pattern.test(message));
|
|
2204
2252
|
}
|
|
@@ -2305,7 +2353,12 @@ var HEAL_RECHECK_MS = 7 * 24 * 60 * 60 * 1e3;
|
|
|
2305
2353
|
var SKIP_ERROR_LOG_CHARS = 4096;
|
|
2306
2354
|
var formatSkipError = (err) => {
|
|
2307
2355
|
let base;
|
|
2308
|
-
|
|
2356
|
+
let isErrorLike = false;
|
|
2357
|
+
try {
|
|
2358
|
+
isErrorLike = err instanceof Error;
|
|
2359
|
+
} catch {
|
|
2360
|
+
}
|
|
2361
|
+
if (isErrorLike || typeof err !== "object" && typeof err !== "function") {
|
|
2309
2362
|
base = safeErrorToString(err);
|
|
2310
2363
|
} else {
|
|
2311
2364
|
try {
|
|
@@ -3997,7 +4050,12 @@ var RetrievalService = class {
|
|
|
3997
4050
|
}
|
|
3998
4051
|
}
|
|
3999
4052
|
} catch (rankerErr) {
|
|
4000
|
-
|
|
4053
|
+
let isErrorLike = false;
|
|
4054
|
+
try {
|
|
4055
|
+
isErrorLike = rankerErr instanceof Error;
|
|
4056
|
+
} catch {
|
|
4057
|
+
}
|
|
4058
|
+
const rankerError = isErrorLike ? rankerErr : new Error(safeErrorToString(rankerErr));
|
|
4001
4059
|
const policy = this.options.vectorRankerFallback ?? "js-cosine";
|
|
4002
4060
|
this.options.onVectorRankerFallback?.({
|
|
4003
4061
|
error: this._sanitizeRankerError(rankerError),
|
|
@@ -4109,12 +4167,21 @@ var RetrievalService = class {
|
|
|
4109
4167
|
}
|
|
4110
4168
|
}
|
|
4111
4169
|
} catch (err) {
|
|
4112
|
-
|
|
4170
|
+
let isErrorLike = false;
|
|
4171
|
+
try {
|
|
4172
|
+
isErrorLike = err instanceof Error;
|
|
4173
|
+
} catch {
|
|
4174
|
+
}
|
|
4175
|
+
const error = isErrorLike ? err : new Error(safeErrorToString(err));
|
|
4113
4176
|
if (rankerShouldRethrow) {
|
|
4114
4177
|
throw error;
|
|
4115
4178
|
}
|
|
4116
4179
|
if (pendingRankerFallbackError) {
|
|
4117
|
-
|
|
4180
|
+
try {
|
|
4181
|
+
error.cause = pendingRankerFallbackError;
|
|
4182
|
+
} catch {
|
|
4183
|
+
this.options.onRetrievalFallback?.(pendingRankerFallbackError);
|
|
4184
|
+
}
|
|
4118
4185
|
pendingRankerFallbackError = void 0;
|
|
4119
4186
|
}
|
|
4120
4187
|
this.options.onRetrievalFallback?.(error);
|
|
@@ -4371,5 +4438,5 @@ var WriteService = class {
|
|
|
4371
4438
|
};
|
|
4372
4439
|
|
|
4373
4440
|
export { BaseRepository, DEFAULT_CHUNK_OVERLAP, DEFAULT_MAX_CHUNK_LENGTH, EmbeddingService, HEAL_BATCH_SIZE, HEAL_RECHECK_MS, HOOK_TIMEOUT_MARKER, ImportExportService, IngestionService, JobManager, MaintenanceService, MetadataRepository, ONTOLOGY_BACKFILL_BATCH_SIZE, ONTOLOGY_BACKFILL_MAX_PROMPT_CHARS, ONTOLOGY_BACKFILL_RECHECK_MS, ONTOLOGY_BACKFILL_SYSTEM_PROMPT, PromptService, PrunePartialFailureError, RetrievalService, SearchService, WikiBusyError, WikiDuplicateHashError, WikiIngestEmptyError, WikiParseError, WikiSourceRefHashCollision, WikiStrictOntologyViolation, WikiTransactionError, WriteService, __privateAdd, __privateGet, __privateSet, chunkText, configureRandomSource, emptyManifest, entitySummaryMetaKey, extractSqliteCode, generateId, normalizeSourceHash, normalizeSourceRef, normalizeTitleKey, parseEmbedding, resolveEdgeDefinitions, resolveNodeType, safeSlice, validateInlineEdges, validateManifest };
|
|
4374
|
-
//# sourceMappingURL=chunk-
|
|
4375
|
-
//# sourceMappingURL=chunk-
|
|
4441
|
+
//# sourceMappingURL=chunk-JQRA6BYQ.mjs.map
|
|
4442
|
+
//# sourceMappingURL=chunk-JQRA6BYQ.mjs.map
|