@equationalapplications/core-llm-wiki 5.5.0 → 5.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -52,6 +52,45 @@ Platform-agnostic TypeScript engine for hybrid LLM memory. Features episodic fac
52
52
  - **Interoperability:** Supports [Open Knowledge Format (OKF) v0.1](https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/okf) import and export via the [llm-wiki OKF profile (llm-wiki/1)](https://github.com/equationalapplications/expo-llm-wiki/blob/main/docs/okf-profile.md).
53
53
  - **Per-entity seeded ontology** — Optional Strict, Emergent, or Off modes govern LLM graph extraction; seed taxonomies per entity and persist typed facts with inline edges.
54
54
 
55
+ ## GraphRAG & Multi-Modal Retrieval
56
+
57
+ `@equationalapplications/core-llm-wiki` exposes three complementary retrieval modes, each addressing a different shape of query:
58
+
59
+ | Mode | API | Best for |
60
+ |---|---|---|
61
+ | **Semantic** (vector cosine) | `wiki.read(entityId, query)` with `embed` configured | Open-ended natural-language questions; "what do I know about X" |
62
+ | **Keyword** (MiniSearch) | `wiki.read(entityId, query)` with `embed` absent or offline | Exact terms, identifiers, names; offline fallback |
63
+ | **GraphRAG** (recursive CTE) | `wiki.traverseGraph(entityId, options)` + `formatGraphContext(result)` | Structural questions; "what connects to X", "everything two hops from this fact", "summarise the people, places, and projects linked to Alice" |
64
+
65
+ The GraphRAG path is structurally distinct: it doesn't rank by relevance to a query string, it walks `llm_wiki_edges` from a known anchor fact. The result is dense and connected — subgraphs, not loose top-K hits.
66
+
67
+ ### Graph traversal APIs
68
+
69
+ ```typescript
70
+ import { WikiMemory, formatGraphContext } from '@equationalapplications/core-llm-wiki';
71
+
72
+ const graph = await wikiMemory.traverseGraph('user-123', {
73
+ sourceId: '<anchor-fact-id>',
74
+ maxDepth: 2,
75
+ direction: 'both', // 'inbound' | 'outbound' | 'both'
76
+ edgeTypes: ['reports_to'], // optional filter
77
+ excludeSourceTypes: ['immutable_document'],
78
+ minTraversalConfidence: 'inferred',
79
+ maxTraversalNodes: 20,
80
+ });
81
+
82
+ const promptContext = formatGraphContext(graph);
83
+ // → dense text block ready for prompt injection
84
+ ```
85
+
86
+ `traverseGraph` runs as a single recursive CTE in SQLite (see the root README's ["The SQL: how traversal works in one query"](../../README.md#the-sql-how-traversal-works-in-one-query) for the query shape). No external graph database.
87
+
88
+ ### Deterministic graph seeding (no LLM)
89
+
90
+ For programmatic pipelines — importing pre-classified data, building a GraphRAG corpus from a CSV, or seed-loading from a JSON file — use `upsertGraph()`. It writes nodes and edges directly under the same `(sourceRef, sourceHash)` ownership semantics as `ingestDocument()`, but skips the LLM extraction step. See [Direct Graph Write](#direct-graph-write) for the canonical signature, including the required `SQLiteAdapter` argument and transactional semantics.
91
+
92
+ This is the GraphRAG seed path: load a corpus, walk it.
93
+
55
94
  ## Installation
56
95
 
57
96
  ```bash
@@ -827,6 +866,100 @@ configureRandomSource(getRandomValues);
827
866
 
828
867
  `@equationalapplications/expo-llm-wiki` does this automatically on import (main entry and `/factory` subpath). If you use `@equationalapplications/core-llm-wiki` directly on React Native without the expo package, you must call `configureRandomSource()` yourself or polyfill `globalThis.crypto.getRandomValues`.
829
868
 
869
+ ## Entity Enumeration
870
+
871
+ List all entities that have stored data in the wiki:
872
+
873
+ ```typescript
874
+ const entityIds = await wikiMemory.listEntityIds();
875
+ // Returns all entity_ids with at least one row (including soft-deleted-only entities)
876
+ // Optional prefix filter: await wikiMemory.listEntityIds({ prefix: 'tier_' });
877
+ ```
878
+
879
+ Use this for maintenance scheduling, multi-entity operations, or discovering which namespaces exist. Includes entities with only soft-deleted rows so `runPrune()` can reclaim orphaned storage.
880
+
881
+ ## Source Reference Enumeration
882
+
883
+ List all documents currently stored for an entity:
884
+
885
+ ```typescript
886
+ const sourceRefs = await wikiMemory.listSourceRefs('user-123');
887
+ // One row per live sourceRef (soft-deleted rows are excluded):
888
+ // Array<{ sourceRef: string; sourceHash: string | null; factCount: number; lastIngestedAt: number }>
889
+ // factCount — number of live facts under that sourceRef
890
+ // lastIngestedAt — Unix timestamp in ms from the most recently updated live entry
891
+ ```
892
+
893
+ Use this to audit stored documents, validate external sync state, or preview the blast radius before `forget()` operations.
894
+
895
+ ## Direct Graph Write
896
+
897
+ Write structured graph data directly without LLM extraction — useful for programmatic fact ingestion, parsers, and deterministic pipelines:
898
+
899
+ ```typescript
900
+ const { nodesWritten, edgesWritten, superseded } = await wikiMemory.upsertGraph('entity-123', {
901
+ sourceRef: 'codebase_main.ts',
902
+ sourceHash: sha256(sourceCode),
903
+ nodes: [
904
+ { id: 'fn_processData', type: 'function', title: 'processData', body: 'Processes user data' },
905
+ { id: 'class_UserService', type: 'class', title: 'UserService', body: 'User management service' },
906
+ ],
907
+ edges: [
908
+ { type: 'calls', sourceId: 'fn_processData', targetId: 'class_UserService' },
909
+ ],
910
+ }, adapter); // SQLiteAdapter from your platform driver — writes join the caller's transaction
911
+ ```
912
+
913
+ `upsertGraph` is "the tail of `ingestDocument` with the middle (LLM extraction) step removed" — it accepts caller-supplied nodes (`{ id, type, title, body? }`) and edges (`{ type, sourceId, targetId, id? }`) and writes them under the same `(sourceRef, sourceHash)` semantics. If a *different* live `sourceRef` already holds the same `sourceHash`, it throws `WikiSourceRefHashCollision`; re-writing the identical `(sourceRef, sourceHash)` is a no-op returning zero counts. The adapter parameter is required so writes participate in the caller's transaction.
914
+
915
+ ## Duplicate Hash Detection
916
+
917
+ Control behavior when a different live `sourceRef` already holds the same `sourceHash`. The option is passed as a third argument to `ingestDocument`, not inside the params object:
918
+
919
+ ```typescript
920
+ await wikiMemory.ingestDocument(
921
+ 'entity-123',
922
+ {
923
+ sourceRef: 'doc.md',
924
+ sourceHash: sha256(content),
925
+ documentChunk: content,
926
+ },
927
+ { onDuplicateHash: 'ingest' } // 'ingest' (default) | 'skip' | 'throw'
928
+ );
929
+ ```
930
+
931
+ - `'ingest'` (default): No duplicate pre-check; extraction proceeds as before this option existed. If a different live `sourceRef` is found holding the same hash at commit time (a concurrent-writer race caught by the source-ref unique index), the call still throws `WikiDuplicateHashError`.
932
+ - `'skip'`: Pre-check before any LLM call; if a different live `sourceRef` already holds the hash, return a zero-chunk result without writing.
933
+ - `'throw'`: Pre-check before any LLM call; throw `WikiDuplicateHashError` (carries the canonical `sourceRef`).
934
+ - The guard only considers **live** references — soft-deleted refs do not trigger it in any mode.
935
+
936
+ ## Batch Change Detection
937
+
938
+ Check multiple documents for changes in one call:
939
+
940
+ ```typescript
941
+ const batch = [
942
+ { sourceRef: 'doc1.md', sourceHash: sha256(content1) },
943
+ { sourceRef: 'doc2.md', sourceHash: sha256(content2) },
944
+ { sourceRef: 'doc3.md', sourceHash: sha256(content3) },
945
+ ];
946
+ const changes = await wikiMemory.hasChanged('entity-123', batch);
947
+ // changes: Array<{ sourceRef: string; changed: boolean; duplicateOf?: string }>
948
+ // duplicateOf — when present, the canonical stored different sourceRef holding
949
+ // the same hash (DB-normalized spelling; sourceRef echoes the raw caller value).
950
+ // Per-document change detection; internally batched across queries
951
+ ```
952
+
953
+ ## Dry-Run Deletion
954
+
955
+ Preview deletion impact without writing:
956
+
957
+ ```typescript
958
+ const preview = await wikiMemory.forget('entity-123', { sourceRef: 'doc.md' }, { dryRun: true });
959
+ // preview: { deleted: { entries: number; tasks: number } }
960
+ // No database writes performed; safe for blast-radius validation
961
+ ```
962
+
830
963
  ## Chunking Utilities
831
964
 
832
965
  `ingestDocument()` splits a document into chunks before extraction. That same chunking is exported as a pure function, so a consumer can reproduce ingest-time chunk boundaries exactly — useful for recovering the passage a fact was extracted from, by re-chunking the source and ranking chunks against the fact's stored embedding.
@@ -1220,7 +1220,12 @@ function extractParsePosition(err) {
1220
1220
  return match ? Number(match[1]) : null;
1221
1221
  }
1222
1222
  function safeErrorToString(e) {
1223
- if (e instanceof Error) {
1223
+ let isErrorLike = false;
1224
+ try {
1225
+ isErrorLike = e instanceof Error;
1226
+ } catch {
1227
+ }
1228
+ if (isErrorLike) {
1224
1229
  const msg = readErrorField(e, "message");
1225
1230
  if (typeof msg === "string" && msg.length > 0) return msg;
1226
1231
  const name = readErrorField(e, "name");
@@ -1245,16 +1250,54 @@ function readErrorField(e, key) {
1245
1250
  }
1246
1251
  }
1247
1252
  function sanitizeRankerError(err, sanitizeRankerErrors) {
1253
+ let isErrorLike = false;
1254
+ try {
1255
+ isErrorLike = err instanceof Error;
1256
+ } catch {
1257
+ }
1248
1258
  if (sanitizeRankerErrors === false) {
1249
- return err instanceof Error ? err : new Error(String(err));
1259
+ return isErrorLike ? err : new Error(safeErrorToString(err));
1260
+ }
1261
+ let errLike = null;
1262
+ if (isErrorLike) errLike = err;
1263
+ let typeName;
1264
+ try {
1265
+ if (errLike) {
1266
+ const rawName = errLike.constructor?.name;
1267
+ typeName = typeof rawName === "string" ? rawName : "Error";
1268
+ } else {
1269
+ typeName = typeof err;
1270
+ }
1271
+ } catch {
1272
+ typeName = isErrorLike ? "Error" : typeof err;
1273
+ }
1274
+ let innerCause;
1275
+ if (errLike) {
1276
+ let cause;
1277
+ try {
1278
+ cause = errLike.cause;
1279
+ } catch {
1280
+ cause = void 0;
1281
+ }
1282
+ if (cause !== void 0) {
1283
+ let causeName;
1284
+ try {
1285
+ const rawCauseName = cause?.constructor?.name;
1286
+ causeName = typeof rawCauseName === "string" ? rawCauseName : typeof cause;
1287
+ } catch {
1288
+ causeName = typeof cause;
1289
+ }
1290
+ innerCause = new Error(`Caused by: ${causeName}`);
1291
+ }
1250
1292
  }
1251
- const typeName = err instanceof Error ? err.constructor?.name ?? "Error" : typeof err;
1252
- const innerCause = err instanceof Error && err.cause !== void 0 ? new Error(`Caused by: ${err.cause?.constructor?.name ?? typeof err.cause}`) : void 0;
1253
1293
  const sanitized = new Error(
1254
1294
  `VectorRanker ${typeName} (message scrubbed for security)`,
1255
1295
  innerCause ? { cause: innerCause } : void 0
1256
1296
  );
1257
- sanitized.name = typeName;
1297
+ try {
1298
+ sanitized.name = typeName;
1299
+ } catch {
1300
+ }
1258
1301
  return sanitized;
1259
1302
  }
1260
1303
  function safeSlice(value, start, end) {
@@ -2198,7 +2241,12 @@ var TRUNCATION_PATTERNS = [
2198
2241
  ];
2199
2242
  var EXCEEDS_LIMIT_PATTERN = /exceed[a-z]*[^.]{0,40}\b(model|context)?[ _-]?limit/i;
2200
2243
  function isTruncationError(err) {
2201
- const message = err instanceof Error ? err.message : String(err ?? "");
2244
+ let message;
2245
+ try {
2246
+ message = err instanceof Error ? err.message : String(err ?? "");
2247
+ } catch {
2248
+ return false;
2249
+ }
2202
2250
  if (EXCEEDS_LIMIT_PATTERN.test(message)) return false;
2203
2251
  return TRUNCATION_PATTERNS.some((pattern) => pattern.test(message));
2204
2252
  }
@@ -2305,7 +2353,12 @@ var HEAL_RECHECK_MS = 7 * 24 * 60 * 60 * 1e3;
2305
2353
  var SKIP_ERROR_LOG_CHARS = 4096;
2306
2354
  var formatSkipError = (err) => {
2307
2355
  let base;
2308
- if (err instanceof Error || typeof err !== "object" && typeof err !== "function") {
2356
+ let isErrorLike = false;
2357
+ try {
2358
+ isErrorLike = err instanceof Error;
2359
+ } catch {
2360
+ }
2361
+ if (isErrorLike || typeof err !== "object" && typeof err !== "function") {
2309
2362
  base = safeErrorToString(err);
2310
2363
  } else {
2311
2364
  try {
@@ -3997,7 +4050,12 @@ var RetrievalService = class {
3997
4050
  }
3998
4051
  }
3999
4052
  } catch (rankerErr) {
4000
- const rankerError = rankerErr instanceof Error ? rankerErr : new Error(String(rankerErr));
4053
+ let isErrorLike = false;
4054
+ try {
4055
+ isErrorLike = rankerErr instanceof Error;
4056
+ } catch {
4057
+ }
4058
+ const rankerError = isErrorLike ? rankerErr : new Error(safeErrorToString(rankerErr));
4001
4059
  const policy = this.options.vectorRankerFallback ?? "js-cosine";
4002
4060
  this.options.onVectorRankerFallback?.({
4003
4061
  error: this._sanitizeRankerError(rankerError),
@@ -4109,12 +4167,21 @@ var RetrievalService = class {
4109
4167
  }
4110
4168
  }
4111
4169
  } catch (err) {
4112
- const error = err instanceof Error ? err : new Error(String(err));
4170
+ let isErrorLike = false;
4171
+ try {
4172
+ isErrorLike = err instanceof Error;
4173
+ } catch {
4174
+ }
4175
+ const error = isErrorLike ? err : new Error(safeErrorToString(err));
4113
4176
  if (rankerShouldRethrow) {
4114
4177
  throw error;
4115
4178
  }
4116
4179
  if (pendingRankerFallbackError) {
4117
- error.cause = pendingRankerFallbackError;
4180
+ try {
4181
+ error.cause = pendingRankerFallbackError;
4182
+ } catch {
4183
+ this.options.onRetrievalFallback?.(pendingRankerFallbackError);
4184
+ }
4118
4185
  pendingRankerFallbackError = void 0;
4119
4186
  }
4120
4187
  this.options.onRetrievalFallback?.(error);
@@ -4371,5 +4438,5 @@ var WriteService = class {
4371
4438
  };
4372
4439
 
4373
4440
  export { BaseRepository, DEFAULT_CHUNK_OVERLAP, DEFAULT_MAX_CHUNK_LENGTH, EmbeddingService, HEAL_BATCH_SIZE, HEAL_RECHECK_MS, HOOK_TIMEOUT_MARKER, ImportExportService, IngestionService, JobManager, MaintenanceService, MetadataRepository, ONTOLOGY_BACKFILL_BATCH_SIZE, ONTOLOGY_BACKFILL_MAX_PROMPT_CHARS, ONTOLOGY_BACKFILL_RECHECK_MS, ONTOLOGY_BACKFILL_SYSTEM_PROMPT, PromptService, PrunePartialFailureError, RetrievalService, SearchService, WikiBusyError, WikiDuplicateHashError, WikiIngestEmptyError, WikiParseError, WikiSourceRefHashCollision, WikiStrictOntologyViolation, WikiTransactionError, WriteService, __privateAdd, __privateGet, __privateSet, chunkText, configureRandomSource, emptyManifest, entitySummaryMetaKey, extractSqliteCode, generateId, normalizeSourceHash, normalizeSourceRef, normalizeTitleKey, parseEmbedding, resolveEdgeDefinitions, resolveNodeType, safeSlice, validateInlineEdges, validateManifest };
4374
- //# sourceMappingURL=chunk-QVY5DFJV.mjs.map
4375
- //# sourceMappingURL=chunk-QVY5DFJV.mjs.map
4441
+ //# sourceMappingURL=chunk-JQRA6BYQ.mjs.map
4442
+ //# sourceMappingURL=chunk-JQRA6BYQ.mjs.map