@equationalapplications/core-llm-wiki 5.4.0 → 5.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -10,6 +10,35 @@ Platform-agnostic TypeScript engine for hybrid LLM memory. Features episodic fac
10
10
 
11
11
  > Inspired by [Andrej Karpathy's LLM Wiki memory spec](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f).
12
12
 
13
+ ## Recent changes
14
+
15
+ ### 5.4.1 — Ingest parse resilience (issue #92)
16
+
17
+ - `parseJsonResponse` now tolerates bare `"` characters in LLM output via a
18
+ container-aware repair pass. The function's signature is unchanged
19
+ (`parseJsonResponse<T>(text: string): T`); the runtime contract widens —
20
+ the new `WikiParseError` type carries `{ tier, position, slice }` instead
21
+ of an opaque `Error`, and `tier: 'repair'` is now reachable for balanced
22
+ invalid payloads (e.g. `{"facts":}`) where the walker found a candidate
23
+ span that `JSON.parse` ultimately rejected.
24
+ - `IngestionService.ingestDocument` no longer rejects the whole call when one
25
+ chunk fails. Sibling chunks commit; the result's `parseFailures[]` records
26
+ per-chunk failures. A new typed error `WikiIngestEmptyError` is thrown when
27
+ every chunk failed. New typed error `WikiParseError` carries `{tier,
28
+ position, slice}`. New result fields: `ingestedChunks`, `failedChunks`,
29
+ `parseFailures?`.
30
+ - Partial-commit semantics: when some chunks fail, the document's
31
+ `(entity, sourceHash) → sourceRef` ownership is **not** recorded and the
32
+ partial rows are stored with `source_hash = NULL`. Subsequent runs with
33
+ the same hash see `hasChanged` return `true` and the failed chunks retry
34
+ on the next pass; the retry's `appendPartialFacts` dedupes against the
35
+ prior partial's surviving rows so titles are never duplicated. On full
36
+ success the next run supersedes everything atomically.
37
+ - `INGEST_SYSTEM_PROMPT` and `ONTOLOGY_BACKFILL_SYSTEM_PROMPT` were tightened
38
+ to call out JSON-escape discipline explicitly.
39
+
40
+ **Hosts must update (host-facing migration)**: a host that today treats `ingestDocument` throwing as the only failure signal will, after 5.4.1, see a successful return with `parseFailures[]` set when a subset of chunks failed. Inspect `result.failedChunks` and surface `result.parseFailures[]` for observability — do not rely on the throw for partial failures. Only `WikiIngestEmptyError` (every chunk failed) still throws. Catch `WikiParseError` explicitly if you previously caught `Error` to log parse failures — the new `tier` field replaces the opaque `message`.
41
+
13
42
  ## Features
14
43
 
15
44
  - **Platform-agnostic** — Zero runtime dependencies; works with any SQLite driver via the `SQLiteAdapter` interface
@@ -23,6 +52,45 @@ Platform-agnostic TypeScript engine for hybrid LLM memory. Features episodic fac
23
52
  - **Interoperability:** Supports [Open Knowledge Format (OKF) v0.1](https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/okf) import and export via the [llm-wiki OKF profile (llm-wiki/1)](https://github.com/equationalapplications/expo-llm-wiki/blob/main/docs/okf-profile.md).
24
53
  - **Per-entity seeded ontology** — Optional Strict, Emergent, or Off modes govern LLM graph extraction; seed taxonomies per entity and persist typed facts with inline edges.
25
54
 
55
+ ## GraphRAG & Multi-Modal Retrieval
56
+
57
+ `@equationalapplications/core-llm-wiki` exposes three complementary retrieval modes, each addressing a different shape of query:
58
+
59
+ | Mode | API | Best for |
60
+ |---|---|---|
61
+ | **Semantic** (vector cosine) | `wiki.read(entityId, query)` with `embed` configured | Open-ended natural-language questions; "what do I know about X" |
62
+ | **Keyword** (MiniSearch) | `wiki.read(entityId, query)` with `embed` absent or offline | Exact terms, identifiers, names; offline fallback |
63
+ | **GraphRAG** (recursive CTE) | `wiki.traverseGraph(entityId, options)` + `formatGraphContext(result)` | Structural questions; "what connects to X", "everything two hops from this fact", "summarise the people, places, and projects linked to Alice" |
64
+
65
+ The GraphRAG path is structurally distinct: it doesn't rank by relevance to a query string, it walks `llm_wiki_edges` from a known anchor fact. The result is dense and connected — subgraphs, not loose top-K hits.
66
+
67
+ ### Graph traversal APIs
68
+
69
+ ```typescript
70
+ import { WikiMemory, formatGraphContext } from '@equationalapplications/core-llm-wiki';
71
+
72
+ const graph = await wikiMemory.traverseGraph('user-123', {
73
+ sourceId: '<anchor-fact-id>',
74
+ maxDepth: 2,
75
+ direction: 'both', // 'inbound' | 'outbound' | 'both'
76
+ edgeTypes: ['reports_to'], // optional filter
77
+ excludeSourceTypes: ['immutable_document'],
78
+ minTraversalConfidence: 'inferred',
79
+ maxTraversalNodes: 20,
80
+ });
81
+
82
+ const promptContext = formatGraphContext(graph);
83
+ // → dense text block ready for prompt injection
84
+ ```
85
+
86
+ `traverseGraph` runs as a single recursive CTE in SQLite (see the root README's ["The SQL: how traversal works in one query"](../../README.md#the-sql-how-traversal-works-in-one-query) for the query shape). No external graph database.
87
+
88
+ ### Deterministic graph seeding (no LLM)
89
+
90
+ For programmatic pipelines — importing pre-classified data, building a GraphRAG corpus from a CSV, or seed-loading from a JSON file — use `upsertGraph()`. It writes nodes and edges directly under the same `(sourceRef, sourceHash)` ownership semantics as `ingestDocument()`, but skips the LLM extraction step. See [Direct Graph Write](#direct-graph-write) for the canonical signature, including the required `SQLiteAdapter` argument and transactional semantics.
91
+
92
+ This is the GraphRAG seed path: load a corpus, walk it.
93
+
26
94
  ## Installation
27
95
 
28
96
  ```bash
@@ -798,6 +866,100 @@ configureRandomSource(getRandomValues);
798
866
 
799
867
  `@equationalapplications/expo-llm-wiki` does this automatically on import (main entry and `/factory` subpath). If you use `@equationalapplications/core-llm-wiki` directly on React Native without the expo package, you must call `configureRandomSource()` yourself or polyfill `globalThis.crypto.getRandomValues`.
800
868
 
869
+ ## Entity Enumeration
870
+
871
+ List all entities that have stored data in the wiki:
872
+
873
+ ```typescript
874
+ const entityIds = await wikiMemory.listEntityIds();
875
+ // Returns all entity_ids with at least one row (including soft-deleted-only entities)
876
+ // Optional prefix filter: await wikiMemory.listEntityIds({ prefix: 'tier_' });
877
+ ```
878
+
879
+ Use this for maintenance scheduling, multi-entity operations, or discovering which namespaces exist. Includes entities with only soft-deleted rows so `runPrune()` can reclaim orphaned storage.
880
+
881
+ ## Source Reference Enumeration
882
+
883
+ List all documents currently stored for an entity:
884
+
885
+ ```typescript
886
+ const sourceRefs = await wikiMemory.listSourceRefs('user-123');
887
+ // One row per live sourceRef (soft-deleted rows are excluded):
888
+ // Array<{ sourceRef: string; sourceHash: string | null; factCount: number; lastIngestedAt: number }>
889
+ // factCount — number of live facts under that sourceRef
890
+ // lastIngestedAt — Unix timestamp in ms from the most recently updated live entry
891
+ ```
892
+
893
+ Use this to audit stored documents, validate external sync state, or preview the blast radius before `forget()` operations.
894
+
895
+ ## Direct Graph Write
896
+
897
+ Write structured graph data directly without LLM extraction — useful for programmatic fact ingestion, parsers, and deterministic pipelines:
898
+
899
+ ```typescript
900
+ const { nodesWritten, edgesWritten, superseded } = await wikiMemory.upsertGraph('entity-123', {
901
+ sourceRef: 'codebase_main.ts',
902
+ sourceHash: sha256(sourceCode),
903
+ nodes: [
904
+ { id: 'fn_processData', type: 'function', title: 'processData', body: 'Processes user data' },
905
+ { id: 'class_UserService', type: 'class', title: 'UserService', body: 'User management service' },
906
+ ],
907
+ edges: [
908
+ { type: 'calls', sourceId: 'fn_processData', targetId: 'class_UserService' },
909
+ ],
910
+ }, adapter); // SQLiteAdapter from your platform driver — writes join the caller's transaction
911
+ ```
912
+
913
+ `upsertGraph` is "the tail of `ingestDocument` with the middle (LLM extraction) step removed" — it accepts caller-supplied nodes (`{ id, type, title, body? }`) and edges (`{ type, sourceId, targetId, id? }`) and writes them under the same `(sourceRef, sourceHash)` semantics. If a *different* live `sourceRef` already holds the same `sourceHash`, it throws `WikiSourceRefHashCollision`; re-writing the identical `(sourceRef, sourceHash)` is a no-op returning zero counts. The adapter parameter is required so writes participate in the caller's transaction.
914
+
915
+ ## Duplicate Hash Detection
916
+
917
+ Control behavior when a different live `sourceRef` already holds the same `sourceHash`. The option is passed as a third argument to `ingestDocument`, not inside the params object:
918
+
919
+ ```typescript
920
+ await wikiMemory.ingestDocument(
921
+ 'entity-123',
922
+ {
923
+ sourceRef: 'doc.md',
924
+ sourceHash: sha256(content),
925
+ documentChunk: content,
926
+ },
927
+ { onDuplicateHash: 'ingest' } // 'ingest' (default) | 'skip' | 'throw'
928
+ );
929
+ ```
930
+
931
+ - `'ingest'` (default): No duplicate pre-check; extraction proceeds as before this option existed. If a different live `sourceRef` is found holding the same hash at commit time (a concurrent-writer race caught by the source-ref unique index), the call still throws `WikiDuplicateHashError`.
932
+ - `'skip'`: Pre-check before any LLM call; if a different live `sourceRef` already holds the hash, return a zero-chunk result without writing.
933
+ - `'throw'`: Pre-check before any LLM call; throw `WikiDuplicateHashError` (carries the canonical `sourceRef`).
934
+ - The guard only considers **live** references — soft-deleted refs do not trigger it in any mode.
935
+
936
+ ## Batch Change Detection
937
+
938
+ Check multiple documents for changes in one call:
939
+
940
+ ```typescript
941
+ const batch = [
942
+ { sourceRef: 'doc1.md', sourceHash: sha256(content1) },
943
+ { sourceRef: 'doc2.md', sourceHash: sha256(content2) },
944
+ { sourceRef: 'doc3.md', sourceHash: sha256(content3) },
945
+ ];
946
+ const changes = await wikiMemory.hasChanged('entity-123', batch);
947
+ // changes: Array<{ sourceRef: string; changed: boolean; duplicateOf?: string }>
948
+ // duplicateOf — when present, the canonical stored different sourceRef holding
949
+ // the same hash (DB-normalized spelling; sourceRef echoes the raw caller value).
950
+ // Per-document change detection; internally batched across queries
951
+ ```
952
+
953
+ ## Dry-Run Deletion
954
+
955
+ Preview deletion impact without writing:
956
+
957
+ ```typescript
958
+ const preview = await wikiMemory.forget('entity-123', { sourceRef: 'doc.md' }, { dryRun: true });
959
+ // preview: { deleted: { entries: number; tasks: number } }
960
+ // No database writes performed; safe for blast-radius validation
961
+ ```
962
+
801
963
  ## Chunking Utilities
802
964
 
803
965
  `ingestDocument()` splits a document into chunks before extraction. That same chunking is exported as a pure function, so a consumer can reproduce ingest-time chunk boundaries exactly — useful for recovering the passage a fact was extracted from, by re-chunking the source and ranking chunks against the fact's stored embedding.