@equationalapplications/core-llm-wiki 5.5.0 → 6.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -10,34 +10,6 @@ Platform-agnostic TypeScript engine for hybrid LLM memory. Features episodic fac
10
10
 
11
11
  > Inspired by [Andrej Karpathy's LLM Wiki memory spec](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f).
12
12
 
13
- ## Recent changes
14
-
15
- ### 5.4.1 — Ingest parse resilience (issue #92)
16
-
17
- - `parseJsonResponse` now tolerates bare `"` characters in LLM output via a
18
- container-aware repair pass. The function's signature is unchanged
19
- (`parseJsonResponse<T>(text: string): T`); the runtime contract widens —
20
- the new `WikiParseError` type carries `{ tier, position, slice }` instead
21
- of an opaque `Error`, and `tier: 'repair'` is now reachable for balanced
22
- invalid payloads (e.g. `{"facts":}`) where the walker found a candidate
23
- span that `JSON.parse` ultimately rejected.
24
- - `IngestionService.ingestDocument` no longer rejects the whole call when one
25
- chunk fails. Sibling chunks commit; the result's `parseFailures[]` records
26
- per-chunk failures. A new typed error `WikiIngestEmptyError` is thrown when
27
- every chunk failed. New typed error `WikiParseError` carries `{tier,
28
- position, slice}`. New result fields: `ingestedChunks`, `failedChunks`,
29
- `parseFailures?`.
30
- - Partial-commit semantics: when some chunks fail, the document's
31
- `(entity, sourceHash) → sourceRef` ownership is **not** recorded and the
32
- partial rows are stored with `source_hash = NULL`. Subsequent runs with
33
- the same hash see `hasChanged` return `true` and the failed chunks retry
34
- on the next pass; the retry's `appendPartialFacts` dedupes against the
35
- prior partial's surviving rows so titles are never duplicated. On full
36
- success the next run supersedes everything atomically.
37
- - `INGEST_SYSTEM_PROMPT` and `ONTOLOGY_BACKFILL_SYSTEM_PROMPT` were tightened
38
- to call out JSON-escape discipline explicitly.
39
-
40
- **Hosts must update (host-facing migration)**: a host that today treats `ingestDocument` throwing as the only failure signal will, after 5.4.1, see a successful return with `parseFailures[]` set when a subset of chunks failed. Inspect `result.failedChunks` and surface `result.parseFailures[]` for observability — do not rely on the throw for partial failures. Only `WikiIngestEmptyError` (every chunk failed) still throws. Catch `WikiParseError` explicitly if you previously caught `Error` to log parse failures — the new `tier` field replaces the opaque `message`.
41
13
 
42
14
  ## Features
43
15
 
@@ -49,9 +21,48 @@ Platform-agnostic TypeScript engine for hybrid LLM memory. Features episodic fac
49
21
  - **Immutable vs mutable facts** — Use `WikiFact.source_type` to distinguish document-sourced facts (`immutable_document`) from derived or user-provided facts (`librarian_inferred`, `user_stated`, `user_confirmed`). Immutable document facts are not rewritten by `runLibrarian()` or `runHeal()` and can only be removed by `forget()` or re-ingesting.
50
22
  - **Full-featured memory** — Facts, tasks, events, maintenance jobs (librarian, heal, reembed, prune)
51
23
  - **Type-safe** — Built with TypeScript, full type exports
52
- - **Interoperability:** Supports [Open Knowledge Format (OKF) v0.1](https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/okf) import and export via the [llm-wiki OKF profile (llm-wiki/1)](https://github.com/equationalapplications/expo-llm-wiki/blob/main/docs/okf-profile.md).
24
+ - **Interoperability:** Supports [Open Knowledge Format (OKF)](https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/okf) v0.1 + v0.2 import and export via the [llm-wiki OKF profiles](https://github.com/equationalapplications/expo-llm-wiki/blob/main/docs/okf-profile.md) (default `llm-wiki/2`, back-compat `llm-wiki/1`).
53
25
  - **Per-entity seeded ontology** — Optional Strict, Emergent, or Off modes govern LLM graph extraction; seed taxonomies per entity and persist typed facts with inline edges.
54
26
 
27
+ ## GraphRAG & Multi-Modal Retrieval
28
+
29
+ `@equationalapplications/core-llm-wiki` exposes three complementary retrieval modes, each addressing a different shape of query:
30
+
31
+ | Mode | API | Best for |
32
+ |---|---|---|
33
+ | **Semantic** (vector cosine) | `wiki.read(entityId, query)` with `embed` configured | Open-ended natural-language questions; "what do I know about X" |
34
+ | **Keyword** (MiniSearch) | `wiki.read(entityId, query)` with `embed` absent or offline | Exact terms, identifiers, names; offline fallback |
35
+ | **GraphRAG** (recursive CTE) | `wiki.traverseGraph(entityId, options)` + `formatGraphContext(result)` | Structural questions; "what connects to X", "everything two hops from this fact", "summarise the people, places, and projects linked to Alice" |
36
+
37
+ The GraphRAG path is structurally distinct: it doesn't rank by relevance to a query string, it walks `llm_wiki_edges` from a known anchor fact. The result is dense and connected — subgraphs, not loose top-K hits.
38
+
39
+ ### Graph traversal APIs
40
+
41
+ ```typescript
42
+ import { WikiMemory, formatGraphContext } from '@equationalapplications/core-llm-wiki';
43
+
44
+ const graph = await wikiMemory.traverseGraph('user-123', {
45
+ sourceId: '<anchor-fact-id>',
46
+ maxDepth: 2,
47
+ direction: 'both', // 'inbound' | 'outbound' | 'both'
48
+ edgeTypes: ['reports_to'], // optional filter
49
+ excludeSourceTypes: ['immutable_document'],
50
+ minTraversalConfidence: 'inferred',
51
+ maxTraversalNodes: 20,
52
+ });
53
+
54
+ const promptContext = formatGraphContext(graph);
55
+ // → dense text block ready for prompt injection
56
+ ```
57
+
58
+ `traverseGraph` runs as a single recursive CTE in SQLite (see the root README's ["The SQL: how traversal works in one query"](../../README.md#the-sql-how-traversal-works-in-one-query) for the query shape). No external graph database.
59
+
60
+ ### Deterministic graph seeding (no LLM)
61
+
62
+ For programmatic pipelines — importing pre-classified data, building a GraphRAG corpus from a CSV, or seed-loading from a JSON file — use `upsertGraph()`. It writes nodes and edges directly under the same `(sourceRef, sourceHash)` ownership semantics as `ingestDocument()`, but skips the LLM extraction step. See [Direct Graph Write](#direct-graph-write) for the canonical signature, including the required `SQLiteAdapter` argument and transactional semantics.
63
+
64
+ This is the GraphRAG seed path: load a corpus, walk it.
65
+
55
66
  ## Installation
56
67
 
57
68
  ```bash
@@ -610,7 +621,7 @@ const result = await wiki.runOntologyBackfill(entityId);
610
621
 
611
622
  ## OKF Import/Export
612
623
 
613
- The core package integrates with `@equationalapplications/core-okf` to seamlessly adapt wiki data dumps to and from Open Knowledge Format (OKF) v0.1 bundles.
624
+ The core package integrates with `@equationalapplications/core-okf` to seamlessly adapt wiki data dumps to and from Open Knowledge Format (OKF) bundles (v0.1 and v0.2; `formatOkfBundle` defaults to the v0.2 / `llm-wiki/2` profile).
614
625
 
615
626
  ### Exporting an OKF Bundle
616
627
 
@@ -827,6 +838,100 @@ configureRandomSource(getRandomValues);
827
838
 
828
839
  `@equationalapplications/expo-llm-wiki` does this automatically on import (main entry and `/factory` subpath). If you use `@equationalapplications/core-llm-wiki` directly on React Native without the expo package, you must call `configureRandomSource()` yourself or polyfill `globalThis.crypto.getRandomValues`.
829
840
 
841
+ ## Entity Enumeration
842
+
843
+ List all entities that have stored data in the wiki:
844
+
845
+ ```typescript
846
+ const entityIds = await wikiMemory.listEntityIds();
847
+ // Returns all entity_ids with at least one row (including soft-deleted-only entities)
848
+ // Optional prefix filter: await wikiMemory.listEntityIds({ prefix: 'tier_' });
849
+ ```
850
+
851
+ Use this for maintenance scheduling, multi-entity operations, or discovering which namespaces exist. Includes entities with only soft-deleted rows so `runPrune()` can reclaim orphaned storage.
852
+
853
+ ## Source Reference Enumeration
854
+
855
+ List all documents currently stored for an entity:
856
+
857
+ ```typescript
858
+ const sourceRefs = await wikiMemory.listSourceRefs('user-123');
859
+ // One row per live sourceRef (soft-deleted rows are excluded):
860
+ // Array<{ sourceRef: string; sourceHash: string | null; factCount: number; lastIngestedAt: number }>
861
+ // factCount — number of live facts under that sourceRef
862
+ // lastIngestedAt — Unix timestamp in ms from the most recently updated live entry
863
+ ```
864
+
865
+ Use this to audit stored documents, validate external sync state, or preview the blast radius before `forget()` operations.
866
+
867
+ ## Direct Graph Write
868
+
869
+ Write structured graph data directly without LLM extraction — useful for programmatic fact ingestion, parsers, and deterministic pipelines:
870
+
871
+ ```typescript
872
+ const { nodesWritten, edgesWritten, superseded } = await wikiMemory.upsertGraph('entity-123', {
873
+ sourceRef: 'codebase_main.ts',
874
+ sourceHash: sha256(sourceCode),
875
+ nodes: [
876
+ { id: 'fn_processData', type: 'function', title: 'processData', body: 'Processes user data' },
877
+ { id: 'class_UserService', type: 'class', title: 'UserService', body: 'User management service' },
878
+ ],
879
+ edges: [
880
+ { type: 'calls', sourceId: 'fn_processData', targetId: 'class_UserService' },
881
+ ],
882
+ }, adapter); // SQLiteAdapter from your platform driver — writes join the caller's transaction
883
+ ```
884
+
885
+ `upsertGraph` is "the tail of `ingestDocument` with the middle (LLM extraction) step removed" — it accepts caller-supplied nodes (`{ id, type, title, body? }`) and edges (`{ type, sourceId, targetId, id? }`) and writes them under the same `(sourceRef, sourceHash)` semantics. If a *different* live `sourceRef` already holds the same `sourceHash`, it throws `WikiSourceRefHashCollision`; re-writing the identical `(sourceRef, sourceHash)` is a no-op returning zero counts. The adapter parameter is required so writes participate in the caller's transaction.
886
+
887
+ ## Duplicate Hash Detection
888
+
889
+ Control behavior when a different live `sourceRef` already holds the same `sourceHash`. The option is passed as a third argument to `ingestDocument`, not inside the params object:
890
+
891
+ ```typescript
892
+ await wikiMemory.ingestDocument(
893
+ 'entity-123',
894
+ {
895
+ sourceRef: 'doc.md',
896
+ sourceHash: sha256(content),
897
+ documentChunk: content,
898
+ },
899
+ { onDuplicateHash: 'ingest' } // 'ingest' (default) | 'skip' | 'throw'
900
+ );
901
+ ```
902
+
903
+ - `'ingest'` (default): No duplicate pre-check; extraction proceeds as before this option existed. If a different live `sourceRef` is found holding the same hash at commit time (a concurrent-writer race caught by the source-ref unique index), the call still throws `WikiDuplicateHashError`.
904
+ - `'skip'`: Pre-check before any LLM call; if a different live `sourceRef` already holds the hash, return a zero-chunk result without writing.
905
+ - `'throw'`: Pre-check before any LLM call; throw `WikiDuplicateHashError` (carries the canonical `sourceRef`).
906
+ - The guard only considers **live** references — soft-deleted refs do not trigger it in any mode.
907
+
908
+ ## Batch Change Detection
909
+
910
+ Check multiple documents for changes in one call:
911
+
912
+ ```typescript
913
+ const batch = [
914
+ { sourceRef: 'doc1.md', sourceHash: sha256(content1) },
915
+ { sourceRef: 'doc2.md', sourceHash: sha256(content2) },
916
+ { sourceRef: 'doc3.md', sourceHash: sha256(content3) },
917
+ ];
918
+ const changes = await wikiMemory.hasChanged('entity-123', batch);
919
+ // changes: Array<{ sourceRef: string; changed: boolean; duplicateOf?: string }>
920
+ // duplicateOf — when present, the canonical stored different sourceRef holding
921
+ // the same hash (DB-normalized spelling; sourceRef echoes the raw caller value).
922
+ // Per-document change detection; internally batched across queries
923
+ ```
924
+
925
+ ## Dry-Run Deletion
926
+
927
+ Preview deletion impact without writing:
928
+
929
+ ```typescript
930
+ const preview = await wikiMemory.forget('entity-123', { sourceRef: 'doc.md' }, { dryRun: true });
931
+ // preview: { deleted: { entries: number; tasks: number } }
932
+ // No database writes performed; safe for blast-radius validation
933
+ ```
934
+
830
935
  ## Chunking Utilities
831
936
 
832
937
  `ingestDocument()` splits a document into chunks before extraction. That same chunking is exported as a pure function, so a consumer can reproduce ingest-time chunk boundaries exactly — useful for recovering the passage a fact was extracted from, by re-chunking the source and ranking chunks against the fact's stored embedding.
@@ -1024,7 +1129,7 @@ The flowchart shows:
1024
1129
  | [@equationalapplications/react-llm-wiki](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/react/README.md) | Persistent episodic memory for Web |
1025
1130
  | [@equationalapplications/prisma-outbox](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/prisma-outbox/README.md) | Sync SQLite outbox events to Prisma |
1026
1131
  | [@equationalapplications/core-llm-tools](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/core-llm-tools/README.md) | Gemini tool schemas and capability injector |
1027
- | [@equationalapplications/core-okf](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/okf/README.md) | Zero-dependency Open Knowledge Format (OKF) v0.1 primitives — parse and produce interoperable knowledge bundles. |
1132
+ | [@equationalapplications/core-okf](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/okf/README.md) | Zero-dependency Open Knowledge Format (OKF) v0.1 + v0.2 primitives — parse and produce interoperable knowledge bundles. |
1028
1133
  | [@equationalapplications/schema-org-llm-wiki](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/schema-org/README.md) | Curated schema.org warm-agent ontology manifest |
1029
1134
 
1030
1135
  ## OKF v0.2 conformance (llm-wiki/2)