@equationalapplications/core-llm-wiki 5.5.0 → 6.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +136 -31
- package/dist/{chunk-QVY5DFJV.mjs → chunk-HREQ6F6Z.mjs} +241 -45
- package/dist/chunk-HREQ6F6Z.mjs.map +1 -0
- package/dist/index.d.mts +18 -3
- package/dist/index.d.ts +18 -3
- package/dist/index.js +241 -42
- package/dist/index.js.map +1 -1
- package/dist/index.mjs +2 -2
- package/dist/index.mjs.map +1 -1
- package/dist/{testing-D0RnZjyW.d.mts → testing-DgqVB29I.d.mts} +72 -6
- package/dist/{testing-D0RnZjyW.d.ts → testing-DgqVB29I.d.ts} +72 -6
- package/dist/testing.d.mts +1 -1
- package/dist/testing.d.ts +1 -1
- package/dist/testing.js +238 -42
- package/dist/testing.js.map +1 -1
- package/dist/testing.mjs +1 -1
- package/package.json +2 -2
- package/dist/chunk-QVY5DFJV.mjs.map +0 -1
package/README.md
CHANGED
|
@@ -10,34 +10,6 @@ Platform-agnostic TypeScript engine for hybrid LLM memory. Features episodic fac
|
|
|
10
10
|
|
|
11
11
|
> Inspired by [Andrej Karpathy's LLM Wiki memory spec](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f).
|
|
12
12
|
|
|
13
|
-
## Recent changes
|
|
14
|
-
|
|
15
|
-
### 5.4.1 — Ingest parse resilience (issue #92)
|
|
16
|
-
|
|
17
|
-
- `parseJsonResponse` now tolerates bare `"` characters in LLM output via a
|
|
18
|
-
container-aware repair pass. The function's signature is unchanged
|
|
19
|
-
(`parseJsonResponse<T>(text: string): T`); the runtime contract widens —
|
|
20
|
-
the new `WikiParseError` type carries `{ tier, position, slice }` instead
|
|
21
|
-
of an opaque `Error`, and `tier: 'repair'` is now reachable for balanced
|
|
22
|
-
invalid payloads (e.g. `{"facts":}`) where the walker found a candidate
|
|
23
|
-
span that `JSON.parse` ultimately rejected.
|
|
24
|
-
- `IngestionService.ingestDocument` no longer rejects the whole call when one
|
|
25
|
-
chunk fails. Sibling chunks commit; the result's `parseFailures[]` records
|
|
26
|
-
per-chunk failures. A new typed error `WikiIngestEmptyError` is thrown when
|
|
27
|
-
every chunk failed. New typed error `WikiParseError` carries `{tier,
|
|
28
|
-
position, slice}`. New result fields: `ingestedChunks`, `failedChunks`,
|
|
29
|
-
`parseFailures?`.
|
|
30
|
-
- Partial-commit semantics: when some chunks fail, the document's
|
|
31
|
-
`(entity, sourceHash) → sourceRef` ownership is **not** recorded and the
|
|
32
|
-
partial rows are stored with `source_hash = NULL`. Subsequent runs with
|
|
33
|
-
the same hash see `hasChanged` return `true` and the failed chunks retry
|
|
34
|
-
on the next pass; the retry's `appendPartialFacts` dedupes against the
|
|
35
|
-
prior partial's surviving rows so titles are never duplicated. On full
|
|
36
|
-
success the next run supersedes everything atomically.
|
|
37
|
-
- `INGEST_SYSTEM_PROMPT` and `ONTOLOGY_BACKFILL_SYSTEM_PROMPT` were tightened
|
|
38
|
-
to call out JSON-escape discipline explicitly.
|
|
39
|
-
|
|
40
|
-
**Hosts must update (host-facing migration)**: a host that today treats `ingestDocument` throwing as the only failure signal will, after 5.4.1, see a successful return with `parseFailures[]` set when a subset of chunks failed. Inspect `result.failedChunks` and surface `result.parseFailures[]` for observability — do not rely on the throw for partial failures. Only `WikiIngestEmptyError` (every chunk failed) still throws. Catch `WikiParseError` explicitly if you previously caught `Error` to log parse failures — the new `tier` field replaces the opaque `message`.
|
|
41
13
|
|
|
42
14
|
## Features
|
|
43
15
|
|
|
@@ -49,9 +21,48 @@ Platform-agnostic TypeScript engine for hybrid LLM memory. Features episodic fac
|
|
|
49
21
|
- **Immutable vs mutable facts** — Use `WikiFact.source_type` to distinguish document-sourced facts (`immutable_document`) from derived or user-provided facts (`librarian_inferred`, `user_stated`, `user_confirmed`). Immutable document facts are not rewritten by `runLibrarian()` or `runHeal()` and can only be removed by `forget()` or re-ingesting.
|
|
50
22
|
- **Full-featured memory** — Facts, tasks, events, maintenance jobs (librarian, heal, reembed, prune)
|
|
51
23
|
- **Type-safe** — Built with TypeScript, full type exports
|
|
52
|
-
- **Interoperability:** Supports [Open Knowledge Format (OKF)
|
|
24
|
+
- **Interoperability:** Supports [Open Knowledge Format (OKF)](https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/okf) v0.1 + v0.2 import and export via the [llm-wiki OKF profiles](https://github.com/equationalapplications/expo-llm-wiki/blob/main/docs/okf-profile.md) (default `llm-wiki/2`, back-compat `llm-wiki/1`).
|
|
53
25
|
- **Per-entity seeded ontology** — Optional Strict, Emergent, or Off modes govern LLM graph extraction; seed taxonomies per entity and persist typed facts with inline edges.
|
|
54
26
|
|
|
27
|
+
## GraphRAG & Multi-Modal Retrieval
|
|
28
|
+
|
|
29
|
+
`@equationalapplications/core-llm-wiki` exposes three complementary retrieval modes, each addressing a different shape of query:
|
|
30
|
+
|
|
31
|
+
| Mode | API | Best for |
|
|
32
|
+
|---|---|---|
|
|
33
|
+
| **Semantic** (vector cosine) | `wiki.read(entityId, query)` with `embed` configured | Open-ended natural-language questions; "what do I know about X" |
|
|
34
|
+
| **Keyword** (MiniSearch) | `wiki.read(entityId, query)` with `embed` absent or offline | Exact terms, identifiers, names; offline fallback |
|
|
35
|
+
| **GraphRAG** (recursive CTE) | `wiki.traverseGraph(entityId, options)` + `formatGraphContext(result)` | Structural questions; "what connects to X", "everything two hops from this fact", "summarise the people, places, and projects linked to Alice" |
|
|
36
|
+
|
|
37
|
+
The GraphRAG path is structurally distinct: it doesn't rank by relevance to a query string, it walks `llm_wiki_edges` from a known anchor fact. The result is dense and connected — subgraphs, not loose top-K hits.
|
|
38
|
+
|
|
39
|
+
### Graph traversal APIs
|
|
40
|
+
|
|
41
|
+
```typescript
|
|
42
|
+
import { WikiMemory, formatGraphContext } from '@equationalapplications/core-llm-wiki';
|
|
43
|
+
|
|
44
|
+
const graph = await wikiMemory.traverseGraph('user-123', {
|
|
45
|
+
sourceId: '<anchor-fact-id>',
|
|
46
|
+
maxDepth: 2,
|
|
47
|
+
direction: 'both', // 'inbound' | 'outbound' | 'both'
|
|
48
|
+
edgeTypes: ['reports_to'], // optional filter
|
|
49
|
+
excludeSourceTypes: ['immutable_document'],
|
|
50
|
+
minTraversalConfidence: 'inferred',
|
|
51
|
+
maxTraversalNodes: 20,
|
|
52
|
+
});
|
|
53
|
+
|
|
54
|
+
const promptContext = formatGraphContext(graph);
|
|
55
|
+
// → dense text block ready for prompt injection
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
`traverseGraph` runs as a single recursive CTE in SQLite (see the root README's ["The SQL: how traversal works in one query"](../../README.md#the-sql-how-traversal-works-in-one-query) for the query shape). No external graph database.
|
|
59
|
+
|
|
60
|
+
### Deterministic graph seeding (no LLM)
|
|
61
|
+
|
|
62
|
+
For programmatic pipelines — importing pre-classified data, building a GraphRAG corpus from a CSV, or seed-loading from a JSON file — use `upsertGraph()`. It writes nodes and edges directly under the same `(sourceRef, sourceHash)` ownership semantics as `ingestDocument()`, but skips the LLM extraction step. See [Direct Graph Write](#direct-graph-write) for the canonical signature, including the required `SQLiteAdapter` argument and transactional semantics.
|
|
63
|
+
|
|
64
|
+
This is the GraphRAG seed path: load a corpus, walk it.
|
|
65
|
+
|
|
55
66
|
## Installation
|
|
56
67
|
|
|
57
68
|
```bash
|
|
@@ -610,7 +621,7 @@ const result = await wiki.runOntologyBackfill(entityId);
|
|
|
610
621
|
|
|
611
622
|
## OKF Import/Export
|
|
612
623
|
|
|
613
|
-
The core package integrates with `@equationalapplications/core-okf` to seamlessly adapt wiki data dumps to and from Open Knowledge Format (OKF) v0.1
|
|
624
|
+
The core package integrates with `@equationalapplications/core-okf` to seamlessly adapt wiki data dumps to and from Open Knowledge Format (OKF) bundles (v0.1 and v0.2; `formatOkfBundle` defaults to the v0.2 / `llm-wiki/2` profile).
|
|
614
625
|
|
|
615
626
|
### Exporting an OKF Bundle
|
|
616
627
|
|
|
@@ -827,6 +838,100 @@ configureRandomSource(getRandomValues);
|
|
|
827
838
|
|
|
828
839
|
`@equationalapplications/expo-llm-wiki` does this automatically on import (main entry and `/factory` subpath). If you use `@equationalapplications/core-llm-wiki` directly on React Native without the expo package, you must call `configureRandomSource()` yourself or polyfill `globalThis.crypto.getRandomValues`.
|
|
829
840
|
|
|
841
|
+
## Entity Enumeration
|
|
842
|
+
|
|
843
|
+
List all entities that have stored data in the wiki:
|
|
844
|
+
|
|
845
|
+
```typescript
|
|
846
|
+
const entityIds = await wikiMemory.listEntityIds();
|
|
847
|
+
// Returns all entity_ids with at least one row (including soft-deleted-only entities)
|
|
848
|
+
// Optional prefix filter: await wikiMemory.listEntityIds({ prefix: 'tier_' });
|
|
849
|
+
```
|
|
850
|
+
|
|
851
|
+
Use this for maintenance scheduling, multi-entity operations, or discovering which namespaces exist. Includes entities with only soft-deleted rows so `runPrune()` can reclaim orphaned storage.
|
|
852
|
+
|
|
853
|
+
## Source Reference Enumeration
|
|
854
|
+
|
|
855
|
+
List all documents currently stored for an entity:
|
|
856
|
+
|
|
857
|
+
```typescript
|
|
858
|
+
const sourceRefs = await wikiMemory.listSourceRefs('user-123');
|
|
859
|
+
// One row per live sourceRef (soft-deleted rows are excluded):
|
|
860
|
+
// Array<{ sourceRef: string; sourceHash: string | null; factCount: number; lastIngestedAt: number }>
|
|
861
|
+
// factCount — number of live facts under that sourceRef
|
|
862
|
+
// lastIngestedAt — Unix timestamp in ms from the most recently updated live entry
|
|
863
|
+
```
|
|
864
|
+
|
|
865
|
+
Use this to audit stored documents, validate external sync state, or preview the blast radius before `forget()` operations.
|
|
866
|
+
|
|
867
|
+
## Direct Graph Write
|
|
868
|
+
|
|
869
|
+
Write structured graph data directly without LLM extraction — useful for programmatic fact ingestion, parsers, and deterministic pipelines:
|
|
870
|
+
|
|
871
|
+
```typescript
|
|
872
|
+
const { nodesWritten, edgesWritten, superseded } = await wikiMemory.upsertGraph('entity-123', {
|
|
873
|
+
sourceRef: 'codebase_main.ts',
|
|
874
|
+
sourceHash: sha256(sourceCode),
|
|
875
|
+
nodes: [
|
|
876
|
+
{ id: 'fn_processData', type: 'function', title: 'processData', body: 'Processes user data' },
|
|
877
|
+
{ id: 'class_UserService', type: 'class', title: 'UserService', body: 'User management service' },
|
|
878
|
+
],
|
|
879
|
+
edges: [
|
|
880
|
+
{ type: 'calls', sourceId: 'fn_processData', targetId: 'class_UserService' },
|
|
881
|
+
],
|
|
882
|
+
}, adapter); // SQLiteAdapter from your platform driver — writes join the caller's transaction
|
|
883
|
+
```
|
|
884
|
+
|
|
885
|
+
`upsertGraph` is "the tail of `ingestDocument` with the middle (LLM extraction) step removed" — it accepts caller-supplied nodes (`{ id, type, title, body? }`) and edges (`{ type, sourceId, targetId, id? }`) and writes them under the same `(sourceRef, sourceHash)` semantics. If a *different* live `sourceRef` already holds the same `sourceHash`, it throws `WikiSourceRefHashCollision`; re-writing the identical `(sourceRef, sourceHash)` is a no-op returning zero counts. The adapter parameter is required so writes participate in the caller's transaction.
|
|
886
|
+
|
|
887
|
+
## Duplicate Hash Detection
|
|
888
|
+
|
|
889
|
+
Control behavior when a different live `sourceRef` already holds the same `sourceHash`. The option is passed as a third argument to `ingestDocument`, not inside the params object:
|
|
890
|
+
|
|
891
|
+
```typescript
|
|
892
|
+
await wikiMemory.ingestDocument(
|
|
893
|
+
'entity-123',
|
|
894
|
+
{
|
|
895
|
+
sourceRef: 'doc.md',
|
|
896
|
+
sourceHash: sha256(content),
|
|
897
|
+
documentChunk: content,
|
|
898
|
+
},
|
|
899
|
+
{ onDuplicateHash: 'ingest' } // 'ingest' (default) | 'skip' | 'throw'
|
|
900
|
+
);
|
|
901
|
+
```
|
|
902
|
+
|
|
903
|
+
- `'ingest'` (default): No duplicate pre-check; extraction proceeds as before this option existed. If a different live `sourceRef` is found holding the same hash at commit time (a concurrent-writer race caught by the source-ref unique index), the call still throws `WikiDuplicateHashError`.
|
|
904
|
+
- `'skip'`: Pre-check before any LLM call; if a different live `sourceRef` already holds the hash, return a zero-chunk result without writing.
|
|
905
|
+
- `'throw'`: Pre-check before any LLM call; throw `WikiDuplicateHashError` (carries the canonical `sourceRef`).
|
|
906
|
+
- The guard only considers **live** references — soft-deleted refs do not trigger it in any mode.
|
|
907
|
+
|
|
908
|
+
## Batch Change Detection
|
|
909
|
+
|
|
910
|
+
Check multiple documents for changes in one call:
|
|
911
|
+
|
|
912
|
+
```typescript
|
|
913
|
+
const batch = [
|
|
914
|
+
{ sourceRef: 'doc1.md', sourceHash: sha256(content1) },
|
|
915
|
+
{ sourceRef: 'doc2.md', sourceHash: sha256(content2) },
|
|
916
|
+
{ sourceRef: 'doc3.md', sourceHash: sha256(content3) },
|
|
917
|
+
];
|
|
918
|
+
const changes = await wikiMemory.hasChanged('entity-123', batch);
|
|
919
|
+
// changes: Array<{ sourceRef: string; changed: boolean; duplicateOf?: string }>
|
|
920
|
+
// duplicateOf — when present, the canonical stored different sourceRef holding
|
|
921
|
+
// the same hash (DB-normalized spelling; sourceRef echoes the raw caller value).
|
|
922
|
+
// Per-document change detection; internally batched across queries
|
|
923
|
+
```
|
|
924
|
+
|
|
925
|
+
## Dry-Run Deletion
|
|
926
|
+
|
|
927
|
+
Preview deletion impact without writing:
|
|
928
|
+
|
|
929
|
+
```typescript
|
|
930
|
+
const preview = await wikiMemory.forget('entity-123', { sourceRef: 'doc.md' }, { dryRun: true });
|
|
931
|
+
// preview: { deleted: { entries: number; tasks: number } }
|
|
932
|
+
// No database writes performed; safe for blast-radius validation
|
|
933
|
+
```
|
|
934
|
+
|
|
830
935
|
## Chunking Utilities
|
|
831
936
|
|
|
832
937
|
`ingestDocument()` splits a document into chunks before extraction. That same chunking is exported as a pure function, so a consumer can reproduce ingest-time chunk boundaries exactly — useful for recovering the passage a fact was extracted from, by re-chunking the source and ranking chunks against the fact's stored embedding.
|
|
@@ -1024,7 +1129,7 @@ The flowchart shows:
|
|
|
1024
1129
|
| [@equationalapplications/react-llm-wiki](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/react/README.md) | Persistent episodic memory for Web |
|
|
1025
1130
|
| [@equationalapplications/prisma-outbox](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/prisma-outbox/README.md) | Sync SQLite outbox events to Prisma |
|
|
1026
1131
|
| [@equationalapplications/core-llm-tools](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/core-llm-tools/README.md) | Gemini tool schemas and capability injector |
|
|
1027
|
-
| [@equationalapplications/core-okf](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/okf/README.md) | Zero-dependency Open Knowledge Format (OKF) v0.1 primitives — parse and produce interoperable knowledge bundles. |
|
|
1132
|
+
| [@equationalapplications/core-okf](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/okf/README.md) | Zero-dependency Open Knowledge Format (OKF) v0.1 + v0.2 primitives — parse and produce interoperable knowledge bundles. |
|
|
1028
1133
|
| [@equationalapplications/schema-org-llm-wiki](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/schema-org/README.md) | Curated schema.org warm-agent ontology manifest |
|
|
1029
1134
|
|
|
1030
1135
|
## OKF v0.2 conformance (llm-wiki/2)
|