@equationalapplications/core-llm-wiki 7.4.0 → 7.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -310,6 +310,36 @@ await wiki.promoteDraft(facts[0].id, 'user-1', { by: 'human:alice' }); // → st
310
310
 
311
311
  `promoteDraft` throws `WikiDraftNotFound` when no live draft with that id exists for the entity. The error is contextless by design. Promotion does not change `updated_at`, so a promoted fact keeps its recency position.
312
312
 
313
+ ## Grounding
314
+
315
+ Opt-in, deterministic evidence check for LLM-authored facts. When on, writers you choose must quote the source they were shown. A fact whose quotes are missing or not found is stored as a `draft` (see [Draft Review](#draft-review)), not rejected. A fact whose quotes all check out is stored `stable` with a `process:grounding-check` verifier, so its `trustTier` is `machine-confirmed`. Default off: 7.x write behavior is unchanged.
316
+
317
+ ```ts
318
+ new WikiMemory(db, {
319
+ llmProvider,
320
+ config: {
321
+ grounding: {
322
+ mode: 'draft', // 'off' (default) | 'draft'
323
+ writers: ['ingest'], // default ['ingest']; also 'librarian', 'heal'
324
+ minEvidenceChars: 20, // shorter quotes count as absent
325
+ maxEvidence: 3, // quotes asked for per fact
326
+ maxEvidenceChars: 300,
327
+ },
328
+ },
329
+ });
330
+ ```
331
+
332
+ - **What counts as source.**
333
+ - Ingest: the chunk text.
334
+ - Librarian: the `summary` of each event in the prompt.
335
+ - Heal: the `summary` of each recent event in the prompt, plus the bodies of non-draft document anchors. When heal is a writer, anchors are shown with their body clipped to 800 characters.
336
+ - Instructions, the ontology manifest, existing facts and identifiers never count, so a model cannot ground a claim by quoting them.
337
+ - **The check.** Both sides are normalized with NFKC, whitespace runs collapse to one space, and matching is case-sensitive. A fact with more than 10 quotes, or any quote not found, fails.
338
+ - **Diagnostics.** `grounding_missing` (reasons `no_evidence`, `evidence_too_short`) and `grounding_failed` (reasons `quote_not_found`, `too_many_quotes`), one per fact, with the new fact's `factId`. Quotes are never included.
339
+ - **`upsertGraph`** nodes are host-supplied and never grounded.
340
+ - **Librarian and heal** synthesize across events, so their pass rates are unknown. Measure them on your own event log before opting them in.
341
+ - Evidence quotes are not stored.
342
+
313
343
  ## Pluggable Vector Retrieval
314
344
 
315
345
  When your entity corpus grows, in-process cosine similarity scoring becomes a bottleneck. The optional **`VectorRanker`** interface lets you delegate semantic ranking to [**sqlite-vec**](https://github.com/asg017/sqlite-vec), [**sqlite-vss**](https://github.com/asg017/sqlite-vss), or an external vector database while `WikiMemory` handles embedding validation, hybrid scoring, and tier-2 row hydration.
@@ -799,6 +829,47 @@ const result = await wiki.runOntologyBackfill(entityId);
799
829
  `config.prompts.ontologyBackfillSystemPrompt` (template may use `{{facts}}`
800
830
  and the ontology placeholders).
801
831
 
832
+ #### Classifier mode (optional)
833
+
834
+ A System-One classifier (for example TypeSafe's Jev, an OpenJev-style open model, or a local ONNX classifier) can type facts during backfill without generating text. Add `classify` to your provider and opt in:
835
+
836
+ ```ts
837
+ const wiki = createWiki(db, {
838
+ llmProvider: { generateText, classify },
839
+ config: { ontology: { backfillClassifier: 'auto', classifyMinConfidence: 0.5 } },
840
+ });
841
+ await wiki.runOntologyBackfill('user-1'); // uses classify
842
+ await wiki.runOntologyBackfill('user-1', { classifier: 'llm' }); // force the generative path
843
+ ```
844
+
845
+ - Providing `classify` changes nothing by itself. The default is `'llm'`.
846
+ - Each untyped fact gets one `choice` question over the entity manifest's node types. Answers below `classifyMinConfidence` (default 0.5) are left untyped and retried after the cooldown.
847
+ - **No edges are proposed in classifier mode** (`edgesAdded: 0`): a classifier cannot extract edge targets. Run with `classifier: 'llm'` when you want edges.
848
+ - Manifests with more than 255 node types, or providers without `classify`, use the generative path.
849
+ - Answers are validated as untrusted. Off-list choices and out-of-range probabilities count toward `failedValidation`. A thrown `classify` counts toward `skipped` and is retried on the next pass.
850
+
851
+ Example adapter for Cloudflare Workers AI's `typesafe/jev`. This is illustrative, not a supported package; check the provider's current API reference before use.
852
+
853
+ ```ts
854
+ const classify: LLMProvider['classify'] = async ({ state, questions }) => {
855
+ const jevQuestions = Object.fromEntries(Object.entries(questions).map(([key, q]) => [key,
856
+ q.kind === 'choice' ? { type: 'choice', instructions: q.instructions, criteria: Object.fromEntries(q.options.map((o) => [o, o])) }
857
+ : q.kind === 'score' ? { type: 'score', instructions: q.instructions, criteria: q.levels }
858
+ : { type: 'noul', instructions: q.instructions },
859
+ ]));
860
+ const res = await env.AI.run('typesafe/jev', { state, questions: jevQuestions });
861
+ const answers = Object.fromEntries(Object.entries(res.answers).map(([key, a]: [string, any]) => [key,
862
+ a.type === 'choice' ? { kind: 'choice', choice: a.choice, confidence: a.confidence, probabilities: a.probabilities }
863
+ : a.type === 'score' ? {
864
+ kind: 'score', score: a.score, confidence: a.confidence,
865
+ probabilities: Object.keys(a.probabilities).sort((x, y) => Number(x) - Number(y)).map((k) => a.probabilities[k]),
866
+ }
867
+ : { kind: 'binary', probability: a.noul },
868
+ ]));
869
+ return { answers };
870
+ };
871
+ ```
872
+
802
873
  ## OKF Import/Export
803
874
 
804
875
  The core package integrates with `@equationalapplications/core-okf` to seamlessly adapt wiki data dumps to and from Open Knowledge Format (OKF) bundles (v0.1 and v0.2; `formatOkfBundle` defaults to the v0.2 / `llm-wiki/2` profile).