@equationalapplications/core-llm-wiki 5.5.1 → 6.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +3 -31
- package/dist/{chunk-JQRA6BYQ.mjs → chunk-HREQ6F6Z.mjs} +168 -39
- package/dist/chunk-HREQ6F6Z.mjs.map +1 -0
- package/dist/index.d.mts +18 -3
- package/dist/index.d.ts +18 -3
- package/dist/index.js +168 -36
- package/dist/index.js.map +1 -1
- package/dist/index.mjs +2 -2
- package/dist/index.mjs.map +1 -1
- package/dist/{testing-D0RnZjyW.d.mts → testing-DgqVB29I.d.mts} +72 -6
- package/dist/{testing-D0RnZjyW.d.ts → testing-DgqVB29I.d.ts} +72 -6
- package/dist/testing.d.mts +1 -1
- package/dist/testing.d.ts +1 -1
- package/dist/testing.js +165 -36
- package/dist/testing.js.map +1 -1
- package/dist/testing.mjs +1 -1
- package/package.json +2 -2
- package/dist/chunk-JQRA6BYQ.mjs.map +0 -1
package/README.md
CHANGED
|
@@ -10,34 +10,6 @@ Platform-agnostic TypeScript engine for hybrid LLM memory. Features episodic fac
|
|
|
10
10
|
|
|
11
11
|
> Inspired by [Andrej Karpathy's LLM Wiki memory spec](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f).
|
|
12
12
|
|
|
13
|
-
## Recent changes
|
|
14
|
-
|
|
15
|
-
### 5.4.1 — Ingest parse resilience (issue #92)
|
|
16
|
-
|
|
17
|
-
- `parseJsonResponse` now tolerates bare `"` characters in LLM output via a
|
|
18
|
-
container-aware repair pass. The function's signature is unchanged
|
|
19
|
-
(`parseJsonResponse<T>(text: string): T`); the runtime contract widens —
|
|
20
|
-
the new `WikiParseError` type carries `{ tier, position, slice }` instead
|
|
21
|
-
of an opaque `Error`, and `tier: 'repair'` is now reachable for balanced
|
|
22
|
-
invalid payloads (e.g. `{"facts":}`) where the walker found a candidate
|
|
23
|
-
span that `JSON.parse` ultimately rejected.
|
|
24
|
-
- `IngestionService.ingestDocument` no longer rejects the whole call when one
|
|
25
|
-
chunk fails. Sibling chunks commit; the result's `parseFailures[]` records
|
|
26
|
-
per-chunk failures. A new typed error `WikiIngestEmptyError` is thrown when
|
|
27
|
-
every chunk failed. New typed error `WikiParseError` carries `{tier,
|
|
28
|
-
position, slice}`. New result fields: `ingestedChunks`, `failedChunks`,
|
|
29
|
-
`parseFailures?`.
|
|
30
|
-
- Partial-commit semantics: when some chunks fail, the document's
|
|
31
|
-
`(entity, sourceHash) → sourceRef` ownership is **not** recorded and the
|
|
32
|
-
partial rows are stored with `source_hash = NULL`. Subsequent runs with
|
|
33
|
-
the same hash see `hasChanged` return `true` and the failed chunks retry
|
|
34
|
-
on the next pass; the retry's `appendPartialFacts` dedupes against the
|
|
35
|
-
prior partial's surviving rows so titles are never duplicated. On full
|
|
36
|
-
success the next run supersedes everything atomically.
|
|
37
|
-
- `INGEST_SYSTEM_PROMPT` and `ONTOLOGY_BACKFILL_SYSTEM_PROMPT` were tightened
|
|
38
|
-
to call out JSON-escape discipline explicitly.
|
|
39
|
-
|
|
40
|
-
**Hosts must update (host-facing migration)**: a host that today treats `ingestDocument` throwing as the only failure signal will, after 5.4.1, see a successful return with `parseFailures[]` set when a subset of chunks failed. Inspect `result.failedChunks` and surface `result.parseFailures[]` for observability — do not rely on the throw for partial failures. Only `WikiIngestEmptyError` (every chunk failed) still throws. Catch `WikiParseError` explicitly if you previously caught `Error` to log parse failures — the new `tier` field replaces the opaque `message`.
|
|
41
13
|
|
|
42
14
|
## Features
|
|
43
15
|
|
|
@@ -49,7 +21,7 @@ Platform-agnostic TypeScript engine for hybrid LLM memory. Features episodic fac
|
|
|
49
21
|
- **Immutable vs mutable facts** — Use `WikiFact.source_type` to distinguish document-sourced facts (`immutable_document`) from derived or user-provided facts (`librarian_inferred`, `user_stated`, `user_confirmed`). Immutable document facts are not rewritten by `runLibrarian()` or `runHeal()` and can only be removed by `forget()` or re-ingesting.
|
|
50
22
|
- **Full-featured memory** — Facts, tasks, events, maintenance jobs (librarian, heal, reembed, prune)
|
|
51
23
|
- **Type-safe** — Built with TypeScript, full type exports
|
|
52
|
-
- **Interoperability:** Supports [Open Knowledge Format (OKF)
|
|
24
|
+
- **Interoperability:** Supports [Open Knowledge Format (OKF)](https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/okf) v0.1 + v0.2 import and export via the [llm-wiki OKF profiles](https://github.com/equationalapplications/expo-llm-wiki/blob/main/docs/okf-profile.md) (default `llm-wiki/2`, back-compat `llm-wiki/1`).
|
|
53
25
|
- **Per-entity seeded ontology** — Optional Strict, Emergent, or Off modes govern LLM graph extraction; seed taxonomies per entity and persist typed facts with inline edges.
|
|
54
26
|
|
|
55
27
|
## GraphRAG & Multi-Modal Retrieval
|
|
@@ -649,7 +621,7 @@ const result = await wiki.runOntologyBackfill(entityId);
|
|
|
649
621
|
|
|
650
622
|
## OKF Import/Export
|
|
651
623
|
|
|
652
|
-
The core package integrates with `@equationalapplications/core-okf` to seamlessly adapt wiki data dumps to and from Open Knowledge Format (OKF) v0.1
|
|
624
|
+
The core package integrates with `@equationalapplications/core-okf` to seamlessly adapt wiki data dumps to and from Open Knowledge Format (OKF) bundles (v0.1 and v0.2; `formatOkfBundle` defaults to the v0.2 / `llm-wiki/2` profile).
|
|
653
625
|
|
|
654
626
|
### Exporting an OKF Bundle
|
|
655
627
|
|
|
@@ -1157,7 +1129,7 @@ The flowchart shows:
|
|
|
1157
1129
|
| [@equationalapplications/react-llm-wiki](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/react/README.md) | Persistent episodic memory for Web |
|
|
1158
1130
|
| [@equationalapplications/prisma-outbox](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/prisma-outbox/README.md) | Sync SQLite outbox events to Prisma |
|
|
1159
1131
|
| [@equationalapplications/core-llm-tools](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/core-llm-tools/README.md) | Gemini tool schemas and capability injector |
|
|
1160
|
-
| [@equationalapplications/core-okf](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/okf/README.md) | Zero-dependency Open Knowledge Format (OKF) v0.1 primitives — parse and produce interoperable knowledge bundles. |
|
|
1132
|
+
| [@equationalapplications/core-okf](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/okf/README.md) | Zero-dependency Open Knowledge Format (OKF) v0.1 + v0.2 primitives — parse and produce interoperable knowledge bundles. |
|
|
1161
1133
|
| [@equationalapplications/schema-org-llm-wiki](https://github.com/equationalapplications/expo-llm-wiki/blob/main/packages/schema-org/README.md) | Curated schema.org warm-agent ontology manifest |
|
|
1162
1134
|
|
|
1163
1135
|
## OKF v0.2 conformance (llm-wiki/2)
|
|
@@ -1481,6 +1481,12 @@ Return ONLY a valid JSON object matching this schema:
|
|
|
1481
1481
|
If no manifest type fits a fact, omit that fact from "classifications" entirely \u2014 do not guess.
|
|
1482
1482
|
When echoing an existing fact's title verbatim into "target_title", preserve every JSON escape sequence (\\", \\n, \\\\, \\/) exactly as it appeared in the input body \u2014 do not strip backslashes, do not add unescaped quotes. Do not return markdown, just raw JSON.`;
|
|
1483
1483
|
|
|
1484
|
+
// src/utils/healConstants.ts
|
|
1485
|
+
var HEAL_MAX_ANCHORS = 50;
|
|
1486
|
+
var HEAL_ANCHORS_PER_CANDIDATE = 4;
|
|
1487
|
+
var HEAL_MAX_FACT_BODY_CHARS_L3 = 4e3;
|
|
1488
|
+
var HEAL_MAX_TASKS = 50;
|
|
1489
|
+
|
|
1484
1490
|
// src/services/PromptService.ts
|
|
1485
1491
|
var PromptService = class {
|
|
1486
1492
|
constructor(globalOverrides) {
|
|
@@ -1548,25 +1554,67 @@ Current Facts:
|
|
|
1548
1554
|
${JSON.stringify(currentFacts, null, 2)}`
|
|
1549
1555
|
};
|
|
1550
1556
|
}
|
|
1551
|
-
|
|
1557
|
+
/**
|
|
1558
|
+
* Heal-prompt level interpretation for the `attemptLevel` ladder.
|
|
1559
|
+
*
|
|
1560
|
+
* Caller contract: `documentAnchors` may be a slice sized for `batch.length`
|
|
1561
|
+
* or a larger set (e.g. a cache hit from `_selectHealAnchors`). This function
|
|
1562
|
+
* applies the prompt-side anchor cap `min(HEAL_MAX_ANCHORS=50, batch.length
|
|
1563
|
+
* * HEAL_ANCHORS_PER_CANDIDATE=4)` so the rendered prompt is bounded
|
|
1564
|
+
* regardless of caller input. `HEAL_MAX_ANCHORS` and
|
|
1565
|
+
* `HEAL_ANCHORS_PER_CANDIDATE` live here too — keeping the formula
|
|
1566
|
+
* co-located with its application avoids a "MaintenanceService policy"
|
|
1567
|
+
* import cycle (`PromptService` is constructed before `MaintenanceService`
|
|
1568
|
+
* exists) and makes the cap testable without a `MaintenanceService`
|
|
1569
|
+
* instance. Task 3 exports the same two constants from `MaintenanceService`
|
|
1570
|
+
* for caller-side overfetch sizing; the values must match.
|
|
1571
|
+
*
|
|
1572
|
+
* Level semantics:
|
|
1573
|
+
* - L0: allTasks + recentEvents + full candidate bodies; anchors re-capped
|
|
1574
|
+
* - L1: drop allTasks; recentEvents present; candidate bodies full
|
|
1575
|
+
* - L2: drop allTasks and recentEvents; candidate bodies full
|
|
1576
|
+
* - L3: drop allTasks and recentEvents; truncate each candidate body to
|
|
1577
|
+
* `bodyTruncationChars` and emit a `degraded` record per truncated fact
|
|
1578
|
+
*/
|
|
1579
|
+
buildHealPrompt(healCandidates, documentAnchors, allTasks, recentEvents, runtimeOverride, attemptLevel, bodyTruncationChars = HEAL_MAX_FACT_BODY_CHARS_L3) {
|
|
1580
|
+
const effectiveTasks = attemptLevel >= 1 ? [] : allTasks;
|
|
1581
|
+
const effectiveEvents = attemptLevel >= 2 ? [] : recentEvents;
|
|
1582
|
+
const maxAnchors = Math.max(1, Math.min(HEAL_MAX_ANCHORS, healCandidates.length * HEAL_ANCHORS_PER_CANDIDATE));
|
|
1583
|
+
const effectiveAnchors = documentAnchors.slice(0, maxAnchors);
|
|
1584
|
+
const { shapedCandidates, degraded } = applyBodyTruncation(
|
|
1585
|
+
healCandidates,
|
|
1586
|
+
attemptLevel,
|
|
1587
|
+
bodyTruncationChars
|
|
1588
|
+
);
|
|
1552
1589
|
const template = runtimeOverride ?? this.globalOverrides?.healSystemPrompt ?? HEAL_SYSTEM_PROMPT;
|
|
1553
1590
|
if (/\{\{\s*healCandidates\s*\}\}/.test(template) || /\{\{\s*documentAnchors\s*\}\}/.test(template) || /\{\{\s*allTasks\s*\}\}/.test(template) || /\{\{\s*recentEvents\s*\}\}/.test(template)) {
|
|
1554
1591
|
return {
|
|
1555
|
-
|
|
1556
|
-
|
|
1592
|
+
prompts: {
|
|
1593
|
+
systemPrompt: this.hydrate(template, {
|
|
1594
|
+
healCandidates: shapedCandidates,
|
|
1595
|
+
documentAnchors: effectiveAnchors,
|
|
1596
|
+
allTasks: effectiveTasks,
|
|
1597
|
+
recentEvents: effectiveEvents
|
|
1598
|
+
}),
|
|
1599
|
+
userPrompt: "Please heal the memory graph."
|
|
1600
|
+
},
|
|
1601
|
+
degraded
|
|
1557
1602
|
};
|
|
1558
1603
|
}
|
|
1559
1604
|
return {
|
|
1560
|
-
|
|
1561
|
-
|
|
1562
|
-
|
|
1605
|
+
prompts: {
|
|
1606
|
+
systemPrompt: template,
|
|
1607
|
+
userPrompt: `Heal Candidates:
|
|
1608
|
+
${JSON.stringify(shapedCandidates, null, 2)}
|
|
1563
1609
|
Document Anchors (DO NOT MODIFY OR DELETE):
|
|
1564
|
-
${JSON.stringify(
|
|
1610
|
+
${JSON.stringify(effectiveAnchors, null, 2)}
|
|
1565
1611
|
All Tasks:
|
|
1566
|
-
${JSON.stringify(
|
|
1612
|
+
${JSON.stringify(effectiveTasks, null, 2)}
|
|
1567
1613
|
Recent Events:
|
|
1568
|
-
${JSON.stringify(
|
|
1614
|
+
${JSON.stringify(effectiveEvents, null, 2)}
|
|
1569
1615
|
The following document anchors are provided for contradiction detection only. Do not include them in \`downgraded\`, \`deleted\`, or \`newFacts\`.`
|
|
1616
|
+
},
|
|
1617
|
+
degraded
|
|
1570
1618
|
};
|
|
1571
1619
|
}
|
|
1572
1620
|
buildOntologyBackfillPrompt(facts, runtimeOverride, ontologyContext) {
|
|
@@ -1586,6 +1634,36 @@ ${JSON.stringify(facts, null, 2)}`
|
|
|
1586
1634
|
};
|
|
1587
1635
|
}
|
|
1588
1636
|
};
|
|
1637
|
+
function applyBodyTruncation(candidates, attemptLevel, bodyTruncationChars) {
|
|
1638
|
+
if (attemptLevel < 3) {
|
|
1639
|
+
return { shapedCandidates: candidates, degraded: [] };
|
|
1640
|
+
}
|
|
1641
|
+
const shapedCandidates = [];
|
|
1642
|
+
const degraded = [];
|
|
1643
|
+
for (const c of candidates) {
|
|
1644
|
+
if (typeof c !== "object" || c === null) {
|
|
1645
|
+
shapedCandidates.push(c);
|
|
1646
|
+
continue;
|
|
1647
|
+
}
|
|
1648
|
+
const fact = c;
|
|
1649
|
+
const body = typeof fact.body === "string" ? fact.body : "";
|
|
1650
|
+
if (body.length <= bodyTruncationChars) {
|
|
1651
|
+
shapedCandidates.push(c);
|
|
1652
|
+
continue;
|
|
1653
|
+
}
|
|
1654
|
+
if (typeof fact.id !== "string") {
|
|
1655
|
+
shapedCandidates.push(c);
|
|
1656
|
+
continue;
|
|
1657
|
+
}
|
|
1658
|
+
const originalBodyChars = body.length;
|
|
1659
|
+
const prefix = safeSlice(body, 0, bodyTruncationChars);
|
|
1660
|
+
const truncatedBodyChars = prefix.length;
|
|
1661
|
+
const truncated = `${prefix}\u2026[truncated at ${truncatedBodyChars} chars, original was ${originalBodyChars}]`;
|
|
1662
|
+
shapedCandidates.push({ ...fact, body: truncated });
|
|
1663
|
+
degraded.push({ id: fact.id, originalBodyChars, truncatedBodyChars });
|
|
1664
|
+
}
|
|
1665
|
+
return { shapedCandidates, degraded };
|
|
1666
|
+
}
|
|
1589
1667
|
|
|
1590
1668
|
// src/utils/chunkingDefaults.ts
|
|
1591
1669
|
var DEFAULT_MAX_CHUNK_LENGTH = 12e3;
|
|
@@ -2240,16 +2318,24 @@ var TRUNCATION_PATTERNS = [
|
|
|
2240
2318
|
/finish[_ ]?reason/i
|
|
2241
2319
|
];
|
|
2242
2320
|
var EXCEEDS_LIMIT_PATTERN = /exceed[a-z]*[^.]{0,40}\b(model|context)?[ _-]?limit/i;
|
|
2243
|
-
function
|
|
2244
|
-
let message;
|
|
2321
|
+
function getErrorMessage(err) {
|
|
2245
2322
|
try {
|
|
2246
|
-
|
|
2323
|
+
return err instanceof Error ? err.message : String(err ?? "");
|
|
2247
2324
|
} catch {
|
|
2248
|
-
return
|
|
2325
|
+
return void 0;
|
|
2249
2326
|
}
|
|
2327
|
+
}
|
|
2328
|
+
function isTruncationError(err) {
|
|
2329
|
+
const message = getErrorMessage(err);
|
|
2330
|
+
if (message === void 0) return false;
|
|
2250
2331
|
if (EXCEEDS_LIMIT_PATTERN.test(message)) return false;
|
|
2251
2332
|
return TRUNCATION_PATTERNS.some((pattern) => pattern.test(message));
|
|
2252
2333
|
}
|
|
2334
|
+
function isConfigError(err) {
|
|
2335
|
+
const message = getErrorMessage(err);
|
|
2336
|
+
if (message === void 0) return false;
|
|
2337
|
+
return EXCEEDS_LIMIT_PATTERN.test(message);
|
|
2338
|
+
}
|
|
2253
2339
|
function initialBatchSize(maxOutputTokens) {
|
|
2254
2340
|
if (!maxOutputTokens || !Number.isFinite(maxOutputTokens) || maxOutputTokens <= 0) {
|
|
2255
2341
|
return DEFAULT_BATCH_SIZE;
|
|
@@ -2291,14 +2377,18 @@ async function runBatched(args) {
|
|
|
2291
2377
|
const single = candidate.slice(0, 1);
|
|
2292
2378
|
return { batch: single, prompts: await buildPrompt(single) };
|
|
2293
2379
|
};
|
|
2294
|
-
const onFailure = async (batch, err) => {
|
|
2295
|
-
if (batch.length
|
|
2296
|
-
if (
|
|
2297
|
-
|
|
2380
|
+
const onFailure = async (batch, err, attemptLevel, fromCall) => {
|
|
2381
|
+
if (batch.length === 1) {
|
|
2382
|
+
if (fromCall && isTruncationError(err) && attemptLevel < 3) {
|
|
2383
|
+
await attempt(batch, void 0, attemptLevel + 1);
|
|
2384
|
+
} else {
|
|
2385
|
+
const reason = fromCall && !isTruncationError(err) && !isConfigError(err) ? "call_error" : "non_convergent";
|
|
2386
|
+
skipped.push({ item: batch[0], reason });
|
|
2298
2387
|
onSkip?.(batch[0], err);
|
|
2299
2388
|
}
|
|
2300
2389
|
return;
|
|
2301
2390
|
}
|
|
2391
|
+
if (fromCall && !isTruncationError(err)) throw err;
|
|
2302
2392
|
const mid = Math.ceil(batch.length / 2);
|
|
2303
2393
|
if (mid < batchSize) batchSize = mid;
|
|
2304
2394
|
let i = 0;
|
|
@@ -2309,23 +2399,22 @@ async function runBatched(args) {
|
|
|
2309
2399
|
i += trimmed.batch.length;
|
|
2310
2400
|
}
|
|
2311
2401
|
};
|
|
2312
|
-
const attempt = async (batch, prebuilt) => {
|
|
2402
|
+
const attempt = async (batch, prebuilt, attemptLevel = 0) => {
|
|
2313
2403
|
if (batch.length === 0) return;
|
|
2314
|
-
const prompts = prebuilt
|
|
2404
|
+
const prompts = prebuilt && attemptLevel === 0 ? prebuilt : await buildPrompt(batch, attemptLevel);
|
|
2315
2405
|
batches++;
|
|
2316
2406
|
let responseText;
|
|
2317
2407
|
try {
|
|
2318
2408
|
responseText = await call(prompts);
|
|
2319
2409
|
} catch (err) {
|
|
2320
|
-
|
|
2321
|
-
await onFailure(batch, err);
|
|
2410
|
+
await onFailure(batch, err, attemptLevel, true);
|
|
2322
2411
|
return;
|
|
2323
2412
|
}
|
|
2324
2413
|
let result;
|
|
2325
2414
|
try {
|
|
2326
2415
|
result = parse(responseText, batch);
|
|
2327
2416
|
} catch (err) {
|
|
2328
|
-
await onFailure(batch, err);
|
|
2417
|
+
await onFailure(batch, err, attemptLevel, false);
|
|
2329
2418
|
return;
|
|
2330
2419
|
}
|
|
2331
2420
|
results.push(result);
|
|
@@ -2334,7 +2423,7 @@ async function runBatched(args) {
|
|
|
2334
2423
|
while (index < items.length) {
|
|
2335
2424
|
const { batch, prompts } = await trim(items.slice(index, index + batchSize));
|
|
2336
2425
|
index += batch.length;
|
|
2337
|
-
await attempt(batch, prompts);
|
|
2426
|
+
await attempt(batch, prompts, 0);
|
|
2338
2427
|
}
|
|
2339
2428
|
return { results, skipped, batches };
|
|
2340
2429
|
}
|
|
@@ -2345,7 +2434,6 @@ var MIN_TOKENS_TO_QUALIFY = 3;
|
|
|
2345
2434
|
var ONTOLOGY_BACKFILL_BATCH_SIZE = 25;
|
|
2346
2435
|
var ONTOLOGY_BACKFILL_MAX_PROMPT_CHARS = 4e4;
|
|
2347
2436
|
var ONTOLOGY_BACKFILL_RECHECK_MS = 7 * 24 * 60 * 60 * 1e3;
|
|
2348
|
-
var HEAL_MAX_ANCHORS = 50;
|
|
2349
2437
|
var HEAL_ANCHOR_SEARCH_OVERFETCH = 4;
|
|
2350
2438
|
var HEAL_MAX_PROMPT_CHARS = 4e4;
|
|
2351
2439
|
var HEAL_BATCH_SIZE = 25;
|
|
@@ -2776,6 +2864,10 @@ var MaintenanceService = class {
|
|
|
2776
2864
|
if (!Number.isInteger(batchSize) || batchSize < 1) {
|
|
2777
2865
|
throw new Error("Invalid batchSize: must be an integer >= 1");
|
|
2778
2866
|
}
|
|
2867
|
+
const bodyTruncationChars = options?.bodyTruncationChars ?? HEAL_MAX_FACT_BODY_CHARS_L3;
|
|
2868
|
+
if (!Number.isInteger(bodyTruncationChars) || bodyTruncationChars < 1) {
|
|
2869
|
+
throw new Error("Invalid bodyTruncationChars: must be an integer >= 1");
|
|
2870
|
+
}
|
|
2779
2871
|
const now = Date.now();
|
|
2780
2872
|
const recheckCutoff = now - HEAL_RECHECK_MS;
|
|
2781
2873
|
const orphanAfterDays = this.options.config?.orphanAfterDays !== void 0 ? this.options.config?.orphanAfterDays : 30;
|
|
@@ -2816,29 +2908,44 @@ var MaintenanceService = class {
|
|
|
2816
2908
|
downgraded: staleDowngradedIds.length,
|
|
2817
2909
|
deleted: orphanedIds.length,
|
|
2818
2910
|
newFactsCreated: 0,
|
|
2819
|
-
skipped:
|
|
2911
|
+
skipped: [],
|
|
2912
|
+
degraded: [],
|
|
2820
2913
|
remaining: counts2.eligible,
|
|
2821
2914
|
deferred: counts2.deferred
|
|
2822
2915
|
};
|
|
2823
2916
|
}
|
|
2824
|
-
const allTasks = await this.taskRepo.findAllPending([entityId]);
|
|
2917
|
+
const allTasks = await this.taskRepo.findAllPending([entityId], HEAL_MAX_TASKS);
|
|
2825
2918
|
const recentEvents = await this.eventRepo.getRecent(entityId, 20);
|
|
2826
2919
|
const toPromptShape = (f) => {
|
|
2827
2920
|
const { embedding: _embedding, embedding_blob: _blob, ...rest } = f;
|
|
2828
2921
|
return { ...rest, tags: typeof rest.tags === "string" ? JSON.parse(rest.tags) : rest.tags };
|
|
2829
2922
|
};
|
|
2830
2923
|
const anchorCache = /* @__PURE__ */ new Map();
|
|
2924
|
+
const degraded = [];
|
|
2831
2925
|
const outcome = await runBatched({
|
|
2832
2926
|
items: healCandidates,
|
|
2833
|
-
buildPrompt: async (batch) => {
|
|
2834
|
-
const documentAnchors = await this._selectHealAnchors(
|
|
2835
|
-
|
|
2927
|
+
buildPrompt: async (batch, attemptLevel = 0) => {
|
|
2928
|
+
const documentAnchors = await this._selectHealAnchors(
|
|
2929
|
+
entityId,
|
|
2930
|
+
batch,
|
|
2931
|
+
// Per-batch anchor cap: `batch.length * HEAL_ANCHORS_PER_CANDIDATE` (capped at
|
|
2932
|
+
// HEAL_MAX_ANCHORS) right-sizes the anchor lookup so a 1-fact batch does
|
|
2933
|
+
// not overfetch the same 200 keyword hits a 25-fact batch once did.
|
|
2934
|
+
// buildHealPrompt applies the matching cap on its side — values must match.
|
|
2935
|
+
Math.min(HEAL_MAX_ANCHORS, batch.length * HEAL_ANCHORS_PER_CANDIDATE),
|
|
2936
|
+
anchorCache
|
|
2937
|
+
);
|
|
2938
|
+
const { prompts, degraded: batchDegraded } = await this.promptService.buildHealPrompt(
|
|
2836
2939
|
batch.map(toPromptShape),
|
|
2837
2940
|
documentAnchors,
|
|
2838
2941
|
allTasks,
|
|
2839
2942
|
recentEvents,
|
|
2840
|
-
promptOverride
|
|
2943
|
+
promptOverride,
|
|
2944
|
+
attemptLevel,
|
|
2945
|
+
bodyTruncationChars
|
|
2841
2946
|
);
|
|
2947
|
+
degraded.push(...batchDegraded);
|
|
2948
|
+
return prompts;
|
|
2842
2949
|
},
|
|
2843
2950
|
call: (prompts) => this.options.llmProvider.generateText(prompts),
|
|
2844
2951
|
parse: (responseText, batch) => {
|
|
@@ -2912,8 +3019,14 @@ var MaintenanceService = class {
|
|
|
2912
3019
|
insertedFacts.push({ id, entity_id: entityId, title: fact.title, body: fact.body, tags: JSON.stringify(fact.tags) });
|
|
2913
3020
|
healFactsForDedupe.push({ id, title: fact.title });
|
|
2914
3021
|
}
|
|
3022
|
+
const callErrorIds = new Set(
|
|
3023
|
+
outcome.skipped.filter((s) => s.reason === "call_error").map((s) => s.item.id)
|
|
3024
|
+
);
|
|
2915
3025
|
await this.entryRepo.markHealChecked(
|
|
2916
|
-
[
|
|
3026
|
+
[
|
|
3027
|
+
...healCandidates.map((f) => f.id).filter((id) => !callErrorIds.has(id)),
|
|
3028
|
+
...insertedFacts.map((f) => f.id)
|
|
3029
|
+
],
|
|
2917
3030
|
entityId,
|
|
2918
3031
|
now,
|
|
2919
3032
|
tx
|
|
@@ -2931,6 +3044,13 @@ var MaintenanceService = class {
|
|
|
2931
3044
|
await this.embeddingService.embedFact(fact);
|
|
2932
3045
|
}
|
|
2933
3046
|
this.searchService.evictCache(entityId);
|
|
3047
|
+
const skippedIds = new Set(outcome.skipped.map(({ item }) => item.id));
|
|
3048
|
+
const healedDegraded = degraded.filter((d) => !skippedIds.has(d.id));
|
|
3049
|
+
for (const d of healedDegraded) {
|
|
3050
|
+
console.warn(
|
|
3051
|
+
`[WikiMemory] heal healed under degraded context ${entityId}/${d.id}: body truncated from ${d.originalBodyChars} to ${d.truncatedBodyChars} chars`
|
|
3052
|
+
);
|
|
3053
|
+
}
|
|
2934
3054
|
let scanned = outcome.skipped.length;
|
|
2935
3055
|
for (const batchResult of outcome.results) scanned += batchResult.batch.length;
|
|
2936
3056
|
const allDowngraded = /* @__PURE__ */ new Set([...staleDowngradedIds, ...safeDowngraded]);
|
|
@@ -2941,7 +3061,8 @@ var MaintenanceService = class {
|
|
|
2941
3061
|
downgraded: allDowngraded.size,
|
|
2942
3062
|
deleted: allDeleted.size,
|
|
2943
3063
|
newFactsCreated: insertedFacts.length,
|
|
2944
|
-
skipped: outcome.skipped.
|
|
3064
|
+
skipped: outcome.skipped.map(({ item, reason }) => ({ id: item.id, reason })),
|
|
3065
|
+
degraded: healedDegraded,
|
|
2945
3066
|
remaining: counts.eligible,
|
|
2946
3067
|
deferred: counts.deferred
|
|
2947
3068
|
};
|
|
@@ -3025,7 +3146,15 @@ var MaintenanceService = class {
|
|
|
3025
3146
|
};
|
|
3026
3147
|
}
|
|
3027
3148
|
if (outcome.skipped.length > 0) {
|
|
3028
|
-
|
|
3149
|
+
const callErrorIds = new Set(
|
|
3150
|
+
outcome.skipped.filter((s) => s.reason === "call_error").map((s) => s.item.id)
|
|
3151
|
+
);
|
|
3152
|
+
await this.entryRepo.markOntologyChecked(
|
|
3153
|
+
outcome.skipped.filter(({ item }) => !callErrorIds.has(item.id)).map(({ item }) => item.id),
|
|
3154
|
+
entityId,
|
|
3155
|
+
now,
|
|
3156
|
+
this.db
|
|
3157
|
+
);
|
|
3029
3158
|
}
|
|
3030
3159
|
this.searchService.evictCache(entityId);
|
|
3031
3160
|
const counts = await this.entryRepo.countUntypedByEntityId(entityId, recheckCutoff);
|
|
@@ -3141,7 +3270,7 @@ var MaintenanceService = class {
|
|
|
3141
3270
|
* batches that reduce to the same query share one lookup. Caller-owned and
|
|
3142
3271
|
* per-pass — see the call site in doRunHeal.
|
|
3143
3272
|
*/
|
|
3144
|
-
async _selectHealAnchors(entityId, batch, cache) {
|
|
3273
|
+
async _selectHealAnchors(entityId, batch, cap = HEAL_MAX_ANCHORS, cache) {
|
|
3145
3274
|
const query = batch.map((f) => f.title).join(" ").trim();
|
|
3146
3275
|
if (!query) return [];
|
|
3147
3276
|
const cached = cache?.get(query);
|
|
@@ -3149,7 +3278,7 @@ var MaintenanceService = class {
|
|
|
3149
3278
|
const hits = this.searchService.searchKeyword(
|
|
3150
3279
|
query,
|
|
3151
3280
|
[entityId],
|
|
3152
|
-
|
|
3281
|
+
cap * HEAL_ANCHOR_SEARCH_OVERFETCH
|
|
3153
3282
|
);
|
|
3154
3283
|
const hitIds = hits.map((h) => h.id);
|
|
3155
3284
|
const anchors = [];
|
|
@@ -3160,7 +3289,7 @@ var MaintenanceService = class {
|
|
|
3160
3289
|
const row = byId.get(id);
|
|
3161
3290
|
if (!row) continue;
|
|
3162
3291
|
anchors.push(row);
|
|
3163
|
-
if (anchors.length >=
|
|
3292
|
+
if (anchors.length >= cap) break;
|
|
3164
3293
|
}
|
|
3165
3294
|
}
|
|
3166
3295
|
cache?.set(query, anchors);
|
|
@@ -4437,6 +4566,6 @@ var WriteService = class {
|
|
|
4437
4566
|
}
|
|
4438
4567
|
};
|
|
4439
4568
|
|
|
4440
|
-
export { BaseRepository, DEFAULT_CHUNK_OVERLAP, DEFAULT_MAX_CHUNK_LENGTH, EmbeddingService, HEAL_BATCH_SIZE, HEAL_RECHECK_MS, HOOK_TIMEOUT_MARKER, ImportExportService, IngestionService, JobManager, MaintenanceService, MetadataRepository, ONTOLOGY_BACKFILL_BATCH_SIZE, ONTOLOGY_BACKFILL_MAX_PROMPT_CHARS, ONTOLOGY_BACKFILL_RECHECK_MS, ONTOLOGY_BACKFILL_SYSTEM_PROMPT, PromptService, PrunePartialFailureError, RetrievalService, SearchService, WikiBusyError, WikiDuplicateHashError, WikiIngestEmptyError, WikiParseError, WikiSourceRefHashCollision, WikiStrictOntologyViolation, WikiTransactionError, WriteService, __privateAdd, __privateGet, __privateSet, chunkText, configureRandomSource, emptyManifest, entitySummaryMetaKey, extractSqliteCode, generateId, normalizeSourceHash, normalizeSourceRef, normalizeTitleKey, parseEmbedding, resolveEdgeDefinitions, resolveNodeType, safeSlice, validateInlineEdges, validateManifest };
|
|
4441
|
-
//# sourceMappingURL=chunk-
|
|
4442
|
-
//# sourceMappingURL=chunk-
|
|
4569
|
+
export { BaseRepository, DEFAULT_CHUNK_OVERLAP, DEFAULT_MAX_CHUNK_LENGTH, EmbeddingService, HEAL_ANCHORS_PER_CANDIDATE, HEAL_BATCH_SIZE, HEAL_MAX_FACT_BODY_CHARS_L3, HEAL_MAX_TASKS, HEAL_RECHECK_MS, HOOK_TIMEOUT_MARKER, ImportExportService, IngestionService, JobManager, MaintenanceService, MetadataRepository, ONTOLOGY_BACKFILL_BATCH_SIZE, ONTOLOGY_BACKFILL_MAX_PROMPT_CHARS, ONTOLOGY_BACKFILL_RECHECK_MS, ONTOLOGY_BACKFILL_SYSTEM_PROMPT, PromptService, PrunePartialFailureError, RetrievalService, SearchService, WikiBusyError, WikiDuplicateHashError, WikiIngestEmptyError, WikiParseError, WikiSourceRefHashCollision, WikiStrictOntologyViolation, WikiTransactionError, WriteService, __privateAdd, __privateGet, __privateSet, chunkText, configureRandomSource, emptyManifest, entitySummaryMetaKey, extractSqliteCode, generateId, normalizeSourceHash, normalizeSourceRef, normalizeTitleKey, parseEmbedding, resolveEdgeDefinitions, resolveNodeType, safeSlice, validateInlineEdges, validateManifest };
|
|
4570
|
+
//# sourceMappingURL=chunk-HREQ6F6Z.mjs.map
|
|
4571
|
+
//# sourceMappingURL=chunk-HREQ6F6Z.mjs.map
|