@yolk_vat-y/dsh-project-memory 0.5.3 → 0.5.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,28 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.5.4 (2026-09-12)
4
+
5
+ ### Changed (indexing: model-free by design; whole-chunk retrieval terms)
6
+
7
+ - **Indexing never calls a model — by construction.** `buildDocEntries(relPath, filePath, opts)` no longer accepts an `llm` parameter, so index-time code cannot reach a model even by accident. `src/doc-pipeline.js`, `src/watch.js`, `src/lazy.js`, `index_doc` and `index_repo` no longer resolve or pass a route. A stub-model run over 289 files / 2141 entries records **0 calls**.
8
+ - **Retrieval terms cover the whole chunk.** The old shape made the whole chunk's searchable body the ≤300-char `summary` (`util/search.js` `weightedFieldText`), so a 3000-char chunk had roughly 90% of its text invisible to `query_memory`. Each doc entry now also carries `terms` — a bounded (≤160), deterministic, stop-word-filtered literal term set covering the **entire chunk** (`src/doc-index.js`, built from the existing `tokenizeRaw`). `summary` stays short for the injection budget; `terms` is search-only.
9
+ - Measured on this repository's docs (179 chunks): reachable unique chunk terms go from **27.3% to 100%**; terms that appear only beyond character 300 go from **0% to 100%** entryHit@5, while already-answerable queries keep their ranking (MRR **0.958** with `terms` vs **0.955** without). Zero model calls.
10
+ - `linkEntries` now includes `terms` in its haystack for the same reason: a symbol mentioned late in a chunk links again.
11
+ - **Existing stores are back-filled automatically.** `index_repo`, `watch` and `lazy` treat a doc record whose entries lack `terms` as needing re-extraction even when the content hash is unchanged (`docEntriesNeedBackfill`), so the coverage fix reaches stores indexed before 0.5.4 without a manual `reindex: true`. It is a one-shot pass: once entries carry `terms`, the hash skip resumes.
12
+ - **Rule-based keywords** (`extractKeywords`: title-weighted top terms) replace the previous placeholder 5–10 keywords, so `keywords` is deterministic and reproducible. Bilingual keyword expansion and self-reported blind spots are intentionally **not** produced at index time: the `blindSpots` field and its rendering are retained for compatibility, but new entries leave it empty.
13
+ - **Recall-time opt-in LLM is routed explicitly and visibly.** `config.llmQueryExpansion` (default off) and `reflection.enabled` (default off) are the only remaining model calls, and both are recall-time and opt-in. They resolve `provider`/`model` through `src/llm-route.js` (config override → session `requestHeader().config` → agent options → last-seen `request/header` route). When no route is available or a call fails, search falls back explicitly and records a one-time `degraded` entry rather than failing silently.
14
+ - **Tests:** `test/doc-index.test.mjs` (new) pins summary ≤300 / terms covering past char 300 / BM25 hitting a post-300 term / zero model calls at index time even with a routed `exec` / deterministic bounded `extractTerms` / the automatic back-fill. `test/llm-route.test.mjs` covers the recall-time routing contract. `test/run-test.mjs` asserts watch indexing makes **zero** model calls and that doc entries are structural.
15
+
16
+ ### Changed (injection semantics + signals)
17
+
18
+ - **The injected block now declares `form: 'snapshot'`, not `'notice'`.** The autocontext block is current state that a later injection supersedes; `notice` means "a one-off account of something that just happened", so the old declaration was semantically wrong for both the transcript UI and the model. The channel is untouched (still appended through `agent/pre-step` as a plugin `user/message`). The host's `ContextFormed` is a discriminated union (`packages/llm/llm/src/message.ts:81-96`), so `snapshot` carries `sections: [{ name, text }]` and must **not** carry `notice`'s `summary`; the section is named `project-memory` and its `text` is exactly the assembled block.
19
+ - **Per-file `reads` / `writes` counters are now collected.** `task.fileMeta[rel]` gained `reads` and `writes` alongside the retained `n` (which stays the read+write total for old tasks). This is **collect-only**: `hotSortFiles`, the resident "editing now" list and the dedupe fingerprint still read only `lastWriteAt` / `lastReadAt`, so no ordering or injection behaviour changes. It starts collecting the one signal a full-index memory plugin cannot observe — what the model actually read.
20
+ - **`query_memory` memory rows carry a status placeholder.** Each doc/symbol row now prints `- status: <value>`, defaulting to `exact` for every entry (the vocabulary is `exact | needs-verify | re-verified | refreshed | demoted`). No state machine is implemented yet; this is a forward-compatible output contract so real values can land later without changing the rendering again.
21
+ - **Tests:** suite 227 → **231** checks. `test/host-contract.test.mjs` now drives a real injection and pins the source shape (`form: 'snapshot'`, `sections[0].name`, no `summary`); `test/taskbridge.test.mjs` pins `reads`/`writes`/`n`; `test/run-test.mjs` pins the `status` placeholder on doc and symbol rows.
22
+
23
+ ### Files
24
+ - src/doc-index.js (new), src/doc-pipeline.js, src/llm.js, src/llm-route.js (new), src/util/text.js, src/util/search.js, src/link.js, src/index.js, src/watch.js, src/lazy.js, src/reflection-pipeline.js, src/auto-inject.js, src/setup/taskbridge.js, src/tools/index-doc.js, src/tools/index-repo.js, src/tools/query-memory.js, src/tools/task-tools.js, src/commands/task-actions.js, test/doc-index.test.mjs (new), test/llm-route.test.mjs, test/host-contract.test.mjs, test/taskbridge.test.mjs, test/run-test.mjs, test/reflection-pipeline.test.mjs, package.json, CHANGELOG.md
25
+
3
26
  ## 0.5.3 (2026-09-10)
4
27
 
5
28
  ### Changed (DSH 0.1.5-rc.1 compatibility)
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@yolk_vat-y/dsh-project-memory",
3
- "version": "0.5.3",
3
+ "version": "0.5.4",
4
4
  "description": "Persistent project memory for dsh agents: index docs (PDF/Markdown/text) and code symbols into a searchable per-workspace store, recall them with cited sources, and keep experience entries (problems -> solutions) searchable on demand.",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
@@ -17,7 +17,7 @@
17
17
  "url": "https://github.com/00080000/dsh-project-memory.git"
18
18
  },
19
19
  "scripts": {
20
- "test": "node test/run-test.mjs && node test/taskbridge.test.mjs && node test/insight-store.test.mjs && node test/reflection-pipeline.test.mjs && node test/auto-inject.test.mjs && node test/insight-actions.test.mjs && node test/host-contract.test.mjs",
20
+ "test": "node test/run-test.mjs && node test/taskbridge.test.mjs && node test/insight-store.test.mjs && node test/reflection-pipeline.test.mjs && node test/auto-inject.test.mjs && node test/insight-actions.test.mjs && node test/host-contract.test.mjs && node test/llm-route.test.mjs && node test/doc-index.test.mjs",
21
21
  "build:client": "tsdown"
22
22
  },
23
23
  "keywords": [
@@ -204,7 +204,15 @@ export function installAutoInject(ctx, config) {
204
204
  if (lastFpBySession.size > LAST_FP_MAX) lastFpBySession.delete(lastFpBySession.keys().next().value)
205
205
  const injectMessage = createUserMessage({
206
206
  content: [{ type: 'text', text: `\n\n${INJECT_MARK} auto-context\n${built.text}` }],
207
- source: { kind: 'plugin', plugin: 'dsh-project-memory', form: 'notice', summary: `记忆注入 ${built.text.length} 字符` },
207
+ // 这一块是「同一生产者后续快照会取代的当前状态」,不是一次性通知。
208
+ // 宿主 ContextFormed 是判别联合:snapshot 必须带 sections(notice 才需要 summary)。
209
+ // 通道不变(仍走 agent/pre-step 追加 user 消息),只修语义。
210
+ source: {
211
+ kind: 'plugin',
212
+ plugin: 'dsh-project-memory',
213
+ form: 'snapshot',
214
+ sections: [{ name: 'project-memory', text: built.text }],
215
+ },
208
216
  })
209
217
  return { ...decision, messages: [...decision.messages, injectMessage] }
210
218
  }
@@ -7,6 +7,7 @@ import { ProjectMemoryStore } from '../store.js'
7
7
  import { projectRootFor, adoptStepsToSession, shouldAdoptToHost } from '../setup/taskbridge.js'
8
8
  import { renderTaskSnapshot, buildTaskPayload } from './tasks.js'
9
9
  import { fireReflect } from '../reflection-pipeline.js'
10
+ import { resolveRoute } from '../llm-route.js'
10
11
 
11
12
  const VERBS = { switch: 'switch', archive: 'archive' }
12
13
 
@@ -138,7 +139,7 @@ export function taskCommandDefinition(config, ctx) {
138
139
  store.save()
139
140
  // 反思钩子:切走旧任务异步收割(默认关)
140
141
  if (prevBound && prevBound !== task.id) {
141
- fireReflect(config, { llm: ctx?.llm }, root, prevBound, 'switch-away')
142
+ fireReflect(config, { llm: ctx?.llm }, root, prevBound, 'switch-away', resolveRoute(session ? { agent: { session } } : undefined, config))
142
143
  }
143
144
  // 反向接管:切换成功后把任务步骤推成宿主 todo/write(dsh 清单跟随)
144
145
  if (shouldAdoptToHost(config)) {
@@ -156,7 +157,7 @@ export function taskCommandDefinition(config, ctx) {
156
157
  task.updatedAt = new Date().toISOString()
157
158
  store.save()
158
159
  // 反思钩子:归档即收割(默认关)
159
- fireReflect(config, { llm: ctx?.llm }, root, taskId, 'archive')
160
+ fireReflect(config, { llm: ctx?.llm }, root, taskId, 'archive', resolveRoute(session ? { agent: { session } } : undefined, config))
160
161
  const note = `已归档: ${describeTask(task)}(select_task 可恢复)`
161
162
  return withTaskSnapshot(config, cwd, sid, store, note)
162
163
  }
@@ -0,0 +1,71 @@
1
+ // 文档索引的结构化词项(纯规则、确定性;索引期不调用任何模型)。
2
+ //
3
+ // 动机:检索只吃 title/keywords/summary/sourcePath(util/search.js 的 weightedFieldText),
4
+ // 而 summary 为了注入预算被压到 300 字符 —— 一个 3000 字符的 chunk 有九成内容检索不到。
5
+ // 这里把「注入用的短摘要」和「检索用的词项」拆开:summary 仍然短,terms 覆盖整个 chunk。
6
+ //
7
+ // 纯规则、确定性、可重放:tokenizeRaw(拉丁词 + CJK bigram)→ 去停用词/纯数字 → 词频排序 → 截断。
8
+ import { tokenizeRaw } from './util/search.js'
9
+
10
+ export const MAX_TERMS = 160
11
+ export const MAX_KEYWORDS = 8
12
+
13
+ const STOPWORDS = new Set([
14
+ 'the', 'and', 'for', 'are', 'but', 'not', 'you', 'all', 'any', 'can', 'had', 'her', 'was', 'one',
15
+ 'our', 'out', 'day', 'get', 'has', 'him', 'his', 'how', 'its', 'new', 'now', 'old', 'see', 'two',
16
+ 'way', 'who', 'boy', 'did', 'use', 'that', 'this', 'with', 'from', 'they', 'will', 'would', 'there',
17
+ 'their', 'what', 'about', 'which', 'when', 'make', 'like', 'time', 'just', 'know', 'take', 'into',
18
+ 'your', 'some', 'them', 'than', 'then', 'only', 'come', 'over', 'also', 'back', 'after', 'other',
19
+ 'many', 'most', 'such', 'even', 'much', 'more', 'been', 'were', 'have', 'each', 'does', 'doing',
20
+ 'should', 'could', 'these', 'those', 'being', 'where', 'while', 'because', 'before', 'between',
21
+ 'under', 'again', 'further', 'once', 'here', 'both', 'few', 'same', 'too', 'very', 'own', 'off',
22
+ 'per', 'via', 'etc', 'see', 'note', 'used', 'using', 'uses', 'may', 'must', 'shall',
23
+ ])
24
+
25
+ /** 一个 token 是否值得作为检索词项。 */
26
+ function isUseful(token) {
27
+ if (token.length < 2) return false
28
+ if (STOPWORDS.has(token)) return false
29
+ if (/^\d+$/.test(token)) return false
30
+ return true
31
+ }
32
+
33
+ /**
34
+ * 整个 chunk 的字面词项(唯一、按词频排序、有上限)。
35
+ * @returns {string[]} 词项数组(已去重)
36
+ */
37
+ export function extractTerms(text, { max = MAX_TERMS } = {}) {
38
+ const counts = new Map()
39
+ for (const token of tokenizeRaw(text)) {
40
+ if (!isUseful(token)) continue
41
+ counts.set(token, (counts.get(token) || 0) + 1)
42
+ }
43
+ return [...counts.entries()]
44
+ .sort((a, b) => b[1] - a[1] || b[0].length - a[0].length || a[0].localeCompare(b[0]))
45
+ .slice(0, max)
46
+ .map(([token]) => token)
47
+ }
48
+
49
+ /**
50
+ * 结构化的检索词项串(用于 entry.terms,并入 BM25 的检索文本)。
51
+ */
52
+ export function extractTermText(text, { max = MAX_TERMS } = {}) {
53
+ return extractTerms(text, { max }).join(' ')
54
+ }
55
+
56
+ /**
57
+ * 规则化关键词:标题重复一次以取得小幅加权(替代此前依赖 LLM 的 5–10 个关键词)。
58
+ */
59
+ export function extractKeywords(title, text, { max = MAX_KEYWORDS } = {}) {
60
+ const source = title ? `${title} ${title} ${text}` : text
61
+ return extractTerms(source, { max })
62
+ }
63
+
64
+ /**
65
+ * 旧 store 的 doc 条目没有 `terms`:需要一次定向回填,否则内容哈希未变的文件会被跳过,
66
+ * 新的检索覆盖不会生效。只对 doc 条目判断,code 条目(符号表)不需要 terms。
67
+ */
68
+ export function docEntriesNeedBackfill(entries) {
69
+ if (!Array.isArray(entries) || entries.length === 0) return true
70
+ return entries.some((e) => e && e.type === 'doc' && typeof e.terms !== 'string')
71
+ }
@@ -2,9 +2,10 @@ import { createHash } from 'node:crypto'
2
2
  import path from 'node:path'
3
3
  import { stat } from 'node:fs/promises'
4
4
  import { looksLikeDump, readTextFile } from './util/fs.js'
5
+ import { summarizeText } from './util/text.js'
5
6
  import { parsePdf } from './parsers/pdfjs-parser.js'
6
7
  import { chunkText } from './chunker.js'
7
- import { extractDocEntry } from './llm.js'
8
+ import { extractKeywords, extractTermText } from './doc-index.js'
8
9
 
9
10
  export async function extractTextFromFile(filePath, { maxFileSizeMb = 50, maxPdfPages = 1000 } = {}) {
10
11
  const ext = path.extname(filePath).toLowerCase()
@@ -21,43 +22,37 @@ export async function extractTextFromFile(filePath, { maxFileSizeMb = 50, maxPdf
21
22
  return readTextFile(filePath, maxFileSizeMb ? maxFileSizeMb * 1024 * 1024 : Infinity)
22
23
  }
23
24
 
24
- const DOC_CONCURRENCY = 4
25
-
26
- export async function buildDocEntries(llm, a, b, c) {
27
- // Backward compatible: old signature (llm, filePath, opts) or new (llm, relPath, filePath, opts)
28
- const [relPath, filePath, opts] = c === undefined ? [a, a, b] : [a, b, c]
25
+ /**
26
+ * 文档分片 → 记忆条目(索引期不调用任何模型:纯规则、确定性、可重放)。
27
+ *
28
+ * 注入用 summary 与检索用 terms 分离:
29
+ * - summary:≤300 字符,进上下文,保持小预算;
30
+ * - terms:整个 chunk 的字面词项,只进 BM25 检索文本,不进注入。
31
+ * 于是「chunk 只有前 300 字符可检索」的旧限制被移除,且没有任何 LLM 调用。
32
+ * 函数签名里刻意没有 llm —— 索引期零 LLM 由构造保证,而不是靠 catch。
33
+ */
34
+ export async function buildDocEntries(relPath, filePath, opts = {}) {
29
35
  const text = await extractTextFromFile(filePath, opts)
30
36
  if (looksLikeDump(text)) return null
31
-
37
+
32
38
  // Compute content hash for update detection
33
39
  const hash = createHash('sha256').update(text).digest('hex').slice(0, 16)
34
-
40
+
35
41
  const chunks = chunkText(text, opts.chunkChars, opts.maxChunks)
36
- const metas = new Array(chunks.length)
37
- let cursor = 0
38
- await Promise.all(
39
- Array.from({ length: Math.min(DOC_CONCURRENCY, chunks.length) }, () =>
40
- (async () => {
41
- while (cursor < chunks.length) {
42
- const i = cursor++
43
- metas[i] = await extractDocEntry(llm, chunks[i], filePath)
44
- }
45
- })(),
46
- ),
47
- )
48
- return metas.map((meta, i) => ({
42
+ return chunks.map((chunk, i) => ({
49
43
  id: `${relativeId(relPath)}#${i}`,
50
44
  sourcePath: relPath,
51
- sourceLine: chunks[i].line,
45
+ sourceLine: chunk.line,
52
46
  type: 'doc',
53
- title: meta.title,
54
- summary: meta.summary,
55
- blindSpots: meta.blindSpots || '',
56
- keywords: meta.keywords,
47
+ title: chunk.title || relPath,
48
+ summary: summarizeText(chunk.text),
49
+ blindSpots: '',
50
+ keywords: extractKeywords(chunk.title, chunk.text),
51
+ terms: extractTermText(chunk.text),
57
52
  hash,
58
53
  }))
59
54
  }
60
55
 
61
56
  function relativeId(filePath) {
62
57
  return String(filePath).replace(/[\\/:\s]/g, '_')
63
- }
58
+ }
package/src/index.js CHANGED
@@ -13,6 +13,7 @@ import { setupTaskbridge } from './setup/taskbridge.js'
13
13
  import { listTasksTool, selectTaskTool, archiveTaskTool, showTaskPanelTool } from './tools/task-tools.js'
14
14
  import { lessonTool } from './tools/lesson-tools.js'
15
15
  import { installAutoInject } from './auto-inject.js'
16
+ import { rememberRoute } from './llm-route.js'
16
17
  import { tasksCommandDefinition } from './commands/tasks.js'
17
18
  import { taskCommandDefinition } from './commands/task-actions.js'
18
19
  import { insightCommandDefinition } from './commands/insight-actions.js'
@@ -29,6 +30,14 @@ export const Config = Schema.object({
29
30
  maxPdfPages: Schema.number().default(1000),
30
31
  llmQueryExpansion: Schema.boolean().default(false),
31
32
  expansionCount: Schema.number().default(6),
33
+ // 辅助 LLM 调用的路由覆写 —— **索引期零 LLM**,这里的路由只服务于召回期可选项
34
+ // (llmQueryExpansion 查询扩展)与反思期(reflection.enabled)。
35
+ // 未设置时:工具调用取当前会话 header 的 provider/model;后台路径取最近一次会话路由(request/header)。
36
+ // 两者都拿不到时明确走回退并记录 degraded(不静默)。
37
+ llm: Schema.object({
38
+ provider: Schema.string(),
39
+ model: Schema.string(),
40
+ }).default({}),
32
41
  lazyIndexing: Schema.boolean().default(true),
33
42
  autoIndexOnFirstUse: Schema.boolean().default(false),
34
43
  watch: Schema.boolean().default(true),
@@ -51,7 +60,7 @@ export const Config = Schema.object({
51
60
  decayDays: Schema.number().default(90),
52
61
  globalFile: Schema.string(),
53
62
  }).default({}),
54
- // PR 1b:反思管线(LLM 消费点,默认关)。产出只写 task 级草稿,见 PLAN §3
63
+ // 反思管线(召回期可选 LLM 消费点,默认关)。产出只写 task 级草稿。
55
64
  reflection: Schema.object({
56
65
  enabled: Schema.boolean().default(false),
57
66
  cooldownMs: Schema.number().default(1800000),
@@ -88,6 +97,10 @@ export function apply(ctx, config) {
88
97
 
89
98
  // TaskBridge:任务实体 + 宿主 todo 同步
90
99
  setupTaskbridge(ctx, config)
100
+ // 跟踪会话路由(request/header 事件),为无会话上下文的召回期可选 LLM(查询扩展 / 反思)兜底
101
+ ctx.on('session/event', (session, event) => {
102
+ if (event?.type === 'request/header') rememberRoute(session)
103
+ })
91
104
  ctx.tools.register(listTasksTool(config))
92
105
  ctx.tools.register(selectTaskTool(config, { llm: ctx.llm, ctx }))
93
106
  ctx.tools.register(archiveTaskTool(config, { llm: ctx.llm, ctx }))
package/src/lazy.js CHANGED
@@ -3,6 +3,7 @@ import { tmpdir } from 'node:os'
3
3
  import { existsSync, readdirSync, statSync } from 'node:fs'
4
4
  import { isSupportedCode, isSupportedDoc, memoryRootFor, readFileForIndex, relativePath, storeKey } from './util/fs.js'
5
5
  import { buildDocEntries } from './doc-pipeline.js'
6
+ import { docEntriesNeedBackfill } from './doc-index.js'
6
7
  import { scanSymbols } from './symbols.js'
7
8
  import { linkEntries } from './link.js'
8
9
  import { ProjectMemoryStore } from './store.js'
@@ -98,7 +99,8 @@ export async function indexFile(ctx, config, filePath, watchManager = null) {
98
99
  } catch {
99
100
  return false
100
101
  }
101
- if (existing && existing.sha256 === hash) return false
102
+ // 旧 store 的 doc 条目缺 terms → 一次性回填(即使哈希未变)
103
+ if (existing && existing.sha256 === hash && !(isSupportedDoc(ext) && docEntriesNeedBackfill(store.entries[rel]))) return false
102
104
 
103
105
  if (watchManager) {
104
106
  watchManager.addRoot(root)
@@ -115,7 +117,7 @@ export async function indexFile(ctx, config, filePath, watchManager = null) {
115
117
  return true
116
118
  })
117
119
  } else {
118
- entries = await buildDocEntries(ctx.llm, rel, filePath, {
120
+ entries = await buildDocEntries(rel, filePath, {
119
121
  chunkChars: config.chunkChars,
120
122
  maxChunks: config.maxChunksPerFile,
121
123
  maxFileSizeMb: config.maxFileSizeMb,
package/src/link.js CHANGED
@@ -28,7 +28,8 @@ export function linkEntries(store) {
28
28
  let links = 0
29
29
  for (const doc of docs) {
30
30
  const linked = new Set()
31
- const haystack = `${doc.title || ''} ${doc.summary || ''} ${doc.keywords ? doc.keywords.join(' ') : ''}`.toLowerCase()
31
+ // 结构词项也参与链接:terms 覆盖整个 chunk,符号在后半段被提及时同样能链上(与检索同源)
32
+ const haystack = `${doc.title || ''} ${doc.summary || ''} ${doc.terms || ''} ${doc.keywords ? doc.keywords.join(' ') : ''}`.toLowerCase()
32
33
  for (const [, entry] of symbolByName) {
33
34
  const hit = entry.re ? entry.re.test(haystack) : haystack.includes(entry.lower)
34
35
  for (const s of entry.syms) {
@@ -0,0 +1,95 @@
1
+ // 辅助 LLM 调用的路由解析(仅召回期可选:查询扩展 / 反思)。
2
+ //
3
+ // 背景:宿主 LlmRuntime.stream 要求 provider/model 必填,缺失即抛 NO_ADAPTER。
4
+ // 本模块负责:
5
+ // 1) 从工具 exec / 会话 / 配置解析出 (provider, model);
6
+ // 2) 把「无法路由 / 调用失败」记成可见的 degraded 记录,而不是静默降级。
7
+ //
8
+ // 解析优先级(与宿主 compaction summarizer 同序,再补一条后台 hint):
9
+ // config.llm 显式覆写 > exec 会话 requestHeader().config > exec agent options >
10
+ // 最近一次 request/header 事件记下的路由(后台 watch/lazy 兜底)。
11
+ //
12
+ // 登记规则:路由只影响辅助调用走哪个模型,不改变检索/排序行为。
13
+
14
+ /** 最近一次在会话里观察到的路由,供无 exec 的后台索引(watch/lazy)兜底。 */
15
+ let routeHint = null
16
+
17
+ /** 无法路由 / 调用失败留下的可见痕迹:code -> { code, reason, at, count }。 */
18
+ const degraded = new Map()
19
+
20
+ function normalize(provider, model, source) {
21
+ if (typeof provider !== 'string' || !provider) return null
22
+ if (typeof model !== 'string' || !model) return null
23
+ return { provider, model, source }
24
+ }
25
+
26
+ function routeFromSession(session) {
27
+ try {
28
+ const config = session?.requestHeader?.()?.config
29
+ return normalize(config?.provider, config?.model, 'session')
30
+ } catch {
31
+ return null
32
+ }
33
+ }
34
+
35
+ /** 记下会话当前路由,作为后台索引的兜底。由 request/header 事件驱动。 */
36
+ export function rememberRoute(session) {
37
+ const route = routeFromSession(session)
38
+ if (route) routeHint = { provider: route.provider, model: route.model }
39
+ }
40
+
41
+ /** 显式配置的路由(config.llm.provider + config.llm.model 同时存在才算)。 */
42
+ function routeFromConfig(config) {
43
+ return normalize(config?.llm?.provider, config?.llm?.model, 'config')
44
+ }
45
+
46
+ /**
47
+ * 解析一次辅助 LLM 调用的路由。
48
+ * @param exec - 工具执行上下文(可含 agent.session / agent.options),可为空。
49
+ * @param config - 插件配置(可含 llm 覆写)。
50
+ * @returns {{provider: string, model: string, source: string}|null}
51
+ */
52
+ export function resolveRoute(exec, config) {
53
+ const configured = routeFromConfig(config)
54
+ if (configured) return configured
55
+
56
+ const fromExecSession = routeFromSession(exec?.agent?.session)
57
+ if (fromExecSession) return fromExecSession
58
+
59
+ const options = exec?.agent?.options
60
+ const fromOptions = normalize(options?.provider, options?.model, 'agent')
61
+ if (fromOptions) return fromOptions
62
+
63
+ if (routeHint) return { ...routeHint, source: 'session-hint' }
64
+ return null
65
+ }
66
+
67
+ /**
68
+ * 记录一次降级(允许降级,不允许静默)。同一 code 只打印一次,避免逐 chunk 刷屏;
69
+ * 计数与最近原因保留下来,供 memory_stats / 测试读取。
70
+ */
71
+ export function noteDegraded(code, reason) {
72
+ const at = new Date().toISOString()
73
+ const prev = degraded.get(code)
74
+ if (prev) {
75
+ prev.count += 1
76
+ prev.at = at
77
+ prev.reason = reason
78
+ return prev
79
+ }
80
+ const record = { code, reason, at, count: 1 }
81
+ degraded.set(code, record)
82
+ console.warn(`[dsh-project-memory] degraded ${code}: ${reason}`)
83
+ return record
84
+ }
85
+
86
+ /** 当前累计的降级记录(快照,按 code 去重)。 */
87
+ export function degradedList() {
88
+ return [...degraded.values()].map((d) => ({ ...d }))
89
+ }
90
+
91
+ /** 测试用:清空路由 hint 与降级记录。 */
92
+ export function resetRouteState() {
93
+ routeHint = null
94
+ degraded.clear()
95
+ }
package/src/llm.js CHANGED
@@ -1,5 +1,7 @@
1
+ // 辅助 LLM 调用(**仅召回期 / 反思期,按需可选**;索引期不调用任何模型)。
2
+ // chatText 是唯一的宿主调用出口:provider/model 必填,缺失时显式抛错,由调用方决定回退并记 degraded。
1
3
  import { BlockAssembler, createUserMessage } from '@deepseek-ai/dsh-llm'
2
- import { tokenize } from './util/search.js'
4
+ import { noteDegraded } from './llm-route.js'
3
5
 
4
6
  function systemMessage(text) {
5
7
  return { role: 'system', content: [{ type: 'text', text }] }
@@ -13,24 +15,19 @@ function textOf(message) {
13
15
  .join('\n')
14
16
  }
15
17
 
16
- const MAX_SUMMARY = 300
17
-
18
- export function summarizeText(text, max = MAX_SUMMARY) {
19
- const flat = String(text || '').replace(/\s+/g, ' ').trim()
20
- if (!flat) return ''
21
- if (flat.length <= max) return flat
22
- const clip = max - 1
23
- const clipped = flat.slice(0, clip)
24
- const lastBreak = Math.max(clipped.lastIndexOf('。'), clipped.lastIndexOf('.'), clipped.lastIndexOf(';'))
25
- return lastBreak > clip * 0.4 ? clipped.slice(0, lastBreak + 1) : clipped + '…'
26
- }
27
-
28
- export async function chatText(llm, system, user, { timeoutMs = 120000 } = {}) {
18
+ export async function chatText(llm, system, user, { timeoutMs = 120000, route } = {}) {
19
+ // provider/model 是宿主 GenerateOptions 的必填项,缺失时 LlmRuntime 抛 NO_ADAPTER。
20
+ // 这里显式失败(由调用方决定回退并记 degraded),不让异常悄悄消失。
21
+ if (!route?.provider || !route?.model) {
22
+ throw new Error('auxiliary LLM call requires an explicit provider/model route')
23
+ }
29
24
  const assembler = new BlockAssembler()
30
25
  const controller = new AbortController()
31
26
  const timer = setTimeout(() => controller.abort(), timeoutMs)
32
27
  try {
33
28
  for await (const chunk of llm.stream({
29
+ provider: route.provider,
30
+ model: route.model,
34
31
  messages: [systemMessage(system), createUserMessage({ content: [{ type: 'text', text: user }] })],
35
32
  signal: controller.signal,
36
33
  })) {
@@ -72,8 +69,13 @@ function parseJson(text, validate) {
72
69
  return null
73
70
  }
74
71
 
75
- export async function expandQuery(llm, query, count = 6) {
72
+ /** 召回期可选的查询扩展(默认关;开启时才需要 provider/model 路由)。 */
73
+ export async function expandQuery(llm, query, count = 6, { route } = {}) {
76
74
  if (!llm) return [query]
75
+ if (!route) {
76
+ noteDegraded('llm.expand.no-route', 'no provider/model route for query expansion; search uses the raw query')
77
+ return [query]
78
+ }
77
79
  const system =
78
80
  'You are a search-query expander for a codebase/document memory search engine. ' +
79
81
  'Given a user query, return a STRICT JSON array of alternative search queries that ' +
@@ -81,57 +83,14 @@ export async function expandQuery(llm, query, count = 6) {
81
83
  'code identifier guesses, and narrower/longer phrasings. Include the original query first. ' +
82
84
  'Output only the JSON array of strings, no fences, no commentary.'
83
85
  try {
84
- const raw = await chatText(llm, system, `Query: "${query}"\n\nReturn the JSON array.`)
86
+ const raw = await chatText(llm, system, `Query: "${query}"\n\nReturn the JSON array.`, { route })
85
87
  const parsed = parseJsonArray(raw)
86
88
  if (Array.isArray(parsed) && parsed.length) {
87
89
  const variants = parsed.map(String).filter((s) => s.trim()).slice(0, count)
88
90
  if (variants.length) return variants
89
91
  }
90
- } catch {
91
- // fall through to the raw query
92
+ } catch (err) {
93
+ noteDegraded('llm.expand.failed', `query expansion LLM call failed: ${err?.message || err}`)
92
94
  }
93
95
  return [query]
94
96
  }
95
-
96
- export async function extractDocEntry(llm, chunk, sourcePath) {
97
- const system =
98
- 'You are a project-documentation indexer. Given a chunk of a project document, ' +
99
- 'return a STRICT JSON object with exactly four fields: ' +
100
- '"title" (short section title, string), ' +
101
- '"summary" (3-5 sentence dense summary of what this section covers, ' +
102
- 'mentioning concrete names, decisions, constraints, and key technical details), ' +
103
- '"blindSpots" (string describing what this summary does NOT cover, ' +
104
- 'e.g. "未覆盖:部署细节、性能基准、v0.2 前 API", empty string if none), ' +
105
- '"keywords" (array of 5-10 searchable strings: ' +
106
- 'cover the document\'s own language AND English equivalents, ' +
107
- 'so a query in either language can match). ' +
108
- 'Do not include markdown fences, do not add commentary, output only the JSON object.'
109
-
110
- const user =
111
- `Document: ${sourcePath}\nSection: ${chunk.title || '(untitled)'}\n\n` +
112
- `Content:\n${chunk.text.slice(0, 6000)}\n\nReturn the JSON object.`
113
-
114
- const fallback = () => ({
115
- title: chunk.title || sourcePath,
116
- summary: summarizeText(chunk.text),
117
- blindSpots: '',
118
- keywords: tokenize(chunk.title).slice(0, 5),
119
- })
120
-
121
- if (!llm) return fallback()
122
-
123
- try {
124
- const raw = await chatText(llm, system, user)
125
- const parsed = parseStructuredJson(raw)
126
- if (!parsed || typeof parsed.summary !== 'string' || !parsed.summary.trim()) return fallback()
127
- const kw = Array.isArray(parsed.keywords) ? parsed.keywords.map(String).filter((k) => k).slice(0, 8) : []
128
- return {
129
- title: typeof parsed.title === 'string' && parsed.title.trim() ? parsed.title.trim() : chunk.title || sourcePath,
130
- summary: summarizeText(parsed.summary.trim()),
131
- blindSpots: typeof parsed.blindSpots === 'string' ? parsed.blindSpots.trim() : '',
132
- keywords: kw.length ? kw : tokenize(chunk.title).slice(0, 5),
133
- }
134
- } catch {
135
- return fallback()
136
- }
137
- }
@@ -10,6 +10,7 @@ import { ProjectMemoryStore } from './store.js'
10
10
  import { memoryRootFor } from './util/fs.js'
11
11
  import { cfgInsight, saveInsight, GlobalStore, defaultGlobalFile } from './insight-store.js'
12
12
  import { chatText, parseStructuredJson } from './llm.js'
13
+ import { noteDegraded } from './llm-route.js'
13
14
 
14
15
  export const REFLECT_SYSTEM =
15
16
  'You distill task work into durable lessons and decisions for project memory. ' +
@@ -57,7 +58,7 @@ function normalizeList(arr, max) {
57
58
  * 对某任务做一次反思(可归档任务照做——归档正是收割时机)。
58
59
  * 全程 try/catch:调用方 fire-and-forget 即可。
59
60
  */
60
- export async function reflectTaskAfter({ config, llm, root, taskId, reason = 'transition' }) {
61
+ export async function reflectTaskAfter({ config, llm, root, taskId, reason = 'transition', route = null }) {
61
62
  const rc = (config && config.reflection) || {}
62
63
  const fallback = { ok: false, skipped: 'disabled' }
63
64
  if (!rc.enabled) return fallback
@@ -67,6 +68,11 @@ export async function reflectTaskAfter({ config, llm, root, taskId, reason = 'tr
67
68
  const task = store.getTask(taskId)
68
69
  const gate = isReflectDue(config, task)
69
70
  if (!gate.due) return { ok: false, skipped: gate.reason }
71
+ // 路由检查放在门控之后:未到期/内容未变时不应留下 no-route 降级噪声
72
+ if (!route) {
73
+ noteDegraded('llm.reflect.no-route', 'no provider/model route for reflection; skipping this reflection')
74
+ return { ok: false, skipped: 'no-route' }
75
+ }
70
76
 
71
77
  const snap = taskReflectionSnapshot(task)
72
78
  const user =
@@ -75,7 +81,7 @@ export async function reflectTaskAfter({ config, llm, root, taskId, reason = 'tr
75
81
  (snap.files ? `Files touched:\n${snap.files.slice(0, 1000)}\n` : '') +
76
82
  `Trigger: ${reason}\n\nReturn the JSON object.`
77
83
 
78
- const raw = await chatText(llm, REFLECT_SYSTEM, user, { timeoutMs: 90000 })
84
+ const raw = await chatText(llm, REFLECT_SYSTEM, user, { timeoutMs: 90000, route })
79
85
  const parsed = parseStructuredJson(raw)
80
86
  if (!parsed) return { ok: false, skipped: 'unparsable' }
81
87
 
@@ -136,10 +142,10 @@ export async function reflectTaskAfter({ config, llm, root, taskId, reason = 'tr
136
142
  }
137
143
 
138
144
  /** 工具/命令层 fire-and-forget:默认关时立刻廉价返回,绝不让错误上抛。 */
139
- export function fireReflect(config, host, root, taskId, reason) {
145
+ export function fireReflect(config, host, root, taskId, reason, route = null) {
140
146
  const rc = (config && config.reflection) || {}
141
147
  if (!rc.enabled) return Promise.resolve({ ok: false, skipped: 'disabled' })
142
- return reflectTaskAfter({ config, llm: host?.llm, root, taskId, reason }).catch((err) => {
148
+ return reflectTaskAfter({ config, llm: host?.llm, root, taskId, reason, route }).catch((err) => {
143
149
  console.error(`[dsh-project-memory] fireReflect error: ${err?.message || err}`)
144
150
  return { ok: false, skipped: 'error' }
145
151
  })
@@ -146,8 +146,14 @@ export function touchTaskFile(task, rel, kind, now) {
146
146
  }
147
147
  }
148
148
  const m = task.fileMeta[rel] || {}
149
- m.n = (m.n || 0) + 1
150
- if (kind === 'write') m.lastWriteAt = now
149
+ m.n = (m.n || 0) + 1 // 兼容旧字段:读+写总数(保留,避免老任务语义漂移)
150
+ // 读/写分计数(先只采集,不消费 —— 排序与注入仍只用 lastWriteAt/lastReadAt)
151
+ if (kind === 'write') {
152
+ m.writes = (m.writes || 0) + 1
153
+ m.lastWriteAt = now
154
+ } else if (kind === 'read') {
155
+ m.reads = (m.reads || 0) + 1
156
+ }
151
157
  m.lastReadAt = now
152
158
  task.fileMeta[rel] = m
153
159
  task.files = hotSortFiles(task.files, task.fileMeta)
package/src/store.js CHANGED
@@ -158,7 +158,7 @@ export class ProjectMemoryStore {
158
158
  }
159
159
 
160
160
  // ---- v0.5 insights:project 级 insight 文档(任务级在 task.insights[]) ----
161
- // 迁移语义(与初版方案的差异,见 PLAN §10 风险 / 收尾说明):
161
+ // 迁移语义(旧版 store 的兼容策略):
162
162
  // v0.4 experience.json 仍由 remember/forget/query_memory 服务,不删除;
163
163
  // 首次加载把旧笔记**复制导入** insights.json(kind: experience, source: migrate),
164
164
  // migratedAt 落盘保证跨进程/崩溃幂等。销毁式收敛放到 recall 统一 PR。
@@ -11,8 +11,8 @@ export function indexDocTool(ctx, config) {
11
11
  name: 'index_doc',
12
12
  description:
13
13
  'Index a project document (PDF, Markdown, txt) into persistent project memory: split into sections, ' +
14
- 'summarize each with the LLM, and store cited summaries (path + line) for later query_memory recall. ' +
15
- 'Re-indexing the same unchanged file is a no-op (content-hash skip).',
14
+ 'store a short cited summary plus a full-chunk literal term index (no LLM at index time), for later ' +
15
+ 'query_memory recall. Re-indexing the same unchanged file is a no-op (content-hash skip).',
16
16
  parameters: {
17
17
  file_path: {
18
18
  type: 'string',
@@ -42,7 +42,7 @@ export function indexDocTool(ctx, config) {
42
42
  return `Skipped (unchanged): ${rel}\nAlready indexed with ${(store.entries[rel] || []).length} entry/entries.`
43
43
  }
44
44
 
45
- const entries = await buildDocEntries(ctx.llm, rel, filePath, {
45
+ const entries = await buildDocEntries(rel, filePath, {
46
46
  chunkChars: config.chunkChars,
47
47
  maxChunks: config.maxChunksPerFile,
48
48
  maxFileSizeMb: config.maxFileSizeMb,
@@ -3,6 +3,7 @@ import path from 'node:path'
3
3
  import { statSync } from 'node:fs'
4
4
  import { isSupportedCode, isSupportedDoc, looksLikeDump, memoryRootFor, readFileForIndex, relativePath, storeKey, walkDir } from '../util/fs.js'
5
5
  import { buildDocEntries } from '../doc-pipeline.js'
6
+ import { docEntriesNeedBackfill } from '../doc-index.js'
6
7
  import { scanSymbols } from '../symbols.js'
7
8
  import { linkEntries } from '../link.js'
8
9
  import { ProjectMemoryStore } from '../store.js'
@@ -40,7 +41,9 @@ export async function indexRepository(ctx, config, root, { reindex = false } = {
40
41
 
41
42
  // 单次读盘:同一 buffer 供哈希与正文使用(不再 sha256OfFile + readFileSync 读两遍)
42
43
  const { hash, buffer } = readFileForIndex(filePath)
43
- if (!reindex && existing && existing.sha256 === hash) {
44
+ // 旧 store 的 doc 条目缺 terms → 即使哈希未变也重抽一次(一次性回填)
45
+ const needsBackfill = isSupportedDoc(ext) && docEntriesNeedBackfill(store.entries[rel])
46
+ if (!reindex && existing && existing.sha256 === hash && !needsBackfill) {
44
47
  skipped++
45
48
  continue
46
49
  }
@@ -57,7 +60,7 @@ export async function indexRepository(ctx, config, root, { reindex = false } = {
57
60
  skipped++
58
61
  continue
59
62
  }
60
- entries = await buildDocEntries(ctx.llm, rel, filePath, {
63
+ entries = await buildDocEntries(rel, filePath, {
61
64
  chunkChars: config.chunkChars,
62
65
  maxChunks: config.maxChunksPerFile,
63
66
  maxFileSizeMb: config.maxFileSizeMb,
@@ -123,7 +126,8 @@ export function indexRepoTool(ctx, config) {
123
126
  return defineTool({
124
127
  name: 'index_repo',
125
128
  description:
126
- 'Index a whole project into persistent memory. Documents (PDF/Markdown/txt) are summarized by the LLM; ' +
129
+ 'Index a whole project into persistent memory. Documents (PDF/Markdown/txt) get a short cited summary plus ' +
130
+ 'a full-chunk literal term index (no LLM at index time); ' +
127
131
  'code files get a zero-token symbol table (function/class names with line numbers). Incremental: only changed ' +
128
132
  'files are re-extracted (content-hash), deleted files are removed from memory. Call once per project, then query_memory.',
129
133
  parameters: {
@@ -3,6 +3,7 @@ import path from 'node:path'
3
3
  import { memoryRootFor, resolveIndexRoot } from '../util/fs.js'
4
4
  import { ProjectMemoryStore, storeOverview } from '../store.js'
5
5
  import { expandQuery } from '../llm.js'
6
+ import { resolveRoute } from '../llm-route.js'
6
7
  import { rankEntriesMergedScored, rankExperienceScored, rankEntriesStreaming } from '../util/search.js'
7
8
  import { truncate } from '../util/text.js'
8
9
 
@@ -49,7 +50,7 @@ export function queryMemoryTool(ctx, config) {
49
50
  const limit = Math.max(1, Math.min(Number(args.limit) || 8, 20))
50
51
 
51
52
  const queries = config.llmQueryExpansion
52
- ? await expandQuery(ctx.llm, args.query, config.expansionCount)
53
+ ? await expandQuery(ctx.llm, args.query, config.expansionCount, { route: resolveRoute(exec, config) })
53
54
  : [args.query]
54
55
  const symbolById = new Map()
55
56
  for (const e of store.allEntries()) {
@@ -77,7 +78,9 @@ export function queryMemoryTool(ctx, config) {
77
78
  summaryLine += `\n- ⚠️ 摘要未覆盖:${e.blindSpots.replace(/^\s*\/\/\s*未覆盖[::]\s*/, '')}。建议读原文 ${absSource}`
78
79
  }
79
80
  }
80
- lines.push(`### ${e.title} (score: ${rel})\n- source: ${absSource}\n${summaryLine}`)
81
+ // 条目状态占位(可证伪状态机落地前恒为 exact)。
82
+ // 先立字段,后续状态机到位时只改值、不改输出契约。
83
+ lines.push(`### ${e.title} (score: ${rel})\n- source: ${absSource}\n- status: ${e.status || 'exact'}\n${summaryLine}`)
81
84
  if (e.type === 'doc' && Array.isArray(e.linkedSymbols) && e.linkedSymbols.length) {
82
85
  const refs = e.linkedSymbols.slice(0, 5).map((id) => {
83
86
  const s = symbolById.get(id)
@@ -4,6 +4,7 @@ import { ProjectMemoryStore } from '../store.js'
4
4
  import { truncate } from '../util/text.js'
5
5
  import { genTaskId, hash8, adoptStepsToSession, shouldAdoptToHost } from '../setup/taskbridge.js'
6
6
  import { fireReflect } from '../reflection-pipeline.js'
7
+ import { resolveRoute } from '../llm-route.js'
7
8
  import { createHash } from 'node:crypto'
8
9
 
9
10
  function sessionIdOf(exec) {
@@ -137,7 +138,7 @@ export function selectTaskTool(config, host) {
137
138
 
138
139
  // 反思钩子:切走旧任务时异步收割(默认关,零成本;失败只记日志)
139
140
  if (prevBound && prevBound !== task.id) {
140
- fireReflect(config, host, root, prevBound, 'switch-away')
141
+ fireReflect(config, host, root, prevBound, 'switch-away', resolveRoute(exec, config))
141
142
  }
142
143
 
143
144
  const card = {
@@ -191,7 +192,7 @@ export function archiveTaskTool(config, host) {
191
192
  task.updatedAt = new Date().toISOString()
192
193
  store.save()
193
194
  // 反思钩子:归档即收割最终教训(默认关)
194
- fireReflect(config, host, root, args.taskId, 'archive')
195
+ fireReflect(config, host, root, args.taskId, 'archive', resolveRoute(exec, config))
195
196
  return truncate(JSON.stringify({ success: true, archived: true, hint: '已归档,select_task 可恢复' }), config.maxOutputChars)
196
197
  },
197
198
  })
@@ -131,6 +131,8 @@ export function weightedFieldText(entry) {
131
131
  for (let i = 0; i < 5; i++) parts.push(entry.title || '')
132
132
  parts.push((entry.keywords || []).join(' '))
133
133
  parts.push(entry.summary || '')
134
+ // 结构词项:覆盖整个 chunk 的字面词项,把「只有前 300 字符可检索」补回来(索引期零 LLM)
135
+ parts.push(entry.terms || '')
134
136
  parts.push(entry.sourcePath || '')
135
137
  return parts.join(' ')
136
138
  }
package/src/util/text.js CHANGED
@@ -2,4 +2,17 @@ export function truncate(text, maxChars) {
2
2
  if (typeof text !== 'string') return String(text)
3
3
  if (text.length <= maxChars) return text
4
4
  return text.slice(0, maxChars) + `\n\n...[truncated at ${maxChars} chars]`
5
+ }
6
+
7
+ export const MAX_SUMMARY = 300
8
+
9
+ /** 注入用短摘要:压平空白,按句子边界截断到 max 字符。纯函数,无 LLM。 */
10
+ export function summarizeText(text, max = MAX_SUMMARY) {
11
+ const flat = String(text || '').replace(/\s+/g, ' ').trim()
12
+ if (!flat) return ''
13
+ if (flat.length <= max) return flat
14
+ const clip = max - 1
15
+ const clipped = flat.slice(0, clip)
16
+ const lastBreak = Math.max(clipped.lastIndexOf('。'), clipped.lastIndexOf('.'), clipped.lastIndexOf(';'))
17
+ return lastBreak > clip * 0.4 ? clipped.slice(0, lastBreak + 1) : clipped + '…'
5
18
  }
package/src/watch.js CHANGED
@@ -2,6 +2,7 @@ import { existsSync, statSync } from 'node:fs'
2
2
  import path from 'node:path'
3
3
  import { isSupportedCode, isSupportedDoc, memoryRootFor, readFileForIndex, relativePath, storeKey, walkDir } from './util/fs.js'
4
4
  import { buildDocEntries } from './doc-pipeline.js'
5
+ import { docEntriesNeedBackfill } from './doc-index.js'
5
6
  import { scanSymbols } from './symbols.js'
6
7
  import { linkEntries } from './link.js'
7
8
  import { ProjectMemoryStore } from './store.js'
@@ -105,14 +106,16 @@ export class WatchManager {
105
106
  // 单次读盘:同一 buffer 供哈希与正文使用
106
107
  const { hash, buffer } = readFileForIndex(filePath)
107
108
  const existing = state.store.fileRecord(rel)
108
- if (existing && existing.sha256 === hash) continue
109
+ // 旧 store 的 doc 条目缺 terms → 一次性回填(即使哈希未变)
110
+ const needsBackfill = isSupportedDoc(ext) && docEntriesNeedBackfill(state.store.entries[rel])
111
+ if (existing && existing.sha256 === hash && !needsBackfill) continue
109
112
 
110
113
  try {
111
114
  let entries
112
115
  if (isSupportedCode(ext)) {
113
116
  entries = scanSymbols(rel, filePath, buffer.toString('utf8'))
114
117
  } else {
115
- entries = await buildDocEntries(this.ctx.llm, rel, filePath, {
118
+ entries = await buildDocEntries(rel, filePath, {
116
119
  chunkChars: this.config.chunkChars,
117
120
  maxChunks: this.config.maxChunksPerFile,
118
121
  maxFileSizeMb: this.config.maxFileSizeMb,