@yolk_vat-y/dsh-project-memory 0.5.3 → 0.5.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +23 -0
- package/package.json +2 -2
- package/src/auto-inject.js +9 -1
- package/src/commands/task-actions.js +3 -2
- package/src/doc-index.js +71 -0
- package/src/doc-pipeline.js +22 -27
- package/src/index.js +14 -1
- package/src/lazy.js +4 -2
- package/src/link.js +2 -1
- package/src/llm-route.js +95 -0
- package/src/llm.js +20 -61
- package/src/reflection-pipeline.js +10 -4
- package/src/setup/taskbridge.js +8 -2
- package/src/store.js +1 -1
- package/src/tools/index-doc.js +3 -3
- package/src/tools/index-repo.js +7 -3
- package/src/tools/query-memory.js +5 -2
- package/src/tools/task-tools.js +3 -2
- package/src/util/search.js +2 -0
- package/src/util/text.js +13 -0
- package/src/watch.js +5 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,28 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.5.4 (2026-09-12)
|
|
4
|
+
|
|
5
|
+
### Changed (indexing: model-free by design; whole-chunk retrieval terms)
|
|
6
|
+
|
|
7
|
+
- **Indexing never calls a model — by construction.** `buildDocEntries(relPath, filePath, opts)` no longer accepts an `llm` parameter, so index-time code cannot reach a model even by accident. `src/doc-pipeline.js`, `src/watch.js`, `src/lazy.js`, `index_doc` and `index_repo` no longer resolve or pass a route. A stub-model run over 289 files / 2141 entries records **0 calls**.
|
|
8
|
+
- **Retrieval terms cover the whole chunk.** The old shape made the whole chunk's searchable body the ≤300-char `summary` (`util/search.js` `weightedFieldText`), so a 3000-char chunk had roughly 90% of its text invisible to `query_memory`. Each doc entry now also carries `terms` — a bounded (≤160), deterministic, stop-word-filtered literal term set covering the **entire chunk** (`src/doc-index.js`, built from the existing `tokenizeRaw`). `summary` stays short for the injection budget; `terms` is search-only.
|
|
9
|
+
- Measured on this repository's docs (179 chunks): reachable unique chunk terms go from **27.3% to 100%**; terms that appear only beyond character 300 go from **0% to 100%** entryHit@5, while already-answerable queries keep their ranking (MRR **0.958** with `terms` vs **0.955** without). Zero model calls.
|
|
10
|
+
- `linkEntries` now includes `terms` in its haystack for the same reason: a symbol mentioned late in a chunk links again.
|
|
11
|
+
- **Existing stores are back-filled automatically.** `index_repo`, `watch` and `lazy` treat a doc record whose entries lack `terms` as needing re-extraction even when the content hash is unchanged (`docEntriesNeedBackfill`), so the coverage fix reaches stores indexed before 0.5.4 without a manual `reindex: true`. It is a one-shot pass: once entries carry `terms`, the hash skip resumes.
|
|
12
|
+
- **Rule-based keywords** (`extractKeywords`: title-weighted top terms) replace the previous placeholder 5–10 keywords, so `keywords` is deterministic and reproducible. Bilingual keyword expansion and self-reported blind spots are intentionally **not** produced at index time: the `blindSpots` field and its rendering are retained for compatibility, but new entries leave it empty.
|
|
13
|
+
- **Recall-time opt-in LLM is routed explicitly and visibly.** `config.llmQueryExpansion` (default off) and `reflection.enabled` (default off) are the only remaining model calls, and both are recall-time and opt-in. They resolve `provider`/`model` through `src/llm-route.js` (config override → session `requestHeader().config` → agent options → last-seen `request/header` route). When no route is available or a call fails, search falls back explicitly and records a one-time `degraded` entry rather than failing silently.
|
|
14
|
+
- **Tests:** `test/doc-index.test.mjs` (new) pins summary ≤300 / terms covering past char 300 / BM25 hitting a post-300 term / zero model calls at index time even with a routed `exec` / deterministic bounded `extractTerms` / the automatic back-fill. `test/llm-route.test.mjs` covers the recall-time routing contract. `test/run-test.mjs` asserts watch indexing makes **zero** model calls and that doc entries are structural.
|
|
15
|
+
|
|
16
|
+
### Changed (injection semantics + signals)
|
|
17
|
+
|
|
18
|
+
- **The injected block now declares `form: 'snapshot'`, not `'notice'`.** The autocontext block is current state that a later injection supersedes; `notice` means "a one-off account of something that just happened", so the old declaration was semantically wrong for both the transcript UI and the model. The channel is untouched (still appended through `agent/pre-step` as a plugin `user/message`). The host's `ContextFormed` is a discriminated union (`packages/llm/llm/src/message.ts:81-96`), so `snapshot` carries `sections: [{ name, text }]` and must **not** carry `notice`'s `summary`; the section is named `project-memory` and its `text` is exactly the assembled block.
|
|
19
|
+
- **Per-file `reads` / `writes` counters are now collected.** `task.fileMeta[rel]` gained `reads` and `writes` alongside the retained `n` (which stays the read+write total for old tasks). This is **collect-only**: `hotSortFiles`, the resident "editing now" list and the dedupe fingerprint still read only `lastWriteAt` / `lastReadAt`, so no ordering or injection behaviour changes. It starts collecting the one signal a full-index memory plugin cannot observe — what the model actually read.
|
|
20
|
+
- **`query_memory` memory rows carry a status placeholder.** Each doc/symbol row now prints `- status: <value>`, defaulting to `exact` for every entry (the vocabulary is `exact | needs-verify | re-verified | refreshed | demoted`). No state machine is implemented yet; this is a forward-compatible output contract so real values can land later without changing the rendering again.
|
|
21
|
+
- **Tests:** suite 227 → **231** checks. `test/host-contract.test.mjs` now drives a real injection and pins the source shape (`form: 'snapshot'`, `sections[0].name`, no `summary`); `test/taskbridge.test.mjs` pins `reads`/`writes`/`n`; `test/run-test.mjs` pins the `status` placeholder on doc and symbol rows.
|
|
22
|
+
|
|
23
|
+
### Files
|
|
24
|
+
- src/doc-index.js (new), src/doc-pipeline.js, src/llm.js, src/llm-route.js (new), src/util/text.js, src/util/search.js, src/link.js, src/index.js, src/watch.js, src/lazy.js, src/reflection-pipeline.js, src/auto-inject.js, src/setup/taskbridge.js, src/tools/index-doc.js, src/tools/index-repo.js, src/tools/query-memory.js, src/tools/task-tools.js, src/commands/task-actions.js, test/doc-index.test.mjs (new), test/llm-route.test.mjs, test/host-contract.test.mjs, test/taskbridge.test.mjs, test/run-test.mjs, test/reflection-pipeline.test.mjs, package.json, CHANGELOG.md
|
|
25
|
+
|
|
3
26
|
## 0.5.3 (2026-09-10)
|
|
4
27
|
|
|
5
28
|
### Changed (DSH 0.1.5-rc.1 compatibility)
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@yolk_vat-y/dsh-project-memory",
|
|
3
|
-
"version": "0.5.
|
|
3
|
+
"version": "0.5.4",
|
|
4
4
|
"description": "Persistent project memory for dsh agents: index docs (PDF/Markdown/text) and code symbols into a searchable per-workspace store, recall them with cited sources, and keep experience entries (problems -> solutions) searchable on demand.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "src/index.js",
|
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
"url": "https://github.com/00080000/dsh-project-memory.git"
|
|
18
18
|
},
|
|
19
19
|
"scripts": {
|
|
20
|
-
"test": "node test/run-test.mjs && node test/taskbridge.test.mjs && node test/insight-store.test.mjs && node test/reflection-pipeline.test.mjs && node test/auto-inject.test.mjs && node test/insight-actions.test.mjs && node test/host-contract.test.mjs",
|
|
20
|
+
"test": "node test/run-test.mjs && node test/taskbridge.test.mjs && node test/insight-store.test.mjs && node test/reflection-pipeline.test.mjs && node test/auto-inject.test.mjs && node test/insight-actions.test.mjs && node test/host-contract.test.mjs && node test/llm-route.test.mjs && node test/doc-index.test.mjs",
|
|
21
21
|
"build:client": "tsdown"
|
|
22
22
|
},
|
|
23
23
|
"keywords": [
|
package/src/auto-inject.js
CHANGED
|
@@ -204,7 +204,15 @@ export function installAutoInject(ctx, config) {
|
|
|
204
204
|
if (lastFpBySession.size > LAST_FP_MAX) lastFpBySession.delete(lastFpBySession.keys().next().value)
|
|
205
205
|
const injectMessage = createUserMessage({
|
|
206
206
|
content: [{ type: 'text', text: `\n\n${INJECT_MARK} auto-context\n${built.text}` }],
|
|
207
|
-
|
|
207
|
+
// 这一块是「同一生产者后续快照会取代的当前状态」,不是一次性通知。
|
|
208
|
+
// 宿主 ContextFormed 是判别联合:snapshot 必须带 sections(notice 才需要 summary)。
|
|
209
|
+
// 通道不变(仍走 agent/pre-step 追加 user 消息),只修语义。
|
|
210
|
+
source: {
|
|
211
|
+
kind: 'plugin',
|
|
212
|
+
plugin: 'dsh-project-memory',
|
|
213
|
+
form: 'snapshot',
|
|
214
|
+
sections: [{ name: 'project-memory', text: built.text }],
|
|
215
|
+
},
|
|
208
216
|
})
|
|
209
217
|
return { ...decision, messages: [...decision.messages, injectMessage] }
|
|
210
218
|
}
|
|
@@ -7,6 +7,7 @@ import { ProjectMemoryStore } from '../store.js'
|
|
|
7
7
|
import { projectRootFor, adoptStepsToSession, shouldAdoptToHost } from '../setup/taskbridge.js'
|
|
8
8
|
import { renderTaskSnapshot, buildTaskPayload } from './tasks.js'
|
|
9
9
|
import { fireReflect } from '../reflection-pipeline.js'
|
|
10
|
+
import { resolveRoute } from '../llm-route.js'
|
|
10
11
|
|
|
11
12
|
const VERBS = { switch: 'switch', archive: 'archive' }
|
|
12
13
|
|
|
@@ -138,7 +139,7 @@ export function taskCommandDefinition(config, ctx) {
|
|
|
138
139
|
store.save()
|
|
139
140
|
// 反思钩子:切走旧任务异步收割(默认关)
|
|
140
141
|
if (prevBound && prevBound !== task.id) {
|
|
141
|
-
fireReflect(config, { llm: ctx?.llm }, root, prevBound, 'switch-away')
|
|
142
|
+
fireReflect(config, { llm: ctx?.llm }, root, prevBound, 'switch-away', resolveRoute(session ? { agent: { session } } : undefined, config))
|
|
142
143
|
}
|
|
143
144
|
// 反向接管:切换成功后把任务步骤推成宿主 todo/write(dsh 清单跟随)
|
|
144
145
|
if (shouldAdoptToHost(config)) {
|
|
@@ -156,7 +157,7 @@ export function taskCommandDefinition(config, ctx) {
|
|
|
156
157
|
task.updatedAt = new Date().toISOString()
|
|
157
158
|
store.save()
|
|
158
159
|
// 反思钩子:归档即收割(默认关)
|
|
159
|
-
fireReflect(config, { llm: ctx?.llm }, root, taskId, 'archive')
|
|
160
|
+
fireReflect(config, { llm: ctx?.llm }, root, taskId, 'archive', resolveRoute(session ? { agent: { session } } : undefined, config))
|
|
160
161
|
const note = `已归档: ${describeTask(task)}(select_task 可恢复)`
|
|
161
162
|
return withTaskSnapshot(config, cwd, sid, store, note)
|
|
162
163
|
}
|
package/src/doc-index.js
ADDED
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
// 文档索引的结构化词项(纯规则、确定性;索引期不调用任何模型)。
|
|
2
|
+
//
|
|
3
|
+
// 动机:检索只吃 title/keywords/summary/sourcePath(util/search.js 的 weightedFieldText),
|
|
4
|
+
// 而 summary 为了注入预算被压到 300 字符 —— 一个 3000 字符的 chunk 有九成内容检索不到。
|
|
5
|
+
// 这里把「注入用的短摘要」和「检索用的词项」拆开:summary 仍然短,terms 覆盖整个 chunk。
|
|
6
|
+
//
|
|
7
|
+
// 纯规则、确定性、可重放:tokenizeRaw(拉丁词 + CJK bigram)→ 去停用词/纯数字 → 词频排序 → 截断。
|
|
8
|
+
import { tokenizeRaw } from './util/search.js'
|
|
9
|
+
|
|
10
|
+
export const MAX_TERMS = 160
|
|
11
|
+
export const MAX_KEYWORDS = 8
|
|
12
|
+
|
|
13
|
+
const STOPWORDS = new Set([
|
|
14
|
+
'the', 'and', 'for', 'are', 'but', 'not', 'you', 'all', 'any', 'can', 'had', 'her', 'was', 'one',
|
|
15
|
+
'our', 'out', 'day', 'get', 'has', 'him', 'his', 'how', 'its', 'new', 'now', 'old', 'see', 'two',
|
|
16
|
+
'way', 'who', 'boy', 'did', 'use', 'that', 'this', 'with', 'from', 'they', 'will', 'would', 'there',
|
|
17
|
+
'their', 'what', 'about', 'which', 'when', 'make', 'like', 'time', 'just', 'know', 'take', 'into',
|
|
18
|
+
'your', 'some', 'them', 'than', 'then', 'only', 'come', 'over', 'also', 'back', 'after', 'other',
|
|
19
|
+
'many', 'most', 'such', 'even', 'much', 'more', 'been', 'were', 'have', 'each', 'does', 'doing',
|
|
20
|
+
'should', 'could', 'these', 'those', 'being', 'where', 'while', 'because', 'before', 'between',
|
|
21
|
+
'under', 'again', 'further', 'once', 'here', 'both', 'few', 'same', 'too', 'very', 'own', 'off',
|
|
22
|
+
'per', 'via', 'etc', 'see', 'note', 'used', 'using', 'uses', 'may', 'must', 'shall',
|
|
23
|
+
])
|
|
24
|
+
|
|
25
|
+
/** 一个 token 是否值得作为检索词项。 */
|
|
26
|
+
function isUseful(token) {
|
|
27
|
+
if (token.length < 2) return false
|
|
28
|
+
if (STOPWORDS.has(token)) return false
|
|
29
|
+
if (/^\d+$/.test(token)) return false
|
|
30
|
+
return true
|
|
31
|
+
}
|
|
32
|
+
|
|
33
|
+
/**
|
|
34
|
+
* 整个 chunk 的字面词项(唯一、按词频排序、有上限)。
|
|
35
|
+
* @returns {string[]} 词项数组(已去重)
|
|
36
|
+
*/
|
|
37
|
+
export function extractTerms(text, { max = MAX_TERMS } = {}) {
|
|
38
|
+
const counts = new Map()
|
|
39
|
+
for (const token of tokenizeRaw(text)) {
|
|
40
|
+
if (!isUseful(token)) continue
|
|
41
|
+
counts.set(token, (counts.get(token) || 0) + 1)
|
|
42
|
+
}
|
|
43
|
+
return [...counts.entries()]
|
|
44
|
+
.sort((a, b) => b[1] - a[1] || b[0].length - a[0].length || a[0].localeCompare(b[0]))
|
|
45
|
+
.slice(0, max)
|
|
46
|
+
.map(([token]) => token)
|
|
47
|
+
}
|
|
48
|
+
|
|
49
|
+
/**
|
|
50
|
+
* 结构化的检索词项串(用于 entry.terms,并入 BM25 的检索文本)。
|
|
51
|
+
*/
|
|
52
|
+
export function extractTermText(text, { max = MAX_TERMS } = {}) {
|
|
53
|
+
return extractTerms(text, { max }).join(' ')
|
|
54
|
+
}
|
|
55
|
+
|
|
56
|
+
/**
|
|
57
|
+
* 规则化关键词:标题重复一次以取得小幅加权(替代此前依赖 LLM 的 5–10 个关键词)。
|
|
58
|
+
*/
|
|
59
|
+
export function extractKeywords(title, text, { max = MAX_KEYWORDS } = {}) {
|
|
60
|
+
const source = title ? `${title} ${title} ${text}` : text
|
|
61
|
+
return extractTerms(source, { max })
|
|
62
|
+
}
|
|
63
|
+
|
|
64
|
+
/**
|
|
65
|
+
* 旧 store 的 doc 条目没有 `terms`:需要一次定向回填,否则内容哈希未变的文件会被跳过,
|
|
66
|
+
* 新的检索覆盖不会生效。只对 doc 条目判断,code 条目(符号表)不需要 terms。
|
|
67
|
+
*/
|
|
68
|
+
export function docEntriesNeedBackfill(entries) {
|
|
69
|
+
if (!Array.isArray(entries) || entries.length === 0) return true
|
|
70
|
+
return entries.some((e) => e && e.type === 'doc' && typeof e.terms !== 'string')
|
|
71
|
+
}
|
package/src/doc-pipeline.js
CHANGED
|
@@ -2,9 +2,10 @@ import { createHash } from 'node:crypto'
|
|
|
2
2
|
import path from 'node:path'
|
|
3
3
|
import { stat } from 'node:fs/promises'
|
|
4
4
|
import { looksLikeDump, readTextFile } from './util/fs.js'
|
|
5
|
+
import { summarizeText } from './util/text.js'
|
|
5
6
|
import { parsePdf } from './parsers/pdfjs-parser.js'
|
|
6
7
|
import { chunkText } from './chunker.js'
|
|
7
|
-
import {
|
|
8
|
+
import { extractKeywords, extractTermText } from './doc-index.js'
|
|
8
9
|
|
|
9
10
|
export async function extractTextFromFile(filePath, { maxFileSizeMb = 50, maxPdfPages = 1000 } = {}) {
|
|
10
11
|
const ext = path.extname(filePath).toLowerCase()
|
|
@@ -21,43 +22,37 @@ export async function extractTextFromFile(filePath, { maxFileSizeMb = 50, maxPdf
|
|
|
21
22
|
return readTextFile(filePath, maxFileSizeMb ? maxFileSizeMb * 1024 * 1024 : Infinity)
|
|
22
23
|
}
|
|
23
24
|
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
25
|
+
/**
|
|
26
|
+
* 文档分片 → 记忆条目(索引期不调用任何模型:纯规则、确定性、可重放)。
|
|
27
|
+
*
|
|
28
|
+
* 注入用 summary 与检索用 terms 分离:
|
|
29
|
+
* - summary:≤300 字符,进上下文,保持小预算;
|
|
30
|
+
* - terms:整个 chunk 的字面词项,只进 BM25 检索文本,不进注入。
|
|
31
|
+
* 于是「chunk 只有前 300 字符可检索」的旧限制被移除,且没有任何 LLM 调用。
|
|
32
|
+
* 函数签名里刻意没有 llm —— 索引期零 LLM 由构造保证,而不是靠 catch。
|
|
33
|
+
*/
|
|
34
|
+
export async function buildDocEntries(relPath, filePath, opts = {}) {
|
|
29
35
|
const text = await extractTextFromFile(filePath, opts)
|
|
30
36
|
if (looksLikeDump(text)) return null
|
|
31
|
-
|
|
37
|
+
|
|
32
38
|
// Compute content hash for update detection
|
|
33
39
|
const hash = createHash('sha256').update(text).digest('hex').slice(0, 16)
|
|
34
|
-
|
|
40
|
+
|
|
35
41
|
const chunks = chunkText(text, opts.chunkChars, opts.maxChunks)
|
|
36
|
-
|
|
37
|
-
let cursor = 0
|
|
38
|
-
await Promise.all(
|
|
39
|
-
Array.from({ length: Math.min(DOC_CONCURRENCY, chunks.length) }, () =>
|
|
40
|
-
(async () => {
|
|
41
|
-
while (cursor < chunks.length) {
|
|
42
|
-
const i = cursor++
|
|
43
|
-
metas[i] = await extractDocEntry(llm, chunks[i], filePath)
|
|
44
|
-
}
|
|
45
|
-
})(),
|
|
46
|
-
),
|
|
47
|
-
)
|
|
48
|
-
return metas.map((meta, i) => ({
|
|
42
|
+
return chunks.map((chunk, i) => ({
|
|
49
43
|
id: `${relativeId(relPath)}#${i}`,
|
|
50
44
|
sourcePath: relPath,
|
|
51
|
-
sourceLine:
|
|
45
|
+
sourceLine: chunk.line,
|
|
52
46
|
type: 'doc',
|
|
53
|
-
title:
|
|
54
|
-
summary:
|
|
55
|
-
blindSpots:
|
|
56
|
-
keywords:
|
|
47
|
+
title: chunk.title || relPath,
|
|
48
|
+
summary: summarizeText(chunk.text),
|
|
49
|
+
blindSpots: '',
|
|
50
|
+
keywords: extractKeywords(chunk.title, chunk.text),
|
|
51
|
+
terms: extractTermText(chunk.text),
|
|
57
52
|
hash,
|
|
58
53
|
}))
|
|
59
54
|
}
|
|
60
55
|
|
|
61
56
|
function relativeId(filePath) {
|
|
62
57
|
return String(filePath).replace(/[\\/:\s]/g, '_')
|
|
63
|
-
}
|
|
58
|
+
}
|
package/src/index.js
CHANGED
|
@@ -13,6 +13,7 @@ import { setupTaskbridge } from './setup/taskbridge.js'
|
|
|
13
13
|
import { listTasksTool, selectTaskTool, archiveTaskTool, showTaskPanelTool } from './tools/task-tools.js'
|
|
14
14
|
import { lessonTool } from './tools/lesson-tools.js'
|
|
15
15
|
import { installAutoInject } from './auto-inject.js'
|
|
16
|
+
import { rememberRoute } from './llm-route.js'
|
|
16
17
|
import { tasksCommandDefinition } from './commands/tasks.js'
|
|
17
18
|
import { taskCommandDefinition } from './commands/task-actions.js'
|
|
18
19
|
import { insightCommandDefinition } from './commands/insight-actions.js'
|
|
@@ -29,6 +30,14 @@ export const Config = Schema.object({
|
|
|
29
30
|
maxPdfPages: Schema.number().default(1000),
|
|
30
31
|
llmQueryExpansion: Schema.boolean().default(false),
|
|
31
32
|
expansionCount: Schema.number().default(6),
|
|
33
|
+
// 辅助 LLM 调用的路由覆写 —— **索引期零 LLM**,这里的路由只服务于召回期可选项
|
|
34
|
+
// (llmQueryExpansion 查询扩展)与反思期(reflection.enabled)。
|
|
35
|
+
// 未设置时:工具调用取当前会话 header 的 provider/model;后台路径取最近一次会话路由(request/header)。
|
|
36
|
+
// 两者都拿不到时明确走回退并记录 degraded(不静默)。
|
|
37
|
+
llm: Schema.object({
|
|
38
|
+
provider: Schema.string(),
|
|
39
|
+
model: Schema.string(),
|
|
40
|
+
}).default({}),
|
|
32
41
|
lazyIndexing: Schema.boolean().default(true),
|
|
33
42
|
autoIndexOnFirstUse: Schema.boolean().default(false),
|
|
34
43
|
watch: Schema.boolean().default(true),
|
|
@@ -51,7 +60,7 @@ export const Config = Schema.object({
|
|
|
51
60
|
decayDays: Schema.number().default(90),
|
|
52
61
|
globalFile: Schema.string(),
|
|
53
62
|
}).default({}),
|
|
54
|
-
//
|
|
63
|
+
// 反思管线(召回期可选 LLM 消费点,默认关)。产出只写 task 级草稿。
|
|
55
64
|
reflection: Schema.object({
|
|
56
65
|
enabled: Schema.boolean().default(false),
|
|
57
66
|
cooldownMs: Schema.number().default(1800000),
|
|
@@ -88,6 +97,10 @@ export function apply(ctx, config) {
|
|
|
88
97
|
|
|
89
98
|
// TaskBridge:任务实体 + 宿主 todo 同步
|
|
90
99
|
setupTaskbridge(ctx, config)
|
|
100
|
+
// 跟踪会话路由(request/header 事件),为无会话上下文的召回期可选 LLM(查询扩展 / 反思)兜底
|
|
101
|
+
ctx.on('session/event', (session, event) => {
|
|
102
|
+
if (event?.type === 'request/header') rememberRoute(session)
|
|
103
|
+
})
|
|
91
104
|
ctx.tools.register(listTasksTool(config))
|
|
92
105
|
ctx.tools.register(selectTaskTool(config, { llm: ctx.llm, ctx }))
|
|
93
106
|
ctx.tools.register(archiveTaskTool(config, { llm: ctx.llm, ctx }))
|
package/src/lazy.js
CHANGED
|
@@ -3,6 +3,7 @@ import { tmpdir } from 'node:os'
|
|
|
3
3
|
import { existsSync, readdirSync, statSync } from 'node:fs'
|
|
4
4
|
import { isSupportedCode, isSupportedDoc, memoryRootFor, readFileForIndex, relativePath, storeKey } from './util/fs.js'
|
|
5
5
|
import { buildDocEntries } from './doc-pipeline.js'
|
|
6
|
+
import { docEntriesNeedBackfill } from './doc-index.js'
|
|
6
7
|
import { scanSymbols } from './symbols.js'
|
|
7
8
|
import { linkEntries } from './link.js'
|
|
8
9
|
import { ProjectMemoryStore } from './store.js'
|
|
@@ -98,7 +99,8 @@ export async function indexFile(ctx, config, filePath, watchManager = null) {
|
|
|
98
99
|
} catch {
|
|
99
100
|
return false
|
|
100
101
|
}
|
|
101
|
-
|
|
102
|
+
// 旧 store 的 doc 条目缺 terms → 一次性回填(即使哈希未变)
|
|
103
|
+
if (existing && existing.sha256 === hash && !(isSupportedDoc(ext) && docEntriesNeedBackfill(store.entries[rel]))) return false
|
|
102
104
|
|
|
103
105
|
if (watchManager) {
|
|
104
106
|
watchManager.addRoot(root)
|
|
@@ -115,7 +117,7 @@ export async function indexFile(ctx, config, filePath, watchManager = null) {
|
|
|
115
117
|
return true
|
|
116
118
|
})
|
|
117
119
|
} else {
|
|
118
|
-
entries = await buildDocEntries(
|
|
120
|
+
entries = await buildDocEntries(rel, filePath, {
|
|
119
121
|
chunkChars: config.chunkChars,
|
|
120
122
|
maxChunks: config.maxChunksPerFile,
|
|
121
123
|
maxFileSizeMb: config.maxFileSizeMb,
|
package/src/link.js
CHANGED
|
@@ -28,7 +28,8 @@ export function linkEntries(store) {
|
|
|
28
28
|
let links = 0
|
|
29
29
|
for (const doc of docs) {
|
|
30
30
|
const linked = new Set()
|
|
31
|
-
|
|
31
|
+
// 结构词项也参与链接:terms 覆盖整个 chunk,符号在后半段被提及时同样能链上(与检索同源)
|
|
32
|
+
const haystack = `${doc.title || ''} ${doc.summary || ''} ${doc.terms || ''} ${doc.keywords ? doc.keywords.join(' ') : ''}`.toLowerCase()
|
|
32
33
|
for (const [, entry] of symbolByName) {
|
|
33
34
|
const hit = entry.re ? entry.re.test(haystack) : haystack.includes(entry.lower)
|
|
34
35
|
for (const s of entry.syms) {
|
package/src/llm-route.js
ADDED
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
// 辅助 LLM 调用的路由解析(仅召回期可选:查询扩展 / 反思)。
|
|
2
|
+
//
|
|
3
|
+
// 背景:宿主 LlmRuntime.stream 要求 provider/model 必填,缺失即抛 NO_ADAPTER。
|
|
4
|
+
// 本模块负责:
|
|
5
|
+
// 1) 从工具 exec / 会话 / 配置解析出 (provider, model);
|
|
6
|
+
// 2) 把「无法路由 / 调用失败」记成可见的 degraded 记录,而不是静默降级。
|
|
7
|
+
//
|
|
8
|
+
// 解析优先级(与宿主 compaction summarizer 同序,再补一条后台 hint):
|
|
9
|
+
// config.llm 显式覆写 > exec 会话 requestHeader().config > exec agent options >
|
|
10
|
+
// 最近一次 request/header 事件记下的路由(后台 watch/lazy 兜底)。
|
|
11
|
+
//
|
|
12
|
+
// 登记规则:路由只影响辅助调用走哪个模型,不改变检索/排序行为。
|
|
13
|
+
|
|
14
|
+
/** 最近一次在会话里观察到的路由,供无 exec 的后台索引(watch/lazy)兜底。 */
|
|
15
|
+
let routeHint = null
|
|
16
|
+
|
|
17
|
+
/** 无法路由 / 调用失败留下的可见痕迹:code -> { code, reason, at, count }。 */
|
|
18
|
+
const degraded = new Map()
|
|
19
|
+
|
|
20
|
+
function normalize(provider, model, source) {
|
|
21
|
+
if (typeof provider !== 'string' || !provider) return null
|
|
22
|
+
if (typeof model !== 'string' || !model) return null
|
|
23
|
+
return { provider, model, source }
|
|
24
|
+
}
|
|
25
|
+
|
|
26
|
+
function routeFromSession(session) {
|
|
27
|
+
try {
|
|
28
|
+
const config = session?.requestHeader?.()?.config
|
|
29
|
+
return normalize(config?.provider, config?.model, 'session')
|
|
30
|
+
} catch {
|
|
31
|
+
return null
|
|
32
|
+
}
|
|
33
|
+
}
|
|
34
|
+
|
|
35
|
+
/** 记下会话当前路由,作为后台索引的兜底。由 request/header 事件驱动。 */
|
|
36
|
+
export function rememberRoute(session) {
|
|
37
|
+
const route = routeFromSession(session)
|
|
38
|
+
if (route) routeHint = { provider: route.provider, model: route.model }
|
|
39
|
+
}
|
|
40
|
+
|
|
41
|
+
/** 显式配置的路由(config.llm.provider + config.llm.model 同时存在才算)。 */
|
|
42
|
+
function routeFromConfig(config) {
|
|
43
|
+
return normalize(config?.llm?.provider, config?.llm?.model, 'config')
|
|
44
|
+
}
|
|
45
|
+
|
|
46
|
+
/**
|
|
47
|
+
* 解析一次辅助 LLM 调用的路由。
|
|
48
|
+
* @param exec - 工具执行上下文(可含 agent.session / agent.options),可为空。
|
|
49
|
+
* @param config - 插件配置(可含 llm 覆写)。
|
|
50
|
+
* @returns {{provider: string, model: string, source: string}|null}
|
|
51
|
+
*/
|
|
52
|
+
export function resolveRoute(exec, config) {
|
|
53
|
+
const configured = routeFromConfig(config)
|
|
54
|
+
if (configured) return configured
|
|
55
|
+
|
|
56
|
+
const fromExecSession = routeFromSession(exec?.agent?.session)
|
|
57
|
+
if (fromExecSession) return fromExecSession
|
|
58
|
+
|
|
59
|
+
const options = exec?.agent?.options
|
|
60
|
+
const fromOptions = normalize(options?.provider, options?.model, 'agent')
|
|
61
|
+
if (fromOptions) return fromOptions
|
|
62
|
+
|
|
63
|
+
if (routeHint) return { ...routeHint, source: 'session-hint' }
|
|
64
|
+
return null
|
|
65
|
+
}
|
|
66
|
+
|
|
67
|
+
/**
|
|
68
|
+
* 记录一次降级(允许降级,不允许静默)。同一 code 只打印一次,避免逐 chunk 刷屏;
|
|
69
|
+
* 计数与最近原因保留下来,供 memory_stats / 测试读取。
|
|
70
|
+
*/
|
|
71
|
+
export function noteDegraded(code, reason) {
|
|
72
|
+
const at = new Date().toISOString()
|
|
73
|
+
const prev = degraded.get(code)
|
|
74
|
+
if (prev) {
|
|
75
|
+
prev.count += 1
|
|
76
|
+
prev.at = at
|
|
77
|
+
prev.reason = reason
|
|
78
|
+
return prev
|
|
79
|
+
}
|
|
80
|
+
const record = { code, reason, at, count: 1 }
|
|
81
|
+
degraded.set(code, record)
|
|
82
|
+
console.warn(`[dsh-project-memory] degraded ${code}: ${reason}`)
|
|
83
|
+
return record
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
/** 当前累计的降级记录(快照,按 code 去重)。 */
|
|
87
|
+
export function degradedList() {
|
|
88
|
+
return [...degraded.values()].map((d) => ({ ...d }))
|
|
89
|
+
}
|
|
90
|
+
|
|
91
|
+
/** 测试用:清空路由 hint 与降级记录。 */
|
|
92
|
+
export function resetRouteState() {
|
|
93
|
+
routeHint = null
|
|
94
|
+
degraded.clear()
|
|
95
|
+
}
|
package/src/llm.js
CHANGED
|
@@ -1,5 +1,7 @@
|
|
|
1
|
+
// 辅助 LLM 调用(**仅召回期 / 反思期,按需可选**;索引期不调用任何模型)。
|
|
2
|
+
// chatText 是唯一的宿主调用出口:provider/model 必填,缺失时显式抛错,由调用方决定回退并记 degraded。
|
|
1
3
|
import { BlockAssembler, createUserMessage } from '@deepseek-ai/dsh-llm'
|
|
2
|
-
import {
|
|
4
|
+
import { noteDegraded } from './llm-route.js'
|
|
3
5
|
|
|
4
6
|
function systemMessage(text) {
|
|
5
7
|
return { role: 'system', content: [{ type: 'text', text }] }
|
|
@@ -13,24 +15,19 @@ function textOf(message) {
|
|
|
13
15
|
.join('\n')
|
|
14
16
|
}
|
|
15
17
|
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
const clip = max - 1
|
|
23
|
-
const clipped = flat.slice(0, clip)
|
|
24
|
-
const lastBreak = Math.max(clipped.lastIndexOf('。'), clipped.lastIndexOf('.'), clipped.lastIndexOf(';'))
|
|
25
|
-
return lastBreak > clip * 0.4 ? clipped.slice(0, lastBreak + 1) : clipped + '…'
|
|
26
|
-
}
|
|
27
|
-
|
|
28
|
-
export async function chatText(llm, system, user, { timeoutMs = 120000 } = {}) {
|
|
18
|
+
export async function chatText(llm, system, user, { timeoutMs = 120000, route } = {}) {
|
|
19
|
+
// provider/model 是宿主 GenerateOptions 的必填项,缺失时 LlmRuntime 抛 NO_ADAPTER。
|
|
20
|
+
// 这里显式失败(由调用方决定回退并记 degraded),不让异常悄悄消失。
|
|
21
|
+
if (!route?.provider || !route?.model) {
|
|
22
|
+
throw new Error('auxiliary LLM call requires an explicit provider/model route')
|
|
23
|
+
}
|
|
29
24
|
const assembler = new BlockAssembler()
|
|
30
25
|
const controller = new AbortController()
|
|
31
26
|
const timer = setTimeout(() => controller.abort(), timeoutMs)
|
|
32
27
|
try {
|
|
33
28
|
for await (const chunk of llm.stream({
|
|
29
|
+
provider: route.provider,
|
|
30
|
+
model: route.model,
|
|
34
31
|
messages: [systemMessage(system), createUserMessage({ content: [{ type: 'text', text: user }] })],
|
|
35
32
|
signal: controller.signal,
|
|
36
33
|
})) {
|
|
@@ -72,8 +69,13 @@ function parseJson(text, validate) {
|
|
|
72
69
|
return null
|
|
73
70
|
}
|
|
74
71
|
|
|
75
|
-
|
|
72
|
+
/** 召回期可选的查询扩展(默认关;开启时才需要 provider/model 路由)。 */
|
|
73
|
+
export async function expandQuery(llm, query, count = 6, { route } = {}) {
|
|
76
74
|
if (!llm) return [query]
|
|
75
|
+
if (!route) {
|
|
76
|
+
noteDegraded('llm.expand.no-route', 'no provider/model route for query expansion; search uses the raw query')
|
|
77
|
+
return [query]
|
|
78
|
+
}
|
|
77
79
|
const system =
|
|
78
80
|
'You are a search-query expander for a codebase/document memory search engine. ' +
|
|
79
81
|
'Given a user query, return a STRICT JSON array of alternative search queries that ' +
|
|
@@ -81,57 +83,14 @@ export async function expandQuery(llm, query, count = 6) {
|
|
|
81
83
|
'code identifier guesses, and narrower/longer phrasings. Include the original query first. ' +
|
|
82
84
|
'Output only the JSON array of strings, no fences, no commentary.'
|
|
83
85
|
try {
|
|
84
|
-
const raw = await chatText(llm, system, `Query: "${query}"\n\nReturn the JSON array
|
|
86
|
+
const raw = await chatText(llm, system, `Query: "${query}"\n\nReturn the JSON array.`, { route })
|
|
85
87
|
const parsed = parseJsonArray(raw)
|
|
86
88
|
if (Array.isArray(parsed) && parsed.length) {
|
|
87
89
|
const variants = parsed.map(String).filter((s) => s.trim()).slice(0, count)
|
|
88
90
|
if (variants.length) return variants
|
|
89
91
|
}
|
|
90
|
-
} catch {
|
|
91
|
-
|
|
92
|
+
} catch (err) {
|
|
93
|
+
noteDegraded('llm.expand.failed', `query expansion LLM call failed: ${err?.message || err}`)
|
|
92
94
|
}
|
|
93
95
|
return [query]
|
|
94
96
|
}
|
|
95
|
-
|
|
96
|
-
export async function extractDocEntry(llm, chunk, sourcePath) {
|
|
97
|
-
const system =
|
|
98
|
-
'You are a project-documentation indexer. Given a chunk of a project document, ' +
|
|
99
|
-
'return a STRICT JSON object with exactly four fields: ' +
|
|
100
|
-
'"title" (short section title, string), ' +
|
|
101
|
-
'"summary" (3-5 sentence dense summary of what this section covers, ' +
|
|
102
|
-
'mentioning concrete names, decisions, constraints, and key technical details), ' +
|
|
103
|
-
'"blindSpots" (string describing what this summary does NOT cover, ' +
|
|
104
|
-
'e.g. "未覆盖:部署细节、性能基准、v0.2 前 API", empty string if none), ' +
|
|
105
|
-
'"keywords" (array of 5-10 searchable strings: ' +
|
|
106
|
-
'cover the document\'s own language AND English equivalents, ' +
|
|
107
|
-
'so a query in either language can match). ' +
|
|
108
|
-
'Do not include markdown fences, do not add commentary, output only the JSON object.'
|
|
109
|
-
|
|
110
|
-
const user =
|
|
111
|
-
`Document: ${sourcePath}\nSection: ${chunk.title || '(untitled)'}\n\n` +
|
|
112
|
-
`Content:\n${chunk.text.slice(0, 6000)}\n\nReturn the JSON object.`
|
|
113
|
-
|
|
114
|
-
const fallback = () => ({
|
|
115
|
-
title: chunk.title || sourcePath,
|
|
116
|
-
summary: summarizeText(chunk.text),
|
|
117
|
-
blindSpots: '',
|
|
118
|
-
keywords: tokenize(chunk.title).slice(0, 5),
|
|
119
|
-
})
|
|
120
|
-
|
|
121
|
-
if (!llm) return fallback()
|
|
122
|
-
|
|
123
|
-
try {
|
|
124
|
-
const raw = await chatText(llm, system, user)
|
|
125
|
-
const parsed = parseStructuredJson(raw)
|
|
126
|
-
if (!parsed || typeof parsed.summary !== 'string' || !parsed.summary.trim()) return fallback()
|
|
127
|
-
const kw = Array.isArray(parsed.keywords) ? parsed.keywords.map(String).filter((k) => k).slice(0, 8) : []
|
|
128
|
-
return {
|
|
129
|
-
title: typeof parsed.title === 'string' && parsed.title.trim() ? parsed.title.trim() : chunk.title || sourcePath,
|
|
130
|
-
summary: summarizeText(parsed.summary.trim()),
|
|
131
|
-
blindSpots: typeof parsed.blindSpots === 'string' ? parsed.blindSpots.trim() : '',
|
|
132
|
-
keywords: kw.length ? kw : tokenize(chunk.title).slice(0, 5),
|
|
133
|
-
}
|
|
134
|
-
} catch {
|
|
135
|
-
return fallback()
|
|
136
|
-
}
|
|
137
|
-
}
|
|
@@ -10,6 +10,7 @@ import { ProjectMemoryStore } from './store.js'
|
|
|
10
10
|
import { memoryRootFor } from './util/fs.js'
|
|
11
11
|
import { cfgInsight, saveInsight, GlobalStore, defaultGlobalFile } from './insight-store.js'
|
|
12
12
|
import { chatText, parseStructuredJson } from './llm.js'
|
|
13
|
+
import { noteDegraded } from './llm-route.js'
|
|
13
14
|
|
|
14
15
|
export const REFLECT_SYSTEM =
|
|
15
16
|
'You distill task work into durable lessons and decisions for project memory. ' +
|
|
@@ -57,7 +58,7 @@ function normalizeList(arr, max) {
|
|
|
57
58
|
* 对某任务做一次反思(可归档任务照做——归档正是收割时机)。
|
|
58
59
|
* 全程 try/catch:调用方 fire-and-forget 即可。
|
|
59
60
|
*/
|
|
60
|
-
export async function reflectTaskAfter({ config, llm, root, taskId, reason = 'transition' }) {
|
|
61
|
+
export async function reflectTaskAfter({ config, llm, root, taskId, reason = 'transition', route = null }) {
|
|
61
62
|
const rc = (config && config.reflection) || {}
|
|
62
63
|
const fallback = { ok: false, skipped: 'disabled' }
|
|
63
64
|
if (!rc.enabled) return fallback
|
|
@@ -67,6 +68,11 @@ export async function reflectTaskAfter({ config, llm, root, taskId, reason = 'tr
|
|
|
67
68
|
const task = store.getTask(taskId)
|
|
68
69
|
const gate = isReflectDue(config, task)
|
|
69
70
|
if (!gate.due) return { ok: false, skipped: gate.reason }
|
|
71
|
+
// 路由检查放在门控之后:未到期/内容未变时不应留下 no-route 降级噪声
|
|
72
|
+
if (!route) {
|
|
73
|
+
noteDegraded('llm.reflect.no-route', 'no provider/model route for reflection; skipping this reflection')
|
|
74
|
+
return { ok: false, skipped: 'no-route' }
|
|
75
|
+
}
|
|
70
76
|
|
|
71
77
|
const snap = taskReflectionSnapshot(task)
|
|
72
78
|
const user =
|
|
@@ -75,7 +81,7 @@ export async function reflectTaskAfter({ config, llm, root, taskId, reason = 'tr
|
|
|
75
81
|
(snap.files ? `Files touched:\n${snap.files.slice(0, 1000)}\n` : '') +
|
|
76
82
|
`Trigger: ${reason}\n\nReturn the JSON object.`
|
|
77
83
|
|
|
78
|
-
const raw = await chatText(llm, REFLECT_SYSTEM, user, { timeoutMs: 90000 })
|
|
84
|
+
const raw = await chatText(llm, REFLECT_SYSTEM, user, { timeoutMs: 90000, route })
|
|
79
85
|
const parsed = parseStructuredJson(raw)
|
|
80
86
|
if (!parsed) return { ok: false, skipped: 'unparsable' }
|
|
81
87
|
|
|
@@ -136,10 +142,10 @@ export async function reflectTaskAfter({ config, llm, root, taskId, reason = 'tr
|
|
|
136
142
|
}
|
|
137
143
|
|
|
138
144
|
/** 工具/命令层 fire-and-forget:默认关时立刻廉价返回,绝不让错误上抛。 */
|
|
139
|
-
export function fireReflect(config, host, root, taskId, reason) {
|
|
145
|
+
export function fireReflect(config, host, root, taskId, reason, route = null) {
|
|
140
146
|
const rc = (config && config.reflection) || {}
|
|
141
147
|
if (!rc.enabled) return Promise.resolve({ ok: false, skipped: 'disabled' })
|
|
142
|
-
return reflectTaskAfter({ config, llm: host?.llm, root, taskId, reason }).catch((err) => {
|
|
148
|
+
return reflectTaskAfter({ config, llm: host?.llm, root, taskId, reason, route }).catch((err) => {
|
|
143
149
|
console.error(`[dsh-project-memory] fireReflect error: ${err?.message || err}`)
|
|
144
150
|
return { ok: false, skipped: 'error' }
|
|
145
151
|
})
|
package/src/setup/taskbridge.js
CHANGED
|
@@ -146,8 +146,14 @@ export function touchTaskFile(task, rel, kind, now) {
|
|
|
146
146
|
}
|
|
147
147
|
}
|
|
148
148
|
const m = task.fileMeta[rel] || {}
|
|
149
|
-
m.n = (m.n || 0) + 1
|
|
150
|
-
|
|
149
|
+
m.n = (m.n || 0) + 1 // 兼容旧字段:读+写总数(保留,避免老任务语义漂移)
|
|
150
|
+
// 读/写分计数(先只采集,不消费 —— 排序与注入仍只用 lastWriteAt/lastReadAt)
|
|
151
|
+
if (kind === 'write') {
|
|
152
|
+
m.writes = (m.writes || 0) + 1
|
|
153
|
+
m.lastWriteAt = now
|
|
154
|
+
} else if (kind === 'read') {
|
|
155
|
+
m.reads = (m.reads || 0) + 1
|
|
156
|
+
}
|
|
151
157
|
m.lastReadAt = now
|
|
152
158
|
task.fileMeta[rel] = m
|
|
153
159
|
task.files = hotSortFiles(task.files, task.fileMeta)
|
package/src/store.js
CHANGED
|
@@ -158,7 +158,7 @@ export class ProjectMemoryStore {
|
|
|
158
158
|
}
|
|
159
159
|
|
|
160
160
|
// ---- v0.5 insights:project 级 insight 文档(任务级在 task.insights[]) ----
|
|
161
|
-
//
|
|
161
|
+
// 迁移语义(旧版 store 的兼容策略):
|
|
162
162
|
// v0.4 experience.json 仍由 remember/forget/query_memory 服务,不删除;
|
|
163
163
|
// 首次加载把旧笔记**复制导入** insights.json(kind: experience, source: migrate),
|
|
164
164
|
// migratedAt 落盘保证跨进程/崩溃幂等。销毁式收敛放到 recall 统一 PR。
|
package/src/tools/index-doc.js
CHANGED
|
@@ -11,8 +11,8 @@ export function indexDocTool(ctx, config) {
|
|
|
11
11
|
name: 'index_doc',
|
|
12
12
|
description:
|
|
13
13
|
'Index a project document (PDF, Markdown, txt) into persistent project memory: split into sections, ' +
|
|
14
|
-
'
|
|
15
|
-
'Re-indexing the same unchanged file is a no-op (content-hash skip).',
|
|
14
|
+
'store a short cited summary plus a full-chunk literal term index (no LLM at index time), for later ' +
|
|
15
|
+
'query_memory recall. Re-indexing the same unchanged file is a no-op (content-hash skip).',
|
|
16
16
|
parameters: {
|
|
17
17
|
file_path: {
|
|
18
18
|
type: 'string',
|
|
@@ -42,7 +42,7 @@ export function indexDocTool(ctx, config) {
|
|
|
42
42
|
return `Skipped (unchanged): ${rel}\nAlready indexed with ${(store.entries[rel] || []).length} entry/entries.`
|
|
43
43
|
}
|
|
44
44
|
|
|
45
|
-
const entries = await buildDocEntries(
|
|
45
|
+
const entries = await buildDocEntries(rel, filePath, {
|
|
46
46
|
chunkChars: config.chunkChars,
|
|
47
47
|
maxChunks: config.maxChunksPerFile,
|
|
48
48
|
maxFileSizeMb: config.maxFileSizeMb,
|
package/src/tools/index-repo.js
CHANGED
|
@@ -3,6 +3,7 @@ import path from 'node:path'
|
|
|
3
3
|
import { statSync } from 'node:fs'
|
|
4
4
|
import { isSupportedCode, isSupportedDoc, looksLikeDump, memoryRootFor, readFileForIndex, relativePath, storeKey, walkDir } from '../util/fs.js'
|
|
5
5
|
import { buildDocEntries } from '../doc-pipeline.js'
|
|
6
|
+
import { docEntriesNeedBackfill } from '../doc-index.js'
|
|
6
7
|
import { scanSymbols } from '../symbols.js'
|
|
7
8
|
import { linkEntries } from '../link.js'
|
|
8
9
|
import { ProjectMemoryStore } from '../store.js'
|
|
@@ -40,7 +41,9 @@ export async function indexRepository(ctx, config, root, { reindex = false } = {
|
|
|
40
41
|
|
|
41
42
|
// 单次读盘:同一 buffer 供哈希与正文使用(不再 sha256OfFile + readFileSync 读两遍)
|
|
42
43
|
const { hash, buffer } = readFileForIndex(filePath)
|
|
43
|
-
|
|
44
|
+
// 旧 store 的 doc 条目缺 terms → 即使哈希未变也重抽一次(一次性回填)
|
|
45
|
+
const needsBackfill = isSupportedDoc(ext) && docEntriesNeedBackfill(store.entries[rel])
|
|
46
|
+
if (!reindex && existing && existing.sha256 === hash && !needsBackfill) {
|
|
44
47
|
skipped++
|
|
45
48
|
continue
|
|
46
49
|
}
|
|
@@ -57,7 +60,7 @@ export async function indexRepository(ctx, config, root, { reindex = false } = {
|
|
|
57
60
|
skipped++
|
|
58
61
|
continue
|
|
59
62
|
}
|
|
60
|
-
entries = await buildDocEntries(
|
|
63
|
+
entries = await buildDocEntries(rel, filePath, {
|
|
61
64
|
chunkChars: config.chunkChars,
|
|
62
65
|
maxChunks: config.maxChunksPerFile,
|
|
63
66
|
maxFileSizeMb: config.maxFileSizeMb,
|
|
@@ -123,7 +126,8 @@ export function indexRepoTool(ctx, config) {
|
|
|
123
126
|
return defineTool({
|
|
124
127
|
name: 'index_repo',
|
|
125
128
|
description:
|
|
126
|
-
'Index a whole project into persistent memory. Documents (PDF/Markdown/txt)
|
|
129
|
+
'Index a whole project into persistent memory. Documents (PDF/Markdown/txt) get a short cited summary plus ' +
|
|
130
|
+
'a full-chunk literal term index (no LLM at index time); ' +
|
|
127
131
|
'code files get a zero-token symbol table (function/class names with line numbers). Incremental: only changed ' +
|
|
128
132
|
'files are re-extracted (content-hash), deleted files are removed from memory. Call once per project, then query_memory.',
|
|
129
133
|
parameters: {
|
|
@@ -3,6 +3,7 @@ import path from 'node:path'
|
|
|
3
3
|
import { memoryRootFor, resolveIndexRoot } from '../util/fs.js'
|
|
4
4
|
import { ProjectMemoryStore, storeOverview } from '../store.js'
|
|
5
5
|
import { expandQuery } from '../llm.js'
|
|
6
|
+
import { resolveRoute } from '../llm-route.js'
|
|
6
7
|
import { rankEntriesMergedScored, rankExperienceScored, rankEntriesStreaming } from '../util/search.js'
|
|
7
8
|
import { truncate } from '../util/text.js'
|
|
8
9
|
|
|
@@ -49,7 +50,7 @@ export function queryMemoryTool(ctx, config) {
|
|
|
49
50
|
const limit = Math.max(1, Math.min(Number(args.limit) || 8, 20))
|
|
50
51
|
|
|
51
52
|
const queries = config.llmQueryExpansion
|
|
52
|
-
? await expandQuery(ctx.llm, args.query, config.expansionCount)
|
|
53
|
+
? await expandQuery(ctx.llm, args.query, config.expansionCount, { route: resolveRoute(exec, config) })
|
|
53
54
|
: [args.query]
|
|
54
55
|
const symbolById = new Map()
|
|
55
56
|
for (const e of store.allEntries()) {
|
|
@@ -77,7 +78,9 @@ export function queryMemoryTool(ctx, config) {
|
|
|
77
78
|
summaryLine += `\n- ⚠️ 摘要未覆盖:${e.blindSpots.replace(/^\s*\/\/\s*未覆盖[::]\s*/, '')}。建议读原文 ${absSource}`
|
|
78
79
|
}
|
|
79
80
|
}
|
|
80
|
-
|
|
81
|
+
// 条目状态占位(可证伪状态机落地前恒为 exact)。
|
|
82
|
+
// 先立字段,后续状态机到位时只改值、不改输出契约。
|
|
83
|
+
lines.push(`### ${e.title} (score: ${rel})\n- source: ${absSource}\n- status: ${e.status || 'exact'}\n${summaryLine}`)
|
|
81
84
|
if (e.type === 'doc' && Array.isArray(e.linkedSymbols) && e.linkedSymbols.length) {
|
|
82
85
|
const refs = e.linkedSymbols.slice(0, 5).map((id) => {
|
|
83
86
|
const s = symbolById.get(id)
|
package/src/tools/task-tools.js
CHANGED
|
@@ -4,6 +4,7 @@ import { ProjectMemoryStore } from '../store.js'
|
|
|
4
4
|
import { truncate } from '../util/text.js'
|
|
5
5
|
import { genTaskId, hash8, adoptStepsToSession, shouldAdoptToHost } from '../setup/taskbridge.js'
|
|
6
6
|
import { fireReflect } from '../reflection-pipeline.js'
|
|
7
|
+
import { resolveRoute } from '../llm-route.js'
|
|
7
8
|
import { createHash } from 'node:crypto'
|
|
8
9
|
|
|
9
10
|
function sessionIdOf(exec) {
|
|
@@ -137,7 +138,7 @@ export function selectTaskTool(config, host) {
|
|
|
137
138
|
|
|
138
139
|
// 反思钩子:切走旧任务时异步收割(默认关,零成本;失败只记日志)
|
|
139
140
|
if (prevBound && prevBound !== task.id) {
|
|
140
|
-
fireReflect(config, host, root, prevBound, 'switch-away')
|
|
141
|
+
fireReflect(config, host, root, prevBound, 'switch-away', resolveRoute(exec, config))
|
|
141
142
|
}
|
|
142
143
|
|
|
143
144
|
const card = {
|
|
@@ -191,7 +192,7 @@ export function archiveTaskTool(config, host) {
|
|
|
191
192
|
task.updatedAt = new Date().toISOString()
|
|
192
193
|
store.save()
|
|
193
194
|
// 反思钩子:归档即收割最终教训(默认关)
|
|
194
|
-
fireReflect(config, host, root, args.taskId, 'archive')
|
|
195
|
+
fireReflect(config, host, root, args.taskId, 'archive', resolveRoute(exec, config))
|
|
195
196
|
return truncate(JSON.stringify({ success: true, archived: true, hint: '已归档,select_task 可恢复' }), config.maxOutputChars)
|
|
196
197
|
},
|
|
197
198
|
})
|
package/src/util/search.js
CHANGED
|
@@ -131,6 +131,8 @@ export function weightedFieldText(entry) {
|
|
|
131
131
|
for (let i = 0; i < 5; i++) parts.push(entry.title || '')
|
|
132
132
|
parts.push((entry.keywords || []).join(' '))
|
|
133
133
|
parts.push(entry.summary || '')
|
|
134
|
+
// 结构词项:覆盖整个 chunk 的字面词项,把「只有前 300 字符可检索」补回来(索引期零 LLM)
|
|
135
|
+
parts.push(entry.terms || '')
|
|
134
136
|
parts.push(entry.sourcePath || '')
|
|
135
137
|
return parts.join(' ')
|
|
136
138
|
}
|
package/src/util/text.js
CHANGED
|
@@ -2,4 +2,17 @@ export function truncate(text, maxChars) {
|
|
|
2
2
|
if (typeof text !== 'string') return String(text)
|
|
3
3
|
if (text.length <= maxChars) return text
|
|
4
4
|
return text.slice(0, maxChars) + `\n\n...[truncated at ${maxChars} chars]`
|
|
5
|
+
}
|
|
6
|
+
|
|
7
|
+
export const MAX_SUMMARY = 300
|
|
8
|
+
|
|
9
|
+
/** 注入用短摘要:压平空白,按句子边界截断到 max 字符。纯函数,无 LLM。 */
|
|
10
|
+
export function summarizeText(text, max = MAX_SUMMARY) {
|
|
11
|
+
const flat = String(text || '').replace(/\s+/g, ' ').trim()
|
|
12
|
+
if (!flat) return ''
|
|
13
|
+
if (flat.length <= max) return flat
|
|
14
|
+
const clip = max - 1
|
|
15
|
+
const clipped = flat.slice(0, clip)
|
|
16
|
+
const lastBreak = Math.max(clipped.lastIndexOf('。'), clipped.lastIndexOf('.'), clipped.lastIndexOf(';'))
|
|
17
|
+
return lastBreak > clip * 0.4 ? clipped.slice(0, lastBreak + 1) : clipped + '…'
|
|
5
18
|
}
|
package/src/watch.js
CHANGED
|
@@ -2,6 +2,7 @@ import { existsSync, statSync } from 'node:fs'
|
|
|
2
2
|
import path from 'node:path'
|
|
3
3
|
import { isSupportedCode, isSupportedDoc, memoryRootFor, readFileForIndex, relativePath, storeKey, walkDir } from './util/fs.js'
|
|
4
4
|
import { buildDocEntries } from './doc-pipeline.js'
|
|
5
|
+
import { docEntriesNeedBackfill } from './doc-index.js'
|
|
5
6
|
import { scanSymbols } from './symbols.js'
|
|
6
7
|
import { linkEntries } from './link.js'
|
|
7
8
|
import { ProjectMemoryStore } from './store.js'
|
|
@@ -105,14 +106,16 @@ export class WatchManager {
|
|
|
105
106
|
// 单次读盘:同一 buffer 供哈希与正文使用
|
|
106
107
|
const { hash, buffer } = readFileForIndex(filePath)
|
|
107
108
|
const existing = state.store.fileRecord(rel)
|
|
108
|
-
|
|
109
|
+
// 旧 store 的 doc 条目缺 terms → 一次性回填(即使哈希未变)
|
|
110
|
+
const needsBackfill = isSupportedDoc(ext) && docEntriesNeedBackfill(state.store.entries[rel])
|
|
111
|
+
if (existing && existing.sha256 === hash && !needsBackfill) continue
|
|
109
112
|
|
|
110
113
|
try {
|
|
111
114
|
let entries
|
|
112
115
|
if (isSupportedCode(ext)) {
|
|
113
116
|
entries = scanSymbols(rel, filePath, buffer.toString('utf8'))
|
|
114
117
|
} else {
|
|
115
|
-
entries = await buildDocEntries(
|
|
118
|
+
entries = await buildDocEntries(rel, filePath, {
|
|
116
119
|
chunkChars: this.config.chunkChars,
|
|
117
120
|
maxChunks: this.config.maxChunksPerFile,
|
|
118
121
|
maxFileSizeMb: this.config.maxFileSizeMb,
|