@yolk_vat-y/dsh-project-memory 0.5.10 → 0.5.12
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +257 -1
- package/LICENSE +1 -1
- package/README.md +40 -31
- package/README.zh-CN.md +40 -31
- package/cordis.patch.yml +1 -1
- package/package.json +3 -2
- package/scripts/bench-load.mjs +47 -0
- package/scripts/bench-synthetic.mjs +2 -4
- package/scripts/bench.mjs +35 -4
- package/src/audit.js +29 -3
- package/src/auto-inject.js +5 -3
- package/src/enhancer.js +181 -126
- package/src/index-pipeline.js +5 -4
- package/src/index.js +1 -1
- package/src/link.js +104 -38
- package/src/store.js +254 -25
- package/src/symbols.js +3 -1
- package/src/tools/index-doc.js +0 -2
- package/src/tools/index-repo.js +15 -6
- package/src/tools/query-memory.js +33 -13
- package/src/util/fs.js +65 -1
- package/src/util/search.js +4 -9
- package/src/util/text.js +16 -0
- package/src/watch.js +80 -18
package/CHANGELOG.md
CHANGED
|
@@ -1,4 +1,260 @@
|
|
|
1
|
-
|
|
1
|
+
## 0.5.12 (2026-09-26)
|
|
2
|
+
|
|
3
|
+
本版五个修复 + 一个新增,主题是记忆的驻留、索引归属,以及两处让 TS 增强层失效 / 失真的缺陷。`npm test` **476 → 539 项 / 26 → 28 个文件**。
|
|
4
|
+
|
|
5
|
+
### 修复:TS 增强队列的 TDZ 会打死宿主(`dsh: fatal load failure`)
|
|
6
|
+
|
|
7
|
+
```
|
|
8
|
+
dsh: fatal load failure: ReferenceError: Cannot access 'p' before initialization
|
|
9
|
+
at enqueueEnhance (enhancer.js) → onFileChanged → WatchManager.pollRoot
|
|
10
|
+
```
|
|
11
|
+
|
|
12
|
+
`enqueueEnhance()` 先启动任务、后登记队列条目,并在 `finally` 里用
|
|
13
|
+
`enhanceQueue.findIndex(q => q.promise === p)` 反查自己。`deepParseWithTS()` 是同步的,文件
|
|
14
|
+
产不出增强符号时任务体走不到 `await`,`finally` 就在同步阶段执行——此时 `const p` 尚未完成
|
|
15
|
+
初始化。队列为空时 `findIndex` 回调不被调用,问题不显形;队列非空时读到 TDZ 的 `p`,抛
|
|
16
|
+
ReferenceError。watch 轮询对每个改动的 `.js` 都在同一个同步循环里调 `onFileChanged`,且丢弃了
|
|
17
|
+
返回的 promise,未处理的 rejection 被宿主当成致命错误。
|
|
18
|
+
|
|
19
|
+
- 队列条目改为先入队,`finally` 按对象身份摘除自己。
|
|
20
|
+
- 三个 fire-and-forget 入口各挂 catch。
|
|
21
|
+
- 顺带修掉同步完成的任务把一条已 resolve 的僵尸条目永久留在队列里的问题:同一路径同内容的
|
|
22
|
+
后续调用会命中它,导致再也不增强。
|
|
23
|
+
- 测试:新增 `test/enhancer.test.mjs`(6 项),`src/enhancer.js` 此前无测试覆盖。
|
|
24
|
+
|
|
25
|
+
### 修复:TS 增强层对 `.js` 文件从未生效(不分平台)
|
|
26
|
+
|
|
27
|
+
`deepParseWithTS()` 调 `ts.createProgram([filePath], {}, host)`,`options` 是空的。
|
|
28
|
+
TypeScript 在**调用 host 之前**就有一道扩展名闸门:`getSourceFileFromReferenceWorker` 先比对
|
|
29
|
+
`getSupportedExtensions(options)`,`.js/.jsx/.mjs/.cjs` 不在其中就直接返回——host 根本不会被问。
|
|
30
|
+
|
|
31
|
+
而插件的 `isTypeScriptFile()` 明确把这四种扩展名列为可增强对象,watch / 懒索引 / `index_repo`
|
|
32
|
+
三条入口也照常喂进去。**结果是 TS 增强层(README 承诺的"推导返回类型、解析泛型、抽取接口")
|
|
33
|
+
对绝大多数真实文件一直没有产出,且没有任何日志。**
|
|
34
|
+
|
|
35
|
+
- `createProgram` 改用 `{ allowJs: true, noLib: true }`。`.js` / `.jsx` / `.mjs` / `.cjs` 从 0 个符号
|
|
36
|
+
恢复到正常产出(`isTypeScriptFile()` 认的 8 种扩展名逐一断言),`.ts` 行为不变。
|
|
37
|
+
- **顺带修掉 L2 对无注解 JS 的文本损伤。** `applyEnhancedSymbols()` 是**就地重写** L1 条目的
|
|
38
|
+
`text`,而旧实现把参数一律渲染成 `${name}: ${type}`:本仓库 51 个 `.js/.mjs` 实测 334 条条目被
|
|
39
|
+
重写,86% 的参数成了 `: any`,52 个带默认值的条目里 36 个把默认值丢了——
|
|
40
|
+
|
|
41
|
+
```
|
|
42
|
+
chunkText(text, chunkChars = 3000, maxChunks = 40) — src/chunker.js:1 ← L1
|
|
43
|
+
chunkText(text: any, chunkChars: any, maxChunks: any): any -- src/chunker.js:1 ← 旧 L2
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
现在参数用**源码文本**(`= 默认值` / `?` / `...rest` / 解构模式都保住),推导出的返回类型只在
|
|
47
|
+
有信息量时才追加(`any` / `unknown` / `{}` 不写)。改后同一份语料:默认值 0 丢失,被重写的条目
|
|
48
|
+
里 51% 的文本与 L1 逐字相同(升级不再是损伤),其余 49% 多出一个真实返回类型(`: boolean` /
|
|
49
|
+
`: void` / `: () => void` / 对象形状)。
|
|
50
|
+
- `.cjs` 的惯用写法 `module.exports.foo = function () {}` 此前产不出符号:它的父节点是
|
|
51
|
+
`BinaryExpression`,不在函数表达式的名字白名单里——而 `isTypeScriptFile()` 明确把 `.cjs` 列为
|
|
52
|
+
可增强对象。现在按 `module.exports.foo` / `exports.foo` / 字符串下标取名字。
|
|
53
|
+
- 增强条目的一行声明改用与 L1 `buildSymbol()` 相同的 ` — ` 分隔(旧为 ` -- `):L2 重写的正是
|
|
54
|
+
L1 的文本,两种分隔符会让同一个文件里的条目混排两种格式。
|
|
55
|
+
- **测试同步加固**:`test/enhancer.test.mjs` 原先的断言是 `existsSync(store.dir)`——store 目录
|
|
56
|
+
只要有任何一次 commit 就会存在,前面几个用例已经把它建好,这条因此永远为真;它验的是
|
|
57
|
+
"跑过了"而不是"处理了"。换成 `store.entries[rel].some(e => e.enhanced)`,并补 `.js` 与
|
|
58
|
+
`.mjs` 的直接断言。固定 `setTimeout` 也改成轮询到条件成立,慢 CI 上不再靠运气。新增扩展名
|
|
59
|
+
覆盖面、文本保真、CJS 断言与静默契约四组(共 13 项;前三组的 11 项在旧代码上为红,静默契约
|
|
60
|
+
那 2 项锁的是"不加日志"这个决定,新旧都绿——它防的是将来有人把告警加回来)。
|
|
61
|
+
|
|
62
|
+
### 修复:TS 增强在 Windows 上因路径写法不同而整层失效
|
|
63
|
+
|
|
64
|
+
compiler host 用 `fileName === filePath` 严格相等来认自己的文件。TypeScript 内部一律走
|
|
65
|
+
`normalizePath`(反斜杠转正斜杠、消解 `.`/`..`、绝对化),而 `path.join` 给出的是平台原生形态。
|
|
66
|
+
Windows 上两边字符串不同 → `getSourceFile` 返回 undefined → `deepParseWithTS` 返回 `[]`。
|
|
67
|
+
POSIX 上路径本就是正斜杠,因此从不现形。
|
|
68
|
+
|
|
69
|
+
2026-09-26 的 Windows CI 就是这样暴露的(唯一失败断言是 enhancer 的「增强结果落进 store」)。
|
|
70
|
+
|
|
71
|
+
- 路径比较改用 `ts.normalizePath`(带 `typeof` 兜底:它在运行时导出但 `typescript.d.ts` 里
|
|
72
|
+
没有声明,而 peer 范围覆盖 TS 5/6/7)。
|
|
73
|
+
- `useCaseSensitiveFileNames` 不再硬编码 `true`,改为跟随 `ts.sys`——TS 在 macOS 上是探测真实
|
|
74
|
+
文件系统而非写死 `false`,跟随它才是正确语义。
|
|
75
|
+
- 预建唯一的 `SourceFile` 并按**对象身份**确认 TS 收下了它,不再用 `program.getSourceFile(filePath)`
|
|
76
|
+
回查(那一步内部会把相对路径拼上 cwd,同样是脆弱点)。
|
|
77
|
+
- 回归断言改成**字符串形态**:反斜杠(Windows 形态)、裸 `./` 段、裸 `..` 段。原先那条用的是
|
|
78
|
+
`path.join(dir, '.', 'shape.ts')`,而 `path.join` 自己就会消解 `.`——它与 `path.join(dir, 'shape.ts')`
|
|
79
|
+
是同一个字符串,断言在旧代码上也是绿的,等于没测。TS 的 `normalizePath` 在**任何平台**都把 `\`
|
|
80
|
+
归一成 `/`,所以这三条在 Linux 上就能判定,不必等 Windows CI。
|
|
81
|
+
- 两条"本该有结果却没有"的路径**保持静默**(host 没认出 root path、`readFileSync` 失败),
|
|
82
|
+
但不再是无从分辨的黑盒:调用方契约(入队 resolve、返回 `[]`)由断言正面锁住,理由写在代码里。
|
|
83
|
+
不做日志是代价与收益的权衡——watch 每 15s 一轮、每个不受支持的 root 都会撞上它,按文件去重
|
|
84
|
+
也只是把"每轮一条"换成"每个文件一条",而"这轮没增强"不会产出任何错误结果。2026-09-26 的
|
|
85
|
+
教训是"静默"藏住了整层失效,但解法是补测试(见上一组断言),不是往终端打字。
|
|
86
|
+
|
|
87
|
+
### 已知限制(本版未做):L2 不加载 TypeScript 的默认 lib
|
|
88
|
+
|
|
89
|
+
compiler host 里的 `getDefaultLibLocation: () => ts.getDefaultLibFilePath({})` 返回的是**文件**
|
|
90
|
+
路径而不是目录,于是默认 lib 从来没有被加载过——显式标注的类型正常,依赖全局类型(`Promise` /
|
|
91
|
+
`Array` / DOM)的一律塌成 `any` / `unknown`:
|
|
92
|
+
|
|
93
|
+
```
|
|
94
|
+
export function f() { return [1, 2, 3].map(n => n * 2) } → (): any (载入 lib 后是 number[])
|
|
95
|
+
export const g = async () => 1 → (): unknown (载入 lib 后是 Promise<number>)
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
本版把它换成**显式**的 `noLib: true` 并在代码里写明理由:行为不变(实测两者产出的符号完全一致),
|
|
99
|
+
但"没生效"不再是一个没人写下来的意外。
|
|
100
|
+
|
|
101
|
+
修它不止一行:目录给对之后 host 还必须按需提供 `lib.d.ts` 及其引用的 6 个文件(`lib.es5` /
|
|
102
|
+
`lib.decorators` / `lib.decorators.legacy` / `lib.dom` / `lib.webworker.importscripts` /
|
|
103
|
+
`lib.scripthost`,共 ~2.1MB 文本)。实测代价:每个 program 重解析 **p50=102ms/文件**;把解析好的
|
|
104
|
+
`SourceFile` 缓存跨 program 复用(或走 `ts.DocumentRegistry`)则是一次性 ~100ms + 每文件 ~1ms。
|
|
105
|
+
这是"性能换质量"的产品选择,留到下一个版本单独做。
|
|
106
|
+
|
|
107
|
+
### 修复:storeCache 的驻留估算量错了对象
|
|
108
|
+
|
|
109
|
+
`estimateResidentBytes()` 用 `Math.max(_residentChars, entries * 2500)` 估算一个 store 的内存占用,
|
|
110
|
+
两个口径都不成立:
|
|
111
|
+
|
|
112
|
+
1. `_residentChars` 是 `loadJson` 读进来的 JSON 文本长度,其中包含 `load()` 随即被
|
|
113
|
+
`stripPersistedDerived` 删掉的字节——老分片仍带 `searchText` / `linkedSymbols`。
|
|
114
|
+
实测一个 store 的读盘文本是真实驻留的 **2.09 倍**。
|
|
115
|
+
2. 估算在 `load()` 内就被调用,而 `searchText` 要到 `allEntries()` 才物化,同一个 store
|
|
116
|
+
在物化前后估出两个不同的数。
|
|
117
|
+
|
|
118
|
+
结果是高估:本可共存的 store 被过早逐出,下一次 `load()` 重读全部分片。在 Windows 上
|
|
119
|
+
(Defender 实时扫描放大每次文件读)尤其明显。
|
|
120
|
+
|
|
121
|
+
`deepseek-harness`(11698 文件 / 70119 条目)实测:
|
|
122
|
+
|
|
123
|
+
| 口径 | load() 后 | 物化 `searchText` 后 |
|
|
124
|
+
|---|---|---|
|
|
125
|
+
| 旧估算 | **167.2MB**(真实 63.5MB) | 167.2MB(真实 97.8MB) |
|
|
126
|
+
| 新估算 | 75.0MB(**+18.2%**) | 104.8MB(**+7.1%**) |
|
|
127
|
+
|
|
128
|
+
- 改为 `_estimateResidentBytes()` 统计内存中的对象图(形状开销 + `Buffer.byteLength()`,
|
|
129
|
+
CJK 在 UTF-8 里 3 字节/字符,用 `String.length` 会再低估一半),按 `_entriesVersion` 与
|
|
130
|
+
新增的 `_materializeVersion` 记忆化——`evictStoreCache` 每次要对整个缓存逐 store 求和。
|
|
131
|
+
代价是版本变化后的一次全量遍历,冷加载 +8~11%。
|
|
132
|
+
- 删除 `_residentChars` 及 `loadJson` 的 `sizeSink` 参数。
|
|
133
|
+
- `keepKey` 豁免(刚加载的 store 即使超预算也保留)保留,但不再静默:单实例超预算时按 root
|
|
134
|
+
告警一次。
|
|
135
|
+
|
|
136
|
+
行为变化:逐出变少,内存上限更接近 256MB 的本意;只有真正超预算时才像以前一样逐出。
|
|
137
|
+
|
|
138
|
+
测试:新增 `test/store-cache.test.mjs`(22 项),`storeCache` 此前无测试覆盖。
|
|
139
|
+
|
|
140
|
+
### 修复:watch 根不再钉住 store,也不在启动时急切读盘
|
|
141
|
+
|
|
142
|
+
`WatchManager` 长期持有每个 watch 根的 store 实例(`this.roots.get(root).store`),而
|
|
143
|
+
`removeRoot()` 的唯一调用方是 `watch_repo --watch false`。两个后果:
|
|
144
|
+
|
|
145
|
+
1. 逐出对 watch 根不省内存——`storeCache` 丢了 key,watcher 还持有那份;
|
|
146
|
+
2. 一旦被逐出,下一次 `load()` 会从盘上再建一份。同一个 root 出现两份互不可见的内存状态,
|
|
147
|
+
各写各的脏分片,存在丢更新的窗口。
|
|
148
|
+
|
|
149
|
+
同一个字段还让 `restorePersisted()` 在启动时对整条 `watchlist` 急切读盘:watch 过多少个根,
|
|
150
|
+
启动就同步读多少个 store 的全部 shard。
|
|
151
|
+
|
|
152
|
+
- `addRoot()` 只登记 `{ snapshot, failures }`;新增 `storeFor(root)` 在 `pollRoot()` 需要时
|
|
153
|
+
按需取,整个 `pollRoot` 期间共用一个实例。
|
|
154
|
+
- 每个 root 全进程只有一个活实例,内存上界回到 `STORE_CACHE_MAX_BYTES`。
|
|
155
|
+
|
|
156
|
+
### 修复:父根不再重复索引嵌套的项目根
|
|
157
|
+
|
|
158
|
+
一个没有项目标记(无 `.git`、无 `package.json`)的容器目录经由「标记缺失 → 回退到会话 cwd」
|
|
159
|
+
成为根后,`walkDir()` 会把整棵树索引一遍,而树里的每个子项目各自都有自己的 store。实测某个
|
|
160
|
+
父根的 store 里 **99.9% 的条目是子根的重复副本**,且是过期快照(部分源文件已从盘上删除)——
|
|
161
|
+
这份 store 没有独有价值,却占用内存,并在父根下 `query_memory` 时返回陈旧内容。
|
|
162
|
+
|
|
163
|
+
- `walkDir()` 新增 `nestedStoreName`:子目录自带 store 时不再往下走,跳过的目录放进返回值的
|
|
164
|
+
`skipped` 字段。判据是子目录下存在 `<子目录>/<memoryDir>/format.json`——只有真正存过盘的
|
|
165
|
+
store 才算数,误建的空目录不作数。与 `findProjectRoot()` 的解析规则一致。
|
|
166
|
+
- `index_repo` 与 watch 轮询都启用;`index_repo` 在结果里列出被跳过的子根。
|
|
167
|
+
- 存量重复自动收敛:被跳过的子树不在 `seen` 里,既有的 `commitFileUpdates(..., { unseen })`
|
|
168
|
+
清理路径会将其移除。升级后首次 `index_repo` 或一次 watch 轮询即生效,不需要迁移脚本。
|
|
169
|
+
|
|
170
|
+
### 新增:父根未命中时列出独立子索引
|
|
171
|
+
|
|
172
|
+
上一节之后,父根不再保存子项目的内容,在父根下 `query_memory` 问子项目必然查不到。现在本根
|
|
173
|
+
记忆层(doc/symbol)无命中时,输出最前面会列出该根下自带 store 的子目录:
|
|
174
|
+
|
|
175
|
+
```
|
|
176
|
+
Note: this root keeps no copy of N nested project(s) — each has its own index.
|
|
177
|
+
Re-run query_memory with `root: <path>` for the one you want:
|
|
178
|
+
- <path>
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
- 触发条件只看记忆层:insight/experience 是全局层,与根无关,不应盖掉这条提示。
|
|
182
|
+
- 提示置于最前,因为 `truncate()` 截的是尾部,而 procedure 类条目单条可超 600 字符。
|
|
183
|
+
- 只影响 `query_memory` 的未命中路径,不涉及注入。
|
|
184
|
+
|
|
185
|
+
**验证**:`npm test` 518 项 / 28 个文件全绿;`npm run typecheck` 通过;`npm run eval:injection`
|
|
186
|
+
逐项不变(命中 14 / 假阳性 0 / 漏召 0,P = R = 1.00)。
|
|
187
|
+
|
|
188
|
+
## 0.5.11 (2026-09-25)
|
|
189
|
+
|
|
190
|
+
### 修复:doc↔symbol 链接不再物化进 entry(真实大仓库的内存与索引开销)
|
|
191
|
+
|
|
192
|
+
`linkedSymbols` 是**跨实体派生关系**(一个 doc chunk 链接到哪些符号,取决于符号表的当前状态),
|
|
193
|
+
旧实现却在索引时把它算好、写进 `entry` 并落盘。实测本仓库自己的 store(11698 文件 / 70119 条目):
|
|
194
|
+
|
|
195
|
+
- 17174 个 chunk 共 **3,948,420** 个链接槽位,只对应 15,456 个符号,单 chunk 最多 2231 个;
|
|
196
|
+
- 唯一的消费者 `query_memory` 只读前 5 个——存了消费量的 **46 倍**;
|
|
197
|
+
- 加载这个 store 的堆占用 **443MB**,链接槽位是其中最大的一块;
|
|
198
|
+
- 每个索引提交点还要对整库做 O(chunks × symbols) 重扫,并靠 `markFile` 标脏追失效,
|
|
199
|
+
否则"文档先索引、符号后到"会永久丢链接。
|
|
200
|
+
|
|
201
|
+
现在链接在**读取期**用每 store 的符号索引解算(`src/link.js` 的 `resolveLinkedSymbols`):
|
|
202
|
+
纯 latin 符号名走倒排表按整词查,CJK/混合名保留原有边界语义的正则回退;成本 O(本 chunk 词数),
|
|
203
|
+
与符号表规模无关。排序改为命中次数 → 名字长度 → id(旧实现交给消费者的前 5 个是符号表插入序,
|
|
204
|
+
即"任意 5 个",这是本次一并修掉的行为)。
|
|
205
|
+
|
|
206
|
+
- **内存**:同一 store 的加载堆占用 **443MB → 129–196MB**,RSS **581MB → ~300MB**;
|
|
207
|
+
- `linkedSymbols` 在加载时从旧 shard 剥离、写入时不再产生;磁盘上的存量 shard 会在该文件
|
|
208
|
+
下次重新索引时自然压实(不主动重写 11698 个 shard);
|
|
209
|
+
- 删除 `linkEntries` 导出,以及 `commitFileUpdates` 的 `link` 参数(它只为"中间批次跳过、
|
|
210
|
+
最后一批统一重建链接"而存在);`enhancer` / `index-doc` 里的链接重建调用一并移除;
|
|
211
|
+
- `query_memory` 的 `references` 输出格式不变,`test/run-test.mjs` 的链接用例改为断言
|
|
212
|
+
解析器行为,并新增"文档先索引、符号后到也能解出"与"limit / 排序"回归。
|
|
213
|
+
|
|
214
|
+
### 修复:其余派生/中间字段与两处无界缓存(体积、内存、轮询)
|
|
215
|
+
|
|
216
|
+
第一轮去掉了链接,这一轮把剩下的放大源和常驻开销一起收掉。同一个 store(11698 文件 /
|
|
217
|
+
70119 条目 / 索引源码 108.5MB)实测:**落盘 361MB → 93MB**(老 store 由下面的自动压实
|
|
218
|
+
收敛;再叠加 type-cache 自愈清理后是 **73MB**),加载 **1017ms → 279ms**,堆占用
|
|
219
|
+
**129–196MB → 101MB**,`recallItems` p50 **53ms → 47ms**。
|
|
220
|
+
|
|
221
|
+
- **`searchText` 不再落盘**(省 22.7MB)。它是 `weightedFieldText` 的纯派生结果,改成
|
|
222
|
+
`allEntries()` 在内存里按需物化;写入时与 `linkedSymbols` 一起剥离(`PERSISTED_DERIVED`)。
|
|
223
|
+
检索语义与输出不变;磁盘与加载解析变少,堆占用基本持平(物化改到首次查询时做)。
|
|
224
|
+
- **符号声明限长成一行**(`oneLineDeclaration`,≤200 字符)。TypeScript enricher 之前把
|
|
225
|
+
interface 的**全部成员**拼进 `typeSig`、再整体落进 `text`/`typeSig`(实测符号条目平均
|
|
226
|
+
1.25KB,其中 `text` 648B),而这两个字段没有任何读取方。现在 interface 只留前 6 个成员 +
|
|
227
|
+
`… +N more`,`typeSig` 不再落盘。代码层 **31% → 19% 源码**,README 声称的"一行声明"
|
|
228
|
+
由此第一次成立。
|
|
229
|
+
- **storeCache 按字节预算 + LRU**。原来只按个数(32),而单个大仓库 store 实测驻留
|
|
230
|
+
130–200MB;现在同时限制条数与估算驻留量(256MB,约 2.5KB/entry),命中会把条目挪到
|
|
231
|
+
队尾,热的不会先被逐出。
|
|
232
|
+
- **删除 type-cache(TS 增强结果缓存)**。它按内容哈希缓存增强结果,但三个增强入口
|
|
233
|
+
(lazy 的 `fs/observed`、watch 轮询、`index_repo`)**都只在"文件已变更并重新索引"之后**
|
|
234
|
+
才触发,此时内容哈希必然是新值——这个缓存永远命中不了。实测本仓库残留 9827 个文件
|
|
235
|
+
(`du` 41MB,内容其实 9.2MB,约 31MB 是 4KB 块开销)。现在 `load()` 会自愈删除该目录;
|
|
236
|
+
文件变更照常触发增强,进程内仍由 `enhanceQueue` 按 (relPath, 内容哈希) 去重。
|
|
237
|
+
- **watch 轮询空闲退避**。轮询是 O(树) 的 walkDir + 逐文件 stat(本仓库实测 58–87ms),
|
|
238
|
+
原来固定一轮、不管有没有改动都在磨 I/O。现在改成递归 `setTimeout`:无变化时翻倍
|
|
239
|
+
退避到最长 2 分钟,任何变化立即回到 `watchInterval`。基准间隔同时 **15s → 30s**
|
|
240
|
+
(仍可配置)。
|
|
241
|
+
- **删除死代码** `rankEntries` / `rankEntriesMerged` / `store.searchEntries`:它们每次调用
|
|
242
|
+
都 `buildBm25` 全库分词(70k 条目实测 p50 1.4s),生产路径没有调用方;相关测试改用线上
|
|
243
|
+
真正跑的 `rankEntriesStreaming` / `rankEntriesMergedScored`。
|
|
244
|
+
|
|
245
|
+
**存量 store 的压实是自动且有界的**:加载时把带派生字段的老分片放进待压实队列,此后任意一次
|
|
246
|
+
`save()`(watch 轮询、索引、写入都会触发)最多补写 `COMPACT_BATCH = 200` 个分片,直到队列
|
|
247
|
+
清空。升级后不需要重新索引,也不会在首次启动时一次性重写整库;想让它立刻跑完,随便索引一次
|
|
248
|
+
即可。为了让压实不产生副作用,IDF 缓存的失效键也从"任何脏写"改成 entries 的变更计数
|
|
249
|
+
(`_entriesVersion`)——IDF 只依赖 entries,只写经验/insight 或只做压实都不该重建它。
|
|
250
|
+
|
|
251
|
+
### 文档
|
|
252
|
+
|
|
253
|
+
- README 的 `npm test` 断言与实测对齐:360 → **476**(核心 184 → 205、host-contract 9 → 10;
|
|
254
|
+
补上此前漏记的 task-view 6 / root-guards 79 / store-gitignore 9);
|
|
255
|
+
- 两份 README 的"交叉链接"机制描述由"索引后挂载到条目"改为"读取期按当前符号表解算";
|
|
256
|
+
- `watchInterval` 文档补充空闲退避语义;"紧凑性"一节改用实测区间(符号稀疏项目 ~0.5%,
|
|
257
|
+
符号密集的 TS monorepo ~19%),不再把 0.5% 当普遍值。
|
|
2
258
|
|
|
3
259
|
## 0.5.10 (2026-09-24)
|
|
4
260
|
|
package/LICENSE
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
MIT License
|
|
2
2
|
|
|
3
|
-
Copyright (c) 2026
|
|
3
|
+
Copyright (c) 2026 00080000 <3388065969@qq.com>
|
|
4
4
|
|
|
5
5
|
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
6
|
of this software and associated documentation files (the "Software"), to deal
|
package/README.md
CHANGED
|
@@ -91,7 +91,7 @@ The design follows four principles:
|
|
|
91
91
|
|
|
92
92
|
- **Volatility** — context is ephemeral; it is lost when a session is compacted.
|
|
93
93
|
- **Persistence** — the **memory** is stored on disk and survives compaction and new sessions.
|
|
94
|
-
- **Compactness** — the code layer stores one declaration line per symbol
|
|
94
|
+
- **Compactness** — the code layer stores one bounded declaration line per symbol (≤200 chars) and the document layer keeps a ≤300-char `summary` plus a bounded `terms` set per chunk. **Derived data is never stored**: doc→symbol links and the BM25 `searchText` are computed at read time. How small the index ends up depends on symbol density and chunk length, so treat these as measurements of specific corpora (2026-09-25), not as guarantees: a code-only Vue app (289 files) lands at **325 bytes/entry ≈ 21% of source**, while a symbol-dense TypeScript monorepo (12,408 files / 106 MB of code + 14 MB of docs) measures **22% of source for the code layer (553 bytes/entry overall)** and **130% for the document layer**. Doc-heavy corpora are the largest per entry: a 274-document workspace (PDFs and Markdown) stores **1,506 bytes/entry**.
|
|
95
95
|
- **Verifiability** — **recalls** carry a `path:line` citation where applicable, so the agent can confirm details against the source.
|
|
96
96
|
|
|
97
97
|
Building the **memory** does not require an upfront scan: files are memorized as the model reads them, so the **memory** grows to cover exactly what has been worked with. Re-reading a file that has not changed is a no-op (content hash), so the **memory** stays fresh with minimal ongoing overhead.
|
|
@@ -117,7 +117,7 @@ The store is per-project and follows the codebase: changed files are re-extracte
|
|
|
117
117
|
Stores created before v0.2.0 (single `entries.json` / `index.json`) migrate automatically and idempotently on first load. Within one dsh process, all tool calls share a single in-memory store per project, so hot-path indexing writes only the shard that changed.
|
|
118
118
|
|
|
119
119
|
- **Incremental** — content hash per file; only changed files are re-extracted.
|
|
120
|
-
- **Cross-linking** —
|
|
120
|
+
- **Cross-linking** — when `query_memory` returns a doc chunk, it resolves the symbols that chunk mentions against the **current** symbol table and appends them as `references`. Links are computed at read time, so they cannot go stale and are not stored in the index (a doc indexed before its symbols still links correctly).
|
|
121
121
|
- **Query expansion** — when `llmQueryExpansion` is on, `query_memory` asks `ctx.llm` to rewrite the query into several variants (synonyms, EN/CN, identifier guesses) and merges BM25 scores across variants; when off, queries never touch the LLM. Indexing itself is model-free: keywords are rule-derived (title-weighted top terms), and doc↔symbol links surface English symbol names from Chinese hits.
|
|
122
122
|
- **Consistency** — the fact layer follows the codebase (hash re-extract / remove-on-delete); the experience layer is retrieval-only with supersede and `forget`. Store writes are serialized per memory directory; the lock is in-process, so avoid running multiple dsh instances against the same project store concurrently.
|
|
123
123
|
|
|
@@ -152,7 +152,7 @@ The workflow panel is collapsible, automatically adapts to dsh and theme plugin
|
|
|
152
152
|
| `lazyIndexing` | true | index files the moment the model reads them (`fs/observed`) |
|
|
153
153
|
| `autoIndexOnFirstUse` | false | full scan of the current working directory on plugin load (opt-in) |
|
|
154
154
|
| `watch` | true | enable the background refresh |
|
|
155
|
-
| `watchInterval` |
|
|
155
|
+
| `watchInterval` | 30 | base poll interval (seconds); idle polls back off up to 2 minutes and reset to this value on any change |
|
|
156
156
|
| `maxScanFiles` | 20000 | hard cap on files per scan pass; a truncated pass is reported and does not remove the entries it did not reach. Set `0` to disable the cap |
|
|
157
157
|
| `maxScanDepth` | 12 | hard cap on directory depth per scan pass. Set `0` to disable |
|
|
158
158
|
| `allowUnsafeRoots` | false | allow **explicit** tool calls (`index_repo`/`watch_repo`/`remember` with a `root`) to target a directory on the excluded list. Automatic paths (lazy indexing, session audit, TaskBridge, `autoIndexOnFirstUse`) stay inert in these directories regardless |
|
|
@@ -211,7 +211,7 @@ Settings live in the plugin's config object. To change them, add an override ent
|
|
|
211
211
|
autoIndexOnFirstUse: false # off: no upfront full scan (default)
|
|
212
212
|
llmQueryExpansion: false # off: do not spend tokens on LLM query expansion (default)
|
|
213
213
|
watch: true # on: background refresh for watched roots (default)
|
|
214
|
-
watchInterval:
|
|
214
|
+
watchInterval: 30 # base poll interval; idle polls back off to at most 2 min
|
|
215
215
|
maxScanFiles: 20000 # per-scan file cap (truncation is reported, never deletes)
|
|
216
216
|
maxScanDepth: 12 # per-scan directory-depth cap
|
|
217
217
|
enableTypeScript: true # on: L2 TS enhancement when TS is installed (default)
|
|
@@ -233,29 +233,35 @@ where `config.yml` contains the same override block.
|
|
|
233
233
|
|
|
234
234
|
## Performance
|
|
235
235
|
|
|
236
|
-
###
|
|
236
|
+
### Measured on real projects (2026-09-25)
|
|
237
237
|
|
|
238
|
-
|
|
239
|
-
|----------|-------|----------|
|
|
240
|
-
| Full cold index | 5,000 files / 20k entries | 269 ms avg (p50 267) |
|
|
241
|
-
| Cold load | 5,000 files | 40 ms |
|
|
242
|
-
| Hot lazy re-index (single file) | 5k files | p50 2.4 ms / max 5.5 ms |
|
|
243
|
-
| query_memory (cached) | 5k files / 20k entries | p50 2.6 ms / p95 5.4 ms |
|
|
244
|
-
| query_memory (cached) | 1k files / 4k entries | p50 0.6 ms / p95 1.6 ms |
|
|
245
|
-
| Full cold index | 10,000 files / 40k entries | 551 ms avg (p50 528) |
|
|
246
|
-
| Cold load | 10,000 files | 90 ms |
|
|
247
|
-
| Hot lazy re-index (single file) | 10k files | p50 5.4 ms / max 9.2 ms |
|
|
238
|
+
Four corpora, one machine (Node 24.19, 20 vCPU, Linux file system), each run twice with the **second, warm-cache run** quoted. "Cold index" is a full index pass (walk + sha256 + extract + commit); "query" runs the shipped scorer over 100 sampled queries; "re-index 1 file" is the watch/lazy hot path.
|
|
248
239
|
|
|
249
|
-
|
|
240
|
+
| Corpus | Files / entries | Cold index | Cold load | Query p50 / p95 | Re-index 1 file | Store content / on disk | Heap after load |
|
|
241
|
+
|--------|-----------------|-----------|-----------|-----------------|-----------------|------------------------|-----------------|
|
|
242
|
+
| Vue 3 + Vite app (code only) | 289 / 2,142 | 283 ms | 5.2 ms | 0.86 / 1.8 ms | 0.4 ms | 0.66 MB / 1.52 MB | 6.1 MB |
|
|
243
|
+
| Docs + PDFs workspace (274 docs) | 286 / 2,120 | 6.5 s | 12.4 ms | 4.6 / 13.9 ms | 0.4 ms | 3.05 MB / 3.63 MB | 9.0 MB |
|
|
244
|
+
| TypeScript monorepo, 3,000-file slice | 3,000 / 17,733 | 2.1 s | 59 ms | 10.2 / 21.4 ms | 1.8 ms | 11.0 MB / 19.0 MB | 22.6 MB |
|
|
245
|
+
| TypeScript monorepo, whole tree | 12,408 / 79,168 | 8.6 s | 239 ms | 45.8 / 89.3 ms | 7.4 ms | 41.7 MB / 74.7 MB | 73.9 MB |
|
|
250
246
|
|
|
251
|
-
|
|
247
|
+
**How it scales.** Re-indexing a changed file costs O(file), not O(corpus) — 0.4–7.4 ms across every corpus above. Cold load (≈19 µs/file), query (≈0.6 µs/entry) and resident heap (≈1.4 KB/entry once the first query materializes `searchText`) grow linearly with the index, which keeps small and mid-size projects in the single-digit-millisecond range.
|
|
252
248
|
|
|
253
|
-
|
|
254
|
-
|---------|-------|---------|------------|-----------|
|
|
255
|
-
| Java Spring Boot backend | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
|
|
256
|
-
| Vue 3 + Vite frontend | 289 | 2,141 | 1.0 MB | ~0.5 KB |
|
|
249
|
+
> Two notes on method: `read+hash` depends on the OS page cache (2.5 s cold vs 0.3 s warm on the 12.4k-file tree), so the warm run is the one quoted; and this benchmark drifts by up to ~20% across days on the same machine, so compare numbers measured in the same session.
|
|
257
250
|
|
|
258
|
-
|
|
251
|
+
### Synthetic Benchmark (Node 24.19, 20 vCPU, Linux file system)
|
|
252
|
+
|
|
253
|
+
| Scenario | Scale | Measured |
|
|
254
|
+
|----------|-------|----------|
|
|
255
|
+
| Full cold index | 5,000 files / 20k entries | 373 ms avg (p50 374) |
|
|
256
|
+
| Cold load | 5,000 files | 56 ms |
|
|
257
|
+
| Hot lazy re-index (single file) | 5k files | p50 2.8 ms / max 3.5 ms |
|
|
258
|
+
| query_memory (cached) | 5k files / 20k entries | p50 3.3 ms / p95 6.9 ms |
|
|
259
|
+
| query_memory (cached) | 1k files / 4k entries | p50 0.7 ms / p95 1.5 ms |
|
|
260
|
+
| Full cold index | 10,000 files / 40k entries | 696 ms avg (p50 668) |
|
|
261
|
+
| Cold load | 10,000 files | 123 ms |
|
|
262
|
+
| Hot lazy re-index (single file) | 10k files | p50 5.9 ms / max 13.5 ms |
|
|
263
|
+
|
|
264
|
+
> Synthetic benchmark: generated code (~4–5 symbols/file), Node 24.19 on 20 vCPU / Linux file system, measured 2026-09-25. Reproduce with `npm run bench:synthetic -- 5000` (harness: `scripts/bench-synthetic.mjs`). Measures pure indexing overhead without LLM calls. query_memory uses the IDF cache + searchText materialized on first use; the first query after a write rebuilds IDF (**142 ms at 40k entries**, 67 ms at 20k, 14 ms at 4k), subsequent queries hit the cache.
|
|
259
265
|
|
|
260
266
|
### Reproduce it on your own project
|
|
261
267
|
|
|
@@ -267,16 +273,17 @@ npm run bench -- /path/to/your/project
|
|
|
267
273
|
node scripts/bench.mjs /path/to/your/project [--json] [--samples 100] [--no-pdf] [--keep]
|
|
268
274
|
```
|
|
269
275
|
|
|
270
|
-
It reports the cold index split into read+hash / extract / commit, cold load, IDF rebuild, cold and hot query latency (p50/p95/max over 100 sampled queries through the shipped scorer), single-file hot re-index, store size
|
|
276
|
+
It reports the cold index split into read+hash / extract / commit, cold load, IDF rebuild, cold and hot query latency (p50/p95/max over 100 sampled queries through the shipped scorer), single-file hot re-index, store content vs on-disk size, resident heap (after load and after the first query), bytes per entry and RSS. Example — the Vue app row above:
|
|
271
277
|
|
|
272
278
|
```
|
|
273
|
-
cold index
|
|
274
|
-
store 1.
|
|
275
|
-
|
|
276
|
-
|
|
279
|
+
cold index 283 ms (read+hash 11 ms · extract 256 ms · commit 14 ms) ← 2nd, warm-cache run
|
|
280
|
+
store 0.66 MB content · 1.52 MB on disk · 325 bytes/entry · cold load 5.2 ms
|
|
281
|
+
memory heap 6.1 MB after load → 6.7 MB after the first query (RSS 62 MB)
|
|
282
|
+
hot query p50 0.86 ms · p95 1.8 ms (2,142 entries)
|
|
283
|
+
re-index 1 file p50 0.4 ms
|
|
277
284
|
```
|
|
278
285
|
|
|
279
|
-
|
|
286
|
+
Pass `--queries your-queries.json` to run the labeled-set method (hit@5 / hit@10 / MRR) against your own project.
|
|
280
287
|
|
|
281
288
|
## Design tradeoffs
|
|
282
289
|
|
|
@@ -288,7 +295,7 @@ Two caveats we would rather state than hide: `read+hash` depends on the OS page
|
|
|
288
295
|
- **Model-facing memory: the agent writes, no human in the loop** — no human approval step: the consumer of this memory is the agent, and agents are usually headless, so memory that only promotes when someone clicks a card would never promote at all. `draft` is a provenance marker plus an evidence threshold, not an approval queue — the one inferring writer, `reflection` (off by default), writes task-level drafts only, and drafts never reach recall or injection.
|
|
289
296
|
- **Full entries returned directly** — no "minimal index first, fetch details in a second call": entries are already compact, so returning them whole is both more verifiable and one round-trip cheaper.
|
|
290
297
|
- **`forget` by query is aggressive; use IDs for precision** — no confirmation prompt, recycle bin, or exact-match-only mode: experience notes are low-risk, high-volume, and retrieval-only, so stale noise hurts more than an over-broad delete. For exact deletion use the ID shown by `query_memory`.
|
|
291
|
-
- **TypeScript enhancement is optional, lazy, and cached** — the L2 TS Compiler API runs asynchronously on a priority queue (P0 `fs/observed`, P1 `watch`, P2 `index_repo`) and caches results by content hash; TS is never required and enhancement never blocks: requiring it would make non-TS projects uninstallable, and blocking would stall `index_repo` on large projects. `npm i -D typescript@5|6` is the entire setup, and a missing TS falls back to the L1 regex scanner.
|
|
298
|
+
- **TypeScript enhancement is optional, lazy, and cached** — the L2 TS Compiler API runs asynchronously on a priority queue (P0 `fs/observed`, P1 `watch`, P2 `index_repo`) and caches results by content hash; TS is never required and enhancement never blocks: requiring it would make non-TS projects uninstallable, and blocking would stall `index_repo` on large projects. `npm i -D typescript@5|6` is the entire setup, and a missing TS falls back to the L1 regex scanner. The **default lib is not loaded** (`noLib`): inference that depends on global types (`Promise`/`Array`/DOM) degrades to `any`/`unknown`, while explicitly annotated types are unaffected.
|
|
292
299
|
- **Subagent sessions are out of scope for now**
|
|
293
300
|
|
|
294
301
|
## Development (for contributors)
|
|
@@ -297,7 +304,7 @@ These commands are for **maintaining the plugin code** — regular users do not
|
|
|
297
304
|
|
|
298
305
|
```bash
|
|
299
306
|
npm install
|
|
300
|
-
npm test #
|
|
307
|
+
npm test # 539 tests (214 core + 16 TaskBridge + 12 insight-store + 9 insight-actions + 8 doc-index + 7 auto-inject + 10 host-contract + 5 reflection + 4 llm-route + 2 client-hints + 10 recall + 14 readiness + 7 insight-derive + 7 readiness-eval + 6 ops + 11 injection-audit + 5 injection-budget + 6 injection-scenarios + 18 bugfix-0.5.7 + 3 client-icons + 10 client-slash + 5 workflow-command + 7 client-session-id + 6 task-view + 79 root-guards + 9 store-gitignore + 22 store-cache + 27 enhancer)
|
|
301
308
|
npm run eval:injection # scenario P/R on the synthetic pool: 14/14 hits, 0 false positives, control group clean
|
|
302
309
|
npm run eval:injection -- --store .dsh-project-memory/insights.json # replay on YOUR store; control group is a hard gate
|
|
303
310
|
npm run selfcheck:triggers # which entries can still push, which declarations are dead (reads your local store)
|
|
@@ -308,4 +315,6 @@ Release notes live in [`CHANGELOG.md`](CHANGELOG.md) and on [GitHub Releases](ht
|
|
|
308
315
|
|
|
309
316
|
## License
|
|
310
317
|
|
|
311
|
-
MIT
|
|
318
|
+
MIT — see [`LICENSE`](LICENSE).
|
|
319
|
+
|
|
320
|
+
Copyright (c) 2026 00080000 <3388065969@qq.com>
|
package/README.zh-CN.md
CHANGED
|
@@ -88,7 +88,7 @@ dsh plugin --profile web add /path/to/dsh-project-memory.tgz
|
|
|
88
88
|
|
|
89
89
|
- **易失性** — 上下文是临时的,会话压缩即丢失。
|
|
90
90
|
- **持久性** — **记忆**存于磁盘,跨压缩与会话保留。
|
|
91
|
-
- **紧凑性** —
|
|
91
|
+
- **紧凑性** — 代码层每个符号只存一行声明(≤200 字符),文档层每个 chunk 保留 ≤300 字符的 `summary` 与有界的 `terms`。**派生数据一律不落盘**:doc→symbol 链接与 BM25 的 `searchText` 都在读取期计算。最终体积取决于符号密度与 chunk 长度,下面是特定语料的实测值(2026-09-25),不是承诺:纯代码的 Vue 应用(289 文件)实测 **325 bytes/条目 ≈ 源码 21%**;符号密集的 TypeScript monorepo(12,408 文件 / 代码 106 MB + 文档 14 MB)实测代码层 **占源码 22%**(整库 553 bytes/条目)、文档层 **占其源码 130%**。文档占比高的语料单条目最大:一个 274 篇文档的工作区(PDF + Markdown)实测 **1,506 bytes/条目**。
|
|
92
92
|
- **可核验性** — **召回**在适用时携带 `路径:行号` 引用,agent 可对照源文件核实。
|
|
93
93
|
|
|
94
94
|
构建**记忆**无需预先全量扫描:文件在模型读取时被记忆,**记忆**恰好覆盖实际处理过的内容。未变更的文件重读是空操作(内容哈希),因此**记忆**的持续维护开销很低。
|
|
@@ -114,7 +114,7 @@ dsh plugin --profile web add /path/to/dsh-project-memory.tgz
|
|
|
114
114
|
v0.2.0 之前创建的库(单文件 `entries.json` / `index.json`)在首次加载时自动幂等迁移。同一个 dsh 进程内,所有工具调用共享每个项目的单一内存 store 实例,热路径索引只写发生变化的那一个分片。
|
|
115
115
|
|
|
116
116
|
- **增量** — 按文件内容哈希,仅重新抽取变更文件。
|
|
117
|
-
- **交叉链接** —
|
|
117
|
+
- **交叉链接** — `query_memory` 返回文档 chunk 时,按**当前**符号表解算它提到的符号,以 `references` 带出。链接在读取期解算、不落盘,因此不会过期(文档先索引、符号后到也能链上),也不占存储。
|
|
118
118
|
- **查询扩展** — `llmQueryExpansion` 开启时,`query_memory` 让 `ctx.llm` 将查询改写为多个变体(同义词、中英、符号名猜测),再跨变体合并 BM25 分数;关闭时查询完全不碰 LLM。索引本身不调用模型:keywords 由规则推导(标题加权词项),doc↔symbol 链接也会从中文命中带出英文符号名。
|
|
119
119
|
- **一致性** — 事实层跟随代码库(哈希重抽 / 删除即移除);经验层仅检索,配合覆盖与 `forget` 机制。每个记忆目录的写入走同步事务 `store.commit(fn)`:fn 内完成校验与变更、成功后才原子落盘,单进程内天然串行;请避免多个 dsh 实例同时写同一项目存储。
|
|
120
120
|
|
|
@@ -149,7 +149,7 @@ TaskPanel (Container)
|
|
|
149
149
|
| `lazyIndexing` | true | 模型读取文件的瞬间即索引(`fs/observed`) |
|
|
150
150
|
| `autoIndexOnFirstUse` | false | 插件加载时对当前工作目录做全量扫描(可选) |
|
|
151
151
|
| `watch` | true | 启用后台刷新 |
|
|
152
|
-
| `watchInterval` |
|
|
152
|
+
| `watchInterval` | 30 | 基础轮询间隔(秒);空闲时逐步退避到最长 2 分钟,一有变化立即回到该值 |
|
|
153
153
|
| `maxScanFiles` | 20000 | 单次扫描的文件数硬上限;被截断时会在报告里说明,且不会删除没扫到的条目。设 `0` 取消上限 |
|
|
154
154
|
| `maxScanDepth` | 12 | 单次扫描的目录深度硬上限。设 `0` 取消 |
|
|
155
155
|
| `allowUnsafeRoots` | false | 允许**显式**工具调用(带 `root` 的 `index_repo`/`watch_repo`/`remember`)指向排除名单上的目录。自动路径(懒索引、会话审计、TaskBridge、`autoIndexOnFirstUse`)无论此项如何都不会越权 |
|
|
@@ -208,7 +208,7 @@ store 建在被索引的目录树里,并且**自我忽略**:它在自己目
|
|
|
208
208
|
autoIndexOnFirstUse: false # 关闭:不做加载时的全量扫描(默认)
|
|
209
209
|
llmQueryExpansion: false # 关闭:不用 LLM 扩展查询,节省 token(默认)
|
|
210
210
|
watch: true # 开启:被监听根目录后台保持新鲜(默认)
|
|
211
|
-
watchInterval:
|
|
211
|
+
watchInterval: 30 # 基础轮询间隔;空闲时退避到最长 2 分钟
|
|
212
212
|
maxScanFiles: 20000 # 单次扫描文件上限(截断会报告,且不会误删旧条目)
|
|
213
213
|
maxScanDepth: 12 # 单次扫描目录深度上限
|
|
214
214
|
enableTypeScript: true # 开启:装了 TS 时启用 L2 语义增强(默认)
|
|
@@ -230,29 +230,35 @@ dsh web --patch ./config.yml
|
|
|
230
230
|
|
|
231
231
|
## 性能
|
|
232
232
|
|
|
233
|
-
###
|
|
233
|
+
### 真实项目实测(2026-09-25)
|
|
234
234
|
|
|
235
|
-
|
|
236
|
-
|------|------|------|
|
|
237
|
-
| 批量冷记忆构建 | 5,000 文件 / 20k 条目 | 269 ms 均值(p50 267)|
|
|
238
|
-
| 冷加载 | 5,000 文件 | 40 ms |
|
|
239
|
-
| 热路径懒记忆 | 单文件重记忆+落盘 | p50 2.4 ms / 最大 5.5 ms (5k) |
|
|
240
|
-
| query_memory (缓存命中) | 5k 文件 / 20k 条目 | p50 2.6 ms / p95 5.4 ms |
|
|
241
|
-
| query_memory (缓存命中) | 1k 文件 / 4k 条目 | p50 0.6 ms / p95 1.6 ms |
|
|
242
|
-
| 批量冷记忆构建 | 10,000 文件 / 40k 条目 | 551 ms 均值(p50 528)|
|
|
243
|
-
| 冷加载 | 10,000 文件 | 90 ms |
|
|
244
|
-
| 热路径懒记忆 | 单文件重记忆+落盘 | p50 5.4 ms / 最大 9.2 ms (10k) |
|
|
235
|
+
四个语料、同一台机器(Node 24.19,20 vCPU,Linux 文件系统),每个跑两遍、引用**第二次(页缓存已热)**的数据。「冷索引」= 完整索引一轮(walk + sha256 + 抽取 + 落盘);「查询」= 线上同一套 scorer 跑 100 条采样;「单文件重索引」= watch / 懒索引热路径。
|
|
245
236
|
|
|
246
|
-
|
|
237
|
+
| 语料 | 文件数 / 条目数 | 冷索引 | 冷加载 | 查询 p50 / p95 | 单文件重索引 | 存储内容 / 落盘 | 加载后堆 |
|
|
238
|
+
|------|----------------|--------|--------|----------------|--------------|----------------|----------|
|
|
239
|
+
| Vue 3 + Vite 应用(纯代码) | 289 / 2,142 | 283 ms | 5.2 ms | 0.86 / 1.8 ms | 0.4 ms | 0.66 MB / 1.52 MB | 6.1 MB |
|
|
240
|
+
| 文档 + PDF 工作区(274 篇文档) | 286 / 2,120 | 6.5 s | 12.4 ms | 4.6 / 13.9 ms | 0.4 ms | 3.05 MB / 3.63 MB | 9.0 MB |
|
|
241
|
+
| TypeScript monorepo,3,000 文件切片 | 3,000 / 17,733 | 2.1 s | 59 ms | 10.2 / 21.4 ms | 1.8 ms | 11.0 MB / 19.0 MB | 22.6 MB |
|
|
242
|
+
| TypeScript monorepo,整棵树 | 12,408 / 79,168 | 8.6 s | 239 ms | 45.8 / 89.3 ms | 7.4 ms | 41.7 MB / 74.7 MB | 73.9 MB |
|
|
247
243
|
|
|
248
|
-
|
|
244
|
+
**扩展形状。** 重索引一个变更文件的成本是 O(文件)、不是 O(语料)——上面每个语料都在 0.4–7.4 ms。冷加载(≈19 µs/文件)、查询(≈0.6 µs/条目)与常驻堆(首次查询物化 `searchText` 后 ≈1.4 KB/条目)随索引规模线性增长,因此小型与中型项目都落在个位数毫秒。
|
|
249
245
|
|
|
250
|
-
|
|
251
|
-
|------|--------|--------|----------|--------|
|
|
252
|
-
| Java Spring Boot 后端 | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
|
|
253
|
-
| Vue 3 + Vite 前端 | 289 | 2,141 | 1.0 MB | ~0.5 KB |
|
|
246
|
+
> 两条口径说明:`read+hash` 受操作系统页缓存影响(12.4k 文件时冷缓存 2.5 s、热缓存 0.3 s),所以引用的是热缓存那一遍;同一台机器上跨天跑同一基准会有约 20% 以内的漂移,请只比较同一会话内测出的数字。
|
|
254
247
|
|
|
255
|
-
|
|
248
|
+
### 合成基准测试(Node 24.19,20 vCPU,Linux 文件系统)
|
|
249
|
+
|
|
250
|
+
| 场景 | 规模 | 实测 |
|
|
251
|
+
|------|------|------|
|
|
252
|
+
| 批量冷记忆构建 | 5,000 文件 / 20k 条目 | 373 ms 均值(p50 374)|
|
|
253
|
+
| 冷加载 | 5,000 文件 | 56 ms |
|
|
254
|
+
| 热路径懒记忆 | 单文件重记忆+落盘 | p50 2.8 ms / 最大 3.5 ms (5k) |
|
|
255
|
+
| query_memory (缓存命中) | 5k 文件 / 20k 条目 | p50 3.3 ms / p95 6.9 ms |
|
|
256
|
+
| query_memory (缓存命中) | 1k 文件 / 4k 条目 | p50 0.7 ms / p95 1.5 ms |
|
|
257
|
+
| 批量冷记忆构建 | 10,000 文件 / 40k 条目 | 696 ms 均值(p50 668)|
|
|
258
|
+
| 冷加载 | 10,000 文件 | 123 ms |
|
|
259
|
+
| 热路径懒记忆 | 单文件重记忆+落盘 | p50 5.9 ms / 最大 13.5 ms (10k) |
|
|
260
|
+
|
|
261
|
+
> 合成基准:生成代码(~4–5 符号/文件),Node 24.19 / 20 vCPU / Linux 文件系统,实测于 2026-09-25。复现命令 `npm run bench:synthetic -- 5000`(脚本 `scripts/bench-synthetic.mjs`)。测量纯索引开销,不含 LLM 调用。query_memory 使用 IDF 缓存 + 首次使用时物化 searchText;写入后的首次查询会重建 IDF(**40k 条目 142 ms**,20k 条目 67 ms,4k 条目 14 ms),后续查询命中缓存。
|
|
256
262
|
|
|
257
263
|
### 自己复现这些数字
|
|
258
264
|
|
|
@@ -264,16 +270,17 @@ npm run bench -- /你的/项目路径
|
|
|
264
270
|
node scripts/bench.mjs /你的/项目路径 [--json] [--samples 100] [--no-pdf] [--keep]
|
|
265
271
|
```
|
|
266
272
|
|
|
267
|
-
输出包含:冷索引(拆成 read+hash / extract / commit 三段)、冷加载、IDF 重建、冷查询与热查询延迟(走线上同一套 scorer,100 条采样报 p50/p95/max
|
|
273
|
+
输出包含:冷索引(拆成 read+hash / extract / commit 三段)、冷加载、IDF 重建、冷查询与热查询延迟(走线上同一套 scorer,100 条采样报 p50/p95/max)、单文件热重索引、存储内容与落盘体积、常驻堆(加载后与首次查询后)、单条目字节数。示例——上表里的 Vue 应用:
|
|
268
274
|
|
|
269
275
|
```
|
|
270
|
-
冷索引
|
|
271
|
-
存储 1.
|
|
272
|
-
|
|
273
|
-
|
|
276
|
+
冷索引 283 ms (read+hash 11 ms · extract 256 ms · commit 14 ms)← 第二次、页缓存已热
|
|
277
|
+
存储 内容 0.66 MB · 落盘 1.52 MB · 325 bytes/条目 · 冷加载 5.2 ms
|
|
278
|
+
内存 加载后堆 6.1 MB → 首次查询后 6.7 MB(RSS 62 MB)
|
|
279
|
+
热查询 p50 0.86 ms · p95 1.8 ms (2,142 条目)
|
|
280
|
+
单文件重索引 p50 0.4 ms
|
|
274
281
|
```
|
|
275
282
|
|
|
276
|
-
|
|
283
|
+
带 `--queries 你的查询集.json` 可以在你自己的项目上跑标注集方法(hit@5 / hit@10 / MRR)。
|
|
277
284
|
|
|
278
285
|
## 设计取舍
|
|
279
286
|
|
|
@@ -285,7 +292,7 @@ node scripts/bench.mjs /你的/项目路径 [--json] [--samples 100] [--no-pdf]
|
|
|
285
292
|
- **面向模型的记忆:agent 自己写,不把人放进回路** — 不要求人工批准:记忆的消费方是 agent,而 agent 通常是无头的,只在有人点卡片时才升级的记忆等于永远不会升级。`draft` 是「来源标记 + 佐证门槛」而不是审批队列——唯一的推断型写入者 `reflection`(默认关闭)只写任务级草稿,草稿不进召回与注入。
|
|
286
293
|
- **直接返回完整条目** — 条目本就紧凑,完整返回更可核验,也少一轮往返。
|
|
287
294
|
- **`forget` 按关键词激进;精确请用 ID** — 不做确认弹窗、回收站或仅精确匹配:经验笔记低风险、高量、仅用于检索,陈旧噪音比误删更伤。精确删用 `query_memory` 输出里的 ID。
|
|
288
|
-
- **TS 增强可选、异步、缓存** — L2 TS Compiler API 在优先级队列异步跑(P0 `fs/observed`、P1 `watch`、P2 `index_repo`),结果按内容哈希缓存;不强制 TS、也不阻塞索引:强制会让非 TS 项目装不上,阻塞会卡死大项目的 `index_repo`;`npm i -D typescript@5|6` 即自动启用,没有 TS 时回退 L1
|
|
295
|
+
- **TS 增强可选、异步、缓存** — L2 TS Compiler API 在优先级队列异步跑(P0 `fs/observed`、P1 `watch`、P2 `index_repo`),结果按内容哈希缓存;不强制 TS、也不阻塞索引:强制会让非 TS 项目装不上,阻塞会卡死大项目的 `index_repo`;`npm i -D typescript@5|6` 即自动启用,没有 TS 时回退 L1 正则。**默认 lib 不加载**(编译期 `noLib`):依赖全局类型(`Promise`/`Array`/DOM)的推导会退化成 `any`/`unknown`,显式标注的类型不受影响。
|
|
289
296
|
- **子代理会话暂不纳入(以后可能做)**
|
|
290
297
|
|
|
291
298
|
## 开发(面向贡献者)
|
|
@@ -294,7 +301,7 @@ node scripts/bench.mjs /你的/项目路径 [--json] [--samples 100] [--no-pdf]
|
|
|
294
301
|
|
|
295
302
|
```bash
|
|
296
303
|
npm install
|
|
297
|
-
npm test #
|
|
304
|
+
npm test # 539 项测试(核心 214 + TaskBridge 16 + insight-store 12 + insight-actions 9 + doc-index 8 + auto-inject 7 + host-contract 10 + reflection 5 + llm-route 4 + client-hints 2 + recall 10 + readiness 14 + insight-derive 7 + readiness-eval 7 + ops 6 + injection-audit 11 + injection-budget 5 + injection-scenarios 6 + bugfix-0.5.7 18 + client-icons 3 + client-slash 10 + workflow-command 5 + client-session-id 7 + task-view 6 + root-guards 79 + store-gitignore 9 + store-cache 22 + enhancer 27)
|
|
298
305
|
npm run eval:injection # 合成池上的场景 P/R:命中 14/14、假阳性 0、对照组零注入
|
|
299
306
|
npm run eval:injection -- --store .dsh-project-memory/insights.json # 用你自己的 store 重放;对照组是硬闸门
|
|
300
307
|
npm run selfcheck:triggers # 哪些条目还推得动、哪些声明是死的(读你本地的 store)
|
|
@@ -305,4 +312,6 @@ npm run bench -- /你的/项目路径 # 对任意项目量索引/查询性能
|
|
|
305
312
|
|
|
306
313
|
## 许可证
|
|
307
314
|
|
|
308
|
-
MIT
|
|
315
|
+
MIT,全文见 [`LICENSE`](LICENSE)。
|
|
316
|
+
|
|
317
|
+
Copyright (c) 2026 00080000 <3388065969@qq.com>
|
package/cordis.patch.yml
CHANGED
package/package.json
CHANGED
|
@@ -1,7 +1,8 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@yolk_vat-y/dsh-project-memory",
|
|
3
|
-
"version": "0.5.
|
|
3
|
+
"version": "0.5.12",
|
|
4
4
|
"description": "Persistent project memory for dsh agents: index docs (PDF/Markdown/text) and code symbols into a searchable per-workspace store, recall them with cited sources, and keep experience entries (problems -> solutions) searchable on demand.",
|
|
5
|
+
"author": "00080000 <3388065969@qq.com>",
|
|
5
6
|
"type": "module",
|
|
6
7
|
"main": "src/index.js",
|
|
7
8
|
"files": [
|
|
@@ -18,7 +19,7 @@
|
|
|
18
19
|
"url": "https://github.com/00080000/dsh-project-memory.git"
|
|
19
20
|
},
|
|
20
21
|
"scripts": {
|
|
21
|
-
"test": "node test/run-test.mjs && node test/taskbridge.test.mjs && node test/insight-store.test.mjs && node test/reflection-pipeline.test.mjs && node test/auto-inject.test.mjs && node test/insight-actions.test.mjs && node test/host-contract.test.mjs && node test/llm-route.test.mjs && node test/doc-index.test.mjs && node test/client-hints.test.mjs && node test/recall.test.mjs && node test/readiness.test.mjs && node test/insight-derive.test.mjs && node test/readiness-eval.test.mjs && node test/ops.test.mjs && node test/injection-audit.test.mjs && node test/injection-budget.test.mjs && node test/injection-scenarios.test.mjs && node test/bugfix-0.5.7.test.mjs && node test/client-icons.test.mjs && node test/client-slash.test.mjs && node test/workflow-command.test.mjs && node test/client-session-id.test.mjs && node test/task-view.test.mjs && node test/root-guards.test.mjs && node test/store-gitignore.test.mjs",
|
|
22
|
+
"test": "node test/run-test.mjs && node test/taskbridge.test.mjs && node test/insight-store.test.mjs && node test/reflection-pipeline.test.mjs && node test/auto-inject.test.mjs && node test/insight-actions.test.mjs && node test/host-contract.test.mjs && node test/llm-route.test.mjs && node test/doc-index.test.mjs && node test/client-hints.test.mjs && node test/recall.test.mjs && node test/readiness.test.mjs && node test/insight-derive.test.mjs && node test/readiness-eval.test.mjs && node test/ops.test.mjs && node test/injection-audit.test.mjs && node test/injection-budget.test.mjs && node test/injection-scenarios.test.mjs && node test/bugfix-0.5.7.test.mjs && node test/client-icons.test.mjs && node test/client-slash.test.mjs && node test/workflow-command.test.mjs && node test/client-session-id.test.mjs && node test/task-view.test.mjs && node test/root-guards.test.mjs && node test/store-gitignore.test.mjs && node test/store-cache.test.mjs && node test/enhancer.test.mjs",
|
|
22
23
|
"eval:injection": "node test/injection-scenarios.test.mjs",
|
|
23
24
|
"typecheck": "tsc -p tsconfig.json && tsc -p tsconfig.client.json",
|
|
24
25
|
"selfcheck:triggers": "node test/injection-scenarios.test.mjs --selfcheck",
|