@yolk_vat-y/dsh-project-memory 0.5.10 → 0.5.11

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,76 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.5.11 (2026-09-25)
4
+
5
+ ### 修复:doc↔symbol 链接不再物化进 entry(真实大仓库的内存与索引开销)
6
+
7
+ `linkedSymbols` 是**跨实体派生关系**(一个 doc chunk 链接到哪些符号,取决于符号表的当前状态),
8
+ 旧实现却在索引时把它算好、写进 `entry` 并落盘。实测本仓库自己的 store(11698 文件 / 70119 条目):
9
+
10
+ - 17174 个 chunk 共 **3,948,420** 个链接槽位,只对应 15,456 个符号,单 chunk 最多 2231 个;
11
+ - 唯一的消费者 `query_memory` 只读前 5 个——存了消费量的 **46 倍**;
12
+ - 加载这个 store 的堆占用 **443MB**,链接槽位是其中最大的一块;
13
+ - 每个索引提交点还要对整库做 O(chunks × symbols) 重扫,并靠 `markFile` 标脏追失效,
14
+ 否则"文档先索引、符号后到"会永久丢链接。
15
+
16
+ 现在链接在**读取期**用每 store 的符号索引解算(`src/link.js` 的 `resolveLinkedSymbols`):
17
+ 纯 latin 符号名走倒排表按整词查,CJK/混合名保留原有边界语义的正则回退;成本 O(本 chunk 词数),
18
+ 与符号表规模无关。排序改为命中次数 → 名字长度 → id(旧实现交给消费者的前 5 个是符号表插入序,
19
+ 即"任意 5 个",这是本次一并修掉的行为)。
20
+
21
+ - **内存**:同一 store 的加载堆占用 **443MB → 129–196MB**,RSS **581MB → ~300MB**;
22
+ - `linkedSymbols` 在加载时从旧 shard 剥离、写入时不再产生;磁盘上的存量 shard 会在该文件
23
+ 下次重新索引时自然压实(不主动重写 11698 个 shard);
24
+ - 删除 `linkEntries` 导出,以及 `commitFileUpdates` 的 `link` 参数(它只为"中间批次跳过、
25
+ 最后一批统一重建链接"而存在);`enhancer` / `index-doc` 里的链接重建调用一并移除;
26
+ - `query_memory` 的 `references` 输出格式不变,`test/run-test.mjs` 的链接用例改为断言
27
+ 解析器行为,并新增"文档先索引、符号后到也能解出"与"limit / 排序"回归。
28
+
29
+ ### 修复:其余派生/中间字段与两处无界缓存(体积、内存、轮询)
30
+
31
+ 第一轮去掉了链接,这一轮把剩下的放大源和常驻开销一起收掉。同一个 store(11698 文件 /
32
+ 70119 条目 / 索引源码 108.5MB)实测:**落盘 361MB → 93MB**(老 store 由下面的自动压实
33
+ 收敛;再叠加 type-cache 自愈清理后是 **73MB**),加载 **1017ms → 279ms**,堆占用
34
+ **129–196MB → 101MB**,`recallItems` p50 **53ms → 47ms**。
35
+
36
+ - **`searchText` 不再落盘**(省 22.7MB)。它是 `weightedFieldText` 的纯派生结果,改成
37
+ `allEntries()` 在内存里按需物化;写入时与 `linkedSymbols` 一起剥离(`PERSISTED_DERIVED`)。
38
+ 检索语义与输出不变;磁盘与加载解析变少,堆占用基本持平(物化改到首次查询时做)。
39
+ - **符号声明限长成一行**(`oneLineDeclaration`,≤200 字符)。TypeScript enricher 之前把
40
+ interface 的**全部成员**拼进 `typeSig`、再整体落进 `text`/`typeSig`(实测符号条目平均
41
+ 1.25KB,其中 `text` 648B),而这两个字段没有任何读取方。现在 interface 只留前 6 个成员 +
42
+ `… +N more`,`typeSig` 不再落盘。代码层 **31% → 19% 源码**,README 声称的"一行声明"
43
+ 由此第一次成立。
44
+ - **storeCache 按字节预算 + LRU**。原来只按个数(32),而单个大仓库 store 实测驻留
45
+ 130–200MB;现在同时限制条数与估算驻留量(256MB,约 2.5KB/entry),命中会把条目挪到
46
+ 队尾,热的不会先被逐出。
47
+ - **删除 type-cache(TS 增强结果缓存)**。它按内容哈希缓存增强结果,但三个增强入口
48
+ (lazy 的 `fs/observed`、watch 轮询、`index_repo`)**都只在"文件已变更并重新索引"之后**
49
+ 才触发,此时内容哈希必然是新值——这个缓存永远命中不了。实测本仓库残留 9827 个文件
50
+ (`du` 41MB,内容其实 9.2MB,约 31MB 是 4KB 块开销)。现在 `load()` 会自愈删除该目录;
51
+ 文件变更照常触发增强,进程内仍由 `enhanceQueue` 按 (relPath, 内容哈希) 去重。
52
+ - **watch 轮询空闲退避**。轮询是 O(树) 的 walkDir + 逐文件 stat(本仓库实测 58–87ms),
53
+ 原来固定一轮、不管有没有改动都在磨 I/O。现在改成递归 `setTimeout`:无变化时翻倍
54
+ 退避到最长 2 分钟,任何变化立即回到 `watchInterval`。基准间隔同时 **15s → 30s**
55
+ (仍可配置)。
56
+ - **删除死代码** `rankEntries` / `rankEntriesMerged` / `store.searchEntries`:它们每次调用
57
+ 都 `buildBm25` 全库分词(70k 条目实测 p50 1.4s),生产路径没有调用方;相关测试改用线上
58
+ 真正跑的 `rankEntriesStreaming` / `rankEntriesMergedScored`。
59
+
60
+ **存量 store 的压实是自动且有界的**:加载时把带派生字段的老分片放进待压实队列,此后任意一次
61
+ `save()`(watch 轮询、索引、写入都会触发)最多补写 `COMPACT_BATCH = 200` 个分片,直到队列
62
+ 清空。升级后不需要重新索引,也不会在首次启动时一次性重写整库;想让它立刻跑完,随便索引一次
63
+ 即可。为了让压实不产生副作用,IDF 缓存的失效键也从"任何脏写"改成 entries 的变更计数
64
+ (`_entriesVersion`)——IDF 只依赖 entries,只写经验/insight 或只做压实都不该重建它。
65
+
66
+ ### 文档
67
+
68
+ - README 的 `npm test` 断言与实测对齐:360 → **476**(核心 184 → 205、host-contract 9 → 10;
69
+ 补上此前漏记的 task-view 6 / root-guards 79 / store-gitignore 9);
70
+ - 两份 README 的"交叉链接"机制描述由"索引后挂载到条目"改为"读取期按当前符号表解算";
71
+ - `watchInterval` 文档补充空闲退避语义;"紧凑性"一节改用实测区间(符号稀疏项目 ~0.5%,
72
+ 符号密集的 TS monorepo ~19%),不再把 0.5% 当普遍值。
73
+
3
74
  ## 0.5.10 (2026-09-24)
4
75
 
5
76
  ### 修复:共享临时目录仍然是可用的记忆根(issue #5 现场复核的补充发现)
package/README.md CHANGED
@@ -91,7 +91,7 @@ The design follows four principles:
91
91
 
92
92
  - **Volatility** — context is ephemeral; it is lost when a session is compacted.
93
93
  - **Persistence** — the **memory** is stored on disk and survives compaction and new sessions.
94
- - **Compactness** — the code layer stores one declaration line per symbol, so code-heavy projects stay near **0.5% of the source** (8.8 MB of source → 49 KB of index in the example project), and **recall** replaces re-reading the full file. The document layer is heavier by design: each chunk keeps a ≤300-char injected `summary`, a bounded `terms` set covering the whole chunk for retrieval, and a precomputed `searchText`. Measured on a docs-only corpus (179 chunks / 225 KB of Markdown): `terms` ≈ **27.5%** of source and the on-disk store ≈ **166%** of source — so on doc-heavy projects budget for roughly the docs themselves, not 0.5%.
94
+ - **Compactness** — the code layer stores one bounded declaration line per symbol (≤200 chars) and the document layer keeps a ≤300-char `summary` plus a bounded `terms` set per chunk. **Derived data is never stored**: doc→symbol links and the BM25 `searchText` are computed at read time. How small the index ends up depends on symbol density, so treat "0.5%" as the sparse end of the range, not a guarantee: a Java/Vue project measured **~0.5% of source** (8.8 MB → 49 KB), while a symbol-dense TypeScript monorepo (11.7k files / 108 MB indexed) measured **~19%** for the code layer and **~106%** for the document layer. On a docs-only corpus (179 chunks / 225 KB of Markdown) `terms` ≈ **27.5%** of source and the on-disk store ≈ **166%** of source — on doc-heavy projects budget for roughly the docs themselves.
95
95
  - **Verifiability** — **recalls** carry a `path:line` citation where applicable, so the agent can confirm details against the source.
96
96
 
97
97
  Building the **memory** does not require an upfront scan: files are memorized as the model reads them, so the **memory** grows to cover exactly what has been worked with. Re-reading a file that has not changed is a no-op (content hash), so the **memory** stays fresh with minimal ongoing overhead.
@@ -117,7 +117,7 @@ The store is per-project and follows the codebase: changed files are re-extracte
117
117
  Stores created before v0.2.0 (single `entries.json` / `index.json`) migrate automatically and idempotently on first load. Within one dsh process, all tool calls share a single in-memory store per project, so hot-path indexing writes only the shard that changed.
118
118
 
119
119
  - **Incremental** — content hash per file; only changed files are re-extracted.
120
- - **Cross-linking** — after indexing, doc summaries are matched against symbol names; matches are attached to the doc entry as `references` and surfaced by `query_memory`.
120
+ - **Cross-linking** — when `query_memory` returns a doc chunk, it resolves the symbols that chunk mentions against the **current** symbol table and appends them as `references`. Links are computed at read time, so they cannot go stale and are not stored in the index (a doc indexed before its symbols still links correctly).
121
121
  - **Query expansion** — when `llmQueryExpansion` is on, `query_memory` asks `ctx.llm` to rewrite the query into several variants (synonyms, EN/CN, identifier guesses) and merges BM25 scores across variants; when off, queries never touch the LLM. Indexing itself is model-free: keywords are rule-derived (title-weighted top terms), and doc↔symbol links surface English symbol names from Chinese hits.
122
122
  - **Consistency** — the fact layer follows the codebase (hash re-extract / remove-on-delete); the experience layer is retrieval-only with supersede and `forget`. Store writes are serialized per memory directory; the lock is in-process, so avoid running multiple dsh instances against the same project store concurrently.
123
123
 
@@ -152,7 +152,7 @@ The workflow panel is collapsible, automatically adapts to dsh and theme plugin
152
152
  | `lazyIndexing` | true | index files the moment the model reads them (`fs/observed`) |
153
153
  | `autoIndexOnFirstUse` | false | full scan of the current working directory on plugin load (opt-in) |
154
154
  | `watch` | true | enable the background refresh |
155
- | `watchInterval` | 15 | poll interval (seconds) |
155
+ | `watchInterval` | 30 | base poll interval (seconds); idle polls back off up to 2 minutes and reset to this value on any change |
156
156
  | `maxScanFiles` | 20000 | hard cap on files per scan pass; a truncated pass is reported and does not remove the entries it did not reach. Set `0` to disable the cap |
157
157
  | `maxScanDepth` | 12 | hard cap on directory depth per scan pass. Set `0` to disable |
158
158
  | `allowUnsafeRoots` | false | allow **explicit** tool calls (`index_repo`/`watch_repo`/`remember` with a `root`) to target a directory on the excluded list. Automatic paths (lazy indexing, session audit, TaskBridge, `autoIndexOnFirstUse`) stay inert in these directories regardless |
@@ -211,7 +211,7 @@ Settings live in the plugin's config object. To change them, add an override ent
211
211
  autoIndexOnFirstUse: false # off: no upfront full scan (default)
212
212
  llmQueryExpansion: false # off: do not spend tokens on LLM query expansion (default)
213
213
  watch: true # on: background refresh for watched roots (default)
214
- watchInterval: 15 # poll interval in seconds
214
+ watchInterval: 30 # base poll interval; idle polls back off to at most 2 min
215
215
  maxScanFiles: 20000 # per-scan file cap (truncation is reported, never deletes)
216
216
  maxScanDepth: 12 # per-scan directory-depth cap
217
217
  enableTypeScript: true # on: L2 TS enhancement when TS is installed (default)
@@ -297,7 +297,7 @@ These commands are for **maintaining the plugin code** — regular users do not
297
297
 
298
298
  ```bash
299
299
  npm install
300
- npm test # 360 tests (184 core + 16 TaskBridge + 12 insight-store + 9 insight-actions + 8 doc-index + 7 auto-inject + 9 host-contract + 5 reflection + 4 llm-route + 2 client-hints + 8 recall + 14 readiness + 7 insight-derive + 7 readiness-eval + 6 ops + 8 injection-audit + 5 injection-budget + 6 injection-scenarios + 18 bugfix-0.5.7 + 3 client-icons + 10 client-slash + 5 workflow-command + 7 client-session-id)
300
+ npm test # 476 tests (205 core + 16 TaskBridge + 12 insight-store + 9 insight-actions + 8 doc-index + 7 auto-inject + 10 host-contract + 5 reflection + 4 llm-route + 2 client-hints + 8 recall + 14 readiness + 7 insight-derive + 7 readiness-eval + 6 ops + 8 injection-audit + 5 injection-budget + 6 injection-scenarios + 18 bugfix-0.5.7 + 3 client-icons + 10 client-slash + 5 workflow-command + 7 client-session-id + 6 task-view + 79 root-guards + 9 store-gitignore)
301
301
  npm run eval:injection # scenario P/R on the synthetic pool: 14/14 hits, 0 false positives, control group clean
302
302
  npm run eval:injection -- --store .dsh-project-memory/insights.json # replay on YOUR store; control group is a hard gate
303
303
  npm run selfcheck:triggers # which entries can still push, which declarations are dead (reads your local store)
package/README.zh-CN.md CHANGED
@@ -88,7 +88,7 @@ dsh plugin --profile web add /path/to/dsh-project-memory.tgz
88
88
 
89
89
  - **易失性** — 上下文是临时的,会话压缩即丢失。
90
90
  - **持久性** — **记忆**存于磁盘,跨压缩与会话保留。
91
- - **紧凑性** — 代码层每个符号只存一行声明,所以代码为主的项目仍约 **0.5% 源码体积**(示例项目中 8.8 MB 源码 → 49 KB 索引),**召回**替代了通读整个文件。文档层按设计更重:每个 chunk 保留 ≤300 字符的注入 `summary`、覆盖整 chunk 的 `terms`,以及预计算的 `searchText`。纯文档语料实测(179 chunk / 225 KB Markdown):`terms` ≈ 源码 **27.5%**,整库落盘 ≈ 源码 **166%**——文档占比高的项目请按「约等于文档本身大小」估,而不是 0.5%。
91
+ - **紧凑性** — 代码层每个符号只存一行声明(≤200 字符),文档层每个 chunk 保留 ≤300 字符的 `summary` 与有界的 `terms`。**派生数据一律不落盘**:doc→symbol 链接与 BM25 的 `searchText` 都在读取期计算。最终体积取决于符号密度,所以「0.5%」是区间里稀疏的那一端、不是承诺:Java/Vue 项目实测约 **0.5% 源码**(8.8 MB → 49 KB),而符号密集的 TypeScript monorepo(11.7k 文件 / 108 MB 索引)实测代码层约 **19%**、文档层约 **106%**。纯文档语料实测(179 chunk / 225 KB Markdown)`terms` ≈ 源码 **27.5%**、整库落盘 ≈ 源码 **166%**——文档占比高的项目请按「约等于文档本身大小」估。
92
92
  - **可核验性** — **召回**在适用时携带 `路径:行号` 引用,agent 可对照源文件核实。
93
93
 
94
94
  构建**记忆**无需预先全量扫描:文件在模型读取时被记忆,**记忆**恰好覆盖实际处理过的内容。未变更的文件重读是空操作(内容哈希),因此**记忆**的持续维护开销很低。
@@ -114,7 +114,7 @@ dsh plugin --profile web add /path/to/dsh-project-memory.tgz
114
114
  v0.2.0 之前创建的库(单文件 `entries.json` / `index.json`)在首次加载时自动幂等迁移。同一个 dsh 进程内,所有工具调用共享每个项目的单一内存 store 实例,热路径索引只写发生变化的那一个分片。
115
115
 
116
116
  - **增量** — 按文件内容哈希,仅重新抽取变更文件。
117
- - **交叉链接** — 索引后将文档摘要与符号名匹配,命中符号以 `references` 挂载到文档条目,由 `query_memory` 带出。
117
+ - **交叉链接** — `query_memory` 返回文档 chunk 时,按**当前**符号表解算它提到的符号,以 `references` 带出。链接在读取期解算、不落盘,因此不会过期(文档先索引、符号后到也能链上),也不占存储。
118
118
  - **查询扩展** — `llmQueryExpansion` 开启时,`query_memory` 让 `ctx.llm` 将查询改写为多个变体(同义词、中英、符号名猜测),再跨变体合并 BM25 分数;关闭时查询完全不碰 LLM。索引本身不调用模型:keywords 由规则推导(标题加权词项),doc↔symbol 链接也会从中文命中带出英文符号名。
119
119
  - **一致性** — 事实层跟随代码库(哈希重抽 / 删除即移除);经验层仅检索,配合覆盖与 `forget` 机制。每个记忆目录的写入走同步事务 `store.commit(fn)`:fn 内完成校验与变更、成功后才原子落盘,单进程内天然串行;请避免多个 dsh 实例同时写同一项目存储。
120
120
 
@@ -149,7 +149,7 @@ TaskPanel (Container)
149
149
  | `lazyIndexing` | true | 模型读取文件的瞬间即索引(`fs/observed`) |
150
150
  | `autoIndexOnFirstUse` | false | 插件加载时对当前工作目录做全量扫描(可选) |
151
151
  | `watch` | true | 启用后台刷新 |
152
- | `watchInterval` | 15 | 轮询间隔(秒) |
152
+ | `watchInterval` | 30 | 基础轮询间隔(秒);空闲时逐步退避到最长 2 分钟,一有变化立即回到该值 |
153
153
  | `maxScanFiles` | 20000 | 单次扫描的文件数硬上限;被截断时会在报告里说明,且不会删除没扫到的条目。设 `0` 取消上限 |
154
154
  | `maxScanDepth` | 12 | 单次扫描的目录深度硬上限。设 `0` 取消 |
155
155
  | `allowUnsafeRoots` | false | 允许**显式**工具调用(带 `root` 的 `index_repo`/`watch_repo`/`remember`)指向排除名单上的目录。自动路径(懒索引、会话审计、TaskBridge、`autoIndexOnFirstUse`)无论此项如何都不会越权 |
@@ -208,7 +208,7 @@ store 建在被索引的目录树里,并且**自我忽略**:它在自己目
208
208
  autoIndexOnFirstUse: false # 关闭:不做加载时的全量扫描(默认)
209
209
  llmQueryExpansion: false # 关闭:不用 LLM 扩展查询,节省 token(默认)
210
210
  watch: true # 开启:被监听根目录后台保持新鲜(默认)
211
- watchInterval: 15 # 轮询间隔(秒)
211
+ watchInterval: 30 # 基础轮询间隔;空闲时退避到最长 2 分钟
212
212
  maxScanFiles: 20000 # 单次扫描文件上限(截断会报告,且不会误删旧条目)
213
213
  maxScanDepth: 12 # 单次扫描目录深度上限
214
214
  enableTypeScript: true # 开启:装了 TS 时启用 L2 语义增强(默认)
@@ -294,7 +294,7 @@ node scripts/bench.mjs /你的/项目路径 [--json] [--samples 100] [--no-pdf]
294
294
 
295
295
  ```bash
296
296
  npm install
297
- npm test # 360 项测试(核心 184 + TaskBridge 16 + insight-store 12 + insight-actions 9 + doc-index 8 + auto-inject 7 + host-contract 9 + reflection 5 + llm-route 4 + client-hints 2 + recall 8 + readiness 14 + insight-derive 7 + readiness-eval 7 + ops 6 + injection-audit 8 + injection-budget 5 + injection-scenarios 6 + bugfix-0.5.7 18 + client-icons 3 + client-slash 10 + workflow-command 5 + client-session-id 7)
297
+ npm test # 476 项测试(核心 205 + TaskBridge 16 + insight-store 12 + insight-actions 9 + doc-index 8 + auto-inject 7 + host-contract 10 + reflection 5 + llm-route 4 + client-hints 2 + recall 8 + readiness 14 + insight-derive 7 + readiness-eval 7 + ops 6 + injection-audit 8 + injection-budget 5 + injection-scenarios 6 + bugfix-0.5.7 18 + client-icons 3 + client-slash 10 + workflow-command 5 + client-session-id 7 + task-view 6 + root-guards 79 + store-gitignore 9)
298
298
  npm run eval:injection # 合成池上的场景 P/R:命中 14/14、假阳性 0、对照组零注入
299
299
  npm run eval:injection -- --store .dsh-project-memory/insights.json # 用你自己的 store 重放;对照组是硬闸门
300
300
  npm run selfcheck:triggers # 哪些条目还推得动、哪些声明是死的(读你本地的 store)
package/cordis.patch.yml CHANGED
@@ -10,7 +10,7 @@
10
10
  lazyIndexing: true
11
11
  autoIndexOnFirstUse: false
12
12
  watch: true
13
- watchInterval: 15
13
+ watchInterval: 30
14
14
  # 危险根护栏(家目录 / 文件系统根 / 系统与包管理器前缀整体扫描会吃满内存)。
15
15
  # 自动路径永远不越权;true 只放开**显式**带 root 的工具调用。
16
16
  allowUnsafeRoots: false
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@yolk_vat-y/dsh-project-memory",
3
- "version": "0.5.10",
3
+ "version": "0.5.11",
4
4
  "description": "Persistent project memory for dsh agents: index docs (PDF/Markdown/text) and code symbols into a searchable per-workspace store, recall them with cited sources, and keep experience entries (problems -> solutions) searchable on demand.",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
@@ -5,7 +5,7 @@
5
5
  * - store.js : load / commit / addExperience(findSupersede) / removeExperience
6
6
  * - util/search.js : buildBm25 / rankEntriesStreaming / rankExperienceScored
7
7
  * - symbols.js : scanSymbols(零 token 符号抽取,无 LLM)
8
- * - link.js : linkEntries(doc↔symbol 交叉链接)
8
+ * - link.js : resolveLinkedSymbols(doc↔symbol 交叉链接,读取期解算)
9
9
  * - insight-store.js: GlobalStore read/write + saveInsight(归一化去重)
10
10
  * - similarity.js : normalizedTokenOverlap(去重阈值判定)
11
11
  *
@@ -21,7 +21,6 @@ import { performance } from 'node:perf_hooks'
21
21
  import { ProjectMemoryStore } from '../src/store.js'
22
22
  import { walkDir, readFileForIndex, relativePath, storeKey, isSupportedCode } from '../src/util/fs.js'
23
23
  import { scanSymbols } from '../src/symbols.js'
24
- import { linkEntries } from '../src/link.js'
25
24
  import { rankEntriesStreaming, rankExperienceScored, makeSearchText } from '../src/util/search.js'
26
25
  import { GlobalStore, saveInsight, normalizeInsight } from '../src/insight-store.js'
27
26
  import { normalizedTokenOverlap } from '../src/similarity.js'
@@ -108,7 +107,6 @@ async function coldIndex(root, config) {
108
107
  const report = store.commit((s) => {
109
108
  for (const u of fileUpdates) s.applyFileUpdate(u.rel, u)
110
109
  for (const rel of Object.keys(s.files)) if (!seen.has(rel)) s.removeFile(rel)
111
- linkEntries(s)
112
110
  return s.stats()
113
111
  })
114
112
  return { report, store }
@@ -146,7 +144,7 @@ async function main() {
146
144
  console.log(`\n[语料生成] ${FILES} 文件写盘耗时 ${genMs.toFixed(0)} ms`)
147
145
 
148
146
  // ---------- 1. 冷索引(完整内核管线,无 LLM)----------
149
- console.log('\n--- 1. 冷索引 index_repo 内核管线(walk + sha256 + scanSymbols + linkEntries + commit)---')
147
+ console.log('\n--- 1. 冷索引 index_repo 内核管线(walk + sha256 + scanSymbols + commit)---')
150
148
  const idxTimes = []
151
149
  let entriesCount = 0
152
150
  for (let r = 0; r < 3; r++) {
package/src/enhancer.js CHANGED
@@ -1,8 +1,8 @@
1
1
  import { createHash } from 'node:crypto'
2
- import { readFileSync, writeFileSync, mkdirSync, existsSync } from 'node:fs'
2
+ import { readFileSync } from 'node:fs'
3
3
  import { join } from 'node:path'
4
4
  import { createRequire } from 'node:module'
5
- import { linkEntries } from './link.js'
5
+ import { oneLineDeclaration } from './util/text.js'
6
6
 
7
7
  const require = createRequire(import.meta.url)
8
8
 
@@ -98,34 +98,11 @@ const PRIORITY = {
98
98
  const enhanceQueue = []
99
99
  let processing = false
100
100
 
101
- function getCacheDirForRoot(root, config) {
102
- return join(root, config.memoryDir || '.dsh-project-memory', 'type-cache')
103
- }
104
-
105
101
  function getCacheKey(content) {
106
102
  const hash = createHash('sha256').update(content).digest('hex').slice(0, 16)
107
103
  return hash
108
104
  }
109
105
 
110
- async function loadTypeCache(cacheDir, key) {
111
- const file = join(cacheDir, `${key}.json`)
112
- if (!existsSync(file)) return null
113
- try {
114
- const data = JSON.parse(readFileSync(file, 'utf8'))
115
- return data
116
- } catch {
117
- return null
118
- }
119
- }
120
-
121
- async function saveTypeCache(cacheDir, key, data) {
122
- try {
123
- if (!existsSync(cacheDir)) mkdirSync(cacheDir, { recursive: true })
124
- const file = join(cacheDir, `${key}.json`)
125
- writeFileSync(file, JSON.stringify(data))
126
- } catch {}
127
- }
128
-
129
106
  export function isTypeScriptFile(filePath) {
130
107
  const ext = filePath.slice(filePath.lastIndexOf('.')).toLowerCase()
131
108
  return ext === '.ts' || ext === '.tsx' || ext === '.js' || ext === '.jsx' ||
@@ -264,11 +241,14 @@ export function deepParseWithTS(filePath, content) {
264
241
  if (!m.name) return null
265
242
  const type = m.type ? getTypeStr(checker.getTypeAtLocation(m.type)) : 'any'
266
243
  return `${m.name.getText()}: ${type}`
267
- }).filter(Boolean).join('; ')
244
+ }).filter(Boolean)
245
+ // 成员列表必须有界:一个大 interface 的全部成员拼起来能到几 KB,而它整条只作为
246
+ // "一行声明"存在。只留前 6 个,其余记数量。
247
+ const shown = members.slice(0, 6)
268
248
  symbols.push({
269
249
  name: node.name.getText(),
270
250
  kind: 'interface',
271
- typeSig: `{ ${members} }`,
251
+ typeSig: `{ ${shown.join('; ')}${members.length > shown.length ? `; … +${members.length - shown.length} more` : ''} }`,
272
252
  line: getLine(node)
273
253
  })
274
254
  } else if (ts.isTypeAliasDeclaration(node)) {
@@ -326,18 +306,8 @@ export function enqueueEnhance(store, relPath, filePath, priority = PRIORITY.BAT
326
306
  const p = (async () => {
327
307
  try {
328
308
  const content = readFileSync(filePath, 'utf8')
329
- const cacheKey = getCacheKey(content)
330
- const cacheDir = getCacheDirForRoot(root || process.cwd(), config)
331
- const cached = await loadTypeCache(cacheDir, cacheKey)
332
- if (cached) {
333
- // Cache hit: persist via store.commit
334
- await store.commit(fn => applyEnhancedSymbols(fn, relPath, cached.symbols))
335
- return
336
- }
337
-
338
309
  const enhanced = deepParseWithTS(filePath, content)
339
310
  if (enhanced?.length) {
340
- await saveTypeCache(cacheDir, cacheKey, { symbols: enhanced })
341
311
  await store.commit(fn => applyEnhancedSymbols(fn, relPath, enhanced))
342
312
  }
343
313
  } catch (err) {
@@ -396,8 +366,10 @@ function applyEnhancedSymbols(fn, relPath, enhanced) {
396
366
  if (!enh || nameOf(e) !== enh.name) return e
397
367
  return {
398
368
  ...e,
399
- text: `${enh.name}${enh.typeSig} -- ${relPath}:${enh.line}`,
400
- typeSig: enh.typeSig,
369
+ // 一行声明(限长)。typeSig 只是构建期的中间量,不落进 entry:它没有读取方,
370
+ // 且 interface 的 typeSig 就是整个类型体,是符号条目变胖的主因。
371
+ text: oneLineDeclaration(`${enh.name}${enh.typeSig} -- ${relPath}:${enh.line}`),
372
+ typeSig: undefined,
401
373
  enhanced: true
402
374
  }
403
375
  })
@@ -420,7 +392,7 @@ function applyEnhancedSymbols(fn, relPath, enhanced) {
420
392
  type: 'symbol',
421
393
  title: `${s.name} (${s.kind})`,
422
394
  keywords: [s.name, s.kind],
423
- text: `${s.name}${s.typeSig} -- ${relPath}:${s.line}`,
395
+ text: oneLineDeclaration(`${s.name}${s.typeSig} -- ${relPath}:${s.line}`),
424
396
  enhanced: true
425
397
  })
426
398
  }
@@ -428,8 +400,7 @@ function applyEnhancedSymbols(fn, relPath, enhanced) {
428
400
  // Write to store.entries via setEntries so the shard is marked dirty and persisted
429
401
  fn.setEntries(relPath, [...mergedEntries, ...newEntries])
430
402
 
431
- // Refresh doc<->symbol links for newly added symbols
432
- if (newEntries.length) linkEntries(fn)
403
+ // doc<->symbol 链接不在这里维护:它是读取期解算的派生关系(见 src/link.js)
433
404
 
434
405
  // Also update fn.files metadata
435
406
  if (fn.files[relPath]) {
@@ -10,7 +10,6 @@ import { isSupportedCode, isSupportedDoc, readFileForIndex } from './util/fs.js'
10
10
  import { buildDocEntries } from './doc-pipeline.js'
11
11
  import { docEntriesNeedBackfill } from './doc-index.js'
12
12
  import { scanSymbols } from './symbols.js'
13
- import { linkEntries } from './link.js'
14
13
 
15
14
  /** 不索引的后缀。 */
16
15
  export const UNSUPPORTED = 'unsupported'
@@ -88,13 +87,16 @@ export function toFileUpdate(rel, plan, record) {
88
87
  }
89
88
 
90
89
  /**
91
- * 一批更新一次性落盘:写入 / 移除 → 清掉本轮未见到的旧条目 → 重建链接。
90
+ * 一批更新一次性落盘:写入 / 移除 → 清掉本轮未见到的旧条目。
92
91
  * 单事务的好处是 store 只 save 一次,watch 每轮不会反复重写。
93
92
  *
93
+ * 不再重建 doc↔symbol 链接:链接是读取期解算的派生关系(见 src/link.js),
94
+ * 所以这里也没有了那个「中间批次跳过链接、最后一批统一做」的 `link` 参数。
95
+ *
94
96
  * @returns {{stale: string[], removed: number}} stale 是 CAS 失败(并发改动)的 rel,
95
97
  * 调用方应让它们保持「未落快照」状态,下一轮重试。
96
98
  */
97
- export function commitFileUpdates(store, { updates, unseen = null, link = true }) {
99
+ export function commitFileUpdates(store, { updates, unseen = null }) {
98
100
  const stale = []
99
101
  let removed = 0
100
102
  store.commit((s) => {
@@ -113,7 +115,6 @@ export function commitFileUpdates(store, { updates, unseen = null, link = true }
113
115
  }
114
116
  }
115
117
  }
116
- if (link) linkEntries(s)
117
118
  })
118
119
  return { stale, removed }
119
120
  }
package/src/index.js CHANGED
@@ -40,7 +40,7 @@ export const Config = Schema.object({
40
40
  lazyIndexing: Schema.boolean().default(true),
41
41
  autoIndexOnFirstUse: Schema.boolean().default(false),
42
42
  watch: Schema.boolean().default(true),
43
- watchInterval: Schema.number().default(15),
43
+ watchInterval: Schema.number().default(30),
44
44
  // 危险根护栏(issue #5):家目录 / 文件系统根 / 系统目录 / 包管理器前缀(/opt/homebrew …)
45
45
  // 整体扫描会吃满内存,默认一律拒绝。**自动路径永不越权**——懒索引、会话审计、任务桥在
46
46
  // 这类目录里始终零副作用;这个开关只放开**显式**工具调用(index_repo / watch_repo /
package/src/link.js CHANGED
@@ -1,54 +1,120 @@
1
+ // doc → symbol 交叉引用:**读取期解算**,不再物化进 entry。
2
+ //
3
+ // 旧实现(≤0.5.10)在每次索引提交时把「本 chunk 命中的所有符号 id」写进
4
+ // `entry.linkedSymbols` 并落盘。实测一个真实 TypeScript 仓库(本仓库的 361MB store):
5
+ // 17174 个 chunk 共 3,948,420 个链接槽位,只对应 15,456 个符号,单 chunk 最多 2231 个;
6
+ // 而唯一的消费者 `query_memory` 只读前 5 个——存了消费量的 46 倍。
7
+ //
8
+ // 更根本的问题是它把**跨实体派生关系**固化进了源记录:doc 链接的有效性取决于符号表的
9
+ // 当前状态,而符号是独立写入的。于是必须追失效(`markFile` 标脏、"文档先索引、符号后到
10
+ // 则链接丢失"),并且每个索引提交点都要对整库做 O(chunks × symbols) 全表重扫。
11
+ //
12
+ // 现在:entry 只存事实;链接在查询时用每 store 的符号索引解算,成本 O(本 chunk 词数),
13
+ // 且天然反映当前符号表——后索引的符号也能链上,`markFile` 那套失效机制随之删除。
14
+
1
15
  const LATIN_NAME = /^[A-Za-z0-9_]+$/
2
16
  const CJK_CHAR = /[\u3400-\u9fff\uf900-\ufaff\u3040-\u309f\u30a0-\u30ff\uac00-\ud7af]/
17
+ /** 与 latin 边界后视 `[a-z0-9_$]` 同字符集:切出的词直接可查倒排表。 */
18
+ const LATIN_WORD = /[a-z0-9_$]+/g
3
19
 
4
20
  function buildMatcher(name) {
5
21
  const lower = name.toLowerCase()
6
22
  if (LATIN_NAME.test(name)) {
7
- return { lower, re: new RegExp(`(?<![a-z0-9_$])${lower}(?![a-z0-9_$])`) }
23
+ return { re: new RegExp(`(?<![a-z0-9_$])${lower}(?![a-z0-9_$])`) }
8
24
  }
9
25
  const escaped = lower.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')
10
26
  // 尾部统一挡 CJK + 字母数字下划线,防止纯 CJK 名误链混合后缀、混合名误链更长后缀
11
- return { lower, re: new RegExp(`${escaped}(?!${CJK_CHAR.source})(?![a-z0-9_$])`) }
27
+ return { re: new RegExp(`${escaped}(?!${CJK_CHAR.source})(?![a-z0-9_$])`) }
12
28
  }
13
29
 
14
- export function linkEntries(store) {
15
- const all = store.allEntries()
16
- const symbols = all.filter((e) => e.type === 'symbol')
17
- const docs = all.filter((e) => e.type === 'doc')
18
- if (!symbols.length || !docs.length) return 0
19
-
20
- const symbolByName = new Map()
21
- for (const s of symbols) {
22
- const name = s.keywords[0]
23
- if (!name || name.length < 3) continue
24
- if (!symbolByName.has(name)) symbolByName.set(name, { syms: [], ...buildMatcher(name) })
25
- symbolByName.get(name).syms.push(s)
30
+ /** 纯 latin 符号名按整词查倒排;其余(CJK / 混合 / 带 `$`)保留正则回退,语义与旧实现一致。 */
31
+ export function buildSymbolIndex(store) {
32
+ const latin = new Map() // lowerName -> symbol[]
33
+ const other = []
34
+ const otherByName = new Map()
35
+ for (const entry of store.allEntries()) {
36
+ if (!entry || entry.type !== 'symbol') continue
37
+ const name = Array.isArray(entry.keywords) ? entry.keywords[0] : undefined
38
+ if (typeof name !== 'string' || name.length < 3) continue
39
+ const key = name.toLowerCase()
40
+ if (LATIN_NAME.test(name)) {
41
+ const bucket = latin.get(key)
42
+ if (bucket) bucket.push(entry)
43
+ else latin.set(key, [entry])
44
+ continue
45
+ }
46
+ const existing = otherByName.get(key)
47
+ if (existing) {
48
+ existing.syms.push(entry)
49
+ continue
50
+ }
51
+ const { re } = buildMatcher(name)
52
+ const record = { reGlobal: new RegExp(re.source, 'g'), syms: [entry] }
53
+ otherByName.set(key, record)
54
+ other.push(record)
26
55
  }
56
+ return { latin, other }
57
+ }
27
58
 
28
- let links = 0
29
- for (const doc of docs) {
30
- const linked = new Set()
31
- // 结构词项也参与链接:terms 覆盖整个 chunk,符号在后半段被提及时同样能链上(与检索同源)
32
- const haystack = `${doc.title || ''} ${doc.summary || ''} ${doc.terms || ''} ${doc.keywords ? doc.keywords.join(' ') : ''}`.toLowerCase()
33
- for (const [, entry] of symbolByName) {
34
- const hit = entry.re ? entry.re.test(haystack) : haystack.includes(entry.lower)
35
- for (const s of entry.syms) {
36
- if (!hit) break
37
- const before = linked.size
38
- linked.add(s.id)
39
- if (linked.size > before) links++
40
- }
41
- }
42
- const before = Array.isArray(doc.linkedSymbols) ? doc.linkedSymbols.join('\u0000') : ''
43
- const next = [...linked]
44
- if (before === next.join('\u0000')) continue
45
- doc.linkedSymbols = next.length ? next : undefined
46
- // 链接是在**符号**落盘那一刻算出来的,此时 doc 的 shard 往往不是脏的;不标脏就只存在于内存,
47
- // 下次进程启动重新加载后链接全部丢失("文档先索引、符号后到"的正常顺序)。
48
- if (typeof store.markFile === 'function' && typeof store.fileRecord === 'function' && doc.sourcePath) {
49
- const record = store.fileRecord(doc.sourcePath)
50
- if (record) store.markFile(doc.sourcePath, record)
59
+ /** 每 store 一份符号索引,按 `store.entriesVersion` 失效(任何 entry 变更即重建)。 */
60
+ const indexCache = new WeakMap()
61
+
62
+ function symbolIndexOf(store) {
63
+ const version = store.entriesVersion
64
+ const cached = indexCache.get(store)
65
+ if (cached && cached.version === version) return cached.index
66
+ const index = buildSymbolIndex(store)
67
+ indexCache.set(store, { version, index })
68
+ return index
69
+ }
70
+
71
+ /**
72
+ * 解算一条 doc entry 提到的符号,按相关度取前 `limit` 个。
73
+ *
74
+ * 判据与旧 linkEntries 同源(title / summary / terms / keywords 四个字段),
75
+ * 排序为:命中次数 → 名字长度 → id(稳定序)。旧实现交给消费者的前 5 个是符号表的
76
+ * 插入序,即「任意 5 个」,这也是本次一并修掉的行为。
77
+ *
78
+ * @param {object} store 带 `allEntries()` 与 `entriesVersion` 的 ProjectMemoryStore
79
+ * @param {object} doc 待解算的 doc entry
80
+ * @param {number} limit 最多返回多少个符号 entry
81
+ * @returns {object[]} 符号 entry 列表(可能为空)
82
+ */
83
+ export function resolveLinkedSymbols(store, doc, limit = 5) {
84
+ if (!store || !doc || doc.type !== 'doc' || !(limit > 0)) return []
85
+ const haystack = `${doc.title || ''} ${doc.summary || ''} ${doc.terms || ''} ${
86
+ Array.isArray(doc.keywords) ? doc.keywords.join(' ') : ''
87
+ }`.toLowerCase()
88
+ if (!haystack.trim()) return []
89
+
90
+ const { latin, other } = symbolIndexOf(store)
91
+ const hits = new Map() // symbol entry -> 命中次数
92
+
93
+ if (latin.size) {
94
+ const counts = new Map() // token -> 在 haystack 中的出现次数
95
+ for (const token of haystack.match(LATIN_WORD) || []) counts.set(token, (counts.get(token) || 0) + 1)
96
+ for (const [token, n] of counts) {
97
+ const bucket = latin.get(token)
98
+ if (!bucket) continue
99
+ for (const symbol of bucket) hits.set(symbol, (hits.get(symbol) || 0) + n)
51
100
  }
52
101
  }
53
- return links
102
+ for (const { reGlobal, syms } of other) {
103
+ reGlobal.lastIndex = 0
104
+ const n = haystack.match(reGlobal)?.length || 0
105
+ if (!n) continue
106
+ for (const symbol of syms) hits.set(symbol, (hits.get(symbol) || 0) + n)
107
+ }
108
+ if (!hits.size) return []
109
+
110
+ const nameOf = (e) => (Array.isArray(e.keywords) && e.keywords[0]) || e.title || ''
111
+ return [...hits.entries()]
112
+ .sort((a, b) => {
113
+ if (b[1] !== a[1]) return b[1] - a[1]
114
+ const delta = nameOf(b[0]).length - nameOf(a[0]).length
115
+ if (delta !== 0) return delta
116
+ return String(a[0].id).localeCompare(String(b[0].id))
117
+ })
118
+ .slice(0, limit)
119
+ .map(([entry]) => entry)
54
120
  }
package/src/store.js CHANGED
@@ -1,7 +1,7 @@
1
1
  import { createHash, randomUUID } from 'node:crypto'
2
- import { mkdirSync, readFileSync, readdirSync, renameSync, statSync, unlinkSync, writeFileSync } from 'node:fs'
2
+ import { mkdirSync, readFileSync, readdirSync, renameSync, rmSync, statSync, unlinkSync, writeFileSync } from 'node:fs'
3
3
  import path from 'node:path'
4
- import { rankEntries, rankExperience, tokenize, tokenizeRaw, extractCjkPhrases, makeSearchText } from './util/search.js'
4
+ import { rankExperience, tokenize, tokenizeRaw, extractCjkPhrases, makeSearchText } from './util/search.js'
5
5
  import { backfillDerivedTriggers } from './readiness.js'
6
6
 
7
7
  const FORMAT_FILE = 'format.json'
@@ -14,9 +14,38 @@ const WATCH_FILE = 'watch.json'
14
14
  const INSIGHTS_FILE = 'insights.json'
15
15
  const SHARDS_DIR = 'shards'
16
16
  const GITIGNORE_FILE = '.gitignore'
17
+ /** ≤0.5.10 的 TS 增强结果缓存目录;0.5.11 起废弃并自愈清理。 */
18
+ const LEGACY_TYPE_CACHE_DIR = 'type-cache'
17
19
 
18
20
  const storeCache = new Map()
21
+ /** 条目数上限(兜底)。 */
19
22
  const STORE_CACHE_MAX = 32
23
+ /**
24
+ * 驻留内存的估算上限。只数"几个 store"是不够的:单个大仓库 store 加载后堆占用实测
25
+ * 130–200MB,32 个就是几 GB。按 store 的条目数与读入文本量估算,约 2.5KB/entry。
26
+ */
27
+ const STORE_CACHE_MAX_BYTES = 256 * 1024 * 1024
28
+ /** 每轮 save 最多补写多少个待压实的老分片(有界,避免升级后一次重写整库)。 */
29
+ const COMPACT_BATCH = 200
30
+
31
+ /** 单个 store 的驻留估算(驻留文本量 vs 条目数换算,取大者)。 */
32
+ function estimateResidentBytes(store) {
33
+ let entries = 0
34
+ for (const list of Object.values(store.entries)) entries += list.length
35
+ return Math.max(store._residentChars || 0, entries * 2500)
36
+ }
37
+
38
+ /** 按 LRU 逐出,直到同时满足条数与字节预算;刚加入的 keepKey 即使超预算也保留。 */
39
+ function evictStoreCache(keepKey) {
40
+ let total = 0
41
+ for (const store of storeCache.values()) total += estimateResidentBytes(store)
42
+ while (storeCache.size > STORE_CACHE_MAX || total > STORE_CACHE_MAX_BYTES) {
43
+ const oldest = storeCache.keys().next().value
44
+ if (oldest === undefined || oldest === keepKey) break
45
+ total -= estimateResidentBytes(storeCache.get(oldest))
46
+ storeCache.delete(oldest)
47
+ }
48
+ }
20
49
 
21
50
  /** 已就"无法迁移的旧 store"告警过的目录:避免每次 load() 都刷一行。 */
22
51
  const migrationWarned = new Set()
@@ -26,13 +55,37 @@ function isRecord(value) {
26
55
  return Boolean(value) && typeof value === 'object' && !Array.isArray(value)
27
56
  }
28
57
 
29
- function loadJson(filePath, fallback) {
58
+ /**
59
+ * 落盘的 entry 只保留**不可推导的事实**。两个派生字段永远不写盘:
60
+ * - `linkedSymbols`:跨实体派生(取决于符号表当前状态),读取期由 src/link.js 解算;
61
+ * - `searchText`:自派生(只依赖本 entry),由 `allEntries()` 在内存里物化。
62
+ * 旧 store 里的这两个字段在加载时剥掉,于是下一次写盘自然压实。
63
+ * 实测两者在一个真实大仓库 store 里合计约 245MB(链接 222.8MB + searchText 22.7MB)。
64
+ */
65
+ const PERSISTED_DERIVED = ['linkedSymbols', 'searchText', 'typeSig']
66
+
67
+ /** 原地剥离派生字段(加载路径用,省掉一次分配)。 */
68
+ function stripPersistedDerived(entry) {
69
+ if (!entry || typeof entry !== 'object') return entry
70
+ for (const key of PERSISTED_DERIVED) if (key in entry) delete entry[key]
71
+ return entry
72
+ }
73
+
74
+ /** 返回不含派生字段的副本(写盘路径用,不能动内存里的对象)。 */
75
+ function withoutPersistedDerived(entry) {
76
+ const out = { ...entry }
77
+ for (const key of PERSISTED_DERIVED) delete out[key]
78
+ return out
79
+ }
80
+
81
+ function loadJson(filePath, fallback, sizeSink) {
30
82
  let raw
31
83
  try {
32
84
  raw = readFileSync(filePath, 'utf8')
33
85
  } catch {
34
86
  return fallback
35
87
  }
88
+ if (sizeSink) sizeSink.bytes += raw.length
36
89
  try {
37
90
  return JSON.parse(raw)
38
91
  } catch {
@@ -84,8 +137,13 @@ export class ProjectMemoryStore {
84
137
  this._dirtyBinding = false
85
138
  this._dirtyWatch = false
86
139
  this._formatWritten = false
87
- this._version = 0
88
140
  this._idfCache = null
141
+ /** entries 的变更计数:IDF 缓存与符号索引(读取期链接)都据此失效。 */
142
+ this._entriesVersion = 0
143
+ /** 待压实的老分片(≤0.5.10 落盘时带 linkedSymbols/searchText);save() 每轮有界补写。 */
144
+ this._compactQueue = new Set()
145
+ /** 从磁盘读入的 JSON 文本量(UTF-16 字符数),作为驻留内存的估算基数。 */
146
+ this._residentChars = 0
89
147
  }
90
148
 
91
149
  load() {
@@ -93,20 +151,21 @@ export class ProjectMemoryStore {
93
151
  const hot = storeCache.get(key)
94
152
  // 缓存命中即返回:`hot === this` 时再读一遍盘会静默丢弃本实例尚未 save() 的变更
95
153
  // (_loadSharded/_loadInsights 会重新赋值 experience/tasks/insights…)。
96
- if (hot) return hot
154
+ if (hot) {
155
+ // LRU:命中挪到队尾,避免"热的先被逐出、冷的常驻"。
156
+ storeCache.delete(key)
157
+ storeCache.set(key, hot)
158
+ return hot
159
+ }
97
160
  this._migrateLegacyIfNeeded()
98
161
  this._loadSharded()
99
162
  this._loadInsights()
163
+ this._removeLegacyTypeCache()
100
164
  // 有些路径只读不写(审计 jsonl 直接写在 store 目录里),所以这里也补一次自我忽略,
101
165
  // 让老版本建出来的 store 在第一次 load 就补上。
102
166
  this.ensureSelfIgnore()
103
167
  storeCache.set(key, this)
104
- // 只保留最近打开的项目:长期跨多项目运行时不至于无限增长(被逐出只是下次重新读盘)
105
- while (storeCache.size > STORE_CACHE_MAX) {
106
- const oldest = storeCache.keys().next().value
107
- if (oldest === key) break
108
- storeCache.delete(oldest)
109
- }
168
+ evictStoreCache(key)
110
169
  return this
111
170
  }
112
171
 
@@ -164,7 +223,7 @@ export class ProjectMemoryStore {
164
223
  }
165
224
  mkdirSync(path.join(this.dir, SHARDS_DIR), { recursive: true })
166
225
  for (const rel of Object.keys(files)) {
167
- writeJsonAtomic(shardRelPath(this.dir, rel), { relPath: rel, record: files[rel], entries: entries[rel] || [] })
226
+ writeJsonAtomic(shardRelPath(this.dir, rel), { relPath: rel, record: files[rel], entries: (entries[rel] || []).map(withoutPersistedDerived) })
168
227
  }
169
228
  writeJsonAtomic(formatPath, { version: 2, layout: 'sharded' })
170
229
  for (const stale of [legacyEntriesPath, path.join(this.dir, INDEX_FILE)]) {
@@ -184,21 +243,45 @@ export class ProjectMemoryStore {
184
243
  } catch {
185
244
  shardNames = []
186
245
  }
246
+ const sizeSink = { bytes: 0 }
187
247
  for (const name of shardNames) {
188
- const shard = loadJson(path.join(this.dir, SHARDS_DIR, name), null)
248
+ const shard = loadJson(path.join(this.dir, SHARDS_DIR, name), null, sizeSink)
189
249
  if (!shard || typeof shard.relPath !== 'string' || !isRecord(shard.record)) continue
190
250
  this.files[shard.relPath] = shard.record
191
251
  // 畸形 shard(entries 被写成对象/null)不能让 allEntries() 在 `for…of` 上抛错,
192
252
  // 否则一个坏文件会拖垮整个进程的每一次读取。
193
- this.entries[shard.relPath] = Array.isArray(shard.entries) ? shard.entries.filter(isRecord) : []
194
- }
195
- this.experience = loadJson(path.join(this.dir, EXPERIENCE_FILE), [])
196
- this.tasks = loadJson(path.join(this.dir, TASKS_FILE), [])
197
- this.binding = loadJson(path.join(this.dir, BINDING_FILE), {})
198
- this.watchlist = loadJson(path.join(this.dir, WATCH_FILE), [])
253
+ const list = Array.isArray(shard.entries) ? shard.entries.filter(isRecord) : []
254
+ // 老分片带着派生字段:内存里立刻剥掉,并排队等 save() 把磁盘上也压实。
255
+ if (list.some((e) => PERSISTED_DERIVED.some((k) => k in e))) this._compactQueue.add(shard.relPath)
256
+ this.entries[shard.relPath] = list.map(stripPersistedDerived)
257
+ }
258
+ this._entriesVersion++
259
+ this.experience = loadJson(path.join(this.dir, EXPERIENCE_FILE), [], sizeSink)
260
+ this.tasks = loadJson(path.join(this.dir, TASKS_FILE), [], sizeSink)
261
+ this.binding = loadJson(path.join(this.dir, BINDING_FILE), {}, sizeSink)
262
+ this.watchlist = loadJson(path.join(this.dir, WATCH_FILE), [], sizeSink)
263
+ this._residentChars = sizeSink.bytes
199
264
  this._formatWritten = existsSafe(path.join(this.dir, FORMAT_FILE))
200
265
  }
201
266
 
267
+ /**
268
+ * 清掉 ≤0.5.10 留下的 `type-cache/` 目录。
269
+ *
270
+ * 它按内容哈希缓存 TS 增强结果,但三个增强入口(lazy 的 `fs/observed`、watch 轮询、
271
+ * `index_repo`)**都只在"文件已变更并重新索引"之后**才触发,此时内容哈希必然是新值——
272
+ * 这个缓存永远命中不了。实测本仓库残留 9827 个文件(`du` 41MB,内容其实 9.2MB,约 31MB
273
+ * 是 4KB 块开销)。这里做一次自愈清理;删的是纯缓存,不丢任何事实。
274
+ */
275
+ _removeLegacyTypeCache() {
276
+ const dir = path.join(this.dir, LEGACY_TYPE_CACHE_DIR)
277
+ if (!existsSafe(dir)) return
278
+ try {
279
+ rmSync(dir, { recursive: true, force: true })
280
+ } catch {
281
+ // 权限/占用导致删不掉也不影响 store 本身
282
+ }
283
+ }
284
+
202
285
  // ---- v0.5 insights:project 级 insight 文档(任务级在 task.insights[]) ----
203
286
  // 迁移语义(旧版 store 的兼容策略):
204
287
  // v0.4 experience.json 仍由 remember/forget/query_memory 服务,不删除;
@@ -314,9 +397,11 @@ export class ProjectMemoryStore {
314
397
 
315
398
  save() {
316
399
  // 没有脏数据就不落盘。watch 每轮对每个根都无条件 commit → save;照旧执行的话,
317
- // 末尾的 `_version++` + `_idfCache = null` 会打在跨实例共享的 store 上,
318
- // 等于每 15 秒清空一次 IDF 缓存,废掉查询侧的 IDF 复用(v0.3.4 的 20x)。
400
+ // 末尾的 ID 缓存失效会打在跨实例共享的 store 上,等于每轮清空一次 IDF 复用(v0.3.4 的 20x)。
401
+ // 待压实队列不算"脏数据",但它需要有界推进,所以也走这条落盘路径。
402
+ const compacting = this._compactQueue.size > 0
319
403
  const dirty =
404
+ compacting ||
320
405
  this._dirtyShards.size > 0 ||
321
406
  this._removedShards.size > 0 ||
322
407
  this._dirtyExperience ||
@@ -335,10 +420,20 @@ export class ProjectMemoryStore {
335
420
  writeJsonAtomic(path.join(this.dir, FORMAT_FILE), { version: 2, layout: 'sharded' })
336
421
  this._formatWritten = true
337
422
  }
423
+ // 存量压实:≤0.5.10 的分片带着派生字段,而那些字段只在分片被重写时才会从磁盘消失。
424
+ // 每轮最多补写 COMPACT_BATCH 个,让升级后的 store 在后续任意一次 save(watch 轮询、
425
+ // 索引、写入)里自动收敛,而不是永远停在旧体积。
426
+ let compactBudget = COMPACT_BATCH
427
+ for (const rel of this._compactQueue) {
428
+ if (compactBudget <= 0) break
429
+ compactBudget--
430
+ this._compactQueue.delete(rel)
431
+ if (this.files[rel]) this._dirtyShards.add(rel)
432
+ }
338
433
  for (const rel of this._dirtyShards) {
339
434
  if (this.files[rel]) {
340
435
  mkdirSync(path.join(this.dir, SHARDS_DIR), { recursive: true })
341
- writeJsonAtomic(shardRelPath(this.dir, rel), { relPath: rel, record: this.files[rel], entries: this.entries[rel] || [] })
436
+ writeJsonAtomic(shardRelPath(this.dir, rel), { relPath: rel, record: this.files[rel], entries: (this.entries[rel] || []).map(withoutPersistedDerived) })
342
437
  } else {
343
438
  this._removedShards.add(rel)
344
439
  }
@@ -374,16 +469,16 @@ export class ProjectMemoryStore {
374
469
  writeJsonAtomic(path.join(this.dir, WATCH_FILE), this.watchlist)
375
470
  this._dirtyWatch = false
376
471
  }
377
- this._version++
378
- this._idfCache = null
379
472
  }
380
473
 
381
474
  getIdfCache() {
382
- if (this._idfCache && this._idfCache.version === this._version) {
475
+ // IDF 只依赖 entries(title/keywords/summary),所以按 entries 的变更计数失效:
476
+ // 只写经验/insight 或只做存量压实的 save 不会再无谓地重建 IDF。
477
+ if (this._idfCache && this._idfCache.version === this._entriesVersion) {
383
478
  return this._idfCache.idf
384
479
  }
385
480
  const idf = this._rebuildIdf()
386
- this._idfCache = { version: this._version, idf }
481
+ this._idfCache = { version: this._entriesVersion, idf }
387
482
  return idf
388
483
  }
389
484
 
@@ -438,11 +533,17 @@ export class ProjectMemoryStore {
438
533
 
439
534
  setEntries(relPath, entries) {
440
535
  if (entries.length) {
441
- const enriched = entries.map((e) => ({ ...e, searchText: makeSearchText(e) }))
536
+ // searchText 在内存里物化(写入/检索都靠它),但 save() 不会把它写盘。
537
+ const enriched = entries.map((e) => {
538
+ const out = stripPersistedDerived({ ...e })
539
+ out.searchText = makeSearchText(out)
540
+ return out
541
+ })
442
542
  this.entries[relPath] = enriched
443
543
  } else {
444
544
  delete this.entries[relPath]
445
545
  }
546
+ this._entriesVersion++
446
547
  this._dirtyShards.add(relPath)
447
548
  this._removedShards.delete(relPath)
448
549
  }
@@ -451,23 +552,34 @@ export class ProjectMemoryStore {
451
552
  if (relPath in this.files) {
452
553
  delete this.files[relPath]
453
554
  delete this.entries[relPath]
555
+ this._entriesVersion++
454
556
  this._dirtyShards.add(relPath)
455
557
  this._removedShards.add(relPath)
456
558
  }
457
559
  }
458
560
 
561
+ /** entries 的变更计数:派生缓存(符号索引)据此失效。 */
562
+ get entriesVersion() {
563
+ return this._entriesVersion
564
+ }
565
+
566
+ /** 还有多少个老分片等着被 save() 压实(0 = 存量已收敛)。 */
567
+ get pendingCompaction() {
568
+ return this._compactQueue.size
569
+ }
570
+
459
571
  allEntries() {
460
572
  const out = []
461
573
  for (const list of Object.values(this.entries)) {
462
- for (const entry of list) out.push(entry)
574
+ for (const entry of list) {
575
+ // searchText 不落盘,首次用到时按需物化并挂在对象上(同一 entry 只算一次)。
576
+ if (entry.searchText === undefined) entry.searchText = makeSearchText(entry)
577
+ out.push(entry)
578
+ }
463
579
  }
464
580
  return out
465
581
  }
466
582
 
467
- searchEntries(query, limit = 8) {
468
- return rankEntries(this.allEntries(), query, limit)
469
- }
470
-
471
583
  addExperience({ problem, solution, sourceFile }) {
472
584
  const existing = this.findSupersede(problem)
473
585
  const now = new Date().toISOString()
package/src/symbols.js CHANGED
@@ -1,4 +1,5 @@
1
1
  import { readFileSync } from 'node:fs'
2
+ import { oneLineDeclaration } from './util/text.js'
2
3
 
3
4
  const JS_LIKE = new Set(['.js', '.mjs', '.cjs', '.ts', '.mts', '.cts', '.jsx', '.tsx'])
4
5
  const PYTHON = new Set(['.py'])
@@ -465,7 +466,8 @@ function buildSymbol(matched, relPath, rawLine, lineNo) {
465
466
  type: 'symbol',
466
467
  title: `${matched.name} (${matched.kind})`,
467
468
  keywords: [matched.name, matched.kind],
468
- text: identity,
469
+ // interface/type 的 typeSig 可能是整个类型体;这里限长成一行,避免符号条目被撑大。
470
+ text: oneLineDeclaration(identity),
469
471
  }
470
472
  }
471
473
 
@@ -3,7 +3,6 @@ import path from 'node:path'
3
3
  import { assertIndexRoot, assertReadableFile, assertSafeRoot, findProjectRoot, memoryRootFor, sessionMemoryRootOrNull, sha256OfFile, storeKey } from '../util/fs.js'
4
4
  import { buildDocEntries } from '../doc-pipeline.js'
5
5
  import { docEntriesNeedBackfill } from '../doc-index.js'
6
- import { linkEntries } from '../link.js'
7
6
  import { ProjectMemoryStore } from '../store.js'
8
7
 
9
8
  export function indexDocTool(ctx, config) {
@@ -76,7 +75,6 @@ export function indexDocTool(ctx, config) {
76
75
  return store.commit((s) => {
77
76
  s.setEntries(rel, entries)
78
77
  s.markFile(rel, { sha256: hash, size, type: 'doc', indexedAt: new Date().toISOString() })
79
- linkEntries(s)
80
78
  const preview = entries
81
79
  .map((e) => ` - ${e.title} @ ${rel}:${e.sourceLine}`)
82
80
  .join('\n')
@@ -33,8 +33,8 @@ export async function indexRepository(ctx, config, root, { reindex = false, allo
33
33
 
34
34
  const flushBatch = () => {
35
35
  if (!batch.length) return
36
- // 中间批次不做 unseen 清理、不重建链接(最后一批统一做)——否则每批都要遍历整个 store。
37
- commitFileUpdates(store, { updates: batch, link: false })
36
+ // 中间批次不做 unseen 清理(最后一批统一做)——否则每批都要遍历整个 store。
37
+ commitFileUpdates(store, { updates: batch })
38
38
  batch = []
39
39
  }
40
40
 
@@ -74,12 +74,11 @@ export async function indexRepository(ctx, config, root, { reindex = false, allo
74
74
  }
75
75
  }
76
76
 
77
- // 收尾:写完最后一批 + 清理本轮未见到的旧条目 + 重建链接。
77
+ // 收尾:写完最后一批 + 清理本轮未见到的旧条目。
78
78
  // 截断时**不做 unseen 清理**:没扫到的文件不等于被删了。
79
79
  const { removed } = commitFileUpdates(store, {
80
80
  updates: batch,
81
81
  unseen: truncated ? null : seen,
82
- link: true,
83
82
  })
84
83
  const stats = store.stats()
85
84
  let report =
@@ -6,6 +6,7 @@ import { expandQuery } from '../llm.js'
6
6
  import { resolveRoute } from '../llm-route.js'
7
7
  import { GlobalStore, cfgInsight, defaultGlobalFile, recordHit } from '../insight-store.js'
8
8
  import { recallItems } from '../recall.js'
9
+ import { resolveLinkedSymbols } from '../link.js'
9
10
  import { truncate } from '../util/text.js'
10
11
  import { stepContent, stepStatus } from '../util/task-view.js'
11
12
  function toAbs(root, rel) {
@@ -59,10 +60,6 @@ export function queryMemoryTool(ctx, config) {
59
60
  const queries = config.llmQueryExpansion
60
61
  ? await expandQuery(ctx.llm, args.query, config.expansionCount, { route: resolveRoute(exec, config) })
61
62
  : [args.query]
62
- const symbolById = new Map()
63
- for (const e of store.allEntries()) {
64
- if (e.type === 'symbol') symbolById.set(e.id, e)
65
- }
66
63
 
67
64
  // 会话绑定决定 task 级 insight 的可见性;global 级始终可见。
68
65
  const sessionId = exec?.agent?.session?.id
@@ -121,12 +118,12 @@ export function queryMemoryTool(ctx, config) {
121
118
  // 条目状态占位(可证伪状态机落地前恒为 exact)。
122
119
  // 先立字段,后续状态机到位时只改值、不改输出契约。
123
120
  lines.push(`### ${e.title} (score: ${rel})\n- source: ${absSource}\n- status: ${e.status || 'exact'}\n${summaryLine}`)
124
- if (e.type === 'doc' && Array.isArray(e.linkedSymbols) && e.linkedSymbols.length) {
125
- const refs = e.linkedSymbols.slice(0, 5).map((id) => {
126
- const s = symbolById.get(id)
127
- return s ? `${s.title} @ ${toAbs(root, s.sourcePath)}:${s.sourceLine}` : id
128
- })
129
- lines.push(`- references: ${refs.join('; ')}`)
121
+ if (e.type === 'doc') {
122
+ // 读取期解算:链接不落盘,按当前符号表排序取前 5(见 src/link.js)
123
+ const refs = resolveLinkedSymbols(store, e, 5).map(
124
+ (s) => `${s.title} @ ${toAbs(root, s.sourcePath)}:${s.sourceLine}`,
125
+ )
126
+ if (refs.length) lines.push(`- references: ${refs.join('; ')}`)
130
127
  }
131
128
  }
132
129
  }
@@ -141,15 +141,10 @@ export function makeSearchText(entry) {
141
141
  return weightedFieldText(entry).toLowerCase()
142
142
  }
143
143
 
144
- export function rankEntries(entries, query, limit = 8) {
145
- const bm25 = buildBm25(entries, weightedFieldText)
146
- const scored = bm25.score(query)
147
- return scored.slice(0, limit).map((r) => r.doc)
148
- }
149
-
150
- export function rankEntriesMerged(entries, queries, limit = 8) {
151
- return rankEntriesMergedScored(entries, queries, limit).map((r) => r.entry)
152
- }
144
+ // 说明:文档/符号的线上检索走 `rankEntriesStreaming`(recall.js),它不重建整库 BM25,
145
+ // 而是用预物化的 `searchText` + 每 store 缓存的 idf。曾经存在的 `rankEntries` /
146
+ // `rankEntriesMerged`(每次调用都 buildBm25 全库分词,70k 条目实测 p50 1.4s)已删除——
147
+ // 它们在生产路径上没有调用方,留着只会被误用。
153
148
 
154
149
  export function rankEntriesMergedScored(entries, queries, limit = 8) {
155
150
  if (!queries.length) return entries.slice(0, limit).map((entry) => ({ entry, score: 0 }))
package/src/util/text.js CHANGED
@@ -20,6 +20,22 @@ export function truncate(text, maxChars) {
20
20
 
21
21
  export const MAX_SUMMARY = 300
22
22
 
23
+ /** 符号声明行的上限。未限长的 interface 类型体会把一行撑到几 KB。 */
24
+ export const MAX_DECLARATION = 200
25
+
26
+ /**
27
+ * 符号的一行声明:压平空白并限长。实测一个真实 TS 仓库里,未限长的类型体正文让
28
+ * 符号 `text` 字段达到 5.6MB(而它当前没有任何消费者);限长后保住"一行声明"的语义,
29
+ * 又不会把代码层体积从 ~0.5% 推到 30%+。
30
+ * @param {unknown} text
31
+ * @param {number} max
32
+ * @returns {string}
33
+ */
34
+ export function oneLineDeclaration(text, max = MAX_DECLARATION) {
35
+ const flat = String(text || '').replace(/\s+/g, ' ').trim()
36
+ return flat.length > max ? `${flat.slice(0, max - 1)}…` : flat
37
+ }
38
+
23
39
  /** 注入用短摘要:压平空白,按句子边界截断到 max 字符。纯函数,无 LLM。 */
24
40
  export function summarizeText(text, max = MAX_SUMMARY) {
25
41
  const flat = String(text || '').replace(/\s+/g, ' ').trim()
package/src/watch.js CHANGED
@@ -79,30 +79,54 @@ export class WatchManager {
79
79
  return this.roots.delete(root)
80
80
  }
81
81
 
82
- start(intervalMs = 15000) {
82
+ start(intervalMs = 30000) {
83
83
  if (this.timer) return
84
- // NaN/undefined 会让 setInterval 退化成 1ms 轮询(Node 只发一条 TimeoutNaNWarning),
85
- // 足以打满事件循环让 agent 无法响应;非法值回退到 15s。
84
+ // NaN/undefined 会让定时器退化成 1ms 轮询(Node 只发一条 TimeoutNaNWarning),
85
+ // 足以打满事件循环让 agent 无法响应;非法值回退到默认 30s。
86
86
  const raw = Number(intervalMs)
87
- const ms = Number.isFinite(raw) ? Math.max(raw, 1000) : 15000
88
- this.timer = setInterval(() => this.poll(), ms)
87
+ this._baseInterval = Number.isFinite(raw) ? Math.max(raw, 1000) : 30000
88
+ // 空闲退避:连续没有变化的轮次把间隔翻倍,最长 2 分钟一轮;任何变化立刻回到 base。
89
+ // 轮询本身是 O(树) 的 walkDir + 逐文件 stat(本仓库一轮实测 58–87ms),
90
+ // 常驻 15s 一轮意味着不管有没有改动都在磨 I/O。
91
+ this._maxInterval = Math.max(this._baseInterval, Math.min(this._baseInterval * 8, 120000))
92
+ this._interval = this._baseInterval
93
+ this._stopped = false
94
+ this._schedule()
95
+ }
96
+
97
+ _schedule() {
98
+ this.timer = setTimeout(() => this._tick(), this._interval)
89
99
  if (this.timer.unref) this.timer.unref()
90
100
  }
91
101
 
102
+ async _tick() {
103
+ this.timer = null
104
+ let changed = false
105
+ try {
106
+ changed = await this.poll()
107
+ } catch {
108
+ // poll() 内部已逐根兜底;这里保证无论发生什么都一定排下一轮,不会静默停掉 watch。
109
+ }
110
+ if (this._stopped) return // 轮询期间被 stop() 了,不要再排下一轮
111
+ this._interval = changed ? this._baseInterval : Math.min(this._interval * 2, this._maxInterval)
112
+ this._schedule()
113
+ }
114
+
92
115
  stop() {
93
- if (this.timer) clearInterval(this.timer)
116
+ this._stopped = true
117
+ if (this.timer) clearTimeout(this.timer)
94
118
  this.timer = null
95
119
  }
96
120
 
97
121
  async poll() {
98
- // setInterval 不等待上一轮:大仓库/文档 LLM 摘要让一轮 >interval 时,
99
- // 轮询会叠加成并发索引,最终打满事件循环。用重入锁让慢轮询自然跳过。
100
- if (this._polling) return
122
+ // 慢轮询期间再次进入直接跳过(递归 setTimeout 下正常不会叠加,保留作防御)。
123
+ if (this._polling) return false
101
124
  this._polling = true
125
+ let changed = false
102
126
  try {
103
127
  for (const [root, state] of this.roots) {
104
128
  try {
105
- await this.pollRoot(root, state)
129
+ if (await this.pollRoot(root, state)) changed = true
106
130
  } catch (err) {
107
131
  console.error(`[dsh-project-memory] watch poll failed for ${root}: ${err.message}`)
108
132
  }
@@ -110,6 +134,7 @@ export class WatchManager {
110
134
  } finally {
111
135
  this._polling = false
112
136
  }
137
+ return changed
113
138
  }
114
139
 
115
140
  async pollRoot(root, state) {
@@ -179,14 +204,15 @@ export class WatchManager {
179
204
  // 单事务写盘;CAS 失败的条目本轮不落快照,下一轮自然重试。
180
205
  // 截断时**不传 unseen**:没扫到的文件不等于被删了,否则一份被上限截掉的树每轮都会
181
206
  // 把自己的记忆删掉一半(先删再下轮重新索引,纯粹的抖动)。
182
- const { stale } = commitFileUpdates(state.store, {
207
+ const { stale, removed } = commitFileUpdates(state.store, {
183
208
  updates,
184
209
  unseen: truncated ? null : seen,
185
- link: updates.length > 0,
186
210
  })
187
211
  const failed = new Set(stale)
212
+ let applied = 0
188
213
  for (const update of updates) {
189
214
  if (failed.has(update.rel)) continue
215
+ applied++
190
216
  state.snapshot[update.rel] = signatures.get(update.rel)
191
217
  if (update.type !== 'code') continue
192
218
  const filePath = path.join(root, update.rel)
@@ -194,5 +220,7 @@ export class WatchManager {
194
220
  onFileChanged(state.store, update.rel, filePath, this.config, root)
195
221
  }
196
222
  }
223
+ // 有变化 → 下一轮回到 base 间隔;纯空转 → 退避。
224
+ return applied > 0 || removed > 0
197
225
  }
198
226
  }