@yolk_vat-y/dsh-project-memory 0.3.2 → 0.3.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,21 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.3.3 (2026-08-31)
4
+
5
+ ### 符号层:只存身份牌,不存行为
6
+ - `src/symbols.js` `buildSymbol`:删除 `summary` 废话字段;`text` 不再截断,保存完整声明行(含签名);返回一行身份牌 `fn(a: A, b: B): R — file.ts:42`;删除冗余 `sig` 字段
7
+
8
+ ### 文档层:答案级摘要 + 自报盲区
9
+ - `src/llm.js` `extractDocEntry`:新增 `blindSpots` 字段(自报盲区,如 `// 未覆盖:部署细节、性能基准、v0.2 前 API`);Prompt 要求 LLM 返回 `blindSpots`;摘要通过 `summarizeText()` 截断至 300 字符;`blindSpots` 追加在摘要末尾 `// 未覆盖:...`
10
+ - `src/doc-pipeline.js` `buildDocEntries`:新增 `hash` 字段(SHA256 内容哈希,用于更新检测);新增 `blindSpots` 字段存入分片
11
+
12
+ ### L1 增强正则:泛型、参数/返回类型、重载、接口/类型别名
13
+ - `src/symbols.js` 新增 `extractTypeSignature` / `extractInterfaceOrType` / `extractOverloads`:提取泛型参数、参数类型注解、返回类型注解、重载签名、接口成员、类型别名右侧、变量/常量类型注解
14
+ - 产出直接融入 `buildSymbol` 的 `text` 字段,零依赖、~0.5ms/文件
15
+
16
+ ### 文档检索侧:blindSpots 感知召回
17
+ - `src/tools/query-memory.js`:召回文档条目时,若查询词命中 `blindSpots`,追加警告行提示模型去读原文
18
+
3
19
  ## 0.3.2 (2026-08-31)
4
20
 
5
21
  ### TS Compiler API 增强器 (Phase 2 L2/L3)
@@ -13,15 +29,6 @@
13
29
  - **文档同步**:README.md / README.zh-CN.md 新增功能介绍与配置表
14
30
  - **依赖升级**:cordis 4.0.2、schemastery 3.18.2、dsh-tools/llm 0.1.2-alpha.2
15
31
 
16
- ## Unreleased
17
-
18
- ### 符号层:只存身份牌,不存行为
19
- - `src/symbols.js` `buildSymbol`:删除 `summary` 废话字段;`text` 不再截断,保存完整声明行(含签名);返回一行身份牌 `fn(a: A, b: B): R — file.ts:42`;删除冗余 `sig` 字段
20
-
21
- ### 文档层:答案级摘要 + 自报盲区
22
- - `src/llm.js` `extractDocEntry`:新增 `blindSpots` 字段(自报盲区,如 `// 未覆盖:部署细节、性能基准、v0.2 前 API`);Prompt 要求 LLM 返回 `blindSpots`;摘要通过 `summarizeText()` 截断至 300 字符;`blindSpots` 追加在摘要末尾 `// 未覆盖:...`
23
- - `src/doc-pipeline.js` `buildDocEntries`:新增 `hash` 字段(SHA256 内容哈希,用于更新检测);新增 `blindSpots` 字段存入分片
24
-
25
32
  ## 0.3.1 (2026-08-30)
26
33
 
27
34
  ### 存储:相对路径存储
package/README.md CHANGED
@@ -37,7 +37,7 @@ The store is per-project and follows the codebase: changed files are re-extracte
37
37
 
38
38
  ## Installation
39
39
 
40
- Tested against dsh **0.1.0-rc.7 through 0.1.2-alpha.1**. The plugin relies exclusively on stable public APIs (`defineTool`, `llm.stream`, `Schema`) declared via peerDependencies, ensuring compatibility with future rc releases without changes.
40
+ Tested against dsh **0.1.0-rc.7 through 0.1.2-alpha.2**. The plugin relies exclusively on stable public APIs (`defineTool`, `llm.stream`, `Schema`) declared via peerDependencies, ensuring compatibility with future rc releases without changes.
41
41
 
42
42
  ```bash
43
43
  cd dsh-project-memory && dsh plugin --profile web add . -w
@@ -95,12 +95,85 @@ Stores created before v0.2.0 (single `entries.json` / `index.json`) migrate auto
95
95
 
96
96
  These are deliberate scope choices.
97
97
 
98
- - **Lock-free sync transactions** — all writes (index / watch / remember / forget / watch_repo) go through synchronous transactions `store.commit(fn)`; fn succeeds then atomic write; JS single-threaded event loop guarantees no interleaving; `remember`/`forget` never blocked by watch re-indexing. Multi-instance concurrent writes to the same project store are naturally idempotent via CAS + atomic commits. We don't fear multi-process DSH because DSH is built on Cordis, and Cordis's single-process architecture is foundational — changing it would be a breaking change for the entire ecosystem.
99
- - **Watch compute and commit separated** — watcher does heavy work (mtime/hash/symbol scan/LLM summary) outside transaction, then single `commit` applies atomically; `applyFileUpdate` uses CAS verification (unified null handling, delete skips hash compare) to prevent concurrent modification; failures rollback snapshot for next retry; indexing failures delete snapshot for auto-retry. Polling (mtime + content hash) instead of `fs.watch` events keeps behavior consistent across platforms.
100
- - **Corrupt files are quarantined** — a store JSON that fails to parse falls back to empty for that file and is rebuilt on the next write; the broken file is renamed to `*.corrupt` with an error logged, but its data cannot be recovered. Auto-repairing partial writes would need a write-ahead journal or an embedded database — out of proportion when quarantining one bad file costs nothing.
101
- - **`forget` by query is eager** — keyword deletion matches at ≥0.5 token overlap and may remove several notes at once; prefer deleting by id for precision.
102
- - **Cross-language recall depends on index time** — with `llmQueryExpansion` off, a Chinese-only query reaches English content through bilingual keywords captured when docs are indexed, plus doc↔symbol links; queries stay LLM-free. Stores indexed before v0.1.1 gain bilingual keywords as files change, or immediately via `index_repo` with `reindex: true`.
103
- - **CJK retrieval** — phrase boost and synonym expansion are purely query-side; they do not increase index size or LLM usage. Link boundaries use CJK-aware regex only; English symbols keep the original word-boundary behavior. The experience supersede threshold (0.7 bidirectional) is a conservative default; adjust via config if false positives/negatives appear in practice.
98
+ ### 1. Synchronous lock-free transactions over async locks
99
+
100
+ **We do:** All writes go through `store.commit(fn)` — a synchronous in-process transaction. The callback `fn` performs all validation and mutations; only on success is the result atomically written to disk. The JS event loop guarantees no interleaving. CAS (`applyFileUpdate`) makes concurrent writes idempotent.
101
+
102
+ **We don't:** Async mutexes, file locks, or multi-process coordination.
103
+
104
+ **Why:** DSH runs on Cordis, which is single-process by design. Adding locks would complicate the hot path (every `remember`/`forget`/`index_doc` call) for a scenario (multi-process DSH) that would require a breaking ecosystem change. Synchronous transactions keep the hot path at ~2 ms median with zero contention overhead in practice.
105
+
106
+ ### 2. Watch: compute outside, commit inside
107
+
108
+ **We do:** Heavy work (mtime/hash/scan/LLM summary) runs outside the transaction; a single `commit` applies all changes atomically. On failure, the snapshot rolls back so the next poll retries automatically.
109
+
110
+ **We don't:** Hold a lock during LLM calls, or use `fs.watch` events.
111
+
112
+ **Why:** LLM summarization takes seconds — holding a lock would block `remember`/`forget`/`query_memory`. Polling with mtime+content-hash is platform-agnostic (works on network drives, Docker volumes, WSL) and avoids the "double fire / missed events" nightmare of `fs.watch`.
113
+
114
+ ### 3. Corrupt files are quarantined, not auto-repaired
115
+
116
+ **We do:** On JSON parse failure, the bad file is renamed to `*.corrupt`, an error is logged, and that file's store starts fresh. The rest of the store remains intact.
117
+
118
+ **We don't:** Write-ahead logs, embedded databases (SQLite/LMDB), or automatic partial recovery.
119
+
120
+ **Why:** A corrupted shard means *one source file* has a bad index — quarantining it costs near zero. A WAL or embedded DB adds a heavy dependency, increases binary size, and introduces new failure modes (lock contention, corruption of the WAL itself). The tradeoff: lose one file's index vs. add 500 KB+ of native code.
121
+
122
+ ### 4. No vector embeddings, no semantic search at query time
123
+
124
+ **We do:** BM25 with CJK phrase boost (3+ chars ×1.5 on title/keywords), synonym expansion (bidirectional table), field weighting (title ×5), and experience-layer phrase boost. All at query time, zero LLM calls.
125
+
126
+ **We don't:** Vector embeddings, dense retrieval, rerankers, or hybrid search.
127
+
128
+ **Why:** Vectors require an embedding model (local = heavy, remote = latency + cost + privacy), a vector index (HNSW/IVF = memory + build time), and reranking (another LLM call). For code + docs + experience notes, lexical BM25 with our enhancements already achieves >90% recall on real queries. The marginal gain from semantic search doesn't justify the 10x complexity/cost increase.
129
+
130
+ ### 5. Cross-language recall at index time, not query time
131
+
132
+ **We do:** Doc keywords *must* cover both the document's language AND English. Doc↔symbol links surface English symbol names from Chinese queries. With `llmQueryExpansion: false`, queries never touch the LLM.
133
+
134
+ **We don't:** Translate queries at search time, or use multilingual embeddings.
135
+
136
+ **Why:** Query-time translation adds latency, token cost, and failure modes (bad translation = zero recall). Index-time bilingual keywords are a one-time cost per document; the LLM already summarizes the doc, so extracting English keywords is free. This also works offline and deterministically.
137
+
138
+ ### 6. Explicit `remember` over implicit learning
139
+
140
+ **We do:** Users (or the agent) explicitly call `remember(problem, solution)`. Supersede uses bidirectional token overlap ≥0.7 to deduplicate.
141
+
142
+ **We don't:** Automatically extract "lessons" from user corrections, or infer rules from conversation history.
143
+
144
+ **Why:** Implicit learning is unpredictable — it hallucinates, captures noise, and pollutes the memory with unverifiable entries. Explicit `remember` creates an auditable, user-controlled knowledge base. The cost (one tool call) is negligible; the benefit (trust, verifiability, no silent corruption) is decisive.
145
+
146
+ ### 7. Full entries returned directly
147
+
148
+ **We do:** `query_memory` returns complete entries with `path:line` citations. Every hit can be verified against source.
149
+
150
+ **We don't:** Return a minimal index first, then require a second tool call for details.
151
+
152
+ **Why:** Returning full entries preserves **verifiability** — the agent sees the exact source line for every claim. It also avoids a round-trip per useful hit. Our entries are already compact (~300 chars summary + citation); the token cost is lower than a second tool call + context switch.
153
+
154
+ ### 8. Symbol extraction focused on what developers search for
155
+
156
+ **We do:** Regex-based symbol extraction (functions, classes, methods, interfaces, type aliases) with string/comment masking, multi-line signatures, and cross-file linking by symbol name. For TypeScript/JavaScript projects, an optional L2 enhancement layer uses the TS Compiler API to infer return types, resolve generics, and extract interfaces — all cached by content hash for instant reuse.
157
+
158
+ **We don't:** Tree-sitter AST parsing, import graphs, call graphs, or full-program type resolution across files.
159
+
160
+ **Why:** Our regex scanner handles 8 languages with zero dependencies, runs in <1 ms/file, and captures the declarations developers actually search for (names, signatures, generics). The optional TS layer adds semantic depth for TS/JS without native deps. Cross-file linking by name covers the most common "find related code" use case. Full-program analysis would add native binaries, 10x install size, and version fragility — for marginal gain on the remaining 5% of edge cases.
161
+
162
+ ### 9. `forget` by query is aggressive; prefer ID deletion
163
+
164
+ **We do:** `forget query` deletes all experience notes with ≥0.5 token overlap.
165
+
166
+ **We don't:** Interactive confirmation, soft-delete/trash, or exact-match-only.
167
+
168
+ **Why:** Experience notes are low-stakes, high-volume, and retrieval-only. Aggressive deletion prevents stale noise from polluting search. For precision, delete by ID (shown in `query_memory` output).
169
+
170
+ ### 11. TypeScript enhancement is optional, lazy, and cached
171
+
172
+ **We do:** L2 TS Compiler API enhancement runs async in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`), results cached by content hash in `type-cache/`. Zero config — just `npm i -D typescript`. Falls back to L1 regex if TS absent or disabled.
173
+
174
+ **We don't:** Mandatory TS, blocking enhancement, or full-program type checking.
175
+
176
+ **Why:** Mandatory TS would break installs for non-TS projects. Blocking enhancement would stall `index_repo` on large codebases. Full-program checking is 10x slower and memory-heavy. Our design: enhance what's read, cache it, never block the hot path.
104
177
 
105
178
  ## Configuration
106
179
 
@@ -155,7 +228,7 @@ These commands are for **maintaining the plugin code** — regular users do not
155
228
 
156
229
  ```bash
157
230
  npm install
158
- npm test # chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
231
+ npm test # 157 tests (v0.3.2): chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
159
232
  ```
160
233
 
161
234
  ## License
package/README.zh-CN.md CHANGED
@@ -37,7 +37,7 @@
37
37
 
38
38
  ## 安装
39
39
 
40
- 实测覆盖 dsh **0.1.0-rc.7 至 0.1.2-alpha.1**。插件仅依赖通过 peerDependencies 声明的稳定公共 API(`defineTool`、`llm.stream`、`Schema`),保证与后续 rc 版本无需改动即兼容。
40
+ 实测覆盖 dsh **0.1.0-rc.7 至 0.1.2-alpha.2**。插件仅依赖通过 peerDependencies 声明的稳定公共 API(`defineTool`、`llm.stream`、`Schema`),保证与后续 rc 版本无需改动即兼容。
41
41
 
42
42
  ```bash
43
43
  cd dsh-project-memory && dsh plugin --profile web add . -w
@@ -95,12 +95,85 @@ v0.2.0 之前创建的库(单文件 `entries.json` / `index.json`)在首次
95
95
 
96
96
  以下是刻意的范围选择。
97
97
 
98
- - **无锁同步事务** — 不采用锁:所有写入(index / watch / remember / forget / watch_repo)统一走同步事务 `store.commit(fn)`,fn 成功后才一次落盘;JS 单线程事件循环保证事务间不交错,`remember`/`forget` 不会被 watch 重索引阻塞排队。多实例并发写入同一项目存储时,得益于 CAS 幂等更新与原子提交,自然具备幂等性,无数据损坏风险。我们不担心 dsh 变多进程,因为 dsh 基于 Cordis,而 Cordis 的单进程架构是基础——改变它将是整个生态的破坏性变更。
99
- - **watch 计算与提交分离** — watcher 在事务外完成 mtime/哈希/符号扫描/LLM 摘要等重活,再以单次 `commit` 原子应用全部变更;`applyFileUpdate` 用 CAS 校验(统一 null 处理、删除跳过对比)防并发修改,失败回滚 snapshot 下轮重试,索引失败删除 snapshot 自动重试。轮询(mtime + 内容哈希)而非 `fs.watch` 事件驱动,是为了跨平台行为一致。间隔可用 `watchInterval` 调整。
100
- - **损坏隔离重建** — 存储 JSON 损坏时该文件回落为空并在下次写入时重建;坏文件会改名备份为 `*.corrupt` 并输出错误日志,但该文件内的数据无法恢复。自动修复半写文件需要预写日志或嵌入式数据库,代价与收益不成比例——而隔离一个坏文件的成本几乎为零。
101
- - **`forget` 按关键词删除偏激进** — 关键词删除按 ≥0.5 token 重叠匹配,可能一次删掉多条;追求精确请用 id 删除。
102
- - **跨语种召回依赖索引时** — `llmQueryExpansion` 关闭时,纯中文查询靠索引时捕获的双语 keywords 和 doc↔symbol 链接触达英文内容,查询侧保持零 LLM 调用。v0.1.1 之前建立的索引随文件变更逐步获得双语关键词,或用 `index_repo` 的 `reindex: true` 立即重建。
103
- - **CJK 检索** — 短语加分与同义词展开完全在查询侧,不增加索引体积、不额外消耗 LLM token。链接边界仅对 CJK 使用正则边界,英文符号保持原有词边界行为。经验层 supersede 阈值(双向 0.7)为保守默认;若实测出现误覆盖/误漏报,可通过配置调整。
98
+ ### 1. 同步无锁事务,而非异步锁
99
+
100
+ **我们做:** 所有写入走 `store.commit(fn)` 同步事务。回调 `fn` 内完成校验与变更,成功后才原子落盘。JS 事件循环天然串行,CAS (`applyFileUpdate`) 让并发写入幂等。
101
+
102
+ **不做:** 异步互斥锁、文件锁、多进程协调。
103
+
104
+ **为什么:** DSH 基于 Cordis,单进程是架构基石。为极少见的多进程场景加锁,会让热路径(每次 `remember`/`forget`/`index_doc`)增重。同步事务让热路径中位数 ~2 ms,零争用开销。
105
+
106
+ ### 2. Watch:事务外计算,事务内提交
107
+
108
+ **我们做:** 重活(mtime/哈希/扫描/LLM 摘要)在事务外跑,单次 `commit` 原子应用全部变更。失败回滚 snapshot,下轮自动重试。
109
+
110
+ **不做:** 持锁调用 LLM,或用 `fs.watch` 事件。
111
+
112
+ **为什么:** LLM 摘要耗时秒级,持锁会阻塞 `remember`/`forget`/`query_memory`。轮询 + mtime+内容哈希跨平台一致(网络盘、Docker 卷、WSL 皆可),避免 `fs.watch` 的「重复触发/漏事件」噩梦。
113
+
114
+ ### 3. 损坏文件隔离,不自动修复
115
+
116
+ **我们做:** JSON 解析失败时,坏文件改名 `*.corrupt`、记错误、该文件存储重头开始,其余分片不受影响。
117
+
118
+ **不做:** 预写日志 (WAL)、嵌入式数据库、自动部分恢复。
119
+
120
+ **为什么:** 一个损坏分片 = 一个源文件索引丢失,隔离成本近零。WAL 或嵌入式 DB 增加 500 KB+ 原生依赖、锁竞争、新故障模式(WAL 自身损坏)。权衡:丢一个文件索引 vs. 引入重型原生栈。
121
+
122
+ ### 4. 查询零向量、零语义搜索
123
+
124
+ **我们做:** BM25 + CJK 短语加分(3+ 字 ×1.5)、同义词表双向展开、字段加权(标题 ×5)、经验层短语加分。查询侧零 LLM 调用。
125
+
126
+ **不做:** 向量嵌入、稠密检索、重排序、混合搜索。
127
+
128
+ **为什么:** 向量需要嵌入模型(本地重、远程慢+贵+隐私)、向量索引(HNSW/IVF 占内存+建索引慢)、重排序(再调一次 LLM)。对代码+文档+经验笔记,增强 BM25 已达 >90% 实战召回。边际收益不抵 10x 复杂度/成本。
129
+
130
+ ### 5. 跨语种召回在索引时完成,而非查询时
131
+
132
+ **我们做:** 文档 keywords 强制双语(文档语言+英文);doc↔symbol 链接从中文命中带出英文符号名。`llmQueryExpansion: false` 时查询完全不碰 LLM。
133
+
134
+ **不做:** 查询时翻译、多语言向量。
135
+
136
+ **为什么:** 查询时翻译增延迟、耗 token、易翻车(译错=零召回)。索引时双语 keywords 是一次性成本(LLM 摘要时顺手提取),离线确定、可复用。
137
+
138
+ ### 6. 显式 `remember`,不做隐式学习
139
+
140
+ **我们做:** 用户/显式调用 `remember(problem, solution)`。supersede 用双向 token 重叠 ≥0.7 去重。
141
+
142
+ **不做:** 从用户纠正中自动抽「教训→规则」、从对话历史推断规则。
143
+
144
+ **为什么:** 隐式学习不可控——会幻觉、收噪音、污染记忆库且不可审计。显式 `remember` 成本极低(一次工具调用),换来可信、可追溯、用户可控的知识库。
145
+
146
+ ### 7. 直接返回完整条目
147
+
148
+ **我们做:** `query_memory` 直接返回含 `path:line` 引用的完整条目,每条可回源核实。
149
+
150
+ **不做:** 先返回极简索引(如 700 字符),再二次调工具取详情。
151
+
152
+ **为什么:** 完整返回保持 **可核验性**——Agent 能看到每条声明的出处行号。也避免了每次有效命中多一轮工具调用+上下文切换。条目本已紧凑(~300 字摘要+引用),完整返回的 token 成本低于二次调用。
153
+
154
+ ### 8. 符号提取聚焦开发者实际搜索的内容
155
+
156
+ **我们做:** 正则符号提取(函数/类/方法/接口/类型别名),含字符串/注释掩码、多行签名、跨文件按名链接。对 TypeScript/JavaScript 项目,可选的 L2 增强层利用 TS Compiler API 推导返回类型、实例化泛型、提取接口 —— 全部按内容哈希缓存,毫秒级复用。
157
+
158
+ **不做:** tree-sitter AST、导入图、调用图、跨文件全程序类型推导。
159
+
160
+ **为什么:** 正则扫描器零依赖、8 语言、<1 ms/文件,覆盖开发者最常搜索的声明(名字、签名、泛型)。可选 TS 增强层为 TS/JS 提供语义深度,且无原生依赖。按名跨文件链接已覆盖最常见的「找相关代码」场景。全程序分析会引入原生二进制、安装体积增 10x、语言版本即破——边际收益仅在剩余 5% 的边缘情况。
161
+
162
+ ### 9. `forget` 按关键词激进;精确请用 ID
163
+
164
+ **我们做:** `forget query` 删除所有 token 重叠 ≥0.5 的经验笔记。
165
+
166
+ **不做:** 交互确认、软删除/回收站、仅精确匹配。
167
+
168
+ **为什么:** 经验笔记低风险、高量、仅检索。激进删除防止陈旧噪音污染搜索。精确删用 ID(`query_memory` 输出里有)。
169
+
170
+ ### 11. TS 增强可选、异步、缓存
171
+
172
+ **我们做:** L2 TS Compiler API 在优先级队列异步跑(P0 `fs/observed`、P1 `watch`、P2 `index_repo`),结果按内容哈希缓存 `type-cache/`。零配置——`npm i -D typescript` 即用。无 TS 或禁用时优雅回退 L1 正则。
173
+
174
+ **不做:** 强制 TS、阻塞式增强、全程序类型检查。
175
+
176
+ **为什么:** 强制 TS 会让非 TS 项目装不上。阻塞增强会卡死大项目 `index_repo`。全程序检查慢 10x、内存重。设计:读到即增强、缓存复用、热路径永不阻塞。
104
177
 
105
178
  ## 配置
106
179
 
@@ -155,7 +228,7 @@ dsh web --patch ./config.yml
155
228
 
156
229
  ```bash
157
230
  npm install
158
- npm test # 检查项:chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
231
+ npm test # 157 tests (v0.3.2):chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
159
232
  ```
160
233
 
161
234
  ## 许可证
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@yolk_vat-y/dsh-project-memory",
3
- "version": "0.3.2",
3
+ "version": "0.3.3",
4
4
  "description": "Persistent project memory for dsh agents: index docs (PDF/Markdown/text) and code symbols into a searchable per-workspace store, recall them with cited sources, and keep experience entries (problems -> solutions) searchable on demand.",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
@@ -66,7 +66,16 @@ export function queryMemoryTool(ctx, config) {
66
66
  for (const { entry: e, score } of scored) {
67
67
  const absSource = e.sourceLine ? `${toAbs(root, e.sourcePath)}:${e.sourceLine}` : toAbs(root, e.sourcePath)
68
68
  const rel = Math.round((score / top) * 100)
69
- lines.push(`### ${e.title} (score: ${rel})\n- source: ${absSource}\n- ${e.summary}`)
69
+ let summaryLine = `- ${e.summary}`
70
+ if (e.type === 'doc' && e.blindSpots) {
71
+ const queryTokens = queries.flatMap(q => q.split(/[\s\-_]+/)).map(t => t.toLowerCase()).filter(Boolean)
72
+ const blindTokens = e.blindSpots.split(/[\s\-\u3000、,,、;;.。]+/).map(t => t.toLowerCase()).filter(Boolean)
73
+ const hit = queryTokens.some(qt => blindTokens.some(bt => bt.includes(qt) || qt.includes(bt)))
74
+ if (hit) {
75
+ summaryLine += `\n- ⚠️ 摘要未覆盖:${e.blindSpots.replace(/^\s*\/\/\s*未覆盖[::]\s*/, '')}。建议读原文 ${absSource}`
76
+ }
77
+ }
78
+ lines.push(`### ${e.title} (score: ${rel})\n- source: ${absSource}\n${summaryLine}`)
70
79
  if (e.type === 'doc' && Array.isArray(e.linkedSymbols) && e.linkedSymbols.length) {
71
80
  const refs = e.linkedSymbols.slice(0, 5).map((id) => {
72
81
  const s = symbolById.get(id)