@yolk_vat-y/dsh-project-memory 0.3.3 → 0.3.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,18 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.3.4 (2026-09-02)
4
+
5
+ ### query_memory 性能优化:流式 TF + IDF 缓存
6
+ - **IDF 缓存**:`store._version` + `store._idfCache`,`save()` 时 `version++` 标记失效,查询时版本命中直接复用,无需重建 BM25 索引
7
+ - **预计算 searchText**:`setEntries` 时预计算 `entry.searchText`(标题×5 + keywords + summary + path 的小写拼接),查询时直接复用,避免重复字符串拼接与 `toLowerCase()`
8
+ - **流式打分**:`rankEntriesStreaming` 单次遍历 entries,用 `countOccurrences()` 字符串计数替代完整 `tokenizeRaw` + TF 表构建,零中间对象分配
9
+ - **性能提升**:5k 文件 / 20k 条目场景 `query_memory` 中位数 **187 ms → 9.3 ms**(20x);1k 文件典型项目 **<1 ms**
10
+
11
+ ### 测试覆盖
12
+ - 新增 `IDF caching & streaming TF` 测试组(7 项):缓存构建、命中、版本失效、流式打分正确性、空查询、searchText 预计算
13
+ - 新增 `query_memory streaming path` 集成测试(2 项):端到端首次查询建缓存、二次查询复用缓存
14
+ - 总测试数 157 → 166 全绿
15
+
3
16
  ## 0.3.3 (2026-08-31)
4
17
 
5
18
  ### 符号层:只存身份牌,不存行为
package/README.md CHANGED
@@ -12,16 +12,45 @@ Persistent project memory for [DeepSeek Harness](https://github.com/deepseek-ai/
12
12
 
13
13
  - **Document memorization** — PDF, Markdown, and plain text files are chunked and summarized by the LLM; each entry carries a `path:line` citation back to the source.
14
14
  - **Code symbol memory** — function, class, and method names with full type signatures (generics, parameters, return types, overloads) are extracted by a dependency-free source scanner (string/comment masking, multi-line signature joining, indentation-aware Python, class-method context), without LLM token usage.
15
- - **Optional TypeScript semantic enhancement** — when `typescript` is installed in the user project (`npm i -D typescript`), the plugin automatically activates a second layer (L2) that uses the TS Compiler API to infer return types, resolve generics, extract interfaces and type aliases, and enrich arrow functions — all asynchronously in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`). Results are cached on disk keyed by file content hash for instant cold-start reuse. Zero config: just install TS and restart dsh. Fully optional; if TS is absent or disabled via `enableTypeScript: false`, the plugin falls back to L1 regex-only extraction.
15
+ - **L1 Enhanced Regex** — zero-dep regex scanner now extracts generics, parameter/return types, overloads, interfaces, and type aliases for all supported languages, producing one-line identity signatures `fn(a: A, b: B): R — file.ts:42`.
16
+ - **Optional TypeScript semantic enhancement (L2/L3)** — when `typescript` is installed in the user project (`npm i -D typescript`), the plugin automatically activates a second layer (L2) that uses the TS Compiler API to infer return types, resolve generics, extract interfaces and type aliases, and enrich arrow functions — all asynchronously in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`). Results are cached on disk keyed by file content hash (L3) for instant cold-start reuse. Zero config: just install TS and restart dsh. Fully optional; if TS is absent or disabled via `enableTypeScript: false`, the plugin falls back to L1 regex-only extraction.
16
17
  - **Automatic refresh** — a background poll (`watch_repo`) detects new or changed files by content hash and re-memorizes only those.
17
18
  - **Read-time memorization** — files are memorized the moment the model actually reads them (`fs/observed`), so the memory is a byproduct of normal work, not a separate upfront scan. Files that are never read are never indexed. The project root is detected by markers (`.git`, `package.json`, …), a README plus source directories, or the file's own directory as a last resort.
18
19
  - **Doc ↔ code cross-linking** — when a document mentions a symbol, the match is recorded as a `reference`; querying a symbol also surfaces the documents that describe it.
19
20
  - **BM25 memory recall** — ranked search over documents, symbols, and experience notes, with optional LLM query expansion to handle vocabulary mismatch. **CJK-optimized**: precise phrase boost (3+ char phrases ×1.5 score on title/keywords match), synonym table (e.g. 数据库连接池 ↔ 连接池 ↔ DB pool), and CJK-aware word boundaries for doc↔symbol linking.
21
+ - **blindSpots-aware recall** — document summaries carry a `blindSpots` field (what the summary explicitly does NOT cover). When a query hits a blind spot, `query_memory` appends a warning pointing the model to read the source file, preventing hallucination from partial summaries.
20
22
  - **Experience notes** — problems → solutions; similar problems supersede instead of duplicating, and notes are returned only when a search matches. The note store is bounded: capacity scales with project size (clamped to 100–2000), and the oldest notes are pruned when the limit is exceeded. **Supersede tightened to bidirectional 0.7 overlap** (was 0.6); **experience `problem` field now participates in CJK phrase boost** for long-tail query recall.
23
+ - **Streaming TF + IDF caching** — query path caches IDF (term inverse frequency) per store version; on cache hit, single-pass streaming scores 20k entries in ~8 ms (5k files) / ~1 ms (1k files) with zero intermediate objects; write path is O(1) version bump.
21
24
  - **Lock-free sync transactions** — all writes (index / watch / remember / forget / watch_repo) go through synchronous transactions `store.commit(fn)`; fn succeeds then atomic write; JS single-threaded event loop guarantees no interleaving; `remember`/`forget` never blocked by watch re-indexing.
22
25
  - **Minimal dependencies** — pure JavaScript; the only runtime dependency is `pdfjs-dist` (PDF text extraction), no native builds required.
23
26
  - **Negligible overhead** — pure in-process operation; cold start <100 ms (5k files), typical project query median 2–3 ms (p99 < 7 ms); bottleneck is LLM summarization and PDF parsing, not the plugin.
24
27
 
28
+ ## Performance
29
+
30
+ ### Synthetic Benchmark (isolated environment, Node 24, Linux)
31
+
32
+ | Scenario | Scale | Measured |
33
+ |----------|-------|----------|
34
+ | Full cold index | 5,000 files / 20k entries | 353 ms |
35
+ | Cold load | 5,000 files | 82 ms |
36
+ | Hot lazy re-index (single file) | 5k files | median 2.3 ms / max 4.0 ms |
37
+ | query_memory (cached) | 5k files / 20k entries | median 9.3 ms / p95 12.6 ms |
38
+ | query_memory (cached) | 1k files / 4k entries | median 1.0 ms / p95 2.0 ms |
39
+ | Full cold index | 10,000 files / 40k entries | 637 ms |
40
+ | Cold load | 10,000 files | 144 ms |
41
+ | Hot lazy re-index (single file) | 10k files | median 4.5 ms / max 10.2 ms |
42
+
43
+ > Synthetic benchmark: generated code (~8 symbols/file), Node 24, Linux native FS, SSD. Measures pure indexing overhead without LLM calls. query_memory benchmark uses IDF cache + precomputed searchText; first query after write rebuilds IDF (~150 ms), subsequent queries hit cache.
44
+
45
+ ### Real Project Storage
46
+
47
+ | Project | Files | Entries | Store Size | Per Entry |
48
+ |---------|-------|---------|------------|-----------|
49
+ | Java Spring Boot backend | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
50
+ | Vue 3 + Vite frontend | 289 | 2,141 | 1.0 MB | ~0.5 KB |
51
+
52
+ > Real projects (Java + Vue), tested on Linux file system (Node 24). Real project entries are smaller than synthetic benchmarks due to lower symbol density and shorter declarations.
53
+
25
54
  ## How it works
26
55
 
27
56
  The design follows four principles:
@@ -37,7 +66,7 @@ The store is per-project and follows the codebase: changed files are re-extracte
37
66
 
38
67
  ## Installation
39
68
 
40
- Tested against dsh **0.1.0-rc.7 through 0.1.2-alpha.2**. The plugin relies exclusively on stable public APIs (`defineTool`, `llm.stream`, `Schema`) declared via peerDependencies, ensuring compatibility with future rc releases without changes.
69
+ Tested against dsh **0.1.0-rc.7 through 0.1.2-alpha.3**. The plugin relies exclusively on stable public APIs (`defineTool`, `llm.stream`, `Schema`) declared via peerDependencies, ensuring compatibility with future rc/alpha releases without changes.
41
70
 
42
71
  ```bash
43
72
  cd dsh-project-memory && dsh plugin --profile web add . -w
@@ -228,7 +257,7 @@ These commands are for **maintaining the plugin code** — regular users do not
228
257
 
229
258
  ```bash
230
259
  npm install
231
- npm test # 157 tests (v0.3.2): chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
260
+ npm test # 157 tests (v0.3.3): chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
232
261
  ```
233
262
 
234
263
  ## License
package/README.zh-CN.md CHANGED
@@ -12,16 +12,45 @@
12
12
 
13
13
  - **文档记忆** — PDF、Markdown、纯文本按块切分并由 LLM 生成摘要,每条记忆携带 `路径:行号` 引用回源文件。
14
14
  - **代码符号记忆** — 通过零依赖的源码扫描器提取函数、类与方法名及完整类型签名(泛型、参数类型、返回类型、重载签名),包含字符串/注释掩码、多行签名续行、Python 缩进感知、类方法上下文,不使用 LLM token。
15
- - **可选 TypeScript 语义增强** — 当用户项目安装了 `typescript`(`npm i -D typescript`),插件自动激活第二层(L2),利用 TS Compiler API 推导返回类型、实例化泛型、提取接口与类型别名、丰富箭头函数签名 —— 全部在优先级队列中异步后台处理(P0:`fs/observed` 读文件瞬间、P1:`watch` 变更后、P2:`index_repo` 批量索引)。结果按文件内容哈希缓存到磁盘,冷启动毫秒级复用。零配置:装 TS 再重启 dsh 即可。完全可选;若无 TS 或设置 `enableTypeScript: false`,回退至 L1 正则提取。
15
+ - **L1 增强正则** — 零依赖正则扫描器现可提取泛型、参数/返回类型、重载、接口、类型别名,产出单行身份签名 `fn(a: A, b: B): R — file.ts:42`。
16
+ - **可选 TypeScript 语义增强 (L2/L3)** — 当用户项目安装了 `typescript`(`npm i -D typescript`),插件自动激活第二层(L2),利用 TS Compiler API 推导返回类型、实例化泛型、提取接口与类型别名、丰富箭头函数签名 —— 全部在优先级队列中异步后台处理(P0:`fs/observed` 读文件瞬间、P1:`watch` 变更后、P2:`index_repo` 批量索引)。结果按文件内容哈希缓存到磁盘(L3),冷启动毫秒级复用。零配置:装 TS 再重启 dsh 即可。完全可选;若无 TS 或设置 `enableTypeScript: false`,回退至 L1 正则提取。
16
17
  - **自动刷新** — `watch_repo` 后台轮询,按内容哈希识别新增或变更文件,仅重记这些文件。
17
18
  - **读到即记忆** — 文件在模型**实际读取的瞬间**被记忆(监听 `fs/observed`),记忆是正常工作的副产品,而非额外的一次全量扫描。从未读过的文件不会被记忆。项目根通过标记(`.git`、`package.json` 等)、README 加源码目录、或兜底到文件所在目录逐级识别。
18
19
  - **文档 ↔ 代码交叉链接** — 文档提及某符号时记录为 `reference`;查询符号时同时带出描述该符号的文档。
19
20
  - **BM25 记忆召回** — 对文档、符号与经验笔记进行排序召回,可选 LLM 查询扩展以应对表述不一致。**CJK 增强**:精确短语乘法加分(3+ 字短语在标题/关键词命中 ×1.5)、同义词表(如 数据库连接池 ↔ 连接池 ↔ DB pool)、CJK 感知的文档↔符号链接边界。
21
+ - **blindSpots 感知召回** — 文档摘要携带 `blindSpots` 字段(明确说明摘要未覆盖的内容)。查询命中盲区时,`query_memory` 追加提示引导模型去读原文,防止半截摘要误导。
20
22
  - **经验笔记** — 记录问题 → 方案;相似问题覆盖而非重复;笔记仅在检索命中时返回。笔记数量有界:容量随项目规模伸缩(钳制在 100–2000),超限时淘汰最旧的笔记。**覆盖阈值收紧为双向 0.7 重叠**(原 0.6);**经验 `problem` 字段现参与 CJK 短语加分**,提升长尾问句召回。
23
+ - **流式 TF + IDF 缓存** — 查询路径按存储版本缓存 IDF(词逆频率);命中时单次流式遍历 20k 条目仅需 ~8 ms(5k 文件) / ~1 ms(1k 文件),零中间对象;写入路径仅 O(1) 版本号递增。
21
24
  - **无锁同步事务** — 不采用锁:所有写入(index / watch / remember / forget / watch_repo)统一走同步事务 `store.commit(fn)`,fn 成功后才一次落盘;JS 单线程事件循环保证事务间不交错,`remember`/`forget` 不会被 watch 重索引阻塞排队。多实例并发写入同一项目存储时,得益于 CAS 幂等更新与原子提交,自然具备幂等性,无数据损坏风险。
22
25
  - **依赖极简** — 纯 JavaScript;唯一运行时依赖是 `pdfjs-dist`(PDF 文本提取),无需原生构建。
23
26
  - **开销可忽略** — 纯进程内操作;冷启动 <100 ms(5k 文件),典型项目查询中位数 2–3 ms(p99 < 7 ms);瓶颈在 LLM 摘要与 PDF 解析,插件本身不阻塞。
24
27
 
28
+ ## 性能
29
+
30
+ ### 合成基准测试(隔离环境,Node 24,Linux)
31
+
32
+ | 场景 | 规模 | 实测 |
33
+ |------|------|------|
34
+ | 批量冷记忆构建 | 5,000 文件 / 20k 条目 | 353 ms |
35
+ | 冷加载 | 5,000 文件 | 82 ms |
36
+ | 热路径懒记忆 | 单文件重记忆+落盘 | 中位数 2.3 ms / 最大 4.0 ms (5k) |
37
+ | query_memory (缓存命中) | 5k 文件 / 20k 条目 | 中位数 9.3 ms / p95 12.6 ms |
38
+ | query_memory (缓存命中) | 1k 文件 / 4k 条目 | 中位数 1.0 ms / p95 2.0 ms |
39
+ | 批量冷记忆构建 | 10,000 文件 / 40k 条目 | 637 ms |
40
+ | 冷加载 | 10,000 文件 | 144 ms |
41
+ | 热路径懒记忆 | 单文件重记忆+落盘 | 中位数 4.5 ms / 最大 10.2 ms (10k) |
42
+
43
+ > 合成基准:生成代码(~8 符号/文件),Node 24,Linux 文件系统,SSD。测量纯索引开销,不含 LLM 调用。query_memory 基准使用 IDF 缓存 + 预计算 searchText;写入后首次查询重建 IDF(~150 ms),后续查询命中缓存。
44
+
45
+ ### 真实项目存储体积
46
+
47
+ | 项目 | 文件数 | 条目数 | 存储体积 | 单条目 |
48
+ |------|--------|--------|----------|--------|
49
+ | Java Spring Boot 后端 | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
50
+ | Vue 3 + Vite 前端 | 289 | 2,141 | 1.0 MB | ~0.5 KB |
51
+
52
+ > 真实项目(Java + Vue),测试于 Linux 文件系统(Node 24)。真实项目单条目体积小于合成基准,因符号密度更低、声明行更短。
53
+
25
54
  ## 工作原理
26
55
 
27
56
  设计遵循四个原则:
@@ -37,7 +66,7 @@
37
66
 
38
67
  ## 安装
39
68
 
40
- 实测覆盖 dsh **0.1.0-rc.7 至 0.1.2-alpha.2**。插件仅依赖通过 peerDependencies 声明的稳定公共 API(`defineTool`、`llm.stream`、`Schema`),保证与后续 rc 版本无需改动即兼容。
69
+ 实测覆盖 dsh **0.1.0-rc.7 至 0.1.2-alpha.3**。插件仅依赖通过 peerDependencies 声明的稳定公共 API(`defineTool`、`llm.stream`、`Schema`),保证与后续 rc/alpha 版本无需改动即兼容。
41
70
 
42
71
  ```bash
43
72
  cd dsh-project-memory && dsh plugin --profile web add . -w
@@ -228,7 +257,7 @@ dsh web --patch ./config.yml
228
257
 
229
258
  ```bash
230
259
  npm install
231
- npm test # 157 tests (v0.3.2):chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
260
+ npm test # 157 tests (v0.3.3):chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
232
261
  ```
233
262
 
234
263
  ## 许可证
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@yolk_vat-y/dsh-project-memory",
3
- "version": "0.3.3",
3
+ "version": "0.3.4",
4
4
  "description": "Persistent project memory for dsh agents: index docs (PDF/Markdown/text) and code symbols into a searchable per-workspace store, recall them with cited sources, and keep experience entries (problems -> solutions) searchable on demand.",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
package/src/store.js CHANGED
@@ -1,7 +1,7 @@
1
1
  import { createHash, randomUUID } from 'node:crypto'
2
2
  import { mkdirSync, readFileSync, readdirSync, renameSync, statSync, unlinkSync, writeFileSync } from 'node:fs'
3
3
  import path from 'node:path'
4
- import { rankEntries, rankExperience, tokenize } from './util/search.js'
4
+ import { rankEntries, rankExperience, tokenize, tokenizeRaw, extractCjkPhrases, makeSearchText } from './util/search.js'
5
5
 
6
6
  const FORMAT_FILE = 'format.json'
7
7
  const INDEX_FILE = 'index.json'
@@ -64,6 +64,8 @@ export class ProjectMemoryStore {
64
64
  this._dirtyExperience = false
65
65
  this._dirtyWatch = false
66
66
  this._formatWritten = false
67
+ this._version = 0
68
+ this._idfCache = null
67
69
  }
68
70
 
69
71
  load() {
@@ -189,6 +191,40 @@ export class ProjectMemoryStore {
189
191
  writeJsonAtomic(path.join(this.dir, WATCH_FILE), this.watchlist)
190
192
  this._dirtyWatch = false
191
193
  }
194
+ this._version++
195
+ this._idfCache = null
196
+ }
197
+
198
+ getIdfCache() {
199
+ if (this._idfCache && this._idfCache.version === this._version) {
200
+ return this._idfCache.idf
201
+ }
202
+ const idf = this._rebuildIdf()
203
+ this._idfCache = { version: this._version, idf }
204
+ return idf
205
+ }
206
+
207
+ _rebuildIdf() {
208
+ const entries = this.allEntries()
209
+ const N = entries.length
210
+ if (N === 0) return {}
211
+ const df = {}
212
+ for (const entry of entries) {
213
+ const text = entry.title || ''
214
+ const keywords = (entry.keywords || []).join(' ')
215
+ const summary = entry.summary || ''
216
+ const combined = `${text} ${text} ${text} ${text} ${text} ${keywords} ${summary}`.toLowerCase()
217
+ const terms = tokenizeRaw(combined)
218
+ const seen = new Set(terms)
219
+ for (const t of seen) {
220
+ df[t] = (df[t] || 0) + 1
221
+ }
222
+ }
223
+ const idf = {}
224
+ for (const [t, df_t] of Object.entries(df)) {
225
+ idf[t] = Math.log(1 + (N - df_t + 0.5) / (df_t + 0.5))
226
+ }
227
+ return idf
192
228
  }
193
229
 
194
230
  addWatch(root) {
@@ -219,7 +255,8 @@ export class ProjectMemoryStore {
219
255
 
220
256
  setEntries(relPath, entries) {
221
257
  if (entries.length) {
222
- this.entries[relPath] = entries
258
+ const enriched = entries.map((e) => ({ ...e, searchText: makeSearchText(e) }))
259
+ this.entries[relPath] = enriched
223
260
  } else {
224
261
  delete this.entries[relPath]
225
262
  }
@@ -3,7 +3,7 @@ import path from 'node:path'
3
3
  import { memoryRootFor, resolveIndexRoot } from '../util/fs.js'
4
4
  import { ProjectMemoryStore, storeOverview } from '../store.js'
5
5
  import { expandQuery } from '../llm.js'
6
- import { rankEntriesMergedScored, rankExperienceScored } from '../util/search.js'
6
+ import { rankEntriesMergedScored, rankExperienceScored, rankEntriesStreaming } from '../util/search.js'
7
7
  import { truncate } from '../util/text.js'
8
8
 
9
9
  function toAbs(root, rel) {
@@ -56,10 +56,12 @@ export function queryMemoryTool(ctx, config) {
56
56
  if (e.type === 'symbol') symbolById.set(e.id, e)
57
57
  }
58
58
 
59
+ const idf = store.getIdfCache()
60
+
59
61
  const lines = []
60
62
  if (type === 'all' || type === 'doc' || type === 'symbol') {
61
63
  const pool = type === 'all' ? store.allEntries() : store.allEntries().filter((e) => e.type === type)
62
- const scored = rankEntriesMergedScored(pool, queries, limit)
64
+ const scored = rankEntriesStreaming(pool, queries, idf, limit)
63
65
  if (scored.length) {
64
66
  const top = scored[0].score || 1
65
67
  lines.push(`## Memory (${type === 'all' ? 'docs + symbols' : type})`)
@@ -8,7 +8,7 @@ const SYNONYMS = new Map([
8
8
  ['db pool', ['数据库连接池', '连接池']],
9
9
  ])
10
10
 
11
- function extractCjkPhrases(text) {
11
+ export function extractCjkPhrases(text) {
12
12
  const phrases = []
13
13
  let run = ''
14
14
  for (const ch of text) {
@@ -135,6 +135,10 @@ export function weightedFieldText(entry) {
135
135
  return parts.join(' ')
136
136
  }
137
137
 
138
+ export function makeSearchText(entry) {
139
+ return weightedFieldText(entry).toLowerCase()
140
+ }
141
+
138
142
  export function rankEntries(entries, query, limit = 8) {
139
143
  const bm25 = buildBm25(entries, weightedFieldText)
140
144
  const scored = bm25.score(query)
@@ -188,4 +192,59 @@ export function rankExperienceScored(items, queryOrQueries, limit = 5) {
188
192
  return [...merged.values()]
189
193
  .sort((a, b) => b.score - a.score)
190
194
  .slice(0, limit)
195
+ }
196
+
197
+ function countOccurrences(text, token) {
198
+ if (!token) return 0
199
+ let count = 0
200
+ let pos = 0
201
+ while ((pos = text.indexOf(token, pos)) !== -1) {
202
+ count++
203
+ pos += token.length
204
+ }
205
+ return count
206
+ }
207
+
208
+ const STREAM_K1 = 1.2
209
+ const STREAM_B = 0.75
210
+
211
+ export function rankEntriesStreaming(entries, queries, idf, limit = 8) {
212
+ if (!queries.length) return entries.slice(0, limit).map((entry) => ({ entry, score: 0 }))
213
+ const merged = new Map()
214
+ for (const query of queries) {
215
+ const { expanded, cjkPhrases } = expandQuery(query)
216
+ const qTokens = new Set()
217
+ for (const term of expanded) {
218
+ for (const tok of tokenizeRaw(term)) qTokens.add(tok)
219
+ }
220
+ const queryTokens = [...qTokens]
221
+ if (!queryTokens.length) continue
222
+ for (const entry of entries) {
223
+ const text = entry.searchText || weightedFieldText(entry).toLowerCase()
224
+ const len = text.length || 1
225
+ let score = 0
226
+ for (const t of queryTokens) {
227
+ const tf = countOccurrences(text, t)
228
+ if (!tf) continue
229
+ const idfVal = idf[t] || 1
230
+ score += idfVal * ((tf * (STREAM_K1 + 1)) / (tf + STREAM_K1 * (1 - STREAM_B + (STREAM_B * len) / 1000)))
231
+ }
232
+ for (const phrase of cjkPhrases) {
233
+ const lowerPhrase = phrase.toLowerCase()
234
+ if (text.includes(lowerPhrase)) {
235
+ score *= 1.5
236
+ }
237
+ }
238
+ if (score > 0) {
239
+ const id = entry.id || entry.sourcePath
240
+ const existing = merged.get(id)
241
+ if (!existing || score > existing.score) {
242
+ merged.set(id, { entry, score })
243
+ }
244
+ }
245
+ }
246
+ }
247
+ return [...merged.values()]
248
+ .sort((a, b) => b.score - a.score)
249
+ .slice(0, limit)
191
250
  }