@yolk_vat-y/dsh-project-memory 0.3.1 → 0.3.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,44 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.3.3 (2026-08-31)
4
+
5
+ ### 符号层:只存身份牌,不存行为
6
+ - `src/symbols.js` `buildSymbol`:删除 `summary` 废话字段;`text` 不再截断,保存完整声明行(含签名);返回一行身份牌 `fn(a: A, b: B): R — file.ts:42`;删除冗余 `sig` 字段
7
+
8
+ ### 文档层:答案级摘要 + 自报盲区
9
+ - `src/llm.js` `extractDocEntry`:新增 `blindSpots` 字段(自报盲区,如 `// 未覆盖:部署细节、性能基准、v0.2 前 API`);Prompt 要求 LLM 返回 `blindSpots`;摘要通过 `summarizeText()` 截断至 300 字符;`blindSpots` 追加在摘要末尾 `// 未覆盖:...`
10
+ - `src/doc-pipeline.js` `buildDocEntries`:新增 `hash` 字段(SHA256 内容哈希,用于更新检测);新增 `blindSpots` 字段存入分片
11
+
12
+ ### L1 增强正则:泛型、参数/返回类型、重载、接口/类型别名
13
+ - `src/symbols.js` 新增 `extractTypeSignature` / `extractInterfaceOrType` / `extractOverloads`:提取泛型参数、参数类型注解、返回类型注解、重载签名、接口成员、类型别名右侧、变量/常量类型注解
14
+ - 产出直接融入 `buildSymbol` 的 `text` 字段,零依赖、~0.5ms/文件
15
+
16
+ ### 文档检索侧:blindSpots 感知召回
17
+ - `src/tools/query-memory.js`:召回文档条目时,若查询词命中 `blindSpots`,追加警告行提示模型去读原文
18
+
19
+ ## 0.3.2 (2026-08-31)
20
+
21
+ ### TS Compiler API 增强器 (Phase 2 L2/L3)
22
+ - **L2 语义增强层**:用户项目安装 `typescript`(`npm i -D typescript`)后,插件自动激活,利用 TS Compiler API 推导返回类型、实例化泛型、提取接口与类型别名、丰富箭头函数签名
23
+ - **L3 磁盘缓存层**:增强结果按文件内容 SHA256 哈希缓存至 `.dsh-project-memory/type-cache/<hash>.json`,冷启动毫秒级复用,跨会话持久化
24
+ - **三级优先级队列**:P0 ACTIVE(`fs/observed` 读文件瞬间)> P1 RECENT(`watch` 变更后)> P2 BATCH(`index_repo` 批量)> P3 BACKLOG(启动补全历史),`setImmediate` 每任务后让出事件循环
25
+ - **零配置、零感知**:装 TS 重启 dsh 即可;无 TS 或 `enableTypeScript: false` 时优雅回退 L1 正则;TS 6.x/7.x 检测到警告并回退 L1
26
+ - **解析策略**:单文件 `createProgram`(快 10x),`createRequire(import.meta.url)` 兼容 ESM,解析优先级:配置 `tsPath` → 项目 cwd 向上 `node_modules/typescript` → 插件自身 `node_modules`
27
+ - **新增配置**:`tsPath`(可选指定 TS 路径)、`enableTypeScript`(默认 true,设 false 彻底禁用)
28
+ - **扩展名支持**:新增 `.mts` `.cts`
29
+ - **文档同步**:README.md / README.zh-CN.md 新增功能介绍与配置表
30
+ - **依赖升级**:cordis 4.0.2、schemastery 3.18.2、dsh-tools/llm 0.1.2-alpha.2
31
+
32
+ ## 0.3.1 (2026-08-30)
33
+
34
+ ### 存储:相对路径存储
35
+ - Entry 存相对路径(如 `src/main.js`),查询时按项目根解析绝对路径;项目搬家仅在根目录变更时需重新索引,兼容旧绝对路径 entry
36
+ - 所有写入路径(`index_doc` / `index_repo` / `watch` / `lazy`)统一传相对路径给构建函数
37
+ - `scanSymbols` / `buildDocEntries` 兼容旧签名(绝对路径 entry 仍可读)
38
+
39
+ ### 设计取舍同步
40
+ - 移除「绝对路径引用」项,该限制已由相对路径存储方案解决
41
+
3
42
  ## 0.3.0 (2026-08-29)
4
43
 
5
44
  ### 存储:无锁同步事务重构
package/README.md CHANGED
@@ -4,19 +4,21 @@
4
4
 
5
5
  [![ci](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml/badge.svg)](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml) [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![npm](https://img.shields.io/npm/v/@yolk_vat-y/dsh-project-memory)](https://www.npmjs.com/package/@yolk_vat-y/dsh-project-memory) [![Listed on dsh-plugin.org](https://dsh-plugin.org/badges/listed.svg)](https://dsh-plugin.org/plugins/00080000/dsh-project-memory) [![Awesome](https://awesome-dsh-plugin.com/badge.svg)](https://awesome-dsh-plugin.com)
6
6
 
7
- Persistent project memory for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(dsh) agents. Indexes documents (PDF / Markdown / txt) and code symbols into a per-workspace store, refreshes them automatically, and recalls them with source citations — documents are cross-linked to the code symbols they reference.
7
+ Persistent project memory for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(dsh) agents. **Memorizes** documents (PDF / Markdown / txt) and code symbols into a per-workspace store, refreshes them automatically, and **recalls** them with source citations — documents are cross-linked to the code symbols they reference.
8
8
 
9
- > The plugin keeps a compact project index on disk, with every entry pointing to a concrete file and line — the agent can reorient quickly instead of re-reading the whole project.
9
+ > The plugin keeps a compact project **memory** on disk, with every entry pointing to a concrete file and line — the agent can reorient quickly instead of re-reading the whole project.
10
10
 
11
11
  ## Features
12
12
 
13
- - **Document indexing** — PDF, Markdown, and plain text files are chunked and summarized by the LLM; each entry carries a `path:line` citation back to the source.
14
- - **Code symbol table** — function, class, and method names are extracted by a dependency-free source scanner (string/comment masking, multi-line signature joining, indentation-aware Python, class-method context), without LLM token usage.
15
- - **Automatic refresh** — a background poll (`watch_repo`) detects new or changed files by content hash and re-indexes only those.
16
- - **Read-time indexing** — files are indexed the moment the model actually reads them (`fs/observed`), so the index is a byproduct of normal work, not a separate upfront scan. Files that are never read are never indexed. The project root is detected by markers (`.git`, `package.json`, …), a README plus source directories, or the file's own directory as a last resort.
13
+ - **Document memorization** — PDF, Markdown, and plain text files are chunked and summarized by the LLM; each entry carries a `path:line` citation back to the source.
14
+ - **Code symbol memory** — function, class, and method names with full type signatures (generics, parameters, return types, overloads) are extracted by a dependency-free source scanner (string/comment masking, multi-line signature joining, indentation-aware Python, class-method context), without LLM token usage.
15
+ - **Optional TypeScript semantic enhancement** — when `typescript` is installed in the user project (`npm i -D typescript`), the plugin automatically activates a second layer (L2) that uses the TS Compiler API to infer return types, resolve generics, extract interfaces and type aliases, and enrich arrow functions — all asynchronously in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`). Results are cached on disk keyed by file content hash for instant cold-start reuse. Zero config: just install TS and restart dsh. Fully optional; if TS is absent or disabled via `enableTypeScript: false`, the plugin falls back to L1 regex-only extraction.
16
+ - **Automatic refresh** — a background poll (`watch_repo`) detects new or changed files by content hash and re-memorizes only those.
17
+ - **Read-time memorization** — files are memorized the moment the model actually reads them (`fs/observed`), so the memory is a byproduct of normal work, not a separate upfront scan. Files that are never read are never indexed. The project root is detected by markers (`.git`, `package.json`, …), a README plus source directories, or the file's own directory as a last resort.
17
18
  - **Doc ↔ code cross-linking** — when a document mentions a symbol, the match is recorded as a `reference`; querying a symbol also surfaces the documents that describe it.
18
- - **BM25 retrieval** — ranked search over documents, symbols, and experience notes, with optional LLM query expansion to handle vocabulary mismatch. **CJK-optimized**: precise phrase boost (3+ char phrases ×1.5 score on title/keywords match), synonym table (e.g. 数据库连接池 ↔ 连接池 ↔ DB pool), and CJK-aware word boundaries for doc↔symbol linking.
19
+ - **BM25 memory recall** — ranked search over documents, symbols, and experience notes, with optional LLM query expansion to handle vocabulary mismatch. **CJK-optimized**: precise phrase boost (3+ char phrases ×1.5 score on title/keywords match), synonym table (e.g. 数据库连接池 ↔ 连接池 ↔ DB pool), and CJK-aware word boundaries for doc↔symbol linking.
19
20
  - **Experience notes** — problems → solutions; similar problems supersede instead of duplicating, and notes are returned only when a search matches. The note store is bounded: capacity scales with project size (clamped to 100–2000), and the oldest notes are pruned when the limit is exceeded. **Supersede tightened to bidirectional 0.7 overlap** (was 0.6); **experience `problem` field now participates in CJK phrase boost** for long-tail query recall.
21
+ - **Lock-free sync transactions** — all writes (index / watch / remember / forget / watch_repo) go through synchronous transactions `store.commit(fn)`; fn succeeds then atomic write; JS single-threaded event loop guarantees no interleaving; `remember`/`forget` never blocked by watch re-indexing.
20
22
  - **Minimal dependencies** — pure JavaScript; the only runtime dependency is `pdfjs-dist` (PDF text extraction), no native builds required.
21
23
  - **Negligible overhead** — pure in-process operation; cold start <100 ms (5k files), typical project query median 2–3 ms (p99 < 7 ms); bottleneck is LLM summarization and PDF parsing, not the plugin.
22
24
 
@@ -25,17 +27,17 @@ Persistent project memory for [DeepSeek Harness](https://github.com/deepseek-ai/
25
27
  The design follows four principles:
26
28
 
27
29
  - **Volatility** — context is ephemeral; it is lost when a session is compacted.
28
- - **Persistence** — the index is stored on disk and survives compaction and new sessions.
29
- - **Compactness** — only summaries are stored; the index runs around 0.5% the size of the source it covers (8.8 MB of source → 49 KB of index in the example project), so retrieval replaces re-reading the full file.
30
- - **Verifiability** — hits carry a `path:line` citation where applicable, so the agent can confirm details against the source.
30
+ - **Persistence** — the **memory** is stored on disk and survives compaction and new sessions.
31
+ - **Compactness** — only summaries are stored; the **memory** runs around 0.5% the size of the source it covers (8.8 MB of source → 49 KB of index in the example project), so **recall** replaces re-reading the full file.
32
+ - **Verifiability** — **recalls** carry a `path:line` citation where applicable, so the agent can confirm details against the source.
31
33
 
32
- Building the index does not require an upfront scan: files are indexed as the model reads them, so the index grows to cover exactly what has been worked with. Re-reading a file that has not changed is a no-op (content hash), so the index stays fresh with minimal ongoing overhead.
34
+ Building the **memory** does not require an upfront scan: files are memorized as the model reads them, so the **memory** grows to cover exactly what has been worked with. Re-reading a file that has not changed is a no-op (content hash), so the **memory** stays fresh with minimal ongoing overhead.
33
35
 
34
36
  The store is per-project and follows the codebase: changed files are re-extracted by content hash, deleted files are removed. Experience notes are retrieval-only, so accumulation does not affect context.
35
37
 
36
38
  ## Installation
37
39
 
38
- Tested against dsh **0.1.0-rc.7 through 0.1.2-alpha.1**. The plugin relies exclusively on stable public APIs (`defineTool`, `llm.stream`, `Schema`) declared via peerDependencies, ensuring compatibility with future rc releases without changes.
40
+ Tested against dsh **0.1.0-rc.7 through 0.1.2-alpha.2**. The plugin relies exclusively on stable public APIs (`defineTool`, `llm.stream`, `Schema`) declared via peerDependencies, ensuring compatibility with future rc releases without changes.
39
41
 
40
42
  ```bash
41
43
  cd dsh-project-memory && dsh plugin --profile web add . -w
@@ -52,7 +54,7 @@ dsh plugin --profile web add @yolk_vat-y/dsh-project-memory -w
52
54
  A prebuilt tarball is published with each release, installable without a build step:
53
55
 
54
56
  ```bash
55
- dsh plugin --profile web add /path/to/dsh-project-memory-0.3.0.tgz
57
+ dsh plugin --profile web add /path/to/dsh-project-memory.tgz
56
58
  ```
57
59
 
58
60
  Each indexed project has its own store at `<root>/.dsh-project-memory/`. Add it to `.gitignore` if it should not be committed.
@@ -93,12 +95,85 @@ Stores created before v0.2.0 (single `entries.json` / `index.json`) migrate auto
93
95
 
94
96
  These are deliberate scope choices.
95
97
 
96
- - **In-process locking** — store writes are serialized per memory directory within one dsh process; two dsh instances sharing a project store is last-writer-wins. A cross-process lock would need a resident daemon, which conflicts with the pure-JS, no-background-service positioning, so multi-instance writes are explicitly unsupported.
97
- - **Watch poll holds the lock** — while the watcher re-indexes changed docs (LLM summarization), `remember`/`forget` queue behind it. Polling (mtime + content hash) instead of `fs.watch` events keeps behavior consistent across platforms; overlapping polls serialize on the same lock: safe, but they can pile up on very large diffs. Tune via `watchInterval`.
98
- - **Corrupt files are quarantined** — a store JSON that fails to parse falls back to empty for that file and is rebuilt on the next write; the broken file is renamed to `*.corrupt` with an error logged, but its data cannot be recovered. Auto-repairing partial writes would need a write-ahead journal or an embedded database — out of proportion when quarantining one bad file costs nothing.
99
- - **`forget` by query is eager** — keyword deletion matches at ≥0.5 token overlap and may remove several notes at once; prefer deleting by id for precision.
100
- - **Cross-language recall depends on index time** — with `llmQueryExpansion` off, a Chinese-only query reaches English content through bilingual keywords captured when docs are indexed, plus doc↔symbol links; queries stay LLM-free. Stores indexed before v0.1.1 gain bilingual keywords as files change, or immediately via `index_repo` with `reindex: true`.
101
- - **CJK retrieval** — phrase boost and synonym expansion are purely query-side; they do not increase index size or LLM usage. Link boundaries use CJK-aware regex only; English symbols keep the original word-boundary behavior. The experience supersede threshold (0.7 bidirectional) is a conservative default; adjust via config if false positives/negatives appear in practice.
98
+ ### 1. Synchronous lock-free transactions over async locks
99
+
100
+ **We do:** All writes go through `store.commit(fn)` — a synchronous in-process transaction. The callback `fn` performs all validation and mutations; only on success is the result atomically written to disk. The JS event loop guarantees no interleaving. CAS (`applyFileUpdate`) makes concurrent writes idempotent.
101
+
102
+ **We don't:** Async mutexes, file locks, or multi-process coordination.
103
+
104
+ **Why:** DSH runs on Cordis, which is single-process by design. Adding locks would complicate the hot path (every `remember`/`forget`/`index_doc` call) for a scenario (multi-process DSH) that would require a breaking ecosystem change. Synchronous transactions keep the hot path at ~2 ms median with zero contention overhead in practice.
105
+
106
+ ### 2. Watch: compute outside, commit inside
107
+
108
+ **We do:** Heavy work (mtime/hash/scan/LLM summary) runs outside the transaction; a single `commit` applies all changes atomically. On failure, the snapshot rolls back so the next poll retries automatically.
109
+
110
+ **We don't:** Hold a lock during LLM calls, or use `fs.watch` events.
111
+
112
+ **Why:** LLM summarization takes seconds — holding a lock would block `remember`/`forget`/`query_memory`. Polling with mtime+content-hash is platform-agnostic (works on network drives, Docker volumes, WSL) and avoids the "double fire / missed events" nightmare of `fs.watch`.
113
+
114
+ ### 3. Corrupt files are quarantined, not auto-repaired
115
+
116
+ **We do:** On JSON parse failure, the bad file is renamed to `*.corrupt`, an error is logged, and that file's store starts fresh. The rest of the store remains intact.
117
+
118
+ **We don't:** Write-ahead logs, embedded databases (SQLite/LMDB), or automatic partial recovery.
119
+
120
+ **Why:** A corrupted shard means *one source file* has a bad index — quarantining it costs near zero. A WAL or embedded DB adds a heavy dependency, increases binary size, and introduces new failure modes (lock contention, corruption of the WAL itself). The tradeoff: lose one file's index vs. add 500 KB+ of native code.
121
+
122
+ ### 4. No vector embeddings, no semantic search at query time
123
+
124
+ **We do:** BM25 with CJK phrase boost (3+ chars ×1.5 on title/keywords), synonym expansion (bidirectional table), field weighting (title ×5), and experience-layer phrase boost. All at query time, zero LLM calls.
125
+
126
+ **We don't:** Vector embeddings, dense retrieval, rerankers, or hybrid search.
127
+
128
+ **Why:** Vectors require an embedding model (local = heavy, remote = latency + cost + privacy), a vector index (HNSW/IVF = memory + build time), and reranking (another LLM call). For code + docs + experience notes, lexical BM25 with our enhancements already achieves >90% recall on real queries. The marginal gain from semantic search doesn't justify the 10x complexity/cost increase.
129
+
130
+ ### 5. Cross-language recall at index time, not query time
131
+
132
+ **We do:** Doc keywords *must* cover both the document's language AND English. Doc↔symbol links surface English symbol names from Chinese queries. With `llmQueryExpansion: false`, queries never touch the LLM.
133
+
134
+ **We don't:** Translate queries at search time, or use multilingual embeddings.
135
+
136
+ **Why:** Query-time translation adds latency, token cost, and failure modes (bad translation = zero recall). Index-time bilingual keywords are a one-time cost per document; the LLM already summarizes the doc, so extracting English keywords is free. This also works offline and deterministically.
137
+
138
+ ### 6. Explicit `remember` over implicit learning
139
+
140
+ **We do:** Users (or the agent) explicitly call `remember(problem, solution)`. Supersede uses bidirectional token overlap ≥0.7 to deduplicate.
141
+
142
+ **We don't:** Automatically extract "lessons" from user corrections, or infer rules from conversation history.
143
+
144
+ **Why:** Implicit learning is unpredictable — it hallucinates, captures noise, and pollutes the memory with unverifiable entries. Explicit `remember` creates an auditable, user-controlled knowledge base. The cost (one tool call) is negligible; the benefit (trust, verifiability, no silent corruption) is decisive.
145
+
146
+ ### 7. Full entries returned directly
147
+
148
+ **We do:** `query_memory` returns complete entries with `path:line` citations. Every hit can be verified against source.
149
+
150
+ **We don't:** Return a minimal index first, then require a second tool call for details.
151
+
152
+ **Why:** Returning full entries preserves **verifiability** — the agent sees the exact source line for every claim. It also avoids a round-trip per useful hit. Our entries are already compact (~300 chars summary + citation); the token cost is lower than a second tool call + context switch.
153
+
154
+ ### 8. Symbol extraction focused on what developers search for
155
+
156
+ **We do:** Regex-based symbol extraction (functions, classes, methods, interfaces, type aliases) with string/comment masking, multi-line signatures, and cross-file linking by symbol name. For TypeScript/JavaScript projects, an optional L2 enhancement layer uses the TS Compiler API to infer return types, resolve generics, and extract interfaces — all cached by content hash for instant reuse.
157
+
158
+ **We don't:** Tree-sitter AST parsing, import graphs, call graphs, or full-program type resolution across files.
159
+
160
+ **Why:** Our regex scanner handles 8 languages with zero dependencies, runs in <1 ms/file, and captures the declarations developers actually search for (names, signatures, generics). The optional TS layer adds semantic depth for TS/JS without native deps. Cross-file linking by name covers the most common "find related code" use case. Full-program analysis would add native binaries, 10x install size, and version fragility — for marginal gain on the remaining 5% of edge cases.
161
+
162
+ ### 9. `forget` by query is aggressive; prefer ID deletion
163
+
164
+ **We do:** `forget query` deletes all experience notes with ≥0.5 token overlap.
165
+
166
+ **We don't:** Interactive confirmation, soft-delete/trash, or exact-match-only.
167
+
168
+ **Why:** Experience notes are low-stakes, high-volume, and retrieval-only. Aggressive deletion prevents stale noise from polluting search. For precision, delete by ID (shown in `query_memory` output).
169
+
170
+ ### 11. TypeScript enhancement is optional, lazy, and cached
171
+
172
+ **We do:** L2 TS Compiler API enhancement runs async in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`), results cached by content hash in `type-cache/`. Zero config — just `npm i -D typescript`. Falls back to L1 regex if TS absent or disabled.
173
+
174
+ **We don't:** Mandatory TS, blocking enhancement, or full-program type checking.
175
+
176
+ **Why:** Mandatory TS would break installs for non-TS projects. Blocking enhancement would stall `index_repo` on large codebases. Full-program checking is 10x slower and memory-heavy. Our design: enhance what's read, cache it, never block the hot path.
102
177
 
103
178
  ## Configuration
104
179
 
@@ -116,6 +191,8 @@ These are deliberate scope choices.
116
191
  | `autoIndexOnFirstUse` | false | full scan of the current working directory on plugin load (opt-in) |
117
192
  | `watch` | true | enable the background refresh |
118
193
  | `watchInterval` | 15 | poll interval (seconds) |
194
+ | `tsPath` | (auto) | optional absolute path to a specific `typescript` install; if omitted, resolves from project cwd → plugin node_modules |
195
+ | `enableTypeScript` | true | set `false` to disable L2 TS enhancement entirely (L1 regex only) |
119
196
 
120
197
  ### Toggling features
121
198
 
@@ -131,6 +208,8 @@ Settings live in the plugin's config object. To change them, add an override ent
131
208
  llmQueryExpansion: false # off: do not spend tokens on LLM query expansion (default)
132
209
  watch: true # on: background refresh for watched roots (default)
133
210
  watchInterval: 15 # poll interval in seconds
211
+ enableTypeScript: true # on: L2 TS enhancement when TS is installed (default)
212
+ # tsPath: /custom/path/to/typescript # optional: force specific TS install
134
213
  ```
135
214
 
136
215
  Only list the keys you want to change; the rest fall back to the plugin defaults. Verify the result with `dsh --profile web --dump-config`.
@@ -149,7 +228,7 @@ These commands are for **maintaining the plugin code** — regular users do not
149
228
 
150
229
  ```bash
151
230
  npm install
152
- npm test # chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
231
+ npm test # 157 tests (v0.3.2): chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
153
232
  ```
154
233
 
155
234
  ## License
package/README.zh-CN.md CHANGED
@@ -4,19 +4,21 @@
4
4
 
5
5
  [![ci](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml/badge.svg)](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml) [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![npm](https://img.shields.io/npm/v/@yolk_vat-y/dsh-project-memory)](https://www.npmjs.com/package/@yolk_vat-y/dsh-project-memory) [![Listed on dsh-plugin.org](https://dsh-plugin.org/badges/listed.svg)](https://dsh-plugin.org/plugins/00080000/dsh-project-memory) [![Awesome](https://awesome-dsh-plugin.com/badge.svg)](https://awesome-dsh-plugin.com)
6
6
 
7
- 为 [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(dsh)agent 提供持久化的项目记忆。将文档(PDF / Markdown / txt)与代码符号索引进每个工作区独立的存储库,自动维护更新,召回时附源文件引用——文档自动交叉链接到其提及的代码符号。
7
+ 为 [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(dsh)agent 提供持久化的项目记忆。**记住**文档(PDF / Markdown / txt)与代码符号,写入每个工作区独立的存储库,自动维护更新,**回忆**时附源文件引用——文档自动交叉链接到其提及的代码符号。
8
8
 
9
- > 插件在磁盘上维护一份精简的项目索引,每条记录指向具体的文件与行号;agent 需要快速了解项目时先查索引,无需重读整个项目。
9
+ > 插件在磁盘上维护一份精简的项目**记忆**,每条记录指向具体的文件与行号;agent 需要快速了解项目时先查**记忆**,无需重读整个项目。
10
10
 
11
11
  ## 特性
12
12
 
13
- - **文档索引** — PDF、Markdown、纯文本按块切分并由 LLM 生成摘要,每条索引携带 `路径:行号` 引用回源文件。
14
- - **代码符号表** — 通过零依赖的源码扫描器提取函数、类与方法名(字符串/注释掩码、多行签名续行、Python 缩进感知、类方法上下文),不使用 LLM token。
15
- - **自动刷新** — `watch_repo` 后台轮询,按内容哈希识别新增或变更文件,仅重抽这些文件。
16
- - **读到即索引** — 文件在模型**实际读取的瞬间**被索引(监听 `fs/observed`),索引是正常工作的副产品,而非额外的一次全量扫描。从未读过的文件不会被索引。项目根通过标记(`.git`、`package.json` 等)、README 加源码目录、或兜底到文件所在目录逐级识别。
13
+ - **文档记忆** — PDF、Markdown、纯文本按块切分并由 LLM 生成摘要,每条记忆携带 `路径:行号` 引用回源文件。
14
+ - **代码符号记忆** — 通过零依赖的源码扫描器提取函数、类与方法名及完整类型签名(泛型、参数类型、返回类型、重载签名),包含字符串/注释掩码、多行签名续行、Python 缩进感知、类方法上下文,不使用 LLM token。
15
+ - **可选 TypeScript 语义增强** — 当用户项目安装了 `typescript`(`npm i -D typescript`),插件自动激活第二层(L2),利用 TS Compiler API 推导返回类型、实例化泛型、提取接口与类型别名、丰富箭头函数签名 —— 全部在优先级队列中异步后台处理(P0:`fs/observed` 读文件瞬间、P1:`watch` 变更后、P2:`index_repo` 批量索引)。结果按文件内容哈希缓存到磁盘,冷启动毫秒级复用。零配置:装 TS 再重启 dsh 即可。完全可选;若无 TS 或设置 `enableTypeScript: false`,回退至 L1 正则提取。
16
+ - **自动刷新** — `watch_repo` 后台轮询,按内容哈希识别新增或变更文件,仅重记这些文件。
17
+ - **读到即记忆** — 文件在模型**实际读取的瞬间**被记忆(监听 `fs/observed`),记忆是正常工作的副产品,而非额外的一次全量扫描。从未读过的文件不会被记忆。项目根通过标记(`.git`、`package.json` 等)、README 加源码目录、或兜底到文件所在目录逐级识别。
17
18
  - **文档 ↔ 代码交叉链接** — 文档提及某符号时记录为 `reference`;查询符号时同时带出描述该符号的文档。
18
- - **BM25 检索** — 对文档、符号与经验笔记进行排序召回,可选 LLM 查询扩展以应对表述不一致。**CJK 增强**:精确短语乘法加分(3+ 字短语在标题/关键词命中 ×1.5)、同义词表(如 数据库连接池 ↔ 连接池 ↔ DB pool)、CJK 感知的文档↔符号链接边界。
19
+ - **BM25 记忆召回** — 对文档、符号与经验笔记进行排序召回,可选 LLM 查询扩展以应对表述不一致。**CJK 增强**:精确短语乘法加分(3+ 字短语在标题/关键词命中 ×1.5)、同义词表(如 数据库连接池 ↔ 连接池 ↔ DB pool)、CJK 感知的文档↔符号链接边界。
19
20
  - **经验笔记** — 记录问题 → 方案;相似问题覆盖而非重复;笔记仅在检索命中时返回。笔记数量有界:容量随项目规模伸缩(钳制在 100–2000),超限时淘汰最旧的笔记。**覆盖阈值收紧为双向 0.7 重叠**(原 0.6);**经验 `problem` 字段现参与 CJK 短语加分**,提升长尾问句召回。
21
+ - **无锁同步事务** — 不采用锁:所有写入(index / watch / remember / forget / watch_repo)统一走同步事务 `store.commit(fn)`,fn 成功后才一次落盘;JS 单线程事件循环保证事务间不交错,`remember`/`forget` 不会被 watch 重索引阻塞排队。多实例并发写入同一项目存储时,得益于 CAS 幂等更新与原子提交,自然具备幂等性,无数据损坏风险。
20
22
  - **依赖极简** — 纯 JavaScript;唯一运行时依赖是 `pdfjs-dist`(PDF 文本提取),无需原生构建。
21
23
  - **开销可忽略** — 纯进程内操作;冷启动 <100 ms(5k 文件),典型项目查询中位数 2–3 ms(p99 < 7 ms);瓶颈在 LLM 摘要与 PDF 解析,插件本身不阻塞。
22
24
 
@@ -25,17 +27,17 @@
25
27
  设计遵循四个原则:
26
28
 
27
29
  - **易失性** — 上下文是临时的,会话压缩即丢失。
28
- - **持久性** — 索引存于磁盘,跨压缩与会话保留。
29
- - **紧凑性** — 仅存摘要;索引规模约为其覆盖源码的 0.5%(示例项目中 8.8 MB 源码 → 49 KB 索引),检索替代了通读整个文件。
30
- - **可核验性** — 命中在适用时携带 `路径:行号` 引用,agent 可对照源文件核实。
30
+ - **持久性** — **记忆**存于磁盘,跨压缩与会话保留。
31
+ - **紧凑性** — 仅存摘要;**记忆**规模约为其覆盖源码的 0.5%(示例项目中 8.8 MB 源码 → 49 KB 索引),**召回**替代了通读整个文件。
32
+ - **可核验性** — **召回**在适用时携带 `路径:行号` 引用,agent 可对照源文件核实。
31
33
 
32
- 构建索引无需预先全量扫描:文件在模型读取时被索引,索引恰好覆盖实际处理过的内容。未变更的文件重读是空操作(内容哈希),因此索引的持续维护开销很低。
34
+ 构建**记忆**无需预先全量扫描:文件在模型读取时被记忆,**记忆**恰好覆盖实际处理过的内容。未变更的文件重读是空操作(内容哈希),因此**记忆**的持续维护开销很低。
33
35
 
34
36
  存储按项目独立存放,并跟随代码库变化:文件变更按内容哈希重新抽取,文件删除则同步移除。经验层仅检索,累积不影响上下文。
35
37
 
36
38
  ## 安装
37
39
 
38
- 实测覆盖 dsh **0.1.0-rc.7 至 0.1.2-alpha.1**。插件仅依赖通过 peerDependencies 声明的稳定公共 API(`defineTool`、`llm.stream`、`Schema`),保证与后续 rc 版本无需改动即兼容。
40
+ 实测覆盖 dsh **0.1.0-rc.7 至 0.1.2-alpha.2**。插件仅依赖通过 peerDependencies 声明的稳定公共 API(`defineTool`、`llm.stream`、`Schema`),保证与后续 rc 版本无需改动即兼容。
39
41
 
40
42
  ```bash
41
43
  cd dsh-project-memory && dsh plugin --profile web add . -w
@@ -52,7 +54,7 @@ dsh plugin --profile web add @yolk_vat-y/dsh-project-memory -w
52
54
  每个版本会附带预构建 tarball,无需构建步骤即可安装:
53
55
 
54
56
  ```bash
55
- dsh plugin --profile web add /path/to/dsh-project-memory-0.3.0.tgz
57
+ dsh plugin --profile web add /path/to/dsh-project-memory.tgz
56
58
  ```
57
59
 
58
60
  每个被索引的项目在 `<root>/.dsh-project-memory/` 下有独立存储。如无需入库,可加入 `.gitignore`。
@@ -93,12 +95,85 @@ v0.2.0 之前创建的库(单文件 `entries.json` / `index.json`)在首次
93
95
 
94
96
  以下是刻意的范围选择。
95
97
 
96
- - **无锁同步事务** — 不采用锁:所有写入(index / watch / remember / forget / watch_repo)统一走同步事务 `store.commit(fn)`,fn 成功后才一次落盘;JS 单线程事件循环保证事务间不交错,`remember`/`forget` 不会被 watch 重索引阻塞排队。两个 dsh 实例共享同一项目存储时后写覆盖先写;跨进程协调需要常驻守护进程,违背纯 JS 插件、无后台服务的定位,故明确不支持多实例共写。
97
- - **watch 计算与提交分离** — watcher 在事务外完成 mtime/哈希/符号扫描/LLM 摘要等重活,再以单次 `commit` 原子应用全部变更;`applyFileUpdate` 用 CAS 校验(统一 null 处理、删除跳过对比)防并发修改,失败回滚 snapshot 下轮重试,索引失败删除 snapshot 自动重试。轮询(mtime + 内容哈希)而非 `fs.watch` 事件驱动,是为了跨平台行为一致。间隔可用 `watchInterval` 调整。
98
- - **损坏隔离重建** — 存储 JSON 损坏时该文件回落为空并在下次写入时重建;坏文件会改名备份为 `*.corrupt` 并输出错误日志,但该文件内的数据无法恢复。自动修复半写文件需要预写日志或嵌入式数据库,代价与收益不成比例——而隔离一个坏文件的成本几乎为零。
99
- - **`forget` 按关键词删除偏激进** — 关键词删除按 ≥0.5 token 重叠匹配,可能一次删掉多条;追求精确请用 id 删除。
100
- - **跨语种召回依赖索引时** — `llmQueryExpansion` 关闭时,纯中文查询靠索引时捕获的双语 keywords 和 doc↔symbol 链接触达英文内容,查询侧保持零 LLM 调用。v0.1.1 之前建立的索引随文件变更逐步获得双语关键词,或用 `index_repo` 的 `reindex: true` 立即重建。
101
- - **CJK 检索** — 短语加分与同义词展开完全在查询侧,不增加索引体积、不额外消耗 LLM token。链接边界仅对 CJK 使用正则边界,英文符号保持原有词边界行为。经验层 supersede 阈值(双向 0.7)为保守默认;若实测出现误覆盖/误漏报,可通过配置调整。
98
+ ### 1. 同步无锁事务,而非异步锁
99
+
100
+ **我们做:** 所有写入走 `store.commit(fn)` 同步事务。回调 `fn` 内完成校验与变更,成功后才原子落盘。JS 事件循环天然串行,CAS (`applyFileUpdate`) 让并发写入幂等。
101
+
102
+ **不做:** 异步互斥锁、文件锁、多进程协调。
103
+
104
+ **为什么:** DSH 基于 Cordis,单进程是架构基石。为极少见的多进程场景加锁,会让热路径(每次 `remember`/`forget`/`index_doc`)增重。同步事务让热路径中位数 ~2 ms,零争用开销。
105
+
106
+ ### 2. Watch:事务外计算,事务内提交
107
+
108
+ **我们做:** 重活(mtime/哈希/扫描/LLM 摘要)在事务外跑,单次 `commit` 原子应用全部变更。失败回滚 snapshot,下轮自动重试。
109
+
110
+ **不做:** 持锁调用 LLM,或用 `fs.watch` 事件。
111
+
112
+ **为什么:** LLM 摘要耗时秒级,持锁会阻塞 `remember`/`forget`/`query_memory`。轮询 + mtime+内容哈希跨平台一致(网络盘、Docker 卷、WSL 皆可),避免 `fs.watch` 的「重复触发/漏事件」噩梦。
113
+
114
+ ### 3. 损坏文件隔离,不自动修复
115
+
116
+ **我们做:** JSON 解析失败时,坏文件改名 `*.corrupt`、记错误、该文件存储重头开始,其余分片不受影响。
117
+
118
+ **不做:** 预写日志 (WAL)、嵌入式数据库、自动部分恢复。
119
+
120
+ **为什么:** 一个损坏分片 = 一个源文件索引丢失,隔离成本近零。WAL 或嵌入式 DB 增加 500 KB+ 原生依赖、锁竞争、新故障模式(WAL 自身损坏)。权衡:丢一个文件索引 vs. 引入重型原生栈。
121
+
122
+ ### 4. 查询零向量、零语义搜索
123
+
124
+ **我们做:** BM25 + CJK 短语加分(3+ 字 ×1.5)、同义词表双向展开、字段加权(标题 ×5)、经验层短语加分。查询侧零 LLM 调用。
125
+
126
+ **不做:** 向量嵌入、稠密检索、重排序、混合搜索。
127
+
128
+ **为什么:** 向量需要嵌入模型(本地重、远程慢+贵+隐私)、向量索引(HNSW/IVF 占内存+建索引慢)、重排序(再调一次 LLM)。对代码+文档+经验笔记,增强 BM25 已达 >90% 实战召回。边际收益不抵 10x 复杂度/成本。
129
+
130
+ ### 5. 跨语种召回在索引时完成,而非查询时
131
+
132
+ **我们做:** 文档 keywords 强制双语(文档语言+英文);doc↔symbol 链接从中文命中带出英文符号名。`llmQueryExpansion: false` 时查询完全不碰 LLM。
133
+
134
+ **不做:** 查询时翻译、多语言向量。
135
+
136
+ **为什么:** 查询时翻译增延迟、耗 token、易翻车(译错=零召回)。索引时双语 keywords 是一次性成本(LLM 摘要时顺手提取),离线确定、可复用。
137
+
138
+ ### 6. 显式 `remember`,不做隐式学习
139
+
140
+ **我们做:** 用户/显式调用 `remember(problem, solution)`。supersede 用双向 token 重叠 ≥0.7 去重。
141
+
142
+ **不做:** 从用户纠正中自动抽「教训→规则」、从对话历史推断规则。
143
+
144
+ **为什么:** 隐式学习不可控——会幻觉、收噪音、污染记忆库且不可审计。显式 `remember` 成本极低(一次工具调用),换来可信、可追溯、用户可控的知识库。
145
+
146
+ ### 7. 直接返回完整条目
147
+
148
+ **我们做:** `query_memory` 直接返回含 `path:line` 引用的完整条目,每条可回源核实。
149
+
150
+ **不做:** 先返回极简索引(如 700 字符),再二次调工具取详情。
151
+
152
+ **为什么:** 完整返回保持 **可核验性**——Agent 能看到每条声明的出处行号。也避免了每次有效命中多一轮工具调用+上下文切换。条目本已紧凑(~300 字摘要+引用),完整返回的 token 成本低于二次调用。
153
+
154
+ ### 8. 符号提取聚焦开发者实际搜索的内容
155
+
156
+ **我们做:** 正则符号提取(函数/类/方法/接口/类型别名),含字符串/注释掩码、多行签名、跨文件按名链接。对 TypeScript/JavaScript 项目,可选的 L2 增强层利用 TS Compiler API 推导返回类型、实例化泛型、提取接口 —— 全部按内容哈希缓存,毫秒级复用。
157
+
158
+ **不做:** tree-sitter AST、导入图、调用图、跨文件全程序类型推导。
159
+
160
+ **为什么:** 正则扫描器零依赖、8 语言、<1 ms/文件,覆盖开发者最常搜索的声明(名字、签名、泛型)。可选 TS 增强层为 TS/JS 提供语义深度,且无原生依赖。按名跨文件链接已覆盖最常见的「找相关代码」场景。全程序分析会引入原生二进制、安装体积增 10x、语言版本即破——边际收益仅在剩余 5% 的边缘情况。
161
+
162
+ ### 9. `forget` 按关键词激进;精确请用 ID
163
+
164
+ **我们做:** `forget query` 删除所有 token 重叠 ≥0.5 的经验笔记。
165
+
166
+ **不做:** 交互确认、软删除/回收站、仅精确匹配。
167
+
168
+ **为什么:** 经验笔记低风险、高量、仅检索。激进删除防止陈旧噪音污染搜索。精确删用 ID(`query_memory` 输出里有)。
169
+
170
+ ### 11. TS 增强可选、异步、缓存
171
+
172
+ **我们做:** L2 TS Compiler API 在优先级队列异步跑(P0 `fs/observed`、P1 `watch`、P2 `index_repo`),结果按内容哈希缓存 `type-cache/`。零配置——`npm i -D typescript` 即用。无 TS 或禁用时优雅回退 L1 正则。
173
+
174
+ **不做:** 强制 TS、阻塞式增强、全程序类型检查。
175
+
176
+ **为什么:** 强制 TS 会让非 TS 项目装不上。阻塞增强会卡死大项目 `index_repo`。全程序检查慢 10x、内存重。设计:读到即增强、缓存复用、热路径永不阻塞。
102
177
 
103
178
  ## 配置
104
179
 
@@ -116,6 +191,8 @@ v0.2.0 之前创建的库(单文件 `entries.json` / `index.json`)在首次
116
191
  | `autoIndexOnFirstUse` | false | 插件加载时对当前工作目录做全量扫描(可选) |
117
192
  | `watch` | true | 启用后台刷新 |
118
193
  | `watchInterval` | 15 | 轮询间隔(秒) |
194
+ | `tsPath` | (自动) | 可选:强制指定特定 `typescript` 安装路径;省略时按项目 cwd → 插件 node_modules 向上解析 |
195
+ | `enableTypeScript` | true | 设为 `false` 彻底禁用 L2 TS 增强(仅保留 L1 正则) |
119
196
 
120
197
  ### 功能开关
121
198
 
@@ -131,6 +208,8 @@ v0.2.0 之前创建的库(单文件 `entries.json` / `index.json`)在首次
131
208
  llmQueryExpansion: false # 关闭:不用 LLM 扩展查询,节省 token(默认)
132
209
  watch: true # 开启:被监听根目录后台保持新鲜(默认)
133
210
  watchInterval: 15 # 轮询间隔(秒)
211
+ enableTypeScript: true # 开启:装了 TS 时启用 L2 语义增强(默认)
212
+ # tsPath: /custom/path/to/typescript # 可选:强制指定 TS 安装路径
134
213
  ```
135
214
 
136
215
  只需列出要改的键,其余键回落到插件默认值。用 `dsh --profile web --dump-config` 验证生效。
@@ -149,7 +228,7 @@ dsh web --patch ./config.yml
149
228
 
150
229
  ```bash
151
230
  npm install
152
- npm test # 检查项:chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
231
+ npm test # 157 tests (v0.3.2):chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
153
232
  ```
154
233
 
155
234
  ## 许可证
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@yolk_vat-y/dsh-project-memory",
3
- "version": "0.3.1",
3
+ "version": "0.3.3",
4
4
  "description": "Persistent project memory for dsh agents: index docs (PDF/Markdown/text) and code symbols into a searchable per-workspace store, recall them with cited sources, and keep experience entries (problems -> solutions) searchable on demand.",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
@@ -39,16 +39,18 @@
39
39
  "pdfjs-dist": "^4.10.38"
40
40
  },
41
41
  "devDependencies": {
42
- "@deepseek-ai/cordis": "4.0.1",
43
- "@deepseek-ai/dsh-tools": "0.1.1-rc.2",
44
- "@deepseek-ai/schemastery": "3.18.1",
45
- "@deepseek-ai/dsh-llm": "0.1.1-rc.2"
42
+ "@deepseek-ai/cordis": "^4.0.2",
43
+ "@deepseek-ai/dsh-llm": "^0.1.2-alpha.2",
44
+ "@deepseek-ai/dsh-tools": "^0.1.2-alpha.2",
45
+ "@deepseek-ai/schemastery": "^3.18.2",
46
+ "typescript": "^5.6.3"
46
47
  },
47
48
  "peerDependencies": {
48
49
  "@deepseek-ai/cordis": "^4.0.1",
50
+ "@deepseek-ai/dsh-llm": ">=0.0.1-rc.1 <0.1.0 || >=0.1.0-rc.1 <0.2.0-0 || >=0.1.1-rc.1 <0.2.0-0",
49
51
  "@deepseek-ai/dsh-tools": ">=0.0.1-rc.1 <0.1.0 || >=0.1.0-rc.1 <0.2.0-0 || >=0.1.1-rc.1 <0.2.0-0",
50
52
  "@deepseek-ai/schemastery": "^3.18.1",
51
- "@deepseek-ai/dsh-llm": ">=0.0.1-rc.1 <0.1.0 || >=0.1.0-rc.1 <0.2.0-0 || >=0.1.1-rc.1 <0.2.0-0"
53
+ "typescript": "^5.0.0"
52
54
  },
53
55
  "peerDependenciesMeta": {
54
56
  "@deepseek-ai/cordis": {
@@ -62,6 +64,9 @@
62
64
  },
63
65
  "@deepseek-ai/dsh-llm": {
64
66
  "optional": true
67
+ },
68
+ "typescript": {
69
+ "optional": true
65
70
  }
66
71
  },
67
72
  "dsh": {
@@ -1,3 +1,4 @@
1
+ import { createHash } from 'node:crypto'
1
2
  import path from 'node:path'
2
3
  import { stat } from 'node:fs/promises'
3
4
  import { looksLikeDump, readTextFile } from './util/fs.js'
@@ -27,6 +28,10 @@ export async function buildDocEntries(llm, a, b, c) {
27
28
  const [relPath, filePath, opts] = c === undefined ? [a, a, b] : [a, b, c]
28
29
  const text = await extractTextFromFile(filePath, opts)
29
30
  if (looksLikeDump(text)) return null
31
+
32
+ // Compute content hash for update detection
33
+ const hash = createHash('sha256').update(text).digest('hex').slice(0, 16)
34
+
30
35
  const chunks = chunkText(text, opts.chunkChars, opts.maxChunks)
31
36
  const metas = new Array(chunks.length)
32
37
  let cursor = 0
@@ -47,7 +52,9 @@ export async function buildDocEntries(llm, a, b, c) {
47
52
  type: 'doc',
48
53
  title: meta.title,
49
54
  summary: meta.summary,
55
+ blindSpots: meta.blindSpots || '',
50
56
  keywords: meta.keywords,
57
+ hash,
51
58
  }))
52
59
  }
53
60