@yolk_vat-y/dsh-project-memory 0.3.3 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,34 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.4.0 (2026-09-02)
4
+
5
+ ### TaskBridge:跨会话开发任务(取代 v1.2 workflow 方案,v2.5 契约落地)
6
+
7
+ - **定位**:模型会话内用宿主 `todo_write` 维护的清单 + 实际读过的文件,自动沉淀为跨会话可续接的任务实体——新会话 `list_tasks` → `select_task` 即接回进度与文件
8
+ - **事件订阅(`session/event`,签名 (session, event))**:`todo/write` → 绑定任务 steps 快照整体覆盖,未绑定会话自动新建任务并绑定;`tool/call`(read/write/edit/read_image,`arguments` 为 JSON 串)→ 绑定任务 files 并集(归一化相对路径、项目外拒绝、上限 100)
9
+ - **标题由模型定**:`select_task(title=任务名)` 先命名再写 todo;自动回退取用户消息最后「:」后的任务段(截断 48 字);续接可 `select_task(taskId, title)` 改名
10
+ - **工具**:`list_tasks` / `select_task`(taskId 精确、自动解归档;title 完全匹配、多候选返回列表、无则新建)/ `archive_task`
11
+ - **用户命令 `/tasks`**(`ctx.commands` 存在时注册,handler 不经模型):任务数、标题、步骤进度、涉及文件、当前会话绑定;快捷键无宿主 API 不做
12
+ - **`query_memory`**:新增 `type:'task'` 检索(title/steps/files);`type:'all'` 结果尾部附一行任务计数提示
13
+ - **存储**:`.dsh-project-memory/tasks.json` + `binding.json`(load 兜底空值、独立于 format v2 布局);容量随项目体积自适应 `fileCount/20` clamp [5,100],超限按 lastActiveAt 归档最旧
14
+ - **边界**:不做步骤↔文件映射(todo 无 id、全量替换,语义上不可靠);不去重(宿主语义);子代理会话无法可靠判定,接受其自动建档(低频)
15
+ - **环境要求**:自动同步需含 `session/event` 事件与 `todo_write` 的 dsh(0.1.2-alpha.x 实测);旧宿主 `ctx.on('session/event')` 不触发时降级——任务工具仍可作纯记录使用
16
+ - **测试**:新增 `test/taskbridge.test.mjs`(5 项:持久化往返/容量裁剪/路径归一化/自动建任务+快照覆盖/tool 文件跟踪边界);既有 166 项测试全绿
17
+ - **文档同步**:README.md / README.zh-CN.md
18
+
19
+ ## 0.3.4 (2026-09-02)
20
+
21
+ ### query_memory 性能优化:流式 TF + IDF 缓存
22
+ - **IDF 缓存**:`store._version` + `store._idfCache`,`save()` 时 `version++` 标记失效,查询时版本命中直接复用,无需重建 BM25 索引
23
+ - **预计算 searchText**:`setEntries` 时预计算 `entry.searchText`(标题×5 + keywords + summary + path 的小写拼接),查询时直接复用,避免重复字符串拼接与 `toLowerCase()`
24
+ - **流式打分**:`rankEntriesStreaming` 单次遍历 entries,用 `countOccurrences()` 字符串计数替代完整 `tokenizeRaw` + TF 表构建,零中间对象分配
25
+ - **性能提升**:5k 文件 / 20k 条目场景 `query_memory` 中位数 **187 ms → 9.3 ms**(20x);1k 文件典型项目 **<1 ms**
26
+
27
+ ### 测试覆盖
28
+ - 新增 `IDF caching & streaming TF` 测试组(7 项):缓存构建、命中、版本失效、流式打分正确性、空查询、searchText 预计算
29
+ - 新增 `query_memory streaming path` 集成测试(2 项):端到端首次查询建缓存、二次查询复用缓存
30
+ - 总测试数 157 → 166 全绿
31
+
3
32
  ## 0.3.3 (2026-08-31)
4
33
 
5
34
  ### 符号层:只存身份牌,不存行为
package/README.md CHANGED
@@ -4,24 +4,54 @@
4
4
 
5
5
  [![ci](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml/badge.svg)](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml) [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![npm](https://img.shields.io/npm/v/@yolk_vat-y/dsh-project-memory)](https://www.npmjs.com/package/@yolk_vat-y/dsh-project-memory) [![Listed on dsh-plugin.org](https://dsh-plugin.org/badges/listed.svg)](https://dsh-plugin.org/plugins/00080000/dsh-project-memory) [![Awesome](https://awesome-dsh-plugin.com/badge.svg)](https://awesome-dsh-plugin.com)
6
6
 
7
- Persistent project memory for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(dsh) agents. **Memorizes** documents (PDF / Markdown / txt) and code symbols into a per-workspace store, refreshes them automatically, and **recalls** them with source citations — documents are cross-linked to the code symbols they reference.
7
+ A persistent **project memory** for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (dsh) agents. Built for software development: the agent's task list (todo_write) and file reads are continuously consolidated into durable cross-session task records, solving context loss. Documents (PDF/Markdown/txt) and code symbols are indexed into a per-workspace store with doc↔symbol cross-links. Experience notes (problem → solution) are deduplicated automatically. All data stored per-project on disk, survives session compaction, recalls with `path:line` citations for verification. Zero external dependencies (only `pdfjs-dist` for PDF text), no vector DB, no native builds.
8
8
 
9
- > The plugin keeps a compact project **memory** on disk, with every entry pointing to a concrete file and line — the agent can reorient quickly instead of re-reading the whole project.
9
+ > The plugin keeps a compact project **memory** on disk, with every entry pointing to a concrete file and line — the agent can reorient quickly instead of re-reading the whole project. Tasks and experience persist across session compactions and handovers.
10
10
 
11
11
  ## Features
12
12
 
13
13
  - **Document memorization** — PDF, Markdown, and plain text files are chunked and summarized by the LLM; each entry carries a `path:line` citation back to the source.
14
14
  - **Code symbol memory** — function, class, and method names with full type signatures (generics, parameters, return types, overloads) are extracted by a dependency-free source scanner (string/comment masking, multi-line signature joining, indentation-aware Python, class-method context), without LLM token usage.
15
- - **Optional TypeScript semantic enhancement** — when `typescript` is installed in the user project (`npm i -D typescript`), the plugin automatically activates a second layer (L2) that uses the TS Compiler API to infer return types, resolve generics, extract interfaces and type aliases, and enrich arrow functions — all asynchronously in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`). Results are cached on disk keyed by file content hash for instant cold-start reuse. Zero config: just install TS and restart dsh. Fully optional; if TS is absent or disabled via `enableTypeScript: false`, the plugin falls back to L1 regex-only extraction.
15
+ - **L1 Enhanced Regex** — zero-dep regex scanner now extracts generics, parameter/return types, overloads, interfaces, and type aliases for all supported languages, producing one-line identity signatures `fn(a: A, b: B): R — file.ts:42`.
16
+ - **Optional TypeScript semantic enhancement (L2/L3)** — when `typescript` is installed in the user project (`npm i -D typescript`), the plugin automatically activates a second layer (L2) that uses the TS Compiler API to infer return types, resolve generics, extract interfaces and type aliases, and enrich arrow functions — all asynchronously in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`). Results are cached on disk keyed by file content hash (L3) for instant cold-start reuse. Zero config: just install TS and restart dsh. Fully optional; if TS is absent or disabled via `enableTypeScript: false`, the plugin falls back to L1 regex-only extraction.
16
17
  - **Automatic refresh** — a background poll (`watch_repo`) detects new or changed files by content hash and re-memorizes only those.
17
18
  - **Read-time memorization** — files are memorized the moment the model actually reads them (`fs/observed`), so the memory is a byproduct of normal work, not a separate upfront scan. Files that are never read are never indexed. The project root is detected by markers (`.git`, `package.json`, …), a README plus source directories, or the file's own directory as a last resort.
18
19
  - **Doc ↔ code cross-linking** — when a document mentions a symbol, the match is recorded as a `reference`; querying a symbol also surfaces the documents that describe it.
19
20
  - **BM25 memory recall** — ranked search over documents, symbols, and experience notes, with optional LLM query expansion to handle vocabulary mismatch. **CJK-optimized**: precise phrase boost (3+ char phrases ×1.5 score on title/keywords match), synonym table (e.g. 数据库连接池 ↔ 连接池 ↔ DB pool), and CJK-aware word boundaries for doc↔symbol linking.
21
+ - **blindSpots-aware recall** — document summaries carry a `blindSpots` field (what the summary explicitly does NOT cover). When a query hits a blind spot, `query_memory` appends a warning pointing the model to read the source file, preventing hallucination from partial summaries.
20
22
  - **Experience notes** — problems → solutions; similar problems supersede instead of duplicating, and notes are returned only when a search matches. The note store is bounded: capacity scales with project size (clamped to 100–2000), and the oldest notes are pruned when the limit is exceeded. **Supersede tightened to bidirectional 0.7 overlap** (was 0.6); **experience `problem` field now participates in CJK phrase boost** for long-tail query recall.
23
+ - **Streaming TF + IDF caching** — query path caches IDF (term inverse frequency) per store version; on cache hit, single-pass streaming scores 20k entries in ~8 ms (5k files) / ~1 ms (1k files) with zero intermediate objects; write path is O(1) version bump.
21
24
  - **Lock-free sync transactions** — all writes (index / watch / remember / forget / watch_repo) go through synchronous transactions `store.commit(fn)`; fn succeeds then atomic write; JS single-threaded event loop guarantees no interleaving; `remember`/`forget` never blocked by watch re-indexing.
25
+ - **TaskBridge: cross-session development tasks** — the plugin watches each session's live todo list (`todo_write` events) and file reads (`tool/call`): progress snapshots (`steps`) and touched files sync into durable per-project task entities. An unbound session that writes a todo auto-creates a task. New sessions continue by `list_tasks` → `select_task` (bind / rename / unarchive); `query_memory` gains `type: 'task'` and appends a task-count hint to `type: 'all'` results. The user-side `/tasks` command shows the task stack, step progress, involved files, and the current session binding. Titles are chosen by the model via `select_task(title=…)` (fallback: the part of your message after the last colon). Capacity is project-size adaptive (`fileCount/20`, clamped 5–100). Storage: `.dsh-project-memory/tasks.json` + `binding.json`. Auto-sync requires a dsh build with session events + `todo_write` (verified on 0.1.2-alpha.x); on older hosts the task tools still work as a plain record list.
22
26
  - **Minimal dependencies** — pure JavaScript; the only runtime dependency is `pdfjs-dist` (PDF text extraction), no native builds required.
23
27
  - **Negligible overhead** — pure in-process operation; cold start <100 ms (5k files), typical project query median 2–3 ms (p99 < 7 ms); bottleneck is LLM summarization and PDF parsing, not the plugin.
24
28
 
29
+ ## Performance
30
+
31
+ ### Synthetic Benchmark (isolated environment, Node 24, Linux)
32
+
33
+ | Scenario | Scale | Measured |
34
+ |----------|-------|----------|
35
+ | Full cold index | 5,000 files / 20k entries | 353 ms |
36
+ | Cold load | 5,000 files | 82 ms |
37
+ | Hot lazy re-index (single file) | 5k files | median 2.3 ms / max 4.0 ms |
38
+ | query_memory (cached) | 5k files / 20k entries | median 9.3 ms / p95 12.6 ms |
39
+ | query_memory (cached) | 1k files / 4k entries | median 1.0 ms / p95 2.0 ms |
40
+ | Full cold index | 10,000 files / 40k entries | 637 ms |
41
+ | Cold load | 10,000 files | 144 ms |
42
+ | Hot lazy re-index (single file) | 10k files | median 4.5 ms / max 10.2 ms |
43
+
44
+ > Synthetic benchmark: generated code (~8 symbols/file), Node 24, Linux native FS, SSD. Measures pure indexing overhead without LLM calls. query_memory benchmark uses IDF cache + precomputed searchText; first query after write rebuilds IDF (~150 ms), subsequent queries hit cache.
45
+
46
+ ### Real Project Storage
47
+
48
+ | Project | Files | Entries | Store Size | Per Entry |
49
+ |---------|-------|---------|------------|-----------|
50
+ | Java Spring Boot backend | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
51
+ | Vue 3 + Vite frontend | 289 | 2,141 | 1.0 MB | ~0.5 KB |
52
+
53
+ > Real projects (Java + Vue), tested on Linux file system (Node 24). Real project entries are smaller than synthetic benchmarks due to lower symbol density and shorter declarations.
54
+
25
55
  ## How it works
26
56
 
27
57
  The design follows four principles:
@@ -37,7 +67,7 @@ The store is per-project and follows the codebase: changed files are re-extracte
37
67
 
38
68
  ## Installation
39
69
 
40
- Tested against dsh **0.1.0-rc.7 through 0.1.2-alpha.2**. The plugin relies exclusively on stable public APIs (`defineTool`, `llm.stream`, `Schema`) declared via peerDependencies, ensuring compatibility with future rc releases without changes.
70
+ Tested against dsh **0.1.0-rc.7 through 0.1.2-alpha.3**. The plugin relies exclusively on stable public APIs (`defineTool`, `llm.stream`, `Schema`) declared via peerDependencies, ensuring compatibility with future rc/alpha releases without changes.
41
71
 
42
72
  ```bash
43
73
  cd dsh-project-memory && dsh plugin --profile web add . -w
@@ -70,6 +100,10 @@ The tools below are **invoked by the agent**, not typed by the user. In the chat
70
100
  | `watch_repo root` | Enable automatic refresh: a background poll detects new/changed files (mtime + content hash) and re-indexes only those. Watched roots persist across plugin restarts. |
71
101
  | `memory_stats root` | Show what the store contains: totals (files / entries / experience notes), last index time, and the per-file list sorted by recency. |
72
102
  | `query_memory query` | BM25 search over docs + symbols + experience, optionally query-expanded by the LLM. Returns ranked hits with relative scores, sources, and doc→symbol references. |
103
+ | `list_tasks` | List task records for the project (archived marked). Call first in a new session before continuing work. |
104
+ | `select_task` | Bind the session to a task so its todo list and file reads sync into it. Exact `taskId`, or exact `title` (multiple matches return candidates; no match creates a new task). Pass `title` with `taskId` to rename. Auto-unarchives. |
105
+ | `archive_task` | Archive a task (hide from default views, exclude from capacity, stop syncing). `select_task` restores it. |
106
+ | `/tasks` (typed by the user, not the model) | Shows the task stack: title, step progress, involved files, and which task the current session is bound to. |
73
107
  | `remember problem solution` | Save an experience note. Similar problems supersede instead of duplicating. |
74
108
  | `forget id_or_query` | Delete stale experience notes. |
75
109
 
@@ -184,6 +218,7 @@ These are deliberate scope choices.
184
218
  | `maxChunksPerFile` | 40 | max chunks per document |
185
219
  | `maxFileSizeMb` | 50 | skip documents (incl. PDF) and code files larger than this (MB) |
186
220
  | `maxOutputChars` | 8000 | cap for `query_memory` result text (chars) |
221
+ | `tasklist.enabled` | true | enable TaskBridge auto-sync (task entities from the session todo list and file reads) |
187
222
  | `maxPdfPages` | 1000 | PDF page cap when pages are not otherwise limited |
188
223
  | `llmQueryExpansion` | false | expand queries via `ctx.llm` before BM25 (off by default to save tokens) |
189
224
  | `expansionCount` | 6 | max expansion variants |
@@ -228,7 +263,7 @@ These commands are for **maintaining the plugin code** — regular users do not
228
263
 
229
264
  ```bash
230
265
  npm install
231
- npm test # 157 tests (v0.3.2): chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
266
+ npm test # 166 tests (v0.4.0) + TaskBridge suite (node test/taskbridge.test.mjs, 5)
232
267
  ```
233
268
 
234
269
  ## License
package/README.zh-CN.md CHANGED
@@ -4,24 +4,54 @@
4
4
 
5
5
  [![ci](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml/badge.svg)](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml) [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![npm](https://img.shields.io/npm/v/@yolk_vat-y/dsh-project-memory)](https://www.npmjs.com/package/@yolk_vat-y/dsh-project-memory) [![Listed on dsh-plugin.org](https://dsh-plugin.org/badges/listed.svg)](https://dsh-plugin.org/plugins/00080000/dsh-project-memory) [![Awesome](https://awesome-dsh-plugin.com/badge.svg)](https://awesome-dsh-plugin.com)
6
6
 
7
- 为 [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(dsh)agent 提供持久化的项目记忆。**记住**文档(PDF / Markdown / txt)与代码符号,写入每个工作区独立的存储库,自动维护更新,**回忆**时附源文件引用——文档自动交叉链接到其提及的代码符号。
7
+ 为 [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(dsh)agent 提供持久化的 **项目开发记忆**。针对项目开发,任务清单与读过的文件自动沉淀为跨会话的任务记录,解决上下文失效;文档(PDF/Markdown/txt)与代码符号写入工作区独立存储,文档自动交叉链接至所提及的代码符号;经验笔记(问题 → 方案)自动去重,避免重复踩坑。所有数据按项目落盘,跨会话压缩与交接保留,召回附带 `路径:行号` 可回源核实。零外部依赖(仅 `pdfjs-dist` 提取 PDF 文本),无向量数据库,无原生构建。
8
8
 
9
- > 插件在磁盘上维护一份精简的项目**记忆**,每条记录指向具体的文件与行号;agent 需要快速了解项目时先查**记忆**,无需重读整个项目。
9
+ > 插件在磁盘上维护一份精简的项目**记忆**,每条记录指向具体的文件与行号;agent 需要快速了解项目时先查**记忆**,无需重读整个项目。任务与经验跨会话压缩与交接保持。
10
10
 
11
11
  ## 特性
12
12
 
13
13
  - **文档记忆** — PDF、Markdown、纯文本按块切分并由 LLM 生成摘要,每条记忆携带 `路径:行号` 引用回源文件。
14
14
  - **代码符号记忆** — 通过零依赖的源码扫描器提取函数、类与方法名及完整类型签名(泛型、参数类型、返回类型、重载签名),包含字符串/注释掩码、多行签名续行、Python 缩进感知、类方法上下文,不使用 LLM token。
15
- - **可选 TypeScript 语义增强** — 当用户项目安装了 `typescript`(`npm i -D typescript`),插件自动激活第二层(L2),利用 TS Compiler API 推导返回类型、实例化泛型、提取接口与类型别名、丰富箭头函数签名 —— 全部在优先级队列中异步后台处理(P0:`fs/observed` 读文件瞬间、P1:`watch` 变更后、P2:`index_repo` 批量索引)。结果按文件内容哈希缓存到磁盘,冷启动毫秒级复用。零配置:装 TS 再重启 dsh 即可。完全可选;若无 TS 或设置 `enableTypeScript: false`,回退至 L1 正则提取。
15
+ - **L1 增强正则** — 零依赖正则扫描器现可提取泛型、参数/返回类型、重载、接口、类型别名,产出单行身份签名 `fn(a: A, b: B): R — file.ts:42`。
16
+ - **可选 TypeScript 语义增强 (L2/L3)** — 当用户项目安装了 `typescript`(`npm i -D typescript`),插件自动激活第二层(L2),利用 TS Compiler API 推导返回类型、实例化泛型、提取接口与类型别名、丰富箭头函数签名 —— 全部在优先级队列中异步后台处理(P0:`fs/observed` 读文件瞬间、P1:`watch` 变更后、P2:`index_repo` 批量索引)。结果按文件内容哈希缓存到磁盘(L3),冷启动毫秒级复用。零配置:装 TS 再重启 dsh 即可。完全可选;若无 TS 或设置 `enableTypeScript: false`,回退至 L1 正则提取。
16
17
  - **自动刷新** — `watch_repo` 后台轮询,按内容哈希识别新增或变更文件,仅重记这些文件。
17
18
  - **读到即记忆** — 文件在模型**实际读取的瞬间**被记忆(监听 `fs/observed`),记忆是正常工作的副产品,而非额外的一次全量扫描。从未读过的文件不会被记忆。项目根通过标记(`.git`、`package.json` 等)、README 加源码目录、或兜底到文件所在目录逐级识别。
18
19
  - **文档 ↔ 代码交叉链接** — 文档提及某符号时记录为 `reference`;查询符号时同时带出描述该符号的文档。
19
20
  - **BM25 记忆召回** — 对文档、符号与经验笔记进行排序召回,可选 LLM 查询扩展以应对表述不一致。**CJK 增强**:精确短语乘法加分(3+ 字短语在标题/关键词命中 ×1.5)、同义词表(如 数据库连接池 ↔ 连接池 ↔ DB pool)、CJK 感知的文档↔符号链接边界。
21
+ - **blindSpots 感知召回** — 文档摘要携带 `blindSpots` 字段(明确说明摘要未覆盖的内容)。查询命中盲区时,`query_memory` 追加提示引导模型去读原文,防止半截摘要误导。
20
22
  - **经验笔记** — 记录问题 → 方案;相似问题覆盖而非重复;笔记仅在检索命中时返回。笔记数量有界:容量随项目规模伸缩(钳制在 100–2000),超限时淘汰最旧的笔记。**覆盖阈值收紧为双向 0.7 重叠**(原 0.6);**经验 `problem` 字段现参与 CJK 短语加分**,提升长尾问句召回。
23
+ - **流式 TF + IDF 缓存** — 查询路径按存储版本缓存 IDF(词逆频率);命中时单次流式遍历 20k 条目仅需 ~8 ms(5k 文件) / ~1 ms(1k 文件),零中间对象;写入路径仅 O(1) 版本号递增。
21
24
  - **无锁同步事务** — 不采用锁:所有写入(index / watch / remember / forget / watch_repo)统一走同步事务 `store.commit(fn)`,fn 成功后才一次落盘;JS 单线程事件循环保证事务间不交错,`remember`/`forget` 不会被 watch 重索引阻塞排队。多实例并发写入同一项目存储时,得益于 CAS 幂等更新与原子提交,自然具备幂等性,无数据损坏风险。
25
+ - **TaskBridge:跨会话开发任务** — 监听会话内宿主 `todo_write` 维护的任务清单与 `tool/call` 读文件:进度快照(steps)与触碰文件自动同步进跨会话的任务实体。未绑定会话写 todo 时自动建档。新会话通过 `list_tasks` → `select_task`(绑定/改名/解归档)续接;`query_memory` 新增 `type:'task'`,`type:'all'` 结果尾部附任务计数提示。用户侧 `/tasks` 命令展示任务栈、步骤进度、涉及文件与当前会话绑定。标题由模型经 `select_task(title=…)` 命名(回退:取消息最后一个「:」后的任务段)。容量随项目体积自适应(fileCount/20,clamp 5–100)。存储:`.dsh-project-memory/tasks.json` + `binding.json`。自动同步需含会话事件与 `todo_write` 的 dsh(0.1.2-alpha.x 实测);旧宿主下降级为纯记录。
22
26
  - **依赖极简** — 纯 JavaScript;唯一运行时依赖是 `pdfjs-dist`(PDF 文本提取),无需原生构建。
23
27
  - **开销可忽略** — 纯进程内操作;冷启动 <100 ms(5k 文件),典型项目查询中位数 2–3 ms(p99 < 7 ms);瓶颈在 LLM 摘要与 PDF 解析,插件本身不阻塞。
24
28
 
29
+ ## 性能
30
+
31
+ ### 合成基准测试(隔离环境,Node 24,Linux)
32
+
33
+ | 场景 | 规模 | 实测 |
34
+ |------|------|------|
35
+ | 批量冷记忆构建 | 5,000 文件 / 20k 条目 | 353 ms |
36
+ | 冷加载 | 5,000 文件 | 82 ms |
37
+ | 热路径懒记忆 | 单文件重记忆+落盘 | 中位数 2.3 ms / 最大 4.0 ms (5k) |
38
+ | query_memory (缓存命中) | 5k 文件 / 20k 条目 | 中位数 9.3 ms / p95 12.6 ms |
39
+ | query_memory (缓存命中) | 1k 文件 / 4k 条目 | 中位数 1.0 ms / p95 2.0 ms |
40
+ | 批量冷记忆构建 | 10,000 文件 / 40k 条目 | 637 ms |
41
+ | 冷加载 | 10,000 文件 | 144 ms |
42
+ | 热路径懒记忆 | 单文件重记忆+落盘 | 中位数 4.5 ms / 最大 10.2 ms (10k) |
43
+
44
+ > 合成基准:生成代码(~8 符号/文件),Node 24,Linux 文件系统,SSD。测量纯索引开销,不含 LLM 调用。query_memory 基准使用 IDF 缓存 + 预计算 searchText;写入后首次查询重建 IDF(~150 ms),后续查询命中缓存。
45
+
46
+ ### 真实项目存储体积
47
+
48
+ | 项目 | 文件数 | 条目数 | 存储体积 | 单条目 |
49
+ |------|--------|--------|----------|--------|
50
+ | Java Spring Boot 后端 | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
51
+ | Vue 3 + Vite 前端 | 289 | 2,141 | 1.0 MB | ~0.5 KB |
52
+
53
+ > 真实项目(Java + Vue),测试于 Linux 文件系统(Node 24)。真实项目单条目体积小于合成基准,因符号密度更低、声明行更短。
54
+
25
55
  ## 工作原理
26
56
 
27
57
  设计遵循四个原则:
@@ -37,7 +67,7 @@
37
67
 
38
68
  ## 安装
39
69
 
40
- 实测覆盖 dsh **0.1.0-rc.7 至 0.1.2-alpha.2**。插件仅依赖通过 peerDependencies 声明的稳定公共 API(`defineTool`、`llm.stream`、`Schema`),保证与后续 rc 版本无需改动即兼容。
70
+ 实测覆盖 dsh **0.1.0-rc.7 至 0.1.2-alpha.3**。插件仅依赖通过 peerDependencies 声明的稳定公共 API(`defineTool`、`llm.stream`、`Schema`),保证与后续 rc/alpha 版本无需改动即兼容。
41
71
 
42
72
  ```bash
43
73
  cd dsh-project-memory && dsh plugin --profile web add . -w
@@ -70,6 +100,10 @@ dsh plugin --profile web add /path/to/dsh-project-memory.tgz
70
100
  | `watch_repo root` | 启用自动刷新:后台轮询检测新增/变更文件(mtime + 内容哈希),仅重抽这些文件。监听的项目在插件重启后自动恢复。 |
71
101
  | `memory_stats root` | 查看记忆库内容:总量(文件 / 条目 / 经验笔记)、最近索引时间,以及按时间排序的逐文件清单。 |
72
102
  | `query_memory query` | 对文档、符号、经验执行 BM25 检索,可选 LLM 查询扩展。返回带相对分数(0-100)、引用与文档→符号链接的排序结果。 |
103
+ | `list_tasks` | 列出本项目任务记录(含归档,带标记)。新会话/续接前先调用。 |
104
+ | `select_task` | 将会话绑定到某任务(此后 todo 清单与读文件同步进该任务)。按 `taskId` 精确绑定,或按 `title` 完全匹配(多个同名返回候选;无则新建)。带 title 可改名;自动解归档。 |
105
+ | `archive_task` | 归档任务(隐藏默认视图、不占容量、停止同步)。`select_task` 可恢复。 |
106
+ | `/tasks`(用户输入,不经模型) | 展示任务栈:标题、步骤进度、涉及文件、当前会话绑定哪套任务。 |
73
107
  | `remember problem solution` | 保存经验笔记。相似问题覆盖而非重复。 |
74
108
  | `forget id_or_query` | 删除过期经验笔记。 |
75
109
 
@@ -184,6 +218,7 @@ v0.2.0 之前创建的库(单文件 `entries.json` / `index.json`)在首次
184
218
  | `maxChunksPerFile` | 40 | 每文档最大块数 |
185
219
  | `maxFileSizeMb` | 50 | 大于该值(MB)的文档(含 PDF)/代码文件跳过 |
186
220
  | `maxOutputChars` | 8000 | `query_memory` 返回文本上限(字符) |
221
+ | `tasklist.enabled` | true | 启用 TaskBridge 自动同步(由会话 todo 清单与文件读取沉淀任务实体) |
187
222
  | `maxPdfPages` | 1000 | 未另行限制时 PDF 的页数上限 |
188
223
  | `llmQueryExpansion` | false | BM25 检索前通过 `ctx.llm` 扩展查询(默认关闭,节省 token) |
189
224
  | `expansionCount` | 6 | 扩展变体上限 |
@@ -228,7 +263,7 @@ dsh web --patch ./config.yml
228
263
 
229
264
  ```bash
230
265
  npm install
231
- npm test # 157 tests (v0.3.2):chunker / symbols / store / tools / BM25 / links / watch / lazy / config / dump / concurrency / restore / size limit
266
+ npm test # 166 tests (v0.4.0) + TaskBridge 套件(node test/taskbridge.test.mjs,5 项)
232
267
  ```
233
268
 
234
269
  ## 许可证
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@yolk_vat-y/dsh-project-memory",
3
- "version": "0.3.3",
3
+ "version": "0.4.0",
4
4
  "description": "Persistent project memory for dsh agents: index docs (PDF/Markdown/text) and code symbols into a searchable per-workspace store, recall them with cited sources, and keep experience entries (problems -> solutions) searchable on demand.",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
@@ -0,0 +1,54 @@
1
+ // /tasks 用户命令:不经模型,直接展示本项目任务总览(非归档 + 归档计数 + 当前会话绑定)。
2
+ import { memoryRootFor } from '../util/fs.js'
3
+ import { ProjectMemoryStore } from '../store.js'
4
+ import { projectRootFor } from '../setup/taskbridge.js'
5
+
6
+ function timeAgo(iso) {
7
+ const diff = Date.now() - new Date(iso).getTime()
8
+ const m = Math.floor(diff / 60000)
9
+ if (m < 1) return '刚刚'
10
+ if (m < 60) return `${m}分钟前`
11
+ const h = Math.floor(m / 60)
12
+ if (h < 24) return `${h}小时前`
13
+ return `${Math.floor(h / 24)}天前`
14
+ }
15
+
16
+ export function tasksCommandDefinition(config) {
17
+ return {
18
+ name: 'tasks',
19
+ description: '查看本项目的任务清单(几套任务、进度、涉及文件、当前绑定)',
20
+ handler: (invocation) => {
21
+ try {
22
+ const cwd = invocation?.agent?.session?.header?.cwd
23
+ const root = projectRootFor(cwd)
24
+ const store = new ProjectMemoryStore(memoryRootFor(root, config.memoryDir)).load()
25
+ const tasks = store.getTasks()
26
+ const active = tasks.filter((t) => !t.archived)
27
+ const archived = tasks.length - active.length
28
+ const sid = invocation?.agent?.session?.id
29
+ const boundId = sid ? store.getBoundTaskId(sid) : null
30
+ const bound = boundId ? tasks.find((t) => t.id === boundId) : null
31
+
32
+ if (!active.length) {
33
+ return { kind: 'success', text: `项目 ${root}\n任务记录: 0 套${archived ? `(归档 ${archived})` : ''}。让模型开始干活并维护 todo 清单后会自动建档。` }
34
+ }
35
+ const lines = [`项目 ${root}`, `任务: ${active.length} 套${archived ? `(归档 ${archived})` : ''}`, '']
36
+ for (const t of active) {
37
+ const done = (t.steps || []).filter((s) => s.status === 'completed').length
38
+ const total = (t.steps || []).length
39
+ const progress = total ? `${done}/${total}` : '无步骤'
40
+ const inProgress = (t.steps || []).find((s) => s.status === 'in_progress')
41
+ const step = inProgress ? ` · 当前: ${inProgress.content}` : ''
42
+ const marker = bound && bound.id === t.id ? '(本会话绑定)' : ''
43
+ const files = (t.files || []).slice(0, 8)
44
+ const fileLine = files.length ? `\n 文件: ${files.join(', ')}${t.files.length > 8 ? ' …' : ''}` : ''
45
+ lines.push(`● ${t.title}${marker} 步骤 ${progress}${step} · ${timeAgo(t.lastActiveAt || t.updatedAt)}${fileLine}`)
46
+ }
47
+ lines.push('', '续接/改名/归档:直接告诉模型(list_tasks / select_task / archive_task)。')
48
+ return { kind: 'success', text: lines.join('\n') }
49
+ } catch (err) {
50
+ return { kind: 'error', text: `[tasks] ${err?.message || err}` }
51
+ }
52
+ },
53
+ }
54
+ }
package/src/index.js CHANGED
@@ -9,6 +9,9 @@ import { statsTool } from './tools/stats.js'
9
9
  import { WatchManager } from './watch.js'
10
10
  import { setupLazyIndexing } from './lazy.js'
11
11
  import { initTypeScript } from './enhancer.js'
12
+ import { setupTaskbridge } from './setup/taskbridge.js'
13
+ import { listTasksTool, selectTaskTool, archiveTaskTool } from './tools/task-tools.js'
14
+ import { tasksCommandDefinition } from './commands/tasks.js'
12
15
 
13
16
  export const name = 'dsh-project-memory'
14
17
  export const inject = ['llm', 'tools']
@@ -28,6 +31,9 @@ export const Config = Schema.object({
28
31
  watchInterval: Schema.number().default(15),
29
32
  tsPath: Schema.string(),
30
33
  enableTypeScript: Schema.boolean().default(true),
34
+ tasklist: Schema.object({
35
+ enabled: Schema.boolean().default(true),
36
+ }).default({}),
31
37
  })
32
38
 
33
39
  export function apply(ctx, config) {
@@ -49,6 +55,21 @@ export function apply(ctx, config) {
49
55
  setupLazyIndexing(ctx, config, watchManager)
50
56
  }
51
57
 
58
+ // TaskBridge:任务实体 + 宿主 todo 同步
59
+ setupTaskbridge(ctx, config)
60
+ ctx.tools.register(listTasksTool(config))
61
+ ctx.tools.register(selectTaskTool(config))
62
+ ctx.tools.register(archiveTaskTool(config))
63
+
64
+ // /tasks 用户命令(宿主 commands 服务存在时注册,feature-detect 降级)
65
+ try {
66
+ ctx.inject(['commands'], (commandsCtx) => {
67
+ commandsCtx.commands.register(tasksCommandDefinition(config))
68
+ })
69
+ } catch (err) {
70
+ console.error(`[dsh-project-memory] /tasks registration skipped: ${err.message}`)
71
+ }
72
+
52
73
  ctx.tools.register(indexDocTool(ctx, config))
53
74
  ctx.tools.register(indexRepoTool(ctx, config))
54
75
  ctx.tools.register(queryMemoryTool(ctx, config))
@@ -0,0 +1,169 @@
1
+ // TaskBridge 运行时:监听会话事件,把宿主 todo/工具调用同步进任务实体。
2
+ // 事实依据:dsh session/event 签名 (session, event);todo/write data.todos;
3
+ // tool/call data.arguments 为 JSON 字符串;fs 工具名 read/write/edit/read_image,参数 file_path。
4
+ import path from 'node:path'
5
+ import { createHash } from 'node:crypto'
6
+ import { memoryRootFor } from '../util/fs.js'
7
+ import { findProjectRoot } from '../lazy.js'
8
+ import { ProjectMemoryStore } from '../store.js'
9
+
10
+ const FS_FILE_TOOLS = new Set(['read', 'write', 'edit', 'read_image'])
11
+ const MAX_FILES_PER_TASK = 100
12
+ const TITLE_MAX = 24
13
+
14
+ export function hash8(text) {
15
+ return createHash('sha256').update(String(text)).digest('hex').slice(0, 8)
16
+ }
17
+
18
+ export function slugifyTitle(text) {
19
+ return String(text || '')
20
+ .toLowerCase()
21
+ .replace(/[^\w\u4e00-\u9fff]+/g, '-')
22
+ .replace(/^-+|-+$/g, '')
23
+ .slice(0, TITLE_MAX)
24
+ }
25
+
26
+ export function genTaskId(projectRoot, title) {
27
+ const slug = slugifyTitle(title) || 'task'
28
+ return `tsk_${hash8(projectRoot)}_${slug}_${Date.now()}`
29
+ }
30
+
31
+ /** 项目根推导:findProjectRoot 期望文件路径,传目录会从父级起跳,故用目录内探针路径。 */
32
+ export function projectRootFor(cwd) {
33
+ const base = cwd || process.cwd()
34
+ return findProjectRoot(path.join(base, '__taskbridge__.probe'))
35
+ }
36
+
37
+ export function taskStoreFor(cwd, config) {
38
+ const root = projectRootFor(cwd)
39
+ return { root, store: new ProjectMemoryStore(memoryRootFor(root, config.memoryDir)).load() }
40
+ }
41
+
42
+ /** 把工具参数里的文件路径归一化为项目相对路径;项目外返回 null。 */
43
+ export function normalizeRelFile(root, file) {
44
+ if (typeof file !== 'string' || !file) return null
45
+ const abs = path.isAbsolute(file) ? file : path.resolve(root, file)
46
+ const rel = path.relative(root, abs)
47
+ if (!rel || rel.startsWith('..') || path.isAbsolute(rel)) return null
48
+ return rel.split(path.sep).join('/')
49
+ }
50
+
51
+ /** content 可能是字符串或 ContentBlock[],统一取首段文本。 */
52
+ export function firstTextOf(content) {
53
+ if (typeof content === 'string') return content
54
+ if (Array.isArray(content)) {
55
+ for (const b of content) {
56
+ if (b && typeof b.text === 'string' && b.text.trim()) return b.text
57
+ }
58
+ }
59
+ return ''
60
+ }
61
+
62
+ function pickTitle(meta, todos) {
63
+ const firstHuman = meta?.firstHuman?.trim()
64
+ if (firstHuman) {
65
+ // 用户消息常带指令前缀("用 todo_write 规划:…"),优先取最后一个":"后的任务段
66
+ const idx = Math.max(firstHuman.lastIndexOf(':'), firstHuman.lastIndexOf(':'))
67
+ const seg = idx >= 0 ? firstHuman.slice(idx + 1).trim() : ''
68
+ if (seg.length >= 2 && seg.length <= 48) return seg
69
+ return firstHuman.slice(0, 48)
70
+ }
71
+ const firstTodo = todos?.[0]?.content
72
+ if (firstTodo && firstTodo.trim()) return firstTodo.trim().slice(0, 48)
73
+ return 'Untitled Task'
74
+ }
75
+
76
+ /**
77
+ * 单条会话事件处理(导出便于测试):user/message 记首条真人文本;
78
+ * todo/write → 已绑定则覆盖 steps,未绑定则自动新建任务并绑定;
79
+ * tool/call(fs 工具)→ 绑定任务 files 并集。
80
+ */
81
+ export function onSessionEvent(config, session, event, meta) {
82
+ const sessionId = session?.id
83
+ const type = event?.type
84
+ if (!sessionId || !type) return
85
+
86
+ if (type === 'user/message') {
87
+ if (!meta.has(sessionId) && event.data?.source?.kind === 'user') {
88
+ const text = firstTextOf(event.data?.content)
89
+ if (text) meta.set(sessionId, { firstHuman: text })
90
+ }
91
+ return
92
+ }
93
+
94
+ const { root, store } = taskStoreFor(session?.header?.cwd, config)
95
+ const now = new Date().toISOString()
96
+
97
+ if (type === 'todo/write') {
98
+ const todos = event.data?.todos
99
+ if (!Array.isArray(todos)) return
100
+ store.commit((s) => {
101
+ let task = s.getBoundTaskId(sessionId) ? s.getTask(s.getBoundTaskId(sessionId)) : null
102
+ if (!task || task.archived) {
103
+ const title = pickTitle(meta.get(sessionId), todos)
104
+ task = {
105
+ id: genTaskId(root, title),
106
+ projectHash: hash8(root),
107
+ projectRoot: root,
108
+ title,
109
+ steps: null,
110
+ files: [],
111
+ archived: false,
112
+ lastSessionId: sessionId,
113
+ createdAt: now,
114
+ updatedAt: now,
115
+ lastActiveAt: now,
116
+ }
117
+ s.addTask(task)
118
+ s.setBinding(sessionId, task.id)
119
+ }
120
+ task.steps = todos
121
+ task.updatedAt = now
122
+ task.lastActiveAt = now
123
+ task.lastSessionId = sessionId
124
+ s._dirtyTasks = true
125
+ })
126
+ return
127
+ }
128
+
129
+ if (type === 'tool/call') {
130
+ const taskId = store.getBoundTaskId(sessionId)
131
+ if (!taskId) return
132
+ const data = event.data
133
+ if (!data?.name || !FS_FILE_TOOLS.has(data.name)) return
134
+ let args
135
+ try {
136
+ args = typeof data.arguments === 'string' ? JSON.parse(data.arguments) : data.arguments
137
+ } catch {
138
+ return
139
+ }
140
+ const raw = typeof args?.file_path === 'string' ? args.file_path : null
141
+ if (!raw) return
142
+ const rel = normalizeRelFile(root, raw)
143
+ if (!rel) return
144
+ store.commit((s) => {
145
+ const task = s.getTask(taskId)
146
+ if (!task || task.archived) return
147
+ if (!task.files.includes(rel)) {
148
+ task.files.push(rel)
149
+ if (task.files.length > MAX_FILES_PER_TASK) task.files.shift()
150
+ }
151
+ task.lastActiveAt = now
152
+ s._dirtyTasks = true
153
+ })
154
+ }
155
+ }
156
+
157
+ /** 插件接线:订阅 session/event;卸载时清会话元数据。 */
158
+ export function setupTaskbridge(ctx, config) {
159
+ if (config.tasklist?.enabled === false) return
160
+ const meta = new Map()
161
+ ctx.on('session/event', (session, event) => {
162
+ try {
163
+ onSessionEvent(config, session, event, meta)
164
+ } catch (err) {
165
+ console.error('[dsh-project-memory] taskbridge:', err?.message)
166
+ }
167
+ })
168
+ ctx.effect(() => () => meta.clear())
169
+ }
package/src/store.js CHANGED
@@ -1,12 +1,14 @@
1
1
  import { createHash, randomUUID } from 'node:crypto'
2
2
  import { mkdirSync, readFileSync, readdirSync, renameSync, statSync, unlinkSync, writeFileSync } from 'node:fs'
3
3
  import path from 'node:path'
4
- import { rankEntries, rankExperience, tokenize } from './util/search.js'
4
+ import { rankEntries, rankExperience, tokenize, tokenizeRaw, extractCjkPhrases, makeSearchText } from './util/search.js'
5
5
 
6
6
  const FORMAT_FILE = 'format.json'
7
7
  const INDEX_FILE = 'index.json'
8
8
  const ENTRIES_FILE = 'entries.json'
9
9
  const EXPERIENCE_FILE = 'experience.json'
10
+ const TASKS_FILE = 'tasks.json'
11
+ const BINDING_FILE = 'binding.json'
10
12
  const WATCH_FILE = 'watch.json'
11
13
  const SHARDS_DIR = 'shards'
12
14
 
@@ -58,12 +60,18 @@ export class ProjectMemoryStore {
58
60
  this.files = {}
59
61
  this.entries = {}
60
62
  this.experience = []
63
+ this.tasks = []
64
+ this.binding = {}
61
65
  this.watchlist = []
62
66
  this._dirtyShards = new Set()
63
67
  this._removedShards = new Set()
64
68
  this._dirtyExperience = false
69
+ this._dirtyTasks = false
70
+ this._dirtyBinding = false
65
71
  this._dirtyWatch = false
66
72
  this._formatWritten = false
73
+ this._version = 0
74
+ this._idfCache = null
67
75
  }
68
76
 
69
77
  load() {
@@ -132,6 +140,8 @@ export class ProjectMemoryStore {
132
140
  this.entries[shard.relPath] = shard.entries || []
133
141
  }
134
142
  this.experience = loadJson(path.join(this.dir, EXPERIENCE_FILE), [])
143
+ this.tasks = loadJson(path.join(this.dir, TASKS_FILE), [])
144
+ this.binding = loadJson(path.join(this.dir, BINDING_FILE), {})
135
145
  this.watchlist = loadJson(path.join(this.dir, WATCH_FILE), [])
136
146
  this._formatWritten = existsSafe(path.join(this.dir, FORMAT_FILE))
137
147
  }
@@ -185,10 +195,52 @@ export class ProjectMemoryStore {
185
195
  writeJsonAtomic(path.join(this.dir, EXPERIENCE_FILE), this.experience)
186
196
  this._dirtyExperience = false
187
197
  }
198
+ if (this._dirtyTasks) {
199
+ writeJsonAtomic(path.join(this.dir, TASKS_FILE), this.tasks)
200
+ this._dirtyTasks = false
201
+ }
202
+ if (this._dirtyBinding) {
203
+ writeJsonAtomic(path.join(this.dir, BINDING_FILE), this.binding)
204
+ this._dirtyBinding = false
205
+ }
188
206
  if (this._dirtyWatch) {
189
207
  writeJsonAtomic(path.join(this.dir, WATCH_FILE), this.watchlist)
190
208
  this._dirtyWatch = false
191
209
  }
210
+ this._version++
211
+ this._idfCache = null
212
+ }
213
+
214
+ getIdfCache() {
215
+ if (this._idfCache && this._idfCache.version === this._version) {
216
+ return this._idfCache.idf
217
+ }
218
+ const idf = this._rebuildIdf()
219
+ this._idfCache = { version: this._version, idf }
220
+ return idf
221
+ }
222
+
223
+ _rebuildIdf() {
224
+ const entries = this.allEntries()
225
+ const N = entries.length
226
+ if (N === 0) return {}
227
+ const df = {}
228
+ for (const entry of entries) {
229
+ const text = entry.title || ''
230
+ const keywords = (entry.keywords || []).join(' ')
231
+ const summary = entry.summary || ''
232
+ const combined = `${text} ${text} ${text} ${text} ${text} ${keywords} ${summary}`.toLowerCase()
233
+ const terms = tokenizeRaw(combined)
234
+ const seen = new Set(terms)
235
+ for (const t of seen) {
236
+ df[t] = (df[t] || 0) + 1
237
+ }
238
+ }
239
+ const idf = {}
240
+ for (const [t, df_t] of Object.entries(df)) {
241
+ idf[t] = Math.log(1 + (N - df_t + 0.5) / (df_t + 0.5))
242
+ }
243
+ return idf
192
244
  }
193
245
 
194
246
  addWatch(root) {
@@ -219,7 +271,8 @@ export class ProjectMemoryStore {
219
271
 
220
272
  setEntries(relPath, entries) {
221
273
  if (entries.length) {
222
- this.entries[relPath] = entries
274
+ const enriched = entries.map((e) => ({ ...e, searchText: makeSearchText(e) }))
275
+ this.entries[relPath] = enriched
223
276
  } else {
224
277
  delete this.entries[relPath]
225
278
  }
@@ -318,11 +371,83 @@ export class ProjectMemoryStore {
318
371
  files: Object.keys(this.files).length,
319
372
  entries: this.allEntries().length,
320
373
  experience: this.experience.length,
374
+ tasks: this.tasks.length,
375
+ tasksActive: this.tasks.filter((t) => !t.archived).length,
376
+ tasksArchived: this.tasks.filter((t) => t.archived).length,
377
+ }
378
+ }
379
+
380
+ // ---- TaskBridge: task entities + per-session binding ----
381
+
382
+ getTasks() {
383
+ return this.tasks
384
+ }
385
+
386
+ getTask(id) {
387
+ return this.tasks.find((t) => t.id === id)
388
+ }
389
+
390
+ addTask(task) {
391
+ this.tasks.push(task)
392
+ this._dirtyTasks = true
393
+ }
394
+
395
+ updateTask(id, updates) {
396
+ const task = this.tasks.find((t) => t.id === id)
397
+ if (!task) return false
398
+ Object.assign(task, updates, { updatedAt: new Date().toISOString() })
399
+ this._dirtyTasks = true
400
+ return true
401
+ }
402
+
403
+ removeTask(id) {
404
+ const idx = this.tasks.findIndex((t) => t.id === id)
405
+ if (idx === -1) return false
406
+ this.tasks.splice(idx, 1)
407
+ for (const sid of Object.keys(this.binding)) {
408
+ if (this.binding[sid] === id) delete this.binding[sid]
409
+ }
410
+ this._dirtyTasks = true
411
+ this._dirtyBinding = true
412
+ return true
413
+ }
414
+
415
+ setBinding(sessionId, taskId) {
416
+ if (!sessionId) return
417
+ this.binding[sessionId] = taskId
418
+ this._dirtyBinding = true
419
+ }
420
+
421
+ getBoundTaskId(sessionId) {
422
+ return sessionId ? this.binding[sessionId] : undefined
423
+ }
424
+
425
+ removeBinding(sessionId) {
426
+ if (sessionId && sessionId in this.binding) {
427
+ delete this.binding[sessionId]
428
+ this._dirtyBinding = true
429
+ }
430
+ }
431
+
432
+ /** 每项目任务数上限随项目体积自适应;超限按 lastActiveAt 归档最旧(只统计非归档)。 */
433
+ pruneTasks() {
434
+ const fileCount = Object.keys(this.files).length
435
+ const maxTasks = Math.min(100, Math.max(5, Math.floor(fileCount / 20)))
436
+ const nonArchived = this.tasks.filter((t) => !t.archived)
437
+ if (nonArchived.length <= maxTasks) return 0
438
+ const ts = (t) => new Date(t.lastActiveAt || t.updatedAt || t.createdAt || 0).getTime()
439
+ const sorted = [...nonArchived].sort((a, b) => ts(a) - ts(b))
440
+ const toArchive = sorted.slice(0, nonArchived.length - maxTasks)
441
+ for (const task of toArchive) {
442
+ task.archived = true
443
+ this._dirtyTasks = true
321
444
  }
445
+ return toArchive.length
322
446
  }
323
447
 
324
448
  commit(fn) {
325
449
  const result = fn(this)
450
+ this.pruneTasks()
326
451
  this.save()
327
452
  return result
328
453
  }
@@ -3,7 +3,7 @@ import path from 'node:path'
3
3
  import { memoryRootFor, resolveIndexRoot } from '../util/fs.js'
4
4
  import { ProjectMemoryStore, storeOverview } from '../store.js'
5
5
  import { expandQuery } from '../llm.js'
6
- import { rankEntriesMergedScored, rankExperienceScored } from '../util/search.js'
6
+ import { rankEntriesMergedScored, rankExperienceScored, rankEntriesStreaming } from '../util/search.js'
7
7
  import { truncate } from '../util/text.js'
8
8
 
9
9
  function toAbs(root, rel) {
@@ -30,8 +30,8 @@ export function queryMemoryTool(ctx, config) {
30
30
  },
31
31
  type: {
32
32
  type: 'string',
33
- enum: ['all', 'doc', 'symbol', 'experience'],
34
- description: 'Which memory layer to search. Default "all".',
33
+ enum: ['all', 'doc', 'symbol', 'experience', 'task'],
34
+ description: 'Which memory layer to search. Default "all". "task" searches task records (title/steps/files).',
35
35
  },
36
36
  limit: {
37
37
  type: 'number',
@@ -56,10 +56,12 @@ export function queryMemoryTool(ctx, config) {
56
56
  if (e.type === 'symbol') symbolById.set(e.id, e)
57
57
  }
58
58
 
59
+ const idf = store.getIdfCache()
60
+
59
61
  const lines = []
60
62
  if (type === 'all' || type === 'doc' || type === 'symbol') {
61
63
  const pool = type === 'all' ? store.allEntries() : store.allEntries().filter((e) => e.type === type)
62
- const scored = rankEntriesMergedScored(pool, queries, limit)
64
+ const scored = rankEntriesStreaming(pool, queries, idf, limit)
63
65
  if (scored.length) {
64
66
  const top = scored[0].score || 1
65
67
  lines.push(`## Memory (${type === 'all' ? 'docs + symbols' : type})`)
@@ -100,6 +102,31 @@ export function queryMemoryTool(ctx, config) {
100
102
  }
101
103
  }
102
104
 
105
+ // TaskBridge: type:'task' 专门查任务记录;type:'all' 尾部附一行任务计数提示
106
+ const tasks = store.tasks || []
107
+ if (type === 'task') {
108
+ const q = (queries[0] || '').toLowerCase()
109
+ const matched = tasks
110
+ .filter((t) => !t.archived && (t.title.toLowerCase().includes(q) || (t.steps || []).some((s) => s.text?.toLowerCase().includes(q)) || (t.files || []).some((f) => f.toLowerCase().includes(q))))
111
+ .slice(0, limit)
112
+ if (!matched.length) {
113
+ return `任务记录: 0 套匹配 "${args.query}"(list_tasks 查看全部,select_task 续做)`
114
+ }
115
+ for (const t of matched) {
116
+ const done = (t.steps || []).filter((s) => s.status === 'completed').length
117
+ const total = (t.steps || []).length
118
+ const stepsText = (t.steps || []).length
119
+ ? (t.steps || []).map((s) => `- [${s.status === 'completed' ? 'x' : s.status === 'in_progress' ? '*' : ' '}] ${s.text}`).join('\n')
120
+ : '(无步骤)'
121
+ const files = (t.files || []).slice(0, 8).join(', ')
122
+ lines.push(`### ${t.title} (${done}/${total} 完成)\n${stepsText}\n- 文件: ${files || '无'}`)
123
+ }
124
+ return truncate(lines.join('\n\n'), config.maxOutputChars)
125
+ }
126
+ if (type === 'all' && tasks.length) {
127
+ lines.push(`任务记录: ${tasks.length} 套(list_tasks 查看,select_task 续做)`)
128
+ }
129
+
103
130
  if (!lines.length) {
104
131
  const overview = storeOverview(store)
105
132
  const hint =
@@ -0,0 +1,164 @@
1
+ import { defineTool } from '@deepseek-ai/dsh-tools'
2
+ import { memoryRootFor, resolveIndexRoot } from '../util/fs.js'
3
+ import { ProjectMemoryStore } from '../store.js'
4
+ import { truncate } from '../util/text.js'
5
+ import { genTaskId, hash8 } from '../setup/taskbridge.js'
6
+ import { createHash } from 'node:crypto'
7
+
8
+ function sessionIdOf(exec) {
9
+ return exec?.agent?.session?.id || exec?.ctx?.session?.id
10
+ }
11
+
12
+ function timeAgo(iso) {
13
+ const diff = Date.now() - new Date(iso).getTime()
14
+ const m = Math.floor(diff / 60000)
15
+ if (m < 1) return '刚刚'
16
+ if (m < 60) return `${m}分钟前`
17
+ const h = Math.floor(m / 60)
18
+ if (h < 24) return `${h}小时前`
19
+ return `${Math.floor(h / 24)}天前`
20
+ }
21
+
22
+ function summaryOf(task) {
23
+ const done = (task.steps || []).filter((s) => s.status === 'completed').length
24
+ const total = (task.steps || []).length
25
+ const files = (task.files || []).length
26
+ const mark = task.archived ? ' [归档]' : ''
27
+ return `### ${task.title}${mark}\n- 进度: ${done}/${total} 完成\n- 文件: ${files} 个\n- 最后活动: ${timeAgo(task.lastActiveAt || task.updatedAt)}`
28
+ }
29
+
30
+ export function listTasksTool(config) {
31
+ return defineTool({
32
+ name: 'list_tasks',
33
+ description:
34
+ 'List all task records in this project (archived included, marked). Call this first in a new session or before continuing work, to see what tasks exist and pick one with select_task.',
35
+ parameters: {
36
+ root: { type: 'string', description: 'Project root. Defaults to current working directory.' },
37
+ },
38
+ output: { schema: { type: 'string' }, render: (_a, v) => [{ type: 'text', text: v }] },
39
+ async execute(args, exec) {
40
+ const root = resolveIndexRoot(exec, args.root)
41
+ const store = new ProjectMemoryStore(memoryRootFor(root, config.memoryDir)).load()
42
+ const tasks = store.getTasks()
43
+ if (!tasks.length) {
44
+ return '任务记录: 0 套(先做一步再建清单,或用 select_task 新建)'
45
+ }
46
+ const lines = ['## 任务记录']
47
+ for (const task of tasks) lines.push(summaryOf(task))
48
+ lines.push(`\n任务记录: ${tasks.length} 套(list_tasks 查看,select_task 续做/切换)`)
49
+ return truncate(lines.join('\n\n'), config.maxOutputChars)
50
+ },
51
+ })
52
+ }
53
+
54
+ export function selectTaskTool(config) {
55
+ return defineTool({
56
+ name: 'select_task',
57
+ description:
58
+ 'Bind the current session to a task record so its todo list and read files sync into it. When starting NEW work, call select_task(title="任务名") FIRST, then maintain your plan with todo_write — you choose the task title (short, no need to restate the user message). taskId: exact id from list_tasks (auto-unarchives; pass title too to rename). title: exact match only; multiple matches return candidates; no match creates a new task. Returns a task card (title, steps, files) to rebuild the session todo from. Call list_tasks first in a new session.',
59
+ parameters: {
60
+ taskId: { type: 'string', description: 'Exact task id, e.g. tsk_ab12cd34_...' },
61
+ title: { type: 'string', description: 'Exact task title; creates a new task when absent' },
62
+ root: { type: 'string', description: 'Project root. Defaults to current working directory.' },
63
+ },
64
+ output: { schema: { type: 'string' }, render: (_a, v) => [{ type: 'text', text: v }] },
65
+ async execute(args, exec) {
66
+ const root = resolveIndexRoot(exec, args.root)
67
+ const store = new ProjectMemoryStore(memoryRootFor(root, config.memoryDir)).load()
68
+ const sid = sessionIdOf(exec)
69
+ const now = new Date().toISOString()
70
+ const error = (msg) => truncate(JSON.stringify({ success: false, error: msg }), config.maxOutputChars)
71
+
72
+ if (!args.taskId && !args.title) {
73
+ return error('需要 taskId 或 title(先 list_tasks;未绑定会话出现 todo 时会自动新建任务)')
74
+ }
75
+
76
+ let task = null
77
+ let hint = ''
78
+ if (args.taskId) {
79
+ task = store.getTask(args.taskId)
80
+ if (!task) return error(`Task not found: ${args.taskId}`)
81
+ if (task.archived) {
82
+ task.archived = false
83
+ hint = '已解归档'
84
+ }
85
+ if (args.title && args.title !== task.title) {
86
+ task.title = args.title // 改名(自动标题可能带指令前缀,续接时可顺手改干净)
87
+ hint = (hint ? hint + ';' : '') + `已改名「${args.title}」`
88
+ }
89
+ } else {
90
+ const matches = store.getTasks().filter((t) => t.title === args.title)
91
+ if (matches.length > 1) {
92
+ const candidates = matches.map((t) => ({ taskId: t.id, title: t.title, updatedAt: t.updatedAt, archived: t.archived }))
93
+ return truncate(JSON.stringify({ success: false, error: '多个同名任务', candidates }), config.maxOutputChars)
94
+ }
95
+ if (matches.length === 1) {
96
+ task = matches[0]
97
+ if (task.archived) {
98
+ task.archived = false
99
+ hint = '已解归档'
100
+ }
101
+ } else {
102
+ task = {
103
+ id: genTaskId(root, args.title),
104
+ projectHash: hash8(root),
105
+ projectRoot: root,
106
+ title: args.title,
107
+ steps: null,
108
+ files: [],
109
+ archived: false,
110
+ lastSessionId: sid,
111
+ createdAt: now,
112
+ updatedAt: now,
113
+ lastActiveAt: now,
114
+ }
115
+ store.addTask(task)
116
+ hint = '已新建任务'
117
+ }
118
+ }
119
+
120
+ if (sid) {
121
+ store.setBinding(sid, task.id)
122
+ hint = (hint ? hint + ';' : '') + '已绑定本会话'
123
+ } else {
124
+ hint = (hint ? hint + ';' : '') + '(拿不到会话 id,未能持久绑定)'
125
+ }
126
+ task.updatedAt = now
127
+ task.lastActiveAt = now
128
+ if (sid) task.lastSessionId = sid
129
+ store.save()
130
+
131
+ const card = {
132
+ taskId: task.id,
133
+ title: task.title,
134
+ steps: task.steps,
135
+ files: task.files,
136
+ hint: hint + ';请据此用 todo_write 重建本会话清单',
137
+ }
138
+ return truncate(JSON.stringify(card), config.maxOutputChars)
139
+ },
140
+ })
141
+ }
142
+
143
+ export function archiveTaskTool(config) {
144
+ return defineTool({
145
+ name: 'archive_task',
146
+ description: 'Archive a task: hide from default views, exclude from capacity, stop syncing. Restore by selecting it again.',
147
+ parameters: {
148
+ taskId: { type: 'string', required: true, description: 'Task id to archive (from list_tasks / select_task)' },
149
+ root: { type: 'string', description: 'Project root. Defaults to current working directory.' },
150
+ },
151
+ output: { schema: { type: 'string' }, render: (_a, v) => [{ type: 'text', text: v }] },
152
+ async execute(args, exec) {
153
+ const root = resolveIndexRoot(exec, args.root)
154
+ const store = new ProjectMemoryStore(memoryRootFor(root, config.memoryDir)).load()
155
+ const task = store.getTask(args.taskId)
156
+ if (!task) return truncate(JSON.stringify({ success: false, error: 'Task not found' }), config.maxOutputChars)
157
+ if (task.archived) return truncate(JSON.stringify({ success: false, error: 'Already archived' }), config.maxOutputChars)
158
+ task.archived = true
159
+ task.updatedAt = new Date().toISOString()
160
+ store.save()
161
+ return truncate(JSON.stringify({ success: true, archived: true, hint: '已归档,select_task 可恢复' }), config.maxOutputChars)
162
+ },
163
+ })
164
+ }
package/src/util/fs.js CHANGED
File without changes
@@ -8,7 +8,7 @@ const SYNONYMS = new Map([
8
8
  ['db pool', ['数据库连接池', '连接池']],
9
9
  ])
10
10
 
11
- function extractCjkPhrases(text) {
11
+ export function extractCjkPhrases(text) {
12
12
  const phrases = []
13
13
  let run = ''
14
14
  for (const ch of text) {
@@ -135,6 +135,10 @@ export function weightedFieldText(entry) {
135
135
  return parts.join(' ')
136
136
  }
137
137
 
138
+ export function makeSearchText(entry) {
139
+ return weightedFieldText(entry).toLowerCase()
140
+ }
141
+
138
142
  export function rankEntries(entries, query, limit = 8) {
139
143
  const bm25 = buildBm25(entries, weightedFieldText)
140
144
  const scored = bm25.score(query)
@@ -188,4 +192,59 @@ export function rankExperienceScored(items, queryOrQueries, limit = 5) {
188
192
  return [...merged.values()]
189
193
  .sort((a, b) => b.score - a.score)
190
194
  .slice(0, limit)
195
+ }
196
+
197
+ function countOccurrences(text, token) {
198
+ if (!token) return 0
199
+ let count = 0
200
+ let pos = 0
201
+ while ((pos = text.indexOf(token, pos)) !== -1) {
202
+ count++
203
+ pos += token.length
204
+ }
205
+ return count
206
+ }
207
+
208
+ const STREAM_K1 = 1.2
209
+ const STREAM_B = 0.75
210
+
211
+ export function rankEntriesStreaming(entries, queries, idf, limit = 8) {
212
+ if (!queries.length) return entries.slice(0, limit).map((entry) => ({ entry, score: 0 }))
213
+ const merged = new Map()
214
+ for (const query of queries) {
215
+ const { expanded, cjkPhrases } = expandQuery(query)
216
+ const qTokens = new Set()
217
+ for (const term of expanded) {
218
+ for (const tok of tokenizeRaw(term)) qTokens.add(tok)
219
+ }
220
+ const queryTokens = [...qTokens]
221
+ if (!queryTokens.length) continue
222
+ for (const entry of entries) {
223
+ const text = entry.searchText || weightedFieldText(entry).toLowerCase()
224
+ const len = text.length || 1
225
+ let score = 0
226
+ for (const t of queryTokens) {
227
+ const tf = countOccurrences(text, t)
228
+ if (!tf) continue
229
+ const idfVal = idf[t] || 1
230
+ score += idfVal * ((tf * (STREAM_K1 + 1)) / (tf + STREAM_K1 * (1 - STREAM_B + (STREAM_B * len) / 1000)))
231
+ }
232
+ for (const phrase of cjkPhrases) {
233
+ const lowerPhrase = phrase.toLowerCase()
234
+ if (text.includes(lowerPhrase)) {
235
+ score *= 1.5
236
+ }
237
+ }
238
+ if (score > 0) {
239
+ const id = entry.id || entry.sourcePath
240
+ const existing = merged.get(id)
241
+ if (!existing || score > existing.score) {
242
+ merged.set(id, { entry, score })
243
+ }
244
+ }
245
+ }
246
+ }
247
+ return [...merged.values()]
248
+ .sort((a, b) => b.score - a.score)
249
+ .slice(0, limit)
191
250
  }