@yolk_vat-y/dsh-project-memory 0.5.8 → 0.5.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,86 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.5.9 (2026-09-24)
4
+
5
+ ### 修复:issue #5 —— 在家目录 / `/opt/homebrew` 自动建索引把 DSH 撑到 OOM
6
+
7
+ 三个根因叠加,缺一条都复现不出报告者的现象:
8
+
9
+ - **`findProjectRoot` 会把任意目录升格成项目根。** 旧实现有两条路:`looksLikeProjectRoot`
10
+ 只要目录里有 `lib`+`include` 之类的"源码目录名"就认(`/opt/homebrew` 正好命中),以及找不到
11
+ 任何标记时兜底返回 `path.dirname(filePath)`(读到家目录下的散文件就把家目录当根)。另外
12
+ ARM Mac 的 Homebrew 是 `git clone` 到 `/opt/homebrew` 的——**那里有 `.git`**,所以单靠
13
+ "只认项目标记"仍然会把整个 Homebrew 判成项目根,危险前缀必须显式列名单。
14
+ - **`.dsh-project-memory` 自己算项目标记,形成自我固化环。** `shadowLog` 默认开、每步写盘前
15
+ `mkdirSync`,所以只要在家目录里启动过一次 dsh,该目录就被建出来;此后读家目录下任意文件都会
16
+ 把根解析回家目录,家目录被**永久**锁成"项目"(并写进 `watch.json`)。
17
+ - **扫描/索引全链路没有上限。** `walkDir` 同步把整棵树攒成一个数组;`index_repo` 攒完全部
18
+ updates 才 commit 一次(峰值 = 整棵树的条目);`DEFAULT_IGNORE` 不含 `Library`/`Documents`/
19
+ `Cellar`;watch 每 15 秒把这一切重跑一遍。
20
+
21
+ 修复:
22
+
23
+ - 新增 `isUnsafeRoot` / `assertSafeRoot` / `resolveSafeIndexRoot`:文件系统根、家目录及其**所有
24
+ 祖先**、共享临时目录、系统与包管理器前缀一律拒绝,且**按平台分组**(POSIX:`/opt/homebrew`、
25
+ `/opt`、`/usr`、`/usr/local`、`/home/linuxbrew`、`/nix` …;Windows:`%SystemRoot%`、
26
+ `%ProgramFiles%`、`%ProgramData%` …)——`C:\opt`/`C:\usr` 是 Windows 上的正常用户目录,
27
+ 套用 POSIX 名单既误伤又自相矛盾。判定是
28
+ **精确匹配**:只拒绝这些目录本身,子目录照常可用(`~/Library/Mobile Documents/…/proj` 仍是合法项目根)。
29
+ - 所有会写盘/扫描的入口统一过这道关:`index_repo`、`watch_repo`、`index_doc`、`remember`/
30
+ `forget`/`lesson`/`query_memory`/`memory_stats`/任务工具、watch 管理器、`autoIndexOnFirstUse`。
31
+ - **项目根推导收敛成 `util/fs.js` 里的唯一策略**,懒索引 / 会话审计 / TaskBridge / 所有工具
32
+ 用同一套顺序:显式 `root` → 显式登记的根(`watch_repo`)→ 最近的 VCS/清单标记 →
33
+ **会话工作目录本身(只要它是安全目录)** → `null`。旧实现是"标记优先,找不到就兜底到文件
34
+ 所在目录",且审计/工具侧直接用原始 cwd、懒索引用标记根,同一个会话会分裂成两个 store。
35
+ 现在"有没有标记"不再决定"能不能用"——只决定"这个根是声明的还是推定的"。没有标记的散目录
36
+ 仍然**不会**被升格:只有会话 cwd(人明确选择的工作位置)才有这个资格。
37
+ - **推定根会通告模型**:根来自无标记的会话工作目录时,首次注入前追加一行"记忆根 = X
38
+ (无项目标记,按会话工作目录推定);若要改,给 index_repo / watch_repo / remember /
39
+ query_memory 传 `root: <dir>`,或在项目目录里重启 dsh"。每个会话一次,与常驻任务卡同类
40
+ (状态声明,不占条目额度),`autoContext.rootNotice: false` 可关。
41
+ - 危险根里**自动路径零副作用**:懒索引不索引、auto-inject 不读 store 也不写审计(不再把
42
+ `.dsh-project-memory` `mkdirSync` 进家目录)、TaskBridge 不建档不记文件;即使被显式调用的工具
43
+ 也会收到可执行的拒绝说明,只有新配置 `allowUnsafeRoots: true` 才放行显式调用。
44
+ - 有界化:`walkDir` 返回 `{files, truncated}` 并支持 `maxScanFiles`(默认 20000)/`maxScanDepth`
45
+ (默认 12);`index_repo` 每 200 个文件分批提交;**截断时不执行"删除本轮未见到条目"的清理**
46
+ (没扫到 ≠ 被删除,否则一份被上限截断的树每轮都会自我清空一半);`DEFAULT_IGNORE` 补入
47
+ `Library`/`Applications`/`Cellar`/`Caskroom`/`Frameworks`/`DerivedData`/`Pods` 等。
48
+ - 自愈:`WatchManager.restorePersisted()` 用新判据剔除历史遗留的污染 watch root(家目录、
49
+ `/opt/homebrew` 等),用户不需要手删 `watch.json`。
50
+ - 无项目的会话不再静默:`/tasks` 等命令给出明确说明,auto-inject 在 stderr 提示一次。
51
+
52
+ 验证:新增 `test/root-guards.test.mjs`(69 项:判定与例外、**用 `path.win32` 在任意平台上
53
+ 模拟验证 Windows 分支**、每个入口的拒绝、历史污染自愈、安全无标记 cwd 照常可用、推定根通告、
54
+ 审计与懒索引用同一个根、截断不误删、auto-inject 静默);
55
+ `test/run-test.mjs` 里依赖旧语义(启发式、兜底、temp ceiling)的断言改为新契约;
56
+ `injection-audit` / `injection-budget` 的夹具根补上项目标记,避免根通告混进它们的计数。
57
+ 全量测试与 `npm run typecheck` 通过。
58
+
59
+ ### 开发工具:typecheck 从「一直不可用」修到可跑,并接入 CI
60
+
61
+ - 根 `tsconfig.json` 的 `include: ["src/**/*"]` 在纯 JS 源码且未开 `allowJs` 时**匹配不到任何输入**,
62
+ `tsc` 只会报 `TS18003`;`tsconfig.client.json` 也从未通过过(`.ts/.tsx` 后缀导入 `TS5097` ×13、
63
+ CSS Modules 无声明 `TS2307`、`session-id.js` 无类型 `TS7016`、`catch (err)` 取 `.message` `TS2339`)。
64
+ 现在:根项目开 `allowJs`(`checkJs: false`,只做解析级校验),客户端项目开
65
+ `allowImportingTsExtensions` + `allowJs`,并补 `src/client/css-modules.d.ts`;新增 `npm run typecheck`,
66
+ CI 在 `npm test` 之前执行。**没有**重建 `client/client.js`:那处收窄是纯类型改动,产物逐字不变。
67
+
68
+ ### 重构:步骤读取收敛为单一纯函数模块(无行为变更,含两处刻意的语义对齐)
69
+
70
+ - **同一套读取逻辑原先抄了 8 份,并且已经漂移。** `commands/tasks.js`、`commands/task-actions.js`、
71
+ `tools/query-memory.js`、`tools/task-tools.js`、`tools/lesson-tools.js`、`setup/taskbridge.js`、
72
+ `auto-inject.js`、`reflection-pipeline.js` 各自实现「取步骤文本 / 归一状态 / 数完成度」:取文本分裂成
73
+ `content ?? text` 与 `content || text` 两种,状态归一分裂成三种(无校验 / `Set` 白名单 / 数组白名单);
74
+ 此外还有 `timeAgo` ×2、`sessionIdOf` ×2、`textOf` ×2、`done/total` 内联 ×2。现在统一到
75
+ `src/util/task-view.js`(`stepContent` / `stepStatus` / `stepProgress` / `timeAgo` / `sessionIdOf`),
76
+ 消息文本抽取并入 `src/util/text.js` 的 `textOf`。
77
+ - 两处刻意的语义决定(**不是笔误**,由 `test/task-view.test.mjs` 钉住):
78
+ 1. `content` 为空串时**不再**回退 `text` —— 显式清空的步骤不该顶出旧文本;
79
+ 2. `/tasks todos` 收到字符串步骤时不再被静默过滤成「0 条」而清空步骤,改为按 `pending` 接受。
80
+ 非法 `status` 一律归一为 `pending`(宿主 `todo/write` 只认那三个合法值)。
81
+ - 验证:真实 store(28 任务 / 150 步骤)的状态取值本就只有三个合法值,所以归一化对现存数据零影响;
82
+ `git stash` 前后各跑一次真实 `/tasks` 快照,输出逐字相同;全量 182 项测试与 `typecheck` 均通过。
83
+
3
84
  ## 0.5.8 (2026-09-22) — 准入的可复现性 + dsh 0.1.7 兼容
4
85
 
5
86
  ### 兼容性修复(dsh 0.1.7 必崩)
package/README.md CHANGED
@@ -9,14 +9,13 @@
9
9
 
10
10
  A persistent **project development memory** for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (dsh) agents. Built specifically for project development, natively integrated with dsh's task system: task lists and files read during a session are automatically persisted as cross-session task records, with tasks ↔ files linked — workflows can be switched and resumed, no need to re-scope the whole project, solving context loss. Documents (PDF/Markdown/txt) and code symbols are stored separately per workspace; documents are automatically cross-linked to the code symbols they mention. Experience notes (problem → solution) are automatically deduplicated, preventing repeated mistakes. All data is stored per project on disk, survives session compaction and handover; recalls include `path:line` citations for source verification. Only one dependency, no vector DB, no native builds.
11
11
 
12
- > The plugin keeps a compact project **memory** on disk, with every entry pointing to a concrete file and line — the agent can reorient quickly instead of re-reading the whole project. Tasks and experience persist across session compactions and handovers.
12
+ > The plugin keeps a compact project **memory** on disk, with every entry pointing to a concrete file and line — the agent can reorient quickly instead of re-reading the whole project.
13
+
14
+ ![Task panel: task list, step progress, and involved files](docs/images/image.png)
13
15
 
14
- ![alt text](docs/images/image.png)
15
- The workflow panel is collapsible, automatically adapts to dsh and theme plugin styles, and offers four card style options to switch between.
16
- ![alt text](docs/images/image-4.png)
17
16
  ## Features
18
17
 
19
- - **TaskBridge: cross-session development tasks** — the session's todo list and the files it touches are persisted as durable per-project task entities, so a workflow can be switched and resumed without re-scoping the project. Associated files are kept in recency-weighted order (a read never outranks a written file), so a resumed session sees where to look first. New sessions continue through `list_tasks` → `select_task`. Work delegated to subagents does not create tasks (see Design tradeoffs). Auto-sync needs a dsh build with session events; on older hosts the task tools still work as a plain record list.
18
+ - **TaskBridge: cross-session development tasks** — the session's todo list and the files it touches are persisted as durable per-project task entities, so a workflow can be switched and resumed without re-scoping the project. Associated files are kept in recency-weighted order (a read never outranks a written file), so a resumed session sees where to look first. New sessions continue through `list_tasks` → `select_task`. Work delegated to subagents does not create tasks (see Known limits and boundaries). Auto-sync needs a dsh build with session events; on older hosts the task tools still work as a plain record list.
20
19
  - **Task panel in dsh web (v0.4.2+)** — draggable cards show a task's steps and files, collapse to a mini-bar, or hide entirely. The panel stays hidden until summoned, syncs in the background on session switch, and does not reopen itself after a page refresh. Render errors are contained, so a panel failure cannot take down the host.
21
20
  - **Panel editing and themes (v0.4.2+)** — bound cards allow inline editing of title and steps and status cycling; unbound cards are read-only. Four visual themes change material, geometry, typeface and density only; colours follow the host.
22
21
  - **Bidirectional task-list sync (v0.4.2+)** — binding a task pushes its steps to the host task list, and panel edits write back through the same code path as model updates. Set `tasklist.syncHostOnAdopt` to false to opt out.
@@ -29,71 +28,11 @@ The workflow panel is collapsible, automatically adapts to dsh and theme plugin
29
28
  - **Experience notes** — problems → solutions, deduplicated by overlap rather than repeated, bounded by project size, and returned only when a search matches.
30
29
  - **Tiered insight memory (lessons / decisions / procedures, v0.5)** — one entity across task, project and global scope. Near-duplicates merge or reinforce; promotion moves an entry between scopes rather than copying it. Use is recorded, so decay and capacity rank by activity rather than by age alone. LLM reflection is off by default and writes task-level drafts only. The panel exposes a per-scope memory view for reviewing and editing entries.
31
30
  - **Triggered injection** — an insight may carry an authored trigger: only `when` triggers, `guard` narrows it, and `prevents` records what breaks without the entry. Corpus text can never trigger an injection; the statistical channel is gated separately.
32
- - **Streaming TF + IDF caching** — the query path caches term weights per store version, scoring 20k entries in p50 2.6 ms / p95 5.4 ms and 4k entries in p50 0.6 ms / p95 1.6 ms. Only a real write drops the cache, so the watch poll never clears one a query just built.
31
+ - **Streaming TF + IDF caching** — the query path caches term weights per store version (measured numbers under Performance). Only a real write drops the cache, so the watch poll never clears one a query just built.
33
32
  - **Lock-free sync transactions** — all writes go through a synchronous transaction, so `remember` and `forget` never queue behind re-indexing. The lock is in-process: avoid pointing two dsh instances at the same store.
34
33
  - **Minimal dependencies** — pure JavaScript; one runtime dependency for PDF text extraction, no native builds.
35
34
  - **Negligible overhead** — memory work is in-process; the bottleneck is document extraction and disk I/O, not scoring.
36
35
 
37
- ## Performance
38
-
39
- ### Synthetic Benchmark (Node 24.19, WSL2 on 20 vCPU, Linux file system)
40
-
41
- | Scenario | Scale | Measured |
42
- |----------|-------|----------|
43
- | Full cold index | 5,000 files / 20k entries | 269 ms avg (p50 267) |
44
- | Cold load | 5,000 files | 40 ms |
45
- | Hot lazy re-index (single file) | 5k files | p50 2.4 ms / max 5.5 ms |
46
- | query_memory (cached) | 5k files / 20k entries | p50 2.6 ms / p95 5.4 ms |
47
- | query_memory (cached) | 1k files / 4k entries | p50 0.6 ms / p95 1.6 ms |
48
- | Full cold index | 10,000 files / 40k entries | 551 ms avg (p50 528) |
49
- | Cold load | 10,000 files | 90 ms |
50
- | Hot lazy re-index (single file) | 10k files | p50 5.4 ms / max 9.2 ms |
51
-
52
- > Synthetic benchmark: generated code (~4–5 symbols/file), Node 24.19 on WSL2 / 20 vCPU / Linux file system, measured 2026-09-14. Reproduce with `npm run bench:synthetic -- 5000` (harness: `scripts/bench-synthetic.mjs`). Measures pure indexing overhead without LLM calls. query_memory uses the IDF cache + precomputed searchText; the first query after a write rebuilds IDF (**106 ms at 40k entries**, 57 ms at 20k, 12 ms at 4k), subsequent queries hit the cache.
53
-
54
- ### Real Project Storage
55
-
56
- | Project | Files | Entries | Store Size | Per Entry |
57
- |---------|-------|---------|------------|-----------|
58
- | Java Spring Boot backend | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
59
- | Vue 3 + Vite frontend | 289 | 2,141 | 1.0 MB | ~0.5 KB |
60
-
61
- > Real projects (Java + Vue), tested on Linux file system (Node 24). Real project entries are smaller than synthetic benchmarks due to lower symbol density and shorter declarations.
62
-
63
- ### Reproduce it on your own project
64
-
65
- Rather than asking you to trust the numbers above, the measurement itself ships with the repository **and with the published npm package** (`scripts/` is part of the tarball). It needs **no dsh instance, no network and no model calls**, and it never touches your project's own store — results go to a temp directory and are removed when it finishes:
66
-
67
- ```bash
68
- npm run bench -- /path/to/your/project
69
- # or, with options:
70
- node scripts/bench.mjs /path/to/your/project [--json] [--samples 100] [--no-pdf] [--keep]
71
- ```
72
-
73
- It reports the cold index split into read+hash / extract / commit, cold load, IDF rebuild, cold and hot query latency (p50/p95/max over 100 sampled queries through the shipped scorer), single-file hot re-index, store size and bytes per entry. Example — our internal Vue project (289 files / 2,141 entries, Node 24, 20 CPU, Linux):
74
-
75
- ```
76
- cold index 253 ms (read+hash 9 ms · extract 229 ms · commit 13 ms) ← 2nd, warm-cache run
77
- store 1.10 MB · 538 bytes/entry · cold load 4.6 ms
78
- hot query p50 0.80 ms · p95 1.35 ms (2,141 entries)
79
- re-index 1 file p50 0.33 ms
80
- ```
81
-
82
- Two caveats we would rather state than hide: `read+hash` depends on the OS page cache — on that corpus the first run spent 787 ms and the second 253 ms, so say which run you quote — and **real projects score slower than the synthetic table above** — on a 3,000-file slice of a large TypeScript repository (15,594 entries) hot queries were p50 7.5 ms, because real declaration text is longer than generated stubs. Pass `--queries your-queries.json` to run the same labeled-set method (hit@5 / hit@10 / MRR) against your own project.
83
-
84
- ## How it works
85
-
86
- The design follows four principles:
87
-
88
- - **Volatility** — context is ephemeral; it is lost when a session is compacted.
89
- - **Persistence** — the **memory** is stored on disk and survives compaction and new sessions.
90
- - **Compactness** — the code layer stores one declaration line per symbol, so code-heavy projects stay near **0.5% of the source** (8.8 MB of source → 49 KB of index in the example project), and **recall** replaces re-reading the full file. The document layer is heavier by design: each chunk keeps a ≤300-char injected `summary`, a bounded `terms` set covering the whole chunk for retrieval, and a precomputed `searchText`. Measured on a docs-only corpus (179 chunks / 225 KB of Markdown): `terms` ≈ **27.5%** of source and the on-disk store ≈ **166%** of source — so on doc-heavy projects budget for roughly the docs themselves, not 0.5%.
91
- - **Verifiability** — **recalls** carry a `path:line` citation where applicable, so the agent can confirm details against the source.
92
-
93
- Building the **memory** does not require an upfront scan: files are memorized as the model reads them, so the **memory** grows to cover exactly what has been worked with. Re-reading a file that has not changed is a no-op (content hash), so the **memory** stays fresh with minimal ongoing overhead.
94
-
95
- The store is per-project and follows the codebase: changed files are re-extracted by content hash, deleted files are removed. Experience notes are retrieval-only, so accumulation does not affect context.
96
-
97
36
  ## Installation
98
37
 
99
38
  The plugin relies exclusively on stable public APIs (`defineTool`, `llm.stream`, `Schema`) declared via peerDependencies, ensuring compatibility with future rc/alpha releases without changes.
@@ -125,8 +64,8 @@ The tools below are **invoked by the agent**, not typed by the user. In the chat
125
64
  | Tool | Purpose |
126
65
  |---|---|
127
66
  | `index_doc file_path` | Index one document (PDF/MD/txt): chunk → deterministic `summary` + whole-chunk `terms` → store with `path:line`. Unchanged files are skipped. |
128
- | `index_repo root` | Index a whole project: docs get deterministic summaries + whole-chunk terms, code files get a zero-token symbol table. Incremental, cleans up deleted files, cross-links docs to symbols. A root that does not exist — including a Windows-style path resolved on Linux/macOS — is rejected before anything is written. |
129
- | `watch_repo root` | Enable automatic refresh: a background poll detects new/changed files (mtime + content hash) and re-indexes only those. Watched roots persist across plugin restarts; a non-existent root, the filesystem root and the shared temp directory are all refused, and roots that disappear are dropped instead of being re-created. |
67
+ | `index_repo root` | Index a whole project: docs get deterministic summaries + whole-chunk terms, code files get a zero-token symbol table. Incremental, cleans up deleted files, cross-links docs to symbols. A root that does not exist — including a Windows-style path resolved on Linux/macOS — is rejected before anything is written, and so is a *dangerous* root (home directory, filesystem root, system / package-manager prefixes): scanning one of those walks hundreds of thousands of files. |
68
+ | `watch_repo root` | Enable automatic refresh: a background poll detects new/changed files (mtime + content hash) and re-indexes only those. Watched roots persist across plugin restarts; a non-existent root and a dangerous root (filesystem root, home directory, shared temp directory, system / package-manager prefixes) are all refused, roots that disappear are dropped instead of being re-created, and a polluted watchlist from an older version is self-healed on startup. |
130
69
  | `memory_stats root` | Show what the store contains: totals (files / entries / experience notes), last index time, and the per-file list sorted by recency. |
131
70
  | `query_memory query` | BM25 search over docs + symbols + experience + insights (lessons / decisions / procedures), optionally query-expanded by the LLM. `type` selects a layer (`all` / `doc` / `symbol` / `experience` / `insight` / `task`). Returns ranked hits with relative scores, sources or insight ids, and doc→symbol references. |
132
71
  | `list_tasks` | List task records for the project (archived marked). Call first in a new session before continuing work. |
@@ -146,6 +85,19 @@ registers **only `/tasks`** and every other action rides it as a sub-verb driven
146
85
  menu duplication is one row. Typing `/tasks` + Enter still executes immediately; typing an argued line
147
86
  such as `/tasks switch x` is no longer recognised as a command — use the card buttons.
148
87
 
88
+ ## How it works
89
+
90
+ The design follows four principles:
91
+
92
+ - **Volatility** — context is ephemeral; it is lost when a session is compacted.
93
+ - **Persistence** — the **memory** is stored on disk and survives compaction and new sessions.
94
+ - **Compactness** — the code layer stores one declaration line per symbol, so code-heavy projects stay near **0.5% of the source** (8.8 MB of source → 49 KB of index in the example project), and **recall** replaces re-reading the full file. The document layer is heavier by design: each chunk keeps a ≤300-char injected `summary`, a bounded `terms` set covering the whole chunk for retrieval, and a precomputed `searchText`. Measured on a docs-only corpus (179 chunks / 225 KB of Markdown): `terms` ≈ **27.5%** of source and the on-disk store ≈ **166%** of source — so on doc-heavy projects budget for roughly the docs themselves, not 0.5%.
95
+ - **Verifiability** — **recalls** carry a `path:line` citation where applicable, so the agent can confirm details against the source.
96
+
97
+ Building the **memory** does not require an upfront scan: files are memorized as the model reads them, so the **memory** grows to cover exactly what has been worked with. Re-reading a file that has not changed is a no-op (content hash), so the **memory** stays fresh with minimal ongoing overhead.
98
+
99
+ The store is per-project and follows the codebase: changed files are re-extracted by content hash, deleted files are removed. Experience notes are retrieval-only, so accumulation does not affect context.
100
+
149
101
  ## Design
150
102
 
151
103
  ```
@@ -179,99 +131,9 @@ TaskPanel (Container)
179
131
  └── TaskComponents (MiniBar, TaskCard — presentational only)
180
132
  ```
181
133
 
182
- ## Design tradeoffs
183
-
184
- These are deliberate scope choices.
185
-
186
- ### 1. Synchronous lock-free transactions over async locks
187
-
188
- **We do:** All writes go through `store.commit(fn)` — a synchronous in-process transaction. The callback `fn` performs all validation and mutations; only on success is the result atomically written to disk. The JS event loop guarantees no interleaving. CAS (`applyFileUpdate`) makes concurrent writes idempotent.
189
-
190
- **We don't:** Async mutexes, file locks, or multi-process coordination.
191
-
192
- **Why:** DSH runs on Cordis, which is single-process by design. Adding locks would complicate the hot path (every `remember`/`forget`/`index_doc` call) for a scenario (multi-process DSH) that would require a breaking ecosystem change. Synchronous transactions keep the hot path at ~2 ms median with zero contention overhead in practice.
193
-
194
- ### 2. Watch: compute outside, commit inside
195
-
196
- **We do:** Heavy work (mtime/hash/scan/parse/PDF extraction) runs outside the transaction; a single `commit` applies all changes atomically. On failure, the snapshot rolls back so the next poll retries automatically.
197
-
198
- **We don't:** Hold a lock during parsing, or use `fs.watch` events.
199
-
200
- **Why:** PDF extraction and large-file parsing take time — holding a lock would block `remember`/`forget`/`query_memory`. Polling with mtime+content-hash is platform-agnostic (works on network drives, Docker volumes, WSL) and avoids the "double fire / missed events" nightmare of `fs.watch`.
201
-
202
- ### 3. Corrupt files are quarantined, not auto-repaired
203
-
204
- **We do:** On JSON parse failure, the bad file is renamed to `*.corrupt`, an error is logged, and that file's store starts fresh. The rest of the store remains intact.
205
-
206
- **We don't:** Write-ahead logs, embedded databases (SQLite/LMDB), or automatic partial recovery.
207
-
208
- **Why:** A corrupted shard means *one source file* has a bad index — quarantining it costs near zero. A WAL or embedded DB adds a heavy dependency, increases binary size, and introduces new failure modes (lock contention, corruption of the WAL itself). The tradeoff: lose one file's index vs. add 500 KB+ of native code.
209
-
210
- ### 4. No vector embeddings, no semantic search at query time
211
-
212
- **We do:** BM25 with CJK phrase boost (3+ chars ×1.5 on title/keywords), synonym expansion (bidirectional table), field weighting (title ×5), and experience-layer phrase boost. All at query time, zero LLM calls.
213
-
214
- **We don't:** Vector embeddings, dense retrieval, rerankers, or hybrid search.
215
-
216
- **Why:** Vectors require an embedding model (local = heavy, remote = latency + cost + privacy), a vector index (HNSW/IVF = memory + build time), and reranking (another LLM call). For the queries this plugin targets, lexical BM25 is already sufficient and measurable: on our benchmark suite (29 queries over a real Vue project) file-level hit@5 is **96.6%**, and 28 of the 29 are exact symbol lookups that lexical search answers essentially always. Whole-chunk `terms` took document-term coverage from **27.3% to 100%** while queries that already worked kept their ranking (MRR **0.958** vs **0.955**). Those figures come from an internal Vue project with a hand-labeled 29-query set, so they are not reproducible outside it — but the **method** now ships as `scripts/bench.mjs --queries <your-set.json>`, so you can run the identical measurement on your own project. The marginal gain from semantic search doesn't justify the 10x complexity/cost increase.
217
-
218
- ### 5. Indexing is deterministic and model-free
219
-
220
- **We do:** Derive keywords with a rule (title-weighted top terms) and build a whole-chunk `terms` set — both deterministic and reproducible. Doc↔symbol links surface English symbol names from Chinese queries, and CJK tokenization keeps cross-language hits working. With `llmQueryExpansion: false`, queries never touch the LLM.
221
-
222
- **We don't:** Call a model at index time to translate or paraphrase a document, and we don't translate queries at search time.
223
-
224
- **Why:** An index-time model call makes indexing slower, non-deterministic and unverifiable — the same document can index differently on two runs. Query-time translation adds latency and a hard failure mode (a bad translation means zero recall). Rules plus symbol linking cover the common cases, work offline, and keep indexing at zero model calls.
225
-
226
- ### 6. Model-facing memory: the agent writes, and no human has to be in the loop
227
-
228
- **We do:** Treat the agent as a first-class writer. `remember` / `save_lesson` write **any scope at any time** (`task` / `project` / `global`) with no human step, and promotion is deterministic and runs inside the ordinary write path: cross-task token-overlap dedupe accumulates `sourceTaskIds`, then `promoteAllTasksToProject` / `promoteProjectToGlobal` move an entry up once its corroboration counts are met (≥2 tasks for project, ≥ `globalPromoteTasks` — 3 by default — for global). Nothing waits on the task panel: a user who never opens the UI still gets a memory that fills, dedupes and graduates.
229
-
230
- **We do (labeling):** Keep inferred content distinguishable from recorded content. The v0.5 `reflection` path (opt-in, **off by default**) is the only writer that infers rather than records: it writes task-scoped drafts stamped `draft: true` / `source: 'reflect'`, and `recall` plus silent injection skip `draft` entries while they remain drafts.
231
-
232
- **We don't:** Require human approval for memory to become useful, or make the UI a step in the write path. `draft` is a **provenance label plus a corroboration threshold**, not an approval queue.
233
-
234
- **Why:** The agent is the consumer and it is usually headless — memory that only graduates when a human clicks a card is memory that never graduates. Labeling keeps the useful half of the caution (inferred ≠ recorded, and unreviewed single-task inference stays out of the prompt) without taxing the normal path. A draft graduates on corroboration: a second task matching it through the model's own writes, or the model writing the same knowledge at project scope, which links the existing entry instead of duplicating it.
235
-
236
- ### 7. Full entries returned directly
237
-
238
- **We do:** `query_memory` returns complete entries with `path:line` citations. Every hit can be verified against source.
239
-
240
- **We don't:** Return a minimal index first, then require a second tool call for details.
241
-
242
- **Why:** Returning full entries preserves **verifiability** — the agent sees the exact source line for every claim. It also avoids a round-trip per useful hit. Our entries are already compact (~300-char summary + citation, plus a search-only `terms` field that never enters the prompt); the token cost is lower than a second tool call + context switch.
243
-
244
- ### 8. Symbol extraction focused on what developers search for
245
-
246
- **We do:** Regex-based symbol extraction (functions, classes, methods, interfaces, type aliases) with string/comment masking, multi-line signatures, and cross-file linking by symbol name. For TypeScript/JavaScript projects, an optional L2 enhancement layer uses the TS Compiler API to infer return types, resolve generics, and extract interfaces — all cached by content hash for instant reuse.
247
-
248
- **We don't:** Tree-sitter AST parsing, import graphs, call graphs, or full-program type resolution across files.
249
-
250
- **Why:** Our regex scanner handles 8 languages with zero dependencies, runs in <1 ms/file, and captures the declarations developers actually search for (names, signatures, generics). The optional TS layer adds semantic depth for TS/JS without native deps. Cross-file linking by name covers the most common "find related code" use case. Full-program analysis would add native binaries, 10x install size, and version fragility — for marginal gain on the remaining 5% of edge cases.
251
-
252
- ### 9. `forget` by query is aggressive; prefer ID deletion
253
-
254
- **We do:** `forget query` deletes all experience notes with ≥0.5 token overlap.
255
-
256
- **We don't:** Interactive confirmation, soft-delete/trash, or exact-match-only.
257
-
258
- **Why:** Experience notes are low-stakes, high-volume, and retrieval-only. Aggressive deletion prevents stale noise from polluting search. For precision, delete by ID (shown in `query_memory` output).
259
-
260
- ### 10. TypeScript enhancement is optional, lazy, and cached
261
-
262
- **We do:** L2 TS Compiler API enhancement runs async in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`), results cached by content hash in `type-cache/`. Zero config — just `npm i -D typescript@5` or `typescript@6`. Falls back to L1 regex if TS absent or disabled.
263
-
264
- **We don't:** Mandatory TS, blocking enhancement, or full-program type checking.
265
-
266
- **Why:** Mandatory TS would break installs for non-TS projects. Blocking enhancement would stall `index_repo` on large codebases. Full-program checking is 10x slower and memory-heavy. Our design: enhance what's read, cache it, never block the hot path.
267
-
268
- ### 11. Subagent sessions are out of scope for now
269
-
270
- **We do:** Exclude sessions spawned as subagents (`origin: 'subagent'` / `delegationDepth > 0`) from auto-creating or binding a task. Their `todo_write` events do not create tasks, and they inherit no task binding.
271
-
272
- **We don't:** Merge a delegated run's steps and files back into the task that spawned it. That is **not designed yet**: there is no parent-link model for delegated work, and the naive version mints one project task per subagent.
134
+ The workflow panel is collapsible, automatically adapts to dsh and theme plugin styles, and offers four card style options to switch between.
273
135
 
274
- **Why:** Every subagent that writes a todo would otherwise create its own task entity, so one fan-out run would flood the task list with ephemeral entries nobody resumes. Excluding them keeps the task list equal to the work the user actually owns. The cost is that a delegation's progress is invisible in the task record; merging it properly (child steps folded into the parent, or a separate delegated-work view) is future work.
136
+ ![Four card styles](docs/images/image-4.png)
275
137
 
276
138
  ## Configuration
277
139
 
@@ -291,11 +153,19 @@ These are deliberate scope choices.
291
153
  | `autoIndexOnFirstUse` | false | full scan of the current working directory on plugin load (opt-in) |
292
154
  | `watch` | true | enable the background refresh |
293
155
  | `watchInterval` | 15 | poll interval (seconds) |
156
+ | `maxScanFiles` | 20000 | hard cap on files per scan pass; a truncated scan is reported and never deletes the entries it did not reach. Set `0` to disable the cap (at your own risk) |
157
+ | `maxScanDepth` | 12 | hard cap on directory depth per scan pass. Set `0` to disable |
158
+ | `allowUnsafeRoots` | false | allow **explicit** tool calls (`index_repo`/`watch_repo`/`remember` with a `root`) to target a dangerous root. Automatic paths (lazy indexing, session audit, TaskBridge, `autoIndexOnFirstUse`) stay inert in these directories regardless |
294
159
  | `tsPath` | (auto) | optional absolute path to a specific `typescript` install; if omitted, resolves from project cwd → plugin node_modules |
295
160
  | `enableTypeScript` | true | set `false` to disable L2 TS enhancement entirely (L1 regex only) |
161
+
162
+ ### Memory and injection knobs
163
+
164
+ | 键 | 默认值 | 含义 |
165
+ |---|---|---|
296
166
  | `insight.*` | dedupOverlap `0.7` · reinforceBand `0.65` · maxProject `100` · maxGlobalProcedures `200` · promoteConfidence `0.7` · globalPromoteTasks `3` · decayDays `90` · `globalFile` (auto) | v0.5 insight dedupe / reinforce / promotion / capacity / archive settings |
297
167
  | `reflection.enabled` | false | v0.5 LLM reflection, **draft-only at task level** (fires on task switch-away / archive). `cooldownMs` `1800000`, `maxLessonsPerReflect` `3`, `maxDecisionsPerReflect` `2` |
298
- | `autoContext.enabled` | true | silent injection wrapper (resident task card + gated items). Inert (full passthrough) until the host exposes a resolvable session cwd; `maxTokens` `400`, `editedMax` `3` (how many recently-written "editing now" files the resident task card shows), `signalMinRatio` `0.5` (a hint must reach half of its layer's top score), `skipEchoSelfTodo` `true` (don't echo the task card back when the model itself maintains the task list with no newer human message; relevant insights still inject), `budgetLog` `off` (budget-drop audit on stderr: `off` silent / `once` at most one line per session / `all` one line per changed dropped set), `reinjectItemsAfter` `0` (cooldown, in pre-steps, before the same insight may be injected again) |
168
+ | `autoContext.enabled` | true | silent injection wrapper (resident task card + gated items). Inert (full passthrough) until the host exposes a resolvable session cwd; `maxTokens` `400`, `editedMax` `3` (how many recently-written "editing now" files the resident task card shows), `signalMinRatio` `0.5` (a hint must reach half of its layer's top score), `skipEchoSelfTodo` `true` (don't echo the task card back when the model itself maintains the task list with no newer human message; relevant insights still inject), `budgetLog` `off` (budget-drop audit on stderr: `off` silent / `once` at most one line per session / `all` one line per changed dropped set), `reinjectItemsAfter` `0` (cooldown, in pre-steps, before the same insight may be injected again), `rootNotice` `true` (when the memory root is inferred from a marker-less working directory, tell the model once where memory lives and how to change it) |
299
169
  | `autoContext.gateCooldownSteps` | 2 | **admission knobs.** Minimum number of pre-steps between two *item* injections (the resident task card is exempt — it is a state snapshot and should update when it changes). This is the main "don't inject often" dial |
300
170
  | `autoContext.maxItemsPerSession` | 12 | hard per-session cap on injected items; the budget is a ceiling, not a target — once exhausted the item channel stays silent |
301
171
  | `autoContext.maxItemCharsPerSession` | 4000 | same, in characters |
@@ -322,6 +192,16 @@ Automatic injection used to be a *retrieval* problem ("which entry is most relat
322
192
 
323
193
  The two most relevant switches are `lazyIndexing` (index a file the moment the model reads it; default on) and `autoIndexOnFirstUse` (full scan of the current working directory on plugin load; default off). Lazily indexed project roots are automatically registered with the watcher, so changed files stay fresh without an explicit `watch_repo`.
324
194
 
195
+ **Where the project root comes from.** One policy, applied identically by lazy indexing, the session audit trail, TaskBridge and every tool: an explicit `root` argument wins; otherwise an explicitly registered root (`watch_repo`); otherwise the nearest ancestor containing a VCS marker (`.git`/`.hg`/`.svn`) or a build/manifest marker (`package.json`, `go.mod`, `Cargo.toml`, `pyproject.toml`, …); otherwise **the session working directory itself, provided it is a safe directory**. So a marker-less scratch folder you started dsh in still gets project memory — the plugin just says so once:
196
+
197
+ ```
198
+ memory root: /Users/me/scratch (inferred from the session working directory; no project marker found).
199
+ If project memory should live elsewhere, pass `root: <dir>` to index_repo / watch_repo / remember / query_memory,
200
+ or restart dsh inside the project directory.
201
+ ```
202
+
203
+ That notice goes out once per session and can be muted with `autoContext.rootNotice: false`. What the plugin will **not** do is promote an arbitrary directory to a project: reading a stray file outside the working directory records nothing, and a dangerous root (filesystem root, your home directory, the shared temp directory or a system / package-manager prefix — `/opt/homebrew` on POSIX, `%SystemRoot%`/`%ProgramFiles%`/`%ProgramData%` on Windows) is refused outright — that is what used to walk an entire home directory and exhaust memory. Sessions whose working directory is one of those run with memory disabled (one stderr line explains why).
204
+
325
205
  Settings live in the plugin's config object. To change them, add an override entry to your profile's `cordis.patch.yml` — for the web profile that is `~/.dsh/profiles/web/cordis.patch.yml`:
326
206
 
327
207
  ```yaml
@@ -332,7 +212,10 @@ Settings live in the plugin's config object. To change them, add an override ent
332
212
  llmQueryExpansion: false # off: do not spend tokens on LLM query expansion (default)
333
213
  watch: true # on: background refresh for watched roots (default)
334
214
  watchInterval: 15 # poll interval in seconds
215
+ maxScanFiles: 20000 # per-scan file cap (truncation is reported, never deletes)
216
+ maxScanDepth: 12 # per-scan directory-depth cap
335
217
  enableTypeScript: true # on: L2 TS enhancement when TS is installed (default)
218
+ # allowUnsafeRoots: false # keep false unless you really want to index a home/system dir explicitly
336
219
  # budgetLog: once # debugging: log budget drops to stderr (default off = silent)
337
220
  # reinjectItemsAfter: 20 # debugging: allow the same insight again after N steps (default 0 = once per session)
338
221
  # tsPath: /custom/path/to/typescript # optional: force specific TS install
@@ -348,6 +231,66 @@ dsh web --patch ./config.yml
348
231
 
349
232
  where `config.yml` contains the same override block.
350
233
 
234
+ ## Performance
235
+
236
+ ### Synthetic Benchmark (Node 24.19, WSL2 on 20 vCPU, Linux file system)
237
+
238
+ | Scenario | Scale | Measured |
239
+ |----------|-------|----------|
240
+ | Full cold index | 5,000 files / 20k entries | 269 ms avg (p50 267) |
241
+ | Cold load | 5,000 files | 40 ms |
242
+ | Hot lazy re-index (single file) | 5k files | p50 2.4 ms / max 5.5 ms |
243
+ | query_memory (cached) | 5k files / 20k entries | p50 2.6 ms / p95 5.4 ms |
244
+ | query_memory (cached) | 1k files / 4k entries | p50 0.6 ms / p95 1.6 ms |
245
+ | Full cold index | 10,000 files / 40k entries | 551 ms avg (p50 528) |
246
+ | Cold load | 10,000 files | 90 ms |
247
+ | Hot lazy re-index (single file) | 10k files | p50 5.4 ms / max 9.2 ms |
248
+
249
+ > Synthetic benchmark: generated code (~4–5 symbols/file), Node 24.19 on WSL2 / 20 vCPU / Linux file system, measured 2026-09-14. Reproduce with `npm run bench:synthetic -- 5000` (harness: `scripts/bench-synthetic.mjs`). Measures pure indexing overhead without LLM calls. query_memory uses the IDF cache + precomputed searchText; the first query after a write rebuilds IDF (**106 ms at 40k entries**, 57 ms at 20k, 12 ms at 4k), subsequent queries hit the cache.
250
+
251
+ ### Real Project Storage
252
+
253
+ | Project | Files | Entries | Store Size | Per Entry |
254
+ |---------|-------|---------|------------|-----------|
255
+ | Java Spring Boot backend | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
256
+ | Vue 3 + Vite frontend | 289 | 2,141 | 1.0 MB | ~0.5 KB |
257
+
258
+ > Real projects (Java + Vue), tested on Linux file system (Node 24). Real project entries are smaller than synthetic benchmarks due to lower symbol density and shorter declarations.
259
+
260
+ ### Reproduce it on your own project
261
+
262
+ Rather than asking you to trust the numbers above, the measurement itself ships with the repository **and with the published npm package** (`scripts/` is part of the tarball). It needs **no dsh instance, no network and no model calls**, and it never touches your project's own store — results go to a temp directory and are removed when it finishes:
263
+
264
+ ```bash
265
+ npm run bench -- /path/to/your/project
266
+ # or, with options:
267
+ node scripts/bench.mjs /path/to/your/project [--json] [--samples 100] [--no-pdf] [--keep]
268
+ ```
269
+
270
+ It reports the cold index split into read+hash / extract / commit, cold load, IDF rebuild, cold and hot query latency (p50/p95/max over 100 sampled queries through the shipped scorer), single-file hot re-index, store size and bytes per entry. Example — our internal Vue project (289 files / 2,141 entries, Node 24, 20 CPU, Linux):
271
+
272
+ ```
273
+ cold index 253 ms (read+hash 9 ms · extract 229 ms · commit 13 ms) ← 2nd, warm-cache run
274
+ store 1.10 MB · 538 bytes/entry · cold load 4.6 ms
275
+ hot query p50 0.80 ms · p95 1.35 ms (2,141 entries)
276
+ re-index 1 file p50 0.33 ms
277
+ ```
278
+
279
+ Two caveats we would rather state than hide: `read+hash` depends on the OS page cache — on that corpus the first run spent 787 ms and the second 253 ms, so say which run you quote — and **real projects score slower than the synthetic table above** — on a 3,000-file slice of a large TypeScript repository (15,594 entries) hot queries were p50 7.5 ms, because real declaration text is longer than generated stubs. Pass `--queries your-queries.json` to run the same labeled-set method (hit@5 / hit@10 / MRR) against your own project.
280
+
281
+ ## Design tradeoffs
282
+
283
+ - **Synchronous lock-free transactions over async locks** — no async mutexes, file locks, or multi-process coordination: DSH runs on Cordis and single-process is an architectural given, so locking for a rare multi-process case would only slow the hot path (every `remember`/`forget`/`index_doc`); synchronous transactions keep that path at ~2 ms median with zero contention.
284
+ - **Watch: compute outside, commit inside** — no lock is held during parsing and `fs.watch` is not used: parsing and PDF extraction are slow, so a held lock would block queries, while polling with mtime + content hash behaves identically on network drives, Docker volumes, and WSL, with none of `fs.watch`'s duplicate-trigger/missed-event failure modes.
285
+ - **Corrupt shards are quarantined, not repaired** — a shard that fails to parse is renamed `*.corrupt` and only that file is re-indexed, leaving every other shard untouched; no WAL or embedded database: those add 500 KB+ of native dependencies, lock contention, and a new failure mode (a corrupt WAL) to avoid losing a single file's index.
286
+ - **No vector embeddings, no semantic search at query time** — no embedding model, vector index (HNSW/IVF), or reranker: lexical retrieval already answers the queries this plugin targets. On a real Vue project with 29 labelled queries, file-level hit@5 is **96.6%**; whole-chunk `terms` lift document term coverage from **27.3% to 100%** with MRR unchanged (**0.958** vs **0.955**). The marginal gain does not justify 10x the complexity, and the method ships with the code — `scripts/bench.mjs --queries your-queries.json` reproduces the same measurement on your own project.
287
+ - **Indexing is deterministic and model-free** — no model at index time and no translation at query time: the former makes two indexings of one document differ, the latter has a hard failure mode (a wrong translation means zero recall); rules plus symbol links already cover the common cases and work offline.
288
+ - **Model-facing memory: the agent writes, no human in the loop** — no human approval step: the consumer of this memory is the agent, and agents are usually headless, so memory that only promotes when someone clicks a card would never promote at all. `draft` is a provenance marker plus an evidence threshold, not an approval queue — the one inferring writer, `reflection` (off by default), writes task-level drafts only, and drafts never reach recall or injection.
289
+ - **Full entries returned directly** — no "minimal index first, fetch details in a second call": entries are already compact, so returning them whole is both more verifiable and one round-trip cheaper.
290
+ - **`forget` by query is aggressive; use IDs for precision** — no confirmation prompt, recycle bin, or exact-match-only mode: experience notes are low-risk, high-volume, and retrieval-only, so stale noise hurts more than an over-broad delete. For exact deletion use the ID shown by `query_memory`.
291
+ - **TypeScript enhancement is optional, lazy, and cached** — the L2 TS Compiler API runs asynchronously on a priority queue (P0 `fs/observed`, P1 `watch`, P2 `index_repo`) and caches results by content hash; TS is never required and enhancement never blocks: requiring it would make non-TS projects uninstallable, and blocking would stall `index_repo` on large projects. `npm i -D typescript@5|6` is the entire setup, and a missing TS falls back to the L1 regex scanner.
292
+ - **Subagent sessions are out of scope for now**
293
+
351
294
  ## Development (for contributors)
352
295
 
353
296
  These commands are for **maintaining the plugin code** — regular users do not need them. Installing the plugin only requires the command in [Installation](#installation).