@yolk_vat-y/dsh-project-memory 0.5.7 → 0.5.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/CHANGELOG.md +188 -0
  2. package/README.md +139 -193
  3. package/README.zh-CN.md +136 -192
  4. package/client/client.js +458 -107
  5. package/client/client.js.map +1 -1
  6. package/cordis.patch.yml +7 -1
  7. package/package.json +4 -6
  8. package/scripts/bench-synthetic.mjs +3 -3
  9. package/scripts/bench.mjs +1 -1
  10. package/src/audit.js +63 -0
  11. package/src/auto-inject.js +177 -34
  12. package/src/client/MemoryView.tsx +12 -4
  13. package/src/client/TaskCommandNode.tsx +1 -1
  14. package/src/client/TaskComponents.tsx +3 -2
  15. package/src/client/TaskPanel.tsx +14 -14
  16. package/src/client/client.ts +93 -21
  17. package/src/client/css-modules.d.ts +12 -0
  18. package/src/client/icons.ts +63 -0
  19. package/src/client/locales.ts +11 -0
  20. package/src/client/session-id.js +88 -0
  21. package/src/client/slash.ts +184 -0
  22. package/src/commands/insight-actions.js +15 -3
  23. package/src/commands/task-actions.js +19 -14
  24. package/src/commands/tasks.js +4 -22
  25. package/src/commands/workflow.js +54 -0
  26. package/src/index.js +51 -9
  27. package/src/insight-store.js +39 -0
  28. package/src/lazy.js +22 -80
  29. package/src/reflection-pipeline.js +2 -1
  30. package/src/setup/taskbridge.js +19 -12
  31. package/src/tools/forget.js +2 -2
  32. package/src/tools/index-doc.js +20 -5
  33. package/src/tools/index-repo.js +50 -17
  34. package/src/tools/lesson-tools.js +3 -6
  35. package/src/tools/query-memory.js +15 -12
  36. package/src/tools/remember.js +2 -2
  37. package/src/tools/stats.js +2 -2
  38. package/src/tools/task-tools.js +6 -20
  39. package/src/tools/watch-repo.js +17 -10
  40. package/src/util/fs.js +351 -16
  41. package/src/util/task-view.js +66 -0
  42. package/src/util/text.js +14 -0
  43. package/src/watch.js +46 -7
package/README.md CHANGED
@@ -4,103 +4,34 @@
4
4
 
5
5
  [English](README.md) | [简体中文](README.zh-CN.md)
6
6
 
7
- [![ci](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml/badge.svg)](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml) [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![npm](https://img.shields.io/npm/v/@yolk_vat-y/dsh-project-memory)](https://www.npmjs.com/package/@yolk_vat-y/dsh-project-memory) [![Listed on dsh-plugin.org](https://dsh-plugin.org/badges/listed.svg)](https://dsh-plugin.org/plugins/00080000/dsh-project-memory) [![Awesome](https://awesome-dsh-plugin.com/badge.svg)](https://awesome-dsh-plugin.com)
7
+ [![ci](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml/badge.svg)](https://github.com/00080000/dsh-project-memory/actions/workflows/ci.yml) [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) [![npm](https://img.shields.io/npm/v/@yolk_vat-y/dsh-project-memory)](https://www.npmjs.com/package/@yolk_vat-y/dsh-project-memory) [![npm downloads](https://img.shields.io/npm/dm/%40yolk_vat-y%2Fdsh-project-memory?style=flat-square&color=orange)](https://www.npmjs.com/package/@yolk_vat-y/dsh-project-memory) [![Listed on dsh-plugin.org](https://dsh-plugin.org/badges/listed.svg)](https://dsh-plugin.org/plugins/00080000/dsh-project-memory) [![Awesome](https://awesome-dsh-plugin.com/badge.svg)](https://awesome-dsh-plugin.com)
8
8
 
9
9
 
10
10
  A persistent **project development memory** for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (dsh) agents. Built specifically for project development, natively integrated with dsh's task system: task lists and files read during a session are automatically persisted as cross-session task records, with tasks ↔ files linked — workflows can be switched and resumed, no need to re-scope the whole project, solving context loss. Documents (PDF/Markdown/txt) and code symbols are stored separately per workspace; documents are automatically cross-linked to the code symbols they mention. Experience notes (problem → solution) are automatically deduplicated, preventing repeated mistakes. All data is stored per project on disk, survives session compaction and handover; recalls include `path:line` citations for source verification. Only one dependency, no vector DB, no native builds.
11
11
 
12
- > The plugin keeps a compact project **memory** on disk, with every entry pointing to a concrete file and line — the agent can reorient quickly instead of re-reading the whole project. Tasks and experience persist across session compactions and handovers.
12
+ > The plugin keeps a compact project **memory** on disk, with every entry pointing to a concrete file and line — the agent can reorient quickly instead of re-reading the whole project.
13
13
 
14
- ![alt text](docs/images/image.png)
15
- The workflow panel is collapsible, automatically adapts to dsh and theme plugin styles, and offers four card style options to switch between.
16
- ![alt text](docs/images/image-4.png)
17
- ## Features
18
-
19
- - **TaskBridge: cross-session development tasks** — the plugin watches each session's live todo list (`todo_write` events) and file reads (`tool/call`): progress snapshots (`steps`) and touched files sync into durable per-project task entities. An unbound session that writes a todo auto-creates a task. Associated files are kept in **recency-weighted order (written/edited first; a read never outranks a written file)** so a resumed session sees at a glance where to look. New sessions continue by `list_tasks` → `select_task` (bind / rename / unarchive); `query_memory` gains `type: 'task'` and appends a task-count hint to `type: 'all'` results. The user-side `/tasks` command shows the task stack, step progress, involved files, and the current session binding. A task named by the model via `select_task(title=…)` keeps that title; one **auto-created** by the first `todo_write` is titled from its **first list entry** (≤48 chars), falling back to the first human message (the part after the last colon), then `Untitled Task`. Sessions spawned as **subagents** are excluded from auto-creation (`origin: 'subagent'` / `delegationDepth > 0`); merging delegated work back into a task is deliberately unbuilt — see §11 in Design tradeoffs. Capacity is project-size adaptive (`fileCount/20`, clamped 5–100). Storage: `.dsh-project-memory/tasks.json` + `binding.json`. Auto-sync requires a dsh build with session events + `todo_write` (verified on 0.1.2-alpha.x, re-verified against the 0.1.5-rc.1 host surface); on older hosts the task tools still work as a plain record list.
20
- - **Task Panel (v0.4.2+): Floating task panel in dsh web** — built on the real dsh web 0.1.5-rc.1 client plugin contract (cordis inject + apply, registered into host `shell.overlay` slot). Draggable cards show steps/files (click to copy path); collapse to a draggable mini-bar; hide completely (summon with `/task` / `/tasks`). Render errors have error boundaries — panel crash no longer takes down the host.
21
- - **Task Panel Behavior** —
22
- - **Default hidden**: panel does not show on dsh web startup
23
- - **Explicit summon**: type `/tasks` or `/task` (list form) to open; model calls `show_task_panel` tool to open
24
- - **Session switch**: only syncs data in background, **does not** auto-open panel
25
- - **Page refresh**: panel stays hidden (UI state `closed` not persisted)
26
- - **Manual close**: click × to fully hide (no mini-bar); reopen requires explicit summon
27
- - **Collapse to mini-bar**: click ↓ to keep draggable top bar; click bar to expand
28
- - **Hide hints**: click ? to suppress every hover tooltip in the panel (drag handle, style/view/minimize/close, rename, step status, copy path, mini-bar, memory view); the preference is stored in localStorage and survives a refresh; the button dims while hints are off — click again to restore
29
- - **Bidirectional task-list sync (host ↔ plugin tasks, v0.4.2+)** — `select_task` or `/task switch` pushes task steps to host `todo/write` so dsh's rendered task list mirrors the plugin's task entity. Config `tasklist.syncHostOnAdopt` (default on) to toggle. Empty `todo/write` means "clear": unbound session clears list without creating junk tasks; bound session clears that task's steps (task retained). Panel edits (step text/status) = write back bound task + push host list, sharing one code path with model `todo_write`. `/task` subcommands: `switch`, `archive`, `unbind`, `rename`, `todos` (invoked by panel buttons/clicks, not the model); `unbind` also clears the host task list above the input.
30
- - **Panel editing & themes (v0.4.2+)** — bound cards: double-click title/step for inline edit (input auto-grows); click step status icon to cycle todo→in-progress→done. Non-bound cards read-only. **Four visual themes** (click folder icon left of title, persisted locally): Native / Glassmorphism / Brutalist / Terminal monospace — only material, geometry, typeface, density change; colors always use dsw alias tokens, follow host light/dark and theme plugins.
31
- - **Document memorization** — PDF, Markdown, and plain text files are chunked and summarized **without any model call**: each entry keeps a ≤300-character `summary` for injection, a bounded (≤160) deterministic, stop-word-filtered `terms` set that covers the **entire chunk** (search-only, so recall is not limited to the opening lines), and a `path:line` citation back to the source. The legacy `blindSpots` field is always empty now that indexing never calls a model; it is kept only so stores written by older versions still load.
32
- - **Code symbol memory (L1 regex)** — a dependency-free scanner extracts functions, classes and methods with full signatures (generics, parameter/return types, overloads) plus interfaces and type aliases across 8 languages, producing one-line identity signatures `fn(a: A, b: B): R — file.ts:42`. It masks strings/comments, joins multi-line signatures, is indentation-aware for Python and carries class-method context — with zero LLM tokens.
33
- - **Optional TypeScript semantic enhancement (L2/L3)** — when `typescript` is installed in the user project (`npm i -D typescript`), the plugin automatically activates a second layer (L2) that uses the TS Compiler API to infer return types, resolve generics, extract interfaces and type aliases, and enrich arrow functions — all asynchronously in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`). Results are cached on disk keyed by file content hash (L3) for instant cold-start reuse. Zero config: just install TS (5.x or 6.x) and restart dsh. Fully optional; if TS is absent or disabled via `enableTypeScript: false`, the plugin falls back to L1 regex-only extraction.
34
- - **Automatic refresh** — a background poll (`watch_repo`) detects new or changed files by content hash and re-memorizes only those.
35
- - **Read-time memorization** — files are memorized the moment the model actually reads them (`fs/observed`), so the memory is a byproduct of normal work, not a separate upfront scan. Files that are never read are never indexed. The project root is detected by markers (`.git`, `package.json`, …), a README plus source directories, or the file's own directory as a last resort.
36
- - **Doc ↔ code cross-linking** — when a document mentions a symbol, the match is recorded as a `reference`; querying a symbol also surfaces the documents that describe it.
37
- - **BM25 memory recall** — ranked search over documents, symbols, and experience notes, with optional LLM query expansion to handle vocabulary mismatch. **CJK-optimized**: precise phrase boost (3+ char phrases ×1.5 score on title/keywords match), synonym table (e.g. 数据库连接池 ↔ 连接池 ↔ DB pool), and CJK-aware word boundaries for doc↔symbol linking.
38
- - **Experience notes** — problems → solutions; similar problems supersede instead of duplicating, and notes are returned only when a search matches. The note store is bounded: capacity scales with project size (clamped to 100–2000), and the oldest notes are pruned when the limit is exceeded. **Supersede tightened to bidirectional 0.7 overlap** (was 0.6); **experience `problem` field now participates in CJK phrase boost** for long-tail query recall.
39
- - **v0.5 tiered insight memory (lessons / decisions / procedures)** — one `insight` entity across three scopes: `task` (private drafts in `tasks.json`), `project` (`.dsh-project-memory/insights.json`), `global` (`~/.config/dsh-project-memory/global.json`). `save_lesson` writes any scope; dedupe is bidirectional token overlap ≥ 0.7 (merge) with a 0.65–0.7 reinforce band; **promotion is a scope change, not a copy** — 2 tasks hitting the same insight promote it to project, 3+ to global. Archive is soft (`archived`), decay/capacity prune archived entries only; writes are filtered for secret/token-shaped content. LLM **reflection is off by default** and only ever writes task-level drafts (`source: reflect`) on task switch-away/archive. Panel gains a Task / Project / Global memory view with approve, promote/demote, archive/restore, delete, edit and a create form (procedures can carry an “as Skill” trigger). Old `experience.json` notes are imported into `insights.json` once, non-destructively. Every kind can carry an authored `trigger`: **only `when` can trigger**, `guard` can only narrow, and `prevents` states what breaks without the entry. `when.ops` are normalized action ids resolved from the tool call itself (`file-write` / `file-delete` / `git-commit` / `release` / `npm-publish` / `render-doc` / `run-bench` / …), `when.writes` are the files this step is about to **write**, `when.intents` are intent words from the human message **after stripping quoted/path references and filenames**. A hit injects the entry deterministically **before the action**. Legacy `keywords` / `symbols` / `actions` / `paths` / `scope` are auto-migrated in memory (actions → `ops`, concrete paths → `writes`, keywords → `intents`, extension/name globs and dead action ids dropped) — an entry left with **no** triggerable member is no longer pushed; run `npm run selfcheck:triggers` to see which ones those are.
40
- - **Streaming TF + IDF caching** — query path caches IDF (term inverse frequency) per store version; on cache hit, single-pass streaming scores 20k entries (5k files) in p50 2.6 ms / p95 5.4 ms — and 4k entries (1k files) in p50 0.6 ms / p95 1.6 ms — with zero intermediate objects. Only a **dirty** write bumps the version and drops the cache — a no-op `save()` returns before touching the disk, so the 15 s watch poll can never clear the cache a query just built.
41
- - **Lock-free sync transactions** — all writes (index / watch / remember / forget / watch_repo) go through synchronous transactions `store.commit(fn)`; fn succeeds then atomic write; the JS single-threaded event loop guarantees no interleaving (**in-process only** — see Consistency); `remember`/`forget` are never blocked by watch re-indexing.
42
- - **Minimal dependencies** — pure JavaScript; the only runtime dependency is `pdfjs-dist` (PDF text extraction), no native builds required.
43
- - **Negligible overhead** — pure in-process operation; a 5k-file store loads in 40 ms, and a cached query over 20k entries is p50 2.6 ms / p95 5.4 ms (4k entries: p50 0.6 ms / p95 1.6 ms); the bottleneck is PDF extraction and disk I/O, not the plugin's scoring.
44
-
45
- ## Performance
46
-
47
- ### Synthetic Benchmark (Node 24.19, WSL2 on 20 vCPU, Linux file system)
48
-
49
- | Scenario | Scale | Measured |
50
- |----------|-------|----------|
51
- | Full cold index | 5,000 files / 20k entries | 269 ms avg (p50 267) |
52
- | Cold load | 5,000 files | 40 ms |
53
- | Hot lazy re-index (single file) | 5k files | p50 2.4 ms / max 5.5 ms |
54
- | query_memory (cached) | 5k files / 20k entries | p50 2.6 ms / p95 5.4 ms |
55
- | query_memory (cached) | 1k files / 4k entries | p50 0.6 ms / p95 1.6 ms |
56
- | Full cold index | 10,000 files / 40k entries | 551 ms avg (p50 528) |
57
- | Cold load | 10,000 files | 90 ms |
58
- | Hot lazy re-index (single file) | 10k files | p50 5.4 ms / max 9.2 ms |
59
-
60
- > Synthetic benchmark: generated code (~4–5 symbols/file), Node 24.19 on WSL2 / 20 vCPU / Linux file system, measured 2026-09-14. Reproduce with `npm run bench:synthetic -- 5000` (harness: `scripts/bench-synthetic.mjs`). Measures pure indexing overhead without LLM calls. query_memory uses the IDF cache + precomputed searchText; the first query after a write rebuilds IDF (**106 ms at 40k entries**, 57 ms at 20k, 12 ms at 4k), subsequent queries hit the cache.
61
-
62
- ### Real Project Storage
63
-
64
- | Project | Files | Entries | Store Size | Per Entry |
65
- |---------|-------|---------|------------|-----------|
66
- | Java Spring Boot backend | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
67
- | Vue 3 + Vite frontend | 289 | 2,141 | 1.0 MB | ~0.5 KB |
68
-
69
- > Real projects (Java + Vue), tested on Linux file system (Node 24). Real project entries are smaller than synthetic benchmarks due to lower symbol density and shorter declarations.
14
+ ![Task panel: task list, step progress, and involved files](docs/images/image.png)
70
15
 
71
- ### Reproduce it on your own project
72
-
73
- Rather than asking you to trust the numbers above, the measurement itself ships with the repository **and with the published npm package** (`scripts/` is part of the tarball). It needs **no dsh instance, no network and no model calls**, and it never touches your project's own store — results go to a temp directory and are removed when it finishes:
74
-
75
- ```bash
76
- npm run bench -- /path/to/your/project
77
- # or, with options:
78
- node scripts/bench.mjs /path/to/your/project [--json] [--samples 100] [--no-pdf] [--keep]
79
- ```
80
-
81
- It reports the cold index split into read+hash / extract / commit, cold load, IDF rebuild, cold and hot query latency (p50/p95/max over 100 sampled queries through the shipped scorer), single-file hot re-index, store size and bytes per entry. Example — our internal Vue project (289 files / 2,141 entries, Node 24, 20 CPU, Linux):
82
-
83
- ```
84
- cold index 253 ms (read+hash 9 ms · extract 229 ms · commit 13 ms) ← 2nd, warm-cache run
85
- store 1.10 MB · 538 bytes/entry · cold load 4.6 ms
86
- hot query p50 0.80 ms · p95 1.35 ms (2,141 entries)
87
- re-index 1 file p50 0.33 ms
88
- ```
89
-
90
- Two caveats we would rather state than hide: `read+hash` depends on the OS page cache — on that corpus the first run spent 787 ms and the second 253 ms, so say which run you quote — and **real projects score slower than the synthetic table above** — on a 3,000-file slice of a large TypeScript repository (15,594 entries) hot queries were p50 7.5 ms, because real declaration text is longer than generated stubs. Pass `--queries your-queries.json` to run the same labeled-set method (hit@5 / hit@10 / MRR) against your own project.
91
-
92
- ## How it works
93
-
94
- The design follows four principles:
95
-
96
- - **Volatility** — context is ephemeral; it is lost when a session is compacted.
97
- - **Persistence** — the **memory** is stored on disk and survives compaction and new sessions.
98
- - **Compactness** — the code layer stores one declaration line per symbol, so code-heavy projects stay near **0.5% of the source** (8.8 MB of source → 49 KB of index in the example project), and **recall** replaces re-reading the full file. The document layer is heavier by design: each chunk keeps a ≤300-char injected `summary`, a bounded `terms` set covering the whole chunk for retrieval, and a precomputed `searchText`. Measured on a docs-only corpus (179 chunks / 225 KB of Markdown): `terms` ≈ **27.5%** of source and the on-disk store ≈ **166%** of source — so on doc-heavy projects budget for roughly the docs themselves, not 0.5%.
99
- - **Verifiability** — **recalls** carry a `path:line` citation where applicable, so the agent can confirm details against the source.
100
-
101
- Building the **memory** does not require an upfront scan: files are memorized as the model reads them, so the **memory** grows to cover exactly what has been worked with. Re-reading a file that has not changed is a no-op (content hash), so the **memory** stays fresh with minimal ongoing overhead.
16
+ ## Features
102
17
 
103
- The store is per-project and follows the codebase: changed files are re-extracted by content hash, deleted files are removed. Experience notes are retrieval-only, so accumulation does not affect context.
18
+ - **TaskBridge: cross-session development tasks** — the session's todo list and the files it touches are persisted as durable per-project task entities, so a workflow can be switched and resumed without re-scoping the project. Associated files are kept in recency-weighted order (a read never outranks a written file), so a resumed session sees where to look first. New sessions continue through `list_tasks` → `select_task`. Work delegated to subagents does not create tasks (see Known limits and boundaries). Auto-sync needs a dsh build with session events; on older hosts the task tools still work as a plain record list.
19
+ - **Task panel in dsh web (v0.4.2+)** — draggable cards show a task's steps and files, collapse to a mini-bar, or hide entirely. The panel stays hidden until summoned, syncs in the background on session switch, and does not reopen itself after a page refresh. Render errors are contained, so a panel failure cannot take down the host.
20
+ - **Panel editing and themes (v0.4.2+)** — bound cards allow inline editing of title and steps and status cycling; unbound cards are read-only. Four visual themes change material, geometry, typeface and density only; colours follow the host.
21
+ - **Bidirectional task-list sync (v0.4.2+)** — binding a task pushes its steps to the host task list, and panel edits write back through the same code path as model updates. Set `tasklist.syncHostOnAdopt` to false to opt out.
22
+ - **Document memory** — PDF, Markdown and plain text are chunked and summarized without a model call. Each entry keeps a short summary for injection, a bounded search-only term set built from the whole chunk so recall is not limited to the opening lines, and a citation back to the source.
23
+ - **Code symbol memory** — a dependency-free scanner extracts functions, classes, methods, interfaces and type aliases with full signatures across 8 languages, one line per declaration. Where `typescript` is installed, an optional second layer infers return types, resolves generics and extracts interfaces; it runs asynchronously, is cached by content hash, and never blocks indexing.
24
+ - **Automatic refresh** — a background poll detects new and changed files by content hash and re-memorizes only those.
25
+ - **Read-time memorization** — a file is memorized the moment the model reads it, so memory is a byproduct of normal work rather than a separate upfront scan. Files that are never read are never indexed.
26
+ - **Doc ↔ code cross-linking** — a document that mentions a symbol is surfaced when that symbol is queried.
27
+ - **BM25 recall** — ranked search over documents, symbols, experience notes and insights, with optional LLM query expansion. Tuned for CJK: phrase boost, a synonym table and CJK-aware boundaries for doc↔symbol linking.
28
+ - **Experience notes** — problems → solutions, deduplicated by overlap rather than repeated, bounded by project size, and returned only when a search matches.
29
+ - **Tiered insight memory (lessons / decisions / procedures, v0.5)** — one entity across task, project and global scope. Near-duplicates merge or reinforce; promotion moves an entry between scopes rather than copying it. Use is recorded, so decay and capacity rank by activity rather than by age alone. LLM reflection is off by default and writes task-level drafts only. The panel exposes a per-scope memory view for reviewing and editing entries.
30
+ - **Triggered injection** — an insight may carry an authored trigger: only `when` triggers, `guard` narrows it, and `prevents` records what breaks without the entry. Corpus text can never trigger an injection; the statistical channel is gated separately.
31
+ - **Streaming TF + IDF caching** — the query path caches term weights per store version (measured numbers under Performance). Only a real write drops the cache, so the watch poll never clears one a query just built.
32
+ - **Lock-free sync transactions** — all writes go through a synchronous transaction, so `remember` and `forget` never queue behind re-indexing. The lock is in-process: avoid pointing two dsh instances at the same store.
33
+ - **Minimal dependencies** — pure JavaScript; one runtime dependency for PDF text extraction, no native builds.
34
+ - **Negligible overhead** — memory work is in-process; the bottleneck is document extraction and disk I/O, not scoring.
104
35
 
105
36
  ## Installation
106
37
 
@@ -133,21 +64,40 @@ The tools below are **invoked by the agent**, not typed by the user. In the chat
133
64
  | Tool | Purpose |
134
65
  |---|---|
135
66
  | `index_doc file_path` | Index one document (PDF/MD/txt): chunk → deterministic `summary` + whole-chunk `terms` → store with `path:line`. Unchanged files are skipped. |
136
- | `index_repo root` | Index a whole project: docs get deterministic summaries + whole-chunk terms, code files get a zero-token symbol table. Incremental, cleans up deleted files, cross-links docs to symbols. A root that does not exist — including a Windows-style path resolved on Linux/macOS — is rejected before anything is written. |
137
- | `watch_repo root` | Enable automatic refresh: a background poll detects new/changed files (mtime + content hash) and re-indexes only those. Watched roots persist across plugin restarts; a non-existent root, the filesystem root and the shared temp directory are all refused, and roots that disappear are dropped instead of being re-created. |
67
+ | `index_repo root` | Index a whole project: docs get deterministic summaries + whole-chunk terms, code files get a zero-token symbol table. Incremental, cleans up deleted files, cross-links docs to symbols. A root that does not exist — including a Windows-style path resolved on Linux/macOS — is rejected before anything is written, and so is a *dangerous* root (home directory, filesystem root, system / package-manager prefixes): scanning one of those walks hundreds of thousands of files. |
68
+ | `watch_repo root` | Enable automatic refresh: a background poll detects new/changed files (mtime + content hash) and re-indexes only those. Watched roots persist across plugin restarts; a non-existent root and a dangerous root (filesystem root, home directory, shared temp directory, system / package-manager prefixes) are all refused, roots that disappear are dropped instead of being re-created, and a polluted watchlist from an older version is self-healed on startup. |
138
69
  | `memory_stats root` | Show what the store contains: totals (files / entries / experience notes), last index time, and the per-file list sorted by recency. |
139
70
  | `query_memory query` | BM25 search over docs + symbols + experience + insights (lessons / decisions / procedures), optionally query-expanded by the LLM. `type` selects a layer (`all` / `doc` / `symbol` / `experience` / `insight` / `task`). Returns ranked hits with relative scores, sources or insight ids, and doc→symbol references. |
140
71
  | `list_tasks` | List task records for the project (archived marked). Call first in a new session before continuing work. |
141
72
  | `select_task` | Bind the session to a task so its todo list and file reads sync into it. Exact `taskId`, or exact `title` (multiple matches return candidates; no match creates a new task). Pass `title` with `taskId` to rename. Auto-unarchives. |
142
73
  | `archive_task` | Archive a task (hide from default views, exclude from capacity, stop syncing). `select_task` restores it. |
143
74
  | `show_task_panel` | Show the task panel in the UI. Call when the user asks to see the task list or when you want to display the panel. |
144
- | `/tasks` (typed by the user, not the model) | Shows the task stack: title, step progress, involved files, and which task the current session is bound to. |
145
- | `/task` (typed by the user, not the model) | Task panel subcommands: `switch` / `archive` / `unbind` / `rename` / `todos`. Invoked by panel buttons/clicks; does not go through the model. |
146
- | `/insight` (typed by the user, not the model) | v0.5 memory view actions (panel buttons): `list [task|project|global]`, `confirm` / `promote` / `demote` / `archive` / `restore` / `delete` `<scope> <id>`, `save <scope> <json>`, `edit <scope> <id> <json>`. |
75
+ | `/tasks` (typed by the user, not the model) | The only user command: shows the task stack (title, step progress, involved files, current session binding) and drives the workflow card. Every other action is a **sub-verb** invoked by card buttons, never typed: `/tasks switch` / `archive` / `unbind` / `rename` / `todos …` (task actions), `/tasks insight list` / `confirm` / `promote` / `demote` / `archive` / `restore` / `delete` / `save` / `edit …` (memory actions). |
147
76
  | `remember problem solution` | Save an experience note. Similar problems supersede instead of duplicating. |
148
77
  | `forget id_or_query` | Delete stale experience notes. |
149
78
  | `save_lesson` (agent tool) | Save a lesson/decision/procedure at task/project/global scope (single insight entity). Near-duplicates merge (≥ 0.7 overlap) or reinforce (0.65–0.7); 2+ tasks hitting the same insight auto-promote task → project, 3+ → global. Params: `title`, `kind`, `scope`, `pattern`/`fix` or `choice`/`reason` or `steps`, `trigger` (`when` = `ops`/`writes`/`intents`, the only trigger surface; `guard` = `paths`/`not_paths`/`hosts`/`tags`, narrowing only; `prevents` = what breaks without it; legacy `keywords`/`symbols`/`actions`/`paths`/`scope` still accepted and auto-migrated), `task_id`, `files`, `symbols`, `confidence`, `root`. |
150
79
 
80
+ `/tasks` appears in the web `/` menu inside an icon-bearing **Workflow** group, offering three view
81
+ entries — Tasks / Project Memory / Global Memory (labelled in the UI language). A host command always
82
+ shows up in the built-in **Commands** section and a plugin cannot hide it (`commands.list()` and
83
+ `commands.execute()` read the same view, and `CommandDefinition` has no hidden flag), so the plugin
84
+ registers **only `/tasks`** and every other action rides it as a sub-verb driven by card buttons: the
85
+ menu duplication is one row. Typing `/tasks` + Enter still executes immediately; typing an argued line
86
+ such as `/tasks switch x` is no longer recognised as a command — use the card buttons.
87
+
88
+ ## How it works
89
+
90
+ The design follows four principles:
91
+
92
+ - **Volatility** — context is ephemeral; it is lost when a session is compacted.
93
+ - **Persistence** — the **memory** is stored on disk and survives compaction and new sessions.
94
+ - **Compactness** — the code layer stores one declaration line per symbol, so code-heavy projects stay near **0.5% of the source** (8.8 MB of source → 49 KB of index in the example project), and **recall** replaces re-reading the full file. The document layer is heavier by design: each chunk keeps a ≤300-char injected `summary`, a bounded `terms` set covering the whole chunk for retrieval, and a precomputed `searchText`. Measured on a docs-only corpus (179 chunks / 225 KB of Markdown): `terms` ≈ **27.5%** of source and the on-disk store ≈ **166%** of source — so on doc-heavy projects budget for roughly the docs themselves, not 0.5%.
95
+ - **Verifiability** — **recalls** carry a `path:line` citation where applicable, so the agent can confirm details against the source.
96
+
97
+ Building the **memory** does not require an upfront scan: files are memorized as the model reads them, so the **memory** grows to cover exactly what has been worked with. Re-reading a file that has not changed is a no-op (content hash), so the **memory** stays fresh with minimal ongoing overhead.
98
+
99
+ The store is per-project and follows the codebase: changed files are re-extracted by content hash, deleted files are removed. Experience notes are retrieval-only, so accumulation does not affect context.
100
+
151
101
  ## Design
152
102
 
153
103
  ```
@@ -160,6 +110,8 @@ The tools below are **invoked by the agent**, not typed by the user. In the chat
160
110
  tasks.json TaskBridge task entities (cross-session)
161
111
  binding.json current session ↔ task binding
162
112
  insights.json v0.5 project-scope insights (lessons/decisions/procedures); v0.4 experience notes imported once, non-destructively
113
+ injection-audit.jsonl one line per real injection (what / why / dropped / budget)
114
+ admission-shadow.jsonl one line **per step**, all scored candidates + features (offline replay, labels)
163
115
  ```
164
116
 
165
117
  Stores created before v0.2.0 (single `entries.json` / `index.json`) migrate automatically and idempotently on first load. Within one dsh process, all tool calls share a single in-memory store per project, so hot-path indexing writes only the shard that changed.
@@ -179,99 +131,9 @@ TaskPanel (Container)
179
131
  └── TaskComponents (MiniBar, TaskCard — presentational only)
180
132
  ```
181
133
 
182
- ## Design tradeoffs
183
-
184
- These are deliberate scope choices.
185
-
186
- ### 1. Synchronous lock-free transactions over async locks
187
-
188
- **We do:** All writes go through `store.commit(fn)` — a synchronous in-process transaction. The callback `fn` performs all validation and mutations; only on success is the result atomically written to disk. The JS event loop guarantees no interleaving. CAS (`applyFileUpdate`) makes concurrent writes idempotent.
189
-
190
- **We don't:** Async mutexes, file locks, or multi-process coordination.
191
-
192
- **Why:** DSH runs on Cordis, which is single-process by design. Adding locks would complicate the hot path (every `remember`/`forget`/`index_doc` call) for a scenario (multi-process DSH) that would require a breaking ecosystem change. Synchronous transactions keep the hot path at ~2 ms median with zero contention overhead in practice.
193
-
194
- ### 2. Watch: compute outside, commit inside
195
-
196
- **We do:** Heavy work (mtime/hash/scan/parse/PDF extraction) runs outside the transaction; a single `commit` applies all changes atomically. On failure, the snapshot rolls back so the next poll retries automatically.
197
-
198
- **We don't:** Hold a lock during parsing, or use `fs.watch` events.
199
-
200
- **Why:** PDF extraction and large-file parsing take time — holding a lock would block `remember`/`forget`/`query_memory`. Polling with mtime+content-hash is platform-agnostic (works on network drives, Docker volumes, WSL) and avoids the "double fire / missed events" nightmare of `fs.watch`.
201
-
202
- ### 3. Corrupt files are quarantined, not auto-repaired
203
-
204
- **We do:** On JSON parse failure, the bad file is renamed to `*.corrupt`, an error is logged, and that file's store starts fresh. The rest of the store remains intact.
205
-
206
- **We don't:** Write-ahead logs, embedded databases (SQLite/LMDB), or automatic partial recovery.
207
-
208
- **Why:** A corrupted shard means *one source file* has a bad index — quarantining it costs near zero. A WAL or embedded DB adds a heavy dependency, increases binary size, and introduces new failure modes (lock contention, corruption of the WAL itself). The tradeoff: lose one file's index vs. add 500 KB+ of native code.
209
-
210
- ### 4. No vector embeddings, no semantic search at query time
211
-
212
- **We do:** BM25 with CJK phrase boost (3+ chars ×1.5 on title/keywords), synonym expansion (bidirectional table), field weighting (title ×5), and experience-layer phrase boost. All at query time, zero LLM calls.
213
-
214
- **We don't:** Vector embeddings, dense retrieval, rerankers, or hybrid search.
215
-
216
- **Why:** Vectors require an embedding model (local = heavy, remote = latency + cost + privacy), a vector index (HNSW/IVF = memory + build time), and reranking (another LLM call). For the queries this plugin targets, lexical BM25 is already sufficient and measurable: on our benchmark suite (29 queries over a real Vue project) file-level hit@5 is **96.6%**, and 28 of the 29 are exact symbol lookups that lexical search answers essentially always. Whole-chunk `terms` took document-term coverage from **27.3% to 100%** while queries that already worked kept their ranking (MRR **0.958** vs **0.955**). Those figures come from an internal Vue project with a hand-labeled 29-query set, so they are not reproducible outside it — but the **method** now ships as `scripts/bench.mjs --queries <your-set.json>`, so you can run the identical measurement on your own project. The marginal gain from semantic search doesn't justify the 10x complexity/cost increase.
217
-
218
- ### 5. Indexing is deterministic and model-free
219
-
220
- **We do:** Derive keywords with a rule (title-weighted top terms) and build a whole-chunk `terms` set — both deterministic and reproducible. Doc↔symbol links surface English symbol names from Chinese queries, and CJK tokenization keeps cross-language hits working. With `llmQueryExpansion: false`, queries never touch the LLM.
221
-
222
- **We don't:** Call a model at index time to translate or paraphrase a document, and we don't translate queries at search time.
223
-
224
- **Why:** An index-time model call makes indexing slower, non-deterministic and unverifiable — the same document can index differently on two runs. Query-time translation adds latency and a hard failure mode (a bad translation means zero recall). Rules plus symbol linking cover the common cases, work offline, and keep indexing at zero model calls.
225
-
226
- ### 6. Model-facing memory: the agent writes, and no human has to be in the loop
227
-
228
- **We do:** Treat the agent as a first-class writer. `remember` / `save_lesson` write **any scope at any time** (`task` / `project` / `global`) with no human step, and promotion is deterministic and runs inside the ordinary write path: cross-task token-overlap dedupe accumulates `sourceTaskIds`, then `promoteAllTasksToProject` / `promoteProjectToGlobal` move an entry up once its corroboration counts are met (≥2 tasks for project, ≥ `globalPromoteTasks` — 3 by default — for global). Nothing waits on the task panel: a user who never opens the UI still gets a memory that fills, dedupes and graduates.
229
-
230
- **We do (labeling):** Keep inferred content distinguishable from recorded content. The v0.5 `reflection` path (opt-in, **off by default**) is the only writer that infers rather than records: it writes task-scoped drafts stamped `draft: true` / `source: 'reflect'`, and `recall` plus silent injection skip `draft` entries while they remain drafts.
231
-
232
- **We don't:** Require human approval for memory to become useful, or make the UI a step in the write path. `draft` is a **provenance label plus a corroboration threshold**, not an approval queue.
233
-
234
- **Why:** The agent is the consumer and it is usually headless — memory that only graduates when a human clicks a card is memory that never graduates. Labeling keeps the useful half of the caution (inferred ≠ recorded, and unreviewed single-task inference stays out of the prompt) without taxing the normal path. A draft graduates on corroboration: a second task matching it through the model's own writes, or the model writing the same knowledge at project scope, which links the existing entry instead of duplicating it.
235
-
236
- ### 7. Full entries returned directly
237
-
238
- **We do:** `query_memory` returns complete entries with `path:line` citations. Every hit can be verified against source.
239
-
240
- **We don't:** Return a minimal index first, then require a second tool call for details.
241
-
242
- **Why:** Returning full entries preserves **verifiability** — the agent sees the exact source line for every claim. It also avoids a round-trip per useful hit. Our entries are already compact (~300-char summary + citation, plus a search-only `terms` field that never enters the prompt); the token cost is lower than a second tool call + context switch.
243
-
244
- ### 8. Symbol extraction focused on what developers search for
245
-
246
- **We do:** Regex-based symbol extraction (functions, classes, methods, interfaces, type aliases) with string/comment masking, multi-line signatures, and cross-file linking by symbol name. For TypeScript/JavaScript projects, an optional L2 enhancement layer uses the TS Compiler API to infer return types, resolve generics, and extract interfaces — all cached by content hash for instant reuse.
247
-
248
- **We don't:** Tree-sitter AST parsing, import graphs, call graphs, or full-program type resolution across files.
249
-
250
- **Why:** Our regex scanner handles 8 languages with zero dependencies, runs in <1 ms/file, and captures the declarations developers actually search for (names, signatures, generics). The optional TS layer adds semantic depth for TS/JS without native deps. Cross-file linking by name covers the most common "find related code" use case. Full-program analysis would add native binaries, 10x install size, and version fragility — for marginal gain on the remaining 5% of edge cases.
251
-
252
- ### 9. `forget` by query is aggressive; prefer ID deletion
253
-
254
- **We do:** `forget query` deletes all experience notes with ≥0.5 token overlap.
255
-
256
- **We don't:** Interactive confirmation, soft-delete/trash, or exact-match-only.
257
-
258
- **Why:** Experience notes are low-stakes, high-volume, and retrieval-only. Aggressive deletion prevents stale noise from polluting search. For precision, delete by ID (shown in `query_memory` output).
259
-
260
- ### 10. TypeScript enhancement is optional, lazy, and cached
261
-
262
- **We do:** L2 TS Compiler API enhancement runs async in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`), results cached by content hash in `type-cache/`. Zero config — just `npm i -D typescript@5` or `typescript@6`. Falls back to L1 regex if TS absent or disabled.
263
-
264
- **We don't:** Mandatory TS, blocking enhancement, or full-program type checking.
265
-
266
- **Why:** Mandatory TS would break installs for non-TS projects. Blocking enhancement would stall `index_repo` on large codebases. Full-program checking is 10x slower and memory-heavy. Our design: enhance what's read, cache it, never block the hot path.
267
-
268
- ### 11. Subagent sessions are out of scope for now
269
-
270
- **We do:** Exclude sessions spawned as subagents (`origin: 'subagent'` / `delegationDepth > 0`) from auto-creating or binding a task. Their `todo_write` events do not create tasks, and they inherit no task binding.
271
-
272
- **We don't:** Merge a delegated run's steps and files back into the task that spawned it. That is **not designed yet**: there is no parent-link model for delegated work, and the naive version mints one project task per subagent.
134
+ The workflow panel is collapsible, automatically adapts to dsh and theme plugin styles, and offers four card style options to switch between.
273
135
 
274
- **Why:** Every subagent that writes a todo would otherwise create its own task entity, so one fan-out run would flood the task list with ephemeral entries nobody resumes. Excluding them keeps the task list equal to the work the user actually owns. The cost is that a delegation's progress is invisible in the task record; merging it properly (child steps folded into the parent, or a separate delegated-work view) is future work.
136
+ ![Four card styles](docs/images/image-4.png)
275
137
 
276
138
  ## Configuration
277
139
 
@@ -291,19 +153,29 @@ These are deliberate scope choices.
291
153
  | `autoIndexOnFirstUse` | false | full scan of the current working directory on plugin load (opt-in) |
292
154
  | `watch` | true | enable the background refresh |
293
155
  | `watchInterval` | 15 | poll interval (seconds) |
156
+ | `maxScanFiles` | 20000 | hard cap on files per scan pass; a truncated scan is reported and never deletes the entries it did not reach. Set `0` to disable the cap (at your own risk) |
157
+ | `maxScanDepth` | 12 | hard cap on directory depth per scan pass. Set `0` to disable |
158
+ | `allowUnsafeRoots` | false | allow **explicit** tool calls (`index_repo`/`watch_repo`/`remember` with a `root`) to target a dangerous root. Automatic paths (lazy indexing, session audit, TaskBridge, `autoIndexOnFirstUse`) stay inert in these directories regardless |
294
159
  | `tsPath` | (auto) | optional absolute path to a specific `typescript` install; if omitted, resolves from project cwd → plugin node_modules |
295
160
  | `enableTypeScript` | true | set `false` to disable L2 TS enhancement entirely (L1 regex only) |
161
+
162
+ ### Memory and injection knobs
163
+
164
+ | 键 | 默认值 | 含义 |
165
+ |---|---|---|
296
166
  | `insight.*` | dedupOverlap `0.7` · reinforceBand `0.65` · maxProject `100` · maxGlobalProcedures `200` · promoteConfidence `0.7` · globalPromoteTasks `3` · decayDays `90` · `globalFile` (auto) | v0.5 insight dedupe / reinforce / promotion / capacity / archive settings |
297
167
  | `reflection.enabled` | false | v0.5 LLM reflection, **draft-only at task level** (fires on task switch-away / archive). `cooldownMs` `1800000`, `maxLessonsPerReflect` `3`, `maxDecisionsPerReflect` `2` |
298
- | `autoContext.enabled` | true | silent injection wrapper (resident task card + gated items). Inert (full passthrough) until the host exposes a resolvable session cwd; `maxTokens` `400`, `editedMax` `3` (how many recently-written "editing now" files the resident task card shows), `signalMinRatio` `0.5` (a hint must reach half of its layer's top score), `skipEchoSelfTodo` `true` (don't echo the task card back when the model itself maintains the task list with no newer human message; relevant insights still inject), `budgetLog` `off` (budget-drop audit on stderr: `off` silent / `once` at most one line per session / `all` one line per changed dropped set), `reinjectItemsAfter` `0` (cooldown, in pre-steps, before the same insight may be injected again) |
168
+ | `autoContext.enabled` | true | silent injection wrapper (resident task card + gated items). Inert (full passthrough) until the host exposes a resolvable session cwd; `maxTokens` `400`, `editedMax` `3` (how many recently-written "editing now" files the resident task card shows), `signalMinRatio` `0.5` (a hint must reach half of its layer's top score), `skipEchoSelfTodo` `true` (don't echo the task card back when the model itself maintains the task list with no newer human message; relevant insights still inject), `budgetLog` `off` (budget-drop audit on stderr: `off` silent / `once` at most one line per session / `all` one line per changed dropped set), `reinjectItemsAfter` `0` (cooldown, in pre-steps, before the same insight may be injected again), `rootNotice` `true` (when the memory root is inferred from a marker-less working directory, tell the model once where memory lives and how to change it) |
299
169
  | `autoContext.gateCooldownSteps` | 2 | **admission knobs.** Minimum number of pre-steps between two *item* injections (the resident task card is exempt — it is a state snapshot and should update when it changes). This is the main "don't inject often" dial |
300
170
  | `autoContext.maxItemsPerSession` | 12 | hard per-session cap on injected items; the budget is a ceiling, not a target — once exhausted the item channel stays silent |
301
171
  | `autoContext.maxItemCharsPerSession` | 4000 | same, in characters |
302
- | `autoContext.hintMinCoverage` | 0.3 | **absolute** floor for the statistical (hint) channel: IDF-weighted share of the query's information mass the entry covers. A ratio-only threshold cannot tell signal from "best of a bad lot" (`relative:1.00` on an unrelated entry) |
172
+ | `autoContext.hintMinCoverage` | 0.45 | **absolute** floor for the statistical (hint) channel: IDF-weighted share of the query's information mass the entry covers. A ratio-only threshold cannot tell signal from "best of a bad lot" (`relative:1.00` on an unrelated entry). Raised from 0.30 in 0.5.8: on a real 43-entry store the control scenario injected 3 unrelated hints at cov 0.32–0.35, because a same-corpus store flattens IDF |
303
173
  | `autoContext.hintMinMatched` | 2 | a hint must share at least this many terms with the query — one generic word ("plugin") is not evidence |
304
174
  | `autoContext.hintMinSupport` | 0.15 | channel-level silence: if less than this share of the query's terms exist anywhere in the corpus, the hint channel says nothing this round — a long sentence that happens to share one word otherwise reports `cov:1.00` |
305
175
  | `autoContext.legacyScope` | `filter` | how to treat a legacy `trigger.scope`: `filter` keeps the old semantics, `ignore` drops it. `npm run selfcheck:triggers` reports entries whose scope values cannot intersect the project tag space |
306
176
  | `autoContext.auditLog` | true | append one JSONL line per **actual** injection to `<root>/.dsh-project-memory/injection-audit.jsonl` (what was injected, why it matched, what was dropped, session budget snapshot); rotates to `.1` past `auditMaxBytes` (`262144`). Silent on any I/O error — never affects the host request |
177
+ | `autoContext.shadowLog` | true | append one JSONL line **per step** (including steps that injected nothing) to `admission-shadow.jsonl`: every scored candidate with its judgement features (`rel` / `coverage` / `matched` / `support` / `terms` / `decision`) plus the step's `query` / `ops` / `writes`. This is what makes a threshold change answerable offline on real history (`decision` shows which gate rejected each candidate). Rotates past `shadowMaxBytes` (`2097152`). Disk only — never enters the prompt, costs no tokens |
178
+ | `autoContext.shadowMaxBytes` | 2097152 | rotation cap for `admission-shadow.jsonl` |
307
179
 
308
180
  ### Injection admission (why it stays quiet)
309
181
 
@@ -314,12 +186,22 @@ Automatic injection used to be a *retrieval* problem ("which entry is most relat
314
186
  - **Ratio *plus* an absolute floor.** The hint channel needs the relative score *and* an IDF-weighted coverage floor *and* at least two shared terms — `relative:1.00` also happens on entries that share nothing with the step.
315
187
  - **Frequency is bounded.** At most one item injection every `gateCooldownSteps`, capped per session by count and characters. The resident task card is exempt (it is a snapshot that should update); the budget is a ceiling, not a target.
316
188
  - **Prefix-cache discipline.** Injections are appended as a user message at the tail of the history, so the cached prefix is never rewritten. What they add is resident *cache-read* tokens, not cache misses; nothing is ever edited in place.
317
- - **It is auditable.** Every real injection appends one line to `injection-audit.jsonl` (reason, dropped candidates, session budget), and `npm run eval:injection` scores 8 labelled scenarios — currently precision 1.00 / recall 1.00 with a clean control group.
189
+ - **It is auditable.** Every real injection appends one line to `injection-audit.jsonl` (reason, dropped candidates, session budget), and `admission-shadow.jsonl` adds one line **per step** — including the steps that correctly injected nothing — with every candidate's features and the gate that rejected it. That second file is what makes a threshold question answerable offline instead of by re-running the agent. `npm run eval:injection` scores 8 labelled scenarios on a **synthetic** pool — currently precision 1.00 / recall 1.00 with a clean control group. That pool is the CI baseline, not evidence about your data: point the same harness at your own store and the control group becomes a **hard gate** (`--store`, exits non-zero on violation). That is how the 0.45 floor was chosen, and how you can re-choose it (`--hint-cov <n>` replays at another floor).
318
190
 
319
191
  ### Toggling features
320
192
 
321
193
  The two most relevant switches are `lazyIndexing` (index a file the moment the model reads it; default on) and `autoIndexOnFirstUse` (full scan of the current working directory on plugin load; default off). Lazily indexed project roots are automatically registered with the watcher, so changed files stay fresh without an explicit `watch_repo`.
322
194
 
195
+ **Where the project root comes from.** One policy, applied identically by lazy indexing, the session audit trail, TaskBridge and every tool: an explicit `root` argument wins; otherwise an explicitly registered root (`watch_repo`); otherwise the nearest ancestor containing a VCS marker (`.git`/`.hg`/`.svn`) or a build/manifest marker (`package.json`, `go.mod`, `Cargo.toml`, `pyproject.toml`, …); otherwise **the session working directory itself, provided it is a safe directory**. So a marker-less scratch folder you started dsh in still gets project memory — the plugin just says so once:
196
+
197
+ ```
198
+ memory root: /Users/me/scratch (inferred from the session working directory; no project marker found).
199
+ If project memory should live elsewhere, pass `root: <dir>` to index_repo / watch_repo / remember / query_memory,
200
+ or restart dsh inside the project directory.
201
+ ```
202
+
203
+ That notice goes out once per session and can be muted with `autoContext.rootNotice: false`. What the plugin will **not** do is promote an arbitrary directory to a project: reading a stray file outside the working directory records nothing, and a dangerous root (filesystem root, your home directory, the shared temp directory or a system / package-manager prefix — `/opt/homebrew` on POSIX, `%SystemRoot%`/`%ProgramFiles%`/`%ProgramData%` on Windows) is refused outright — that is what used to walk an entire home directory and exhaust memory. Sessions whose working directory is one of those run with memory disabled (one stderr line explains why).
204
+
323
205
  Settings live in the plugin's config object. To change them, add an override entry to your profile's `cordis.patch.yml` — for the web profile that is `~/.dsh/profiles/web/cordis.patch.yml`:
324
206
 
325
207
  ```yaml
@@ -330,7 +212,10 @@ Settings live in the plugin's config object. To change them, add an override ent
330
212
  llmQueryExpansion: false # off: do not spend tokens on LLM query expansion (default)
331
213
  watch: true # on: background refresh for watched roots (default)
332
214
  watchInterval: 15 # poll interval in seconds
215
+ maxScanFiles: 20000 # per-scan file cap (truncation is reported, never deletes)
216
+ maxScanDepth: 12 # per-scan directory-depth cap
333
217
  enableTypeScript: true # on: L2 TS enhancement when TS is installed (default)
218
+ # allowUnsafeRoots: false # keep false unless you really want to index a home/system dir explicitly
334
219
  # budgetLog: once # debugging: log budget drops to stderr (default off = silent)
335
220
  # reinjectItemsAfter: 20 # debugging: allow the same insight again after N steps (default 0 = once per session)
336
221
  # tsPath: /custom/path/to/typescript # optional: force specific TS install
@@ -346,15 +231,76 @@ dsh web --patch ./config.yml
346
231
 
347
232
  where `config.yml` contains the same override block.
348
233
 
234
+ ## Performance
235
+
236
+ ### Synthetic Benchmark (Node 24.19, WSL2 on 20 vCPU, Linux file system)
237
+
238
+ | Scenario | Scale | Measured |
239
+ |----------|-------|----------|
240
+ | Full cold index | 5,000 files / 20k entries | 269 ms avg (p50 267) |
241
+ | Cold load | 5,000 files | 40 ms |
242
+ | Hot lazy re-index (single file) | 5k files | p50 2.4 ms / max 5.5 ms |
243
+ | query_memory (cached) | 5k files / 20k entries | p50 2.6 ms / p95 5.4 ms |
244
+ | query_memory (cached) | 1k files / 4k entries | p50 0.6 ms / p95 1.6 ms |
245
+ | Full cold index | 10,000 files / 40k entries | 551 ms avg (p50 528) |
246
+ | Cold load | 10,000 files | 90 ms |
247
+ | Hot lazy re-index (single file) | 10k files | p50 5.4 ms / max 9.2 ms |
248
+
249
+ > Synthetic benchmark: generated code (~4–5 symbols/file), Node 24.19 on WSL2 / 20 vCPU / Linux file system, measured 2026-09-14. Reproduce with `npm run bench:synthetic -- 5000` (harness: `scripts/bench-synthetic.mjs`). Measures pure indexing overhead without LLM calls. query_memory uses the IDF cache + precomputed searchText; the first query after a write rebuilds IDF (**106 ms at 40k entries**, 57 ms at 20k, 12 ms at 4k), subsequent queries hit the cache.
250
+
251
+ ### Real Project Storage
252
+
253
+ | Project | Files | Entries | Store Size | Per Entry |
254
+ |---------|-------|---------|------------|-----------|
255
+ | Java Spring Boot backend | 1,254 | 7,335 | 6.7 MB | ~0.9 KB |
256
+ | Vue 3 + Vite frontend | 289 | 2,141 | 1.0 MB | ~0.5 KB |
257
+
258
+ > Real projects (Java + Vue), tested on Linux file system (Node 24). Real project entries are smaller than synthetic benchmarks due to lower symbol density and shorter declarations.
259
+
260
+ ### Reproduce it on your own project
261
+
262
+ Rather than asking you to trust the numbers above, the measurement itself ships with the repository **and with the published npm package** (`scripts/` is part of the tarball). It needs **no dsh instance, no network and no model calls**, and it never touches your project's own store — results go to a temp directory and are removed when it finishes:
263
+
264
+ ```bash
265
+ npm run bench -- /path/to/your/project
266
+ # or, with options:
267
+ node scripts/bench.mjs /path/to/your/project [--json] [--samples 100] [--no-pdf] [--keep]
268
+ ```
269
+
270
+ It reports the cold index split into read+hash / extract / commit, cold load, IDF rebuild, cold and hot query latency (p50/p95/max over 100 sampled queries through the shipped scorer), single-file hot re-index, store size and bytes per entry. Example — our internal Vue project (289 files / 2,141 entries, Node 24, 20 CPU, Linux):
271
+
272
+ ```
273
+ cold index 253 ms (read+hash 9 ms · extract 229 ms · commit 13 ms) ← 2nd, warm-cache run
274
+ store 1.10 MB · 538 bytes/entry · cold load 4.6 ms
275
+ hot query p50 0.80 ms · p95 1.35 ms (2,141 entries)
276
+ re-index 1 file p50 0.33 ms
277
+ ```
278
+
279
+ Two caveats we would rather state than hide: `read+hash` depends on the OS page cache — on that corpus the first run spent 787 ms and the second 253 ms, so say which run you quote — and **real projects score slower than the synthetic table above** — on a 3,000-file slice of a large TypeScript repository (15,594 entries) hot queries were p50 7.5 ms, because real declaration text is longer than generated stubs. Pass `--queries your-queries.json` to run the same labeled-set method (hit@5 / hit@10 / MRR) against your own project.
280
+
281
+ ## Design tradeoffs
282
+
283
+ - **Synchronous lock-free transactions over async locks** — no async mutexes, file locks, or multi-process coordination: DSH runs on Cordis and single-process is an architectural given, so locking for a rare multi-process case would only slow the hot path (every `remember`/`forget`/`index_doc`); synchronous transactions keep that path at ~2 ms median with zero contention.
284
+ - **Watch: compute outside, commit inside** — no lock is held during parsing and `fs.watch` is not used: parsing and PDF extraction are slow, so a held lock would block queries, while polling with mtime + content hash behaves identically on network drives, Docker volumes, and WSL, with none of `fs.watch`'s duplicate-trigger/missed-event failure modes.
285
+ - **Corrupt shards are quarantined, not repaired** — a shard that fails to parse is renamed `*.corrupt` and only that file is re-indexed, leaving every other shard untouched; no WAL or embedded database: those add 500 KB+ of native dependencies, lock contention, and a new failure mode (a corrupt WAL) to avoid losing a single file's index.
286
+ - **No vector embeddings, no semantic search at query time** — no embedding model, vector index (HNSW/IVF), or reranker: lexical retrieval already answers the queries this plugin targets. On a real Vue project with 29 labelled queries, file-level hit@5 is **96.6%**; whole-chunk `terms` lift document term coverage from **27.3% to 100%** with MRR unchanged (**0.958** vs **0.955**). The marginal gain does not justify 10x the complexity, and the method ships with the code — `scripts/bench.mjs --queries your-queries.json` reproduces the same measurement on your own project.
287
+ - **Indexing is deterministic and model-free** — no model at index time and no translation at query time: the former makes two indexings of one document differ, the latter has a hard failure mode (a wrong translation means zero recall); rules plus symbol links already cover the common cases and work offline.
288
+ - **Model-facing memory: the agent writes, no human in the loop** — no human approval step: the consumer of this memory is the agent, and agents are usually headless, so memory that only promotes when someone clicks a card would never promote at all. `draft` is a provenance marker plus an evidence threshold, not an approval queue — the one inferring writer, `reflection` (off by default), writes task-level drafts only, and drafts never reach recall or injection.
289
+ - **Full entries returned directly** — no "minimal index first, fetch details in a second call": entries are already compact, so returning them whole is both more verifiable and one round-trip cheaper.
290
+ - **`forget` by query is aggressive; use IDs for precision** — no confirmation prompt, recycle bin, or exact-match-only mode: experience notes are low-risk, high-volume, and retrieval-only, so stale noise hurts more than an over-broad delete. For exact deletion use the ID shown by `query_memory`.
291
+ - **TypeScript enhancement is optional, lazy, and cached** — the L2 TS Compiler API runs asynchronously on a priority queue (P0 `fs/observed`, P1 `watch`, P2 `index_repo`) and caches results by content hash; TS is never required and enhancement never blocks: requiring it would make non-TS projects uninstallable, and blocking would stall `index_repo` on large projects. `npm i -D typescript@5|6` is the entire setup, and a missing TS falls back to the L1 regex scanner.
292
+ - **Subagent sessions are out of scope for now**
293
+
349
294
  ## Development (for contributors)
350
295
 
351
296
  These commands are for **maintaining the plugin code** — regular users do not need them. Installing the plugin only requires the command in [Installation](#installation).
352
297
 
353
298
  ```bash
354
299
  npm install
355
- npm test # 331 tests (184 core + 16 TaskBridge + 11 insight-store + 9 insight-actions + 8 doc-index + 7 auto-inject + 9 host-contract + 5 reflection + 4 llm-route + 2 client-hints + 8 recall + 14 readiness + 7 insight-derive + 6 readiness-eval + 6 ops + 6 injection-audit + 5 injection-budget + 6 injection-scenarios + 18 bugfix-0.5.7)
356
- npm run eval:injection # scenario P/R: 14/14 hits, 0 false positives, control group clean
357
- npm run selfcheck:triggers # which entries can still push, which declarations are dead
300
+ npm test # 360 tests (184 core + 16 TaskBridge + 12 insight-store + 9 insight-actions + 8 doc-index + 7 auto-inject + 9 host-contract + 5 reflection + 4 llm-route + 2 client-hints + 8 recall + 14 readiness + 7 insight-derive + 7 readiness-eval + 6 ops + 8 injection-audit + 5 injection-budget + 6 injection-scenarios + 18 bugfix-0.5.7 + 3 client-icons + 10 client-slash + 5 workflow-command + 7 client-session-id)
301
+ npm run eval:injection # scenario P/R on the synthetic pool: 14/14 hits, 0 false positives, control group clean
302
+ npm run eval:injection -- --store .dsh-project-memory/insights.json # replay on YOUR store; control group is a hard gate
303
+ npm run selfcheck:triggers # which entries can still push, which declarations are dead (reads your local store)
358
304
  npm run bench -- /path/to/project # index/query performance on any project — no dsh needed
359
305
  ```
360
306