@yolk_vat-y/dsh-project-memory 0.5.4 → 0.5.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +128 -0
- package/README.md +98 -42
- package/README.zh-CN.md +98 -42
- package/client/client.js +111 -92
- package/client/client.js.map +1 -1
- package/package.json +7 -2
- package/scripts/bench-store.mjs +113 -0
- package/scripts/bench-synthetic.mjs +364 -0
- package/scripts/bench.mjs +396 -0
- package/src/audit.js +87 -0
- package/src/auto-inject.js +367 -51
- package/src/client/MemoryView.tsx +4 -2
- package/src/client/TaskComponents.tsx +7 -6
- package/src/client/TaskPanel.module.css +7 -0
- package/src/client/TaskPanel.tsx +12 -8
- package/src/client/task-ui-store.ts +11 -1
- package/src/index.js +19 -2
- package/src/insight-store.js +26 -4
- package/src/ops.js +210 -0
- package/src/parsers/pdfjs-parser.js +3 -0
- package/src/readiness.js +468 -0
- package/src/recall.js +217 -0
- package/src/store.js +20 -1
- package/src/tools/index-repo.js +7 -1
- package/src/tools/lesson-tools.js +31 -4
- package/src/tools/query-memory.js +61 -23
- package/src/tools/watch-repo.js +10 -1
- package/src/util/fs.js +41 -0
- package/src/watch.js +28 -4
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,133 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.5.6 (2026-09-16) — injection admission (lessons/decisions/procedures stop arriving by coincidence)
|
|
4
|
+
|
|
5
|
+
### Changed (only what you are about to *do* can trigger an injection)
|
|
6
|
+
|
|
7
|
+
- **Trigger matching was a substring test over a haystack of everything the step had seen** (`humanText + tool arguments`, i.e. file contents included). Measured on a real session: 10 injections, **0 of them useful**, 5921 characters appended permanently — including a "recover deleted files" procedure pulled in by the literal string `dcterms` (it contains `rm`) and an "arXiv fetching" note pulled in by the `.pptx` file extension. A labelled 8-scenario evaluation set (now `npm run eval:injection`) scored **precision 0.48 / recall 0.72**.
|
|
8
|
+
- **A trigger now has a shape: `when` (ops / writes / intents) triggers, `guard` only narrows, `prevents` states what breaks without the entry.** `ops` are normalized action ids resolved from the tool call itself (`src/ops.js`: `file-write` / `file-delete` / `git-commit` / `release` / `npm-publish` / `render-doc` / `run-bench` / …); `writes` are the files this step is about to write; `intents` are human-message words matched only **after** stripping quoted spans, paths and filenames, and only above a minimum length/word-boundary rule (so `ppt` no longer matches `pptx`).
|
|
9
|
+
- **Corpus text can no longer trigger anything.** File contents are out of the haystack entirely, extension/name globs (`*.pptx`, `README*`) are ignored outright, and machine identifiers (`web_fetch`, `LAYOUT_16x9`) are filtered out of intent words.
|
|
10
|
+
- **Legacy triggers migrate in memory, idempotently and without rewriting your data**: `actions` → `when.ops` (dead ids mapped where possible, otherwise recorded), concrete `paths` → `when.writes`, `keywords` → `when.intents`, extension/name globs dropped, and weak actions (`git-add` / `git-commit`) dropped when the entry already has precise paths. `npm run selfcheck:triggers` reports what changed — against the live stores: **20 pushable / 4 pull-only, 8 dead action ids, 10 dropped globs, 3 dropped weak actions**.
|
|
11
|
+
- **Result on the evaluation set: precision 1.00, recall 1.00, and the control scenario (rename a `.pptx` timestamp, with the filename in quotes) injects nothing at all.**
|
|
12
|
+
|
|
13
|
+
### Changed (hint channel: relative *and* absolute, plus a frequency budget)
|
|
14
|
+
|
|
15
|
+
- **The statistical channel required only "half of this layer's top score"**, which a ranking satisfies even when the top score is itself noise — measured `relative:1.00` on entries sharing nothing with the step. It now also needs an **IDF-weighted coverage floor** (`hintMinCoverage`, default `0.3`) **and** at least two shared terms (`hintMinMatched`, default `2`). Query text is the human message *plus this step's write targets* — raw tool arguments are no longer a query.
|
|
16
|
+
- **Injections are now events, not a heartbeat**: `gateCooldownSteps` (default `2`) bounds how often the item channel may speak, and `maxItemsPerSession` / `maxItemCharsPerSession` cap the session. The resident task card is exempt (it is a state snapshot), and the budget is a ceiling rather than a target — when nothing clears the gates, nothing is injected.
|
|
17
|
+
- **Cache discipline is explicit**: injections are appended as a user message at the tail of history, so the cached prefix is never rewritten; the cost they add is resident cache-read tokens, not cache misses. Nothing is edited in place.
|
|
18
|
+
|
|
19
|
+
### Added (observability: this change is measured, not asserted)
|
|
20
|
+
|
|
21
|
+
- **`injection-audit.jsonl`** — one JSONL line per *actual* injection under `<root>/.dsh-project-memory/`, recording what went in, why it matched (`op:` / `write:` / `intent:` / hint coverage), what lost the budget (`budget` / `quota` / `cooldown` / `coverage:` / `thin:`), and the session budget snapshot. Rotation at `auditMaxBytes`; every failure is swallowed so the host request is never affected. Config: `autoContext.auditLog`.
|
|
22
|
+
- **`npm run eval:injection`** — 8 labelled scenarios (publish, trigger-debugging, stale source, benchmark, deleted doc, interview prep, and the clean control) over a 24-entry synthetic pool whose triggers are copied verbatim from the real ones. Asserts the ratchet, precision ≥ 0.90 and a clean control group. `--store <insights.json>` replays a real store; `--selfcheck` prints the trigger audit.
|
|
23
|
+
- **New suites**: `test/ops.test.mjs` (the action plane, including a regression for reading argument *values* rather than key names), `test/injection-budget.test.mjs` (cooldown/caps actually silence, 6 steps with 6 fresh entries → 3 injections), `test/injection-audit.test.mjs` (JSONL shape, rotation, silent failure, end-to-end wiring through `agent/pre-step`). Suite is now **313 tests**.
|
|
24
|
+
- **`trigger.when` / `trigger.guard` / `trigger.prevents`** are declared in the `save_lesson` schema, so new entries can be authored in the new shape.
|
|
25
|
+
|
|
26
|
+
## 0.5.5 (2026-09-15)
|
|
27
|
+
|
|
28
|
+
### Fixed (the README promised a benchmark the npm package did not contain)
|
|
29
|
+
|
|
30
|
+
- **`scripts/` was missing from `package.json`'s `files` whitelist, so the published tarball shipped a README that told users to run a script it did not include.** `README.md` / `README.zh-CN.md` are inside the package, and their *Performance → Reproduce it on your own project* section (plus the Development block) documents `npm run bench -- /path/to/your/project` and `node scripts/bench.mjs` — but `npm pack` contained only `src/`, `client/`, `cordis.patch.yml` and the three documents: no `scripts/bench.mjs`, no `scripts/bench-synthetic.mjs`. Anyone installing the plugin from npm and following the reproduce instructions found the file absent, which is exactly the "trust the numbers" position that section exists to avoid. `scripts` is now in `files`, so both harnesses ship with the package; they still need no dsh instance, no network and no model calls, since they import only `src/`. Verified with `npm pack --dry-run` (tarball now lists `scripts/bench.mjs`, `scripts/bench-synthetic.mjs`, `scripts/bench-store.mjs`).
|
|
31
|
+
- **`bench/` stays local-only** — it is git-ignored and has never been tracked. It holds the raw measurement logs (`bench/results/`) and the scratch profilers; the synthetic harness it used to contain is `scripts/bench-synthetic.mjs`. Note this is a *packaging* fix only: those harnesses were already on GitHub, so "the README references a script that is not in the repository" was never true for the git remote — it was true only for the npm tarball.
|
|
32
|
+
- **Both READMEs advertised a stale test count.** They still said `280 tests` (the pre-0.5.5 number) while the three unreleased fixes above took the suite to **286**; the figures are now the measured ones (core 180 → 184, host-contract 7 → 9, everything else unchanged).
|
|
33
|
+
|
|
34
|
+
### Fixed (silent injection re-sent memory that was already in context)
|
|
35
|
+
|
|
36
|
+
- **Dedupe was whole-block, so a benign change anywhere re-sent every item.** The only guard was a fingerprint of the whole injected block, and that block changes for reasons that have nothing to do with the memory being new: the resident task card advances, the observed tool-call window slides, and `fitBody` re-truncates the last item whenever the remaining budget shifts. Measured on a real 66-step session: **20 injections carrying 41 item-sends for 9 unique items — 78% of them re-sends**, including the same 1732-char procedure three times and another 1008-char one six times.
|
|
37
|
+
- **Items are now remembered per session** (`sessionId → insightId → { hash, step }`) and an unchanged, already-injected item is filtered out of `buildInjection` **before** budget scheduling, so the budget it used to occupy goes to new items instead. The hash is over the injected body, so editing an insight re-injects it immediately. Only items that actually made it into an appended message are recorded — a budget-dropped item was never seen and must stay eligible. Per-session memory is capped (600 items) like the other session tables.
|
|
38
|
+
- **Why "once per session" is safe as the default:** the host session history is append-only — `agent-loop` has no compaction, and `agent/inbox/spliced` only touches the *pending* inbox. A re-send is therefore pure duplication. `autoContext.reinjectItemsAfter` (default `0`) exists for setups where history is trimmed outside the plugin: `N > 0` allows the same item again once `N` pre-steps have passed (it still has to change the block to be observable, since an identical block is already in history).
|
|
39
|
+
- **Regression:** two handler-level checks drive one session across steps and assert that a task-card advance no longer re-sends the procedure, that a fully unchanged step injects nothing at all, and that `reinjectItemsAfter: 2` re-injects on the third step; plus one schema check for the default. Suite 283 → **286** checks.
|
|
40
|
+
|
|
41
|
+
### Changed (the budget-drop audit is now opt-in — the user's terminal stays clean)
|
|
42
|
+
|
|
43
|
+
- **`auto-inject` no longer prints `degraded: … insights kept out by budget` unless the operator asks for it.** The line exists so a budget drop can never be *silent*, but the terminal is the **user's** surface, and a drop is **normal priority-based degradation, not a failure**: the resident task card and authored triggers take the 1200-char budget first, and lower-priority hints lose out. On a real host one `dsh web` start printed five such lines (four from a single busy project store, one global), which reads as "the plugin is broken" — exactly the first impression that gets a plugin uninstalled. New `autoContext.budgetLog` selects the level: `off` (**default**, never prints), `once` (at most one line per session, on the first drop), `all` (one line per changed dropped set — the previous behaviour). Accounting itself is unchanged in every mode: the per-session dedupe and the 200-session cap still run, so switching the log on later still reflects the current state. `README.md` / `README.zh-CN.md` document the key in both the `autoContext` row and the `cordis.patch.yml` example.
|
|
44
|
+
- **Regression:** the degraded check now drives two steps of one session (two different dropped sets) and asserts: silent by default, silent for `off`, ≥2 lines for `all`, exactly 1 line for `once`. Three schema checks were added as well — the *resolved default* is `off`, an explicit `budgetLog` survives `Config` and reaches the engine, and an unknown value is rejected — because the last config-default bug in this file was exactly this class (a key the engine reads but the schema does not declare). Suite 280 → **283** checks.
|
|
45
|
+
|
|
46
|
+
### Fixed (the silent-injection hint gate was dead in production)
|
|
47
|
+
|
|
48
|
+
- **A config schema default pinned the `legacy` hint gate on, making the shipped relative threshold dead code.** `autoContext.relevanceMin` carried `default(0.25)`, so `cfgEngine()` always saw a number and therefore always chose the legacy branch — `normalizedTokenOverlap >= 0.25`, the very judgement PR2 replaced because natural-language messages measure 0.014–0.057 against an insight and can never pass it. The documented default (`signalMinRatio: 0.5`, BM25 relative to the layer's top score) was unreachable in a real host, so only authored triggers could ever inject. The default is removed (the key remains, for explicitly opting back into the old behaviour), and `entryMaxInsights` / `signalMinRatio` / `editedMax` / `skipEchoSelfTodo` are now declared in the schema rather than being read by the engine but absent from it. The test suite could not catch this because every readiness test builds its config object by hand and never goes through `Config`.
|
|
49
|
+
- **Regression:** 4 new checks asserting that the *resolved default* config yields `relevanceMin: null` and `signalMinRatio: 0.5`, that an explicit `relevanceMin` still restores the legacy gate, and that the engine-only keys round-trip through the schema. Suite 276 → **280**.
|
|
50
|
+
|
|
51
|
+
### Added (standalone benchmark — `scripts/bench.mjs`)
|
|
52
|
+
|
|
53
|
+
- **`npm run bench -- <projectPath>` measures the shipped index/query path on a real project, with no dsh instance, no network and no model calls.** It walks the project, builds its store in a temp directory (never the project's own `.dsh-project-memory`), and reports: cold index split into read+hash / extract / commit; cold load measured from a *copied* store (so the in-process store cache cannot fake it); IDF rebuild, cold and hot query latency through the real scorer (p50/p95/max over `--samples`, default 100); single-file hot re-index sampled **evenly across file sizes**; store size and bytes/entry; and whole-chunk `terms` versus the ≤300-char summary. Flags: `--json`, `--no-pdf`, `--keep`, `--max-files` (default 20000), `--queries <labeled-set.json>` (runs the hit@5 / hit@10 / MRR method on your own corpus).
|
|
54
|
+
- **The READMEs now point at it** (`Performance → Reproduce it on your own project`, plus the Development block) instead of asking readers to trust the tables. The labeled-set figures in Design tradeoff #4 are scoped to their internal corpus, with the method shipped; the docs also state the two caveats found while building it — `read+hash` is OS-page-cache sensitive, and real projects score slower than the synthetic corpus.
|
|
55
|
+
- **Retired the local-only harnesses it replaces** (`test/ablation*.mjs`, `test/doc-coverage*.mjs`, `test/gen-queries.mjs`, `test/simulate-agent.mjs`, `test/queries.json`) together with the dead `.gitignore` entries for them and for the August debug scaffolding.
|
|
56
|
+
- **The synthetic harness ships too.** `bench/microbench.mjs` moved to `scripts/bench-synthetic.mjs` (`npm run bench:synthetic -- 5000`), so the synthetic table is reproducible as well; `bench/` itself stays local-only.
|
|
57
|
+
- **Every published number re-measured on 2026-09-14** (Node 24.19, WSL2, 20 vCPU): 5k-file cold index 269 ms avg (p50 267), cold load 40 ms, cached query over 20k entries p50 2.6 / p95 5.4 ms, hot single-file re-index p50 2.4 ms, 10k-file cold index 551 ms avg, cold load 90 ms, IDF rebuild 106 ms at 40k entries (57 ms at 20k, 12 ms at 4k), symbol scan 12.9–19.2 µs/file. Table captions now name the hardware and the reproduce command; the real-project example quotes the warm-cache run and states the cold-cache first run (787 ms vs 253 ms) instead of hiding it. `README.md`, `README.zh-CN.md` and the local interview cheat sheet were updated together so no two documents quote different numbers.
|
|
58
|
+
|
|
59
|
+
### Docs (README realigned with shipped behaviour)
|
|
60
|
+
|
|
61
|
+
- **`README.md` / `README.zh-CN.md` re-synced with the code.** Removed claims that documented superseded behaviour: the client-plugin contract version (0.1.2-rc.1 → **0.1.5-rc.1**, bumped in v0.5.3); the auto-created task title rule (v0.5.3 reversed it to **first todo entry → first human message**, and added the subagent exception); and the write path (`O(1) version bump` → only a **dirty** write bumps the version, so the 15 s watch poll cannot clear the IDF cache). `blindSpots` is now described as always empty rather than "empty for newly indexed documents"; the IDF-rebuild note carries its scale; `index_repo` / `watch_repo` document the missing-root guards. De-duplicated two feature bullets that restated each other (document `terms`; L1 regex vs. symbol memory). The Chinese README no longer claims multi-instance writes are safe — it now carries the same in-process-only lock caveat as the English one.
|
|
62
|
+
- **Two design-tradeoff entries corrected.** (a) "Subagent sessions never auto-create tasks" was stated as a feature; it is a deliberate non-goal, so it moved to a new tradeoff **§11 "Subagent sessions are out of scope for now"** ("not designed yet", with the cost stated) and the feature bullet now only points at it. (b) **§6 "Explicit `remember` over implicit learning" was factually wrong**: v0.5 ships `reflection-pipeline.js`, an opt-in LLM path that *does* infer lessons/decisions from the conversation on task switch-away/archive, and promotion runs automatically inside normal writes. §6 is now **"Model-facing memory: the agent writes, and no human has to be in the loop"** and states the actual policy: any scope writable by the agent at any time, promotion deterministic on `sourceTaskIds` corroboration (≥2 → project, ≥ `globalPromoteTasks` → global), and **`draft` described as a provenance label plus a corroboration threshold — not an approval queue** (a draft graduates on corroboration, which is intended; the panel is an optional inspection surface, not a step in the write path).
|
|
63
|
+
|
|
64
|
+
### Fixed (console noise: watching the temp dir, PDF warnings, repeat failure logs)
|
|
65
|
+
|
|
66
|
+
- **The shared temp directory is no longer watchable.** `WatchManager.addRoot()`, `restorePersisted()` and `watch_repo` now refuse a root that is the filesystem root or `os.tmpdir()` (`isUnwatchableRoot()` in `src/util/fs.js`), and a persisted entry for one is self-healed out of the watchlist on the next start. This was the root cause of a real spam loop: reading any file directly under `/tmp` made `findProjectRoot` fall back to `/tmp` as the project root, `src/lazy.js` then auto-registered it as a watch root, and every 15 s poll re-indexed every file left under `/tmp` — including the test suite's deliberately broken `pm-pdf-*/big.pdf` fixtures, which printed `Invalid PDF structure` on every poll and re-created `/tmp/.dsh-project-memory` forever. Subdirectories of the temp dir (a project that genuinely lives there) still work.
|
|
67
|
+
- **pdf.js warnings are silenced** (`verbosity: 0` added to `PDFJS_OPTIONS`). `Warning: Indexing all PDF objects` is pdf.js's own recovery path for a broken or non-linearized PDF. It is not actionable from here, and a failed parse already reports its own error.
|
|
68
|
+
- **A permanently failing file is reported once, not every poll.** `watch` still retries a failed index every interval by design (so a file that gets fixed is picked up), but it now remembers the last reported error per file (`state.failures`) and stays quiet until the message changes or the file indexes successfully.
|
|
69
|
+
- **Tests:** suite 269 → **276** checks: a repeatedly failing file logs once, `watch_repo` / `addRoot` refuse the shared temp dir, the refusal is not persisted, `addRoot` refuses the filesystem root, and `restorePersisted` drops an unwatchable persisted root.
|
|
70
|
+
|
|
71
|
+
### Fixed (watch / save hot path)
|
|
72
|
+
|
|
73
|
+
- **A no-op `save()` no longer rewrites the store.** `commit()` always called `save()`, and `save()` ran `mkdirSync`, bumped `_version` and cleared the IDF cache even when nothing had changed. Since `watch` calls `commit()` for every watched root every 15 s, this (a) wiped the query-side IDF cache on every poll — cancelling the v0.3.4 IDF-reuse optimization whenever `watch` was on — and (b) **re-created the memory directory of any persisted watch root that no longer existed**, so a deleted root came back every 15 s. `save()` now returns before touching the disk unless something is dirty; stale `*.tmp` cleanup still runs unconditionally.
|
|
74
|
+
- **Dead watch roots are dropped instead of re-created.** `WatchManager.addRoot()` now refuses a root that does not exist, `restorePersisted()` removes such roots from the persisted watchlist on startup, and `watch_repo` rejects a non-existent root rather than persisting it.
|
|
75
|
+
- **Tests:** suite 231 → **235** checks. New: `no-op save() keeps the IDF cache` / `dirty save() invalidates the IDF cache`, plus `restorePersisted skips a root that no longer exists`, `a dead root is dropped from the persisted watchlist` and `addRoot refuses a non-existent root`.
|
|
76
|
+
|
|
77
|
+
### Fixed (Task Panel / 「隐藏提示信息」开关)
|
|
78
|
+
|
|
79
|
+
- **The hints toggle now hides every hover tooltip, not just two.** `showHints` was only wired to the step-edit hint and the drag-cancel button; the drag handle, panel-style / view-cycle / minimize / close buttons, card title rename hint, step status tooltip, file-path copy hint, mini-bar hint and the memory-view hints all ignored it — so clicking the button changed almost nothing. Every `title` hint in the panel now honours the switch (the toggle's own tooltip is kept so it stays discoverable).
|
|
80
|
+
- **The preference lives in the UI store and persists.** `showHints` moved from component-local state to `dsh-pm-task-panel-ui` (localStorage, default on), so it survives refreshes and remounts instead of silently resetting to “show”; the button dims and sets `aria-pressed` while hints are off, giving feedback without hovering.
|
|
81
|
+
|
|
82
|
+
### Fixed (index_repo root validation)
|
|
83
|
+
|
|
84
|
+
- **`index_repo` no longer builds a store for a root that does not exist.** Both the tool and `autoIndexOnFirstUse` only ran `path.resolve()` first: on Linux/macOS a Windows-style path such as `D:\project\foo` resolved to the relative `<cwd>/D:\project\foo`, and the store's `mkdirSync` then created that literal directory together with its `.dsh-project-memory` (verified: `new ProjectMemoryStore(dir).load().commit(…)` creates every missing parent). `assertIndexRoot()` (`src/util/fs.js`) now rejects a missing or non-directory root before anything is written, and names the Windows-path case in the error message. `index_doc`'s file check and `watch_repo`'s `existsSync` guard are unchanged.
|
|
85
|
+
- **Tests:** `test/doc-index.test.mjs` gains 1 check (suite 7 → 8, total 237 → **238**): a missing root and a Windows-style path are both rejected with zero filesystem side effects.
|
|
86
|
+
|
|
87
|
+
### Added (unified recall: the insight layer becomes searchable — PR1)
|
|
88
|
+
|
|
89
|
+
- **`query_memory` now searches insights.** `type: 'insight'` was added, and `type: 'all'` appends an `## Insights (lessons / decisions / procedures)` section. Before this, lessons/decisions/procedures had **no retrieval path at all**: `util/search.js` had no insight ranker and the tool's `type` enum had no insight value, so a lesson could only reach an agent through silent injection. Verified against this repository's own store: the "public face" lesson (`ins_3dcc7458`) now returns at score 100 for a natural-language query.
|
|
90
|
+
- **New `src/recall.js` — one retrieval core, ranked per layer.** Every layer gets an adapter to the same entry contract (`title` / `keywords` / `summary` / `terms` / `sourcePath` — the shape `weightedFieldText` already scores), its own top-k, and a layer prior. Cross-layer order is only ever a *hint* (`flat`); it is never spliced into one score, because the doc/symbol streamer and the in-memory BM25 have different scales. `query_memory` routes doc / symbol / experience / insight through this single call, so the experience layer no longer carries its own scoring path inside the tool.
|
|
91
|
+
- **Scope visibility follows the session binding.** `project` / `global` insights are visible to every session; `task`-level insights only to a session bound to that task (`select_task`). Drafts stay out of recall. The v0.4 `experience.json` → `insights.json` migration shadow (`kind: 'experience'`, `source: 'migrate'`) is excluded from the insight layer so a note is not listed once as an insight and again as an experience — convergence the store's own comment had deferred to this PR.
|
|
92
|
+
- **Normative-heading prior.** A doc chunk whose heading matches the multilingual normative set (必须 / 禁止 / 铁律 / rules / must / must not / do not …) is boosted ×1.5 within the doc layer: deterministic, model-free, no new dependency. This is the "「必须遵守」never surfaced" case — `test/recall.test.mjs` first pins today's term-frequency ordering, then asserts the prior flips it.
|
|
93
|
+
- **Invariants held:** indexing still never calls a model (no index-time LLM), recall adds no LLM call (`llmQueryExpansion` remains opt-in, default off), and the plugin still has exactly one runtime dependency.
|
|
94
|
+
- **Tests:** new `test/recall.test.mjs` (8 checks); suite 238 → **246**.
|
|
95
|
+
|
|
96
|
+
### Added (readiness: triggers for every insight kind, two windows — PR2)
|
|
97
|
+
|
|
98
|
+
- **`trigger` is no longer procedure-only.** Any lesson / decision / procedure may carry `trigger: { keywords, symbols, actions, paths, scope }`. A hit injects the entry **before the action**, in full, at priority 1 — deterministic, auditable (`reasons[].why`), and independent of wording. `actions` are normalized ids (`git-commit`, `npm-publish`, `go-public`, …) matched by a small lexicon that also reads **human intent** ("提交 / 发包 / 公开"), so the block arrives while the model is still planning instead of after it acted.
|
|
99
|
+
- **Two windows, one matcher.** The pre-emptive window matches the turn's human text; the reactive window matches `tool/call` arguments observed through `session/event`. DSH has no pre-tool interceptor (`tool/call` is appended after the call is decided), so these are the only two honest windows — both feed one `ReadinessContext`.
|
|
100
|
+
- **The hint gate became scale-free.** It used to be `normalizedTokenOverlap(humanMessage, insight) >= 0.25`, normalized by the **larger** token set, so a real message measured 0.014–0.057 and could never fire. It is now the recall core's BM25 plus a **relative** threshold (`autoContext.signalMinRatio`, default `0.5` = at least half of the layer's top score), with zero-score candidates always dropped. The old judgement stays reachable only by explicitly setting `relevanceMin` (undocumented knob; profiles that set it keep the old behaviour).
|
|
101
|
+
- **Budget is a schedule, not a number.** Resident task card → trigger hits (full) → hints (truncated). Every decision is recorded as `reasons` (why injected) or `dropped` (why not), so silence is never the only signal. Procedures still inject **only** through the trigger channel, so a statistical hit cannot bypass `trigger.scope` profile filtering.
|
|
102
|
+
- **Measured against this repository's own store:** the original commit-turn message ("…你顺便提交一下…") now injects the public-face lesson via `hint:relative:1.00`; reworded and relying on an observed `git commit`, only an authored `trigger.actions` reaches it (`trigger:action:git-commit`) — the two channels are complementary, not redundant.
|
|
103
|
+
- **Tests:** new `test/readiness.test.mjs` (10 checks); suite 246 → **256**.
|
|
104
|
+
|
|
105
|
+
### Added (write path + measurement: derived triggers, v1→v2, degraded, labeled eval — PR3)
|
|
106
|
+
|
|
107
|
+
- **Derived triggers are additive and never force-inject.** `deriveTrigger()` builds `{ keywords, actions, paths }` from an entry's own text with the same deterministic lexicon and path scanner the readiness matcher uses (bounded: ≤ 6 / ≤ 4 / ≤ 6), and stores it as `triggerDerived`. It only feeds the entry's search text, so it improves recall; whether an entry is injected **deterministically** is still decided solely by the authored `trigger`. This is the "explicit over implicit" trade-off carried through to injection.
|
|
108
|
+
- **`insights.json` v1 → v2, lazily and idempotently.** On load, entries without `triggerDerived` are back-filled in memory (pure, model-free, no field rewriting); the next commit writes `version: 2`. A second pass is a no-op.
|
|
109
|
+
- **The write path accepts the full trigger.** `normalizeInsight` no longer whitelists `keywords`/`symbols`/`scope` only — it keeps `actions` and `paths` too, and a trigger consisting *only* of `actions`/`paths` is no longer dropped. `save_lesson`'s schema exposes both, and its descriptions now say "any kind", not "procedure".
|
|
110
|
+
- **Silence is not the only signal.** When the budget keeps entries out, `installAutoInject` logs one `degraded` record per session and drop-set (`… kept out by budget — id(reason), …`) instead of silently shrinking the injection.
|
|
111
|
+
- **The relative threshold is now a measured parameter, and the measurement ships.** `test/readiness-eval.test.mjs` holds a labeled set (8 cases: trigger-on-intent, trigger-on-observed-action, path glob, keywords, procedure scope filter, strong-vs-weak hint, no-match, budget scheduling) and sweeps `signalMinRatio`:
|
|
112
|
+
|
|
113
|
+
| signalMinRatio | precision | recall |
|
|
114
|
+
|---|---|---|
|
|
115
|
+
| 0.20 | 0.78 | 1.00 |
|
|
116
|
+
| 0.35 | 1.00 | 1.00 |
|
|
117
|
+
| 0.50 (shipped) | 1.00 | 1.00 |
|
|
118
|
+
| 0.70 | 1.00 | 1.00 |
|
|
119
|
+
|
|
120
|
+
The shipped default sits in the safe zone with margin, and 0.20 visibly admits weak matches — so the number is justified by data rather than chosen by feel. The table is the **9-case** set measured after the hint-precision fix below (the `0.20` figure was 0.86 on the earlier 8-case set; adding "缩写巧合(PR)不得命中" turns one more row into a genuine weak match, `7/9 = 0.78`). It lives under `test/` (CI-enforced, shipped) rather than the git-ignored `bench/`, deliberately: a threshold that only exists on one machine is not a threshold.
|
|
121
|
+
- **Tests:** new `test/insight-derive.test.mjs` (6 checks) and `test/readiness-eval.test.mjs` (4 checks, including the sweep table); suite 256 → **266**.
|
|
122
|
+
|
|
123
|
+
### Fixed (readiness hint precision: query source, acronym noise, stub truncation)
|
|
124
|
+
|
|
125
|
+
- **The readiness query is the last *human* message again.** `lastUserText()` accepted any `role: 'user'` message — and the injection itself is appended as a user message, so the previous step's injected block became the next step's query (**self-query**). Visible symptom: the injected set kept changing in a way the human message could not explain. It now prefers `source.kind === 'user'` and keeps the sourceless fallback for older hosts and tests.
|
|
126
|
+
- **A 1–2 character latin token is no longer hint evidence.** Measured on this repository's own store, the message "PR 干嘛…" put an unrelated "automatic security PR" lesson at the top of the insight layer on the strength of the acronym `PR` alone, i.e. `relative:1.00`. `hintQueryText()` drops *standalone* 1–2 character latin tokens; path-shaped tokens are preserved so `src/util/fs.js` still contributes `src`/`util`. Authored triggers are untouched: deterministic matching stays literal.
|
|
127
|
+
- **No stub truncations.** A body that cannot fit its channel's useful minimum (trigger 48 chars, hint 120 chars) is dropped and recorded instead of emitting `- [ins_…] procedure ` — a fragment that costs tokens and conveys nothing.
|
|
128
|
+
- **Eval set grew to 9 cases** (added "缩写巧合(PR)不得命中"). The shipped `signalMinRatio 0.5` still scores `precision 1.00 / recall 1.00`; the `0.20` column now shows two *genuine* weak matches (`h-weak`, and `a-pr` on 检查), which is precisely the evidence that 0.5 — not 0.2 — is the right shipped point. Measured after the fix on the real store, the same message yields `[trigger:重启, hint:提交]` instead of `[trigger:重启, hint:PR, hint:提交]`.
|
|
129
|
+
- **Tests:** `test/readiness.test.mjs` 10 → 13 checks; suite 266 → **269**.
|
|
130
|
+
|
|
3
131
|
## 0.5.4 (2026-09-12)
|
|
4
132
|
|
|
5
133
|
### Changed (indexing: model-free by design; whole-chunk retrieval terms)
|
package/README.md
CHANGED
|
@@ -16,8 +16,8 @@ The workflow panel is collapsible, automatically adapts to dsh and theme plugin
|
|
|
16
16
|

|
|
17
17
|
## Features
|
|
18
18
|
|
|
19
|
-
- **TaskBridge: cross-session development tasks** — the plugin watches each session's live todo list (`todo_write` events) and file reads (`tool/call`): progress snapshots (`steps`) and touched files sync into durable per-project task entities. An unbound session that writes a todo auto-creates a task. Associated files are kept in **recency-weighted order (written/edited first; a read never outranks a written file)** so a resumed session sees at a glance where to look. New sessions continue by `list_tasks` → `select_task` (bind / rename / unarchive); `query_memory` gains `type: 'task'` and appends a task-count hint to `type: 'all'` results. The user-side `/tasks` command shows the task stack, step progress, involved files, and the current session binding.
|
|
20
|
-
- **Task Panel (v0.4.2+): Floating task panel in dsh web** — built on the real dsh web 0.1.
|
|
19
|
+
- **TaskBridge: cross-session development tasks** — the plugin watches each session's live todo list (`todo_write` events) and file reads (`tool/call`): progress snapshots (`steps`) and touched files sync into durable per-project task entities. An unbound session that writes a todo auto-creates a task. Associated files are kept in **recency-weighted order (written/edited first; a read never outranks a written file)** so a resumed session sees at a glance where to look. New sessions continue by `list_tasks` → `select_task` (bind / rename / unarchive); `query_memory` gains `type: 'task'` and appends a task-count hint to `type: 'all'` results. The user-side `/tasks` command shows the task stack, step progress, involved files, and the current session binding. A task named by the model via `select_task(title=…)` keeps that title; one **auto-created** by the first `todo_write` is titled from its **first list entry** (≤48 chars), falling back to the first human message (the part after the last colon), then `Untitled Task`. Sessions spawned as **subagents** are excluded from auto-creation (`origin: 'subagent'` / `delegationDepth > 0`); merging delegated work back into a task is deliberately unbuilt — see §11 in Design tradeoffs. Capacity is project-size adaptive (`fileCount/20`, clamped 5–100). Storage: `.dsh-project-memory/tasks.json` + `binding.json`. Auto-sync requires a dsh build with session events + `todo_write` (verified on 0.1.2-alpha.x, re-verified against the 0.1.5-rc.1 host surface); on older hosts the task tools still work as a plain record list.
|
|
20
|
+
- **Task Panel (v0.4.2+): Floating task panel in dsh web** — built on the real dsh web 0.1.5-rc.1 client plugin contract (cordis inject + apply, registered into host `shell.overlay` slot). Draggable cards show steps/files (click to copy path); collapse to a draggable mini-bar; hide completely (summon with `/task` / `/tasks`). Render errors have error boundaries — panel crash no longer takes down the host.
|
|
21
21
|
- **Task Panel Behavior** —
|
|
22
22
|
- **Default hidden**: panel does not show on dsh web startup
|
|
23
23
|
- **Explicit summon**: type `/tasks` or `/task` (list form) to open; model calls `show_task_panel` tool to open
|
|
@@ -25,40 +25,39 @@ The workflow panel is collapsible, automatically adapts to dsh and theme plugin
|
|
|
25
25
|
- **Page refresh**: panel stays hidden (UI state `closed` not persisted)
|
|
26
26
|
- **Manual close**: click × to fully hide (no mini-bar); reopen requires explicit summon
|
|
27
27
|
- **Collapse to mini-bar**: click ↓ to keep draggable top bar; click bar to expand
|
|
28
|
+
- **Hide hints**: click ? to suppress every hover tooltip in the panel (drag handle, style/view/minimize/close, rename, step status, copy path, mini-bar, memory view); the preference is stored in localStorage and survives a refresh; the button dims while hints are off — click again to restore
|
|
28
29
|
- **Bidirectional task-list sync (host ↔ plugin tasks, v0.4.2+)** — `select_task` or `/task switch` pushes task steps to host `todo/write` so dsh's rendered task list mirrors the plugin's task entity. Config `tasklist.syncHostOnAdopt` (default on) to toggle. Empty `todo/write` means "clear": unbound session clears list without creating junk tasks; bound session clears that task's steps (task retained). Panel edits (step text/status) = write back bound task + push host list, sharing one code path with model `todo_write`. `/task` subcommands: `switch`, `archive`, `unbind`, `rename`, `todos` (invoked by panel buttons/clicks, not the model); `unbind` also clears the host task list above the input.
|
|
29
30
|
- **Panel editing & themes (v0.4.2+)** — bound cards: double-click title/step for inline edit (input auto-grows); click step status icon to cycle todo→in-progress→done. Non-bound cards read-only. **Four visual themes** (click folder icon left of title, persisted locally): Native / Glassmorphism / Brutalist / Terminal monospace — only material, geometry, typeface, density change; colors always use dsw alias tokens, follow host light/dark and theme plugins.
|
|
30
|
-
- **Document memorization** — PDF, Markdown, and plain text files are chunked and summarized
|
|
31
|
-
- **Code symbol memory** —
|
|
32
|
-
- **L1 Enhanced Regex** — zero-dep regex scanner now extracts generics, parameter/return types, overloads, interfaces, and type aliases for all supported languages, producing one-line identity signatures `fn(a: A, b: B): R — file.ts:42`.
|
|
31
|
+
- **Document memorization** — PDF, Markdown, and plain text files are chunked and summarized **without any model call**: each entry keeps a ≤300-character `summary` for injection, a bounded (≤160) deterministic, stop-word-filtered `terms` set that covers the **entire chunk** (search-only, so recall is not limited to the opening lines), and a `path:line` citation back to the source. The legacy `blindSpots` field is always empty now that indexing never calls a model; it is kept only so stores written by older versions still load.
|
|
32
|
+
- **Code symbol memory (L1 regex)** — a dependency-free scanner extracts functions, classes and methods with full signatures (generics, parameter/return types, overloads) plus interfaces and type aliases across 8 languages, producing one-line identity signatures `fn(a: A, b: B): R — file.ts:42`. It masks strings/comments, joins multi-line signatures, is indentation-aware for Python and carries class-method context — with zero LLM tokens.
|
|
33
33
|
- **Optional TypeScript semantic enhancement (L2/L3)** — when `typescript` is installed in the user project (`npm i -D typescript`), the plugin automatically activates a second layer (L2) that uses the TS Compiler API to infer return types, resolve generics, extract interfaces and type aliases, and enrich arrow functions — all asynchronously in a priority queue (P0 on `fs/observed`, P1 on `watch`, P2 on `index_repo`). Results are cached on disk keyed by file content hash (L3) for instant cold-start reuse. Zero config: just install TS (5.x or 6.x) and restart dsh. Fully optional; if TS is absent or disabled via `enableTypeScript: false`, the plugin falls back to L1 regex-only extraction.
|
|
34
34
|
- **Automatic refresh** — a background poll (`watch_repo`) detects new or changed files by content hash and re-memorizes only those.
|
|
35
35
|
- **Read-time memorization** — files are memorized the moment the model actually reads them (`fs/observed`), so the memory is a byproduct of normal work, not a separate upfront scan. Files that are never read are never indexed. The project root is detected by markers (`.git`, `package.json`, …), a README plus source directories, or the file's own directory as a last resort.
|
|
36
36
|
- **Doc ↔ code cross-linking** — when a document mentions a symbol, the match is recorded as a `reference`; querying a symbol also surfaces the documents that describe it.
|
|
37
37
|
- **BM25 memory recall** — ranked search over documents, symbols, and experience notes, with optional LLM query expansion to handle vocabulary mismatch. **CJK-optimized**: precise phrase boost (3+ char phrases ×1.5 score on title/keywords match), synonym table (e.g. 数据库连接池 ↔ 连接池 ↔ DB pool), and CJK-aware word boundaries for doc↔symbol linking.
|
|
38
|
-
- **blindSpots-aware recall** — document summaries carry a `blindSpots` field (what the summary explicitly does NOT cover). When a query hits a blind spot, `query_memory` appends a warning pointing the model to read the source file, preventing hallucination from partial summaries.
|
|
39
38
|
- **Experience notes** — problems → solutions; similar problems supersede instead of duplicating, and notes are returned only when a search matches. The note store is bounded: capacity scales with project size (clamped to 100–2000), and the oldest notes are pruned when the limit is exceeded. **Supersede tightened to bidirectional 0.7 overlap** (was 0.6); **experience `problem` field now participates in CJK phrase boost** for long-tail query recall.
|
|
40
|
-
- **v0.5 tiered insight memory (lessons / decisions / procedures)** — one `insight` entity across three scopes: `task` (private drafts in `tasks.json`), `project` (`.dsh-project-memory/insights.json`), `global` (`~/.config/dsh-project-memory/global.json`). `save_lesson` writes any scope; dedupe is bidirectional token overlap ≥ 0.7 (merge) with a 0.65–0.7 reinforce band; **promotion is a scope change, not a copy** — 2 tasks hitting the same insight promote it to project, 3+ to global. Archive is soft (`archived`), decay/capacity prune archived entries only; writes are filtered for secret/token-shaped content. LLM **reflection is off by default** and only ever writes task-level drafts (`source: reflect`) on task switch-away/archive. Panel gains a Task / Project / Global memory view with approve, promote/demote, archive/restore, delete, edit and a create form (procedures can carry an “as Skill” trigger). Old `experience.json` notes are imported into `insights.json` once, non-destructively.
|
|
41
|
-
- **Streaming TF + IDF caching** — query path caches IDF (term inverse frequency) per store version; on cache hit, single-pass streaming scores 20k entries in
|
|
42
|
-
- **Lock-free sync transactions** — all writes (index / watch / remember / forget / watch_repo) go through synchronous transactions `store.commit(fn)`; fn succeeds then atomic write; JS single-threaded event loop guarantees no interleaving; `remember`/`forget` never blocked by watch re-indexing.
|
|
39
|
+
- **v0.5 tiered insight memory (lessons / decisions / procedures)** — one `insight` entity across three scopes: `task` (private drafts in `tasks.json`), `project` (`.dsh-project-memory/insights.json`), `global` (`~/.config/dsh-project-memory/global.json`). `save_lesson` writes any scope; dedupe is bidirectional token overlap ≥ 0.7 (merge) with a 0.65–0.7 reinforce band; **promotion is a scope change, not a copy** — 2 tasks hitting the same insight promote it to project, 3+ to global. Archive is soft (`archived`), decay/capacity prune archived entries only; writes are filtered for secret/token-shaped content. LLM **reflection is off by default** and only ever writes task-level drafts (`source: reflect`) on task switch-away/archive. Panel gains a Task / Project / Global memory view with approve, promote/demote, archive/restore, delete, edit and a create form (procedures can carry an “as Skill” trigger). Old `experience.json` notes are imported into `insights.json` once, non-destructively. Every kind can carry an authored `trigger`: **only `when` can trigger**, `guard` can only narrow, and `prevents` states what breaks without the entry. `when.ops` are normalized action ids resolved from the tool call itself (`file-write` / `file-delete` / `git-commit` / `release` / `npm-publish` / `render-doc` / `run-bench` / …), `when.writes` are the files this step is about to **write**, `when.intents` are intent words from the human message **after stripping quoted/path references and filenames**. A hit injects the entry deterministically **before the action**. Legacy `keywords` / `symbols` / `actions` / `paths` / `scope` are auto-migrated in memory (actions → `ops`, concrete paths → `writes`, keywords → `intents`, extension/name globs and dead action ids dropped) — an entry left with **no** triggerable member is no longer pushed; run `npm run selfcheck:triggers` to see which ones those are.
|
|
40
|
+
- **Streaming TF + IDF caching** — query path caches IDF (term inverse frequency) per store version; on cache hit, single-pass streaming scores 20k entries (5k files) in p50 2.6 ms / p95 5.4 ms — and 4k entries (1k files) in p50 0.6 ms / p95 1.6 ms — with zero intermediate objects. Only a **dirty** write bumps the version and drops the cache — a no-op `save()` returns before touching the disk, so the 15 s watch poll can never clear the cache a query just built.
|
|
41
|
+
- **Lock-free sync transactions** — all writes (index / watch / remember / forget / watch_repo) go through synchronous transactions `store.commit(fn)`; fn succeeds then atomic write; the JS single-threaded event loop guarantees no interleaving (**in-process only** — see Consistency); `remember`/`forget` are never blocked by watch re-indexing.
|
|
43
42
|
- **Minimal dependencies** — pure JavaScript; the only runtime dependency is `pdfjs-dist` (PDF text extraction), no native builds required.
|
|
44
|
-
- **Negligible overhead** — pure in-process operation;
|
|
43
|
+
- **Negligible overhead** — pure in-process operation; a 5k-file store loads in 40 ms, and a cached query over 20k entries is p50 2.6 ms / p95 5.4 ms (4k entries: p50 0.6 ms / p95 1.6 ms); the bottleneck is PDF extraction and disk I/O, not the plugin's scoring.
|
|
45
44
|
|
|
46
45
|
## Performance
|
|
47
46
|
|
|
48
|
-
### Synthetic Benchmark (
|
|
47
|
+
### Synthetic Benchmark (Node 24.19, WSL2 on 20 vCPU, Linux file system)
|
|
49
48
|
|
|
50
49
|
| Scenario | Scale | Measured |
|
|
51
50
|
|----------|-------|----------|
|
|
52
|
-
| Full cold index | 5,000 files / 20k entries |
|
|
53
|
-
| Cold load | 5,000 files |
|
|
54
|
-
| Hot lazy re-index (single file) | 5k files |
|
|
55
|
-
| query_memory (cached) | 5k files / 20k entries |
|
|
56
|
-
| query_memory (cached) | 1k files / 4k entries |
|
|
57
|
-
| Full cold index | 10,000 files / 40k entries |
|
|
58
|
-
| Cold load | 10,000 files |
|
|
59
|
-
| Hot lazy re-index (single file) | 10k files |
|
|
51
|
+
| Full cold index | 5,000 files / 20k entries | 269 ms avg (p50 267) |
|
|
52
|
+
| Cold load | 5,000 files | 40 ms |
|
|
53
|
+
| Hot lazy re-index (single file) | 5k files | p50 2.4 ms / max 5.5 ms |
|
|
54
|
+
| query_memory (cached) | 5k files / 20k entries | p50 2.6 ms / p95 5.4 ms |
|
|
55
|
+
| query_memory (cached) | 1k files / 4k entries | p50 0.6 ms / p95 1.6 ms |
|
|
56
|
+
| Full cold index | 10,000 files / 40k entries | 551 ms avg (p50 528) |
|
|
57
|
+
| Cold load | 10,000 files | 90 ms |
|
|
58
|
+
| Hot lazy re-index (single file) | 10k files | p50 5.4 ms / max 9.2 ms |
|
|
60
59
|
|
|
61
|
-
> Synthetic benchmark: generated code (~4–5 symbols/file), Node 24
|
|
60
|
+
> Synthetic benchmark: generated code (~4–5 symbols/file), Node 24.19 on WSL2 / 20 vCPU / Linux file system, measured 2026-09-14. Reproduce with `npm run bench:synthetic -- 5000` (harness: `scripts/bench-synthetic.mjs`). Measures pure indexing overhead without LLM calls. query_memory uses the IDF cache + precomputed searchText; the first query after a write rebuilds IDF (**106 ms at 40k entries**, 57 ms at 20k, 12 ms at 4k), subsequent queries hit the cache.
|
|
62
61
|
|
|
63
62
|
### Real Project Storage
|
|
64
63
|
|
|
@@ -69,13 +68,34 @@ The workflow panel is collapsible, automatically adapts to dsh and theme plugin
|
|
|
69
68
|
|
|
70
69
|
> Real projects (Java + Vue), tested on Linux file system (Node 24). Real project entries are smaller than synthetic benchmarks due to lower symbol density and shorter declarations.
|
|
71
70
|
|
|
71
|
+
### Reproduce it on your own project
|
|
72
|
+
|
|
73
|
+
Rather than asking you to trust the numbers above, the measurement itself ships with the repository **and with the published npm package** (`scripts/` is part of the tarball). It needs **no dsh instance, no network and no model calls**, and it never touches your project's own store — results go to a temp directory and are removed when it finishes:
|
|
74
|
+
|
|
75
|
+
```bash
|
|
76
|
+
npm run bench -- /path/to/your/project
|
|
77
|
+
# or, with options:
|
|
78
|
+
node scripts/bench.mjs /path/to/your/project [--json] [--samples 100] [--no-pdf] [--keep]
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
It reports the cold index split into read+hash / extract / commit, cold load, IDF rebuild, cold and hot query latency (p50/p95/max over 100 sampled queries through the shipped scorer), single-file hot re-index, store size and bytes per entry. Example — our internal Vue project (289 files / 2,141 entries, Node 24, 20 CPU, Linux):
|
|
82
|
+
|
|
83
|
+
```
|
|
84
|
+
cold index 253 ms (read+hash 9 ms · extract 229 ms · commit 13 ms) ← 2nd, warm-cache run
|
|
85
|
+
store 1.10 MB · 538 bytes/entry · cold load 4.6 ms
|
|
86
|
+
hot query p50 0.80 ms · p95 1.35 ms (2,141 entries)
|
|
87
|
+
re-index 1 file p50 0.33 ms
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
Two caveats we would rather state than hide: `read+hash` depends on the OS page cache — on that corpus the first run spent 787 ms and the second 253 ms, so say which run you quote — and **real projects score slower than the synthetic table above** — on a 3,000-file slice of a large TypeScript repository (15,594 entries) hot queries were p50 7.5 ms, because real declaration text is longer than generated stubs. Pass `--queries your-queries.json` to run the same labeled-set method (hit@5 / hit@10 / MRR) against your own project.
|
|
91
|
+
|
|
72
92
|
## How it works
|
|
73
93
|
|
|
74
94
|
The design follows four principles:
|
|
75
95
|
|
|
76
96
|
- **Volatility** — context is ephemeral; it is lost when a session is compacted.
|
|
77
97
|
- **Persistence** — the **memory** is stored on disk and survives compaction and new sessions.
|
|
78
|
-
- **Compactness** —
|
|
98
|
+
- **Compactness** — the code layer stores one declaration line per symbol, so code-heavy projects stay near **0.5% of the source** (8.8 MB of source → 49 KB of index in the example project), and **recall** replaces re-reading the full file. The document layer is heavier by design: each chunk keeps a ≤300-char injected `summary`, a bounded `terms` set covering the whole chunk for retrieval, and a precomputed `searchText`. Measured on a docs-only corpus (179 chunks / 225 KB of Markdown): `terms` ≈ **27.5%** of source and the on-disk store ≈ **166%** of source — so on doc-heavy projects budget for roughly the docs themselves, not 0.5%.
|
|
79
99
|
- **Verifiability** — **recalls** carry a `path:line` citation where applicable, so the agent can confirm details against the source.
|
|
80
100
|
|
|
81
101
|
Building the **memory** does not require an upfront scan: files are memorized as the model reads them, so the **memory** grows to cover exactly what has been worked with. Re-reading a file that has not changed is a no-op (content hash), so the **memory** stays fresh with minimal ongoing overhead.
|
|
@@ -112,11 +132,11 @@ The tools below are **invoked by the agent**, not typed by the user. In the chat
|
|
|
112
132
|
|
|
113
133
|
| Tool | Purpose |
|
|
114
134
|
|---|---|
|
|
115
|
-
| `index_doc file_path` | Index one document (PDF/MD/txt): chunk →
|
|
116
|
-
| `index_repo root` | Index a whole project: docs get
|
|
117
|
-
| `watch_repo root` | Enable automatic refresh: a background poll detects new/changed files (mtime + content hash) and re-indexes only those. Watched roots persist across plugin restarts. |
|
|
135
|
+
| `index_doc file_path` | Index one document (PDF/MD/txt): chunk → deterministic `summary` + whole-chunk `terms` → store with `path:line`. Unchanged files are skipped. |
|
|
136
|
+
| `index_repo root` | Index a whole project: docs get deterministic summaries + whole-chunk terms, code files get a zero-token symbol table. Incremental, cleans up deleted files, cross-links docs to symbols. A root that does not exist — including a Windows-style path resolved on Linux/macOS — is rejected before anything is written. |
|
|
137
|
+
| `watch_repo root` | Enable automatic refresh: a background poll detects new/changed files (mtime + content hash) and re-indexes only those. Watched roots persist across plugin restarts; a non-existent root, the filesystem root and the shared temp directory are all refused, and roots that disappear are dropped instead of being re-created. |
|
|
118
138
|
| `memory_stats root` | Show what the store contains: totals (files / entries / experience notes), last index time, and the per-file list sorted by recency. |
|
|
119
|
-
| `query_memory query` | BM25 search over docs + symbols + experience, optionally query-expanded by the LLM. Returns ranked hits with relative scores, sources, and doc→symbol references. |
|
|
139
|
+
| `query_memory query` | BM25 search over docs + symbols + experience + insights (lessons / decisions / procedures), optionally query-expanded by the LLM. `type` selects a layer (`all` / `doc` / `symbol` / `experience` / `insight` / `task`). Returns ranked hits with relative scores, sources or insight ids, and doc→symbol references. |
|
|
120
140
|
| `list_tasks` | List task records for the project (archived marked). Call first in a new session before continuing work. |
|
|
121
141
|
| `select_task` | Bind the session to a task so its todo list and file reads sync into it. Exact `taskId`, or exact `title` (multiple matches return candidates; no match creates a new task). Pass `title` with `taskId` to rename. Auto-unarchives. |
|
|
122
142
|
| `archive_task` | Archive a task (hide from default views, exclude from capacity, stop syncing). `select_task` restores it. |
|
|
@@ -126,7 +146,7 @@ The tools below are **invoked by the agent**, not typed by the user. In the chat
|
|
|
126
146
|
| `/insight` (typed by the user, not the model) | v0.5 memory view actions (panel buttons): `list [task|project|global]`, `confirm` / `promote` / `demote` / `archive` / `restore` / `delete` `<scope> <id>`, `save <scope> <json>`, `edit <scope> <id> <json>`. |
|
|
127
147
|
| `remember problem solution` | Save an experience note. Similar problems supersede instead of duplicating. |
|
|
128
148
|
| `forget id_or_query` | Delete stale experience notes. |
|
|
129
|
-
| `save_lesson` (agent tool) | Save a lesson/decision/procedure at task/project/global scope (single insight entity). Near-duplicates merge (≥ 0.7 overlap) or reinforce (0.65–0.7); 2+ tasks hitting the same insight auto-promote task → project, 3+ → global. Params: `title`, `kind`, `scope`, `pattern`/`fix` or `choice`/`reason` or `steps`/`trigger`, `task_id`, `files`, `symbols`, `confidence`, `root`. |
|
|
149
|
+
| `save_lesson` (agent tool) | Save a lesson/decision/procedure at task/project/global scope (single insight entity). Near-duplicates merge (≥ 0.7 overlap) or reinforce (0.65–0.7); 2+ tasks hitting the same insight auto-promote task → project, 3+ → global. Params: `title`, `kind`, `scope`, `pattern`/`fix` or `choice`/`reason` or `steps`, `trigger` (`when` = `ops`/`writes`/`intents`, the only trigger surface; `guard` = `paths`/`not_paths`/`hosts`/`tags`, narrowing only; `prevents` = what breaks without it; legacy `keywords`/`symbols`/`actions`/`paths`/`scope` still accepted and auto-migrated), `task_id`, `files`, `symbols`, `confidence`, `root`. |
|
|
130
150
|
|
|
131
151
|
## Design
|
|
132
152
|
|
|
@@ -146,7 +166,7 @@ Stores created before v0.2.0 (single `entries.json` / `index.json`) migrate auto
|
|
|
146
166
|
|
|
147
167
|
- **Incremental** — content hash per file; only changed files are re-extracted.
|
|
148
168
|
- **Cross-linking** — after indexing, doc summaries are matched against symbol names; matches are attached to the doc entry as `references` and surfaced by `query_memory`.
|
|
149
|
-
- **Query expansion** — when `llmQueryExpansion` is on, `query_memory` asks `ctx.llm` to rewrite the query into several variants (synonyms, EN/CN, identifier guesses) and merges BM25 scores across variants; when off, queries never touch the LLM.
|
|
169
|
+
- **Query expansion** — when `llmQueryExpansion` is on, `query_memory` asks `ctx.llm` to rewrite the query into several variants (synonyms, EN/CN, identifier guesses) and merges BM25 scores across variants; when off, queries never touch the LLM. Indexing itself is model-free: keywords are rule-derived (title-weighted top terms), and doc↔symbol links surface English symbol names from Chinese hits.
|
|
150
170
|
- **Consistency** — the fact layer follows the codebase (hash re-extract / remove-on-delete); the experience layer is retrieval-only with supersede and `forget`. Store writes are serialized per memory directory; the lock is in-process, so avoid running multiple dsh instances against the same project store concurrently.
|
|
151
171
|
|
|
152
172
|
## Architecture (Task Panel)
|
|
@@ -173,11 +193,11 @@ These are deliberate scope choices.
|
|
|
173
193
|
|
|
174
194
|
### 2. Watch: compute outside, commit inside
|
|
175
195
|
|
|
176
|
-
**We do:** Heavy work (mtime/hash/scan/
|
|
196
|
+
**We do:** Heavy work (mtime/hash/scan/parse/PDF extraction) runs outside the transaction; a single `commit` applies all changes atomically. On failure, the snapshot rolls back so the next poll retries automatically.
|
|
177
197
|
|
|
178
|
-
**We don't:** Hold a lock during
|
|
198
|
+
**We don't:** Hold a lock during parsing, or use `fs.watch` events.
|
|
179
199
|
|
|
180
|
-
**Why:**
|
|
200
|
+
**Why:** PDF extraction and large-file parsing take time — holding a lock would block `remember`/`forget`/`query_memory`. Polling with mtime+content-hash is platform-agnostic (works on network drives, Docker volumes, WSL) and avoids the "double fire / missed events" nightmare of `fs.watch`.
|
|
181
201
|
|
|
182
202
|
### 3. Corrupt files are quarantined, not auto-repaired
|
|
183
203
|
|
|
@@ -193,23 +213,25 @@ These are deliberate scope choices.
|
|
|
193
213
|
|
|
194
214
|
**We don't:** Vector embeddings, dense retrieval, rerankers, or hybrid search.
|
|
195
215
|
|
|
196
|
-
**Why:** Vectors require an embedding model (local = heavy, remote = latency + cost + privacy), a vector index (HNSW/IVF = memory + build time), and reranking (another LLM call). For
|
|
216
|
+
**Why:** Vectors require an embedding model (local = heavy, remote = latency + cost + privacy), a vector index (HNSW/IVF = memory + build time), and reranking (another LLM call). For the queries this plugin targets, lexical BM25 is already sufficient and measurable: on our benchmark suite (29 queries over a real Vue project) file-level hit@5 is **96.6%**, and 28 of the 29 are exact symbol lookups that lexical search answers essentially always. Whole-chunk `terms` took document-term coverage from **27.3% to 100%** while queries that already worked kept their ranking (MRR **0.958** vs **0.955**). Those figures come from an internal Vue project with a hand-labeled 29-query set, so they are not reproducible outside it — but the **method** now ships as `scripts/bench.mjs --queries <your-set.json>`, so you can run the identical measurement on your own project. The marginal gain from semantic search doesn't justify the 10x complexity/cost increase.
|
|
197
217
|
|
|
198
|
-
### 5.
|
|
218
|
+
### 5. Indexing is deterministic and model-free
|
|
199
219
|
|
|
200
|
-
**We do:**
|
|
220
|
+
**We do:** Derive keywords with a rule (title-weighted top terms) and build a whole-chunk `terms` set — both deterministic and reproducible. Doc↔symbol links surface English symbol names from Chinese queries, and CJK tokenization keeps cross-language hits working. With `llmQueryExpansion: false`, queries never touch the LLM.
|
|
201
221
|
|
|
202
|
-
**We don't:**
|
|
222
|
+
**We don't:** Call a model at index time to translate or paraphrase a document, and we don't translate queries at search time.
|
|
203
223
|
|
|
204
|
-
**Why:**
|
|
224
|
+
**Why:** An index-time model call makes indexing slower, non-deterministic and unverifiable — the same document can index differently on two runs. Query-time translation adds latency and a hard failure mode (a bad translation means zero recall). Rules plus symbol linking cover the common cases, work offline, and keep indexing at zero model calls.
|
|
205
225
|
|
|
206
|
-
### 6.
|
|
226
|
+
### 6. Model-facing memory: the agent writes, and no human has to be in the loop
|
|
207
227
|
|
|
208
|
-
**We do:**
|
|
228
|
+
**We do:** Treat the agent as a first-class writer. `remember` / `save_lesson` write **any scope at any time** (`task` / `project` / `global`) with no human step, and promotion is deterministic and runs inside the ordinary write path: cross-task token-overlap dedupe accumulates `sourceTaskIds`, then `promoteAllTasksToProject` / `promoteProjectToGlobal` move an entry up once its corroboration counts are met (≥2 tasks for project, ≥ `globalPromoteTasks` — 3 by default — for global). Nothing waits on the task panel: a user who never opens the UI still gets a memory that fills, dedupes and graduates.
|
|
209
229
|
|
|
210
|
-
**We
|
|
230
|
+
**We do (labeling):** Keep inferred content distinguishable from recorded content. The v0.5 `reflection` path (opt-in, **off by default**) is the only writer that infers rather than records: it writes task-scoped drafts stamped `draft: true` / `source: 'reflect'`, and `recall` plus silent injection skip `draft` entries while they remain drafts.
|
|
211
231
|
|
|
212
|
-
**
|
|
232
|
+
**We don't:** Require human approval for memory to become useful, or make the UI a step in the write path. `draft` is a **provenance label plus a corroboration threshold**, not an approval queue.
|
|
233
|
+
|
|
234
|
+
**Why:** The agent is the consumer and it is usually headless — memory that only graduates when a human clicks a card is memory that never graduates. Labeling keeps the useful half of the caution (inferred ≠ recorded, and unreviewed single-task inference stays out of the prompt) without taxing the normal path. A draft graduates on corroboration: a second task matching it through the model's own writes, or the model writing the same knowledge at project scope, which links the existing entry instead of duplicating it.
|
|
213
235
|
|
|
214
236
|
### 7. Full entries returned directly
|
|
215
237
|
|
|
@@ -217,7 +239,7 @@ These are deliberate scope choices.
|
|
|
217
239
|
|
|
218
240
|
**We don't:** Return a minimal index first, then require a second tool call for details.
|
|
219
241
|
|
|
220
|
-
**Why:** Returning full entries preserves **verifiability** — the agent sees the exact source line for every claim. It also avoids a round-trip per useful hit. Our entries are already compact (~300
|
|
242
|
+
**Why:** Returning full entries preserves **verifiability** — the agent sees the exact source line for every claim. It also avoids a round-trip per useful hit. Our entries are already compact (~300-char summary + citation, plus a search-only `terms` field that never enters the prompt); the token cost is lower than a second tool call + context switch.
|
|
221
243
|
|
|
222
244
|
### 8. Symbol extraction focused on what developers search for
|
|
223
245
|
|
|
@@ -243,6 +265,14 @@ These are deliberate scope choices.
|
|
|
243
265
|
|
|
244
266
|
**Why:** Mandatory TS would break installs for non-TS projects. Blocking enhancement would stall `index_repo` on large codebases. Full-program checking is 10x slower and memory-heavy. Our design: enhance what's read, cache it, never block the hot path.
|
|
245
267
|
|
|
268
|
+
### 11. Subagent sessions are out of scope for now
|
|
269
|
+
|
|
270
|
+
**We do:** Exclude sessions spawned as subagents (`origin: 'subagent'` / `delegationDepth > 0`) from auto-creating or binding a task. Their `todo_write` events do not create tasks, and they inherit no task binding.
|
|
271
|
+
|
|
272
|
+
**We don't:** Merge a delegated run's steps and files back into the task that spawned it. That is **not designed yet**: there is no parent-link model for delegated work, and the naive version mints one project task per subagent.
|
|
273
|
+
|
|
274
|
+
**Why:** Every subagent that writes a todo would otherwise create its own task entity, so one fan-out run would flood the task list with ephemeral entries nobody resumes. Excluding them keeps the task list equal to the work the user actually owns. The cost is that a delegation's progress is invisible in the task record; merging it properly (child steps folded into the parent, or a separate delegated-work view) is future work.
|
|
275
|
+
|
|
246
276
|
## Configuration
|
|
247
277
|
|
|
248
278
|
| Key | Default | Meaning |
|
|
@@ -265,7 +295,26 @@ These are deliberate scope choices.
|
|
|
265
295
|
| `enableTypeScript` | true | set `false` to disable L2 TS enhancement entirely (L1 regex only) |
|
|
266
296
|
| `insight.*` | dedupOverlap `0.7` · reinforceBand `0.65` · maxProject `100` · maxGlobalProcedures `200` · promoteConfidence `0.7` · globalPromoteTasks `3` · decayDays `90` · `globalFile` (auto) | v0.5 insight dedupe / reinforce / promotion / capacity / archive settings |
|
|
267
297
|
| `reflection.enabled` | false | v0.5 LLM reflection, **draft-only at task level** (fires on task switch-away / archive). `cooldownMs` `1800000`, `maxLessonsPerReflect` `3`, `maxDecisionsPerReflect` `2` |
|
|
268
|
-
| `autoContext.enabled` | true |
|
|
298
|
+
| `autoContext.enabled` | true | silent injection wrapper (resident task card + gated items). Inert (full passthrough) until the host exposes a resolvable session cwd; `maxTokens` `400`, `editedMax` `3` (how many recently-written "editing now" files the resident task card shows), `signalMinRatio` `0.5` (a hint must reach half of its layer's top score), `skipEchoSelfTodo` `true` (don't echo the task card back when the model itself maintains the task list with no newer human message; relevant insights still inject), `budgetLog` `off` (budget-drop audit on stderr: `off` silent / `once` at most one line per session / `all` one line per changed dropped set), `reinjectItemsAfter` `0` (cooldown, in pre-steps, before the same insight may be injected again) |
|
|
299
|
+
| `autoContext.gateCooldownSteps` | 2 | **admission knobs.** Minimum number of pre-steps between two *item* injections (the resident task card is exempt — it is a state snapshot and should update when it changes). This is the main "don't inject often" dial |
|
|
300
|
+
| `autoContext.maxItemsPerSession` | 12 | hard per-session cap on injected items; the budget is a ceiling, not a target — once exhausted the item channel stays silent |
|
|
301
|
+
| `autoContext.maxItemCharsPerSession` | 4000 | same, in characters |
|
|
302
|
+
| `autoContext.hintMinCoverage` | 0.3 | **absolute** floor for the statistical (hint) channel: IDF-weighted share of the query's information mass the entry covers. A ratio-only threshold cannot tell signal from "best of a bad lot" (`relative:1.00` on an unrelated entry) |
|
|
303
|
+
| `autoContext.hintMinMatched` | 2 | a hint must share at least this many terms with the query — one generic word ("plugin") is not evidence |
|
|
304
|
+
| `autoContext.hintMinSupport` | 0.15 | channel-level silence: if less than this share of the query's terms exist anywhere in the corpus, the hint channel says nothing this round — a long sentence that happens to share one word otherwise reports `cov:1.00` |
|
|
305
|
+
| `autoContext.legacyScope` | `filter` | how to treat a legacy `trigger.scope`: `filter` keeps the old semantics, `ignore` drops it. `npm run selfcheck:triggers` reports entries whose scope values cannot intersect the project tag space |
|
|
306
|
+
| `autoContext.auditLog` | true | append one JSONL line per **actual** injection to `<root>/.dsh-project-memory/injection-audit.jsonl` (what was injected, why it matched, what was dropped, session budget snapshot); rotates to `.1` past `auditMaxBytes` (`262144`). Silent on any I/O error — never affects the host request |
|
|
307
|
+
|
|
308
|
+
### Injection admission (why it stays quiet)
|
|
309
|
+
|
|
310
|
+
Automatic injection used to be a *retrieval* problem ("which entry is most related to this text?"), which is total — a ranking always returns something, so noise was structural. It is now an **admission** problem ("is this step about to cross a boundary I have been burned by?"), with silence as the default:
|
|
311
|
+
|
|
312
|
+
- **Only `when` triggers**, and it is a low-dimensional typed signal: normalized `ops`, the files this step is about to **write**, and intent words from the human message *after* stripping quoted/path references. Corpus text — raw tool arguments, file contents, filenames — can never trigger anything.
|
|
313
|
+
- **`guard` only narrows.** Extension/name globs (`*.pptx`, `README*`) are ignored outright: they can only lie, never narrow.
|
|
314
|
+
- **Ratio *plus* an absolute floor.** The hint channel needs the relative score *and* an IDF-weighted coverage floor *and* at least two shared terms — `relative:1.00` also happens on entries that share nothing with the step.
|
|
315
|
+
- **Frequency is bounded.** At most one item injection every `gateCooldownSteps`, capped per session by count and characters. The resident task card is exempt (it is a snapshot that should update); the budget is a ceiling, not a target.
|
|
316
|
+
- **Prefix-cache discipline.** Injections are appended as a user message at the tail of the history, so the cached prefix is never rewritten. What they add is resident *cache-read* tokens, not cache misses; nothing is ever edited in place.
|
|
317
|
+
- **It is auditable.** Every real injection appends one line to `injection-audit.jsonl` (reason, dropped candidates, session budget), and `npm run eval:injection` scores 8 labelled scenarios — currently precision 1.00 / recall 1.00 with a clean control group.
|
|
269
318
|
|
|
270
319
|
### Toggling features
|
|
271
320
|
|
|
@@ -282,6 +331,8 @@ Settings live in the plugin's config object. To change them, add an override ent
|
|
|
282
331
|
watch: true # on: background refresh for watched roots (default)
|
|
283
332
|
watchInterval: 15 # poll interval in seconds
|
|
284
333
|
enableTypeScript: true # on: L2 TS enhancement when TS is installed (default)
|
|
334
|
+
# budgetLog: once # debugging: log budget drops to stderr (default off = silent)
|
|
335
|
+
# reinjectItemsAfter: 20 # debugging: allow the same insight again after N steps (default 0 = once per session)
|
|
285
336
|
# tsPath: /custom/path/to/typescript # optional: force specific TS install
|
|
286
337
|
```
|
|
287
338
|
|
|
@@ -301,9 +352,14 @@ These commands are for **maintaining the plugin code** — regular users do not
|
|
|
301
352
|
|
|
302
353
|
```bash
|
|
303
354
|
npm install
|
|
304
|
-
npm test
|
|
355
|
+
npm test # 313 tests (184 core + 16 TaskBridge + 11 insight-store + 9 insight-actions + 8 doc-index + 7 auto-inject + 9 host-contract + 5 reflection + 4 llm-route + 2 client-hints + 8 recall + 14 readiness + 7 insight-derive + 6 readiness-eval + 6 ops + 6 injection-audit + 5 injection-budget + 6 injection-scenarios)
|
|
356
|
+
npm run eval:injection # scenario P/R: 14/14 hits, 0 false positives, control group clean
|
|
357
|
+
npm run selfcheck:triggers # which entries can still push, which declarations are dead
|
|
358
|
+
npm run bench -- /path/to/project # index/query performance on any project — no dsh needed
|
|
305
359
|
```
|
|
306
360
|
|
|
361
|
+
Release notes live in [`CHANGELOG.md`](CHANGELOG.md) and on [GitHub Releases](https://github.com/00080000/dsh-project-memory/releases).
|
|
362
|
+
|
|
307
363
|
## License
|
|
308
364
|
|
|
309
365
|
MIT
|