@yolk_vat-y/dsh-project-memory 0.5.5 → 0.5.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,74 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.5.7 (2026-09-20) — bug-fix release
4
+
5
+ 发布前审计在 313 项全绿下发现并修复以下缺陷,新增 18 项回归测试(共 331)。
6
+
7
+ ### 数据完整性
8
+
9
+ - 旧库迁移:`index.json` 损坏时会连带删除完好的 `entries.json`(静默清空整个 store)
10
+ - `store.load()` 不幂等,重复调用丢掉未落盘的变更
11
+ - 符号后到时 doc↔symbol 链接不落盘;畸形 shard/insight 条目会让所有读取抛错
12
+
13
+ ### 召回与注入
14
+
15
+ - 提示通道覆盖率语料与 BM25 排序不一致:只匹配 `fix` 的查询会让整条提示通道沉默
16
+ - `recallItems` 改为每层各自 top-k(文档不再挤掉符号层);任务级 draft 不再泄漏
17
+ - 未加引号的 CJK 文件名不再触发 `when.intents`
18
+ - `when.writes` 纳入人类消息里的路径(动手前可命中;读取路径不算)
19
+ - 会话条目字符额度不再截断常驻任务卡;`entryOn:false` 与显式 `0` 生效;`fitBody` 边界
20
+
21
+ ### insight
22
+
23
+ - `applyDecay` 判据不可达,`decayDays` 完全失效
24
+ - 提升/降级同样收口 `maxProject` / `maxGlobalProcedures`
25
+ - 跨层移动保留 `hitCount`/`createdAt`/`triggerDerived`;global 也回填派生 trigger
26
+
27
+ ### 索引
28
+
29
+ - watch 不再每轮重读重哈希所有未变文件;`index_doc` 补 `terms` 回填
30
+ - `readTextFile` 先 stat 再读;chunker `sourceLine` 不再漂移
31
+ - `export`/多行 interface 与 type 别名可被 L1 扫到;扩展名大小写不敏感
32
+ - 损坏 PDF 销毁 loading task;TS 增强器 type 别名与关闭开关
33
+ - scoped 依赖 tag 修正;畸形依赖不再清空全部 tags
34
+
35
+ ### 工具与命令
36
+
37
+ - `query_memory(type:'task')` 读错步骤字段;空查询被拒绝
38
+ - 任务文件超限淘汰最冷文件;任务 id 补随机后缀
39
+ - `/insight edit` 合并 trigger 而非重建;`/task rename|todos` 保留内部空白
40
+ - 写类工具校验 root;`invocationContext` 补 `agent.session` 降级
41
+
42
+ ### 重构(已在 main 上,无行为变更)
43
+
44
+ - 索引逻辑收敛为 `src/index-pipeline.js`;auto-inject 会话状态收敛为 `InjectionSessions`
45
+ - 命令处理器共用 `invocationContext` / `fencedJson`
46
+
47
+ 未修的低优先级发现见 `RELEASE-NOTES-0.5.7.md` 的 backlog。
48
+
49
+ ## 0.5.6 (2026-09-16) — injection admission (lessons/decisions/procedures stop arriving by coincidence)
50
+
51
+ ### Changed (only what you are about to *do* can trigger an injection)
52
+
53
+ - **Trigger matching was a substring test over a haystack of everything the step had seen** (`humanText + tool arguments`, i.e. file contents included). Measured on a real session: 10 injections, **0 of them useful**, 5921 characters appended permanently — including a "recover deleted files" procedure pulled in by the literal string `dcterms` (it contains `rm`) and an "arXiv fetching" note pulled in by the `.pptx` file extension. A labelled 8-scenario evaluation set (now `npm run eval:injection`) scored **precision 0.48 / recall 0.72**.
54
+ - **A trigger now has a shape: `when` (ops / writes / intents) triggers, `guard` only narrows, `prevents` states what breaks without the entry.** `ops` are normalized action ids resolved from the tool call itself (`src/ops.js`: `file-write` / `file-delete` / `git-commit` / `release` / `npm-publish` / `render-doc` / `run-bench` / …); `writes` are the files this step is about to write; `intents` are human-message words matched only **after** stripping quoted spans, paths and filenames, and only above a minimum length/word-boundary rule (so `ppt` no longer matches `pptx`).
55
+ - **Corpus text can no longer trigger anything.** File contents are out of the haystack entirely, extension/name globs (`*.pptx`, `README*`) are ignored outright, and machine identifiers (`web_fetch`, `LAYOUT_16x9`) are filtered out of intent words.
56
+ - **Legacy triggers migrate in memory, idempotently and without rewriting your data**: `actions` → `when.ops` (dead ids mapped where possible, otherwise recorded), concrete `paths` → `when.writes`, `keywords` → `when.intents`, extension/name globs dropped, and weak actions (`git-add` / `git-commit`) dropped when the entry already has precise paths. `npm run selfcheck:triggers` reports what changed — against the live stores: **20 pushable / 4 pull-only, 8 dead action ids, 10 dropped globs, 3 dropped weak actions**.
57
+ - **Result on the evaluation set: precision 1.00, recall 1.00, and the control scenario (rename a `.pptx` timestamp, with the filename in quotes) injects nothing at all.**
58
+
59
+ ### Changed (hint channel: relative *and* absolute, plus a frequency budget)
60
+
61
+ - **The statistical channel required only "half of this layer's top score"**, which a ranking satisfies even when the top score is itself noise — measured `relative:1.00` on entries sharing nothing with the step. It now also needs an **IDF-weighted coverage floor** (`hintMinCoverage`, default `0.3`) **and** at least two shared terms (`hintMinMatched`, default `2`). Query text is the human message *plus this step's write targets* — raw tool arguments are no longer a query.
62
+ - **Injections are now events, not a heartbeat**: `gateCooldownSteps` (default `2`) bounds how often the item channel may speak, and `maxItemsPerSession` / `maxItemCharsPerSession` cap the session. The resident task card is exempt (it is a state snapshot), and the budget is a ceiling rather than a target — when nothing clears the gates, nothing is injected.
63
+ - **Cache discipline is explicit**: injections are appended as a user message at the tail of history, so the cached prefix is never rewritten; the cost they add is resident cache-read tokens, not cache misses. Nothing is edited in place.
64
+
65
+ ### Added (observability: this change is measured, not asserted)
66
+
67
+ - **`injection-audit.jsonl`** — one JSONL line per *actual* injection under `<root>/.dsh-project-memory/`, recording what went in, why it matched (`op:` / `write:` / `intent:` / hint coverage), what lost the budget (`budget` / `quota` / `cooldown` / `coverage:` / `thin:`), and the session budget snapshot. Rotation at `auditMaxBytes`; every failure is swallowed so the host request is never affected. Config: `autoContext.auditLog`.
68
+ - **`npm run eval:injection`** — 8 labelled scenarios (publish, trigger-debugging, stale source, benchmark, deleted doc, interview prep, and the clean control) over a 24-entry synthetic pool whose triggers are copied verbatim from the real ones. Asserts the ratchet, precision ≥ 0.90 and a clean control group. `--store <insights.json>` replays a real store; `--selfcheck` prints the trigger audit.
69
+ - **New suites**: `test/ops.test.mjs` (the action plane, including a regression for reading argument *values* rather than key names), `test/injection-budget.test.mjs` (cooldown/caps actually silence, 6 steps with 6 fresh entries → 3 injections), `test/injection-audit.test.mjs` (JSONL shape, rotation, silent failure, end-to-end wiring through `agent/pre-step`). Suite is now **313 tests**.
70
+ - **`trigger.when` / `trigger.guard` / `trigger.prevents`** are declared in the `save_lesson` schema, so new entries can be authored in the new shape.
71
+
3
72
  ## 0.5.5 (2026-09-15)
4
73
 
5
74
  ### Fixed (the README promised a benchmark the npm package did not contain)
@@ -89,12 +158,12 @@
89
158
 
90
159
  | signalMinRatio | precision | recall |
91
160
  |---|---|---|
92
- | 0.20 | 0.86 | 1.00 |
161
+ | 0.20 | 0.78 | 1.00 |
93
162
  | 0.35 | 1.00 | 1.00 |
94
163
  | 0.50 (shipped) | 1.00 | 1.00 |
95
164
  | 0.70 | 1.00 | 1.00 |
96
165
 
97
- The shipped default sits in the safe zone with margin, and 0.20 visibly admits a weak match — so the number is justified by data rather than chosen by feel. It lives under `test/` (CI-enforced, shipped) rather than the git-ignored `bench/`, deliberately: a threshold that only exists on one machine is not a threshold.
166
+ The shipped default sits in the safe zone with margin, and 0.20 visibly admits weak matches — so the number is justified by data rather than chosen by feel. The table is the **9-case** set measured after the hint-precision fix below (the `0.20` figure was 0.86 on the earlier 8-case set; adding "缩写巧合(PR)不得命中" turns one more row into a genuine weak match, `7/9 = 0.78`). It lives under `test/` (CI-enforced, shipped) rather than the git-ignored `bench/`, deliberately: a threshold that only exists on one machine is not a threshold.
98
167
  - **Tests:** new `test/insight-derive.test.mjs` (6 checks) and `test/readiness-eval.test.mjs` (4 checks, including the sweep table); suite 256 → **266**.
99
168
 
100
169
  ### Fixed (readiness hint precision: query source, acronym noise, stub truncation)
package/README.md CHANGED
@@ -36,7 +36,7 @@ The workflow panel is collapsible, automatically adapts to dsh and theme plugin
36
36
  - **Doc ↔ code cross-linking** — when a document mentions a symbol, the match is recorded as a `reference`; querying a symbol also surfaces the documents that describe it.
37
37
  - **BM25 memory recall** — ranked search over documents, symbols, and experience notes, with optional LLM query expansion to handle vocabulary mismatch. **CJK-optimized**: precise phrase boost (3+ char phrases ×1.5 score on title/keywords match), synonym table (e.g. 数据库连接池 ↔ 连接池 ↔ DB pool), and CJK-aware word boundaries for doc↔symbol linking.
38
38
  - **Experience notes** — problems → solutions; similar problems supersede instead of duplicating, and notes are returned only when a search matches. The note store is bounded: capacity scales with project size (clamped to 100–2000), and the oldest notes are pruned when the limit is exceeded. **Supersede tightened to bidirectional 0.7 overlap** (was 0.6); **experience `problem` field now participates in CJK phrase boost** for long-tail query recall.
39
- - **v0.5 tiered insight memory (lessons / decisions / procedures)** — one `insight` entity across three scopes: `task` (private drafts in `tasks.json`), `project` (`.dsh-project-memory/insights.json`), `global` (`~/.config/dsh-project-memory/global.json`). `save_lesson` writes any scope; dedupe is bidirectional token overlap ≥ 0.7 (merge) with a 0.65–0.7 reinforce band; **promotion is a scope change, not a copy** — 2 tasks hitting the same insight promote it to project, 3+ to global. Archive is soft (`archived`), decay/capacity prune archived entries only; writes are filtered for secret/token-shaped content. LLM **reflection is off by default** and only ever writes task-level drafts (`source: reflect`) on task switch-away/archive. Panel gains a Task / Project / Global memory view with approve, promote/demote, archive/restore, delete, edit and a create form (procedures can carry an “as Skill” trigger). Old `experience.json` notes are imported into `insights.json` once, non-destructively. Every kind can carry an authored `trigger` (`keywords` / `symbols` / `actions` / `paths`): a hit injects the entry **before the action**, deterministically — procedure-only in v0.5, all kinds since the readiness layer.
39
+ - **v0.5 tiered insight memory (lessons / decisions / procedures)** — one `insight` entity across three scopes: `task` (private drafts in `tasks.json`), `project` (`.dsh-project-memory/insights.json`), `global` (`~/.config/dsh-project-memory/global.json`). `save_lesson` writes any scope; dedupe is bidirectional token overlap ≥ 0.7 (merge) with a 0.65–0.7 reinforce band; **promotion is a scope change, not a copy** — 2 tasks hitting the same insight promote it to project, 3+ to global. Archive is soft (`archived`), decay/capacity prune archived entries only; writes are filtered for secret/token-shaped content. LLM **reflection is off by default** and only ever writes task-level drafts (`source: reflect`) on task switch-away/archive. Panel gains a Task / Project / Global memory view with approve, promote/demote, archive/restore, delete, edit and a create form (procedures can carry an “as Skill” trigger). Old `experience.json` notes are imported into `insights.json` once, non-destructively. Every kind can carry an authored `trigger`: **only `when` can trigger**, `guard` can only narrow, and `prevents` states what breaks without the entry. `when.ops` are normalized action ids resolved from the tool call itself (`file-write` / `file-delete` / `git-commit` / `release` / `npm-publish` / `render-doc` / `run-bench` / …), `when.writes` are the files this step is about to **write**, `when.intents` are intent words from the human message **after stripping quoted/path references and filenames**. A hit injects the entry deterministically **before the action**. Legacy `keywords` / `symbols` / `actions` / `paths` / `scope` are auto-migrated in memory (actions → `ops`, concrete paths → `writes`, keywords → `intents`, extension/name globs and dead action ids dropped) — an entry left with **no** triggerable member is no longer pushed; run `npm run selfcheck:triggers` to see which ones those are.
40
40
  - **Streaming TF + IDF caching** — query path caches IDF (term inverse frequency) per store version; on cache hit, single-pass streaming scores 20k entries (5k files) in p50 2.6 ms / p95 5.4 ms — and 4k entries (1k files) in p50 0.6 ms / p95 1.6 ms — with zero intermediate objects. Only a **dirty** write bumps the version and drops the cache — a no-op `save()` returns before touching the disk, so the 15 s watch poll can never clear the cache a query just built.
41
41
  - **Lock-free sync transactions** — all writes (index / watch / remember / forget / watch_repo) go through synchronous transactions `store.commit(fn)`; fn succeeds then atomic write; the JS single-threaded event loop guarantees no interleaving (**in-process only** — see Consistency); `remember`/`forget` are never blocked by watch re-indexing.
42
42
  - **Minimal dependencies** — pure JavaScript; the only runtime dependency is `pdfjs-dist` (PDF text extraction), no native builds required.
@@ -146,7 +146,7 @@ The tools below are **invoked by the agent**, not typed by the user. In the chat
146
146
  | `/insight` (typed by the user, not the model) | v0.5 memory view actions (panel buttons): `list [task|project|global]`, `confirm` / `promote` / `demote` / `archive` / `restore` / `delete` `<scope> <id>`, `save <scope> <json>`, `edit <scope> <id> <json>`. |
147
147
  | `remember problem solution` | Save an experience note. Similar problems supersede instead of duplicating. |
148
148
  | `forget id_or_query` | Delete stale experience notes. |
149
- | `save_lesson` (agent tool) | Save a lesson/decision/procedure at task/project/global scope (single insight entity). Near-duplicates merge (≥ 0.7 overlap) or reinforce (0.65–0.7); 2+ tasks hitting the same insight auto-promote task → project, 3+ → global. Params: `title`, `kind`, `scope`, `pattern`/`fix` or `choice`/`reason` or `steps`, `trigger` (`keywords`/`symbols`/`actions`/`paths`/`scope` — any kind; a hit injects the entry before the action), `task_id`, `files`, `symbols`, `confidence`, `root`. |
149
+ | `save_lesson` (agent tool) | Save a lesson/decision/procedure at task/project/global scope (single insight entity). Near-duplicates merge (≥ 0.7 overlap) or reinforce (0.65–0.7); 2+ tasks hitting the same insight auto-promote task → project, 3+ → global. Params: `title`, `kind`, `scope`, `pattern`/`fix` or `choice`/`reason` or `steps`, `trigger` (`when` = `ops`/`writes`/`intents`, the only trigger surface; `guard` = `paths`/`not_paths`/`hosts`/`tags`, narrowing only; `prevents` = what breaks without it; legacy `keywords`/`symbols`/`actions`/`paths`/`scope` still accepted and auto-migrated), `task_id`, `files`, `symbols`, `confidence`, `root`. |
150
150
 
151
151
  ## Design
152
152
 
@@ -295,7 +295,26 @@ These are deliberate scope choices.
295
295
  | `enableTypeScript` | true | set `false` to disable L2 TS enhancement entirely (L1 regex only) |
296
296
  | `insight.*` | dedupOverlap `0.7` · reinforceBand `0.65` · maxProject `100` · maxGlobalProcedures `200` · promoteConfidence `0.7` · globalPromoteTasks `3` · decayDays `90` · `globalFile` (auto) | v0.5 insight dedupe / reinforce / promotion / capacity / archive settings |
297
297
  | `reflection.enabled` | false | v0.5 LLM reflection, **draft-only at task level** (fires on task switch-away / archive). `cooldownMs` `1800000`, `maxLessonsPerReflect` `3`, `maxDecisionsPerReflect` `2` |
298
- | `autoContext.enabled` | true | v0.5 silent injection wrapper (entry block + relevance). Inert (full passthrough) until the host exposes a resolvable session cwd; `maxTokens` `400`, `editedMax` `3` (how many recently-written "editing now" files the resident task card shows), `signalMinRatio` `0.5` (a hint must reach half of its layer's top score), `skipEchoSelfTodo` `true` (don't echo the task card back when the model itself maintains the task list with no newer human message; relevant insights still inject), `budgetLog` `off` (budget-drop audit on stderr: `off` silent / `once` at most one line per session / `all` one line per changed dropped set. Injection is priority-scheduled, so dropping low-priority entries when the budget runs out is **normal degradation, not a failure** — hence the default keeps the user's terminal clean), `reinjectItemsAfter` `0` (cooldown, in pre-steps, before the same insight may be injected again; `0` = an unchanged item is never re-injected in the same session, because the injected message stays in the session history) |
298
+ | `autoContext.enabled` | true | silent injection wrapper (resident task card + gated items). Inert (full passthrough) until the host exposes a resolvable session cwd; `maxTokens` `400`, `editedMax` `3` (how many recently-written "editing now" files the resident task card shows), `signalMinRatio` `0.5` (a hint must reach half of its layer's top score), `skipEchoSelfTodo` `true` (don't echo the task card back when the model itself maintains the task list with no newer human message; relevant insights still inject), `budgetLog` `off` (budget-drop audit on stderr: `off` silent / `once` at most one line per session / `all` one line per changed dropped set), `reinjectItemsAfter` `0` (cooldown, in pre-steps, before the same insight may be injected again) |
299
+ | `autoContext.gateCooldownSteps` | 2 | **admission knobs.** Minimum number of pre-steps between two *item* injections (the resident task card is exempt — it is a state snapshot and should update when it changes). This is the main "don't inject often" dial |
300
+ | `autoContext.maxItemsPerSession` | 12 | hard per-session cap on injected items; the budget is a ceiling, not a target — once exhausted the item channel stays silent |
301
+ | `autoContext.maxItemCharsPerSession` | 4000 | same, in characters |
302
+ | `autoContext.hintMinCoverage` | 0.3 | **absolute** floor for the statistical (hint) channel: IDF-weighted share of the query's information mass the entry covers. A ratio-only threshold cannot tell signal from "best of a bad lot" (`relative:1.00` on an unrelated entry) |
303
+ | `autoContext.hintMinMatched` | 2 | a hint must share at least this many terms with the query — one generic word ("plugin") is not evidence |
304
+ | `autoContext.hintMinSupport` | 0.15 | channel-level silence: if less than this share of the query's terms exist anywhere in the corpus, the hint channel says nothing this round — a long sentence that happens to share one word otherwise reports `cov:1.00` |
305
+ | `autoContext.legacyScope` | `filter` | how to treat a legacy `trigger.scope`: `filter` keeps the old semantics, `ignore` drops it. `npm run selfcheck:triggers` reports entries whose scope values cannot intersect the project tag space |
306
+ | `autoContext.auditLog` | true | append one JSONL line per **actual** injection to `<root>/.dsh-project-memory/injection-audit.jsonl` (what was injected, why it matched, what was dropped, session budget snapshot); rotates to `.1` past `auditMaxBytes` (`262144`). Silent on any I/O error — never affects the host request |
307
+
308
+ ### Injection admission (why it stays quiet)
309
+
310
+ Automatic injection used to be a *retrieval* problem ("which entry is most related to this text?"), which is total — a ranking always returns something, so noise was structural. It is now an **admission** problem ("is this step about to cross a boundary I have been burned by?"), with silence as the default:
311
+
312
+ - **Only `when` triggers**, and it is a low-dimensional typed signal: normalized `ops`, the files this step is about to **write**, and intent words from the human message *after* stripping quoted/path references. Corpus text — raw tool arguments, file contents, filenames — can never trigger anything.
313
+ - **`guard` only narrows.** Extension/name globs (`*.pptx`, `README*`) are ignored outright: they can only lie, never narrow.
314
+ - **Ratio *plus* an absolute floor.** The hint channel needs the relative score *and* an IDF-weighted coverage floor *and* at least two shared terms — `relative:1.00` also happens on entries that share nothing with the step.
315
+ - **Frequency is bounded.** At most one item injection every `gateCooldownSteps`, capped per session by count and characters. The resident task card is exempt (it is a snapshot that should update); the budget is a ceiling, not a target.
316
+ - **Prefix-cache discipline.** Injections are appended as a user message at the tail of the history, so the cached prefix is never rewritten. What they add is resident *cache-read* tokens, not cache misses; nothing is ever edited in place.
317
+ - **It is auditable.** Every real injection appends one line to `injection-audit.jsonl` (reason, dropped candidates, session budget), and `npm run eval:injection` scores 8 labelled scenarios — currently precision 1.00 / recall 1.00 with a clean control group.
299
318
 
300
319
  ### Toggling features
301
320
 
@@ -333,10 +352,14 @@ These commands are for **maintaining the plugin code** — regular users do not
333
352
 
334
353
  ```bash
335
354
  npm install
336
- npm test # 286 tests (184 core + 16 TaskBridge + 11 insight-store + 9 insight-actions + 8 doc-index + 7 auto-inject + 9 host-contract + 5 reflection + 4 llm-route + 2 client-hints + 8 recall + 13 readiness + 6 insight-derive + 4 readiness-eval)
355
+ npm test # 331 tests (184 core + 16 TaskBridge + 11 insight-store + 9 insight-actions + 8 doc-index + 7 auto-inject + 9 host-contract + 5 reflection + 4 llm-route + 2 client-hints + 8 recall + 14 readiness + 7 insight-derive + 6 readiness-eval + 6 ops + 6 injection-audit + 5 injection-budget + 6 injection-scenarios + 18 bugfix-0.5.7)
356
+ npm run eval:injection # scenario P/R: 14/14 hits, 0 false positives, control group clean
357
+ npm run selfcheck:triggers # which entries can still push, which declarations are dead
337
358
  npm run bench -- /path/to/project # index/query performance on any project — no dsh needed
338
359
  ```
339
360
 
361
+ Release notes live in [`CHANGELOG.md`](CHANGELOG.md) and on [GitHub Releases](https://github.com/00080000/dsh-project-memory/releases).
362
+
340
363
  ## License
341
364
 
342
365
  MIT
package/README.zh-CN.md CHANGED
@@ -37,7 +37,7 @@
37
37
  - **文档 ↔ 代码交叉链接** — 文档提及某符号时记录为 `reference`;查询符号时同时带出描述该符号的文档。
38
38
  - **BM25 记忆召回** — 对文档、符号与经验笔记进行排序召回,可选 LLM 查询扩展以应对表述不一致。**CJK 增强**:精确短语乘法加分(3+ 字短语在标题/关键词命中 ×1.5)、同义词表(如 数据库连接池 ↔ 连接池 ↔ DB pool)、CJK 感知的文档↔符号链接边界。
39
39
  - **经验笔记** — 记录问题 → 方案;相似问题覆盖而非重复;笔记仅在检索命中时返回。笔记数量有界:容量随项目规模伸缩(钳制在 100–2000),超限时淘汰最旧的笔记。**覆盖阈值收紧为双向 0.7 重叠**(原 0.6);**经验 `problem` 字段现参与 CJK 短语加分**,提升长尾问句召回。
40
- - **v0.5 分层 insight 记忆(教训 / 决策 / 流程)** — 一个 `insight` 实体贯穿三级:`task`(任务私有草稿,存 `tasks.json`)、`project`(`.dsh-project-memory/insights.json`)、`global`(`~/.config/dsh-project-memory/global.json`)。`save_lesson` 三级可写;去重采用双向 token overlap ≥ 0.7(合并)外加 0.65–0.7 近重复强化带;**提升 = scope 字段变更而非复制**——同一 insight 被 2 个任务命中升 project、3+ 升 global。归档为软删(`archived`),容量/衰减只清归档区;写盘前过滤密钥/token 形态内容。LLM **反思默认关闭**,且只产任务级草稿(`source: reflect`,触发于任务切走/归档时)。面板新增 Task / Project / Global 记忆视图:审核、提升/降级、归档/恢复、删除、编辑与新建表单(procedure 可带"作为 Skill"触发关键词)。旧 `experience.json` 笔记**非破坏**导入 `insights.json` 一次。所有 kind 都可带 authored `trigger`(`keywords` / `symbols` / `actions` / `paths`):命中即**在动手前**确定性注入——v0.5 仅 procedure,就绪层起覆盖全部 kind。
40
+ - **v0.5 分层 insight 记忆(教训 / 决策 / 流程)** — 一个 `insight` 实体贯穿三级:`task`(任务私有草稿,存 `tasks.json`)、`project`(`.dsh-project-memory/insights.json`)、`global`(`~/.config/dsh-project-memory/global.json`)。`save_lesson` 三级可写;去重采用双向 token overlap ≥ 0.7(合并)外加 0.65–0.7 近重复强化带;**提升 = scope 字段变更而非复制**——同一 insight 被 2 个任务命中升 project、3+ 升 global。归档为软删(`archived`),容量/衰减只清归档区;写盘前过滤密钥/token 形态内容。LLM **反思默认关闭**,且只产任务级草稿(`source: reflect`,触发于任务切走/归档时)。面板新增 Task / Project / Global 记忆视图:审核、提升/降级、归档/恢复、删除、编辑与新建表单(procedure 可带"作为 Skill"触发关键词)。旧 `experience.json` 笔记**非破坏**导入 `insights.json` 一次。所有 kind 都可带 authored `trigger`:**只有 `when` 能触发**,`guard` 只能收窄,`prevents` 说明不知道这条会做错什么。`when.ops` 是从工具调用本身解析出的归一动作 id(`file-write` / `file-delete` / `git-commit` / `release` / `npm-publish` / `render-doc` / `run-bench` …),`when.writes` 是这一步**要写**的文件,`when.intents` 是**剥离引号/路径/文件名之后**的人类意图词。命中即**在动手前**确定性注入。旧字段 `keywords`/`symbols`/`actions`/`paths`/`scope` 会在内存里自动迁移(actions→ops、具体路径→writes、keywords→intents,扩展名与泛名 glob、死 action 值一律丢弃);迁移后**没有任何可触发成员**的条目不再被推送,用 `npm run selfcheck:triggers` 看是哪些。
41
41
  - **流式 TF + IDF 缓存** — 查询路径按存储版本缓存 IDF(词逆频率);命中时单次流式遍历 20k 条目(5k 文件)为 p50 2.6 ms / p95 5.4 ms,4k 条目(1k 文件)为 p50 0.6 ms / p95 1.6 ms,零中间对象。只有**真正脏了**的写入才递增版本号并清空缓存——无变更时 `save()` 在碰盘前直接返回,因此 15 秒一轮的 watch 轮询不会把查询刚建好的 IDF 缓存清掉。
42
42
  - **无锁同步事务** — 不采用锁:所有写入(index / watch / remember / forget / watch_repo)统一走同步事务 `store.commit(fn)`,fn 成功后才一次落盘;JS 单线程事件循环保证事务间不交错,`remember`/`forget` 不会被 watch 重索引阻塞排队。全部写入在**进程内**串行;CAS 幂等更新保证同一文件的重复写入不会写坏。但这里**没有跨进程文件锁**——请勿让多个 dsh 实例同时写同一项目存储(见「设计」的一致性一节)。
43
43
  - **依赖极简** — 纯 JavaScript;唯一运行时依赖是 `pdfjs-dist`(PDF 文本提取),无需原生构建。
@@ -147,7 +147,7 @@ dsh plugin --profile web add /path/to/dsh-project-memory.tgz
147
147
  | `/insight`(用户输入,不经模型) | v0.5 记忆视图动作(面板按钮触发):`list [task|project|global]`、`confirm` / `promote` / `demote` / `archive` / `restore` / `delete` `<scope> <id>`、`save <scope> <json>`、`edit <scope> <id> <json>`。 |
148
148
  | `remember problem solution` | 保存经验笔记。相似问题覆盖而非重复。 |
149
149
  | `forget id_or_query` | 删除过期经验笔记。 |
150
- | `save_lesson`(模型工具) | 在 task/project/global 任一作用域保存教训/决策/流程(单一 insight 实体)。近重复按双向 overlap ≥ 0.7 合并、0.65–0.7 强化;同一 insight 被 2+ 任务命中自动 task→project、3+ → global。参数:`title`、`kind`、`scope`、`pattern`/`fix` 或 `choice`/`reason` 或 `steps`、`trigger`(`keywords`/`symbols`/`actions`/`paths`/`scope`,所有 kind 通用,命中即在动手前注入)、`task_id`、`files`、`symbols`、`confidence`、`root`。 |
150
+ | `save_lesson`(模型工具) | 在 task/project/global 任一作用域保存教训/决策/流程(单一 insight 实体)。近重复按双向 overlap ≥ 0.7 合并、0.65–0.7 强化;同一 insight 被 2+ 任务命中自动 task→project、3+ → global。参数:`title`、`kind`、`scope`、`pattern`/`fix` 或 `choice`/`reason` 或 `steps`、`trigger`(`when` = `ops`/`writes`/`intents`,唯一触发面;`guard` = `paths`/`not_paths`/`hosts`/`tags`,只能收窄;`prevents` = 准入条件;旧 `keywords`/`symbols`/`actions`/`paths`/`scope` 仍接受并自动迁移)、`task_id`、`files`、`symbols`、`confidence`、`root`。 |
151
151
 
152
152
  ## 设计
153
153
 
@@ -295,6 +295,25 @@ TaskPanel (Container)
295
295
  | `insight.*` | dedupOverlap `0.7` · reinforceBand `0.65` · maxProject `100` · maxGlobalProcedures `200` · promoteConfidence `0.7` · globalPromoteTasks `3` · decayDays `90` · `globalFile`(自动) | v0.5 insight 去重/强化/提升/容量/归档设置 |
296
296
  | `reflection.enabled` | false | v0.5 LLM 反思,**只写任务级草稿**(触发于任务切走/归档)。`cooldownMs` `1800000`、`maxLessonsPerReflect` `3`、`maxDecisionsPerReflect` `2` |
297
297
  | `autoContext.enabled` | true | v0.5 静默注入包装(entry 常驻块 + relevance)。宿主无法解析会话 cwd 时完全透传(零副作用);`maxTokens` `400`、`editedMax` `3`(resident 任务卡显示最近"编辑中"文件数)、`signalMinRatio` `0.5`(提示至少要达到该层最高分的一半)、`skipEchoSelfTodo` `true`(模型自己写/维护任务清单后、无新人类消息时不回声任务卡,省 token;相关 insights 仍注入)、`budgetLog` `off`(预算丢弃审计写到 stderr:`off` 静默 / `once` 每会话最多一行 / `all` 丢弃组合每变化一次一行。注入按优先级排程,预算不够时丢掉低优先级条目属于**正常降级而非故障**,所以默认不占用用户终端)、`reinjectItemsAfter` `0`(同一条 insight 重复注入的冷却步数;`0` = 正文没变就不在本会话内再注入——注入消息留在会话历史里,重发只是重复占位) |
298
+ | `autoContext.gateCooldownSteps` | 2 | **准入旋钮**:两次*条目*注入之间至少隔几步(常驻任务卡不受限——它是状态快照,内容变了就该更新)。这是"别频繁注入"的主旋钮 |
299
+ | `autoContext.maxItemsPerSession` | 12 | 每会话条目注入条数硬上限;预算是上限不是目标,用尽后条目通道持续沉默 |
300
+ | `autoContext.maxItemCharsPerSession` | 4000 | 同上,按字符计 |
301
+ | `autoContext.hintMinCoverage` | 0.3 | 提示通道的**绝对**下限:条目覆盖了查询多少 IDF 加权信息量。只用相对阈值分不出"有信号"和"矮子里拔将军"(实测无关条目也拿 `relative:1.00`) |
302
+ | `autoContext.hintMinMatched` | 2 | 提示还必须至少共享这么多个词:单个通用词("插件")不构成证据 |
303
+ | `autoContext.hintMinSupport` | 0.15 | 通道级沉默:查询里能在语料中找到对应的词占比低于此值时,提示通道本轮整体不出声——否则一句只碰巧共享一个词的长句子会报出 `cov:1.00` |
304
+ | `autoContext.legacyScope` | `filter` | 旧 `trigger.scope` 的处理:`filter` 保留旧语义,`ignore` 丢弃。`npm run selfcheck:triggers` 会列出 scope 值与项目画像 tag 空间不可能相交的条目 |
305
+ | `autoContext.auditLog` | true | 每次**真实**注入往 `<root>/.dsh-project-memory/injection-audit.jsonl` 追加一行(注入了什么、为什么命中、丢了什么、会话额度快照);超过 `auditMaxBytes`(`262144`)轮转 `.1`。任何 IO 失败都静默,绝不影响宿主请求 |
306
+
307
+ ### 注入的准入化(为什么它保持安静)
308
+
309
+ 自动注入过去是个**检索**问题("哪条记忆和这段文本最相关")——而检索是全函数,排序永远有答案,所以噪声是结构性的。现在它是个**准入**问题("这一步是否即将跨过我踩过坑的边界"),默认沉默:
310
+
311
+ - **只有 `when` 能触发**,且是低维的类型化信号:归一 `ops`、这一步**要写**的文件、**剥离引用之后**的人类意图词。语料(原始工具参数、文件正文、文件名)永远不能触发任何东西。
312
+ - **`guard` 只收窄**。扩展名/泛名 glob(`*.pptx`、`README*`)被硬性忽略:它们只能撒谎,不能收窄。
313
+ - **相对分 + 绝对下限**。提示通道要同时满足相对分、IDF 加权覆盖率下限、以及至少两个共同词——`relative:1.00` 也会出现在和这一步毫无关系的条目上。
314
+ - **频率有上限**。每 `gateCooldownSteps` 步最多一次条目注入,每会话还有条数与字符上限;常驻任务卡不受限(它是快照),预算是上限不是目标。
315
+ - **前缀缓存纪律**。注入以 user 消息追加在历史尾部,缓存前缀永不被改写;它带来的是常驻的 cache-read token,不是缓存失效;没有任何内容被原地改写。
316
+ - **可审计**。每次真实注入往 `injection-audit.jsonl` 落一行(原因、被丢弃的候选、会话额度),`npm run eval:injection` 跑 8 个标注场景——当前精确率 1.00 / 召回率 1.00,对照组零注入。
298
317
 
299
318
  ### 功能开关
300
319
 
@@ -332,10 +351,14 @@ dsh web --patch ./config.yml
332
351
 
333
352
  ```bash
334
353
  npm install
335
- npm test # 286 项测试(核心 184 + TaskBridge 16 + insight-store 11 + insight-actions 9 + doc-index 8 + auto-inject 7 + host-contract 9 + reflection 5 + llm-route 4 + client-hints 2 + recall 8 + readiness 13 + insight-derive 6 + readiness-eval 4)
354
+ npm test # 331 项测试(核心 184 + TaskBridge 16 + insight-store 11 + insight-actions 9 + doc-index 8 + auto-inject 7 + host-contract 9 + reflection 5 + llm-route 4 + client-hints 2 + recall 8 + readiness 14 + insight-derive 7 + readiness-eval 6 + ops 6 + injection-audit 6 + injection-budget 5 + injection-scenarios 6 + bugfix-0.5.7 18)
355
+ npm run eval:injection # 场景 P/R:命中 14/14、假阳性 0、对照组零注入
356
+ npm run selfcheck:triggers # 哪些条目还推得动、哪些声明是死的
336
357
  npm run bench -- /你的/项目路径 # 对任意项目量索引/查询性能,不需要 dsh
337
358
  ```
338
359
 
360
+ 发布说明见 [`CHANGELOG.md`](CHANGELOG.md) 与 [GitHub Releases](https://github.com/00080000/dsh-project-memory/releases)。
361
+
339
362
  ## 许可证
340
363
 
341
364
  MIT
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@yolk_vat-y/dsh-project-memory",
3
- "version": "0.5.5",
3
+ "version": "0.5.7",
4
4
  "description": "Persistent project memory for dsh agents: index docs (PDF/Markdown/text) and code symbols into a searchable per-workspace store, recall them with cited sources, and keep experience entries (problems -> solutions) searchable on demand.",
5
5
  "type": "module",
6
6
  "main": "src/index.js",
@@ -18,7 +18,9 @@
18
18
  "url": "https://github.com/00080000/dsh-project-memory.git"
19
19
  },
20
20
  "scripts": {
21
- "test": "node test/run-test.mjs && node test/taskbridge.test.mjs && node test/insight-store.test.mjs && node test/reflection-pipeline.test.mjs && node test/auto-inject.test.mjs && node test/insight-actions.test.mjs && node test/host-contract.test.mjs && node test/llm-route.test.mjs && node test/doc-index.test.mjs && node test/client-hints.test.mjs && node test/recall.test.mjs && node test/readiness.test.mjs && node test/insight-derive.test.mjs && node test/readiness-eval.test.mjs",
21
+ "test": "node test/run-test.mjs && node test/taskbridge.test.mjs && node test/insight-store.test.mjs && node test/reflection-pipeline.test.mjs && node test/auto-inject.test.mjs && node test/insight-actions.test.mjs && node test/host-contract.test.mjs && node test/llm-route.test.mjs && node test/doc-index.test.mjs && node test/client-hints.test.mjs && node test/recall.test.mjs && node test/readiness.test.mjs && node test/insight-derive.test.mjs && node test/readiness-eval.test.mjs && node test/ops.test.mjs && node test/injection-audit.test.mjs && node test/injection-budget.test.mjs && node test/injection-scenarios.test.mjs && node test/bugfix-0.5.7.test.mjs",
22
+ "eval:injection": "node test/injection-scenarios.test.mjs",
23
+ "selfcheck:triggers": "node test/injection-scenarios.test.mjs --selfcheck",
22
24
  "bench": "node scripts/bench.mjs",
23
25
  "bench:synthetic": "node scripts/bench-synthetic.mjs",
24
26
  "build:client": "tsdown"
package/src/audit.js ADDED
@@ -0,0 +1,87 @@
1
+ // 注入审计(PLAN S0):把每次**真实进入上下文**的注入决策落成 JSONL。
2
+ //
3
+ // 为什么需要:注入的有效性无法在线逐条归因——宿主只给 turn 级聚合(cacheReadTokens /
4
+ // inputTokens),拿不到"这一条记忆值多少"。所以"当时注入了哪几条、因为什么命中、同一轮
5
+ // 哪些被预算挤掉"必须落盘才可离线复盘。缺了它,任何 trigger / 阈值改动都只能靠感觉。
6
+ //
7
+ // 契约(与 auto-inject 同款安全约定):
8
+ // - 只在确有决策(本条真的进了上下文)时写一行:dropped 单独出现不写,否则每步刷屏;
9
+ // - 任何异常(目录不可写、磁盘满、记录序列化失败)一律吞掉,绝不影响宿主请求;
10
+ // - 单文件超限就地轮转一份 `.1`,不引入新的清理线程或后台任务。
11
+ import { appendFileSync, mkdirSync, renameSync, statSync } from 'node:fs'
12
+ import path from 'node:path'
13
+
14
+ export const AUDIT_FILE = 'injection-audit.jsonl'
15
+ const DEFAULT_MAX_BYTES = 256 * 1024
16
+
17
+ /** 审计配置:默认开(观测是这个插件唯一的仪表盘),显式 `auditLog: false` 才关。 */
18
+ export function cfgAudit(config) {
19
+ const c = (config && config.autoContext) || {}
20
+ return {
21
+ enabled: c.auditLog !== false,
22
+ maxBytes: typeof c.auditMaxBytes === 'number' && c.auditMaxBytes > 0 ? c.auditMaxBytes : DEFAULT_MAX_BYTES,
23
+ }
24
+ }
25
+
26
+ export function auditFileFor(memoryDir) {
27
+ return path.join(memoryDir, AUDIT_FILE)
28
+ }
29
+
30
+ /**
31
+ * 一次注入的审计记录(纯函数,可单测)。字段刻意保持"能直接回答两个问题":
32
+ * 注入了什么(injected)、为什么没注入别的(dropped)。
33
+ * @param {object} input { sessionId, root, step, text, labels, reasons, dropped }
34
+ */
35
+ export function auditRecordFrom(input) {
36
+ const reasons = Array.isArray(input.reasons) ? input.reasons : []
37
+ const dropped = Array.isArray(input.dropped) ? input.dropped : []
38
+ return {
39
+ at: new Date().toISOString(),
40
+ session: input.sessionId || null,
41
+ root: input.root || null,
42
+ step: typeof input.step === 'number' ? input.step : null,
43
+ chars: typeof input.text === 'string' ? input.text.length : 0,
44
+ labels: Array.isArray(input.labels) ? [...input.labels] : [],
45
+ injected: reasons.map((r) => ({
46
+ id: r.id,
47
+ channel: r.channel,
48
+ why: r.why,
49
+ chars: typeof r.chars === 'number' ? r.chars : null,
50
+ })),
51
+ dropped: dropped.map((d) => ({ id: d.id, channel: d.channel, reason: d.reason })),
52
+ // 会话级额度快照与"本轮为什么沉默":注入频率本身是可观测指标,不该只能靠感觉。
53
+ budget: input.budget || null,
54
+ silence: input.silence || null,
55
+ }
56
+ }
57
+
58
+ /**
59
+ * 追加一行审计。返回是否写入成功(调用方不必关心,仅供测试断言)。
60
+ * @param {string} memoryDir 项目记忆目录(`<root>/.dsh-project-memory`)
61
+ * @param {object} record {@link auditRecordFrom} 的产物
62
+ * @param {{enabled?: boolean, maxBytes?: number}} [cfg] {@link cfgAudit} 的产物
63
+ */
64
+ export function appendInjectionAudit(memoryDir, record, cfg) {
65
+ const c = cfg || {}
66
+ if (c.enabled === false) return false
67
+ if (!memoryDir || !record) return false
68
+ try {
69
+ mkdirSync(memoryDir, { recursive: true })
70
+ const file = auditFileFor(memoryDir)
71
+ rotateIfOversized(file, c.maxBytes || DEFAULT_MAX_BYTES)
72
+ appendFileSync(file, `${JSON.stringify(record)}\n`)
73
+ return true
74
+ } catch {
75
+ // 审计是旁路:写不进去不能影响这一轮注入,更不能影响宿主的请求
76
+ return false
77
+ }
78
+ }
79
+
80
+ function rotateIfOversized(file, maxBytes) {
81
+ try {
82
+ if (statSync(file).size < maxBytes) return
83
+ renameSync(file, `${file}.1`)
84
+ } catch {
85
+ // 文件不存在(首次写入)或轮转失败:下一次 append 会重新建文件;失败就不轮转
86
+ }
87
+ }