@a9i5k4/dsh-auto-memory 2.5.3 → 3.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +171 -1
- package/README.zh-CN.md +171 -1
- package/docs/CONTRIBUTORS.html +471 -0
- package/docs/HANDOFF-CRITERIA.md +92 -0
- package/docs/INTEGRATION-ANALYSIS.md +350 -348
- package/docs/USER-GUIDE.en.md +56 -1
- package/docs/USER-GUIDE.zh-CN.md +57 -2
- package/docs/internal/ACCEPT-35-LIVE.md +143 -0
- package/docs/internal/ACCEPTANCE-20260914.md +90 -0
- package/docs/internal/ARCH-REVIEW-BRIEF.md +411 -0
- package/docs/internal/ARCH-REVIEW-REQUEST.md +201 -0
- package/docs/internal/ARCH-REVIEW-ROUND2.md +169 -0
- package/docs/internal/ARCH-REVIEW-ROUND3.md +206 -0
- package/docs/internal/AUDIT-WB-GRAPH-FULL-20260916.md +314 -0
- package/docs/internal/CONCURRENCY-INVESTIGATION-20260917.md +192 -0
- package/docs/internal/CROSS-SESSION-SEARCH-PATH-DECISION.md +72 -0
- package/docs/internal/CROSS-SESSION-SEARCH-RESEARCH.md +131 -0
- package/docs/internal/DECISIONS-20260914-SESSION.md +269 -0
- package/docs/internal/DESIGN-P1-STATE-COMMIT-20260915.md +219 -0
- package/docs/internal/DIRECTION-CHECK-WB-GRAPH-20260916.md +132 -0
- package/docs/internal/FEEDBACK-TO-DSHAPI-RELAY.md +13 -0
- package/docs/internal/GH-DISCUSSION-5732-COMMENT.md +74 -0
- package/docs/internal/GPT-ACCEPTANCE-PROMPT-20260916.md +352 -0
- package/docs/internal/GPT-REVIEW-PROMPT.md +216 -0
- package/docs/internal/GROUP-WEBHOOK-SETUP.md +33 -0
- package/docs/internal/KICKOFF-P0.md +254 -0
- package/docs/internal/MASTER-PLAN-3.0.md +411 -0
- package/docs/internal/MEMORY-MUTATION-AND-INDEX-DESIGN.md +85 -0
- package/docs/internal/MERGE-CONFLICT-SCAN-20260914.md +222 -0
- package/docs/internal/PENDING-FIXES-20260916.md +289 -0
- package/docs/internal/RAG-KARPATHY-PROGRAM.md +229 -0
- package/docs/internal/REPORT-P0-NIGHTLY.md +212 -0
- package/docs/internal/REPORT-P5-ACCEPTANCE.md +31 -0
- package/docs/internal/REPORT-WB-GRAPH-NIGHTLY.md +153 -0
- package/docs/internal/REVIEW-WB-GRAPH-SELF.md +81 -0
- package/docs/internal/ROADMAP-20260917-WEEK.md +305 -0
- package/docs/internal/ROADMAP.md +106 -0
- package/docs/internal/RUN-P0-NIGHTLY.md +227 -0
- package/docs/internal/S10-CONSTRUCTION-HANDOFF-20260917.md +175 -0
- package/docs/internal/S10-GAPS-PLAIN-20260917.md +125 -0
- package/docs/internal/SEMANTIC-ARCHITECTURE-SPEC.md +360 -0
- package/docs/internal/SESSION-FILE-REPAIR-PROTOCOL.md +90 -0
- package/docs/internal/THREE-LAYER-CONTRACT.md +210 -0
- package/docs/internal/TODO-BACKLOG.md +263 -142
- package/docs/internal/TODO-GRAPH.html +715 -0
- package/docs/internal/TODO-GRAPH.html.bak-20260914-v2 +493 -0
- package/docs/internal/TODO-GRAPH.html.bak-20260915-alsfix +710 -0
- package/docs/internal/TODO-GRAPH.html.bak-20260915-p1 +710 -0
- package/docs/internal/TODO-GRAPH.html.bak-20260915-p6a-rev +703 -0
- package/docs/internal/TODO-GRAPH.html.bak-20260915-wshint +710 -0
- package/docs/internal/TODO-GRAPH.html.bak-20260916-batch +715 -0
- package/docs/internal/WB-FORMAT-CONVENTION.md +112 -0
- package/docs/internal/WB-GRAPH-DECISIONS-20260914.md +71 -0
- package/docs/internal/reviews/CLAIM-VERIFICATION-20260914.md +56 -0
- package/docs/internal/reviews/PLAN-gpt6astra-round2-20260914.md +787 -0
- package/docs/internal/reviews/REVIEW-gpt6astra-20260914.md +112 -0
- package/docs/internal/reviews/ROUND3-REVIEW-INTEGRATION-20260914.md +230 -0
- package/docs/prompts/M8-3-enable-verify.md +49 -49
- package/lib/acceptance.js +71 -0
- package/lib/activation-host.js +90 -9
- package/lib/activation-inbox.js +25 -7
- package/lib/board-mode.js +30 -0
- package/lib/client.js +878 -75
- package/lib/context-bridge.js +2 -2
- package/lib/context-host.js +70 -6
- package/lib/engine-identity.js +149 -0
- package/lib/engine-switch.js +247 -0
- package/lib/episodic-store.js +11 -10
- package/lib/evidence-store.js +2 -2
- package/lib/fact-store.js +1 -1
- package/lib/fs-retry.js +46 -0
- package/lib/index.js +1987 -153
- package/lib/intent-clean-safe.js +40 -0
- package/lib/intent-clean.js +12 -16
- package/lib/l0-extract.js +263 -149
- package/lib/l0-index-sync.js +195 -0
- package/lib/l0-index.js +349 -239
- package/lib/ledger-criteria.js +142 -0
- package/lib/m7-index-sync-host.js +65 -4
- package/lib/m7-wire.js +3 -3
- package/lib/memory-anchor.js +56 -1
- package/lib/memory-envelope.js +252 -0
- package/lib/memory-hub.js +14 -4
- package/lib/memory-mutation.js +246 -0
- package/lib/memory-writer.js +204 -24
- package/lib/procedure-observation.js +48 -0
- package/lib/procedure-store.js +34 -17
- package/lib/python-setup.js +1 -1
- package/lib/rerank-host.js +160 -0
- package/lib/rules-layer.js +261 -0
- package/lib/semantic-js.js +15 -0
- package/lib/shadow-retrieval.js +3 -3
- package/lib/state-commit.js +245 -0
- package/lib/subagent-gc.js +4 -8
- package/lib/tier-layer-inject.js +650 -0
- package/lib/tier0-catalog.js +693 -0
- package/lib/water-window.js +263 -186
- package/lib/wb-contract.js +495 -0
- package/lib/wb-sidecar.js +839 -0
- package/lib/ws-overview-rank.js +2 -2
- package/package.json +1 -1
- package/python/m7_embedding_v1.py +5 -5
- package/python/worker_semantic_v1.py +17 -6
- package/python/worker_v1.py +38 -4
package/README.md
CHANGED
|
@@ -53,7 +53,27 @@
|
|
|
53
53
|
</details>
|
|
54
54
|
|
|
55
55
|
<p align="center">
|
|
56
|
-
<a href="README.zh-CN.md"
|
|
56
|
+
<a href="./README.zh-CN.md"><img alt="中文" src="https://img.shields.io/badge/%E4%B8%AD%E6%96%87-switch-lightgrey?style=for-the-badge"></a>
|
|
57
|
+
<a href="./README.md"><img alt="English" src="https://img.shields.io/badge/English-current-blue?style=for-the-badge"></a>
|
|
58
|
+
</p>
|
|
59
|
+
|
|
60
|
+
<p align="center">
|
|
61
|
+
<a href="https://www.npmjs.com/package/@a9i5k4/dsh-auto-memory"><img alt="npm" src="https://img.shields.io/npm/v/@a9i5k4/dsh-auto-memory"></a>
|
|
62
|
+
<a href="LICENSE"><img alt="License" src="https://img.shields.io/badge/License-BSD--3--Clause-yellow.svg"></a>
|
|
63
|
+
<img alt="Runtime dependencies" src="https://img.shields.io/badge/runtime%20deps-0-brightgreen">
|
|
64
|
+
<img alt="Platform" src="https://img.shields.io/badge/Platform-Windows%20%7C%20macOS%20%7C%20Linux-lightgrey">
|
|
65
|
+
</p>
|
|
66
|
+
|
|
67
|
+
<p align="center">
|
|
68
|
+
<code>pnpm add @a9i5k4/dsh-auto-memory</code>
|
|
69
|
+
</p>
|
|
70
|
+
|
|
71
|
+
<p align="center">
|
|
72
|
+
<a href="docs/USER-GUIDE.en.md"><strong>📖 User guide</strong></a> ·
|
|
73
|
+
<a href="docs/USER-GUIDE.zh-CN.md"><strong>📖 用户手册</strong></a> ·
|
|
74
|
+
<a href="CHANGELOG.md">Changelog</a> ·
|
|
75
|
+
<a href="https://htmlpreview.github.io/?https://github.com/Aik358/dsh-auto-memory/blob/main/docs/CONTRIBUTORS.html">Contributors & Sponsors</a> ·
|
|
76
|
+
<a href="https://qm.qq.com/q/v7Asxn6vPa">QQ group</a>
|
|
57
77
|
</p>
|
|
58
78
|
|
|
59
79
|
---
|
|
@@ -83,6 +103,7 @@ Now we push this route to its last missing piece — when the context fills, she
|
|
|
83
103
|
| **External memory inheritance** | Memories from WorkBuddy / CodeBuddy / Claude Code / Codex are scanned, importable, per-source managed |
|
|
84
104
|
| **Production-grade hygiene** | Write gate (mojibake/stutter/JSON-injection blocking) + dirty-token scanner + credentials never enter prompts |
|
|
85
105
|
| **Astra-style context management** | A filling context no longer collapses into one summary — four-part handoff notes carry work across windows, full history stays searchable, the agent retrieves on demand (on by default, threshold 0.75) |
|
|
106
|
+
| **No cross-talk between workspaces** | Open several workspaces or sessions at once and each keeps its own recall decisions and index cache. Clicking into one never disturbs the one that's running |
|
|
86
107
|
| **Model-agnostic** | No vendor lock, no tier lock: any model on DSH works out of the box — lexical 0GB floor, built-in ~130MB semantic tier, advanced 563MB |
|
|
87
108
|
| **Portable memory** | Everything lives on your own disk; memories scan in from other AI tools, every entry has an evidence chain — auditable, deletable. Memory belongs to you, not to any vendor |
|
|
88
109
|
|
|
@@ -177,6 +198,98 @@ Ask, and she answers: `memory_recall` returns a layered summary list first (each
|
|
|
177
198
|
|
|
178
199
|
The panel's Workspace tab draws all of this as a mind map: workspaces at the center, memory topics as branches, dashed lines for cross-workspace shares; draggable, zoomable, click a card for details. **Your memory has a shape for the first time.**
|
|
179
200
|
|
|
201
|
+
### Retrieval isn't "dump all memory in" — three tiers, descending (OpenViking-style)
|
|
202
|
+
|
|
203
|
+
This borrows **OpenViking**'s tiering idea, recalibrated against real corpus measurements in this repo. The rule is **descend tier by tier — never all at once**:
|
|
204
|
+
|
|
205
|
+
| Tier | Content | Budget | When it appears |
|
|
206
|
+
|---|---|---|---|
|
|
207
|
+
| **Tier-0 · Catalog** (index layer) | One line per entry = title · conclusion · layer · status · date | ≤ `B0` = **800 tokens** (resident, ~30 entries) | **The norm** — only this tier is injected by default |
|
|
208
|
+
| **Tier-1 · Digests** (candidate layer) | ≤ `L1` = **140 chars** each, top `K` = **8** | 8 × 140 = 1120 chars | Only when the catalog under-hits |
|
|
209
|
+
| **Tier-2 · Source chunks** (evidence layer) | `chunkId = hash(memoryId, digest, index)` | ≤ `B2` = **2400 chars** / call | Only when evidence is needed |
|
|
210
|
+
|
|
211
|
+
Flow: `Tier-0 out first → narrow → descend to Tier-1 only if short → fetch Tier-2 only for evidence`. Single-turn injection still respects `injectBudgetChars` (default 8000 chars).
|
|
212
|
+
|
|
213
|
+
**Why it's built this way (measured, not guessed)**:
|
|
214
|
+
- **The native corpus is smaller than you'd think**: source text p90 is only **1196 chars**, max **1692**. So OpenViking's `L0 → L1(2k) → L2` middle step can be dropped — jumping from a 140-char digest straight to a ≤2400 source chunk is an acceptable span.
|
|
215
|
+
- **Pure priority makes tiering collapse**: on real corpus, the `project` layer's 77 chunks **consumed the entire 800-token budget**, leaving `whiteboard` / `user` / `log` with **zero** entries — "tiered" degenerating into single-tier. Hence **per-layer quotas** (e.g. `project ≤ 60% · B0`) — a hard rule, not a suggestion.
|
|
216
|
+
- **Reconciliation**: spot-checking real `expand` output against corpus length — `mem_d55f8e8e` reported 1499 chars ↔ corpus 1499 ✅, `mem_5a7f779a` reported 1403 ↔ 1403 ✅.
|
|
217
|
+
|
|
218
|
+
### The "Karpathy module" — the whiteboard *is* the corpus: feeding the memory flow graph straight to the AI
|
|
219
|
+
|
|
220
|
+
This line started from a sentence: **"dsh graph is exactly the Karpathy module I wanted."** The problem it solves isn't "is retrieval accurate" — it's **the shape of memory**. Traditional RAG rediscovers from zero on every query, accumulating nothing; the Karpathy-style answer is to let a model incrementally maintain a persistent wiki (entity pages, concept pages, cross-references, contradiction flags) navigated by `index.md` + `log.md`.
|
|
221
|
+
|
|
222
|
+
What this repo shipped isn't "yet another wiki" — it found an interface that needs **zero new machinery**: **pages *are* the corpus (S10.1)**.
|
|
223
|
+
|
|
224
|
+
> **A whiteboard card carries an anchor, and the anchor is what the Tier-0 catalog slices by.** The slicing order is 「anchor → heading → top-level item → whole-file fallback」, so an anchored card becomes its own entry in the **guidance layer injected every turn** — no new pipeline required.
|
|
225
|
+
|
|
226
|
+
```markdown
|
|
227
|
+
### Card title
|
|
228
|
+
<!-- memory:mem_<32hex> -->
|
|
229
|
+
```
|
|
230
|
+
|
|
231
|
+
**How the anchor is computed (content-addressed, recomputable)**:
|
|
232
|
+
`mem_` + first 32 hex chars of `sha256(workspaceKey + '\0' + page relative path + '\0' + card title)`.
|
|
233
|
+
The whiteboard text is one of **five sources** (`user` / `project` / `log` / `whiteboard` / `reflection`) feeding the Tier-0 catalog, and it holds a **reserved floor quota** (`whiteboard` and `user` each reserve `floorRatio·maxTokens`) — so the `project` layer can't swallow the whole 800-token budget and lock the whiteboard out.
|
|
234
|
+
|
|
235
|
+
> **⚠️ Implementation boundary (stated plainly, not oversold)**: the anchor→corpus path is currently wired **only for injection** (Tier-0 catalog, via `add('whiteboard', …)` in `index.js`).
|
|
236
|
+
> The `memory_recall` **retrieval** corpus still contains only four source classes — daily logs / reflections / project notes / user-level memory — **the whiteboard is not yet wired in** (see the four `pushL0` call sites and the four `semSources` entries in `lib/index.js`).
|
|
237
|
+
> In other words: whiteboard content **is injected every turn**, but `memory_recall` **still cannot search it**. This is a contract-layer item defined but not yet implemented, and it is outside the scope of this change.
|
|
238
|
+
|
|
239
|
+
The payoff cuts both ways:
|
|
240
|
+
- **The kanban/whiteboard stops being "a view for humans only"** — it is simultaneously a **source of injection corpus**. Goals, criteria and conclusions you write on the board are visible to the AI in the guidance layer on the next turn;
|
|
241
|
+
- It **also eases the "two sources of truth" problem**: content exists once, on the board; the guidance layer is derived from it, so you can't get "board says A, index says B".
|
|
242
|
+
|
|
243
|
+
**The six Karpathy-style landing items** (S10.1–S10.6):
|
|
244
|
+
|
|
245
|
+
| Item | What it does |
|
|
246
|
+
|---|---|
|
|
247
|
+
| **S10.1 Pages are the corpus** | Whiteboard cards carry anchors ⇒ sliced by the Tier-0 guidance layer and injected every turn (**retrieval side not yet wired — see the boundary note above**) |
|
|
248
|
+
| **S10.2 Index auto-derived** | `index` is **derived** from pages (link + one line + `layer`/`status`) and must **not** be written twice; the derived result must match Tier-0 catalog entries |
|
|
249
|
+
| **S10.3 Lint completed** | Four zero-token classes: orphan entries / stale / mentioned-but-no-page / missing cross-reference. **Only "contradiction detection" needs an LLM, and it must be user-triggered — never on the automatic path**; lint **reports, never auto-fixes** |
|
|
250
|
+
| **S10.4 No state machine** | The whiteboard is a **view layer**; state belongs to memory entries' `layer` + `status`. Adding a state machine = a violation |
|
|
251
|
+
| **S10.5 Answers flow back** | The conclusion of any search/analysis must be one-click depositable as: ① a new whiteboard card (anchored) ② a memory entry ③ a handoff ledger line. **Any conclusion that "lives only in the conversation" is a process failure** |
|
|
252
|
+
| **S10.6 Human/model split** | `<!-- model -->` / `<!-- user -->` sections; a full rewrite must carry the user section back verbatim — **after a model rewrite, the user section is preserved byte-for-byte** |
|
|
253
|
+
|
|
254
|
+
**Sequencing discipline (a lesson actually paid for)**: the **contract layer (format / anchors / index derivation) must come before the RAG substrate** — it determines the shape of the corpus; build RAG first and you redo the corpus. The UI layer (board rendering / interaction) comes *after* RAG instead — "it's only a view".
|
|
255
|
+
|
|
256
|
+
**A path explicitly rejected**: Grep agentic ("model-driven glob/grep beats everything"). Its premises are **multi-turn LLM tool calls every round** (a token multiplier) and a corpus that is **exactly token-matchable** — but natural-language memory has no literal string to grep. We took only the phrase "whether to search should be judged intelligently", and **handed the default judgement to a local linear classifier, fv2 (0 tokens)**.
|
|
257
|
+
|
|
258
|
+
### Activation eligibility: judging "should this be recalled", not "how similar is it"
|
|
259
|
+
|
|
260
|
+
This is the most technical — and most misunderstood — part of the system. **Semantic relevance ≠ activation eligibility**: material can be very similar to the current topic and still be unhelpful to inject right now. We split the decision into two separately measurable targets: **semantic relevance** and **activation eligibility**.
|
|
261
|
+
|
|
262
|
+
**The failure mode we named and measured: the echo trap.**
|
|
263
|
+
When a user restates a memory ("you said X, right?"), that memory's semantic score is **necessarily high** — yet injecting it now is redundant. On 86 human gold labels: the highest suppress-class score (the noodle echo at **0.6507** and its variant 0.6254) **exceeds every activate positive (max 0.5914)**. In other words — **the threshold that "looks safest" lands squarely on echoes**.
|
|
264
|
+
|
|
265
|
+
**The two-arm echo rule**: echo = 「query and top-1 candidate form a near-duplicate restatement」∧「declarative mood」∧「no recall intent」.
|
|
266
|
+
Lexical arm (bigram containment ≥ θ) and semantic arm (`denseTop ≥ 0.75`) combine with **OR** — either arm alone was falsified as insufficient; combined, they hit zero false negatives and zero false positives across the 86 gold labels.
|
|
267
|
+
|
|
268
|
+
**Three falsified shortcuts** (negative results, published in the paper):
|
|
269
|
+
1. Lexical containment **alone cannot** detect echoes — the echo-suppress group's median containment (0.273) is *lower* than the activate group's (0.462), because questions naturally share the target's terminology;
|
|
270
|
+
2. Pure text tri-classification **cannot** judge eligibility (macroF1 0.494);
|
|
271
|
+
3. **A global upfront echo veto harms explicit follow-ups** — moving it out of the global front and into the proactive lane improved the best operating point from precision 0.818 / recall 0.237 / 1 violation to **1.000 / 0.289 / 0 violations**.
|
|
272
|
+
|
|
273
|
+
**Calibration and feature weights** (reproducible numbers):
|
|
274
|
+
- Sigmoid calibration lifted intent-head accuracy **0.744 → 0.872** and Brier **0.227 → 0.131** (58 gold);
|
|
275
|
+
- LR coefficients of the deployable feature set: `mark` (question/recall markers) **+1.64** ≫ `containment` **+0.94** > `intentProb` **+0.58** > `margin` **+0.27** ≫ `denseTop` **−0.38**.
|
|
276
|
+
- **In one line**: "**is this a question**" matters an order of magnitude more than "**how similar is it**".
|
|
277
|
+
|
|
278
|
+
### Three deployment tiers: the size–quality curve (all measured)
|
|
279
|
+
|
|
280
|
+
| Tier | Size | Runtime | L2 R@5 | Role |
|
|
281
|
+
|---|---|---|---|---|
|
|
282
|
+
| Lexical BM25 (`lexical_pre_v2`) | **0** | Pure JS | 0.200 | The always-available floor and final fallback |
|
|
283
|
+
| **JS semantic tier** (transformers.js + `multilingual-e5-small` q8) | ~**130MB** | Node-side ONNX | **0.850** | The standard tier, yours on `npm install` |
|
|
284
|
+
| **Python advanced tier** (BGE-M3 int8 ONNX 563MB / fp32 2.3GB) | optional 2nd tier | sidecar | **0.925** | Quality champion, enabled on demand in the wizard |
|
|
285
|
+
|
|
286
|
+
JS tier key numbers: model load **679ms**, query encode **3.8ms**, full rebuild of 251 entries **5.5s**.
|
|
287
|
+
int8 key numbers: head-to-head with fp32 **R@5 delta = 0.000**, MRR gap 0.007 (noise level), mean vector cosine 0.975, encode speedup **6×** (44s vs 262s), single-query p50 **16ms**.
|
|
288
|
+
⇒ **Quantization loss is zero in ranking terms**, so the 563MB tier can replace fp32 as the default.
|
|
289
|
+
The e5-small vs BGE-M3 gap (0.85 vs 0.925) concentrates on hard-negative twin pairs — the small model is still far better than pure lexical (**+65pt**), but adversarial near-neighbour discrimination is a **capacity problem, not a protocol problem**.
|
|
290
|
+
|
|
291
|
+
> Every conclusion comes from reproducible experiments and is frozen into an engineering decision ledger (D1–D11): [retrieval model selection](docs/M7-RESEARCH-PAPER.md) · [activation policy v2](docs/M7-ACTIVATION-V2-PAPER.md) · [embedding benchmark](docs/M7-EMBEDDING-BENCHMARK.md) · [held-out human gold evaluation](docs/M7-ACTIVATION-V2-HOLDEDOUT-EVAL.md) (67 human-scored items: actPrecision **0.917** / harmful injections **0** / echo layer **7/7**).
|
|
292
|
+
|
|
180
293
|
---
|
|
181
294
|
|
|
182
295
|
## How she reminds
|
|
@@ -209,6 +322,28 @@ She can also look things back up herself: `memory_search` queries the full archi
|
|
|
209
322
|
|
|
210
323
|
Token water-level awareness completes it: as the window fills, she suggests opening a new window and handing off, instead of silently compressing. The window is the host's territory — she midwifes the handoff, and never decides for the host.
|
|
211
324
|
|
|
325
|
+
### Why handoff isn't "write a summary" — the engineering of whiteboard and ledger
|
|
326
|
+
|
|
327
|
+
**The problem**: asking an LLM to produce a paragraph of "what happened before" for a new window looks simple but degrades — each compaction loses a layer, and after a few rounds the handoff material is out of sync with reality, **with nobody able to tell that it is**. So the rule here is: **handoff material must not get to speak for itself either — it has to be traceable, decidable, and regression-tested.**
|
|
328
|
+
|
|
329
|
+
Three entities, each minding one job:
|
|
330
|
+
|
|
331
|
+
| Entity | What it is | Where it lives |
|
|
332
|
+
|---|---|---|
|
|
333
|
+
| **Handoff ledger** | A fixed **four-part** shape: task state / goals / approaches tried and why they failed / progress and next step | `handoff/handoff-*.md` |
|
|
334
|
+
| **Whiteboard** (PLAN.md) | The project's **whole-picture map**: a human-readable planning snapshot; old versions are archived on rewrite | `handoff/PLAN.md` |
|
|
335
|
+
| **Anchor** | Every record carries `<!-- memory:mem_<32hex> -->` — **identity addressing**, not positional addressing | inside the memory file |
|
|
336
|
+
|
|
337
|
+
**Two hard engineering constraints**:
|
|
338
|
+
1. **Ledger quality is a hard gate, not a style suggestion.** The four headings must match **verbatim** (a wrong heading breaks downstream weighted truncation and the injection-side parser), each section has a minimum length, and a bad one is rejected outright — the very shape of the section you're reading was formed by that gate.
|
|
339
|
+
2. **Anchors stop append-writes from piercing the file.** `appendAnchoredRecord()` parses the existing file first: any non-clean state **fails closed** (neither appends nor rewrites); if reserved syntax appears in the body, the write is **refused on the spot with a line number** — rather than "succeeding" and then making the whole file permanently unwritable from the next write onward.
|
|
340
|
+
|
|
341
|
+
**The kanban board is not a separate UI** — it is **another view over the same ledgers and whiteboard**: ledger entries laid out as a swimlane matrix (goals / in progress / failures and detours / archived).
|
|
342
|
+
|
|
343
|
+
> A real recorded misstep: board v1 squeezed 92 files into 92 cards and left three swimlanes permanently empty. **The root cause was not "too few cards" but a slicing granularity off by one level** (slicing by file instead of by the sections inside), compounded by a missing `break` in `sectionOf` that collapsed every document into its last heading. v2 aligned granularity to "section".
|
|
344
|
+
|
|
345
|
+
**Where the deep end lives**: the cross-window continuity internals are under `docs/internal/` — the [three-tier contract](docs/internal/THREE-LAYER-CONTRACT.md) (Tier 0/1/2 budgets and acceptance predicates), the [semantic architecture spec](docs/internal/SEMANTIC-ARCHITECTURE-SPEC.md) (clauses S1–S10 and stage gates), and the [RAG + Karpathy program](docs/internal/RAG-KARPATHY-PROGRAM.md) (six-step pipeline × three stage lines as a build map).
|
|
346
|
+
|
|
212
347
|
---
|
|
213
348
|
|
|
214
349
|
## How she moves in
|
|
@@ -354,6 +489,23 @@ Config file `~/.dsh/dsh-auto-memory.json` (everything adjustable in the Settings
|
|
|
354
489
|
- **Centralized storage**: all workspace memory under one root (`~/.dsh/memory/workspaces/`), readable from any session
|
|
355
490
|
- **30-day distillation**: old logs are AI-distilled into project notes; originals archived, nothing lost
|
|
356
491
|
|
|
492
|
+
### The 3.0 rebuild (invisible to you — but every recall goes through it)
|
|
493
|
+
|
|
494
|
+
Most of this release adds no new buttons. It changes *what makes a memory trustworthy*. Eight mechanisms, all shipped and measured:
|
|
495
|
+
|
|
496
|
+
| Mechanism | The problem it kills |
|
|
497
|
+
|---|---|
|
|
498
|
+
| **Write-side gate** | Reserved-syntax filtering moved *into the write primitives* instead of being detected after the fact. Content carrying a reserved marker is refused on the spot **with a line number** — no more "one bad line makes the whole file permanently unwritable" |
|
|
499
|
+
| **Decision ledger** | Every "should I recall this?" call becomes a reviewable ledger entry; five-grade scoring (A/P/S/H/E) flows back into policy. Not a log — an **auditable chain of decisions** |
|
|
500
|
+
| **Single-source state commit** | The memory index version (miv) converges on one source, killing the "one store, two version numbers" class of recompute-and-miss bugs |
|
|
501
|
+
| **Concurrent atomic writes** | On Windows a `rename` hitting an external file handle throws `EPERM`. Now it retries with backoff, and on final failure it **preserves the full candidate snapshot** (`recoveryPath`) for forensics instead of destroying already-rendered content |
|
|
502
|
+
| **Engine identity gate** | The JS and Python semantic implementations are **mutually exclusive identities**: whichever you pick is the one that runs — never shadowing, never cross-triggering |
|
|
503
|
+
| **True incremental embedding** | Only changed records get re-embedded, with reuse ordering — instead of recomputing the whole store |
|
|
504
|
+
| **Bounded rerank window** | Optional rerank tiers (off/fast/enthusiast): the clock starts at enqueue, 60s expiry with no renewal, LRU ≤16, yields when busy — background work never slows the live conversation |
|
|
505
|
+
| **Multi-workspace / multi-session isolation** | With several workspaces and sessions open at once, each keeps its own gate decisions, index cache, and degradation state. **Clicking into a workspace can no longer disturb the one that's actually running** |
|
|
506
|
+
|
|
507
|
+
> Each of the eight ships with a regression suite and a mutation demo (revert the mechanism to its old behaviour and the tests must genuinely go red). Engineering detail lives in [`docs/internal/`](docs/internal/).
|
|
508
|
+
|
|
357
509
|
---
|
|
358
510
|
|
|
359
511
|
## UI gallery
|
|
@@ -465,8 +617,26 @@ Papers were authored by the autonomous engineering agent (ZCode / GLM); all conc
|
|
|
465
617
|
|
|
466
618
|
Community contributors:
|
|
467
619
|
|
|
620
|
+
- [@Minervaowl7](https://github.com/Minervaowl7) — the most prolific contributor: 15 PRs + 8 issues covering workspace-overview log-date anchoring, auto-continuation host hardening, and recovery-candidate lifecycle ([#16](https://github.com/Aik358/dsh-auto-memory/issues/16)–[#53](https://github.com/Aik358/dsh-auto-memory/pull/53))
|
|
621
|
+
- [@JIE42393](https://github.com/JIE42393) — 7 issues on panel behaviour, recall quality and configuration edge cases ([#15](https://github.com/Aik358/dsh-auto-memory/issues/15), [#26](https://github.com/Aik358/dsh-auto-memory/issues/26), [#30](https://github.com/Aik358/dsh-auto-memory/issues/30), [#41](https://github.com/Aik358/dsh-auto-memory/issues/41)–[#43](https://github.com/Aik358/dsh-auto-memory/issues/43), [#45](https://github.com/Aik358/dsh-auto-memory/issues/45))
|
|
622
|
+
- [@Fishsb](https://github.com/Fishsb) — 3 issues on memory recall and injection behaviour ([#18](https://github.com/Aik358/dsh-auto-memory/issues/18)–[#20](https://github.com/Aik358/dsh-auto-memory/issues/20))
|
|
623
|
+
- [@messiahyl](https://github.com/messiahyl) — 2 issues ([#8](https://github.com/Aik358/dsh-auto-memory/issues/8), [#9](https://github.com/Aik358/dsh-auto-memory/issues/9))
|
|
468
624
|
- [@ProperSAMA](https://github.com/ProperSAMA) — panel readability fix for DSH Desktop enhanced mode (transparent/Mica materials) + entry-button anti-occlusion & outside-click/Esc close ([PR #12](https://github.com/Aik358/dsh-auto-memory/pull/12))
|
|
469
625
|
- [@nkh0472](https://github.com/nkh0472) — unattended/batch workflow hardening feedback that drove the welcome tour and per-feature switches ([Issue #10](https://github.com/Aik358/dsh-auto-memory/issues/10))
|
|
626
|
+
- [@fei009009](https://github.com/fei009009) — pull request ([#29](https://github.com/Aik358/dsh-auto-memory/pull/29))
|
|
627
|
+
- [@alexchenzl](https://github.com/alexchenzl) ([#6](https://github.com/Aik358/dsh-auto-memory/issues/6)) · [@ALuoXue](https://github.com/ALuoXue) ([#2](https://github.com/Aik358/dsh-auto-memory/issues/2)) · [@Architectxz](https://github.com/Architectxz) ([#1](https://github.com/Aik358/dsh-auto-memory/issues/1)) · [@cuohua](https://github.com/cuohua) ([#40](https://github.com/Aik358/dsh-auto-memory/issues/40)) · [@eclgo](https://github.com/eclgo) ([#13](https://github.com/Aik358/dsh-auto-memory/issues/13)) · [@jeffsui](https://github.com/jeffsui) ([#39](https://github.com/Aik358/dsh-auto-memory/issues/39)) · [@lhbsaa](https://github.com/lhbsaa) ([#3](https://github.com/Aik358/dsh-auto-memory/issues/3)) · [@moonltppt](https://github.com/moonltppt) ([#14](https://github.com/Aik358/dsh-auto-memory/issues/14)) · [@swtseaman](https://github.com/swtseaman) ([#21](https://github.com/Aik358/dsh-auto-memory/issues/21)) · [@xiaochaZ](https://github.com/xiaochaZ) ([#38](https://github.com/Aik358/dsh-auto-memory/issues/38)) · [@zjj871114037](https://github.com/zjj871114037) ([#7](https://github.com/Aik358/dsh-auto-memory/issues/7)) — bug reports and feature requests
|
|
628
|
+
|
|
629
|
+
Full credits, including infrastructure sponsors: **[Contributors & Sponsors](https://htmlpreview.github.io/?https://github.com/Aik358/dsh-auto-memory/blob/main/docs/CONTRIBUTORS.html)**
|
|
630
|
+
|
|
631
|
+
---
|
|
632
|
+
|
|
633
|
+
## Sponsors
|
|
634
|
+
|
|
635
|
+
Development resources for this project are partly provided by:
|
|
636
|
+
|
|
637
|
+
- **[DSH API](https://api.dshapi.icu/)** — API relay station providing the model endpoints used for development, testing, and the semantic-engine research behind the M-series features. Thank you for keeping the lights on.
|
|
638
|
+
|
|
639
|
+
Infrastructure and API-quota sponsors are listed on the **[Contributors & Sponsors](https://htmlpreview.github.io/?https://github.com/Aik358/dsh-auto-memory/blob/main/docs/CONTRIBUTORS.html)** page. If you would like to support the project, open an issue or join the QQ group.
|
|
470
640
|
|
|
471
641
|
---
|
|
472
642
|
|
package/README.zh-CN.md
CHANGED
|
@@ -53,7 +53,27 @@
|
|
|
53
53
|
</details>
|
|
54
54
|
|
|
55
55
|
<p align="center">
|
|
56
|
-
<
|
|
56
|
+
<a href="./README.zh-CN.md"><img alt="中文" src="https://img.shields.io/badge/%E4%B8%AD%E6%96%87-%E5%BD%93%E5%89%8D-blue?style=for-the-badge"></a>
|
|
57
|
+
<a href="./README.md"><img alt="English" src="https://img.shields.io/badge/English-switch-lightgrey?style=for-the-badge"></a>
|
|
58
|
+
</p>
|
|
59
|
+
|
|
60
|
+
<p align="center">
|
|
61
|
+
<a href="https://www.npmjs.com/package/@a9i5k4/dsh-auto-memory"><img alt="npm" src="https://img.shields.io/npm/v/@a9i5k4/dsh-auto-memory"></a>
|
|
62
|
+
<a href="LICENSE"><img alt="License" src="https://img.shields.io/badge/License-BSD--3--Clause-yellow.svg"></a>
|
|
63
|
+
<img alt="Runtime dependencies" src="https://img.shields.io/badge/runtime%20deps-0-brightgreen">
|
|
64
|
+
<img alt="Platform" src="https://img.shields.io/badge/Platform-Windows%20%7C%20macOS%20%7C%20Linux-lightgrey">
|
|
65
|
+
</p>
|
|
66
|
+
|
|
67
|
+
<p align="center">
|
|
68
|
+
<code>pnpm add @a9i5k4/dsh-auto-memory</code>
|
|
69
|
+
</p>
|
|
70
|
+
|
|
71
|
+
<p align="center">
|
|
72
|
+
<a href="docs/USER-GUIDE.zh-CN.md"><strong>📖 用户手册</strong></a> ·
|
|
73
|
+
<a href="docs/USER-GUIDE.en.md"><strong>📖 User guide</strong></a> ·
|
|
74
|
+
<a href="CHANGELOG.md">更新日志</a> ·
|
|
75
|
+
<a href="https://htmlpreview.github.io/?https://github.com/Aik358/dsh-auto-memory/blob/main/docs/CONTRIBUTORS.html">贡献者与赞助</a> ·
|
|
76
|
+
<a href="https://qm.qq.com/q/v7Asxn6vPa">QQ 交流群</a>
|
|
57
77
|
</p>
|
|
58
78
|
|
|
59
79
|
---
|
|
@@ -83,6 +103,7 @@ dsh-auto-memory 从第一天就不信这件事只能如此。她把记忆放在
|
|
|
83
103
|
| **外部记忆继承** | WorkBuddy / CodeBuddy / Claude Code / Codex 的历史记忆可扫描、导入、按源管理 |
|
|
84
104
|
| **生产级卫生** | 写入门禁(乱码/复读/JSON 注入拦截)+ 脏 token 扫描 + 凭证永不进提示词 |
|
|
85
105
|
| **Astra 式上下文管理** | 上下文将满不再压成一段摘要——四段式交接笔记跨窗口续命,全量历史归档可搜,Agent 按需检索(默认开启,阈值 0.75) |
|
|
106
|
+
| **多工作区不串线** | 同时开多个工作区 / 多个会话,各自的唤起判据与索引缓存互不覆盖——你点哪个工作区,都不影响正在跑的那个 |
|
|
86
107
|
| **模型无关** | 不锁厂商、不锁档位:DSH 上任何模型即装即得,词法 0GB 保底、内置语义 ~130MB、进阶 563MB |
|
|
87
108
|
| **记忆可携带** | 全部存在你自己的盘上;跨 AI 工具扫描导入,每条有证据链、可审计、可删除——记忆属于你,不属于任何厂商 |
|
|
88
109
|
|
|
@@ -177,6 +198,98 @@ dsh-auto-memory 从第一天就不信这件事只能如此。她把记忆放在
|
|
|
177
198
|
|
|
178
199
|
面板「工作区」页签把这一切画成一张关系图:中心是工作区,分支是记忆主题,虚线是跨区共享;可拖拽、可缩放、点卡片看详情。**你的记忆第一次有了形状。**
|
|
179
200
|
|
|
201
|
+
### 检索不是"把记忆全塞进去"——三层下探(OpenViking 式)
|
|
202
|
+
|
|
203
|
+
这里参考了 **OpenViking** 的分层思路,但按本仓语料的实测尺寸重新标定。核心是**逐层下探、不同时给**:
|
|
204
|
+
|
|
205
|
+
| 层 | 内容 | 预算 | 何时出现 |
|
|
206
|
+
|---|---|---|---|
|
|
207
|
+
| **Tier-0 · 目录**(指引层) | 每条 1 行 = 标题 · 结论 · 层 · 状态 · 日期 | ≤ `B0` = **800 token**(常驻,约 30 条) | **常态**:默认只注入这一层 |
|
|
208
|
+
| **Tier-1 · 摘要**(候选层) | 每条 ≤ `L1` = **140 字符**,取 `K` = **8** 条 | 8 × 140 = 1120 字符 | 目录命中不足才下探 |
|
|
209
|
+
| **Tier-2 · 原文块**(证据层) | `chunkId = hash(记忆ID, 内容摘要, 序号)` | ≤ `B2` = **2400 字符** / 次 | 需要证据才取 |
|
|
210
|
+
|
|
211
|
+
流程:`Tier-0 先出 → 缩窄 → 命中不足才下探 Tier-1 → 需要证据才取 Tier-2`。单轮注入总长仍受 `injectBudgetChars`(默认 8000 字符)约束。
|
|
212
|
+
|
|
213
|
+
**为什么这么设计(实测,不是拍脑袋)**:
|
|
214
|
+
- **原生语料比想象的小**:原文本身 p90 只有 **1196 字符**、max **1692**。所以 OpenViking 那种 `L0 → L1(2k) → L2` 的中间档可以省掉——从 140 字符的摘要直接跳到 ≤2400 的原文块,跨度可接受。
|
|
215
|
+
- **纯优先级会让分层退化**:真实语料实测,`project` 层 77 个块能把 800 token **全占**,`whiteboard` / `user` / `log` 一条都进不来——"分层"就成了单层。因此**每层有配额**(如 `project ≤ 60% · B0`),这是硬规则不是建议。
|
|
216
|
+
- **对账口径**:抽真实 `expand` 与语料长度比对,`mem_d55f8e8e` 报 1499 字符 ↔ 语料 1499 ✅,`mem_5a7f779a` 报 1403 ↔ 1403 ✅。
|
|
217
|
+
|
|
218
|
+
### 「Karpathy 模块」——白板即语料:把记忆流程图直接喂给 AI
|
|
219
|
+
|
|
220
|
+
这条线的起点是一句原话:**「dsh graph 正是我想要的 Karpathy 模块」**。它要解决的不是"检索得准不准",而是**记忆的形态**:传统 RAG 每次查询都从零重新发现,没有知识积累;Karpathy 式的做法是让模型渐进维护一个持久 wiki(实体页 / 概念页 / 交叉引用 / 矛盾标注),靠 `index.md` + `log.md` 导航。
|
|
221
|
+
|
|
222
|
+
本仓的落地不是"再写一套 wiki",而是找到了一个**零新机制**的接口——**页面即语料(S10.1)**:
|
|
223
|
+
|
|
224
|
+
> **白板卡片带锚点,凭锚点被 Tier-0 目录切条**——切分顺序是「锚点 → 标题 → 顶层条目 → 整文件兜底」,所以带锚点的卡片会作为独立条目进入**每轮注入的指引层**,不需要任何新管线。
|
|
225
|
+
|
|
226
|
+
```markdown
|
|
227
|
+
### 卡片标题
|
|
228
|
+
<!-- memory:mem_<32hex> -->
|
|
229
|
+
```
|
|
230
|
+
|
|
231
|
+
**锚点怎么算(内容寻址,可复算)**:
|
|
232
|
+
`mem_` + `sha256(workspaceKey + '\0' + 页面相对路径 + '\0' + 卡片标题)` 的前 32 位十六进制。
|
|
233
|
+
白板文本作为**五个来源之一**(`user` / `project` / `log` / `whiteboard` / `reflection`)进入 Tier-0 目录,并有**保底配额**(`whiteboard` 与 `user` 各保底 `floorRatio·maxTokens`)——防止 `project` 层把 800 token 全占、白板一条都进不来。
|
|
234
|
+
|
|
235
|
+
> **⚠️ 实现边界(如实标注,不夸大)**:锚点→语料这条目前**只打通了「注入」路径**(Tier-0 目录,`index.js` 的 `add('whiteboard', …)`)。
|
|
236
|
+
> **`memory_recall` 的检索语料目前仍只含四类**——日志 / 反思 / 项目笔记 / 用户级记忆,**白板尚未接入**(见 `lib/index.js` 中 `pushL0` 的四个调用点与 `semSources` 的四个来源)。
|
|
237
|
+
> 也就是说:白板内容**每轮会被注入**,但**主动 `memory_recall` 还搜不到它**。这是契约层已定、实现待补的一项缺口,不在本轮改动范围内。
|
|
238
|
+
|
|
239
|
+
这个设计的收益是双向的:
|
|
240
|
+
- **看板/白板不再只是"给人看的视图"**,它同时是**注入语料的来源**——你在看板上记下的目标、判据、结论,AI 下一轮就能在指引层看到;
|
|
241
|
+
- **顺带缓解"双状态源"**:内容只有白板一份,指引层是它的派生,不会出现"白板说 A、索引说 B"。
|
|
242
|
+
|
|
243
|
+
**Karpathy 式 wiki 的六条落地**(S10.1–S10.6):
|
|
244
|
+
|
|
245
|
+
| 条款 | 做什么 |
|
|
246
|
+
|---|---|
|
|
247
|
+
| **S10.1 页面即语料** | 白板卡片带锚点 ⇒ 被 Tier-0 指引层切条并每轮注入(**检索侧尚未接入,见上方边界说明**) |
|
|
248
|
+
| **S10.2 索引自动生成** | `index` 由页面**派生**(链接 + 一句话 + `layer`/`status`),与白板**不得各写一份**;派生结果须与 Tier-0 目录条目一致 |
|
|
249
|
+
| **S10.3 lint 补齐** | 零 token 四类:孤立条目 / 陈旧 / 被提及却无独立页 / 缺交叉引用。**只有"矛盾检测"需要 LLM,且必须用户点一下才跑,不得进自动路径**;lint **只报告不自动改** |
|
|
250
|
+
| **S10.4 不建状态机** | 白板是**视图层**;状态归记忆条目的 `layer` + `status`。新增状态机 = 违规 |
|
|
251
|
+
| **S10.5 答案归档回流** | 一次检索/分析的结论必须能**一键**沉淀为:① 白板新卡(带锚点)② 记忆条目 ③ handoff 账本一条。**任何"只活在对话里"的结论都算流程不合格** |
|
|
252
|
+
| **S10.6 人机分区** | `<!-- model -->` / `<!-- user -->` 分区,整篇重写必须原样带回用户段——**模型重写后用户段逐字节保留** |
|
|
253
|
+
|
|
254
|
+
**顺序纪律(踩过的坑)**:**契约层(格式 / 锚点 / 索引派生)必须先于 RAG 底层**——它决定语料形状;先做 RAG 就得对语料重做一遍。界面层(看板渲染 / 交互)反而排在 RAG 之后,"它只是视图"。
|
|
255
|
+
|
|
256
|
+
**被明确否决的路线**:Grep agentic("模型驱动 glob/grep 打败一切")。前提是**每轮多轮 LLM 工具调用**(token 乘数)且语料**精确 token 可匹配**——而自然语言记忆没有可 grep 的字面。只吸收了那句「要不要搜由智能判断」,**默认判断交给本地线性分类器 fv2(0 token)**。
|
|
257
|
+
|
|
258
|
+
### 唤起度:判断"该不该想起",而不是"有多像"
|
|
259
|
+
|
|
260
|
+
这是整个系统最技术、也最容易被误解的一点。**语义相关 ≠ 唤起必要**——材料与当前话题很像,不代表现在注入它有帮助。我们把决策拆成两个可分别度量的目标:**语义相关性**与**唤起必要性**。
|
|
261
|
+
|
|
262
|
+
**发现并命名的失效模式:回声陷阱(echo trap)**。
|
|
263
|
+
当用户复述了某条记忆("你说过 X 对吧"),该记忆的语义分**必然很高**——但此时注入它是冗余的。在 86 条人工金标上的实测分布:suppress 类的最高分(面条回声 **0.6507** 及其变体 0.6254)**超过全部 activate 正例(max 0.5914)**。也就是说——**"看起来最安全"的阈值,恰好稳定地踩中回声**。
|
|
264
|
+
|
|
265
|
+
**回声的双臂规则**:回声 = 「查询与 top-1 候选构成近重复复述」∧「陈述句式」∧「无回忆意图」。
|
|
266
|
+
词面臂(bigram containment ≥ θ)与语义臂(`denseTop ≥ 0.75`)**取或**——任一单臂都被证伪不足,组合后在 86 gold 上零漏报零误报。
|
|
267
|
+
|
|
268
|
+
**三条被证伪的捷径**(写进论文的负面结果):
|
|
269
|
+
1. 词面覆盖率**单独不可判**回声——echo-suppress 组的 containment 中位数(0.273)反而**低于** activate 组(0.462),因为问句天然共享目标术语;
|
|
270
|
+
2. 纯文本三分类**不可判**必要性(macroF1 0.494);
|
|
271
|
+
3. **全局前置回声否决会误伤显式追问**——把它从全局前置移入 proactive lane 后,同一数据最优格由 precision 0.818 / recall 0.237 / 越界 1 改善为 **1.000 / 0.289 / 越界 0**。
|
|
272
|
+
|
|
273
|
+
**校准与特征权重**(可直接复现的数字):
|
|
274
|
+
- sigmoid 校准把意图头准确率 **0.744 → 0.872**、Brier **0.227 → 0.131**(58 gold);
|
|
275
|
+
- 可部署特征集的 LR 系数:`mark`(疑问/回忆标记) **+1.64** ≫ `containment` **+0.94** > `intentProb` **+0.58** > `margin` **+0.27** ≫ `denseTop` **−0.38**。
|
|
276
|
+
- **结论一句话**:「**是否在问**」比「**有多像**」重要一个数量级。
|
|
277
|
+
|
|
278
|
+
### 三级部署:体积—质量曲线(都实测过)
|
|
279
|
+
|
|
280
|
+
| 层 | 体积 | 运行时 | L2 R@5 | 定位 |
|
|
281
|
+
|---|---|---|---|---|
|
|
282
|
+
| 词法 BM25(`lexical_pre_v2`) | **0** | 纯 JS | 0.200 | 人人可用的基础层与最终回退 |
|
|
283
|
+
| **JS 标准语义层**(transformers.js + `multilingual-e5-small` q8) | ~**130MB** | Node 内 ONNX | **0.850** | npm 安装即得的标准层 |
|
|
284
|
+
| **Python 进阶层**(BGE-M3 int8 ONNX 563MB / fp32 2.3GB) | 二档可选 | sidecar | **0.925** | 效果冠军,向导按需启用 |
|
|
285
|
+
|
|
286
|
+
JS 层关键数字:模型加载 **679ms**、查询编码 **3.8ms**、251 条全库重建 **5.5s**。
|
|
287
|
+
int8 关键数字:与 fp32 同口径 head-to-head **R@5 delta = 0.000**、MRR 差 0.007(噪声级)、向量余弦均值 0.975、编码提速 **6×**(44s vs 262s)、单查询 p50 **16ms**。
|
|
288
|
+
⇒ **量化损失在排序意义上为零**,所以 563MB 档可作为 fp32 的默认替代。
|
|
289
|
+
e5-small 与 BGE-M3 的差距(0.85 vs 0.925)集中在 hard-negative 双子对——小模型仍显著优于纯词法(**+65pt**),但对抗式近邻区分是**容量问题,不是协议问题**。
|
|
290
|
+
|
|
291
|
+
> 全部结论来自可复现实验并冻结为工程决策台账(D1–D11):[检索选型研究](docs/M7-RESEARCH-PAPER.md) · [激活策略 v2 技术报告](docs/M7-ACTIVATION-V2-PAPER.md) · [嵌入基准](docs/M7-EMBEDDING-BENCHMARK.md) · [Held-out 人工金标验收](docs/M7-ACTIVATION-V2-HOLDEDOUT-EVAL.md)(67 条人工打分:actPrecision **0.917** / 有害注入 **0** / echo 层 **7/7**)。
|
|
292
|
+
|
|
180
293
|
---
|
|
181
294
|
|
|
182
295
|
## 她怎么提醒
|
|
@@ -209,6 +322,28 @@ dsh-auto-memory 从第一天就不信这件事只能如此。她把记忆放在
|
|
|
209
322
|
|
|
210
323
|
配 token 水位感知:快满时,她提示你开新窗口交接,而不是默默压缩。窗口是宿主的领地——她只助产交接,从不越权替宿主做决定。
|
|
211
324
|
|
|
325
|
+
### 交接为什么不是"写一段摘要"——白板与账本的工程
|
|
326
|
+
|
|
327
|
+
**问题**:让 LLM 生成一段"之前的进展"交给新窗口,看起来简单,实际会退化——每次压缩都丢一层,几轮之后交接材料与真实状态脱节,而**没人能发现它脱节了**。所以这里的做法是:**交接材料也不能自说自话,它得是可追溯、可判定、可回归的**。
|
|
328
|
+
|
|
329
|
+
三件实体,各管一段:
|
|
330
|
+
|
|
331
|
+
| 实体 | 是什么 | 落盘位置 |
|
|
332
|
+
|---|---|---|
|
|
333
|
+
| **交接账本**(handoff) | 固定**四段式**:任务状态 / 目标 / 已试方案与失败原因 / 进度与下一步 | `handoff/handoff-*.md` |
|
|
334
|
+
| **白板**(PLAN.md) | 项目**全貌图**:人能读的规划快照,改版时旧版自动归档 | `handoff/PLAN.md` |
|
|
335
|
+
| **锚点**(memory anchor) | 每条记录带 `<!-- memory:mem_<32hex> -->`,**身份寻址**而非位置寻址 | 记忆文件内 |
|
|
336
|
+
|
|
337
|
+
**两条关键工程约束**:
|
|
338
|
+
1. **账本质量是硬门,不是文风建议**。四段标题必须**逐字匹配**(标题错会让后续的权重化截断失效、注入端解析失败),每段有最小长度,写好直接拒收——本节开头那段「任务状态 / 目标 / 已试方案与失败原因 / 进度与下一步」就是被这个门塑形的结果。
|
|
339
|
+
2. **锚点让追加写入不会击穿文件**。`appendAnchoredRecord()` 在写入前先解析既有文件:**非 clean 状态一律 fail closed**(既不追加也不改写);正文里一旦出现保留语法,**写入当场被拒并报出行号**——而不是写入"成功"、从下一次开始整个文件永久拒写。
|
|
340
|
+
|
|
341
|
+
**看板**(Kanban / 泳道图)不是另做一个 UI,而是**同一批账本与白板的另一种视图**:把账本条目按泳道(目标 / 进行中 / 失败与弯路 / 归档)铺成矩阵。
|
|
342
|
+
|
|
343
|
+
> 一个真实的踩坑记录:看板 v1 曾把 92 个文件压成 92 张卡、三条泳道恒空——**根因不是"卡太少",而是切分粒度错了一个层级**(按文件切,而不是按内容里的小节切),外加 `sectionOf` 循环缺 `break` 导致整篇归一到最后一个标题。v2 才把粒度对齐到「小节」。
|
|
344
|
+
|
|
345
|
+
**主视觉与子代理的边界**:跨窗口续命的细节在 `docs/internal/` —— [三层契约](docs/internal/THREE-LAYER-CONTRACT.md)(Tier 0/1/2 的预算与验收判据)、[语义架构规范](docs/internal/SEMANTIC-ARCHITECTURE-SPEC.md)(条款 S1–S10 与阶段门)、[RAG + Karpathy 攻关细则](docs/internal/RAG-KARPATHY-PROGRAM.md)(六步链路 × 三条阶段线的施工图)。
|
|
346
|
+
|
|
212
347
|
---
|
|
213
348
|
|
|
214
349
|
## 她怎么搬家
|
|
@@ -356,6 +491,23 @@ cd ~/.dsh/profiles/web && pnpm up @a9i5k4/dsh-auto-memory
|
|
|
356
491
|
- **集中式存储**:全部工作区记忆收在 `~/.dsh/memory/workspaces/` 一个根下,任意会话可读
|
|
357
492
|
- **30 天蒸馏**:旧日志由 AI 蒸馏进项目笔记,原文归档不丢失
|
|
358
493
|
|
|
494
|
+
### 3.0 底层重建(用户看不见,但每次唤起都经过它)
|
|
495
|
+
|
|
496
|
+
这一版的大部分工作不产生新按钮——它们改的是"记忆凭什么被相信"。八项底层机制全部落地并实测:
|
|
497
|
+
|
|
498
|
+
| 机制 | 解决的问题 |
|
|
499
|
+
|---|---|
|
|
500
|
+
| **写入侧门禁** | 保留语法过滤前移到**写入原语**(而非事后检测)。含保留标记的正文会被当场拒绝并给出**行号**,不再出现"一条脏正文让整个文件永久拒写" |
|
|
501
|
+
| **判据账本** | 每次「要不要唤起」的决策落成可复核的账本条目,五档评分(A/P/S/H/E)回流成策略——不是日志,是**可审计的判据链** |
|
|
502
|
+
| **状态单源提交** | 记忆索引版本(miv)收敛到单一来源,消除"同一个库两个版本号"导致的重算与漏判 |
|
|
503
|
+
| **并发原子写** | Windows 下 `rename` 遇外部句柄占用会抛 `EPERM`——现在带退避重试,且失败时**保全完整候选快照**(`recoveryPath`)供人工取证,不再把已渲染好的内容一起销毁 |
|
|
504
|
+
| **引擎身份门** | JS 与 Python 两套语义实现**身份互斥**:选了谁就是谁,绝不互相顶替、不互相联动 |
|
|
505
|
+
| **真增量嵌入** | 只对变化的记录重新嵌入并复用顺序,而不是整库重算 |
|
|
506
|
+
| **精排有界窗口** | 可选精排档位(off/fast/enthusiast),入队起算 60s 到期不续命、LRU ≤16、忙则让路——后台重活永远不拖慢当前对话 |
|
|
507
|
+
| **多工作区 / 多会话隔离** | 同时开多个工作区、多个会话时,各自的唤起判据、索引缓存、降级状态**互不覆盖**——你点进哪个工作区,都不影响正在跑的那个 |
|
|
508
|
+
|
|
509
|
+
> 这八项都带回归套件与变异演示(把机制改回旧行为,测试必须真红)。工程细节见 [`docs/internal/`](docs/internal/)。
|
|
510
|
+
|
|
359
511
|
---
|
|
360
512
|
|
|
361
513
|
## 界面速览
|
|
@@ -467,8 +619,26 @@ DeepSeek Harness (Node, 127.0.0.1:3080)
|
|
|
467
619
|
|
|
468
620
|
社区贡献者:
|
|
469
621
|
|
|
622
|
+
- [@Minervaowl7](https://github.com/Minervaowl7) — 贡献最活跃:15 个 PR + 8 个 issue,覆盖工作区概览日志日期锚定、自动续跑宿主加固、恢复候选生命周期等([#16](https://github.com/Aik358/dsh-auto-memory/issues/16)–[#53](https://github.com/Aik358/dsh-auto-memory/pull/53))
|
|
623
|
+
- [@JIE42393](https://github.com/JIE42393) — 7 个 issue,覆盖面板行为、召回质量与配置边界([#15](https://github.com/Aik358/dsh-auto-memory/issues/15)、[#26](https://github.com/Aik358/dsh-auto-memory/issues/26)、[#30](https://github.com/Aik358/dsh-auto-memory/issues/30)、[#41](https://github.com/Aik358/dsh-auto-memory/issues/41)–[#43](https://github.com/Aik358/dsh-auto-memory/issues/43)、[#45](https://github.com/Aik358/dsh-auto-memory/issues/45))
|
|
624
|
+
- [@Fishsb](https://github.com/Fishsb) — 3 个 issue,关于记忆召回与注入行为([#18](https://github.com/Aik358/dsh-auto-memory/issues/18)–[#20](https://github.com/Aik358/dsh-auto-memory/issues/20))
|
|
625
|
+
- [@messiahyl](https://github.com/messiahyl) — 2 个 issue([#8](https://github.com/Aik358/dsh-auto-memory/issues/8)、[#9](https://github.com/Aik358/dsh-auto-memory/issues/9))
|
|
470
626
|
- [@ProperSAMA](https://github.com/ProperSAMA) — DSH Desktop 增强模式(透明/Mica 材质)面板可读性修复 + 入口按钮防遮挡与外点/Esc 关闭([PR #12](https://github.com/Aik358/dsh-auto-memory/pull/12))
|
|
471
627
|
- [@nkh0472](https://github.com/nkh0472) — 无人值守/批处理场景加固反馈,推动了欢迎向导与功能开关化([Issue #10](https://github.com/Aik358/dsh-auto-memory/issues/10))
|
|
628
|
+
- [@fei009009](https://github.com/fei009009) — 提交 PR([#29](https://github.com/Aik358/dsh-auto-memory/pull/29))
|
|
629
|
+
- [@alexchenzl](https://github.com/alexchenzl)([#6](https://github.com/Aik358/dsh-auto-memory/issues/6))· [@ALuoXue](https://github.com/ALuoXue)([#2](https://github.com/Aik358/dsh-auto-memory/issues/2))· [@Architectxz](https://github.com/Architectxz)([#1](https://github.com/Aik358/dsh-auto-memory/issues/1))· [@cuohua](https://github.com/cuohua)([#40](https://github.com/Aik358/dsh-auto-memory/issues/40))· [@eclgo](https://github.com/eclgo)([#13](https://github.com/Aik358/dsh-auto-memory/issues/13))· [@jeffsui](https://github.com/jeffsui)([#39](https://github.com/Aik358/dsh-auto-memory/issues/39))· [@lhbsaa](https://github.com/lhbsaa)([#3](https://github.com/Aik358/dsh-auto-memory/issues/3))· [@moonltppt](https://github.com/moonltppt)([#14](https://github.com/Aik358/dsh-auto-memory/issues/14))· [@swtseaman](https://github.com/swtseaman)([#21](https://github.com/Aik358/dsh-auto-memory/issues/21))· [@xiaochaZ](https://github.com/xiaochaZ)([#38](https://github.com/Aik358/dsh-auto-memory/issues/38))· [@zjj871114037](https://github.com/zjj871114037)([#7](https://github.com/Aik358/dsh-auto-memory/issues/7))— 问题反馈与功能建议
|
|
630
|
+
|
|
631
|
+
完整名单(含基础设施赞助):**[贡献者与赞助](https://htmlpreview.github.io/?https://github.com/Aik358/dsh-auto-memory/blob/main/docs/CONTRIBUTORS.html)**
|
|
632
|
+
|
|
633
|
+
---
|
|
634
|
+
|
|
635
|
+
## 赞助
|
|
636
|
+
|
|
637
|
+
本项目部分开发资源由以下方提供:
|
|
638
|
+
|
|
639
|
+
- **[DSH API](https://api.dshapi.icu/)** — API 中转站,为本项目的开发、测试以及 M 系列语义引擎研究提供模型端点。感谢一路同行。
|
|
640
|
+
|
|
641
|
+
基础设施与 API 额度赞助方列在 **[贡献者与赞助](https://htmlpreview.github.io/?https://github.com/Aik358/dsh-auto-memory/blob/main/docs/CONTRIBUTORS.html)** 页面。如希望支持本项目,欢迎提 issue 或加入 QQ 交流群。
|
|
472
642
|
|
|
473
643
|
---
|
|
474
644
|
|