@kenz1117/dsh-engram 0.7.4 → 0.7.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.en.md CHANGED
@@ -73,15 +73,17 @@ Sensory and emotional dimensions (smell, temperature, emotional weight) are deli
73
73
  - **Capture redaction**: inbound content passes a regex scrub for common secrets and credentials (sk- API keys, Bearer, AWS AKIA, GitHub tokens, PEM private keys, password/token assignments) and matched fragments become `[REDACTED:<type>]`.
74
74
  - **Recall placeholder (anti-echo-chamber)**: memory-recall tool output inside captured slices is replaced with `[engram memory result omitted from capture: <tool>]`, and the extraction prompt states that restating existing memory is not new information, breaking the memory self-reinforcement loop.
75
75
  - **Multi-query retrieval**: `engram_search` can use the aux LLM to rewrite the query into ≤3 complementary queries, retrieving each and fusing them with cross-query RRF plus a per-query floor; a failed rewrite degrades to the single query (`queryRewrite: false` disables).
76
+ - **Evidence gate (search → assess)**: a hit means "relevant", not "enough to answer". Every retrieval registers an in-process batch (latest 20 per session, released when the session ends) and prints `ref=…` per row plus the batch id; `engram_assess` may only cite refs of that batch, and `sufficient` is enforced by code — all three must hold (the model claims adequacy, at least one valid evidence ref, `nextStrategy=answer`) or the verdict is insufficient with the strategy rewritten back to retrieval. Verdicts and rejected refs go to the audit log, visible under the "retrieval" category in the curator log.
76
77
  - **Portable data**: `engram_export` exports Markdown / JSON files in one step, with a redacted-view variant (secondary scrubbing + 40-char preview truncation, share-safe). `engram_mirror` writes an Obsidian / Logseq-friendly mirror directory (one Markdown per memory with YAML frontmatter + `[[id]]` backlinks), turning the palace into a human-readable private knowledge base.
77
78
  - **AGI Architecture Exploration (dsh-market · AGI Architecture)**: listed in dsh-market's "AGI Architecture Exploration" category as a cognitive-science reframe of agent long-term memory — memory palace (imagery labels + room placards), corridor topology (force-directed graph), closure questions (the ingest prompt nudges the model to ask clarifying questions), consolidation merging (heuristic dedup + cosine similarity) — sitting alongside MemGPT / Letta in the "agent memory architecture" conversation.
78
79
 
79
- ## Tools (16, narrow parameters)
80
+ ## Tools (17, narrow parameters)
80
81
 
81
82
  | Tool | Purpose |
82
83
  |---|---|
83
84
  | `engram_save` | Save (automatic contradiction-candidate detection when embeddings are available); `items` array saves ≤10 in one call with shared scrubbing and in-batch dedup, one failure not blocking the rest (`count`/`items`/`failed` summary); `placard` attaches a marker (scored unique · distinctive · dated; low scores get a rewrite hint) |
84
- | `engram_search` | Hybrid semantic + keyword retrieval (hits reinforce confidence); `room` routes through the corridor — search inside one room only; the top 5 hits carry same-room neighbouring-slot cues |
85
+ | `engram_search` | Hybrid semantic + keyword retrieval (hits reinforce confidence); `room` routes through the corridor — search inside one room only; the top 5 hits carry same-room neighbouring-slot cues. Output rows carry `id=` and `ref=`; the tail carries the batch id |
86
+ | `engram_assess` | Evidence gate: before answering, judge whether the retrieved content suffices. Submit `batchId` + ≤8 `evidenceRefs` (only `ref=` values from that batch's output) + `missing` + `nextStrategy`; the code requires all three for `sufficient` (claimed adequate, at least one valid in-batch evidence ref, `nextStrategy=answer`), otherwise it rules insufficient and rewrites the strategy back to retrieval; refs from other batches are rejected and listed, and the verdict is written to the audit log |
85
87
  | `engram_timeline` | Timeline browsing: creation-time descending by default; `order: 'tour'` follows the fixed tour route's slot order instead (output carries palace coordinates, entries off-route sorted last) so the agent can re-walk the route |
86
88
  | `engram_update` | Correct an entry (supersedes chain); can re-attach a `placard` marker too |
87
89
  | `engram_forget` | Forget (soft-delete, restorable) |
@@ -108,7 +110,7 @@ Optional configuration (cordis.yml):
108
110
  dbDir: '~/.dsh/engram' # store and model-cache root directory
109
111
  injectProfile: true # inject the user memory profile at session start
110
112
  profileTopN: 8 # injection entry cap (1-64)
111
- injectTokenBudget: 1024 # injection token budget (128-8192, estimated ceil(len/4); over-budget entries degrade to index lines)
113
+ injectTokenBudget: 1024 # injection token budget (128-8192; CJK text counted at 1.5 tokens/char, other text at 4 chars/token; over-budget entries degrade to index lines)
112
114
  modelCacheDir: '~/.dsh/engram/models' # embedding model cache directory
113
115
  hfEndpoint: 'https://huggingface.co' # model download endpoint; set a mirror behind restricted networks
114
116
  ingest: 'off' # automatic capture: off | light (user messages only, ≤2/turn) | eager (assistant messages too, ≤5/turn)
@@ -163,6 +165,10 @@ Since v0.7.3 there is a dedicated **"History backfill" tab**: you choose every i
163
165
 
164
166
  Since v0.7.4 the panel gets a second pass of polish: Today is rearranged to "tour proposal on the left, room directory + due today + refurb list on the right", with roomier rows in the tour proposal and the room directory; the Room exhibition toolbar becomes one row of search + status + sort with the room filter as its own wrapping chip row; the curator log merges its counts and category filter into a single toolbar and marks every row with a category dot (writes / capture / retrieval / organize); history backfill's rules, estimate, and run blocks are separated by hairlines instead of a tinted estimate box; and section spacing now comes solely from the container gap, removing the asymmetry where a heading hugged the card above but sat far from the one below.
165
167
 
168
+ Since v0.7.5: two underlying fixes plus one new capability. ① Profile-injection token estimation is now CJK-aware (CJK at 1.5 tokens/char, other text at 4 chars/token); the old "length / 4" undercounted Chinese by more than 4×, so Chinese users' injections routinely exceeded `injectTokenBudget` by 17%–50% — the trailing `+N more` counter line now counts against the budget too, so the injection never exceeds its promise. ② A new **evidence gate** `engram_assess` (17 tools): a hit means "relevant", not "enough to answer". `engram_search` registers an in-process evidence batch per call (latest 20 per session, released when the session ends) and prints `ref=` per row plus the batch id; `engram_assess` may only cite refs of that batch and `sufficient` is enforced by code — the model's claim, at least one valid evidence ref, and `nextStrategy=answer` must all hold, otherwise the verdict is insufficient with the strategy rewritten back to retrieval; refs from other batches are rejected and listed, and the verdict goes to the audit log (visible under the "retrieval" category of the curator log). ③ The curator log gains the missing op labels and detail formatting (consolidation / review answer / slot assigned / slots backfilled / room opened) instead of raw op names and JSON.
169
+
170
+ Since v0.7.6 scopes stop being hardcoded and the project palace follows the workspace. ① **Per-candidate palace on capture**: extraction now emits `scope` — content tied to the current project/repo (tech choices, project conventions, architecture decisions) goes to that project store, while cross-project material (personal preferences, habits, the user's own history) goes to the private store. Live capture (previous turn, final turn on dispose, pending replay) and history backfill share the same rule, and backfill files each session into **its own cwd's** store. Idempotency keys (`ingest-done` / `ingest-pending`) stay in the private store, decoupled from where memories land, so live and backfill share one `(session, turn)` key set and never re-ingest across paths. ② **Project palace follows the workspace**: `GET /api/engram/workspaces` lists the selectable project palaces (host workspace registry → cwds seen in session headers → plugin process directory, deduped by store name, so several worktrees of one repo share a palace) and every project-scoped endpoint takes `?project=<dbName>` (unknown selectors answer 404); `engram_save/search` and friends resolve their project store from the current session's cwd. In the panel, a new row under the three scope segments shows the **active workspace chip plus a workspace picker** (default "follow the current workspace", pinnable to one workspace) and switching workspaces switches every view. ③ Without git, the project store name changes from "first 12 chars of cwd" (which collided for sibling directories) to a **full cwd sha256**, with the old store renamed on startup; and unloading the plugin now closes every store connection (no more locked `.db` files on Windows).
171
+
166
172
  ## Development
167
173
 
168
174
  ```sh
package/README.md CHANGED
@@ -59,29 +59,31 @@ dsh plugin --profile web add @kenz1117/dsh-engram
59
59
  ## 特性
60
60
 
61
61
  - **跨会话记忆**:会话开始注入用户画像摘要(条数 + token 预算双重上限,可配),Agent 天然"记得"你是谁、在做什么;工具检索跨会话召回历史事实。
62
- - **双层分库**:`user.db` 全局共享;`project-<hash>.db` 按 git origin 标识隔离(无 git 时回退工作目录编码,旧库自动迁移)——个人偏好跟人走,项目约定跟仓库走。
62
+ - **双层分库**:`user.db` 全局共享;`project-<hash>.db` 按 git origin 标识隔离(无 git 时回退 cwd 全量哈希,旧命名库自动迁移)——个人偏好跟人走,项目约定跟仓库走。**项目宫殿随工作区切换**:管理面板「项目」scope 默认跟随 GUI 当前选中的工作区(Header 显示工作区标题与路径),也可在下拉里固定到某个工作区;`engram_save/search` 等工具的 project 读写同样按当前会话 cwd 归属,与面板同一口径。
63
63
  - **混合检索**:FTS5(unicode61 + 中文 2-gram 预切词)与本地向量(`Xenova/bge-small-zh-v1.5`,512 维,q8)RRF 融合 + 关系边一跳扩展 + 新鲜度/命中次数乘性排序 boost;嵌入模型离线运行,下载失败自动降级纯关键词并显式标记。
64
64
  - **记忆宫殿信息架构**(v0.7.2+):四原则全部落进核心路径,而非展示层皮肤——**位置当索引**(写入按 kind 分房并钉「房间#桩位」坐标,房间容量 9,满员开新房,桩位只增不回收);**固定路线定顺序**(`tour_routes` append-only,`engram_tour mode=fixed` 按桩位顺序走全宫);**骨架长期复用**(同一 topic 永远落在同一房间同一序号,顺序提取而非重新检索);**标记独一无二**(门牌规则 0-1 评分:全库唯一 +0.4 / 带日期锚点 +0.3 / 同房前 6 字不重复 +0.3,低分进翻新清单;库内 active 记忆少于 8 条时不扫描,避免小库噪声)。
65
65
  - **走廊路由检索**:画像里附房间目录,`engram_search` 支持 `room` 参数——先决定进哪个房间,再在房内检索,而非一上来全库 RRF。检索命中 top5 附同房间相邻桩位 id 作为编码特异性线索。
66
66
  - **检索练习闭环**(间隔重复):`engram_review_queue` 只给宫殿坐标与门牌线索、**不给正文**,迫使模型先主动回忆;`engram_review` 揭示核对,`engram_report grade`(0-5)自评推进 SM-2 调度(1 → 6 → round(prev × ease) 天,失败重置,ease 下限 1.3)。进入复习调度的条目**不再参与自动衰减**——命运由回忆结果决定。会话开始注入会提示今日待回忆条数。
67
- - **历史会话回填**(v0.7.3+):把 dsh 已持久化的历史会话逐轮提炼进宫殿——按**每个会话自己的 cwd** 写进对应项目库(不串库),复用实时摄取的节流/脱敏/防回声,靠 (会话, 轮次) 幂等键支持中断续跑;回填条目不进今日复习队列(避免一次性回填淹没「今日待回忆」)。**导入规则由你选**(时间窗 / 单会话轮数 / 总轮数上限 / 辅助模型 / 是否含子代理·种子·无 cwd 会话),先估算(零成本、不调 LLM)再执行;设置页有独立的**「历史回填」tab**,可看进度(含**跳过原因分布**)与暂停续做。辅助调用默认**复用你当前在用的模型**(历史日志里记的是当年的 provider/model,在当前环境可能已不可用),也可在面板「辅助模型」下拉里从宿主已注册的 provider/model 中直接指定。
67
+ - **历史会话回填**(v0.7.3+):把 dsh 已持久化的历史会话逐轮提炼进宫殿——默认按**每个会话自己的 cwd** 写进对应项目库(不串库),同样逐条判作用域(跨项目通用的个人偏好落私人库),复用实时摄取的节流/脱敏/防回声,靠 (会话, 轮次) 幂等键支持中断续跑(键固定在私人库,与实时路径共用);回填条目不进今日复习队列(避免一次性回填淹没「今日待回忆」)。**导入规则由你选**(时间窗 / 单会话轮数 / 总轮数上限 / 辅助模型 / 是否含子代理·种子·无 cwd 会话),先估算(零成本、不调 LLM)再执行;设置页有独立的**「历史回填」tab**,可看进度(含**跳过原因分布**)与暂停续做。辅助调用默认**复用你当前在用的模型**(历史日志里记的是当年的 provider/model,在当前环境可能已不可用),也可在面板「辅助模型」下拉里从宿主已注册的 provider/model 中直接指定。
68
68
  - **知识飞轮**:摄取/保存 → 矛盾候选(写入时高相似近邻建 `contradicts` 边并报告,模型/用户裁决)→ 命中强化(confidence +0.05)→ 蒸馏(同主题簇合并为高层规律、supersedes 取代链、置信度继承)→ 衰减(低重要性且长期未访问归档,可恢复)。
69
- - **自动摄取**(`ingest` 配置开启时):新一轮第一步从会话日志提取上一轮的候选事实,会话结束时补摄取最后一轮(失败留 pending 键,下次会话自动补做,幂等不重复),低 confidence 写入并按嵌入去重——不说"记住"也能攒记忆。
69
+ - **自动摄取**(`ingest` 配置开启时):新一轮第一步从会话日志提取上一轮的候选事实,会话结束时补摄取最后一轮(失败留 pending 键,下次会话自动补做,幂等不重复),低 confidence 写入并按嵌入去重——不说"记住"也能攒记忆。**逐条判宫殿**:提炼时同步判定作用域,只跟当前项目/仓库有关的(技术选型、项目约定、架构决策)进当前会话 cwd 对应的项目库,跨项目通用的(个人偏好、习惯、本人经历)进私人库——偏好跟人走、约定跟仓库走。
70
70
  - **来源审计**:每条记忆记录来源会话、轮次与事件 seq,`engram_review` 完整回查来源链、取代链、矛盾与操作日志;全部写入/修改/遗忘/蒸馏/衰减入操作日志表。
71
- - **Web 管理面板**(v0.7.0+):设置页「记忆库」tab 分五个视图——今日(速览条:记忆 / 开放 / 清晰度 + 近 7 天计数 + 健康分环;下面是今日待回忆、待翻新、健康分构成、房间目录、入殿导航)、宫殿陈展(筛选 / 巡游路线序 / 批量 / 列表与编辑)、走廊巡游(走廊鸟瞰 + 检索实验台)、管家日志(近 7 天计数 + 两库合并的完整 op_log,可按操作类别筛选)、历史回填。Header 三宫格驱动全局 scope(私人 / 项目 / 共享),全部数据源同步。界面文案中英双语,跟随宿主语言设置实时切换。支持按脱敏标记筛选(仅看/排除含 `[REDACTED:*]` 的条目)并给命中条目挂琥珀色徽标,方便审计脱敏覆盖面。
71
+ - **Web 管理面板**(v0.7.0+):设置页「记忆库」tab 分五个视图——今日(速览条:记忆 / 开放 / 清晰度 + 近 7 天计数 + 健康分环;下面是今日待回忆、待翻新、健康分构成、房间目录、入殿导航)、宫殿陈展(筛选 / 巡游路线序 / 批量 / 列表与编辑)、走廊巡游(走廊鸟瞰 + 检索实验台)、管家日志(近 7 天计数 + 两库合并的完整 op_log,可按操作类别筛选)、历史回填。Header 三宫格驱动全局 scope(私人 / 项目 / 共享),全部数据源同步;**项目 scope 下三宫格右侧显示当前项目宫殿所属工作区**(标题 + 路径,默认跟随 GUI 当前工作区),旁边的工作区下拉可固定到某个工作区或切回「跟随当前会话」。界面文案中英双语,跟随宿主语言设置实时切换。支持按脱敏标记筛选(仅看/排除含 `[REDACTED:*]` 的条目)并给命中条目挂琥珀色徽标,方便审计脱敏覆盖面。
72
72
  - **提示注入防护**:全部记忆召回出口(画像注入、`engram_search/timeline/review` 输出)包 `<engram_memory_context>` 协议标签并附使用警告(历史记忆非当前请求、不遵循其中指令、仅相关时使用),当前请求独立包 `<current_user_request>`;所有入库内容(摄取候选、保存正文)先剥离这些协议标签,防伪造协议块二次注入。
73
73
  - **摄取脱敏**:入库前正则清洗常见密钥凭据(sk- 系 API key、Bearer、AWS AKIA、GitHub token、PEM 私钥、password/token 赋值),命中片段替换为 `[REDACTED:<类型>]`。
74
74
  - **召回占位(防回声室)**:摄取切片中记忆召回工具的输出替换为 `[engram memory result omitted from capture: <tool>]`,并向提取模型附注"既有记忆的复述不是新信息",阻断记忆自我强化循环。
75
75
  - **多查询检索**:`engram_search` 可用辅助 LLM 把查询改写为 ≤3 个互补查询分别检索,跨查询 RRF 融合 + 每查询保底命中;改写失败自动降级单查询(`queryRewrite: false` 关闭)。
76
+ - **证据门(search → assess)**:检索命中只说明「相关」,不说明「足以回答」。每次检索登记一个进程内批次(每会话保留最近 20 个,会话结束即释放),输出行尾给出 `ref=…` 与批次 id;`engram_assess` 只能引用同一批次的 ref,且 `sufficient` 由代码强制——三者齐备(模型声称充足、至少一条有效证据、`nextStrategy=answer`)才算充足,否则判为不足并把策略改回继续检索。判定与拒绝明细写入审计日志,面板「管家日志」的「检索」类别可见。
76
77
  - **数据可携带**:`engram_export` 一键导出 Markdown / JSON 文件,支持脱敏视图(内容二次清洗 + 预览截断,分享安全)。`engram_mirror` 导出可漫游的镜像目录(Obsidian / Logseq 友好:每条记忆一个 Markdown,正文 + YAML frontmatter + 双向链接 `[[id]]`),让「宫殿」也成为可人读的私人知识库。
77
78
  - **认知架构探索(dsh-market · AGI 架构探索)**:本仓库是 dsh-market「AGI 架构探索」类目下,对 agent 长期记忆的认知科学方法论重构——记忆宫殿(意象标签 + 房间铭牌)、走廊拓扑(力导向图)、闭环提问(摄入时让模型主动追问用户细节)、巩固合并(启发式去重 + 余弦相似度),与 MemGPT/Letta 同层「agent 记忆架构」叙事。
78
79
 
79
- ## 工具(16 个,窄参数)
80
+ ## 工具(17 个,窄参数)
80
81
 
81
82
  | 工具 | 作用 |
82
83
  |---|---|
83
- | `engram_save` | 保存(嵌入可用时自动做矛盾候选检测);支持 `items` 数组单次批量保存 ≤10 条,统一清洗/批量内去重,单条失败不影响其余(`count`/`items`/`failed` 汇总返回);`placard` 挂门牌(按唯一·差异化·带日期评分,低分附改写建议) |
84
- | `engram_search` | 语义 + 关键词混合检索(命中强化置信度);`room` 参数做走廊路由——只在指定房间内检索;命中 top5 附同房相邻桩位线索 |
84
+ | `engram_save` | 保存(嵌入可用时自动做矛盾候选检测);支持 `items` 数组单次批量保存 ≤10 条,统一清洗/批量内去重,单条失败不影响其余(`count`/`items`/`failed` 汇总返回);`placard` 挂门牌(按唯一·差异化·带日期评分,低分附改写建议)。`scope=project` 落当前会话 cwd 对应的项目宫殿 |
85
+ | `engram_search` | 语义 + 关键词混合检索(命中强化置信度);`room` 参数做走廊路由——只在指定房间内检索;命中 top5 附同房相邻桩位线索。输出行尾给 `id=` 与 `ref=`,末尾给批次 id。`scope=project` 查当前会话 cwd 对应的项目宫殿 |
86
+ | `engram_assess` | 证据门:作答前判定「检索到的内容是否足以回答」。提交 `batchId` + ≤8 条 `evidenceRefs`(只能取该批次输出里的 `ref=`)+ `missing` + `nextStrategy`;代码强制 `sufficient` 需同时满足「声称充足」「至少一条属于本批次的有效证据」「nextStrategy=answer」,否则判为不足并把策略改回继续检索;非本批次的 ref 会被拒绝并列出,判定写入审计日志 |
85
87
  | `engram_timeline` | 时间线浏览:默认按创建时间倒序;`order: 'tour'` 改按固定巡游路线桩位顺序(输出附宫殿坐标,未上路线者排末尾),让 agent 也能沿固定路线复述 |
86
88
  | `engram_update` | 修正(supersedes 取代链);可同时改挂 `placard` 门牌 |
87
89
  | `engram_forget` | 遗忘(软删可恢复) |
@@ -108,7 +110,7 @@ dsh plugin --profile web add @kenz1117/dsh-engram
108
110
  dbDir: '~/.dsh/engram' # 分库与模型缓存根目录
109
111
  injectProfile: true # 会话开始注入用户画像摘要
110
112
  profileTopN: 8 # 注入条数上限(1-64)
111
- injectTokenBudget: 1024 # 注入 token 预算(128-8192,估算 ceil(len/4),超预算条目降级为索引行)
113
+ injectTokenBudget: 1024 # 注入 token 预算(128-8192,中文按 1.5 token/字、其余按 4 字符/token 估算,超预算条目降级为索引行)
112
114
  modelCacheDir: '~/.dsh/engram/models' # 嵌入模型缓存目录
113
115
  hfEndpoint: 'https://huggingface.co' # 模型下载端点,网络受限可配镜像
114
116
  ingest: 'off' # 自动摄取:off | light(仅用户消息,每轮≤2条)| eager(含助手消息,每轮≤5条)
@@ -146,9 +148,10 @@ dsh plugin --profile web add @kenz1117/dsh-engram
146
148
  └─ 设置页「记忆库」tab ◀────────────────── ─── 回环 API /api/engram/*(写操作校验回环 Origin)
147
149
  ```
148
150
 
149
- - **双层分库**:`user.db` 全局共享;`project-<hash>.db` 按 git origin URL 归一化哈希命名(`git@github.com:a/b.git` 与 `https://github.com/a/b` 同库;worktree 沿指针解析到主仓库 origin);无 git 或无 origin 时回退工作目录编码命名,旧的 cwd 命名库在启动时自动 rename 迁移(新旧并存则不动并告警)。
151
+ - **双层分库**:`user.db` 全局共享;`project-<hash>.db` 按 git origin URL 归一化哈希命名(`git@github.com:a/b.git` 与 `https://github.com/a/b` 同库;worktree 沿指针解析到主仓库 origin);无 git 或无 origin 时按 **cwd 全量 sha256** 命名(v0.7.6 起;旧的「cwd 前 12 字符」截断命名会同前缀撞库,启动时自动 rename 迁移,新旧并存则不动并告警)。
152
+ - **项目宫殿路由**:`GET /api/engram/workspaces` 给出可选项目宫殿清单(来源 = 宿主 `workspaceRegistry` 工作区 → 会话 header 里出现过的 cwd → 插件进程目录兜底,按分库名去重;同仓库的多个 worktree 共用一个宫殿);所有项目 scope 的接口与工具读写接受选择器(HTTP `?project=<dbName>` / POST body `project`,工具用当前会话 `session.header.cwd`),未知选择器 HTTP 回 404、工具回退进程目录。面板「项目」scope 默认跟随 GUI 当前工作区并可固定到某个工作区。
150
153
  - **宫殿结构(目录即房间,路径即路线)**:**主厅** = 每轮注入的常驻核心画像(少而稳,每次都在场);**走廊** = 画像里附的房间目录 + `engram_search room` 参数(先定房间,再检索);**房间** = 按 kind 分房(事实厅 / 偏好阁 / 决策堂 / 往事廊 / 技法坊),容量 9,满员开「房名-2」;**门牌** = 每条记忆的 `placard` 铭牌,受唯一·差异化·带日期纪律评分。存量库首次打开时自动补排桩(幂等,开新房会告警提醒人工命名)。
151
- - **自动摄取**(`ingest` 开启时):新一轮第一步从会话日志提取上一轮的候选事实;会话结束(session/disposed)补摄取最后一轮,5 秒超时,失败/超时把 pending 键写入操作日志,下次会话首步自动重放补做;已摄取的 (会话, 轮次) 幂等去重。读取源是会话日志;辅助调用的请求审计走插件自身操作日志,不向会话日志 append 未知事件。候选以低 confidence 写入并按嵌入去重。
154
+ - **自动摄取**(`ingest` 开启时):新一轮第一步从会话日志提取上一轮的候选事实;会话结束(session/disposed)补摄取最后一轮,5 秒超时,失败/超时把 pending 键写入操作日志,下次会话首步自动重放补做;已摄取的 (会话, 轮次) 幂等去重(键固定在私人库,实时与历史回填共用一份,跨路径不重复)。读取源是会话日志;辅助调用的请求审计走插件自身操作日志,不向会话日志 append 未知事件。候选以低 confidence 写入并按**目标分库内**的嵌入近邻去重。**作用域逐条判**:提炼输出带 `scope`,`project` 落当前会话 cwd 的项目库、`user` 落私人库;无 cwd 的会话只能进私人库(此类会话不做逐条判定,避免标记与实际分库不符)。历史回填沿用同一判定,默认按各会话自己的 cwd 落项目库。
152
155
  - **来源链**:每条记忆记录来源会话、轮次与事件 seq,`engram_review` 可完整回查;操作日志表记录全部写入/修改/遗忘/蒸馏/衰减。
153
156
  - **嵌入离线**:模型首次使用需联网下载(q8 约 50MB,端点可配镜像),此后完全离线;失败时插件照常工作,检索降级纯关键词并显式标记。
154
157
  - **界面本地化**:client 半经宿主 locale 服务注册 zh/en 词典,跟随宿主语言设置实时切换;状态/种类等数据枚举仅在显示层映射,存储值保持英文。
@@ -163,6 +166,10 @@ v0.7.3 起新增独立的**「历史回填」tab**:导入规则全部由你选
163
166
 
164
167
  v0.7.4 起继续打磨面板细节:今日视图重排为「左入殿导航 · 右房间目录、今日待回忆、翻新清单」,入殿导航与房间目录加大行间距;宫殿陈展工具栏改「搜索 + 状态 + 排序」一行、房间筛选独立成可换行 chips;管家日志把计数与类别筛选合成一条工具条,并给每行加类别色点(落成 / 发掘 / 检索 / 整理);历史回填的规则、估算、执行三段改用分隔线切块;区域间距统一由容器间距给出,消除「标题贴住上方卡片、下方却过松」的不对称。
165
168
 
169
+ v0.7.5 起是两处底层修正加一层新能力。① 画像注入的 token 估算改为 CJK 感知(中文按 1.5 token/字、其余按 4 字符/token):此前按长度除以 4 会把中文低估四倍以上,中文用户的实际注入长期超出 `injectTokenBudget` 约 17%–50%;末尾 `+N more` 计数行也纳入预算,注入总量不再超承诺。② 新增**证据门** `engram_assess`(工具 17 个):检索命中只说明「相关」,不说明「足以回答」;`engram_search` 每次登记一个进程内证据批次(每会话保留最近 20 个,会话结束即释放),输出每行带 `ref=`、末尾带批次 id;`engram_assess` 只能引用同一批次的 ref,且 `sufficient` 由代码强制——声称充足、至少一条有效证据、`nextStrategy=answer` 三者齐备才算充足,否则判为不足并把策略改回继续检索;不属于该批次的 ref 会被拒绝并列出,判定写入审计日志(管家日志「检索」类别可见)。③ 管家日志补齐 op 词典与明细格式化(闭馆整理 / 复习答题 / 排桩 / 批量排桩 / 开新房),不再显示英文原名与原始 JSON。
170
+
171
+ v0.7.6 起把「作用域」从写死改成逐条判、并让项目宫殿跟着工作区走。① **摄取逐条判宫殿**:提炼时同步判定 `scope`——只跟当前项目/仓库有关的(技术选型、项目约定、架构决策)进该项目库,跨项目通用的(个人偏好、习惯、本人经历)进私人库;实时摄取(上一轮 / 会话结束末轮 / 待补做重放)与历史回填同一口径,回填按**每条会话自己的 cwd** 落库。幂等键(`ingest-done` / `ingest-pending`)固定在私人库,与写入落点解耦,所以实时与回填共用一份 (会话, 轮次) 键、跨路径不重复摄取。② **项目宫殿随工作区切换**:新增 `GET /api/engram/workspaces`(来源 = 宿主工作区注册表 → 会话 header 里出现过的 cwd → 插件进程目录兜底,按分库名去重,同仓库多 worktree 共用一个宫殿),所有项目 scope 的接口接受 `?project=<dbName>` 选择器,未知选择器回 404;`engram_save/search` 等工具的 project 读写改按当前会话 cwd 归属。面板的「项目」作用域在作用域三宫格下方新增一行:**当前工作区 chip + 工作区下拉**(默认「跟随当前工作区」,可临时固定到某个工作区),切工作区即整体切换。③ 无 git 时的项目分库名从「cwd 前 12 字符编码」(同前缀目录会撞库)改为 **cwd 全量 sha256**,旧命名库启动时自动 rename 迁移;插件卸载时关闭所有分库连接(Windows 上不再锁住 .db)。
172
+
166
173
  ## 开发
167
174
 
168
175
  ```sh
@@ -182,7 +189,7 @@ pnpm bundle
182
189
 
183
190
  #### Token effect
184
191
 
185
- 画像注入为条件性固定成本(受条数上限与 token 预算双重约束);工具 schema 为常驻成本(16 个窄参数工具)。
192
+ 画像注入为条件性固定成本(受条数上限与 token 预算双重约束);工具 schema 为常驻成本(17 个窄参数工具)。
186
193
 
187
194
  #### KV Cache effect
188
195