@yandy0725/pi-memory 2.1.1 → 2.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +26 -10
- package/README.zh.md +26 -10
- package/index.ts +32 -5
- package/package.json +1 -1
- package/src/config.ts +17 -11
- package/src/entry-index.ts +2 -2
- package/src/inject.ts +136 -7
package/README.md
CHANGED
|
@@ -6,6 +6,14 @@ Aligned with Claude Code's auto memory mechanism: **one memory = one file**, a `
|
|
|
6
6
|
|
|
7
7
|
> ## ⚠️ Breaking changes
|
|
8
8
|
>
|
|
9
|
+
> **In 2.3.0:**
|
|
10
|
+
>
|
|
11
|
+
> - **The injected index window is now the newest 50 lines / 16 KiB.** `memIndexInjectMaxLines` 200 → 50 and `memIndexInjectMaxBytes` 25600 → 16384. The write capacity is unchanged (200 lines / 25600 bytes), so memories older than the newest 50 index lines no longer reach the system prompt — they stay reachable through auto-surfacing and the `memory` tool. Configs that already set `memIndexInjectMax*` are unaffected.
|
|
12
|
+
>
|
|
13
|
+
> **In 2.2.0:**
|
|
14
|
+
>
|
|
15
|
+
> - **`extractMemories.enabled` now defaults to `false`.** Per-turn extraction is opt-in: once enabled, every turn ends with a headless model call. Configs that already set `"extractMemories": { "enabled": true }` are unaffected.
|
|
16
|
+
>
|
|
9
17
|
> **In 2.1.0:**
|
|
10
18
|
>
|
|
11
19
|
> - **Models must be configured explicitly.** There is no shipped default and no parent-model fallback: `defaults.model` (or a per-task `model`) must exist and be resolvable, or `session_start` reports a config error and initialises **nothing**. See [Model configuration](#model-configuration).
|
|
@@ -26,7 +34,7 @@ Aligned with Claude Code's auto memory mechanism: **one memory = one file**, a `
|
|
|
26
34
|
- **`memory_index` prompt section, frozen per session** ⭐ — the index goes into `event.systemPromptOptions.sections["memory_index"]` and its value **does not change for the rest of the session**; only compaction re-reads it from disk. Because pi diffs sections and appends nothing when they are unchanged, the system prompt stays byte-identical turn after turn and the provider's prefix cache keeps hitting. `resume` / `fork` / `reload` replay the **recorded** value from the transcript instead of reading disk, so restoring a session does not rewrite its head.
|
|
27
35
|
- **Injection sanitising** — everything injected (index lines, surfaced entry bodies and names) has invisible/bidi characters stripped and `<` `>` escaped, so a memory can never forge `</relevant_memories>`, `<system>`, `<project_instructions>`, `<active_agent …>` or `<memory_index>`. Sanitising happens **at injection time only**: your files on disk are never rewritten (they stay readable and hand-editable).
|
|
28
36
|
- **Auto-surfacing** ⭐ — on every user turn a lightweight side query selects up to `maxFiles` **entries** (selected from `description` alone) and injects their bodies inside `<relevant_memories>`. Already-injected files are deduplicated per session; the manifest is served from an in-process `mtime` cache, so a turn costs one `readdir` plus one `stat` per file. Disabled inside subagents.
|
|
29
|
-
- **Extract memories** ⭐ — after each run an async headless agent receives a **structured rendering of the whole conversation** (every user message in full, assistant text and tool calls, tool results with error flags), not just two messages. It writes through the same `memory` primitives, under a whole-round logical lock it never waits for: if a dream is running, that turn is simply skipped.
|
|
37
|
+
- **Extract memories** ⭐ — after each run an async headless agent receives a **structured rendering of the whole conversation** (every user message in full, assistant text and tool calls, tool results with error flags), not just two messages. It writes through the same `memory` primitives, under a whole-round logical lock it never waits for: if a dream is running, that turn is simply skipped. This feature is **off by default** — set `extractMemories.enabled: true` to turn it on.
|
|
30
38
|
- **`/dream`** — a headless consolidation agent (Orient → Gather Signal → Consolidate → Prune & Index) that merges duplicates, resolves contradictions, renames entries and rebuilds the index. It has **no raw file access**: it only gets the seven `memory` actions, holds the logical lock for the whole round, and snapshots the entire directory on entry.
|
|
31
39
|
- **Dream nudge** — after N sessions or N hours a notification suggests `/dream`.
|
|
32
40
|
- **`/memory`** — full status (switch, directory, index capacity, entry count, last dream, lock state including the holder), plus `unlock`.
|
|
@@ -93,7 +101,7 @@ File names are derived from `name` (unsafe characters replaced, 100-byte cap, `-
|
|
|
93
101
|
- [Test command](Test-command.md) — run npm test, not npm run test
|
|
94
102
|
```
|
|
95
103
|
|
|
96
|
-
Writes are **surgical**: only the target line changes, hand-written headings, groups and comments are preserved byte-for-byte, and line order is stable. The one exception is line endings: CRLF (or lone CR) is normalised to LF before parsing, so the first write to a CRLF file rewrites it with LF.
|
|
104
|
+
Writes are **surgical**: only the target line changes, hand-written headings, groups and comments are preserved byte-for-byte, and line order is stable. The one exception is line endings: CRLF (or lone CR) is normalised to LF before parsing, so the first write to a CRLF file rewrites it with LF. The injected index uses the same normalisation — a CRLF file never sends `\r` into the system prompt.
|
|
97
105
|
|
|
98
106
|
### Memory types
|
|
99
107
|
|
|
@@ -108,7 +116,9 @@ Writes are **surgical**: only the target line changes, hand-written headings, gr
|
|
|
108
116
|
|
|
109
117
|
The index holds at most `memIndexMaxLines` (200) non-empty lines and `memIndexMaxBytes` (25600) bytes. Those 200 lines are **index lines, not memories**: `rebuildIndex` guarantees at least one header line — an existing hand-written header is kept verbatim (trailing blank lines before the first entry are dropped), otherwise it writes `# Memory Index` — and hand-written headings, groups and comments count too. A rebuilt index therefore holds at most about **199 memories per project directory** (fewer if you keep hand-written headings). Exceeding the limit does **not** fail the write: the write succeeds and the tool returns an actionable warning telling the model to merge or drop entries (everything past the limit is invisible on the next load).
|
|
110
118
|
|
|
111
|
-
|
|
119
|
+
**What reaches the model is a separate, smaller window:** `memIndexInjectMaxLines` / `memIndexInjectMaxBytes` (default 50 lines / 16384 bytes). The window is taken from the **newest** end of the index, i.e. from the **bottom of the file**: `memory(action="add")` appends, and `/dream`'s `rebuild_index` re-sorts by `modified`. Two exceptions matter — `memory(action="replace")` rewrites its line **in place**, so an edited older memory keeps its position (and can stay outside the window) until the next `/dream` re-sort; and hand-written headings, groups or notes at the top of the file are positional rather than chronological, so a curated top block is the first thing the window drops. A full index therefore injects the **50 newest memories** (the `# Memory Index` header and its blank line fall outside the window once it truncates); an index of 48 memories or fewer is injected whole. Omitted memories are **not lost**: auto-surfacing, `memory(action="search")` and `/dream` consolidation still see them. `/memory` prints both budgets apart: `Index:` is the write capacity (disk truth), `Inject:` is the window that goes into the system prompt.
|
|
120
|
+
|
|
121
|
+
This is why `/dream` is no longer optional housekeeping — it is **capacity management**. Two thresholds matter: prompt visibility ends at the injection window, so consolidate **before you pass ~48 memories** if you want every entry in the system prompt, while 199 is only the hard write limit past which the index itself has to shrink.
|
|
112
122
|
|
|
113
123
|
## Configuration
|
|
114
124
|
|
|
@@ -120,8 +130,8 @@ Create `memory.json` in the agent directory (`~/.pi/agent/memory.json`) or the p
|
|
|
120
130
|
"memoryDir": "~/.pi/memory",
|
|
121
131
|
"memIndexMaxLines": 200,
|
|
122
132
|
"memIndexMaxBytes": 25600,
|
|
123
|
-
"memIndexInjectMaxLines":
|
|
124
|
-
"memIndexInjectMaxBytes":
|
|
133
|
+
"memIndexInjectMaxLines": 50,
|
|
134
|
+
"memIndexInjectMaxBytes": 16384,
|
|
125
135
|
"lock": { "timeoutMs": 5000, "snapshotKeep": 5 },
|
|
126
136
|
"defaults": { "model": "provider/model-id", "sessionPersistence": { "enabled": false } },
|
|
127
137
|
"dream": { "nudgeAfterSessions": 5, "nudgeAfterHours": 24, "thinkLevel": "high" },
|
|
@@ -134,7 +144,7 @@ Create `memory.json` in the agent directory (`~/.pi/agent/memory.json`) or the p
|
|
|
134
144
|
"maxInjectionBytes": 10240
|
|
135
145
|
},
|
|
136
146
|
"extractMemories": {
|
|
137
|
-
"enabled":
|
|
147
|
+
"enabled": false,
|
|
138
148
|
"thinkLevel": "high",
|
|
139
149
|
"maxContextTokens": 2000,
|
|
140
150
|
"maxToolResultChars": 500,
|
|
@@ -151,8 +161,8 @@ Create `memory.json` in the agent directory (`~/.pi/agent/memory.json`) or the p
|
|
|
151
161
|
| `memoryDir` | `~/.pi/memory` | Root directory for all memory data |
|
|
152
162
|
| `memIndexMaxLines` | `200` | Write capacity: max non-empty lines in `MEMORY.md` (the `# Memory Index` header and hand-written headings count too, so this is not exactly the memory count) |
|
|
153
163
|
| `memIndexMaxBytes` | `25600` | Write capacity: max bytes of `MEMORY.md` |
|
|
154
|
-
| `memIndexInjectMaxLines` | `
|
|
155
|
-
| `memIndexInjectMaxBytes` | `
|
|
164
|
+
| `memIndexInjectMaxLines` | `50` | Injection window: max lines of the index put into the `memory_index` section. The window keeps the **newest** lines and drops the **oldest** ones — the index is pure chronological order, so a smaller window never hides the memory you just wrote. **`0` (either key) injects no index at all** — the `memory_index` section stays empty |
|
|
165
|
+
| `memIndexInjectMaxBytes` | `16384` | Injection window: max bytes of the index section (older lines are dropped first, with a `[truncated: …]` marker at the **top**) |
|
|
156
166
|
| `lock.timeoutMs` | `5000` | How long a write waits for the logical lock (single primitive) or the cross-process `.lock`. Also the upper bound `session_shutdown` waits for in-flight writes |
|
|
157
167
|
| `lock.snapshotKeep` | `5` | Rollback points kept in `.backups/` (directories named `migrate-*` — whole-directory snapshots from an earlier 1.x migration, whose `originals/` subdirectory holds the pre-2.0 topic files — are never pruned) |
|
|
158
168
|
| `defaults.model` | `— (required)` | Shared model for dream / extract / side query. **No default**: every task that will run must resolve a model, otherwise `session_start` fails (see [Model configuration](#model-configuration)). A per-task `model` overrides it |
|
|
@@ -172,7 +182,7 @@ Create `memory.json` in the agent directory (`~/.pi/agent/memory.json`) or the p
|
|
|
172
182
|
| `autoSurfacing.maxEntryBytes` | `3072` | ⭐ Max bytes of a single injected entry body (truncated). Replaces 1.x's `maxTopicBytes`, which is ignored |
|
|
173
183
|
| `autoSurfacing.maxInjectionBytes` | `10240` | ⭐ Max total bytes of injected content per turn |
|
|
174
184
|
| `autoSurfacing.sessionPersistence.*` | inherits `defaults` | Persist side-query sessions to disk |
|
|
175
|
-
| `extractMemories.enabled` | `
|
|
185
|
+
| `extractMemories.enabled` | `false` | ⭐ Enable per-turn memory extraction. **Off by default** (opt-in): once enabled, every turn ends with a headless model call |
|
|
176
186
|
| `extractMemories.model` | — | ⭐ Model for the extraction agent. Falls back to `defaults.model`; required unless `defaults.model` is set (must be resolvable, no parent-model fallback) |
|
|
177
187
|
| `extractMemories.thinkLevel` | `"high"` | ⭐ Thinking effort for extraction |
|
|
178
188
|
| `extractMemories.maxContextTokens` | `2000` | ⭐ Budget for the rendered conversation (`× 4` characters; the middle is trimmed first, head and tail are kept, user messages are dropped last) |
|
|
@@ -207,7 +217,7 @@ Fix `memory.json` and restart the session — the config is read once at session
|
|
|
207
217
|
|---|---|
|
|
208
218
|
| `session_start` | Load config → validate the required models (a failure means nothing is initialised) → resolve the memory directory → **pick the index value and freeze it** (disk for `startup`/`new`; the recorded transcript value for `resume`/`fork`/`reload`) → register the `memory` tool (once, five actions) → rebuild the manifest cache → dream nudge |
|
|
209
219
|
| `before_agent_start` | Write the frozen value into `sections["memory_index"]` (**unconditionally, every turn**), then auto-surfacing (main session, not a subagent) |
|
|
210
|
-
| `agent_end` |
|
|
220
|
+
| `agent_end` | With `extractMemories.enabled`, fire the async extractor; notify `Extracted N memories.` when it wrote something, or `Extract failed: …` once per session |
|
|
211
221
|
| `session_compact` | Clear the injected-file set **and re-read the index from disk** — the only in-session refresh point |
|
|
212
222
|
| `session_shutdown` | Wait for in-flight writes to finish (bounded by `lock.timeoutMs`) so a quit does not leave a stale `.lock` |
|
|
213
223
|
|
|
@@ -228,6 +238,8 @@ If the host pi is older than the sections API, pi-memory falls back to appending
|
|
|
228
238
|
|
|
229
239
|
### Extract memories
|
|
230
240
|
|
|
241
|
+
> **Off by default.** Set `"extractMemories": { "enabled": true }` in `memory.json` first (config is read once at `session_start`, so a restart is needed), and make sure `extractMemories.model` or `defaults.model` resolves.
|
|
242
|
+
|
|
231
243
|
The extractor receives a structured rendering instead of a lossy two-message summary:
|
|
232
244
|
|
|
233
245
|
```
|
|
@@ -303,12 +315,16 @@ Status output:
|
|
|
303
315
|
Memory: enabled
|
|
304
316
|
Dir: /home/you/.pi/memory/git/github.com__owner__repo
|
|
305
317
|
Index: 38/200 lines, 2841/25600 bytes, 1 unrecognized lines
|
|
318
|
+
Inject: 39/50 lines, 2841/16384 bytes
|
|
306
319
|
Entries: 37
|
|
320
|
+
Modules: dream=on(provider/model-a) extractMemories=off autoSurfacing=on(provider/model-b)
|
|
307
321
|
Last dream: 2026-10-01T22:10:04.882Z
|
|
308
322
|
Lock: free
|
|
309
323
|
```
|
|
310
324
|
|
|
311
325
|
- `Index` uses the **write** capacity (`memIndexMax*`) and reports how many non-empty lines could not be parsed as index lines (the `# Memory Index` header and hand-written headings count). CRLF (or lone CR) line endings are normalised to LF before parsing, and the next write emits LF too, so a `MEMORY.md` re-saved by a Windows editor does **not** raise this count.
|
|
326
|
+
- `Inject` uses the **injection** window (`memIndexInjectMax*`) and counts the window's lines and bytes — the index text that goes into the `memory_index` section, taken from the **newest** end (the truncation marker itself is not counted). It is computed by the same window code that produces the injected value, so the two cannot drift. Note the two lines count different things: `Index` counts **non-empty** lines, `Inject` counts **every** line of the window, so on a canonical index (LF endings, trailing newline, one blank line after the header) `Inject` reports one more line than `Index` and the same byte count. The value in the system prompt is **frozen for the session** (see [Why the index is frozen](#why-the-index-is-frozen)): a memory written after `session_start` appears in `Index` immediately but in `Inject` only after compaction or in the next session.
|
|
327
|
+
- `Modules` reports the activation state of the three model-driven features as `on(<effective model>)` / `off`. The effective model is the task's own `model`, otherwise `defaults.model`. `dream` has no switch of its own — it is on whenever the memory system is enabled.
|
|
312
328
|
- `Lock` is `free`, `held by <op> (pid N on <hostname>, started <ISO>)`, or `unreadable — run /memory unlock`. `/memory unlock` shows the same holder line in its confirmation prompt.
|
|
313
329
|
- In a session started with `enabled: false`, nothing is initialized at boot: `/memory` reports `Memory: disabled` plus `Dir: not initialized — set "enabled": true in memory.json and restart`, there is no way to enable it mid-session, and `/memory unlock` still works without a store.
|
|
314
330
|
- If a required model is missing or cannot be resolved, nothing is initialized and `/memory` reports `Memory: misconfigured` and `Dir: not initialized`, followed by one `- <error>` line per problem. The same errors are shown as an error notification at session start.
|
package/README.zh.md
CHANGED
|
@@ -6,6 +6,14 @@ pi coding agent 的文件系统持久记忆层。把项目知识(事实、偏
|
|
|
6
6
|
|
|
7
7
|
> ## ⚠️ 破坏性变更
|
|
8
8
|
>
|
|
9
|
+
> **2.3.0:**
|
|
10
|
+
>
|
|
11
|
+
> - **注入的索引窗口改为「最新的 50 行 / 16 KiB」。** `memIndexInjectMaxLines` 200 → 50、`memIndexInjectMaxBytes` 25600 → 16384。写入口径不变(200 行 / 25600 字节),因此比「最新 50 行」更旧的记忆不再进 system prompt —— 它们仍可由 auto-surfacing 与 `memory` 工具检索。已经显式配置 `memIndexInjectMax*` 的用户不受影响。
|
|
12
|
+
>
|
|
13
|
+
> **2.2.0:**
|
|
14
|
+
>
|
|
15
|
+
> - **`extractMemories.enabled` 默认改为 `false`。** 每轮自动提取现在是 opt-in:开启后每轮结束都会跑一次 headless 模型调用。已经显式写了 `"extractMemories": { "enabled": true }` 的配置不受影响。
|
|
16
|
+
>
|
|
9
17
|
> **2.1.0:**
|
|
10
18
|
>
|
|
11
19
|
> - **模型必须显式配置。** 没有内置默认值,也没有父会话模型回退:`defaults.model`(或 per-task `model`)必须存在且可解析,否则 `session_start` 会报配置错误并且**什么都不初始化**。详见[模型配置](#模型配置)。
|
|
@@ -26,7 +34,7 @@ pi coding agent 的文件系统持久记忆层。把项目知识(事实、偏
|
|
|
26
34
|
- **`memory_index` section,会话内冻结** ⭐ —— 索引写进 `event.systemPromptOptions.sections["memory_index"]`,其值在**整个会话内不再变化**,只有 compaction 会从磁盘重读。pi 对 sections 做 diff,值没变就一条消息都不追加,于是 system prompt 逐轮逐字节相同,provider 的 prefix cache 一直命中。`resume` / `fork` / `reload` 用 transcript 里的**录制值**重放,而不是读磁盘,因此恢复会话不会改写它的头部。
|
|
27
35
|
- **注入净化** —— 所有注入内容(索引行、浮现的 entry 正文与 name)都会剥离不可见/bidi 字符并转义 `<` `>`,因此记忆无法伪造 `</relevant_memories>`、`<system>`、`<project_instructions>`、`<active_agent …>` 或 `<memory_index>`。净化**只发生在注入时**:磁盘上的文件永远不会被改写(保持可读、可手工编辑)。
|
|
28
36
|
- **自动浮现(auto-surfacing)** ⭐ —— 每个用户回合由一次轻量侧查询挑出至多 `maxFiles` 条 **entry**(只看 `description`),把正文注入 `<relevant_memories>`。同一会话内按文件名去重;清单来自进程内的 `mtime` 缓存,每回合只付一次 `readdir` + 每文件一次 `stat`。子 agent 中不启用。
|
|
29
|
-
- **自动提取(extract memories)** ⭐ —— 每轮结束后一个异步 headless agent 拿到的是**整轮对话的结构化渲染**(user 消息全文、assistant 文本与 tool_call、tool_result 及其错误标记),而不是两条消息。它经同一套 `memory` 原语写入,并且**从不排队等锁**:dream
|
|
37
|
+
- **自动提取(extract memories)** ⭐ —— 每轮结束后一个异步 headless agent 拿到的是**整轮对话的结构化渲染**(user 消息全文、assistant 文本与 tool_call、tool_result 及其错误标记),而不是两条消息。它经同一套 `memory` 原语写入,并且**从不排队等锁**:dream 正在整轮持锁时,本回合直接跳过。该功能**默认关闭**,需要显式设 `extractMemories.enabled: true`。
|
|
30
38
|
- **`/dream`** —— headless 整理 agent(Orient → Gather Signal → Consolidate → Prune & Index),合并重复、消解矛盾、改名、重建索引。它**没有裸文件权限**:只有七个 `memory` action,整轮持有逻辑锁,进入时先对整个目录拍一次快照。
|
|
31
39
|
- **Dream 提醒** —— 距上次 dream 超过 N 个会话或 N 小时后提示 `/dream`。
|
|
32
40
|
- **`/memory`** —— 完整状态(开关、目录、索引容量、entry 数、上次 dream、锁状态含持有者),以及 `unlock`。
|
|
@@ -108,7 +116,9 @@ staging 的 SSH 用 2222 端口,密钥在 ~/.ssh/staging。
|
|
|
108
116
|
|
|
109
117
|
索引上限是 `memIndexMaxLines`(200)个非空行与 `memIndexMaxBytes`(25600)字节。这 200 行是**索引行,不是记忆条数**:`rebuildIndex` 至少保证一行头部(已有手写头部时原样保留 —— 首个条目之前的末尾空行会被去掉;否则写 `# Memory Index`),手写的标题、分组、注释同样占额度。因此重建后的索引最多约 **199 条记忆**(每个项目目录;若保留手写标题则更少)。超限时写入**不会失败**:写入照样成功,工具把一条可操作的警告回给模型,让它去合并或删除条目(超出上限的部分下次加载时不可见)。
|
|
110
118
|
|
|
111
|
-
|
|
119
|
+
**真正进模型的是另一个更小的窗口**:`memIndexInjectMaxLines` / `memIndexInjectMaxBytes`(默认 50 行 / 16384 字节)。窗口取索引的**最新**一端 —— 即**文件底部**:`memory(action="add")` 追加到末尾、`/dream` 的 `rebuild_index` 按 `modified` 重排。两种例外值得知道:`memory(action="replace")` 是**原地**重写那一行,被改写的老记忆会保持原位(可能就留在窗口外)直到下次 `/dream` 重排;而文件顶部的手写标题/分组/注释是**位置**语义而不是时间语义,索引一旦超过 50 行,先被丢出窗口的正是这块手工整理的内容。因此写满的索引恰好注入**最新的 50 条记忆**(窗口一旦截断,`# Memory Index` 头行与它下面的空行就落在窗口外);48 条及以下则整份注入。被略过的记忆**没有丢**:auto-surfacing、`memory(action="search")` 与 `/dream` 整理都还能看到它们。`/memory` 用两行区分两套口径:`Index:` 是写入口径(磁盘真相),`Inject:` 是进 system prompt 的窗口。
|
|
120
|
+
|
|
121
|
+
这也是 `/dream` 不再是「可选的整理」而是**容量管理必需**的原因。两个阀值要分开看:prompt 可见性止于注入窗口,想让每条记忆都进 system prompt,就在**接近 48 条之前**整理;199 只是写入口径的硬上限,过了它索引自身就必须缩小。
|
|
112
122
|
|
|
113
123
|
## 配置
|
|
114
124
|
|
|
@@ -120,8 +130,8 @@ staging 的 SSH 用 2222 端口,密钥在 ~/.ssh/staging。
|
|
|
120
130
|
"memoryDir": "~/.pi/memory",
|
|
121
131
|
"memIndexMaxLines": 200,
|
|
122
132
|
"memIndexMaxBytes": 25600,
|
|
123
|
-
"memIndexInjectMaxLines":
|
|
124
|
-
"memIndexInjectMaxBytes":
|
|
133
|
+
"memIndexInjectMaxLines": 50,
|
|
134
|
+
"memIndexInjectMaxBytes": 16384,
|
|
125
135
|
"lock": { "timeoutMs": 5000, "snapshotKeep": 5 },
|
|
126
136
|
"defaults": { "model": "provider/model-id", "sessionPersistence": { "enabled": false } },
|
|
127
137
|
"dream": { "nudgeAfterSessions": 5, "nudgeAfterHours": 24, "thinkLevel": "high" },
|
|
@@ -134,7 +144,7 @@ staging 的 SSH 用 2222 端口,密钥在 ~/.ssh/staging。
|
|
|
134
144
|
"maxInjectionBytes": 10240
|
|
135
145
|
},
|
|
136
146
|
"extractMemories": {
|
|
137
|
-
"enabled":
|
|
147
|
+
"enabled": false,
|
|
138
148
|
"thinkLevel": "high",
|
|
139
149
|
"maxContextTokens": 2000,
|
|
140
150
|
"maxToolResultChars": 500,
|
|
@@ -151,8 +161,8 @@ staging 的 SSH 用 2222 端口,密钥在 ~/.ssh/staging。
|
|
|
151
161
|
| `memoryDir` | `~/.pi/memory` | 所有记忆数据的根目录 |
|
|
152
162
|
| `memIndexMaxLines` | `200` | 写入口径:`MEMORY.md` 的最大非空行数(`# Memory Index` 头行与手写标题同样占额度,所以并不等于记忆条数) |
|
|
153
163
|
| `memIndexMaxBytes` | `25600` | 写入口径:`MEMORY.md` 的最大字节数 |
|
|
154
|
-
| `memIndexInjectMaxLines` | `
|
|
155
|
-
| `memIndexInjectMaxBytes` | `
|
|
164
|
+
| `memIndexInjectMaxLines` | `50` | 注入口径:放进 `memory_index` section 的最大行数。窗口保留**最新**的行、丢弃**最旧**的行 —— 索引是纯时间序,窗口再小也不会藏住你刚写完的那条。**任一键写 `0` = 完全不注入索引**(section 保持空值) |
|
|
165
|
+
| `memIndexInjectMaxBytes` | `16384` | 注入口径:section 的最大字节数(优先丢最旧的行,截断标记在**开头**) |
|
|
156
166
|
| `lock.timeoutMs` | `5000` | 单次原语等逻辑锁 / 等跨进程 `.lock` 的上限。同时也是 `session_shutdown` 等在途写入的上限 |
|
|
157
167
|
| `lock.snapshotKeep` | `5` | `.backups/` 保留的回滚点数量(`migrate-` 前缀的目录永不裁剪 —— 它们是旧版迁移留下的整目录快照,`originals/` 子目录里装着 2.0 之前的 topic 原文) |
|
|
158
168
|
| `defaults.model` | —(必需) | 三个子任务的共享模型。**没有默认值**:会执行的任务必须能解析出模型,否则启动失败(见[模型配置](#模型配置))。per-task 覆盖它 |
|
|
@@ -172,7 +182,7 @@ staging 的 SSH 用 2222 端口,密钥在 ~/.ssh/staging。
|
|
|
172
182
|
| `autoSurfacing.maxEntryBytes` | `3072` | ⭐ 单条 entry 正文的注入字节上限(超出截断)。取代 1.x 的 `maxTopicBytes`(旧键已失效) |
|
|
173
183
|
| `autoSurfacing.maxInjectionBytes` | `10240` | ⭐ 每回合注入内容的总字节上限 |
|
|
174
184
|
| `autoSurfacing.sessionPersistence.*` | 继承 `defaults` | 把侧查询会话落盘 |
|
|
175
|
-
| `extractMemories.enabled` | `
|
|
185
|
+
| `extractMemories.enabled` | `false` | ⭐ 开启每轮自动提取。**默认关闭**(opt-in):开启后每轮结束都会跑一次 headless 模型调用 |
|
|
176
186
|
| `extractMemories.model` | — | ⭐ 提取 agent 用的模型。回退 `defaults.model`;没有 `defaults.model` 时必填(必须可解析,不回退父会话模型) |
|
|
177
187
|
| `extractMemories.thinkLevel` | `"high"` | ⭐ 提取的思考强度 |
|
|
178
188
|
| `extractMemories.maxContextTokens` | `2000` | ⭐ 渲染后对话的预算(`× 4` 个字符;超出时先裁中段、首尾优先保留,user 消息最后才动) |
|
|
@@ -207,7 +217,7 @@ headless 会话默认落在 `<项目记忆目录>/sessions/` —— 在项目记
|
|
|
207
217
|
|---|---|
|
|
208
218
|
| `session_start` | 加载配置 → 校验必需模型(失败即什么都不初始化)→ 解析记忆目录 → **确定索引值并冻结**(`startup`/`new` 读磁盘;`resume`/`fork`/`reload` 重放 transcript 取录制值)→ 注册 `memory` 工具(仅首次,5 个 action)→ 重建清单缓存 → dream 提醒检查 |
|
|
209
219
|
| `before_agent_start` | 把冻结值写进 `sections["memory_index"]`(**每一轮、无条件**),然后做 auto-surfacing(主会话且非子 agent) |
|
|
210
|
-
| `agent_end` |
|
|
220
|
+
| `agent_end` | `extractMemories.enabled` 开启时触发异步 extract;写入成功通知 `Extracted N memories.`,失败通知 `Extract failed: …`(每会话一次) |
|
|
211
221
|
| `session_compact` | 清空已注入集合,**并从磁盘重读索引** —— 会话内唯一的刷新点 |
|
|
212
222
|
| `session_shutdown` | 等在途写入收尾(上限 `lock.timeoutMs`),避免退出时留下 stale 的 `.lock` |
|
|
213
223
|
|
|
@@ -228,6 +238,8 @@ pi 用有序的 sections 构建 system prompt,只对**值发生变化**的 sec
|
|
|
228
238
|
|
|
229
239
|
### 自动提取
|
|
230
240
|
|
|
241
|
+
> **默认关闭。** 先在 `memory.json` 里设 `"extractMemories": { "enabled": true }`(配置只在 `session_start` 读一次,改动需重启会话),并确保 `extractMemories.model` 或 `defaults.model` 可解析。
|
|
242
|
+
|
|
231
243
|
extract 拿到的是结构化渲染,而不是有损的两条消息摘要:
|
|
232
244
|
|
|
233
245
|
```
|
|
@@ -303,12 +315,16 @@ memory(action: "add" | "replace" | "remove" | "list" | "search",
|
|
|
303
315
|
Memory: enabled
|
|
304
316
|
Dir: /home/you/.pi/memory/git/github.com__owner__repo
|
|
305
317
|
Index: 38/200 lines, 2841/25600 bytes, 1 unrecognized lines
|
|
318
|
+
Inject: 39/50 lines, 2841/16384 bytes
|
|
306
319
|
Entries: 37
|
|
320
|
+
Modules: dream=on(provider/model-a) extractMemories=off autoSurfacing=on(provider/model-b)
|
|
307
321
|
Last dream: 2026-10-01T22:10:04.882Z
|
|
308
322
|
Lock: free
|
|
309
323
|
```
|
|
310
324
|
|
|
311
|
-
- `Index` 用**写入**口径(`memIndexMax*`),并报告索引里有多少非空行解析不出(`# Memory Index` 头行与手写标题会计入)。CRLF(以及单独的 CR)行尾在解析前就被归一为 LF,下一次写入也一律输出 LF,因此被 Windows 编辑器改过行尾的 `MEMORY.md`
|
|
325
|
+
- `Index` 用**写入**口径(`memIndexMax*`),并报告索引里有多少非空行解析不出(`# Memory Index` 头行与手写标题会计入)。CRLF(以及单独的 CR)行尾在解析前就被归一为 LF,下一次写入也一律输出 LF,因此被 Windows 编辑器改过行尾的 `MEMORY.md` **不会**推高这个计数。注入侧同样做归一:CRLF 文件不会把 `\r` 送进 system prompt。
|
|
326
|
+
- `Inject` 用**注入**口径(`memIndexInjectMax*`),统计窗口内的行数与字节数 —— 即真正会进 `memory_index` section 的索引文本(截断标记本身不计入)。它与真正注入的值由同一份窗口代码算出来,不可能漂移。注意两行的口径不同:`Index` 数的是**非空**行,`Inject` 数的是窗口内的**全部**行,所以规范索引(LF 行尾、以换行结尾、头部后有且仅有一个空行)下 `Inject` 会比 `Index` 多一行而字节数相同。system prompt 里的值是**会话内冻结**的(见[为什么索引是冻结的](#为什么索引是冻结的)):`session_start` 之后写入的记忆会立刻出现在 `Index`,但要等 compaction 或下一个会话才出现在 `Inject`。
|
|
327
|
+
- `Modules` 报三个模型驱动功能的激活状态:`on(<生效模型>)` / `off`。生效模型 = 该任务自己的 `model`,没有则用 `defaults.model`。`dream` 没有独立开关 —— memory 系统启用它就可用。
|
|
312
328
|
- `Lock` 有三种:`free`、`held by <op> (pid N on <hostname>, started <ISO>)`、`unreadable — run /memory unlock`。`/memory unlock` 的确认框会显示同一行持有者信息。
|
|
313
329
|
- 以 `enabled: false` 启动的会话在启动时不初始化任何东西:`/memory` 报两行(`Memory: disabled` + `Dir: not initialized — set "enabled": true in memory.json and restart`);会话中途无法开启;`/memory unlock` 不需要 store 也能用。
|
|
314
330
|
- 必需模型缺失或解析不出时不初始化任何东西,`/memory` 报 `Memory: misconfigured` + `Dir: not initialized` + 每行一条 `- <error>`;同样的错误在 session_start 时以 error 通知出现。
|
package/index.ts
CHANGED
|
@@ -2,13 +2,13 @@ import { unlink } from "node:fs/promises";
|
|
|
2
2
|
import { join } from "node:path";
|
|
3
3
|
import type { ExtensionAPI, ExtensionContext, ExtensionUIContext } from "@earendil-works/pi-coding-agent";
|
|
4
4
|
import { SessionManager } from "@earendil-works/pi-coding-agent";
|
|
5
|
-
import { loadConfig, modelConfigErrors, requiredModel, type MemoryConfig, type SessionPersistenceConfig } from "./src/config";
|
|
5
|
+
import { loadConfig, modelConfigErrors, requiredModel, taskModel, type MemoryConfig, type SessionPersistenceConfig } from "./src/config";
|
|
6
6
|
import { runDream } from "./src/dream";
|
|
7
7
|
import { indexCapacity, parseEntryIndex } from "./src/entry-index";
|
|
8
8
|
import { runExtract } from "./src/extract";
|
|
9
9
|
import { readLockStatus, type LockInfo } from "./src/fs-lock";
|
|
10
10
|
import { readRecordedMemoryIndex } from "./src/index-source";
|
|
11
|
-
import { applyIndexSection, buildIndexSection, buildInjection, injectSurfacedContent, runSideQuery, scanEntries } from "./src/inject";
|
|
11
|
+
import { applyIndexSection, buildIndexSection, buildInjection, indexInjectionCapacity, injectSurfacedContent, runSideQuery, scanEntries } from "./src/inject";
|
|
12
12
|
import {
|
|
13
13
|
createMemoryTool,
|
|
14
14
|
DREAM_ACTIONS,
|
|
@@ -68,6 +68,28 @@ function countInjectedBlocks(content: string): number {
|
|
|
68
68
|
return Math.max(0, content.split("\n## ").length - 1);
|
|
69
69
|
}
|
|
70
70
|
|
|
71
|
+
/**
|
|
72
|
+
* `/memory` 的模块激活状态一行:`dream=on(model) extractMemories=off autoSurfacing=on(model)`。
|
|
73
|
+
*
|
|
74
|
+
* `dream` 没有独立开关 —— memory 系统启用(能走到状态分支)它就可用;另外两个直接反映各自的
|
|
75
|
+
* `enabled`。模型是**生效值**(per-task 优先,其次 `defaults.model`),关闭的模块不显示模型:
|
|
76
|
+
* 用户问「为什么没生效」时答案在开关上,而不在模型上。
|
|
77
|
+
*
|
|
78
|
+
* 导出供测试直接覆盖:正常会话里「模型解析不出」那条路径走不到(`session_start` 会先拦成
|
|
79
|
+
* misconfigured),只能直接调纯函数。
|
|
80
|
+
*/
|
|
81
|
+
export function moduleStatusLine(cfg: MemoryConfig): string {
|
|
82
|
+
return (["dream", "extractMemories", "autoSurfacing"] as const)
|
|
83
|
+
.map((task) => {
|
|
84
|
+
// `?.` 是防御:`deepMerge` 会把用户写的 `"extractMemories": null` 原样带进来,
|
|
85
|
+
// 那时 `cfg[task].enabled` 会抛错,把整个 `/memory` 命令带崩。
|
|
86
|
+
const enabled = task === "dream" ? true : cfg[task]?.enabled;
|
|
87
|
+
if (!enabled) return `${task}=off`;
|
|
88
|
+
return `${task}=on(${taskModel(cfg, task) ?? "no model"})`;
|
|
89
|
+
})
|
|
90
|
+
.join(" ");
|
|
91
|
+
}
|
|
92
|
+
|
|
71
93
|
/** `<op> (pid N on <hostname>, started <ISO>)` —— `/memory` 的 Lock 行与 unlock 确认框共用同一份描述。 */
|
|
72
94
|
function describeHolder(holder: LockInfo): string {
|
|
73
95
|
return `${holder.op} (pid ${holder.pid} on ${holder.hostname}, started ${holder.startedAt})`;
|
|
@@ -550,16 +572,21 @@ export default function (pi: ExtensionAPI) {
|
|
|
550
572
|
ctx.ui.notify(lines.join("\n"), "info");
|
|
551
573
|
return;
|
|
552
574
|
}
|
|
553
|
-
//
|
|
554
|
-
//
|
|
555
|
-
//
|
|
575
|
+
// 两套口径都要报:`Index:` 是写入口径(用户要知道「还能不能写」),`Inject:` 是注入口径
|
|
576
|
+
// —— 窗口只取索引**最新**的 50 行,与写入口径解耦,只报前者会让用户以为 Entries 全在
|
|
577
|
+
// prompt 里。`indexInjectionCapacity` 与真正注入共用同一个窗口核心,数字不可能漂移。
|
|
578
|
+
// unrecognized 是索引里非空但解析不了的行数 —— 手写标题/分组/被 Windows 编辑器改坏的行
|
|
579
|
+
// 都在这里露出来。
|
|
556
580
|
const indexRaw = await activeStore.readIndex();
|
|
557
581
|
const cap = indexCapacity(indexRaw, config.memIndexMaxLines, config.memIndexMaxBytes);
|
|
582
|
+
const inject = indexInjectionCapacity(indexRaw, config.memIndexInjectMaxLines, config.memIndexInjectMaxBytes);
|
|
558
583
|
const summary = [
|
|
559
584
|
`Memory: ${config.enabled ? "enabled" : "disabled"}`,
|
|
560
585
|
`Dir: ${dir}`,
|
|
561
586
|
`Index: ${cap.lineCount}/${config.memIndexMaxLines} lines, ${cap.byteLength}/${config.memIndexMaxBytes} bytes, ${parseEntryIndex(indexRaw).unrecognized} unrecognized lines`,
|
|
587
|
+
`Inject: ${inject.lineCount}/${config.memIndexInjectMaxLines} lines, ${inject.byteLength}/${config.memIndexInjectMaxBytes} bytes`,
|
|
562
588
|
`Entries: ${(await activeStore.listEntries()).length}`,
|
|
589
|
+
`Modules: ${moduleStatusLine(config)}`,
|
|
563
590
|
`Last dream: ${(await readDreamMeta(dir))?.lastDreamAt ?? "never"}`,
|
|
564
591
|
`Lock: ${await lockStatusLine(dir)}`,
|
|
565
592
|
].join("\n");
|
package/package.json
CHANGED
package/src/config.ts
CHANGED
|
@@ -35,6 +35,7 @@ export interface AutoSurfacingConfig {
|
|
|
35
35
|
}
|
|
36
36
|
|
|
37
37
|
export interface ExtractMemoriesConfig {
|
|
38
|
+
/** 每轮自动提取。opt-in:默认 `false` —— 开启后每轮结束都会跑一次 headless 模型调用。 */
|
|
38
39
|
enabled: boolean;
|
|
39
40
|
model?: string;
|
|
40
41
|
thinkLevel: ThinkLevel;
|
|
@@ -55,12 +56,13 @@ export interface MemoryConfig {
|
|
|
55
56
|
/** Write capacity: max bytes of serialized MEMORY.md index. */
|
|
56
57
|
memIndexMaxBytes: number;
|
|
57
58
|
/**
|
|
58
|
-
* 注入截断:`memory_index` section
|
|
59
|
-
*
|
|
60
|
-
*
|
|
59
|
+
* 注入截断:`memory_index` section 最多带多少**行**索引(默认 50)。
|
|
60
|
+
* 窗口取索引**最新**的一段(索引是纯时间序,见 `truncateIndexForInjection`),所以被丢掉的
|
|
61
|
+
* 永远是**最旧**的记忆。写入口径(`memIndexMaxLines`)刻意**不随**它收紧:索引里能留更多条目,
|
|
62
|
+
* 超窗口的部分只靠 auto-surfacing / `memory` 工具检索,不进 system prompt。
|
|
61
63
|
*/
|
|
62
64
|
memIndexInjectMaxLines: number;
|
|
63
|
-
/** 注入截断:`memory_index` section
|
|
65
|
+
/** 注入截断:`memory_index` section 最多带多少**字节**(默认 16384 ≈ 50 条中文索引行的实测上界)。 */
|
|
64
66
|
memIndexInjectMaxBytes: number;
|
|
65
67
|
/**
|
|
66
68
|
* 两级锁的参数(spec §5.2)。结构与 `StoreConfig["lock"]` 逐字一致,因此可以原样传给
|
|
@@ -92,9 +94,10 @@ export const DEFAULT_CONFIG: MemoryConfig = {
|
|
|
92
94
|
memoryDir: join(homedir(), CONFIG_DIR_NAME, "memory"),
|
|
93
95
|
memIndexMaxLines: 200,
|
|
94
96
|
memIndexMaxBytes: 25600,
|
|
95
|
-
//
|
|
96
|
-
|
|
97
|
-
|
|
97
|
+
// 注入口径独立于写入口径(v2.3.0 起):窗口只取索引**最新**的 50 行。
|
|
98
|
+
// 16384 B ≈ 50 条中文索引行的实测上界(本机样本 221–321 B/行),保证「行数」才是真正生效的上限。
|
|
99
|
+
memIndexInjectMaxLines: 50,
|
|
100
|
+
memIndexInjectMaxBytes: 16384,
|
|
98
101
|
lock: { timeoutMs: 5000, snapshotKeep: 5 },
|
|
99
102
|
dream: { nudgeAfterSessions: 5, nudgeAfterHours: 24, thinkLevel: "high" },
|
|
100
103
|
sessionSearch: { maxSessions: 10, maxMatches: 5 },
|
|
@@ -106,7 +109,8 @@ export const DEFAULT_CONFIG: MemoryConfig = {
|
|
|
106
109
|
maxInjectionBytes: 10240,
|
|
107
110
|
},
|
|
108
111
|
extractMemories: {
|
|
109
|
-
|
|
112
|
+
// 每轮结束都要跑一次 headless 模型调用,代价必须由用户显式承担:默认关闭。
|
|
113
|
+
enabled: false,
|
|
110
114
|
thinkLevel: "high",
|
|
111
115
|
maxContextTokens: 2000,
|
|
112
116
|
maxToolResultChars: 500,
|
|
@@ -117,9 +121,11 @@ export const DEFAULT_CONFIG: MemoryConfig = {
|
|
|
117
121
|
/** 需要显式模型的子任务。顺序固定:dream → extractMemories → autoSurfacing(校验信息按此顺序输出)。 */
|
|
118
122
|
export type ModelTask = "dream" | "extractMemories" | "autoSurfacing";
|
|
119
123
|
|
|
120
|
-
/** 某任务的模型值:per-task 优先,其次共享的 defaults.model
|
|
121
|
-
function taskModel(cfg: MemoryConfig, task: ModelTask): string | undefined {
|
|
122
|
-
|
|
124
|
+
/** 某任务的模型值:per-task 优先,其次共享的 defaults.model。`/memory` 的模块状态行也用它。 */
|
|
125
|
+
export function taskModel(cfg: MemoryConfig, task: ModelTask): string | undefined {
|
|
126
|
+
// `?.`:`deepMerge` 把用户写的 `"dream": null` 原样带进来 —— 那时应该报「没有模型」,
|
|
127
|
+
// 而不是抛 TypeError,把启动校验变成一句看不懂的初始化失败。
|
|
128
|
+
return cfg[task]?.model ?? cfg.defaults?.model;
|
|
123
129
|
}
|
|
124
130
|
|
|
125
131
|
/**
|
package/src/entry-index.ts
CHANGED
|
@@ -17,7 +17,7 @@ export interface ParsedEntryIndex {
|
|
|
17
17
|
}
|
|
18
18
|
|
|
19
19
|
/**
|
|
20
|
-
*
|
|
20
|
+
* 索引唯一的按行拆分入口(解析、写入与**注入**必须共用它,否则同一个 `lineNo` 在三处含义不同)。
|
|
21
21
|
*
|
|
22
22
|
* CRLF(以及老 Mac 的单独 CR)必须在这里归一:JS 的 `.` 不匹配 `\r`,且无 `m` 标志的 `$`
|
|
23
23
|
* 也无法在一个 `\r` 之前成立 —— `LINE_RE` 对 CRLF 行会**完全匹配不上**,于是每一行都被计成
|
|
@@ -25,7 +25,7 @@ export interface ParsedEntryIndex {
|
|
|
25
25
|
* 归一之后写回仍只输出 `\n`(`joinLines`),因此对 CRLF 文件的首次写入会把它转成 LF ——
|
|
26
26
|
* 这是有意的,它消除的是「混合行尾」这个更糟的中间态。
|
|
27
27
|
*/
|
|
28
|
-
function splitLines(raw: string): string[] {
|
|
28
|
+
export function splitLines(raw: string): string[] {
|
|
29
29
|
const text = raw.includes("\r") ? raw.replace(/\r\n?/g, "\n") : raw;
|
|
30
30
|
if (text.length === 0) return [];
|
|
31
31
|
const lines = text.split("\n");
|
package/src/inject.ts
CHANGED
|
@@ -2,13 +2,22 @@ import type { ModelRegistry } from "@earendil-works/pi-coding-agent";
|
|
|
2
2
|
import { runHeadlessAgent } from "./agent-runner";
|
|
3
3
|
import type { SessionPersistenceConfig, ThinkLevel } from "./config";
|
|
4
4
|
import type { EntryType } from "./entry-file";
|
|
5
|
+
import { splitLines } from "./entry-index";
|
|
5
6
|
import { MEMORY_INDEX_SECTION } from "./index-source";
|
|
6
7
|
import type { MemoryStore } from "./memory-store";
|
|
7
8
|
import { sanitizeForInjection } from "./sanitize";
|
|
8
9
|
|
|
9
10
|
/**
|
|
10
|
-
*
|
|
11
|
-
*
|
|
11
|
+
* entry 正文的截断标记。与 `INDEX_TRUNCATION_MARKER` **刻意分开**:这里截的是
|
|
12
|
+
* `<relevant_memories>` 里的一条正文,说「memory index」会指错对象。
|
|
13
|
+
*/
|
|
14
|
+
export const ENTRY_TRUNCATION_MARKER = "[truncated: memory entry exceeds injection limit]";
|
|
15
|
+
|
|
16
|
+
/**
|
|
17
|
+
* entry 正文用的「按行 + 按字节」截断:**保留开头**,从尾部削。
|
|
18
|
+
*
|
|
19
|
+
* 索引不进这里 —— 它的窗口取**最新**的一段(`truncateIndexForInjection`),方向正好相反:
|
|
20
|
+
* 一起用会让窗口永远只保留最旧的记忆。
|
|
12
21
|
*/
|
|
13
22
|
export function truncateForInjection(
|
|
14
23
|
content: string,
|
|
@@ -23,15 +32,135 @@ export function truncateForInjection(
|
|
|
23
32
|
truncated = true;
|
|
24
33
|
}
|
|
25
34
|
if (Buffer.byteLength(out, "utf8") > maxBytes) {
|
|
26
|
-
|
|
27
|
-
while (Buffer.byteLength(cut, "utf8") > maxBytes && cut.length > 0) cut = cut.slice(0, -1);
|
|
28
|
-
out = cut;
|
|
35
|
+
out = cutToBytes(out, maxBytes);
|
|
29
36
|
truncated = true;
|
|
30
37
|
}
|
|
31
|
-
if (truncated) out += `\n
|
|
38
|
+
if (truncated) out += `\n${ENTRY_TRUNCATION_MARKER}`;
|
|
32
39
|
return { ok: !truncated, content: out, truncated };
|
|
33
40
|
}
|
|
34
41
|
|
|
42
|
+
/**
|
|
43
|
+
* 按字符从尾部削到字节上限。
|
|
44
|
+
*
|
|
45
|
+
* 顺手丢掉被切一半的代理对:孤立高代理是**无效 UTF-16**,任何编码器都会把它变成 U+FFFD ——
|
|
46
|
+
* 宁可少一个字符,也不要给模型塞一个替换字符。
|
|
47
|
+
*/
|
|
48
|
+
function cutToBytes(text: string, maxBytes: number): string {
|
|
49
|
+
let cut = text;
|
|
50
|
+
while (Buffer.byteLength(cut, "utf8") > maxBytes && cut.length > 0) cut = cut.slice(0, -1);
|
|
51
|
+
const last = cut.charCodeAt(cut.length - 1);
|
|
52
|
+
return cut.length > 0 && last >= 0xd800 && last <= 0xdbff ? cut.slice(0, -1) : cut;
|
|
53
|
+
}
|
|
54
|
+
|
|
55
|
+
/**
|
|
56
|
+
* 注入截断标记。放在窗口**开头**:索引是纯时间序,被丢掉的永远是**最旧**的行 ——
|
|
57
|
+
* 标记若追加在末尾,读起来像「这条之后的都被切了」,而事实恰好相反。
|
|
58
|
+
*/
|
|
59
|
+
export const INDEX_TRUNCATION_MARKER = "[truncated: memory index exceeds injection limit; older entries omitted]";
|
|
60
|
+
|
|
61
|
+
/**
|
|
62
|
+
* 索引文本 → 窗口用的行数组。
|
|
63
|
+
*
|
|
64
|
+
* 行拆分**共用 `entry-index` 的 `splitLines`**(先做行尾归一):`/memory` 的 `Index:` 与
|
|
65
|
+
* `Inject:` 必须活在同一个「行」的定义上,否则 CRLF / 单独 CR 的文件上两行行数会互相矛盾。
|
|
66
|
+
* 额外摘掉**尾部全部空行**:`rebuildIndex` 把它们当噪音,但它们照样白吃窗口行数。
|
|
67
|
+
*/
|
|
68
|
+
function indexWindowLines(content: string): string[] {
|
|
69
|
+
const lines = splitLines(content);
|
|
70
|
+
while (lines.length > 0 && lines[lines.length - 1].trim() === "") lines.pop();
|
|
71
|
+
return lines;
|
|
72
|
+
}
|
|
73
|
+
|
|
74
|
+
interface IndexWindow {
|
|
75
|
+
/** 窗口内的行,保持文件原顺序(旧 → 新)。 */
|
|
76
|
+
lines: string[];
|
|
77
|
+
/** 窗口的字节数(行间换行算在内;标记与末尾换行不算)。 */
|
|
78
|
+
byteLength: number;
|
|
79
|
+
/** 是否有行被丢弃(含「只保留了最新一行的字节前缀」这种退化)。 */
|
|
80
|
+
truncated: boolean;
|
|
81
|
+
}
|
|
82
|
+
|
|
83
|
+
/** 从尾部(最新)向头部累积,直到行数或字节数触顶。`truncateIndexForInjection` 与
|
|
84
|
+
* `indexInjectionCapacity` 共用它 —— 两处口径不可能漂移。 */
|
|
85
|
+
function indexWindow(lines: string[], maxLines: number, maxBytes: number): IndexWindow {
|
|
86
|
+
const kept: string[] = [];
|
|
87
|
+
let byteLength = 0;
|
|
88
|
+
let degenerated = false;
|
|
89
|
+
for (let i = lines.length - 1; i >= 0; i--) {
|
|
90
|
+
const line = lines[i];
|
|
91
|
+
const lineBytes = Buffer.byteLength(line, "utf8");
|
|
92
|
+
const cost = kept.length === 0 ? lineBytes : lineBytes + 1;
|
|
93
|
+
if (kept.length >= maxLines || byteLength + cost > maxBytes) break;
|
|
94
|
+
kept.unshift(line);
|
|
95
|
+
byteLength += cost;
|
|
96
|
+
}
|
|
97
|
+
if (kept.length === 0 && lines.length > 0 && maxLines > 0 && maxBytes > 0) {
|
|
98
|
+
// 退化:最新一行自身就超预算 —— 保留它的字节前缀。空 section 比截断的一行更糟:
|
|
99
|
+
// 模型连「这里本来有索引」都看不到。两个非正守卫是内部防御(调用方已先挡掉非正预算)。
|
|
100
|
+
const prefix = cutToBytes(lines[lines.length - 1], maxBytes);
|
|
101
|
+
kept.push(prefix);
|
|
102
|
+
byteLength = Buffer.byteLength(prefix, "utf8");
|
|
103
|
+
// 行数没变(1 → 1),但内容确实被削过:必须算截断,否则 `truncateIndexForInjection`
|
|
104
|
+
// 会以为「没丢行」而返回未截断的原文,标记也不出现。
|
|
105
|
+
degenerated = true;
|
|
106
|
+
}
|
|
107
|
+
return { lines: kept, byteLength, truncated: degenerated || kept.length < lines.length };
|
|
108
|
+
}
|
|
109
|
+
|
|
110
|
+
/** 行数组 → 注入文本(行间与末尾都是 LF);空窗口不补换行,只留标记。 */
|
|
111
|
+
function joinWindowLines(lines: string[]): string {
|
|
112
|
+
return lines.length === 0 ? "" : `${lines.join("\n")}\n`;
|
|
113
|
+
}
|
|
114
|
+
|
|
115
|
+
/**
|
|
116
|
+
* 索引的注入窗口:保留**最新**的 maxLines 行 / maxBytes 字节。
|
|
117
|
+
*
|
|
118
|
+
* - 窗口内保持文件原顺序(旧 → 新),**不反转**:形如「索引末尾的一段连续片段」;
|
|
119
|
+
* - 超出字节预算时继续丢窗口内**最旧**的一端;
|
|
120
|
+
* - 截断时在开头加 `INDEX_TRUNCATION_MARKER`;
|
|
121
|
+
* - 未超预算时**逐字节返回原文**(空索引也是)—— 小索引的注入值与改动前完全一致,
|
|
122
|
+
* 会话内冻结的值也不会因为这次改动而漂移;唯一的例外是 CRLF / 单独 CR 的文件:行拆分
|
|
123
|
+
* 先做归一,这时返回的是 LF 文本(`\r` 不进 system prompt)。
|
|
124
|
+
* - 预算非正(`memIndexInjectMaxLines/Bytes` 写 0)= **不注入**:返回空值而不是裸标记 ——
|
|
125
|
+
* 标记的含义是「索引被截断了」,与「按配置不注入」对模型的含义完全不同。
|
|
126
|
+
*/
|
|
127
|
+
export function truncateIndexForInjection(
|
|
128
|
+
content: string,
|
|
129
|
+
maxLines: number,
|
|
130
|
+
maxBytes: number,
|
|
131
|
+
): { ok: boolean; content: string; truncated: boolean } {
|
|
132
|
+
if (maxLines <= 0 || maxBytes <= 0) return { ok: true, content: "", truncated: false };
|
|
133
|
+
|
|
134
|
+
const window = indexWindow(indexWindowLines(content), maxLines, maxBytes);
|
|
135
|
+
if (!window.truncated) {
|
|
136
|
+
return { ok: true, content: content.includes("\r") ? joinWindowLines(window.lines) : content, truncated: false };
|
|
137
|
+
}
|
|
138
|
+
return { ok: false, content: `${INDEX_TRUNCATION_MARKER}\n${joinWindowLines(window.lines)}`, truncated: true };
|
|
139
|
+
}
|
|
140
|
+
|
|
141
|
+
/**
|
|
142
|
+
* `/memory` 的注入口径。复用 `indexWindow`,所以报出来的数字与实际注入的值不可能不一致。
|
|
143
|
+
*
|
|
144
|
+
* 口径:行数 = 窗口内的索引行(**不含**截断标记那一行);字节数 = 窗口正文的 UTF-8 字节
|
|
145
|
+
* (行间换行与末尾换行算在内,**不含**截断标记)。未截断且文件是 LF 时,它与 `Index:`
|
|
146
|
+
* (写入口径,同样用 `splitLines`)对同一份文件报的字节数相同;真实 section 在截断时还多
|
|
147
|
+
* 出「标记 + 一个换行」。
|
|
148
|
+
*/
|
|
149
|
+
export function indexInjectionCapacity(
|
|
150
|
+
content: string,
|
|
151
|
+
maxLines: number,
|
|
152
|
+
maxBytes: number,
|
|
153
|
+
): { lineCount: number; byteLength: number; truncated: boolean } {
|
|
154
|
+
if (maxLines <= 0 || maxBytes <= 0) return { lineCount: 0, byteLength: 0, truncated: false };
|
|
155
|
+
|
|
156
|
+
const window = indexWindow(indexWindowLines(content), maxLines, maxBytes);
|
|
157
|
+
return {
|
|
158
|
+
lineCount: window.lines.length,
|
|
159
|
+
byteLength: window.lines.length > 0 ? window.byteLength + 1 : 0,
|
|
160
|
+
truncated: window.truncated,
|
|
161
|
+
};
|
|
162
|
+
}
|
|
163
|
+
|
|
35
164
|
export function buildInjection(systemPrompt: string, snapshot: string): string {
|
|
36
165
|
if (!snapshot) return systemPrompt;
|
|
37
166
|
return `${systemPrompt}\n\n${snapshot}`;
|
|
@@ -53,7 +182,7 @@ export function buildInjection(systemPrompt: string, snapshot: string): string {
|
|
|
53
182
|
*/
|
|
54
183
|
export async function buildIndexSection(store: MemoryStore, maxLines: number, maxBytes: number): Promise<string> {
|
|
55
184
|
const raw = await store.readIndex();
|
|
56
|
-
const { content } =
|
|
185
|
+
const { content } = truncateIndexForInjection(raw, maxLines, maxBytes);
|
|
57
186
|
return sanitizeForInjection(content);
|
|
58
187
|
}
|
|
59
188
|
|