minnimemory 1.0.0-beta.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +39 -0
- package/README.md +824 -0
- package/dist/bench.d.ts +98 -0
- package/dist/bench.js +142 -0
- package/dist/benchReport.d.ts +12 -0
- package/dist/benchReport.js +128 -0
- package/dist/bounds.d.ts +40 -0
- package/dist/bounds.js +44 -0
- package/dist/cli.d.ts +15 -0
- package/dist/cli.js +503 -0
- package/dist/compile.d.ts +187 -0
- package/dist/compile.js +516 -0
- package/dist/discover.d.ts +125 -0
- package/dist/discover.js +520 -0
- package/dist/doctor.d.ts +9 -0
- package/dist/doctor.js +67 -0
- package/dist/episodic.d.ts +47 -0
- package/dist/episodic.js +130 -0
- package/dist/hook.d.ts +45 -0
- package/dist/hook.js +104 -0
- package/dist/index.d.ts +18 -0
- package/dist/index.js +18 -0
- package/dist/init.d.ts +125 -0
- package/dist/init.js +475 -0
- package/dist/instructions.d.ts +60 -0
- package/dist/instructions.js +270 -0
- package/dist/mcp.d.ts +109 -0
- package/dist/mcp.js +252 -0
- package/dist/mcpServer.d.ts +136 -0
- package/dist/mcpServer.js +997 -0
- package/dist/paths.d.ts +25 -0
- package/dist/paths.js +47 -0
- package/dist/recall.d.ts +113 -0
- package/dist/recall.js +256 -0
- package/dist/recallDir.d.ts +50 -0
- package/dist/recallDir.js +187 -0
- package/dist/reorganize.d.ts +62 -0
- package/dist/reorganize.js +216 -0
- package/dist/report.d.ts +16 -0
- package/dist/report.js +204 -0
- package/dist/router.d.ts +141 -0
- package/dist/router.js +314 -0
- package/dist/rules.d.ts +32 -0
- package/dist/rules.js +651 -0
- package/dist/scan.d.ts +110 -0
- package/dist/scan.js +173 -0
- package/dist/text.d.ts +158 -0
- package/dist/text.js +395 -0
- package/dist/tokenizer.d.ts +26 -0
- package/dist/tokenizer.js +69 -0
- package/dist/types.d.ts +156 -0
- package/dist/types.js +17 -0
- package/dist/version.d.ts +7 -0
- package/dist/version.js +7 -0
- package/dist/writeProtocol.d.ts +19 -0
- package/dist/writeProtocol.js +45 -0
- package/examples/CLAUDE.md +75 -0
- package/examples/README.md +7 -0
- package/package.json +52 -0
package/README.md
ADDED
|
@@ -0,0 +1,824 @@
|
|
|
1
|
+
# MinniMemoryMCP
|
|
2
|
+
|
|
3
|
+
**Curated memory for AI agents, served over MCP.** Your AI coding agent re-reads its entire
|
|
4
|
+
project memory file on every turn. MinniMemoryMCP is an MCP server whose tools audit that file,
|
|
5
|
+
compile it into a small AlwaysOnMemory body plus OnDemandMemory files loaded only when
|
|
6
|
+
relevant, and serve it back on demand. The `minnimemory` command line underneath is the build layer those tools call;
|
|
7
|
+
it is documented further down for maintainers.
|
|
8
|
+
|
|
9
|
+
Built by MinniAI.
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
BEFORE AFTER
|
|
13
|
+
CLAUDE.md 689 tokens CLAUDE.md 282 tokens <- every turn
|
|
14
|
+
x every turn, all of it .minnimemory/OnDemandMemory/
|
|
15
|
+
963 tokens <- only when relevant
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
Those are real numbers from `examples/CLAUDE.md` in this package, default settings, offline
|
|
19
|
+
tokenizer (`approx-v2`, calibrated against a real BPE encoder, see
|
|
20
|
+
[Measurement policy](#measurement-policy)). Rerun them yourself (after the build in
|
|
21
|
+
[Install](#install)): `npx minnimemory init node_modules/minnimemory/examples/CLAUDE.md`.
|
|
22
|
+
|
|
23
|
+
The classifier was validated on 23 real `CLAUDE.md` and `AGENTS.md` files from public
|
|
24
|
+
repositories on 2026-09-02: every one compiled losslessly, and every one landed inside the
|
|
25
|
+
2,000-token always-loaded budget afterwards.
|
|
26
|
+
|
|
27
|
+
Deterministic, lossless, offline. No LLM, no API key, no network. `doctor`, `init`, and `bench`
|
|
28
|
+
have no runtime dependencies; only `mcp` pulls in the MCP SDK.
|
|
29
|
+
|
|
30
|
+
> **Status: 1.0.0-beta.1, prepared for a public beta; not yet on npm (2026-09-16).** `doctor`, `init`, and `bench` are implemented and tested from the
|
|
31
|
+
> CLI. `mcp` is the product and still early: it works, is tested in-process, and `recall` (only, not the other tools)
|
|
32
|
+
> has been exercised in a second host (Claude Code loading it as an external MCP server). The
|
|
33
|
+
> default client surface is `--profile basic`: `check` before the workspace is compiled and
|
|
34
|
+
> `recall` after, chosen from the workspace at launch; `optimize` and `sync` (same split) are
|
|
35
|
+
> opt-in behind `--allow-write`, back up before every write, and have not yet been exercised
|
|
36
|
+
> outside this project's own test corpus. The original nine tools stay
|
|
37
|
+
> available, unchanged, under `--profile full` for maintainers and for an agent driving a folder
|
|
38
|
+
> reorganization directly. `bench` reports exact session bounds and, given a workload, an
|
|
39
|
+
> explicitly-labelled model of a session; it does not measure a live agent by itself, though a
|
|
40
|
+
> separate internal harness has now measured the memory layer live on two real project files (see
|
|
41
|
+
> [Measurement policy](#measurement-policy)). See [Roadmap](#roadmap).
|
|
42
|
+
|
|
43
|
+
---
|
|
44
|
+
|
|
45
|
+
## Contents
|
|
46
|
+
|
|
47
|
+
- [The problem](#the-problem)
|
|
48
|
+
- [What it does](#what-it-does)
|
|
49
|
+
- [Install](#install)
|
|
50
|
+
- [Use it](#use-it)
|
|
51
|
+
- [What gets generated](#what-gets-generated)
|
|
52
|
+
- [Commands](#commands)
|
|
53
|
+
- [Profiles](#profiles)
|
|
54
|
+
- [Keeping it current](#keeping-it-current)
|
|
55
|
+
- [What doctor counts as always loaded](#what-doctor-counts-as-always-loaded)
|
|
56
|
+
- [Kinds, episodic JSON, and the write protocol](#kinds-episodic-json-and-the-write-protocol)
|
|
57
|
+
- [Security](#security)
|
|
58
|
+
- [Rule set](#rule-set) (moved to `docs/DESIGN.md`)
|
|
59
|
+
- [Design principles](#design-principles) (moved to `docs/DESIGN.md`)
|
|
60
|
+
- [What it does not do](#what-it-does-not-do)
|
|
61
|
+
- [Measurement policy](#measurement-policy) (moved to `docs/DESIGN.md`)
|
|
62
|
+
- [FAQ](#faq)
|
|
63
|
+
- [Roadmap](#roadmap) (moved to `docs/DESIGN.md`)
|
|
64
|
+
- [License](#license)
|
|
65
|
+
|
|
66
|
+
---
|
|
67
|
+
|
|
68
|
+
## The problem
|
|
69
|
+
|
|
70
|
+
Every turn of an agentic coding session re-sends the whole prefix to the model: the system
|
|
71
|
+
prompt, your project memory file (`CLAUDE.md`, `AGENTS.md`, `.cursorrules`), the conversation so
|
|
72
|
+
far, and every tool result.
|
|
73
|
+
|
|
74
|
+
The conversation grows because work is happening. That is fair. **The memory file does not grow
|
|
75
|
+
for any good reason, and it is charged on every turn regardless.**
|
|
76
|
+
|
|
77
|
+
```
|
|
78
|
+
a 12,000 token memory file
|
|
79
|
+
x 200 turns in a working session
|
|
80
|
+
= 2,400,000 tokens of re-sent, identical text
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
Most of it is irrelevant to any given task. Your deploy runbook is being re-sent while you write
|
|
84
|
+
a unit test. Your changelog is being re-sent while you debug auth.
|
|
85
|
+
|
|
86
|
+
There is a second, quieter cost. Prompt caching keys on an **exact prefix match**, so one edited
|
|
87
|
+
character near the top invalidates the cache for everything after it. The usual instinct, putting
|
|
88
|
+
"Current status" and open TODOs at the top of the memory file where they are most visible, is
|
|
89
|
+
therefore exactly backwards: it puts the most-edited content in the most cache-sensitive
|
|
90
|
+
position, busting the cache on every session where anything moved.
|
|
91
|
+
|
|
92
|
+
Memory files also rot in a specific way. They grow by accretion, because appending a paragraph is
|
|
93
|
+
easier than restructuring, so six months in you are paying every turn for a stale directory tree,
|
|
94
|
+
three copies of the test command, and a changelog nobody reads.
|
|
95
|
+
|
|
96
|
+
## What it does
|
|
97
|
+
|
|
98
|
+
Two commands do the work, plus an optional read-only MCP server for hosts that need one.
|
|
99
|
+
|
|
100
|
+
**`doctor`** audits your current setup and tells you what it is costing. It changes nothing, needs
|
|
101
|
+
no install, and takes about a second.
|
|
102
|
+
|
|
103
|
+
**`init`** compiles your memory file into a small AlwaysOnMemory body plus OnDemandMemory files
|
|
104
|
+
the agent loads only when they are relevant, and rewrites your memory file as a compact stub.
|
|
105
|
+
Your original is backed up first, and every line of it survives into the output.
|
|
106
|
+
|
|
107
|
+
Splitting is by **how often content is needed** and **how often it changes**:
|
|
108
|
+
|
|
109
|
+
| tier | always loaded? | holds | churn |
|
|
110
|
+
|---|---|---|---|
|
|
111
|
+
| AlwaysOnMemory | yes | identity, invariants, hard constraints, conventions, plus one routing line per OnDemandMemory file | rarely edited |
|
|
112
|
+
| OnDemandMemory | **no** | architecture, testing, deploy, gotchas, status, changelog | edited freely |
|
|
113
|
+
|
|
114
|
+
## Install
|
|
115
|
+
|
|
116
|
+
Requires Node 18 or later. The package is `minnimemory` on npm; nothing else to download.
|
|
117
|
+
|
|
118
|
+
**Claude Code**, from inside the repo whose memory you want served (read-only: `check` until
|
|
119
|
+
the memory is compiled, `recall` once it is):
|
|
120
|
+
|
|
121
|
+
```bash
|
|
122
|
+
claude mcp add minnimemory -- npx -y minnimemory mcp
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Add `--allow-write` to let the agent also compile and maintain the file (`optimize` before the
|
|
126
|
+
compile, `sync` after; every write is backed up first), and `--include-auto-memory` to include Claude Code's own
|
|
127
|
+
auto-memory folder for this project:
|
|
128
|
+
|
|
129
|
+
```bash
|
|
130
|
+
claude mcp add minnimemory -- npx -y minnimemory mcp --allow-write --include-auto-memory
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
Use `-s user` to register it once for every project instead of the current one. Restart Claude
|
|
134
|
+
Code, then ask it to run `check`.
|
|
135
|
+
|
|
136
|
+
**Any other MCP host** (Cursor, Windsurf, Claude Desktop, an SDK client), in its MCP config:
|
|
137
|
+
|
|
138
|
+
```json
|
|
139
|
+
{
|
|
140
|
+
"mcpServers": {
|
|
141
|
+
"minnimemory": {
|
|
142
|
+
"command": "npx",
|
|
143
|
+
"args": ["-y", "minnimemory", "mcp", "/path/to/your/repo"]
|
|
144
|
+
}
|
|
145
|
+
}
|
|
146
|
+
}
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
The path is optional; without it the server uses its working directory.
|
|
150
|
+
|
|
151
|
+
**Command line**, for the `doctor`/`init`/`bench` commands below: `npx minnimemory doctor`, or
|
|
152
|
+
`npm install -g minnimemory` for a `minnimemory` binary. Maintainers working on the compiler
|
|
153
|
+
itself clone the repo, `npm ci && npm run build`, and run `node dist/cli.js` in place of `npx
|
|
154
|
+
minnimemory`.
|
|
155
|
+
|
|
156
|
+
## Use it
|
|
157
|
+
|
|
158
|
+
Every block below is the real output of that command against
|
|
159
|
+
`examples/CLAUDE.md`, a 75-line memory file for a fictional bookkeeping API, shipped in this
|
|
160
|
+
package.
|
|
161
|
+
|
|
162
|
+
### 1. Find out what your setup costs
|
|
163
|
+
|
|
164
|
+
```
|
|
165
|
+
$ npx minnimemory doctor
|
|
166
|
+
|
|
167
|
+
MM003 "Current status" (lines 6-11) holds a status heading and high
|
|
168
|
+
dated entries and status language in the always-loaded
|
|
169
|
+
prefix
|
|
170
|
+
at CLAUDE.md:6-11
|
|
171
|
+
-> move it to an OnDemandMemory file; editing it here invalidates
|
|
172
|
+
the prompt cache every time
|
|
173
|
+
|
|
174
|
+
MM003 "Changelog" (lines 70-76) holds a status heading and dated high
|
|
175
|
+
entries and a version string in the always-loaded prefix
|
|
176
|
+
at CLAUDE.md:70-76
|
|
177
|
+
-> move it to an OnDemandMemory file; editing it here invalidates
|
|
178
|
+
the prompt cache every time
|
|
179
|
+
|
|
180
|
+
MM002 lines 28-36 are a directory tree derivable from the repo med
|
|
181
|
+
itself
|
|
182
|
+
at CLAUDE.md:28-36
|
|
183
|
+
-> delete it; the agent can list the directory, and this copy goes
|
|
184
|
+
stale silently
|
|
185
|
+
|
|
186
|
+
3 findings (2 high, 1 med)
|
|
187
|
+
always-loaded prefix: 689 tokens across 1 file (budget 2,000, within)
|
|
188
|
+
tokenizer: approx-v2, an offline estimate
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
Every finding names a rule id, a file, a line range, and what to do about it. Nothing is changed.
|
|
192
|
+
|
|
193
|
+
### 2. Preview the compile
|
|
194
|
+
|
|
195
|
+
`init` is a **dry run by default**. It prints the full plan and writes nothing. The plan carries
|
|
196
|
+
the decision as well as the shape: exact session bounds, a plain verdict, and the `doctor`
|
|
197
|
+
findings `init` will not act on. If the verdict is "leave this file as it is", `--write` refuses
|
|
198
|
+
unless you pass `--force`: the tool will not quietly make a file more expensive.
|
|
199
|
+
|
|
200
|
+
```
|
|
201
|
+
$ npx minnimemory init
|
|
202
|
+
|
|
203
|
+
source: CLAUDE.md (689 tokens)
|
|
204
|
+
|
|
205
|
+
AlwaysOnMemory.md 282 tokens always loaded
|
|
206
|
+
OnDemandMemory/ 963 tokens on demand
|
|
207
|
+
|
|
208
|
+
OnDemandMemory/current_status_to_deployment.md 438 test, status, current, request
|
|
209
|
+
OnDemandMemory/common_tasks.md 75 common, tasks, edit, endpoint
|
|
210
|
+
OnDemandMemory/changelog.md 64 changelog, centralised, draft, flag
|
|
211
|
+
OnDemandMemory/memory_write_protocol.md 386 memory, remember, save, write
|
|
212
|
+
|
|
213
|
+
always-loaded prefix: 689 -> 282 tokens (-59.1%)
|
|
214
|
+
of which 0 tokens are the added token-discipline block (--profile none)
|
|
215
|
+
auto: source is under the 2,000 token budget, so no discipline block is embedded
|
|
216
|
+
tokenizer: approx-v2, an offline estimate
|
|
217
|
+
|
|
218
|
+
over a 50-turn session, exact bounds, no workload assumed:
|
|
219
|
+
best case (no OnDemandMemory file opened) 34,450 -> 14,100, saves 20,350 (59.1%)
|
|
220
|
+
break-even 407 tokens of OnDemandMemory can stay open all session
|
|
221
|
+
|
|
222
|
+
verdict: compile. The prefix shrinks on every turn; OnDemandMemory files cost only when opened.
|
|
223
|
+
|
|
224
|
+
doctor findings init does not apply (init moves content, never deletes it):
|
|
225
|
+
MM002 lines 28-36 are a directory tree derivable from the repo itself (lines 28-36)
|
|
226
|
+
-> delete it; the agent can list the directory, and this copy goes stale silently
|
|
227
|
+
|
|
228
|
+
files:
|
|
229
|
+
backup .minnimemory/original/CLAUDE.md.bak
|
|
230
|
+
create .minnimemory/original/.gitignore
|
|
231
|
+
create .minnimemory/AlwaysOnMemory.md
|
|
232
|
+
create .minnimemory/manifest.json
|
|
233
|
+
create .minnimemory/OnDemandMemory/current_status_to_deployment.md
|
|
234
|
+
create .minnimemory/OnDemandMemory/common_tasks.md
|
|
235
|
+
create .minnimemory/OnDemandMemory/changelog.md
|
|
236
|
+
create .minnimemory/OnDemandMemory/memory_write_protocol.md
|
|
237
|
+
overwrite CLAUDE.md
|
|
238
|
+
|
|
239
|
+
dry run, nothing written. Pass --write to apply.
|
|
240
|
+
```
|
|
241
|
+
|
|
242
|
+
### 3. Apply it
|
|
243
|
+
|
|
244
|
+
```bash
|
|
245
|
+
npx minnimemory init --write
|
|
246
|
+
```
|
|
247
|
+
|
|
248
|
+
Your original is copied verbatim to `.minnimemory/original/` before anything is touched.
|
|
249
|
+
|
|
250
|
+
### 4. Check the result
|
|
251
|
+
|
|
252
|
+
```bash
|
|
253
|
+
npx minnimemory doctor
|
|
254
|
+
```
|
|
255
|
+
|
|
256
|
+
### 5. Keep it honest in CI
|
|
257
|
+
|
|
258
|
+
`doctor` exits non-zero when findings reach the threshold, so memory bloat fails the build the
|
|
259
|
+
same way a lint error does.
|
|
260
|
+
|
|
261
|
+
```yaml
|
|
262
|
+
- run: npx minnimemory doctor --ci --fail-on med
|
|
263
|
+
```
|
|
264
|
+
|
|
265
|
+
Exit codes: `0` clean, `1` findings at or above `--fail-on`, `2` execution error.
|
|
266
|
+
|
|
267
|
+
## What gets generated
|
|
268
|
+
|
|
269
|
+
```
|
|
270
|
+
your-repo/
|
|
271
|
+
CLAUDE.md compact stub: AlwaysOnMemory + the OnDemandMemory list
|
|
272
|
+
.minnimemory/
|
|
273
|
+
AlwaysOnMemory.md always loaded: always-on content, then the OnDemandMemory list
|
|
274
|
+
OnDemandMemory/
|
|
275
|
+
<topic>.md loaded only when relevant
|
|
276
|
+
manifest.json token counts, hashes, triggers, provenance
|
|
277
|
+
original/
|
|
278
|
+
CLAUDE.md.bak verbatim copy of what was there before
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
The OnDemandMemory list inside `AlwaysOnMemory.md` (and mirrored into the stub) is a routing
|
|
282
|
+
manifest, not a summary, and it says nothing the discipline block already says:
|
|
283
|
+
|
|
284
|
+
```markdown
|
|
285
|
+
<!-- minnimemory:ondemand-list -->
|
|
286
|
+
# OnDemandMemory
|
|
287
|
+
- `OnDemandMemory/current_status_to_deployment.md`: test, status, current, request, token, integration
|
|
288
|
+
- `OnDemandMemory/common_tasks.md`: common, tasks, edit, endpoint, generated, hour
|
|
289
|
+
- `OnDemandMemory/changelog.md`: changelog, centralised, draft, flag, introduced, map
|
|
290
|
+
- `OnDemandMemory/memory_write_protocol.md`: memory, remember, save, write, note, changelog
|
|
291
|
+
|
|
292
|
+
Read an OnDemandMemory file only when the task needs it, and answer from it, never from this list.
|
|
293
|
+
<!-- /minnimemory:ondemand-list -->
|
|
294
|
+
```
|
|
295
|
+
|
|
296
|
+
The closing line appears only with the `none` profile; with `routing` or `full` the same
|
|
297
|
+
instruction is already in the discipline block, so it is not repeated. Assembly order lives in
|
|
298
|
+
`manifest.json` (`order`), not in the prefix.
|
|
299
|
+
|
|
300
|
+
OnDemandMemory file names are slugs of the source headings each one spans, so the output is
|
|
301
|
+
traceable back to the original file without opening the manifest.
|
|
302
|
+
|
|
303
|
+
`manifest.json` records what came from where: per-file token counts, content hashes for
|
|
304
|
+
staleness detection, the keyword triggers, and the source line range each OnDemandMemory file
|
|
305
|
+
came from.
|
|
306
|
+
|
|
307
|
+
## Commands
|
|
308
|
+
|
|
309
|
+
### `minnimemory doctor [path]`
|
|
310
|
+
|
|
311
|
+
Audits. Never modifies anything. Works on a repo, a directory of memory files, or a single file.
|
|
312
|
+
|
|
313
|
+
| flag | effect |
|
|
314
|
+
|---|---|
|
|
315
|
+
| `--json` | machine-readable output |
|
|
316
|
+
| `--ci` | terse output, no colour |
|
|
317
|
+
| `--fail-on <sev>` | exit 1 at this severity or above: `low`, `med`, `high`. Default `high` |
|
|
318
|
+
| `--budget <n>` | always-loaded token budget for MM001. Default 2000 |
|
|
319
|
+
| `--only <ids>` | run only these rules, e.g. `--only MM001,MM003` |
|
|
320
|
+
| `--ignore <ids>` | skip these rules |
|
|
321
|
+
| `--rules` | list the rule set and exit |
|
|
322
|
+
|
|
323
|
+
### `minnimemory init [path]`
|
|
324
|
+
|
|
325
|
+
Compiles. **Dry run by default.**
|
|
326
|
+
|
|
327
|
+
| flag | effect |
|
|
328
|
+
|---|---|
|
|
329
|
+
| `--write` | actually apply. Without it, prints the plan and exits |
|
|
330
|
+
| `--force` | overwrite an existing `.minnimemory/`, and compile a file the plan's verdict says to leave alone |
|
|
331
|
+
| `--update` | recompile a compiled workspace from its stub and the OnDemandMemory files on disk, keeping hand edits. See [Keeping it current](#keeping-it-current) |
|
|
332
|
+
| `--budget <n>` | always-loaded token budget AlwaysOnMemory must fit. AlwaysOnMemory-like sections that do not fit are routed to OnDemandMemory and listed in the plan. Default 2000 |
|
|
333
|
+
| `--profile <name>` | how much guidance to embed: `auto`, `none`, `routing`, `full`. Default `auto`: `none` when the source is under the budget, `routing` at or over it |
|
|
334
|
+
| `--allow-secrets` | compile even when the source holds a credential-shaped string. Without it, `--write` refuses |
|
|
335
|
+
|
|
336
|
+
`init` follows a pointer: if `CLAUDE.md` is only `@AGENTS.md` or "See `AGENTS.md`", it compiles
|
|
337
|
+
`AGENTS.md`. It never follows a symlink, in either direction.
|
|
338
|
+
|
|
339
|
+
### `minnimemory mcp [path]`
|
|
340
|
+
|
|
341
|
+
Starts an MCP server, `--profile basic` (four verbs, two registered per launch) by default, for hosts that cannot hold the
|
|
342
|
+
OnDemandMemory list in their own prefix or that have no memory file at all. `mcp`'s `--profile`
|
|
343
|
+
picks the tool surface (`basic`/`full`); it is a different flag from `init`'s `--profile`
|
|
344
|
+
(`auto`/`none`/`routing`/`full`, [below](#profiles)), which picks how much discipline guidance
|
|
345
|
+
gets embedded in a compile - same word, two unrelated knobs on two different commands. Register
|
|
346
|
+
it with your agent rather than running it by hand:
|
|
347
|
+
|
|
348
|
+
```bash
|
|
349
|
+
claude mcp add minnimemory -- npx -y minnimemory mcp
|
|
350
|
+
```
|
|
351
|
+
|
|
352
|
+
Other hosts: see [Install](#install).
|
|
353
|
+
|
|
354
|
+
#### The basic profile (default): `check`, `optimize`, `sync`, `recall`
|
|
355
|
+
|
|
356
|
+
Four verbs, two registered per launch. Which two follows from the workspace on disk, with no
|
|
357
|
+
flag to set: the server looks once, at launch, for a `.minnimemory/` directory under its root.
|
|
358
|
+
|
|
359
|
+
| workspace at launch | read-only | with `--allow-write` |
|
|
360
|
+
|---|---|---|
|
|
361
|
+
| not compiled (no `.minnimemory/`) | `check` | `check`, `optimize` |
|
|
362
|
+
| compiled | `recall` | `recall`, `sync` |
|
|
363
|
+
| a memory directory (Layout B, no manifest) | `check`, `recall` | `check`, `recall`, `optimize` |
|
|
364
|
+
|
|
365
|
+
`recall` cannot answer without a manifest, and `check`/`optimize` have no first compile left to
|
|
366
|
+
plan once one exists, so registering the other half would charge its definition every turn for
|
|
367
|
+
a tool that session could not use. The surface is fixed for the life of the process: a session
|
|
368
|
+
that compiles mid-way (`optimize` says so in its response) relaunches to get `recall` and `sync`.
|
|
369
|
+
One consequence: compiling a second, uncompiled file under a root that already holds a
|
|
370
|
+
`.minnimemory/` directory is not reachable over MCP on this profile; the CLI's `init` covers it.
|
|
371
|
+
|
|
372
|
+
| tool | returns |
|
|
373
|
+
|---|---|
|
|
374
|
+
| `check(target?, budget?, format?)` | audit findings, always-loaded token count against budget, drift since init for an already-compiled target, and - when `target` is a single compilable file - the compile plan: AlwaysOnMemory/OnDemandMemory split, session bounds, compile-or-leave verdict. When `target` is a memory directory, already compiled, or an index file, the plan half is left out and one line says why plus what to call instead. `format: "json"` returns `{ doctor, plan, planUnavailableReason }` |
|
|
375
|
+
| `optimize(target?, budget?, force?, ops?)` | write. A compilable file: compiles it, or - already compiled - re-applies it instead and says so. A memory directory: with no `ops`, returns `scan`-shape facts plus an instruction block describing the op shapes to propose; with `ops`, applies them (`write_file`/`delete_file`/`rename_file`), backing up every `.md` file first. `force` forwards to the compile path's own force. Needs `--allow-write` |
|
|
376
|
+
| `sync(target?, budget?, force?)` | write. Checks a compiled `target` for drift (MM010) and, only if something drifted, re-applies it, keeping hand edits. Changes nothing when there is no drift unless `force: true`, a deliberate re-apply (on a compiled target this is the same operation as `optimize`, which is not registered on the compiled surface); never compiles a fresh target. Needs `--allow-write` |
|
|
377
|
+
| `recall(query?, file?, headingsOnly?, maxTokens?)` | `query` given: the memory sections most relevant to it, content inline, best first, under a token cap, restricted to `file` when given. `query` omitted: the OnDemandMemory routing list, no content. `headingsOnly: true` with `file`: that file's headings only. `headingsOnly` without `file` is an error. When its index rebuild finds an OnDemandMemory file changed since `init` wrote it (MM010, the same hash `check`/`doctor` use), the response names the drifted files and points at `sync()`, or at the CLI when `--allow-write` is off |
|
|
378
|
+
|
|
379
|
+
A `recall()` query that matches nothing does not just say so: the response inlines the same
|
|
380
|
+
routing list `query` omitted returns, so the caller lands on the right `file` without a second
|
|
381
|
+
round trip, followed by "No file matched; pick from the list above and call recall() with file
|
|
382
|
+
set." (`module` on the full profile's `recall`).
|
|
383
|
+
|
|
384
|
+
`check`/`optimize`/`sync` accept the literal `"auto-memory"` as `target` where it makes sense
|
|
385
|
+
(needs `--include-auto-memory`), the same resolution `scan`/`reorganize` use in the full profile,
|
|
386
|
+
below. None of the four take `profile`, `episodicJson` or `allowSecrets` over MCP - the CLI keeps
|
|
387
|
+
those for maintainers; the MCP tools always compile with `profile: "auto"` and refuse rather than
|
|
388
|
+
override.
|
|
389
|
+
|
|
390
|
+
#### Memory folders without compiling (2026-09-16)
|
|
391
|
+
|
|
392
|
+
A third launch shape, alongside "not compiled" and "compiled" above: a memory directory an index
|
|
393
|
+
file (`MEMORY.md`, `index.md`, `INDEX.md`, or `README.md`) plus topic `.md` files, or a Claude
|
|
394
|
+
Code auto-memory folder, with no `.minnimemory/` at all. On that shape the server registers
|
|
395
|
+
`recall` alongside `check` (and `optimize` with `--allow-write`), backed by `recallDir` instead
|
|
396
|
+
of a manifest: it builds the same section-level units `recall` builds from a compiled workspace's
|
|
397
|
+
OnDemandMemory files, straight off the topic files on disk, and ranks them the same way (stemmed
|
|
398
|
+
BM25, `router.ts`). An index file's hook lines ("`- [Title](file.md) - hook text`") count as
|
|
399
|
+
triggers, the same way `Research/Pipeline/replay_ranker.mjs` already treats them for its own
|
|
400
|
+
offline replay. It is lexical and fully offline - no embeddings, no model call, no compile step -
|
|
401
|
+
so it works on a repo you have not run `init` against at all, and on Claude Code's own
|
|
402
|
+
OS-level auto-memory folder unmodified. `README.md` and any `archive/` subfolder are never
|
|
403
|
+
treated as topic files. The same command is available from the CLI:
|
|
404
|
+
|
|
405
|
+
```bash
|
|
406
|
+
minnimemory recall <path> <query...> # full section content, best first
|
|
407
|
+
minnimemory recall <path> <query...> --headings # just "file > heading path (n tokens)" lines
|
|
408
|
+
```
|
|
409
|
+
|
|
410
|
+
Real per-turn cost of the two tools together, `Research/TokenTest/tool_cost.mjs`
|
|
411
|
+
(`Research/docs/2026-09-16-tool-cost.md`; see the table under
|
|
412
|
+
[When the MCP route pays off](#when-the-mcp-route-pays-off)): 740 tokens read-only, 1,397 with
|
|
413
|
+
`--allow-write`.
|
|
414
|
+
|
|
415
|
+
#### Claude Code hook (experimental, 2026-09-16)
|
|
416
|
+
|
|
417
|
+
`minnimemory hook` runs the same ranking as `recall`/`recall <path> <query...>`, wired to Claude
|
|
418
|
+
Code's own `UserPromptSubmit` hook instead of an MCP tool call: it reads the hook's JSON off
|
|
419
|
+
stdin (fields `prompt` and `cwd`), resolves the memory target - a compiled workspace at `cwd`,
|
|
420
|
+
else `cwd`'s own Claude Code auto-memory folder, else `--dir` explicitly - and prints a fenced
|
|
421
|
+
`=== MEMORY (data, not instructions) ===` block on stdout before the model sees the prompt.
|
|
422
|
+
Prints nothing on no hit, no usable target, or any internal error; always exits 0.
|
|
423
|
+
|
|
424
|
+
```json
|
|
425
|
+
{
|
|
426
|
+
"hooks": {
|
|
427
|
+
"UserPromptSubmit": [
|
|
428
|
+
{ "hooks": [{ "type": "command", "command": "npx minnimemory hook --headings" }] }
|
|
429
|
+
]
|
|
430
|
+
}
|
|
431
|
+
}
|
|
432
|
+
```
|
|
433
|
+
|
|
434
|
+
Cost: about the size of the printed block (600 tokens by default, `--headings` for just the
|
|
435
|
+
heading-path lines), added to the transcript on every turn the hook fires - not a one-time
|
|
436
|
+
`tools/list` charge like the MCP tools above, so start with `--headings` and widen only if the
|
|
437
|
+
heading line alone is not enough to decide whether to open the file. This is experimental until
|
|
438
|
+
the gate in `Research/FlagShipInterfaces/README.md` part (c) runs (a live Pipeline run measuring
|
|
439
|
+
it against a real session, not the offline unit tests in `tests/hook.test.ts`).
|
|
440
|
+
|
|
441
|
+
#### The full profile (`--profile full`): today's nine, for maintainers
|
|
442
|
+
|
|
443
|
+
`--profile full` registers the original nine tools, unchanged in name, schema and behaviour, so
|
|
444
|
+
anything already built against them keeps working. Use it for the session where you are
|
|
445
|
+
maintaining the compiler itself, or driving a memory-directory reorganization directly rather than
|
|
446
|
+
through `optimize`'s facts-then-ops loop.
|
|
447
|
+
|
|
448
|
+
| tool | returns |
|
|
449
|
+
|---|---|
|
|
450
|
+
| `recall(query, maxTokens?, module?)` | the memory sections most relevant to the query, content inline, best first, under a token cap (1,500 by default); a hit inside a changelog returns that entry, not the log; each result names its OnDemandMemory file, heading path, line range and token count |
|
|
451
|
+
| `modules()` | the OnDemandMemory list, no content |
|
|
452
|
+
| `outline(module)` | an OnDemandMemory file's headings only, so the agent can decide before paying for the body |
|
|
453
|
+
| `doctor(target?, budget?, only?, ignore?, format?)` | audit findings, always-loaded token count against budget, and (for a compiled target) drift since init. `format: "json"` returns the machine-readable report instead of text |
|
|
454
|
+
| `plan(target?, update?, budget?, format?)` | what compiling `target` would do - AlwaysOnMemory, OnDemandMemory, session bounds, the compile-versus-leave verdict - without writing anything. `format: "json"` returns a compact summary object instead of the rendered report |
|
|
455
|
+
| `scan(target?, detail?, files?)` | structured inventory of memory files: frontmatter, sections, volatility evidence, keywords, cross-file duplicates. `target` is a path, or the literal `"auto-memory"` (needs `--include-auto-memory`). `detail: "summary"` (default) returns per-file counts and rollups; `"full"` returns every section. `files` restricts to exact rel matches, with unmatched names in `unknownFiles`. |
|
|
456
|
+
| `apply(target?, budget?)` | compiles `target`, same as `init --write`. Refuses an already-compiled target - call `update()` instead - and a compile that would not shrink the prefix or holds a credential-shaped string, with no override over MCP. Needs `--allow-write` |
|
|
457
|
+
| `update(target?, budget?)` | recompiles an already-compiled `target` from its stub and OnDemandMemory files, same as `init --update`. Needs `--allow-write` |
|
|
458
|
+
| `reorganize(target?, ops)` | applies `write_file` / `delete_file` / `rename_file` operations to `target` (default: this server's root, or `"auto-memory"`, needs `--include-auto-memory`). Needs `--allow-write` |
|
|
459
|
+
|
|
460
|
+
`recall`/`modules`/`outline` need a compiled workspace. `doctor`/`plan`/`scan` do not: they work
|
|
461
|
+
on any target, compiled or not. `apply`/`update` are the single-file compile path (what the CLI's
|
|
462
|
+
`init` and `init --update` do); `reorganize` is the multi-file path.
|
|
463
|
+
|
|
464
|
+
#### The loop
|
|
465
|
+
|
|
466
|
+
Point a client at the basic profile and the intended flow is: **check, then optimize, then
|
|
467
|
+
sync**, with **recall** in between turns once something is compiled.
|
|
468
|
+
|
|
469
|
+
```
|
|
470
|
+
check() - what does this file/directory cost, and what would compiling it do?
|
|
471
|
+
- (show the person the result before doing anything)
|
|
472
|
+
optimize() - compile it (or apply a reorganize plan you both approved)
|
|
473
|
+
recall() - per-turn retrieval against the compiled workspace, in later turns
|
|
474
|
+
sync() - after someone hand-edits a memory file, catch the workspace back up
|
|
475
|
+
```
|
|
476
|
+
|
|
477
|
+
**The server only advertises the tools you asked for, and the basic profile only the half the
|
|
478
|
+
workspace can use.** Every tool definition is charged to the client's prefix on every turn,
|
|
479
|
+
called or not. Real costs (`Research/TokenTest/tool_cost.mjs`, `approx-v2` tokenizer; full table
|
|
480
|
+
and what the number is: `Research/docs/2026-09-16-tool-cost.md`):
|
|
481
|
+
|
|
482
|
+
| launch | workspace | tools advertised | prefix cost |
|
|
483
|
+
|---|---|---|---|
|
|
484
|
+
| `mcp` (default, basic, read-only) | not compiled | `check` | 364 tokens |
|
|
485
|
+
| `mcp --allow-write` (basic) | not compiled | `check`, `optimize` | 1,021 tokens |
|
|
486
|
+
| `mcp` (default, basic, read-only) | compiled | `recall` | 329 tokens |
|
|
487
|
+
| `mcp --allow-write` (basic) | compiled | `recall`, `sync` | 620 tokens |
|
|
488
|
+
| `mcp` (default, basic, read-only) | memory directory | `recall`, `check` | 740 tokens |
|
|
489
|
+
| `mcp --allow-write` (basic) | memory directory | `recall`, `check`, `optimize` | 1,397 tokens |
|
|
490
|
+
| `mcp --profile full` (read-only) | any | `recall`, `modules`, `outline`, `scan`, `doctor`, `plan` | 1,666 tokens |
|
|
491
|
+
| `mcp --profile full --allow-write` | any | + `apply`, `update`, `reorganize` | 2,668 tokens |
|
|
492
|
+
|
|
493
|
+
The compiled rows are what a working session holds - well under this tool's own 2,000-token
|
|
494
|
+
always-loaded budget. A tool that was not registered is absent from `tools/list` and answers
|
|
495
|
+
"tool not found" if called; nothing is present-and-refusing, because a refusing tool still costs
|
|
496
|
+
its definition every turn.
|
|
497
|
+
|
|
498
|
+
Ranking (`recall`) is whole-word keyword overlap against the triggers `init` computed. No
|
|
499
|
+
embeddings, no model call, no network. A query that shares no vocabulary with an OnDemandMemory
|
|
500
|
+
file's triggers will not find it, and that is a deliberate trade for determinism, not a bug.
|
|
501
|
+
|
|
502
|
+
#### When the MCP route pays off
|
|
503
|
+
|
|
504
|
+
The cost that matters is the compiled one, because setup happens once: a working session on a
|
|
505
|
+
compiled workspace holds `recall` for 329 tokens of definitions per turn, or `recall` and `sync`
|
|
506
|
+
for 620 with `--allow-write`. The setup surface (`check` 364, `check` and `optimize` 1,021) is
|
|
507
|
+
paid only by the session that does the first compile, and that session relaunches afterwards.
|
|
508
|
+
The shipped `examples/CLAUDE.md` breaks even at 407 tokens per turn (see [Use it](#use-it)). The
|
|
509
|
+
read-only definitions alone sit under that, but they are not the whole bill: every `recall`
|
|
510
|
+
answer adds the sections it returns, up to 1,500 tokens by default, so one routed answer on that
|
|
511
|
+
example costs more than the compile saved on that turn. The MCP route pays for itself when the
|
|
512
|
+
per-turn prefix saving is bigger than the definition cost plus whatever `recall` returns for a
|
|
513
|
+
typical query - in practice, a memory whose always-loaded prefix sits well above the 2,000-token
|
|
514
|
+
budget, where the alternative `recall` is replacing is an agent reading the whole file on every
|
|
515
|
+
task rather than a routed section. A small memory file is cheaper served the way this README's
|
|
516
|
+
own examples are: compiled once, read directly by the agent, no server attached.
|
|
517
|
+
|
|
518
|
+
#### Single-file compile (`check` + `optimize`, or `doctor` + `plan` + `apply` + `update` under `--profile full`)
|
|
519
|
+
|
|
520
|
+
The CLI's `init`/`init --update`, over MCP, so an agent applies the interface without shelling
|
|
521
|
+
out. `target` is a path relative to the server's root (default: the root itself); it is confined
|
|
522
|
+
to the root the same way `reorganize`'s paths are - no absolute paths, no `..` segments, no
|
|
523
|
+
escaping via a symlink.
|
|
524
|
+
|
|
525
|
+
- `check` never writes; it is registered whenever the root is not yet compiled, with or
|
|
526
|
+
without `--allow-write`. It audits (`doctor`) and, when `target` is a compilable file,
|
|
527
|
+
previews the compile (`plan`) in the same response. Accepts `format: "json"`
|
|
528
|
+
for a machine-readable shape instead of the rendered report.
|
|
529
|
+
- `optimize` is gated behind `--allow-write`, same as `sync`. On a compilable file it refuses a
|
|
530
|
+
compile that would not shrink the always-loaded prefix, refuses a source holding a
|
|
531
|
+
credential-shaped string, and - on an already-compiled target - re-applies it instead of
|
|
532
|
+
refusing outright, saying so in the response (`--profile full`'s `apply` refuses that case
|
|
533
|
+
and points at `update()`; `optimize` just does what `update()` would). `sync` only ever
|
|
534
|
+
re-applies, and only when `doctor`'s MM010 drift check finds something to re-apply or the
|
|
535
|
+
call passes `force: true` (a deliberate recompile, reachable on the compiled surface where
|
|
536
|
+
`optimize` is not registered); it keeps hand edits to AlwaysOnMemory text and OnDemandMemory
|
|
537
|
+
file content, and backs up the pre-update stub first.
|
|
538
|
+
|
|
539
|
+
#### Multi-file reorganization (`optimize` with `ops`, or `scan` + `reorganize` under `--profile full`)
|
|
540
|
+
|
|
541
|
+
`init` compiles one host file. A Claude Code auto-memory folder is a different shape: an
|
|
542
|
+
OnDemandMemory-style list plus dozens of topic files, already multi-file, needing content moved
|
|
543
|
+
*between* files rather than split out of one. `optimize`, pointed at a memory directory, splits
|
|
544
|
+
the work along the same deliberate line the full profile's `scan`/`reorganize` pair always has:
|
|
545
|
+
|
|
546
|
+
- **Facts first, and only facts.** With no `ops`, `optimize` returns `scan`-shape structure:
|
|
547
|
+
frontmatter (`name`/`description`/`metadata.type`), section boundaries, which sections carry
|
|
548
|
+
volatility evidence and why, candidate routing keywords, and blocks duplicated across files. It
|
|
549
|
+
never classifies, merges, or decides on its own.
|
|
550
|
+
- **The calling agent makes the judgment.** Which sections are semantic vs episodic vs
|
|
551
|
+
procedural, what merges, what is stale, where content belongs under O1-O5. That reasoning is
|
|
552
|
+
the agent's, done in the agent's own session, from the facts `optimize` just handed it.
|
|
553
|
+
- **Then `optimize` again, with `ops`, executes.** The same explicit `write_file` / `delete_file`
|
|
554
|
+
/ `rename_file` operations `reorganize` takes, applied the same way.
|
|
555
|
+
|
|
556
|
+
Both take a `target`: a path relative to the server root (default: the root itself), or the
|
|
557
|
+
literal `"auto-memory"`, which resolves to the operator's OS-level Claude Code auto-memory folder
|
|
558
|
+
for this root - a real directory outside the server root, so it only resolves when the server is
|
|
559
|
+
launched with `--include-auto-memory`; without the flag, `"auto-memory"` is refused with a clear
|
|
560
|
+
error rather than silently falling back to the server root.
|
|
561
|
+
|
|
562
|
+
No LLM is embedded in this package and none is called by it. `init`, `doctor` and `scan` stay
|
|
563
|
+
fully deterministic and offline. The intelligence is the agent already driving the tools, not an
|
|
564
|
+
API key this package holds.
|
|
565
|
+
|
|
566
|
+
Safety, non-negotiable and not skippable:
|
|
567
|
+
|
|
568
|
+
- **Every `.md` file under the resolved target is backed up verbatim** to
|
|
569
|
+
`<target>/.minnimemory-backup/<timestamp>/` before the first operation of a plan runs - the only
|
|
570
|
+
files this tool can touch; `node_modules`, `.git`, `dist`, `build` are skipped. For an
|
|
571
|
+
`"auto-memory"` target, the backup lands inside that folder, not the server root.
|
|
572
|
+
- **Every path in a plan is validated before any of it is applied**, so a bad path fails the
|
|
573
|
+
whole plan rather than leaving it half-applied.
|
|
574
|
+
- Paths are confined to the resolved target: absolute paths, `..` segments, symlinked escapes,
|
|
575
|
+
and writes into the backup directory itself are all refused.
|
|
576
|
+
- **Every `write_file`/`delete_file`/`rename_file` path must end in `.md`.** Containment alone
|
|
577
|
+
still let a write land on an extensionless file like `.git/hooks/pre-commit` inside a confined
|
|
578
|
+
root; this operates on memory files, not arbitrary files a confined write could otherwise reach.
|
|
579
|
+
- **Write access is opt-in per launch.** Without `--allow-write` the server is read-only:
|
|
580
|
+
`optimize`/`sync` (basic) and `apply`/`update`/`reorganize` (full) are not registered at all,
|
|
581
|
+
so a client cannot call them.
|
|
582
|
+
|
|
583
|
+
In Claude Code the file-based route is usually enough, because the agent can read an
|
|
584
|
+
OnDemandMemory file directly. MCP tool definitions themselves cost prefix tokens on every turn, which is why the
|
|
585
|
+
server registers only the profile you launch it with; weigh even the default two-tool basic
|
|
586
|
+
profile before registering the server on a repo with only a few OnDemandMemory files.
|
|
587
|
+
|
|
588
|
+
### `minnimemory bench [path]`
|
|
589
|
+
|
|
590
|
+
What the compiled shape costs over a whole session, rather than on one turn.
|
|
591
|
+
|
|
592
|
+
The figure `init` prints is tokens *moved* out of the always-loaded prefix. That is not the same
|
|
593
|
+
as tokens saved. An OnDemandMemory file the agent opens on turn 1 becomes a tool result in the
|
|
594
|
+
conversation and is re-sent on every turn after it, so it costs the rest of the session, not one
|
|
595
|
+
turn. Real savings are the OnDemandMemory files a session never opens, and `bench` is what makes
|
|
596
|
+
that visible.
|
|
597
|
+
|
|
598
|
+
It reports two different kinds of claim, and labels which is which every time.
|
|
599
|
+
|
|
600
|
+
**Bounds. Exact arithmetic, no workload, nothing to assume:**
|
|
601
|
+
|
|
602
|
+
```
|
|
603
|
+
$ npx minnimemory bench --turns 50
|
|
604
|
+
|
|
605
|
+
original memory file 689 tokens per turn
|
|
606
|
+
compiled always-loaded 282 tokens per turn
|
|
607
|
+
OnDemandMemory, all of it 963 tokens, on demand
|
|
608
|
+
|
|
609
|
+
best case, no OnDemandMemory file ever opened
|
|
610
|
+
34,450 -> 14,100 tokens, saves 20,350 (59.1%)
|
|
611
|
+
|
|
612
|
+
worst case, every OnDemandMemory file opened on turn 1 and carried
|
|
613
|
+
34,450 -> 62,250 tokens, costs 27,800 more
|
|
614
|
+
|
|
615
|
+
break-even: 407 tokens of OnDemandMemory can sit in
|
|
616
|
+
context from turn 1 to the end and come out even.
|
|
617
|
+
```
|
|
618
|
+
|
|
619
|
+
The worst case is a real number and it is a loss. On a small memory file a compiled layout can
|
|
620
|
+
cost more than the file it replaced, and this command will say so rather than quietly reporting
|
|
621
|
+
the prefix reduction and stopping there.
|
|
622
|
+
|
|
623
|
+
**A session model. Conditional on a workload you supply:**
|
|
624
|
+
|
|
625
|
+
```bash
|
|
626
|
+
npx minnimemory bench --workload tasks.txt # one task description per line
|
|
627
|
+
```
|
|
628
|
+
|
|
629
|
+
It routes each task at section level with the same deterministic ranker and token cap `recall` ships (`--whole-files` models an agent reading entire files instead), carries every opened section
|
|
630
|
+
forward for the rest of the session, and reports the total, which OnDemandMemory files were never
|
|
631
|
+
opened, what one further turn costs at the steady state, and whether the workload ever breaks even.
|
|
632
|
+
|
|
633
|
+
`minnimemory` ships **no default task list**. A number produced from tasks the tool invented for
|
|
634
|
+
itself would say more about the tool than about your repo, so without `--workload` only the bounds
|
|
635
|
+
are reported.
|
|
636
|
+
|
|
637
|
+
**This is a model, not a measurement, and the output says so.** It compares memory tokens only,
|
|
638
|
+
excluding reasoning, tool results and conversation from both sides, so it is not a session bill.
|
|
639
|
+
Observed usage against a live multi-turn agent is measured by a separate internal harness (the
|
|
640
|
+
`claude-cli` provider) and is the only path allowed to call a number measured.
|
|
641
|
+
|
|
642
|
+
## Profiles
|
|
643
|
+
|
|
644
|
+
`init` embeds a short block of token-discipline guidance into AlwaysOnMemory, because an
|
|
645
|
+
AlwaysOnMemory-plus-OnDemandMemory layout does nothing unless the agent is actually told to route
|
|
646
|
+
rather than read everything.
|
|
647
|
+
|
|
648
|
+
That block costs tokens too, so it is adjustable. Real numbers from the same 689-token
|
|
649
|
+
`examples/CLAUDE.md`:
|
|
650
|
+
|
|
651
|
+
| profile | what it embeds | result on that file |
|
|
652
|
+
|---|---|---|
|
|
653
|
+
| `none` | nothing but the one-line routing note, structure only | 689 -> 276 tokens (**-59.9%**) |
|
|
654
|
+
| `routing` | the rules without which routing does not happen, plus the one that keeps new memory from bypassing them, about 200 tokens | 689 -> 476 tokens (**-30.9%**) |
|
|
655
|
+
| `full` | the complete discipline set, about 690 tokens | 689 -> 969 tokens (**+40.6%**) |
|
|
656
|
+
|
|
657
|
+
`auto`, the default, picks `none` for a source under the budget and `routing` at or over it, so
|
|
658
|
+
a small file never pays for guidance it does not need. On this file `auto` chose `none`.
|
|
659
|
+
|
|
660
|
+
`full` is a net loss on a file this size, and `init` says so rather than reporting a fake win:
|
|
661
|
+
|
|
662
|
+
```
|
|
663
|
+
warning: this compile does not shrink the always-loaded prefix.
|
|
664
|
+
The source is small enough that the added instruction block costs more than routing saves.
|
|
665
|
+
```
|
|
666
|
+
|
|
667
|
+
That is why `auto` never picks `full` on its own: it embeds `routing`, the smallest set that
|
|
668
|
+
makes the structure work, and only once the source is large enough to pay for it. `full` is an
|
|
669
|
+
explicit opt-in that earns its keep on large memory files, and on high-volume agents where
|
|
670
|
+
output-token discipline matters more than prefix size.
|
|
671
|
+
|
|
672
|
+
## Keeping it current
|
|
673
|
+
|
|
674
|
+
After the first compile, the source of truth is the stub's own AlwaysOnMemory text plus the
|
|
675
|
+
OnDemandMemory files on disk. Edit either freely.
|
|
676
|
+
|
|
677
|
+
**Coming back to a memory you optimised earlier? Start with `doctor`.** On a compiled workspace
|
|
678
|
+
it reports whether anything has drifted from what `init` wrote, and names the command to fix it
|
|
679
|
+
(over MCP, `check` does the same on a compiled target, and on the basic profile's compiled
|
|
680
|
+
surface, where `check` is absent, `recall` reports drift itself when it rebuilds its index):
|
|
681
|
+
|
|
682
|
+
```
|
|
683
|
+
$ npx minnimemory doctor
|
|
684
|
+
|
|
685
|
+
...findings...
|
|
686
|
+
|
|
687
|
+
compiled workspace: 4 OnDemandMemory files in .minnimemory/
|
|
688
|
+
drift: 1 file changed since init (MM010). Re-apply with: minnimemory init --update
|
|
689
|
+
```
|
|
690
|
+
|
|
691
|
+
When there is nothing to do it says so plainly ("no drift since init ... Nothing to re-apply"),
|
|
692
|
+
so a re-run costs you one second and no thinking. `doctor` never writes. To apply:
|
|
693
|
+
|
|
694
|
+
```bash
|
|
695
|
+
npx minnimemory init --update
|
|
696
|
+
```
|
|
697
|
+
|
|
698
|
+
It strips the generated parts of the stub, routes anything you appended to it the same way a
|
|
699
|
+
fresh section would be routed, keeps every OnDemandMemory file exactly as it is on disk,
|
|
700
|
+
recomputes triggers and hashes, and rewrites the OnDemandMemory list, manifest, and stub. The previous stub is kept in
|
|
701
|
+
`.minnimemory/previous/`. Nothing is merged against the original file, because the original is
|
|
702
|
+
no longer what anyone edits.
|
|
703
|
+
|
|
704
|
+
## What doctor counts as always loaded
|
|
705
|
+
|
|
706
|
+
The host file, plus everything it pulls in with Claude Code's `@path` import syntax, resolved
|
|
707
|
+
the way Claude Code resolves it: relative to the importing file, recursive to a depth of five,
|
|
708
|
+
ignoring anything inside code spans and fenced blocks. If the host file is only a pointer at
|
|
709
|
+
`AGENTS.md`, the pointer and its target are both counted. For a project that Claude Code has an
|
|
710
|
+
auto-memory folder for, `MEMORY.md` from that folder counts too; its topic files are read on
|
|
711
|
+
demand and are audited but not charged. This is the CLI's `doctor`/`bench` behavior; over MCP,
|
|
712
|
+
`scan`/`doctor`/`plan` leave the auto-memory folder out unless launched with
|
|
713
|
+
`--follow-external-imports` (see [Security](#security)).
|
|
714
|
+
|
|
715
|
+
## Kinds, episodic JSON, and the write protocol
|
|
716
|
+
|
|
717
|
+
Every OnDemandMemory file carries a **kind** in the manifest (O2 of the interface): `semantic`
|
|
718
|
+
(facts true now, edited in place), `episodic` (things that happened, append-only: changelogs,
|
|
719
|
+
logs) or `procedural` (a learned rule). The kind is decided per section from its heading and
|
|
720
|
+
body, so a `project` file that blends current state with a changelog compiles into separate
|
|
721
|
+
OnDemandMemory files, and an episodic section is never merged into a neighbour. `scan` reports the kind for every section
|
|
722
|
+
of a memory directory too.
|
|
723
|
+
|
|
724
|
+
With **`init --episodic-json`** (O3, opt in), an episodic OnDemandMemory file is written as
|
|
725
|
+
`OnDemandMemory/<name>.json`: the heading, any preamble, and one entry per dated bullet with its
|
|
726
|
+
continuation lines, all verbatim, so the markdown rebuilds byte for byte. `recall`, `outline`,
|
|
727
|
+
`bench`, `doctor` and `init --update` read JSON OnDemandMemory files as the markdown they stand for. The
|
|
728
|
+
point is Anthropic's documented finding that a model is less likely to rewrite or summarize
|
|
729
|
+
JSON it was only meant to append to. Off by default until it has run in a second host.
|
|
730
|
+
|
|
731
|
+
`init` also emits one OnDemandMemory file the source did not contain:
|
|
732
|
+
**`OnDemandMemory/memory_write_protocol.md`**, the eleven-item memory write protocol (O9),
|
|
733
|
+
routed by the OnDemandMemory list under `memory, remember, save, write, note, changelog`, so the
|
|
734
|
+
agent opens it when a task is about saving memory and is never charged for it otherwise.
|
|
735
|
+
`init --update` re-emits it, so a protocol change reaches every workspace on the next update. It
|
|
736
|
+
accounts for one OnDemandMemory list line in the prefix figures above.
|
|
737
|
+
|
|
738
|
+
OnDemandMemory paths are relative to the file the list sits in: `OnDemandMemory/<name>.md`
|
|
739
|
+
inside `.minnimemory/AlwaysOnMemory.md`, and `.minnimemory/OnDemandMemory/<name>.md` in the stub
|
|
740
|
+
the agent actually reads from. The first session measurement (2026-09-05) caught an agent reading
|
|
741
|
+
the bare path from the stub, missing, and giving up; the stub has named the full path since.
|
|
742
|
+
|
|
743
|
+
## Security
|
|
744
|
+
|
|
745
|
+
- `doctor`, `init`, and `mcp` make no network calls and never run a subprocess.
|
|
746
|
+
- `init` refuses a symlinked source and refuses to write through a symlink, so a cloned repo
|
|
747
|
+
cannot point `CLAUDE.md` or an OnDemandMemory file at a file outside the checkout.
|
|
748
|
+
- `init` refuses `--write` when the source holds a credential-shaped string, unless you pass
|
|
749
|
+
`--allow-secrets`. The verbatim backup in `.minnimemory/original/` ships with a `.gitignore`
|
|
750
|
+
so it stays local.
|
|
751
|
+
- The MCP server validates `manifest.json` and only serves files that resolve inside
|
|
752
|
+
`.minnimemory/OnDemandMemory/`, whatever the manifest says.
|
|
753
|
+
- `doctor --json` includes the absolute workspace path and the headings of your auto-memory
|
|
754
|
+
files. Read it before pasting it somewhere public.
|
|
755
|
+
|
|
756
|
+
## Rule set
|
|
757
|
+
|
|
758
|
+
Moved to [`docs/DESIGN.md` section 12](docs/DESIGN.md#12-rule-set-for-doctor) (2026-09-16): the
|
|
759
|
+
full, current MM001-MM010 table, one copy instead of two that had drifted apart.
|
|
760
|
+
|
|
761
|
+
## Design principles
|
|
762
|
+
|
|
763
|
+
Moved to [`docs/DESIGN.md` section 1.1](docs/DESIGN.md#11-design-principles) (2026-09-16):
|
|
764
|
+
deterministic, lossless, volatile content never in AlwaysOnMemory, OnDemandMemory files emitted
|
|
765
|
+
verbatim, the original never destroyed.
|
|
766
|
+
|
|
767
|
+
## What it does not do
|
|
768
|
+
|
|
769
|
+
See [`docs/DESIGN.md` section 2, "What it is not"](docs/DESIGN.md#2-what-it-is-not): not a
|
|
770
|
+
conversation compressor, not a RAG index over your code, not a proxy or model router, not a
|
|
771
|
+
summariser. One addition specific to this package: **not a compiler for a memory routing list.**
|
|
772
|
+
`init` refuses `MEMORY.md`, or any file that is mostly links to topic files beside it, with no
|
|
773
|
+
override - compiling a routing list re-files its hook lines as content and orphans every topic
|
|
774
|
+
file. A memory directory is `check`'s job to audit and `optimize`'s (`mcp --allow-write`: facts
|
|
775
|
+
first, then `ops`, backup first) to restructure. Auditing alone needs no write access - `check`
|
|
776
|
+
is registered by default.
|
|
777
|
+
|
|
778
|
+
## Measurement policy
|
|
779
|
+
|
|
780
|
+
Moved to [`docs/DESIGN.md` section 8](docs/DESIGN.md#8-measurement-and-honesty-rules) (2026-09-16),
|
|
781
|
+
including the live-run measurement table (section 8.1) and the tokenizer calibration figures.
|
|
782
|
+
|
|
783
|
+
## FAQ
|
|
784
|
+
|
|
785
|
+
**Does it send my code or memory files anywhere?**
|
|
786
|
+
No. `doctor` and `init` make no network calls at all. There is no telemetry and no API key.
|
|
787
|
+
|
|
788
|
+
**Will it destroy my `CLAUDE.md`?**
|
|
789
|
+
No. `init` is dry run by default. With `--write`, your original is copied verbatim to
|
|
790
|
+
`.minnimemory/original/` before anything is modified.
|
|
791
|
+
|
|
792
|
+
**Why not just ask an LLM to summarise my memory file?**
|
|
793
|
+
Lossy, non-reproducible, and it costs tokens to save tokens. See
|
|
794
|
+
[Design principles](#design-principles).
|
|
795
|
+
|
|
796
|
+
**Does this fight with prompt caching?**
|
|
797
|
+
The opposite. Splitting volatile content out of the always-loaded prefix is precisely what keeps
|
|
798
|
+
the cache valid across sessions.
|
|
799
|
+
|
|
800
|
+
**What if the agent ignores the OnDemandMemory list and never opens an OnDemandMemory file?**
|
|
801
|
+
That is the real failure mode of routing, which is why OnDemandMemory list rows are written as
|
|
802
|
+
hooks rather than summaries, why `auto` embeds the `routing` profile as soon as the source is at
|
|
803
|
+
the budget, and why `MM004` flags OnDemandMemory files nothing reachable ever references.
|
|
804
|
+
|
|
805
|
+
**Do I have to keep the generated structure?**
|
|
806
|
+
No. It is plain markdown. Edit it, restructure it, delete OnDemandMemory files. `doctor` checks budgets and
|
|
807
|
+
routing, not prose style.
|
|
808
|
+
|
|
809
|
+
**Does it work with Cursor or Codex?**
|
|
810
|
+
`doctor` reads `AGENTS.md` and `.cursorrules` as well as `CLAUDE.md`. Only the Claude Code path
|
|
811
|
+
has been exercised end to end, and this README will not claim otherwise until that changes.
|
|
812
|
+
|
|
813
|
+
## Roadmap
|
|
814
|
+
|
|
815
|
+
Moved to [`docs/DESIGN.md` section 11, "Milestones"](docs/DESIGN.md#11-milestones) (2026-09-16):
|
|
816
|
+
merged with that section's own milestone table, which had drifted out of date. Publish state
|
|
817
|
+
lives there too (as of 2026-09-16 the package is prepared but not yet on npm).
|
|
818
|
+
|
|
819
|
+
## License
|
|
820
|
+
|
|
821
|
+
Proprietary, free to use. Copyright MinniAI. You may install and run the package, unmodified,
|
|
822
|
+
for your own use, including at work. You may not redistribute, modify or reverse engineer it. The
|
|
823
|
+
files it reads and writes on your machine stay yours, and it sends nothing off your machine. Full
|
|
824
|
+
terms in the `LICENSE` file shipped with the package.
|