nexusmem 0.10.4 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,562 +1,605 @@
1
- # NexusMem
2
-
3
- [![CI](https://github.com/yaminbkk/NexusMem/actions/workflows/ci.yml/badge.svg)](https://github.com/yaminbkk/NexusMem/actions/workflows/ci.yml)
4
- [![npm](https://img.shields.io/npm/v/nexusmem)](https://www.npmjs.com/package/nexusmem)
5
- [![npm downloads](https://img.shields.io/npm/dm/nexusmem)](https://www.npmjs.com/package/nexusmem)
6
- [![License: MIT](https://img.shields.io/badge/license-MIT-informational)](LICENSE)
7
- ![Node](https://img.shields.io/badge/node-%3E%3D22-brightgreen)
8
- [![yaminbkk/NexusMem MCP server](https://glama.ai/mcp/servers/yaminbkk/NexusMem/badges/score.svg)](https://glama.ai/mcp/servers/yaminbkk/NexusMem)
9
- [![Listed on AiList](https://hifriendbot.com/ai-list/badge/nexusmem.svg)](https://hifriendbot.com/ai-list/nexusmem/)
10
-
11
- ![NexusMem: init, sync --github, and a query against this repo's own history — surfacing a real issue, the PR that closed it, and the commits it shipped](docs/demo.gif)
12
-
13
- Your coding agent can read `git log`. It cannot read the four things you tried last Tuesday that
14
- didn't work.
15
-
16
- NexusMem records what actually happened on your machine (shell commands and their exit codes, git
17
- history down to the patch of each changed file, project docs, optionally your assistant transcripts)
18
- into a local SQLite database, and
19
- serves back a ranked, token-budgeted slice of it on demand. Everything stays on disk. No account, no
20
- cloud, no telemetry.
21
-
22
- The shell history is the part worth caring about. Git tells an agent what shipped. Shell history
23
- tells it what was attempted, in what order, and which commands exited non-zero. That information
24
- exists nowhere else, and it disappears when your terminal scrollback rolls over.
25
-
26
- ![How NexusMem works: git commits, shell commands with their real exit codes, docs, and GitHub threads flow into a local SQLite index, which returns a ranked, token-budgeted slice of that history so your coding agent answers with a real citation instead of a guess](docs/how-it-works.svg)
27
-
28
- **Contents:** [Try it](#try-it) · [Exact shell capture](#optional-exact-shell-capture) ·
29
- [Failure → fix chains](#failure--fix-chains-opt-in) · [How retrieval works](#how-retrieval-works) ·
30
- [Session summaries](#session-summaries-optional-local-model) ·
31
- [GitHub issues & PRs](#github-issues--prs-optional) · [Use it from an agent](#use-it-from-an-agent)
32
- · [What it costs you](#what-it-costs-you) · [Staleness & provenance](#staleness--provenance) ·
33
- [Where it breaks](#where-it-breaks) · [Commands](#commands) · [Cross-project recall](#recall-across-projects)
34
- · [On disk](#on-disk) · [Development](#development)
35
-
36
- ## Try it
37
-
38
- From inside any git repository:
39
-
40
- ```
41
- npx nexusmem init
42
- npx nexusmem sync
43
- ```
44
-
45
- Then ask it something. Real output from this repository, top 2 of 5 hits:
46
-
47
- ```
48
- $ nexusmem query "windows spawn failure"
49
-
50
- Relevant history for: windows spawn failure
51
-
52
- - 2026-08-09 [observed] fix: distinguish a failed git spawn from "not a git repository"
53
- readRepoInfo collapsed three unrelated failures into one error: git running and reporting
54
- the path is not a work tree, git not being installed, and the process failing to spawn at
55
- all. Dogfooding hit the third case in two separate sessions...
56
- - 2026-08-09 [authored] README.md — Before a tagged release
57
- - [ ] Retry on transient process-spawn failures on Windows
58
- ```
59
-
60
- `[observed]`/`[authored]` is the provenance tag (see [Staleness & provenance](#staleness--provenance))
61
- — a commit is a directly observed event, a doc section is a written claim that could go stale.
62
-
63
- A commit and a docs section, ranked against each other, inside whatever token budget you gave it.
64
- Nothing was summarized by a model on the way out; the ranker just decided what not to send. (One
65
- optional source, session summaries, does run a local model — but at ingest time, never on the way
66
- out. What you query is always stored text.)
67
-
68
- For a sense of what actually accumulates, here is `nexusmem status` on this repo after two days:
69
-
70
- ```
71
- 527 node(s) 2026-08-08 .. 2026-08-09
72
- 321 shell_command
73
- 130 conversation_turn
74
- 60 doc_section
75
- 16 git_commit
76
- ```
77
-
78
- Sixteen commits. Three hundred and twenty-one shell commands. The commits were already retrievable
79
- by any agent with a terminal. The rest was not.
80
-
81
- That `conversation_turn` row only appears because this corpus was synced with `--conversation`.
82
- Assistant transcripts are the one source that is off by default and stays off until you opt in, since
83
- they are the likeliest place for a pasted credential to be sitting. A default install indexes git
84
- commits, their diffs, shell and docs.
85
-
86
- Requirements: Node 22 or newer, and git. Node 20 will not work, because `better-sqlite3` ships no
87
- prebuilt binary for it and Node 20 went end-of-life in April 2026. Ollama is optional and only
88
- affects semantic search (see below).
89
-
90
- ## Optional: exact shell capture
91
-
92
- Scraped history files (PSReadLine, `.bash_history`, `.zsh_history`) give you command text and not
93
- much else. The hook gives you working directory, exit code and a real timestamp:
94
-
95
- ```bash
96
- nexusmem hook install
97
- ```
98
-
99
- It wraps your existing PowerShell prompt rather than replacing it, is idempotent, and
100
- `nexusmem hook remove` undoes it cleanly.
101
-
102
- Exit codes are what make this worth installing. A failed command is a stronger signal than a
103
- successful one, and without the hook there is no way to tell them apart.
104
-
105
- ## Failure → fix chains (opt-in)
106
-
107
- ```bash
108
- nexusmem sync --link-failures
109
- ```
110
-
111
- After a normal sync, this walks every failed `shell_command` (non-zero exit code) and looks for
112
- whatever later resolved it, using two independent heuristics: a later command in the same project
113
- and working directory, exact same normalized text, that exited `0` within 24h (**same-command
114
- retry**); and, separately, the best full-text match among nearby conversation turns or session
115
- summaries, requiring every significant word of the failing command to appear, not just one
116
- (**conversation bridge**). A failure can be linked by either, both, or neither.
117
-
118
- Both links are surfaced in query results. The conversation-bridge heuristic originally matched on
119
- any shared word, and dogfooding against this repo's own real history found it wrong on roughly half
120
- its links — a shared word as generic as "npm" was enough to link an unrelated discussion. Requiring
121
- every significant word fixed that: re-dogfooded against the same corpus, every resulting link (the
122
- full set produced, not a sample) checked out correct on manual review of the full text, not just the
123
- summary.
124
-
125
- When a linked failure appears in a result set, its fix rides along immediately after it, inheriting
126
- the failure's own relevance score rather than needing to match the query on its own merits. That is
127
- the point: a query about why something failed shouldn't need to separately guess the words used in
128
- whatever fixed it. This works across projects too — `query --all-projects` chains a failure to its
129
- fix using whichever project's own database recorded the link, since links are always local to the
130
- project they were found in.
131
-
132
- ```
133
- $ nexusmem query "why did npm whoami fail"
134
-
135
- - 2026-08-12 [observed] shell: npm whoami (exit 1)
136
- - 2026-08-12 [observed] shell: npm login (exit 0) -- linked as the fix
137
- ```
138
-
139
- ## How retrieval works
140
-
141
- Every source normalizes to the same `MemoryNode` shape, so a commit, a shell command and a docs
142
- section compete on equal terms. Retrieval runs BM25 over FTS5 and, if an embedding model is
143
- reachable, a vector search over `sqlite-vec`, fused with Reciprocal Rank Fusion on rank *position*
144
- only, never raw scores — a BM25 cost and a vector distance live on unrelated, unbounded scales, and
145
- position is the only thing they agree on.
146
-
147
- Ranking then multiplies three factors:
148
-
149
- ```
150
- score = relevance × signal^0.215 × recency^0.288
151
- ```
152
-
153
- `relevance` comes from the query. `signal` (a `fix:` commit outranks a `chore:`; a failed command
154
- outranks a successful one) and `recency` are priors that hold before any query exists. Each factor is
155
- floored into `[floor, 1]` rather than `[0, 1]`, so one weak dimension can't zero out a strong match.
156
-
157
- The exponents bound how far signal and recency, *together*, may overturn relevance: at most a 2× gap
158
- across their whole range, applied jointly rather than per-prior. That's deliberate — the score
159
- *multiplies* the two priors, so capping each at 2× separately still let the pair overturn 4×, and
160
- that hit hardest on fresh, high-signal commits made during an active working day. The bug that
161
- exposed this: two unrelated same-day `fix:` commits outranked the docs section that actually answered
162
- the query. See [`retrieval/rank.ts`](src/retrieval/rank.ts) for the full derivation.
163
-
164
- Without Ollama, vector search is skipped and you get BM25 only — fully supported, not a degraded
165
- state; `sync` and `query` both succeed and simply do less.
166
-
167
- ## Session summaries (optional, local model)
168
-
169
- With `sources.session.enabled`, each finished session becomes one distilled node next to the raw
170
- exchanges — what was decided and why, rather than forty individual turns. It runs a local Ollama
171
- chat model (`qwen2.5:3b` by default); nothing is downloaded automatically and nothing leaves the
172
- machine.
173
-
174
- ```bash
175
- nexusmem scan-session --dry-run
176
- ```
177
-
178
- That prints the exact prompt a session would produce, after redaction and budget trimming, without
179
- calling the model.
180
-
181
- Three things bound the cost. A session is only summarized once it has been quiet for
182
- `settleMinutes` (default 30), so a session in progress is not re-summarized on every sync. The
183
- prompt is hashed, and an unchanged hash skips the model entirely — on this repo a steady-state sync
184
- of 14 summarized sessions takes 0.25s and makes no model calls. And `maxSessions` (default 10) caps
185
- how many reach the model per run; the rest are reported as queued and picked up next sync.
186
-
187
- **What it is actually like, measured on 14 real sessions with `qwen2.5:3b`.** The summaries
188
- themselves are good: decisions with their reasons, in the shape the prompt asks for. Titles are less
189
- reliable — the model returned a usable one about a third of the time, and otherwise produced a
190
- conversational preamble, a stray bullet, or a bare "Summary of the Session". Those are rejected and
191
- the title falls back to the first line of the question that opened the session, which is always
192
- specific even when it is not elegant. Compliance was worst on long sessions and on transcripts not
193
- in English. A larger model (`qwen2.5:7b`) is the lever if the titles matter to you; set
194
- `sources.session.model`.
195
-
196
- ## GitHub issues & PRs (optional)
197
-
198
- With `sources.github.enabled`, each issue and PR on this repo's github.com remote becomes one node
199
- (title, opening post and every comment, folded together) alongside the raw discourse a discussion
200
- already leaves in shell/conversation history. Off by default — not for a sensitivity reason, but
201
- because it's the first source with a real external dependency: it reads via the `gh` CLI, so it
202
- needs `gh` installed and authenticated (`gh auth login`), and it makes live network calls instead of
203
- only reading what's already on disk. A repo with no github.com remote, or an unauthenticated `gh`,
204
- is a silent no-op either way.
205
-
206
- ```bash
207
- nexusmem scan-github
208
- ```
209
-
210
- previews the nodes a sync would produce, same as the other `scan-*` commands. `maxThreads` (default
211
- 100) and `maxCommentsPerThread` (default 100) bound one sync's cost; `since` is tracked as its own
212
- cursor, so a repeat sync only re-reads threads that changed. Dogfooded against this repo's own 14
213
- real issues/PRs: ingest took under a second, and a query for "labelled retrieval regression corpus"
214
- correctly ranked issue #8 — the one that asked for it — first.
215
-
216
- ## Use it from an agent
217
-
218
- ```json
219
- {
220
- "mcpServers": {
221
- "nexusmem": {
222
- "command": "npx",
223
- "args": ["-y", "nexusmem", "mcp"]
224
- }
225
- }
226
- }
227
- ```
228
-
229
- Three tools over stdio: `search_memory` returns the packed context block, `sync_project` ingests, and
230
- `get_status` reports what is currently remembered. Each takes an explicit `projectRoot`, because an
231
- MCP tool call carries no shell working directory. `sync_project` runs `init` for you if the
232
- repository has not been set up.
233
-
234
- ## What it costs you
235
-
236
- Two numbers get conflated in tools like this, so they are kept apart here.
237
-
238
- **Packer efficiency** is how much the ranker trims from its own candidate set. On this repository's
239
- corpus it runs 81–84%. It is useful for tuning the ranker and useless as a claim about your bill,
240
- because the baseline is hypothetical: without NexusMem those candidates were never going into your
241
- context window in the first place.
242
-
243
- **End-to-end saving** compares the packed context NexusMem actually sends against reading, in full,
244
- the same files its own ranking identified as relevant to the query. Measured with
245
- [`scripts/benchmark.ts`](scripts/benchmark.ts) (`npm run bench`), which anyone who clones this repo
246
- and points it at a synced corpus can re-run from scratch:
247
-
248
- | Corpus | Commits | Query set | vs. full file content | vs. `git log -p` on those files |
249
- | --- | --- | --- | --- | --- |
250
- | This repo | 62 | 16 real prompts, verbatim from this project's own history | 95% (median 94%) | 98% (median 97%) |
251
- | [`vitejs/vite`](https://github.com/vitejs/vite) | 9,567 | 16, mechanically sampled — see below | 99% (median 98%) | ~100% (median ~100%) — see caveat |
252
-
253
- Both clear the original >70% target ("cut API token spend versus sending full context"), and the vite
254
- run is the first measurement at the scale that target was always described as applying to.
255
-
256
- **Read the methodology before quoting either number — it's a narrower claim than it looks:**
257
-
258
- - **Graded against NexusMem's own ranking**, not an outside answer key: the file set is whichever
259
- files the packed nodes for that query touch. This measures what the pack step saves once retrieval
260
- already picked a candidate set; it doesn't independently verify that set was the right one.
261
- - **Query sets are mechanical, not cherry-picked** (see `scripts/benchmark.ts`): vite's is an even
262
- sample of well-explained `fix`/`feat`/`perf`/`refactor` commits plus rationale-bearing doc headings;
263
- this repo's reuses real historical prompts verbatim, several of which are broad task instructions
264
- rather than narrow questions — part of why its number sits below vite's.
265
- - **`git log -p` baselines can be enormous** — one vite query's baseline hit 7.5M tokens because a
266
- file in its resolved set has that much history. At that scale, "just read the file's history
267
- instead" stops being a viable alternative at all.
268
- - **Supersedes the old ~40% figure**, which was hand-tallied from two hand-picked queries against
269
- this repo alone with an unstated baseline. Not wrong, just underspecified — this replaces it with a
270
- stated method and a script that reproduces it.
271
-
272
- One thing that is not a percentage: shell commands and conversation turns have no cheap `grep`
273
- equivalent. Without something recording them, they are gone, not merely more expensive to find.
274
-
275
- For how these numbers compare to a similar tool's own claims, see
276
- [`docs/competitor-comparison.md`](docs/competitor-comparison.md) (vs. projectmem) and
277
- [`docs/competitor-comparison-yesmem.md`](docs/competitor-comparison-yesmem.md) (vs. YesMem, including
278
- native Windows support vs. its documented WSL2 requirement).
279
-
280
- Latency on a ~530-node corpus, warm, p50 over 10 runs:
281
-
282
- | Operation | |
283
- | --- | --- |
284
- | BM25 retrieval (FTS5) | ~1.1 ms |
285
- | Vector KNN (`sqlite-vec`) | ~3.2 ms |
286
- | Fuse, rank, pack | ~0.6 ms |
287
- | Query embedding (local Ollama) | ~55–77 ms |
288
- | **End-to-end hybrid** | **~56 ms** |
289
-
290
- All the SQLite work totals about 5 ms. The embedding call is the only thing on this path worth
291
- optimizing, and it is somebody else's process.
292
-
293
- ## Staleness & provenance
294
-
295
- Two things a memory layer needs and this one only partly has: a way to tell an observed fact from a
296
- guess, and a way to retire a conclusion once something contradicts it. This section is what exists
297
- and what doesn't.
298
-
299
- Every node carries a `provenance`, a four-tier trust hierarchy set once per collector at ingest
300
- time: `observed` (a commit that landed, a shell command's real exit code) > `authored` (a doc
301
- section — a human's own written claim) > `recorded` (a conversation turn — verbatim, but talk about
302
- events rather than the events) > `derived` (a session summary — a model's distillation). The tier is
303
- shown as a tag on every query result and decays retrieval weight — the lower the trust, the faster a
304
- node fades from ranking as it ages. The ordering is the design claim; the exact decay ratios are
305
- judgment calls, not measured optima.
306
-
307
- ```bash
308
- nexusmem stale
309
- ```
310
-
311
- Lists non-`observed` nodes old enough (45+ days by default) that nothing has confirmed they still
312
- hold — a heuristic on age and provenance, not on content. It writes nothing; you decide which
313
- candidates are actually wrong. Any candidate the SLM has already flagged (see below) is decorated
314
- with its standing `likely superseded by` suggestion — reading those costs nothing, so the plain
315
- command stays instant and offline.
316
-
317
- ```bash
318
- nexusmem mark-stale <oldNodeId> --supersedes <newNodeId>
319
- ```
320
-
321
- Links `newNodeId` as the replacement for `oldNodeId`. The ranker down-weights the old node from then
322
- on (it stays queryable, just usually loses to its replacement) — nothing is deleted, unlike `forget`.
323
-
324
- ```bash
325
- nexusmem stale --check-contradictions
326
- ```
327
-
328
- For each candidate, finds the most similar newer node (local embedding search) and asks a local SLM
329
- (Ollama, `qwen2.5:3b` by default) whether it actually contradicts the older one — real content
330
- comparison, not just age. A match is printed as `likely superseded by <id> <title> -- <reason>`
331
- under the candidate. Every judgment (either verdict) is memoized, so a judged pair is never sent to
332
- the model again; nothing else is written — `supersedes` stays yours to set via `mark-stale`.
333
-
334
- **This also runs automatically during `sync`** — at most 3 new judgments per run (configurable via
335
- the `contradictions` block in `.nexusmem/config.json`; set `autoCheck: false` to turn it off), only
336
- when the embedding provider was reachable anyway, and free on repeat syncs thanks to the
337
- memoization. New and open suggestions show up in the sync summary, `nexusmem status` (a `flagged`
338
- line), and plain `nexusmem stale`.
339
-
340
- **What this doesn't do:** it is one small model's yes/no judgment on one older/newer pair, not a
341
- verified fact — treat a match as a lead to check, not a conclusion. It also only ever compares a
342
- candidate against nodes *found by embedding similarity*; a contradiction from an unrelated-sounding
343
- node would never surface. Comprehensive contradiction detection (not just for the pair the vector
344
- search happens to surface) is still an open problem, and nothing here supersedes a node on its own.
345
-
346
- `provenance` is a separate question from `trust_state`: provenance says where a claim came from,
347
- never whether anyone checked it.
348
-
349
- ```bash
350
- nexusmem review <nodeId> --verify
351
- nexusmem review <nodeId> --reject
352
- ```
353
-
354
- Records your own verdict on one node, independent of the SLM contradiction checker above (which only
355
- ever writes a suggestion, never a verdict). `--reject` down-weights the node in ranking — same
356
- demote-not-delete rule as `mark-stale`, it stays queryable, just usually loses to better matches —
357
- and both verdicts are shown as a `[verified]`/`[rejected]` tag on every query result that returns the
358
- node afterward. `--verify` is a label only; it does not boost ranking. Every node starts `candidate`
359
- (untagged) until reviewed, and a re-sync never overwrites a verdict once one is set.
360
-
361
- ## Where it breaks
362
-
363
- - **Shell history without the hook is unscoped.** Scraped history has no directory context, so it is
364
- attributed to whichever repository you ran `sync` from. Bounded to a tail window, and an
365
- approximation rather than a guarantee.
366
- - **Japanese and Chinese depend on the vector pass.** FTS5's `unicode61` tokenizer splits on
367
- whitespace, so languages without space boundaries get no useful BM25 recall.
368
- - **Rebasing strands nodes.** Rewritten history leaves nodes for unreachable commits. They describe
369
- real events so they are not wrong, but a targeted prune does not exist yet. `sync --rebuild`
370
- forces a clean re-scan.
371
- - **Multi-line PowerShell input is read as separate commands.** A function typed across several lines
372
- at the prompt is not reconstructed.
373
- - **Scrape-fallback ids drift** if the history file is trimmed from the front between syncs.
374
- Installing the hook fixes this.
375
- - **Session-summary titles depend on the model following instructions**, and a 3B model often does
376
- not. The fallback keeps them specific rather than generic, but see the section above for what to
377
- expect.
378
- - **Changing the embedding model re-embeds everything.** Vectors from two models are not comparable
379
- and `nodes_vec` records no per-row provenance, so `sync` drops the lot and rebuilds rather than
380
- ranking across a mixture. It says so when it happens. Nodes are untouched and BM25 keeps working
381
- throughout.
382
- - **Diff indexing is bounded, and deliberately lossy.** A first sync indexes the patches of the most
383
- recent 200 commits (later syncs only walk `cursor..HEAD`); merge commits contribute none, since
384
- their patch exists only in a combined format this parser does not read; and binaries, lockfiles and
385
- build output are skipped so a dependency bump cannot bury the corpus. All of it is still recorded
386
- as a `git_commit` node. A patch longer than `limits.maxBodyChars` is truncated, so the tail of a
387
- very large change is not indexed. The caps live under `sources.diff` in `config.json`.
388
- - **Cross-project recall favours breadth.** Each repository's hits are fused by rank, so a project
389
- whose best match is mediocre still contributes a rank-1 item, and rank 1 is worth the same in
390
- every list. Adding a repository that has little to say about your question still pushes a few of
391
- its results into the budget. Signal, recency and the budget are what hold that in check; there is
392
- no per-project quality weight.
393
- - **The project registry is an index, not a source of truth.** It can point at a database that has
394
- moved or been deleted; those are reported and skipped, never silently pruned, because an
395
- unmounted drive is not a deleted project.
396
- - **Conversation chunking is unevaluated.** Splitting long replies at heading boundaries measurably
397
- helped, but it has never been tested systematically.
398
- - **A chunked node's sibling count in one result is capped, not tuned.** `conversation_turn` and
399
- `doc_section` both split one reply or file into several nodes; at most 2 of them may appear
400
- together in a packed result. Found live: a query for "token" returned 9 of its top 12 hits as
401
- different pieces of one heavily-sectioned reply, crowding out the node that actually answered it.
402
- The cap of 2 is a judgement call, not a measured optimum, same as the ranking priors' budget above.
403
- - **The size of the prior budget is a judgement call, not a measured optimum.** Priors are now
404
- bounded jointly rather than one at a time, which closed a real 4× hole (see the ranking section),
405
- but the 2× budget itself has never been tuned against a labelled relevance set — there isn't one.
406
- It is a defensible constant, not a result. What is measured is the direction: on four real queries
407
- against this repo's own memory, switching to the joint cap moved the section that answered the
408
- question up in three of them (the rationale section for "why BM25 before vector search" went from
409
- rank 4 to rank 1) and displaced no query's correct top hit.
410
- - **`forget` is per-repository, not global.** Its deny-list lives in the one `.nexusmem/memory.db` it
411
- ran against (plus that repo's stale prior identities, same scope `--prune-source` already uses). A
412
- value that leaked into shell history from several repositories needs `forget` run once per repo —
413
- there is no shared, machine-wide deny-list across every project you have synced.
414
- - **A deny-list doesn't survive a clone or restore on its own.** `.nexusmem/` is gitignored by design,
415
- so `deny_list` never travels with `git clone`/`git push` — while git history itself, the thing a
416
- fresh `sync` re-derives from, is fully portable and copied by every clone. A teammate's fresh
417
- checkout, a new machine, or a restored backup starts with zero protection: the forgotten value comes
418
- right back on the first sync. Confirmed live 2026-08-17, not just a theoretical read of the code.
419
- `forget --export <path>` / `forget --import <path>` close this: export writes the active entries to
420
- a plaintext JSON file you move through a channel you control (never git — the file is exactly as
421
- sensitive as the value it holds), and import re-applies them in the new checkout, deleting any
422
- copies that already synced back in. It is deliberately manual, not automatic on every `sync`.
423
-
424
- ## Commands
425
-
426
- `init`, `sync`, `query <text>` (add `--as-of <date>` for a bi-temporal read, see below), `status` (add
427
- `--share` for a plain-text summary worth pasting somewhere), `projects`, `mcp`, `forget <value>`,
428
- `stale` (add `--check-contradictions` for a local-SLM content check, see above), `mark-stale
429
- <nodeId> --supersedes <newNodeId>`, `review <nodeId> --verify|--reject` (record a human verdict on
430
- one node, see above), `precheck` (advisory — warns about staged files with an unresolved past
431
- failure or high recent churn; exits 0 unless `--strict`), `hook install|remove|status` (the
432
- PowerShell exit-code hook), `hook git install|remove|status` (a git pre-commit hook that runs
433
- `precheck` before each commit), and `hook git-post install|remove|status` (a git post-commit hook
434
- that runs a full `sync`, including embedding, in the background after each commit — detached, so it
435
- never makes `git commit` itself wait; a burst of commits coalesces into one sync via `sync --auto`'s
436
- lock instead of piling up).
437
-
438
- There are also eight dry-run previews (`scan-git`, `scan-diff`, `scan-shell`, `scan-docs`,
439
- `scan-conversation`, `scan-session`, `scan-github`, `scan-structure`) that write nothing and print
440
- what ingestion *would* produce — nodes and their signal scores for the first seven, import-graph
441
- edges for `scan-structure`. That is the intended way to tune scoring against a real repository
442
- before committing to a change. Add `--json` to pipe them somewhere.
443
-
444
- Every command takes `-C <path>` to target another repository. On `sync`, `--conversation` and
445
- `--github` opt their (opt-in) sources in for one run without persisting it, `--no-embed` skips the
446
- vector pass, `--link-failures` builds the failure → fix chains described above, and `--rebuild`
447
- drops the project's nodes and re-ingests from scratch.
448
-
449
- `sync --prune-source <name>` deletes an entire collector source (e.g. `shell:pwsh`); `forget <value>`
450
- is the finer-grained complement — it deletes every node matching one exact string (or `--regex`
451
- pattern) *and* writes a standing deny-list entry so the value can never be re-ingested, even by a
452
- later `sync --rebuild` re-reading the append-only shell-hook log or a full transcript scan. Every
453
- removal leaves a hash-only tombstone, never the forgotten content itself. Both are dry-run by
454
- default; `--yes` confirms. `forget --list` shows active entries; `forget --export <path>` /
455
- `forget --import <path>` carry them to another checkout of the same repo (see the limitation above).
456
- See [`docs/forget-mechanism.md`](docs/forget-mechanism.md) for why this exists.
457
-
458
- ## Recall across projects
459
-
460
- `query --all-projects` searches every repository you have run NexusMem in, not just the current one,
461
- and tags each result with the repository it came from:
462
-
463
- ```
464
- $ nexusmem query --all-projects "why was the retry budget raised"
465
- scope 2 project(s): NexusMem, uploader
466
-
467
- - 2026-08-12 [observed] [uploader] fix: raise the retry budget after the S3 upload timeouts
468
- - 2026-08-12 [observed] [uploader] retry.ts @ 8d0f98b — fix: raise the retry budget after the S3 upload timeouts
469
- @@ -1 +1 @@
470
- -export const RETRY_BUDGET = 3;
471
- +export const RETRY_BUDGET = 5;
472
- - 2026-08-09 [observed] [NexusMem] fix(git): retry a transient failure to spawn git
473
- ```
474
-
475
- ## Bi-temporal reads
476
-
477
- Every node carries two clocks: `ts`, the event's own time ("what happened then"), and `created_at`,
478
- the moment the store actually recorded it ("what did the store hold then") — normally the same
479
- question, but not when a sync runs late, a backfill lands weeks after the events it describes, or a
480
- teammate's clone catches up all at once. `query`/`search_memory` answer the first by default;
481
- `--as-of <date>` switches to the second:
482
-
483
- ```bash
484
- nexusmem query "why does the ranker cap joint priors" --as-of 2026-08-10
485
- ```
486
-
487
- Only nodes recorded at or before that instant are considered, even if the events they describe are
488
- older still. There is no equivalent write — this is a read-time filter over `created_at`, not a
489
- snapshot or a way to query a value that has since changed, since nodes are write-once (see
490
- [Staleness & provenance](#staleness--provenance) for what changes a node's *weight*, not its
491
- record).
492
-
493
- Databases stay per-repository — there is no shared global store, and deleting one repo's
494
- `.nexusmem/` still removes exactly that repo's memory. What makes the others findable is a plain
495
- index at `~/.nexusmem/projects.json`, written by `init` and refreshed by every `sync`. `nexusmem
496
- projects` shows what is in it, and `--prune` forgets entries whose database is gone.
497
-
498
- Ranking across repositories uses reciprocal rank fusion per project rather than raw BM25, because a
499
- BM25 cost is computed against its own corpus and means different things in a 50-node and a
500
- 50,000-node database. The trade is stated in *Where it breaks*.
501
-
502
- The MCP `search_memory` tool takes the same switch as `allProjects: true`.
503
-
504
- ## On disk
505
-
506
- ```
507
- <repo>/.nexusmem/
508
- .gitignore '*' — the workspace ignores itself, so init never edits a file it doesn't own
509
- config.json validated on read; a corrupt config fails loudly rather than silently
510
- memory.db SQLite in WAL mode
511
-
512
- ~/.nexusmem/
513
- projects.json which repositories exist, for cross-project recall; a corrupt one reads as empty
514
- shell-history.jsonl the hook's log, if you installed it
515
- ```
516
-
517
- `NEXUSMEM_HOME` overrides the user-scoped directory.
518
-
519
- Node ids are content-addressed from `sha256(projectId + kind + naturalKey)`, so running `sync` twice
520
- cannot produce duplicates and ingestion stays correct even if a cursor is lost. Project identity
521
- comes from the normalized origin URL when there is one, falling back to the absolute path, so two
522
- clones of the same repo share one memory namespace.
523
-
524
- Deleting `.nexusmem/` loses nothing that `sync` cannot rebuild.
525
-
526
- ## Status
527
-
528
- Ingestion, hybrid retrieval, budgeted packing and the MCP server all work and are covered by 701
529
- tests running on Linux and Windows across Node 22 and 24.
530
-
531
- ## Development
532
-
533
- ```bash
534
- npm install
535
- npm run typecheck
536
- npm test
537
- npm run build
538
- ```
539
-
540
- Tests are behavioral rather than snapshot-based, and several are regressions tied to specific
541
- observed failures. `tests/git-errors.test.ts` injects a fake `spawn` to exercise the Windows
542
- process-spawn faults, which cannot be provoked on demand.
543
-
544
- Bug fixes, test coverage, and small focused features are welcome — see
545
- [CONTRIBUTING.md](CONTRIBUTING.md) for the workflow, or browse issues labeled
546
- [`good first issue`](https://github.com/yaminbkk/nexusmem/labels/good%20first%20issue)
547
- for something scoped and self-contained to start with.
548
-
549
- ## On how this was built
550
-
551
- This started as an experiment in whether a local context-memory engine for coding agents was viable,
552
- prototyped with Claude Code. The code was written through AI-assisted workflows; the architecture,
553
- the design decisions and the specifications were human-directed.
554
-
555
- That is worth stating plainly because it should change how you read the code, not whether you trust
556
- it. Audits, corrections and PRs are genuinely welcome, and the commit history is deliberately
557
- detailed about *why* things are the way they are, including the times an earlier assumption turned
558
- out to be wrong.
559
-
560
- ## License
561
-
562
- MIT
1
+ # NexusMem
2
+
3
+ [![CI](https://github.com/yaminbkk/NexusMem/actions/workflows/ci.yml/badge.svg)](https://github.com/yaminbkk/NexusMem/actions/workflows/ci.yml)
4
+ [![npm](https://img.shields.io/npm/v/nexusmem)](https://www.npmjs.com/package/nexusmem)
5
+ [![npm downloads](https://img.shields.io/npm/dm/nexusmem)](https://www.npmjs.com/package/nexusmem)
6
+ [![License: MIT](https://img.shields.io/badge/license-MIT-informational)](LICENSE)
7
+ ![Node](https://img.shields.io/badge/node-%3E%3D22-brightgreen)
8
+ [![yaminbkk/NexusMem MCP server](https://glama.ai/mcp/servers/yaminbkk/NexusMem/badges/score.svg)](https://glama.ai/mcp/servers/yaminbkk/NexusMem)
9
+ [![Listed on AiList](https://hifriendbot.com/ai-list/badge/nexusmem.svg)](https://hifriendbot.com/ai-list/nexusmem/)
10
+
11
+ ![NexusMem: init, sync --github, and a query against this repo's own history — surfacing a real issue, the PR that closed it, and the commits it shipped](docs/demo.gif)
12
+
13
+ Your coding agent can read `git log`. It cannot read the four things you tried last Tuesday that
14
+ didn't work.
15
+
16
+ NexusMem records what actually happened on your machine (shell commands and their exit codes, git
17
+ history down to the patch of each changed file, project docs, optionally your assistant transcripts)
18
+ into a local SQLite database, and
19
+ serves back a ranked, token-budgeted slice of it on demand. Everything stays on disk. No account, no
20
+ cloud, no telemetry.
21
+
22
+ The shell history is the part worth caring about. Git tells an agent what shipped. Shell history
23
+ tells it what was attempted, in what order, and which commands exited non-zero. That information
24
+ exists nowhere else, and it disappears when your terminal scrollback rolls over.
25
+
26
+ ![How NexusMem works: git commits, shell commands with their real exit codes, docs, and GitHub threads flow into a local SQLite index, which returns a ranked, token-budgeted slice of that history so your coding agent answers with a real citation instead of a guess](docs/how-it-works.svg)
27
+
28
+ **Contents:** [Try it](#try-it) · [Exact shell capture](#optional-exact-shell-capture) ·
29
+ [Failure → fix chains](#failure--fix-chains-opt-in) · [How retrieval works](#how-retrieval-works) ·
30
+ [Session summaries](#session-summaries-optional-local-model) ·
31
+ [GitHub issues & PRs](#github-issues--prs-optional) · [Use it from an agent](#use-it-from-an-agent)
32
+ · [What it costs you](#what-it-costs-you) · [Staleness & provenance](#staleness--provenance) ·
33
+ [Where it breaks](#where-it-breaks) · [Commands](#commands) · [Cross-project recall](#recall-across-projects)
34
+ · [On disk](#on-disk) · [Development](#development)
35
+
36
+ ## Try it
37
+
38
+ From inside any git repository:
39
+
40
+ ```
41
+ npx nexusmem init
42
+ npx nexusmem sync
43
+ ```
44
+
45
+ Then ask it something. Real output from this repository, top 2 of 5 hits:
46
+
47
+ ```
48
+ $ nexusmem query "windows spawn failure"
49
+
50
+ Relevant history for: windows spawn failure
51
+
52
+ - 2026-08-09 [observed] fix: distinguish a failed git spawn from "not a git repository"
53
+ readRepoInfo collapsed three unrelated failures into one error: git running and reporting
54
+ the path is not a work tree, git not being installed, and the process failing to spawn at
55
+ all. Dogfooding hit the third case in two separate sessions...
56
+ - 2026-08-09 [authored] README.md — Before a tagged release
57
+ - [ ] Retry on transient process-spawn failures on Windows
58
+ ```
59
+
60
+ `[observed]`/`[authored]` is the provenance tag (see [Staleness & provenance](#staleness--provenance))
61
+ — a commit is a directly observed event, a doc section is a written claim that could go stale.
62
+
63
+ A commit and a docs section, ranked against each other, inside whatever token budget you gave it.
64
+ Nothing was summarized by a model on the way out; the ranker just decided what not to send. (One
65
+ optional source, session summaries, does run a local model — but at ingest time, never on the way
66
+ out. What you query is always stored text.)
67
+
68
+ For a sense of what actually accumulates, here is `nexusmem status` on this repo after two days:
69
+
70
+ ```
71
+ 527 node(s) 2026-08-08 .. 2026-08-09
72
+ 321 shell_command
73
+ 130 conversation_turn
74
+ 60 doc_section
75
+ 16 git_commit
76
+ ```
77
+
78
+ Sixteen commits. Three hundred and twenty-one shell commands. The commits were already retrievable
79
+ by any agent with a terminal. The rest was not.
80
+
81
+ That `conversation_turn` row only appears because this corpus was synced with `--conversation`.
82
+ Assistant transcripts are the one source that is off by default and stays off until you opt in, since
83
+ they are the likeliest place for a pasted credential to be sitting. A default install indexes git
84
+ commits, their diffs, shell and docs.
85
+
86
+ Requirements: Node 22 or newer, and git. Node 20 will not work, because `better-sqlite3` ships no
87
+ prebuilt binary for it and Node 20 went end-of-life in April 2026. Ollama is optional and only
88
+ affects semantic search (see below).
89
+
90
+ ## Optional: exact shell capture
91
+
92
+ Scraped history files (PSReadLine, `.bash_history`, `.zsh_history`) give you command text and not
93
+ much else. The hook gives you working directory, exit code and a real timestamp:
94
+
95
+ ```bash
96
+ nexusmem hook install
97
+ ```
98
+
99
+ It wraps your existing PowerShell prompt rather than replacing it, is idempotent, and
100
+ `nexusmem hook remove` undoes it cleanly.
101
+
102
+ Exit codes are what make this worth installing. A failed command is a stronger signal than a
103
+ successful one, and without the hook there is no way to tell them apart.
104
+
105
+ ## Optional: what your coding agent tried
106
+
107
+ The shell hook above only sees commands *you* type. Commands your coding agent runs go through its
108
+ own tool, so they never reach an interactive prompt — which means the dead ends an agent hits, the
109
+ thing worth remembering most, were exactly what NexusMem could not see.
110
+
111
+ ```bash
112
+ nexusmem agent install # writes Claude Code's hooks; --project keeps it to this repo
113
+ ```
114
+
115
+ Restart Claude Code afterwards. From then on, in any repository NexusMem has been run in:
116
+
117
+ - **it records** each command the agent runs, its outcome, and the files edited just before it, so an
118
+ attempt reads "edited `a.ts`, then `npm test` failed" rather than "`npm test` failed twice";
119
+ - **when a command fails**, it says — to the agent, at that moment — whether that exact command has
120
+ failed here before, what was edited each time, and what eventually fixed it;
121
+ - **when a session starts**, it syncs in the background and names commands that failed with no
122
+ recorded fix.
123
+
124
+ It stays quiet the rest of the time. There is no match, no message; at most one note per failure per
125
+ session and five per session; no embedding, model or network call on that path. If NexusMem is
126
+ missing or broken the hook prints nothing and exits 0, and the agent carries on as if it were not
127
+ installed.
128
+
129
+ Commands are redacted before they are written, exactly like the shell hook, and matching is done on
130
+ a hash of the raw command so two commands differing only by a secret are never confused.
131
+
132
+ `nexusmem agent status` shows whether the hooks are installed; `nexusmem agent remove` takes out
133
+ NexusMem's own entries and leaves any other tool's hooks alone.
134
+
135
+ Install from the same environment Claude Code runs in. What gets written into the settings file is a
136
+ literal command — an absolute path to a Node executable and an absolute path into this copy of
137
+ NexusMem — and Claude Code hands that string to a shell in *its* environment. Install inside WSL
138
+ while Claude Code is the Windows binary (or the reverse) and the hooks are configured, look
139
+ installed, and capture nothing: the paths do not resolve on the side that runs them, and a hook that
140
+ cannot start produces no error anywhere. `nexusmem agent status` names any path in the installed
141
+ command that does not exist on the machine you ask from, which is also what it says after the
142
+ NexusMem it points at is moved, upgraded to a different location, or uninstalled.
143
+
144
+ ## Failure → fix chains (opt-in)
145
+
146
+ ```bash
147
+ nexusmem sync --link-failures
148
+ ```
149
+
150
+ After a normal sync, this walks every failed `shell_command` (non-zero exit code) and looks for
151
+ whatever later resolved it, using two independent heuristics: a later command in the same project
152
+ and working directory, exact same normalized text, that exited `0` within 24h (**same-command
153
+ retry**); and, separately, the best full-text match among nearby conversation turns or session
154
+ summaries, requiring every significant word of the failing command to appear, not just one
155
+ (**conversation bridge**). A failure can be linked by either, both, or neither.
156
+
157
+ Both links are surfaced in query results. The conversation-bridge heuristic originally matched on
158
+ any shared word, and dogfooding against this repo's own real history found it wrong on roughly half
159
+ its links — a shared word as generic as "npm" was enough to link an unrelated discussion. Requiring
160
+ every significant word fixed that: re-dogfooded against the same corpus, every resulting link (the
161
+ full set produced, not a sample) checked out correct on manual review of the full text, not just the
162
+ summary.
163
+
164
+ When a linked failure appears in a result set, its fix rides along immediately after it, inheriting
165
+ the failure's own relevance score rather than needing to match the query on its own merits. That is
166
+ the point: a query about why something failed shouldn't need to separately guess the words used in
167
+ whatever fixed it. This works across projects too — `query --all-projects` chains a failure to its
168
+ fix using whichever project's own database recorded the link, since links are always local to the
169
+ project they were found in.
170
+
171
+ ```
172
+ $ nexusmem query "why did npm whoami fail"
173
+
174
+ - 2026-08-12 [observed] shell: npm whoami (exit 1)
175
+ - 2026-08-12 [observed] shell: npm login (exit 0) -- linked as the fix
176
+ ```
177
+
178
+ ## How retrieval works
179
+
180
+ Every source normalizes to the same `MemoryNode` shape, so a commit, a shell command and a docs
181
+ section compete on equal terms. Retrieval runs BM25 over FTS5 and, if an embedding model is
182
+ reachable, a vector search over `sqlite-vec`, fused with Reciprocal Rank Fusion on rank *position*
183
+ only, never raw scores — a BM25 cost and a vector distance live on unrelated, unbounded scales, and
184
+ position is the only thing they agree on.
185
+
186
+ Ranking then multiplies three factors:
187
+
188
+ ```
189
+ score = relevance × signal^0.215 × recency^0.288
190
+ ```
191
+
192
+ `relevance` comes from the query. `signal` (a `fix:` commit outranks a `chore:`; a failed command
193
+ outranks a successful one) and `recency` are priors that hold before any query exists. Each factor is
194
+ floored into `[floor, 1]` rather than `[0, 1]`, so one weak dimension can't zero out a strong match.
195
+
196
+ The exponents bound how far signal and recency, *together*, may overturn relevance: at most a 2× gap
197
+ across their whole range, applied jointly rather than per-prior. That's deliberate — the score
198
+ *multiplies* the two priors, so capping each at 2× separately still let the pair overturn 4×, and
199
+ that hit hardest on fresh, high-signal commits made during an active working day. The bug that
200
+ exposed this: two unrelated same-day `fix:` commits outranked the docs section that actually answered
201
+ the query. See [`retrieval/rank.ts`](src/retrieval/rank.ts) for the full derivation.
202
+
203
+ Without Ollama, vector search is skipped and you get BM25 only — fully supported, not a degraded
204
+ state; `sync` and `query` both succeed and simply do less.
205
+
206
+ ## Session summaries (optional, local model)
207
+
208
+ With `sources.session.enabled`, each finished session becomes one distilled node next to the raw
209
+ exchanges — what was decided and why, rather than forty individual turns. It runs a local Ollama
210
+ chat model (`qwen2.5:3b` by default); nothing is downloaded automatically and nothing leaves the
211
+ machine.
212
+
213
+ ```bash
214
+ nexusmem scan-session --dry-run
215
+ ```
216
+
217
+ That prints the exact prompt a session would produce, after redaction and budget trimming, without
218
+ calling the model.
219
+
220
+ Three things bound the cost. A session is only summarized once it has been quiet for
221
+ `settleMinutes` (default 30), so a session in progress is not re-summarized on every sync. The
222
+ prompt is hashed, and an unchanged hash skips the model entirely — on this repo a steady-state sync
223
+ of 14 summarized sessions takes 0.25s and makes no model calls. And `maxSessions` (default 10) caps
224
+ how many reach the model per run; the rest are reported as queued and picked up next sync.
225
+
226
+ **What it is actually like, measured on 14 real sessions with `qwen2.5:3b`.** The summaries
227
+ themselves are good: decisions with their reasons, in the shape the prompt asks for. Titles are less
228
+ reliable — the model returned a usable one about a third of the time, and otherwise produced a
229
+ conversational preamble, a stray bullet, or a bare "Summary of the Session". Those are rejected and
230
+ the title falls back to the first line of the question that opened the session, which is always
231
+ specific even when it is not elegant. Compliance was worst on long sessions and on transcripts not
232
+ in English. A larger model (`qwen2.5:7b`) is the lever if the titles matter to you; set
233
+ `sources.session.model`.
234
+
235
+ ## GitHub issues & PRs (optional)
236
+
237
+ With `sources.github.enabled`, each issue and PR on this repo's github.com remote becomes one node
238
+ (title, opening post and every comment, folded together) alongside the raw discourse a discussion
239
+ already leaves in shell/conversation history. Off by default — not for a sensitivity reason, but
240
+ because it's the first source with a real external dependency: it reads via the `gh` CLI, so it
241
+ needs `gh` installed and authenticated (`gh auth login`), and it makes live network calls instead of
242
+ only reading what's already on disk. A repo with no github.com remote, or an unauthenticated `gh`,
243
+ is a silent no-op either way.
244
+
245
+ ```bash
246
+ nexusmem scan-github
247
+ ```
248
+
249
+ previews the nodes a sync would produce, same as the other `scan-*` commands. `maxThreads` (default
250
+ 100) and `maxCommentsPerThread` (default 100) bound one sync's cost; `since` is tracked as its own
251
+ cursor, so a repeat sync only re-reads threads that changed. Dogfooded against this repo's own 14
252
+ real issues/PRs: ingest took under a second, and a query for "labelled retrieval regression corpus"
253
+ correctly ranked issue #8 — the one that asked for it — first.
254
+
255
+ ## Use it from an agent
256
+
257
+ ```json
258
+ {
259
+ "mcpServers": {
260
+ "nexusmem": {
261
+ "command": "npx",
262
+ "args": ["-y", "nexusmem", "mcp"]
263
+ }
264
+ }
265
+ }
266
+ ```
267
+
268
+ Three tools over stdio: `search_memory` returns the packed context block, `sync_project` ingests, and
269
+ `get_status` reports what is currently remembered. Each takes an explicit `projectRoot`, because an
270
+ MCP tool call carries no shell working directory. `sync_project` runs `init` for you if the
271
+ repository has not been set up.
272
+
273
+ ## What it costs you
274
+
275
+ Two numbers get conflated in tools like this, so they are kept apart here.
276
+
277
+ **Packer efficiency** is how much the ranker trims from its own candidate set. On this repository's
278
+ corpus it runs 81–84%. It is useful for tuning the ranker and useless as a claim about your bill,
279
+ because the baseline is hypothetical: without NexusMem those candidates were never going into your
280
+ context window in the first place.
281
+
282
+ **End-to-end saving** compares the packed context NexusMem actually sends against reading, in full,
283
+ the same files its own ranking identified as relevant to the query. Measured with
284
+ [`scripts/benchmark.ts`](scripts/benchmark.ts) (`npm run bench`), which anyone who clones this repo
285
+ and points it at a synced corpus can re-run from scratch:
286
+
287
+ | Corpus | Commits | Query set | vs. full file content | vs. `git log -p` on those files |
288
+ | --- | --- | --- | --- | --- |
289
+ | This repo | 62 | 16 real prompts, verbatim from this project's own history | 95% (median 94%) | 98% (median 97%) |
290
+ | [`vitejs/vite`](https://github.com/vitejs/vite) | 9,567 | 16, mechanically sampled — see below | 99% (median 98%) | ~100% (median ~100%) — see caveat |
291
+
292
+ Both clear the original >70% target ("cut API token spend versus sending full context"), and the vite
293
+ run is the first measurement at the scale that target was always described as applying to.
294
+
295
+ **Read the methodology before quoting either number — it's a narrower claim than it looks:**
296
+
297
+ - **Graded against NexusMem's own ranking**, not an outside answer key: the file set is whichever
298
+ files the packed nodes for that query touch. This measures what the pack step saves once retrieval
299
+ already picked a candidate set; it doesn't independently verify that set was the right one.
300
+ - **Query sets are mechanical, not cherry-picked** (see `scripts/benchmark.ts`): vite's is an even
301
+ sample of well-explained `fix`/`feat`/`perf`/`refactor` commits plus rationale-bearing doc headings;
302
+ this repo's reuses real historical prompts verbatim, several of which are broad task instructions
303
+ rather than narrow questions — part of why its number sits below vite's.
304
+ - **`git log -p` baselines can be enormous** — one vite query's baseline hit 7.5M tokens because a
305
+ file in its resolved set has that much history. At that scale, "just read the file's history
306
+ instead" stops being a viable alternative at all.
307
+ - **Supersedes the old ~40% figure**, which was hand-tallied from two hand-picked queries against
308
+ this repo alone with an unstated baseline. Not wrong, just underspecified — this replaces it with a
309
+ stated method and a script that reproduces it.
310
+
311
+ One thing that is not a percentage: shell commands and conversation turns have no cheap `grep`
312
+ equivalent. Without something recording them, they are gone, not merely more expensive to find.
313
+
314
+ For how these numbers compare to a similar tool's own claims, see
315
+ [`docs/competitor-comparison.md`](docs/competitor-comparison.md) (vs. projectmem) and
316
+ [`docs/competitor-comparison-yesmem.md`](docs/competitor-comparison-yesmem.md) (vs. YesMem, including
317
+ native Windows support vs. its documented WSL2 requirement).
318
+
319
+ Latency on a ~530-node corpus, warm, p50 over 10 runs:
320
+
321
+ | Operation | |
322
+ | --- | --- |
323
+ | BM25 retrieval (FTS5) | ~1.1 ms |
324
+ | Vector KNN (`sqlite-vec`) | ~3.2 ms |
325
+ | Fuse, rank, pack | ~0.6 ms |
326
+ | Query embedding (local Ollama) | ~55–77 ms |
327
+ | **End-to-end hybrid** | **~56 ms** |
328
+
329
+ All the SQLite work totals about 5 ms. The embedding call is the only thing on this path worth
330
+ optimizing, and it is somebody else's process.
331
+
332
+ ## Staleness & provenance
333
+
334
+ Two things a memory layer needs and this one only partly has: a way to tell an observed fact from a
335
+ guess, and a way to retire a conclusion once something contradicts it. This section is what exists
336
+ and what doesn't.
337
+
338
+ Every node carries a `provenance`, a four-tier trust hierarchy set once per collector at ingest
339
+ time: `observed` (a commit that landed, a shell command's real exit code) > `authored` (a doc
340
+ section — a human's own written claim) > `recorded` (a conversation turn — verbatim, but talk about
341
+ events rather than the events) > `derived` (a session summary — a model's distillation). The tier is
342
+ shown as a tag on every query result and decays retrieval weight — the lower the trust, the faster a
343
+ node fades from ranking as it ages. The ordering is the design claim; the exact decay ratios are
344
+ judgment calls, not measured optima.
345
+
346
+ ```bash
347
+ nexusmem stale
348
+ ```
349
+
350
+ Lists non-`observed` nodes old enough (45+ days by default) that nothing has confirmed they still
351
+ hold — a heuristic on age and provenance, not on content. It writes nothing; you decide which
352
+ candidates are actually wrong. Any candidate the SLM has already flagged (see below) is decorated
353
+ with its standing `likely superseded by` suggestion — reading those costs nothing, so the plain
354
+ command stays instant and offline.
355
+
356
+ ```bash
357
+ nexusmem mark-stale <oldNodeId> --supersedes <newNodeId>
358
+ ```
359
+
360
+ Links `newNodeId` as the replacement for `oldNodeId`. The ranker down-weights the old node from then
361
+ on (it stays queryable, just usually loses to its replacement) — nothing is deleted, unlike `forget`.
362
+
363
+ ```bash
364
+ nexusmem stale --check-contradictions
365
+ ```
366
+
367
+ For each candidate, finds the most similar newer node (local embedding search) and asks a local SLM
368
+ (Ollama, `qwen2.5:3b` by default) whether it actually contradicts the older one — real content
369
+ comparison, not just age. A match is printed as `likely superseded by <id> <title> -- <reason>`
370
+ under the candidate. Every judgment (either verdict) is memoized, so a judged pair is never sent to
371
+ the model again; nothing else is written — `supersedes` stays yours to set via `mark-stale`.
372
+
373
+ **This also runs automatically during `sync`** — at most 3 new judgments per run (configurable via
374
+ the `contradictions` block in `.nexusmem/config.json`; set `autoCheck: false` to turn it off), only
375
+ when the embedding provider was reachable anyway, and free on repeat syncs thanks to the
376
+ memoization. New and open suggestions show up in the sync summary, `nexusmem status` (a `flagged`
377
+ line), and plain `nexusmem stale`.
378
+
379
+ **What this doesn't do:** it is one small model's yes/no judgment on one older/newer pair, not a
380
+ verified fact — treat a match as a lead to check, not a conclusion. It also only ever compares a
381
+ candidate against nodes *found by embedding similarity*; a contradiction from an unrelated-sounding
382
+ node would never surface. Comprehensive contradiction detection (not just for the pair the vector
383
+ search happens to surface) is still an open problem, and nothing here supersedes a node on its own.
384
+
385
+ `provenance` is a separate question from `trust_state`: provenance says where a claim came from,
386
+ never whether anyone checked it.
387
+
388
+ ```bash
389
+ nexusmem review <nodeId> --verify
390
+ nexusmem review <nodeId> --reject
391
+ ```
392
+
393
+ Records your own verdict on one node, independent of the SLM contradiction checker above (which only
394
+ ever writes a suggestion, never a verdict). `--reject` down-weights the node in ranking — same
395
+ demote-not-delete rule as `mark-stale`, it stays queryable, just usually loses to better matches —
396
+ and both verdicts are shown as a `[verified]`/`[rejected]` tag on every query result that returns the
397
+ node afterward. `--verify` is a label only; it does not boost ranking. Every node starts `candidate`
398
+ (untagged) until reviewed, and a re-sync never overwrites a verdict once one is set.
399
+
400
+ ## Where it breaks
401
+
402
+ - **Shell history without the hook is unscoped.** Scraped history has no directory context, so it is
403
+ attributed to whichever repository you ran `sync` from. Bounded to a tail window, and an
404
+ approximation rather than a guarantee.
405
+ - **Japanese and Chinese depend on the vector pass.** FTS5's `unicode61` tokenizer splits on
406
+ whitespace, so languages without space boundaries get no useful BM25 recall.
407
+ - **Rebasing strands nodes.** Rewritten history leaves nodes for unreachable commits. They describe
408
+ real events so they are not wrong, but a targeted prune does not exist yet. `sync --rebuild`
409
+ forces a clean re-scan.
410
+ - **Multi-line PowerShell input is read as separate commands.** A function typed across several lines
411
+ at the prompt is not reconstructed.
412
+ - **Scrape-fallback ids drift** if the history file is trimmed from the front between syncs.
413
+ Installing the hook fixes this.
414
+ - **Session-summary titles depend on the model following instructions**, and a 3B model often does
415
+ not. The fallback keeps them specific rather than generic, but see the section above for what to
416
+ expect.
417
+ - **Changing the embedding model re-embeds everything.** Vectors from two models are not comparable
418
+ and `nodes_vec` records no per-row provenance, so `sync` drops the lot and rebuilds rather than
419
+ ranking across a mixture. It says so when it happens. Nodes are untouched and BM25 keeps working
420
+ throughout.
421
+ - **Diff indexing is bounded, and deliberately lossy.** A first sync indexes the patches of the most
422
+ recent 200 commits (later syncs only walk `cursor..HEAD`); merge commits contribute none, since
423
+ their patch exists only in a combined format this parser does not read; and binaries, lockfiles and
424
+ build output are skipped so a dependency bump cannot bury the corpus. All of it is still recorded
425
+ as a `git_commit` node. A patch longer than `limits.maxBodyChars` is truncated, so the tail of a
426
+ very large change is not indexed. The caps live under `sources.diff` in `config.json`.
427
+ - **Cross-project recall favours breadth.** Each repository's hits are fused by rank, so a project
428
+ whose best match is mediocre still contributes a rank-1 item, and rank 1 is worth the same in
429
+ every list. Adding a repository that has little to say about your question still pushes a few of
430
+ its results into the budget. Signal, recency and the budget are what hold that in check; there is
431
+ no per-project quality weight.
432
+ - **The project registry is an index, not a source of truth.** It can point at a database that has
433
+ moved or been deleted; those are reported and skipped, never silently pruned, because an
434
+ unmounted drive is not a deleted project.
435
+ - **Conversation chunking is unevaluated.** Splitting long replies at heading boundaries measurably
436
+ helped, but it has never been tested systematically.
437
+ - **A chunked node's sibling count in one result is capped, not tuned.** `conversation_turn` and
438
+ `doc_section` both split one reply or file into several nodes; at most 2 of them may appear
439
+ together in a packed result. Found live: a query for "token" returned 9 of its top 12 hits as
440
+ different pieces of one heavily-sectioned reply, crowding out the node that actually answered it.
441
+ The cap of 2 is a judgement call, not a measured optimum, same as the ranking priors' budget above.
442
+ - **The size of the prior budget is a judgement call, not a measured optimum.** Priors are now
443
+ bounded jointly rather than one at a time, which closed a real 4× hole (see the ranking section),
444
+ but the 2× budget itself has never been tuned against a labelled relevance set — there isn't one.
445
+ It is a defensible constant, not a result. What is measured is the direction: on four real queries
446
+ against this repo's own memory, switching to the joint cap moved the section that answered the
447
+ question up in three of them (the rationale section for "why BM25 before vector search" went from
448
+ rank 4 to rank 1) and displaced no query's correct top hit.
449
+ - **`forget` is per-repository, not global.** Its deny-list lives in the one `.nexusmem/memory.db` it
450
+ ran against (plus that repo's stale prior identities, same scope `--prune-source` already uses). A
451
+ value that leaked into shell history from several repositories needs `forget` run once per repo —
452
+ there is no shared, machine-wide deny-list across every project you have synced.
453
+ - **A deny-list doesn't survive a clone or restore on its own.** `.nexusmem/` is gitignored by design,
454
+ so `deny_list` never travels with `git clone`/`git push` — while git history itself, the thing a
455
+ fresh `sync` re-derives from, is fully portable and copied by every clone. A teammate's fresh
456
+ checkout, a new machine, or a restored backup starts with zero protection: the forgotten value comes
457
+ right back on the first sync. Confirmed live 2026-08-17, not just a theoretical read of the code.
458
+ `forget --export <path>` / `forget --import <path>` close this: export writes the active entries to
459
+ a plaintext JSON file you move through a channel you control (never git — the file is exactly as
460
+ sensitive as the value it holds), and import re-applies them in the new checkout, deleting any
461
+ copies that already synced back in. It is deliberately manual, not automatic on every `sync`.
462
+
463
+ ## Commands
464
+
465
+ `init`, `sync`, `query <text>` (add `--as-of <date>` for a bi-temporal read, see below), `status` (add
466
+ `--share` for a plain-text summary worth pasting somewhere), `projects`, `mcp`, `forget <value>`,
467
+ `scrub-secrets` (re-redacts secrets that versions before 0.10.5 already stored — database rows, FTS
468
+ index, embeddings and the shell hook log; dry-run unless `--yes`, which backs each database up first;
469
+ `--all-projects` covers every registered repository),
470
+ `stale` (add `--check-contradictions` for a local-SLM content check, see above), `mark-stale
471
+ <nodeId> --supersedes <newNodeId>`, `review <nodeId> --verify|--reject` (record a human verdict on
472
+ one node, see above), `precheck` (advisory — warns about staged files with an unresolved past
473
+ failure or high recent churn; exits 0 unless `--strict`), `hook install|remove|status` (the
474
+ PowerShell exit-code hook), `hook git install|remove|status` (a git pre-commit hook that runs
475
+ `precheck` before each commit), and `hook git-post install|remove|status` (a git post-commit hook
476
+ that runs a full `sync`, including embedding, in the background after each commit — detached, so it
477
+ never makes `git commit` itself wait; a burst of commits coalesces into one sync via `sync --auto`'s
478
+ lock instead of piling up). `agent install|remove|status` manages the Claude Code hooks described
479
+ above; `agent recall` and `agent session-start` are run by those hooks, not by hand.
480
+
481
+ There are also eight dry-run previews (`scan-git`, `scan-diff`, `scan-shell`, `scan-docs`,
482
+ `scan-conversation`, `scan-session`, `scan-github`, `scan-structure`) that write nothing and print
483
+ what ingestion *would* produce — nodes and their signal scores for the first seven, import-graph
484
+ edges for `scan-structure`. That is the intended way to tune scoring against a real repository
485
+ before committing to a change. Add `--json` to pipe them somewhere.
486
+
487
+ Every command takes `-C <path>` to target another repository. On `sync`, `--conversation` and
488
+ `--github` opt their (opt-in) sources in for one run without persisting it, `--no-embed` skips the
489
+ vector pass, `--link-failures` builds the failure → fix chains described above, and `--rebuild`
490
+ drops the project's nodes and re-ingests from scratch.
491
+
492
+ `sync --prune-source <name>` deletes an entire collector source (e.g. `shell:pwsh`); `forget <value>`
493
+ is the finer-grained complement — it deletes every node matching one exact string (or `--regex`
494
+ pattern) *and* writes a standing deny-list entry so the value can never be re-ingested, even by a
495
+ later `sync --rebuild` re-reading the append-only shell-hook log or a full transcript scan. Every
496
+ removal leaves a hash-only tombstone, never the forgotten content itself. Both are dry-run by
497
+ default; `--yes` confirms. `forget --list` shows active entries; `forget --export <path>` /
498
+ `forget --import <path>` carry them to another checkout of the same repo (see the limitation above).
499
+ See [`docs/forget-mechanism.md`](docs/forget-mechanism.md) for why this exists.
500
+
501
+ ## Recall across projects
502
+
503
+ `query --all-projects` searches every repository you have run NexusMem in, not just the current one,
504
+ and tags each result with the repository it came from:
505
+
506
+ ```
507
+ $ nexusmem query --all-projects "why was the retry budget raised"
508
+ scope 2 project(s): NexusMem, uploader
509
+
510
+ - 2026-08-12 [observed] [uploader] fix: raise the retry budget after the S3 upload timeouts
511
+ - 2026-08-12 [observed] [uploader] retry.ts @ 8d0f98b — fix: raise the retry budget after the S3 upload timeouts
512
+ @@ -1 +1 @@
513
+ -export const RETRY_BUDGET = 3;
514
+ +export const RETRY_BUDGET = 5;
515
+ - 2026-08-09 [observed] [NexusMem] fix(git): retry a transient failure to spawn git
516
+ ```
517
+
518
+ ## Bi-temporal reads
519
+
520
+ Every node carries two clocks: `ts`, the event's own time ("what happened then"), and `created_at`,
521
+ the moment the store actually recorded it ("what did the store hold then") — normally the same
522
+ question, but not when a sync runs late, a backfill lands weeks after the events it describes, or a
523
+ teammate's clone catches up all at once. `query`/`search_memory` answer the first by default;
524
+ `--as-of <date>` switches to the second:
525
+
526
+ ```bash
527
+ nexusmem query "why does the ranker cap joint priors" --as-of 2026-08-10
528
+ ```
529
+
530
+ Only nodes recorded at or before that instant are considered, even if the events they describe are
531
+ older still. There is no equivalent write — this is a read-time filter over `created_at`, not a
532
+ snapshot or a way to query a value that has since changed, since nodes are write-once (see
533
+ [Staleness & provenance](#staleness--provenance) for what changes a node's *weight*, not its
534
+ record).
535
+
536
+ Databases stay per-repository — there is no shared global store, and deleting one repo's
537
+ `.nexusmem/` still removes exactly that repo's memory. What makes the others findable is a plain
538
+ index at `~/.nexusmem/projects.json`, written by `init` and refreshed by every `sync`. `nexusmem
539
+ projects` shows what is in it, and `--prune` forgets entries whose database is gone.
540
+
541
+ Ranking across repositories uses reciprocal rank fusion per project rather than raw BM25, because a
542
+ BM25 cost is computed against its own corpus and means different things in a 50-node and a
543
+ 50,000-node database. The trade is stated in *Where it breaks*.
544
+
545
+ The MCP `search_memory` tool takes the same switch as `allProjects: true`.
546
+
547
+ ## On disk
548
+
549
+ ```
550
+ <repo>/.nexusmem/
551
+ .gitignore '*' — the workspace ignores itself, so init never edits a file it doesn't own
552
+ config.json validated on read; a corrupt config fails loudly rather than silently
553
+ memory.db SQLite in WAL mode
554
+
555
+ ~/.nexusmem/
556
+ projects.json which repositories exist, for cross-project recall; a corrupt one reads as empty
557
+ shell-history.jsonl the hook's log, if you installed it
558
+ ```
559
+
560
+ `NEXUSMEM_HOME` overrides the user-scoped directory.
561
+
562
+ Node ids are content-addressed from `sha256(projectId + kind + naturalKey)`, so running `sync` twice
563
+ cannot produce duplicates and ingestion stays correct even if a cursor is lost. Project identity
564
+ comes from the normalized origin URL when there is one, falling back to the absolute path, so two
565
+ clones of the same repo share one memory namespace.
566
+
567
+ Deleting `.nexusmem/` loses nothing that `sync` cannot rebuild.
568
+
569
+ ## Status
570
+
571
+ Ingestion, hybrid retrieval, budgeted packing and the MCP server all work and are covered by 701
572
+ tests running on Linux and Windows across Node 22 and 24.
573
+
574
+ ## Development
575
+
576
+ ```bash
577
+ npm install
578
+ npm run typecheck
579
+ npm test
580
+ npm run build
581
+ ```
582
+
583
+ Tests are behavioral rather than snapshot-based, and several are regressions tied to specific
584
+ observed failures. `tests/git-errors.test.ts` injects a fake `spawn` to exercise the Windows
585
+ process-spawn faults, which cannot be provoked on demand.
586
+
587
+ Bug fixes, test coverage, and small focused features are welcome — see
588
+ [CONTRIBUTING.md](CONTRIBUTING.md) for the workflow, or browse issues labeled
589
+ [`good first issue`](https://github.com/yaminbkk/nexusmem/labels/good%20first%20issue)
590
+ for something scoped and self-contained to start with.
591
+
592
+ ## On how this was built
593
+
594
+ This started as an experiment in whether a local context-memory engine for coding agents was viable,
595
+ prototyped with Claude Code. The code was written through AI-assisted workflows; the architecture,
596
+ the design decisions and the specifications were human-directed.
597
+
598
+ That is worth stating plainly because it should change how you read the code, not whether you trust
599
+ it. Audits, corrections and PRs are genuinely welcome, and the commit history is deliberately
600
+ detailed about *why* things are the way they are, including the times an earlier assumption turned
601
+ out to be wrong.
602
+
603
+ ## License
604
+
605
+ MIT