nexusmem 0.10.4 → 0.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +768 -651
- package/LICENSE +21 -21
- package/README.md +605 -562
- package/dist/cli/agent-hook.js +450 -0
- package/dist/cli/agent-hook.js.map +1 -0
- package/dist/cli/index.js +1904 -291
- package/dist/cli/index.js.map +1 -1
- package/dist/cli/recorder.js +225 -0
- package/dist/cli/recorder.js.map +1 -0
- package/package.json +71 -70
package/CHANGELOG.md
CHANGED
|
@@ -1,16 +1,128 @@
|
|
|
1
|
-
# Changelog
|
|
2
|
-
|
|
3
|
-
Notable changes per published version. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/);
|
|
4
|
-
this project uses [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
5
|
-
|
|
6
|
-
Tags were added retroactively on 2026-08-11 and point at the exact commits the npm tarballs were
|
|
7
|
-
built from, matched by publish timestamp: `v0.1.0` → `67a4776`, `v0.1.1` → `809e62c`,
|
|
8
|
-
`v0.1.2` → `b22a3b0`.
|
|
9
|
-
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
Notable changes per published version. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/);
|
|
4
|
+
this project uses [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
5
|
+
|
|
6
|
+
Tags were added retroactively on 2026-08-11 and point at the exact commits the npm tarballs were
|
|
7
|
+
built from, matched by publish timestamp: `v0.1.0` → `67a4776`, `v0.1.1` → `809e62c`,
|
|
8
|
+
`v0.1.2` → `b22a3b0`.
|
|
9
|
+
|
|
10
10
|
## [Unreleased]
|
|
11
11
|
|
|
12
12
|
No unreleased changes yet.
|
|
13
13
|
|
|
14
|
+
## [0.11.0] — 2026-09-27
|
|
15
|
+
|
|
16
|
+
### Added
|
|
17
|
+
|
|
18
|
+
- Ambient memory for coding agents: `nexusmem agent install` wires NexusMem into Claude Code's own
|
|
19
|
+
hooks, so what the agent tried is recorded and what it already tried is surfaced without anyone
|
|
20
|
+
asking for either.
|
|
21
|
+
- **Capture.** Every Bash command Claude Code runs, and every file it edits, is recorded with its
|
|
22
|
+
outcome. Commands the agent runs are invisible to the shell hook, which only sees interactive
|
|
23
|
+
prompts — this is what makes an agent's own dead ends memorable at all. Edits are attached to the
|
|
24
|
+
command that follows them, so an attempt reads "edited a.ts, then npm test failed".
|
|
25
|
+
- **Recall on failure.** When a command fails, NexusMem injects a short note only if that exact
|
|
26
|
+
command failed here before, naming what was edited each time and what eventually fixed it. No
|
|
27
|
+
match means no output.
|
|
28
|
+
- **Session start.** Starts a background sync, then names commands that failed with no recorded
|
|
29
|
+
fix — and, for one already fixed, which file(s) the fix touched — in about 150 tokens. A
|
|
30
|
+
repository with nothing unresolved gets nothing.
|
|
31
|
+
- Budgeted: at most one note per failure per session, five per session, and no embedding, model or
|
|
32
|
+
network call on the path. If NexusMem is missing, broken or slow, the hook prints nothing and
|
|
33
|
+
exits 0, and the agent carries on.
|
|
34
|
+
- Commands the agent ran are redacted by the same recorder pattern as the shell hook: the raw text
|
|
35
|
+
is hashed and redacted in memory, and only redacted text is written. Correlation matches on the
|
|
36
|
+
hash, so two commands that differ only by a secret are never treated as the same command.
|
|
37
|
+
- `nexusmem agent status` and `nexusmem agent remove` manage the hooks; `remove` takes only
|
|
38
|
+
NexusMem's own entries. Installs into user settings by default, or `--project` for
|
|
39
|
+
`.claude/settings.local.json`, never a settings file a team would commit.
|
|
40
|
+
- `nexusmem agent status` also lists any path in the installed command that does not exist on this
|
|
41
|
+
machine. Installing from one environment and running Claude Code in another (WSL vs. the Windows
|
|
42
|
+
binary) writes paths the executing shell cannot resolve, and the hook then captures nothing with
|
|
43
|
+
no error anywhere; the same line appears once the NexusMem it points at is moved or uninstalled.
|
|
44
|
+
- Agent actions are ingested by every `sync`, and their arrival now runs failure→fix correlation
|
|
45
|
+
without `--link-failures`.
|
|
46
|
+
|
|
47
|
+
### Changed
|
|
48
|
+
|
|
49
|
+
- Failure→fix correlation treats an agent attempt as files changed, execution and result, rather
|
|
50
|
+
than as a command string. When an agent-recorded command fails and the identical command later
|
|
51
|
+
passes with nothing edited in between, that pass is reported as unexplained instead of being
|
|
52
|
+
linked as the fix — the pass is real, the explanation is not. Human shell history records no files,
|
|
53
|
+
so its linking is unchanged.
|
|
54
|
+
|
|
55
|
+
### Fixed
|
|
56
|
+
|
|
57
|
+
- A command piped into `head`, `tail` or `tee` (`node check.js 2>&1 | head -100`) reports the
|
|
58
|
+
filter's exit code to Claude Code's hook, not the command it fed — `PostToolUse` fired even when
|
|
59
|
+
the piped command failed loudly, and the run was recorded as a false `ok`. Recorded as `unknown`
|
|
60
|
+
instead, and a repeat of the same command surfaces an honestly-worded note rather than either a
|
|
61
|
+
guessed pass or a guessed failure. A pipe into anything else (`| grep`, `| jq`, ...) is left
|
|
62
|
+
untouched — nothing about its own exit code is known here.
|
|
63
|
+
- The per-session injection quota — "explain a failure once per session" — was a plain read
|
|
64
|
+
(`shouldInject`) then a later write (`markInjected`), racy across the separate OS processes Claude
|
|
65
|
+
Code actually spawns per tool call: confirmed live, 3 of 20 concurrent invocations printed the same
|
|
66
|
+
recall instead of 1. Replaced with a single lock-protected read-modify-write that never blocks —
|
|
67
|
+
contention or a stale lock both decline rather than wait.
|
|
68
|
+
|
|
69
|
+
## [0.10.5] — 2026-09-11
|
|
70
|
+
|
|
71
|
+
Security release. Existing installs need action after upgrading — see "Action required" below.
|
|
72
|
+
|
|
73
|
+
### Security
|
|
74
|
+
|
|
75
|
+
- Secret redaction missed keys whose secret word is part of a longer identifier: `DB_PASSWORD=x`,
|
|
76
|
+
`OPENAI_API_KEY=x`, `AUTH_TOKEN=x`, `CLIENT_SECRET=x` were kept verbatim (the rule's leading `\b`
|
|
77
|
+
never matches after `_`), as were values shorter than 8 characters or containing characters such
|
|
78
|
+
as `@ ! # $`. This affected `scan-shell` output, stored shell and conversation nodes, session
|
|
79
|
+
summaries, the FTS index, embeddings, and `search_memory` / `query` results.
|
|
80
|
+
- `shell_command` nodes stored the raw command in `meta.command`, so a typed secret was written to
|
|
81
|
+
`memory.db` even when the title showed it redacted, and was printed by `scan-shell --json` and
|
|
82
|
+
`precheck`. `meta.command` is now redacted; `meta.commandHash` keeps the raw command's hash so node
|
|
83
|
+
ids and project-id reconciliation are unchanged.
|
|
84
|
+
- The installed shell hooks appended every command raw to `~/.nexusmem/shell-history.jsonl`. Hooks now
|
|
85
|
+
pass each command to a NexusMem recorder over stdin; the recorder hashes and redacts it in memory
|
|
86
|
+
and writes only the redacted command. It writes no temp files and prints nothing, and it drops any
|
|
87
|
+
event it cannot process instead of writing it raw. Hooks installed by earlier versions keep writing
|
|
88
|
+
raw commands until reinstalled: `hook status` reports them as outdated, and `sync` redacts what they
|
|
89
|
+
write and warns.
|
|
90
|
+
- Redaction now also covers credentials in connection URIs (`postgres://user:pass@host`),
|
|
91
|
+
`Authorization:` headers and bearer tokens, secret-named options whose value is the next argument
|
|
92
|
+
(`--password x`), and `mysql -pX`, `sshpass -p`, `mongosh -p`, `redis-cli -a`, `curl -u user:pass`.
|
|
93
|
+
|
|
94
|
+
### Added
|
|
95
|
+
|
|
96
|
+
- `nexusmem scrub-secrets [--all-projects] [--yes] [--no-embed]` removes secrets that earlier versions
|
|
97
|
+
already stored: shell, conversation, session-summary and code-diff rows, contradiction reasons, the
|
|
98
|
+
FTS index, embeddings, and the shell hook log. Dry run by default. `--yes` takes a backup of each
|
|
99
|
+
database first and prints its path, redacts rows in place (ids unchanged), drops and re-embeds
|
|
100
|
+
affected vectors, rebuilds the FTS index, runs `VACUUM`, and truncates the WAL. Safe to run
|
|
101
|
+
repeatedly. Existing `memory.db.backup-*` files are listed, never modified or deleted.
|
|
102
|
+
|
|
103
|
+
### Action required
|
|
104
|
+
|
|
105
|
+
1. Upgrade every NexusMem install (global CLI, MCP server configs, the post-commit hook's `nexusmem`),
|
|
106
|
+
then restart MCP clients and editors so no older version keeps writing.
|
|
107
|
+
2. Run `nexusmem hook install` again in each shell you installed the hook for; `nexusmem hook status`
|
|
108
|
+
should say `installed`, not `OUTDATED`.
|
|
109
|
+
3. Rotate any credential typed at a shell prompt or pasted into a captured conversation in the
|
|
110
|
+
affected forms. Scrubbing cannot recall copies already returned to an LLM, backed up, or synced.
|
|
111
|
+
4. Run `nexusmem scrub-secrets --all-projects`, review, then re-run with `--yes`.
|
|
112
|
+
5. Delete the `memory.db.backup-*` files NexusMem lists (they predate redaction) once you have
|
|
113
|
+
verified the scrubbed database.
|
|
114
|
+
6. Do not use `nexusmem forget <secret>` for this: it stores the value itself in the database's
|
|
115
|
+
deny-list, and the shell hook captures a bare argument unredacted.
|
|
116
|
+
|
|
117
|
+
### Known limitations
|
|
118
|
+
|
|
119
|
+
- NexusMem cannot change your shell's own history files (PSReadLine `ConsoleHost_history.txt`,
|
|
120
|
+
`~/.bash_history`, `~/.zsh_history`); they still contain whatever you typed.
|
|
121
|
+
- Not yet redacted: a token used as a URL username (`https://TOKEN@host`), `openssl -passin`,
|
|
122
|
+
`keytool -storepass`, `ssh-keygen -N`, secret-bearing keys without a secret word (`STRIPE_KEY`,
|
|
123
|
+
`DB_PW`), secrets passed as plain positional arguments, and commit messages, docs and GitHub
|
|
124
|
+
threads (never redacted).
|
|
125
|
+
|
|
14
126
|
## [0.10.4] — 2026-09-07
|
|
15
127
|
|
|
16
128
|
### Added
|
|
@@ -29,645 +141,650 @@ No unreleased changes yet.
|
|
|
29
141
|
evidence is excluded from failure-to-fix correlation.
|
|
30
142
|
- Schema V13 adds `nodes.capture_mode` and nullable `nodes.source_ts`. Existing approximate
|
|
31
143
|
shell timestamps and document mtimes migrate to a null `source_ts`; other existing event timestamps are retained.
|
|
32
|
-
|
|
33
|
-
## [0.10.3] — 2026-08-31
|
|
34
|
-
|
|
35
|
-
### Added
|
|
36
|
-
|
|
37
|
-
- `hook git-post install|remove|status`: an opt-in git post-commit hook that runs a full
|
|
38
|
-
`nexusmem sync` (including embedding) in the background after every commit. Detached (`nohup ... &`),
|
|
39
|
-
so it never makes `git commit` itself wait; output goes to `.nexusmem/post-commit-sync.log` instead
|
|
40
|
-
of the terminal. `sync` gained a matching `--auto` flag (used only by this hook) that skips instead
|
|
41
|
-
of running if another `--auto` sync already holds an advisory lock over the project's workspace dir —
|
|
42
|
-
a burst of commits (e.g. a rebase) coalesces into one sync instead of piling up. The log file is
|
|
43
|
-
reset once it passes 2000 lines so it can't grow forever, and a skipped/coalesced run is always
|
|
44
|
-
reported there, even though the hook always passes `--quiet`.
|
|
45
|
-
|
|
46
|
-
### Known limitations
|
|
47
|
-
|
|
48
|
-
- A burst of commits fast enough to launch genuinely overlapping `sync --auto` processes can
|
|
49
|
-
interleave two runs' text mid-line in `.nexusmem/post-commit-sync.log` — the advisory lock
|
|
50
|
-
serializes the actual database writes (never at risk), not who gets to write to this diagnostic
|
|
51
|
-
log file. Cosmetic only; not planned to be fixed unless it turns out to matter in practice.
|
|
52
|
-
|
|
53
|
-
## [0.10.2] — 2026-08-30
|
|
54
|
-
|
|
55
|
-
### Security
|
|
56
|
-
|
|
57
|
-
- `shell_command` nodes never ran through the same secret-redaction pass `conversation_turn` and
|
|
58
|
-
`code_diff` nodes already get: a command like `export API_KEY=...` landed verbatim in `title`/`body`,
|
|
59
|
-
which is exactly what the FTS index and vector embeddings are built from — a secret typed at a
|
|
60
|
-
prompt could resurface later through `search_memory`/`nexusmem query`. Fixed by running `redact()`
|
|
61
|
-
over the command before it reaches those fields. `meta.command` is kept raw on purpose: project-id
|
|
62
|
-
reconciliation and failure/fix correlation both hash or exact-match against the real command text,
|
|
63
|
-
and redacting that copy too would have silently orphaned nodes on a project-id migration.
|
|
64
|
-
|
|
65
|
-
## [0.10.1] — 2026-08-30
|
|
66
|
-
|
|
67
|
-
### Added
|
|
68
|
-
|
|
69
|
-
- `nodes.retrieved_count`/`nodes.last_retrieved_at` (schema V12): retrieval outcomes are now recorded
|
|
70
|
-
instead of silently discarded. Bumped once per completed `nexusmem query`/MCP `search_memory` call,
|
|
71
|
-
for every node actually packed into the returned context (both CLI and MCP share one pipeline, so
|
|
72
|
-
both are covered from a single call site). Not yet folded into `rank.ts`'s score formula — every
|
|
73
|
-
existing ranking factor was tuned against real dogfooded queries and validated with `npm run eval`
|
|
74
|
-
before being trusted, and a new factor needs the same treatment. This release is the instrumentation
|
|
75
|
-
half only; confirmed eval-neutral (MRR 0.943 / Recall@5 0.964, unchanged).
|
|
76
|
-
|
|
77
|
-
### Fixed
|
|
78
|
-
|
|
79
|
-
- `nexusmem precheck`'s "What already failed here" warning implied the failure was located in the
|
|
80
|
-
flagged file, but the match is against basename word-tokens (`tokensForFile`) — a failing `npm run
|
|
81
|
-
precheck` flagged every file whose name contained that word, not just the one actually responsible.
|
|
82
|
-
Reworded to state what's actually true ("commands naming this file failed").
|
|
83
|
-
- The high-churn warning in `nexusmem precheck` was rendered as a `WARN`, the same severity as the
|
|
84
|
-
dogfooded failure-correlation warning, despite `HIGH_CHURN_THRESHOLD` being an admitted, untuned
|
|
85
|
-
guess. Demoted to a lower-severity `note`, explicitly labeled as an untuned heuristic, so it no
|
|
86
|
-
longer reads as equally trustworthy.
|
|
87
|
-
|
|
88
|
-
## [0.10.0] — 2026-08-29
|
|
89
|
-
|
|
90
|
-
### Added
|
|
91
|
-
|
|
92
|
-
- New opt-in `github` source: `sources.github.enabled`, `nexusmem scan-github`, `sync --github`.
|
|
93
|
-
Reads issue/PR threads (title, body, comments) from this repo's github.com remote via the `gh`
|
|
94
|
-
CLI, one `github_thread` node per thread. First source with a real external dependency (needs `gh`
|
|
95
|
-
installed and authenticated) rather than reading only what's already on disk; a missing remote or
|
|
96
|
-
an unreachable `gh` is a silent no-op, matching how an unreachable Ollama degrades elsewhere.
|
|
97
|
-
|
|
98
|
-
## [0.9.1] — 2026-08-29
|
|
99
|
-
|
|
100
|
-
Three ranking/retrieval correctness fixes, found and validated against a new 28-case retrieval-quality
|
|
101
|
-
eval harness (`npm run eval`, dev-only, not part of the published package).
|
|
102
|
-
|
|
103
|
-
### Fixed
|
|
104
|
-
|
|
105
|
-
- `nodes_vec` (vector search) over-fetched `k` globally, then filtered by `project_id` after the join —
|
|
106
|
-
a heuristic, not a guarantee. A sparse project sharing `memory.db` with a much larger one could have
|
|
107
|
-
every true nearest neighbour fall outside the over-fetch window and get silently dropped (reproduced:
|
|
108
|
-
495 rows in one project + 5 in another, global `k=50` surfaced 0 of the 5). Schema V11 gives `nodes_vec`
|
|
109
|
-
a `PARTITION KEY` on `project_id` (`sqlite-vec` 0.1.9+), pushing the equality filter into the k-NN
|
|
110
|
-
search itself so cross-project exactness is now guaranteed, not probable. Migrates existing databases
|
|
111
|
-
automatically on next `sync`/`query`. `--as-of` date filtering still over-fetches — only the
|
|
112
|
-
`project_id` dimension was ever a correctness guarantee.
|
|
113
|
-
- `MAX_PRIOR_OVERTURN` (the cap on how far a recency/signal prior may overturn relevance) raised
|
|
114
|
-
`2 → 2.4`, the highest value that both improves eval MRR (0.924→0.943) and still passes every legacy
|
|
115
|
-
regression test guarding against the original same-day-fix-commits bug this constant exists to bound.
|
|
116
|
-
- A commit's `code_diff` siblings (one node per changed file, all sharing the same `ts`) could crowd a
|
|
117
|
-
packed result and bury that commit's own `git_commit` node — its answer — several ranks down. Now
|
|
118
|
-
capped per commit, matching the existing `conversation_turn`/`doc_section` family cap.
|
|
119
|
-
- `package.json`/`tsconfig.json` diffs were ranking above more relevant results when a changed file's
|
|
120
|
-
path happened to echo its commit's conventional-commit scope (title is weighted 10x body, so the scope
|
|
121
|
-
word counted twice). Down-weighted as mechanical wiring, same reasoning `TEST_PATHS` already applies to
|
|
122
|
-
tests. Only affects nodes ingested from here on — existing manifest/config diffs need `sync --rebuild`
|
|
123
|
-
to get the corrected signal retroactively.
|
|
124
|
-
|
|
125
|
-
## [0.9.0] — 2026-08-27
|
|
126
|
-
|
|
127
|
-
Closes the three remaining mechanisms from a recurring external review (Simon Strandgaard, Agent
|
|
128
|
-
Memory Atlas): trust_state, bi-temporal reads, and a dismiss verb for standing suggestions. Also
|
|
129
|
-
fixes two real Windows-specific correctness bugs found dogfooding.
|
|
130
|
-
|
|
131
|
-
### Added
|
|
132
|
-
|
|
133
|
-
- `nodes.trust_state` (`candidate` | `verified` | `rejected`, schema V10) and `nexusmem review
|
|
134
|
-
<nodeId> --verify|--reject`: a human verdict on a node, independent of the SLM contradiction
|
|
135
|
-
checker. `--reject` down-weights ranking (harsher than the existing supersede penalty, since a
|
|
136
|
-
human said no directly); `--verify` is a label only, no ranking boost. Both render as a tag in
|
|
137
|
-
packed context. Never deletes — same demote-not-delete rule as `mark-stale`.
|
|
138
|
-
- `--as-of <date>` on `nexusmem query` and the MCP `search_memory` tool: restricts both the BM25
|
|
139
|
-
and vector arms to nodes recorded at or before that instant via `created_at` (record time),
|
|
140
|
-
independent of how old the events themselves (`ts`, event time) are. Read-only — there is no
|
|
141
|
-
equivalent write, and nodes stay write-once.
|
|
142
|
-
- `nexusmem stale --dismiss`: silences a contradiction suggestion the user disagrees with, without
|
|
143
|
-
fabricating a `supersedes` relationship. Previously a wrong `--check-contradictions` verdict had
|
|
144
|
-
no way to be rejected — it re-decorated every future `stale`/`sync` run forever.
|
|
145
|
-
- `--prune-source`/`--prune-stale-shell` now record a `mutation_audit` row on their `--yes` path,
|
|
146
|
-
matching what `forget` already did. The dry-run preview stays a pure read.
|
|
147
|
-
|
|
148
|
-
### Fixed
|
|
149
|
-
|
|
150
|
-
- `git rev-parse` failures reported as "error launching git: Access is denied." (Git for Windows'
|
|
151
|
-
launcher shim failing to exec `git.exe` under handle/AV contention) were treated as git's own
|
|
152
|
-
verdict instead of a transient spawn failure, turning a one-off environment hiccup into a hard
|
|
153
|
-
failure. Now retried the same as the other known transient-spawn classes.
|
|
154
|
-
- The PowerShell prompt hook read only `$LASTEXITCODE`, which cmdlets (`Remove-Item`, `Copy-Item`,
|
|
155
|
-
…) never set — every failing cmdlet was silently logged as a success, and a stale
|
|
156
|
-
`$LASTEXITCODE` from an earlier native command could misattribute to later successful cmdlet
|
|
157
|
-
runs. Now reads `$?` alongside `$LASTEXITCODE`, before anything else can overwrite `$?`.
|
|
158
|
-
|
|
159
|
-
## [0.8.0] — 2026-08-22
|
|
160
|
-
|
|
161
|
-
### Added
|
|
162
|
-
|
|
163
|
-
- Provenance widened from 2 tiers to a 4-tier trust hierarchy: `observed` (commits, diffs, shell
|
|
164
|
-
exit codes) > `authored` (doc sections — human-written claims) > `recorded` (conversation turns —
|
|
165
|
-
verbatim discourse) > `derived` (session summaries — a model's distillation). Schema V7 backfills
|
|
166
|
-
existing databases by kind. The ranker decays each tier at its own rate (lower trust fades
|
|
167
|
-
faster); the ordering is the design claim, the exact ratios are documented judgment calls.
|
|
168
|
-
- Contradiction checking now runs automatically during `sync`: at most 3 new SLM judgments per run
|
|
169
|
-
(`contradictions.maxPerSync`; `contradictions.autoCheck: false` turns it off), gated on the
|
|
170
|
-
embedding provider already being reachable. Every judgment — either verdict — is memoized in a new
|
|
171
|
-
`contradiction_checks` table (schema V8), so a judged pair is never sent to the model again and
|
|
172
|
-
repeat syncs converge to zero model calls. Measured live against this repo's own database: 5.4s
|
|
173
|
-
for the first 10 judgments, 0.5s for the identical re-run. Still suggest-only: nothing ever writes
|
|
174
|
-
`supersedes` automatically.
|
|
175
|
-
- Standing suggestions now surface everywhere without a model in the loop: plain `nexusmem stale`
|
|
176
|
-
decorates flagged candidates from the memoized judgments (instant, offline), `nexusmem status`
|
|
177
|
-
gains a `flagged` line, and the sync summary reports new/open suggestion counts.
|
|
178
|
-
- `stale --check-contradictions` reuses each candidate's stored embedding instead of re-embedding
|
|
179
|
-
it, and stops after two consecutive null SLM replies so a down provider costs at most two timeouts
|
|
180
|
-
rather than one per candidate.
|
|
181
|
-
|
|
182
|
-
## [0.7.0] — 2026-08-21
|
|
183
|
-
|
|
184
|
-
### Added
|
|
185
|
-
|
|
186
|
-
- `nexusmem stale --check-contradictions`: for each stale candidate, finds the most similar newer
|
|
187
|
-
node (local embedding search) and asks a local SLM (Ollama, `qwen2.5:3b` by default) whether it
|
|
188
|
-
actually contradicts the older one, instead of only surfacing by age. Suggest-only — nothing is
|
|
189
|
-
written, same as plain `stale`. Live-dogfooded against this repo's own real database and Ollama
|
|
190
|
-
instance; found and fixed a real gap along the way (below) before the feature surfaced anything
|
|
191
|
-
useful.
|
|
192
|
-
- Import graph: bare Python imports with no leading dot (`import foo`, `from foo import bar`)
|
|
193
|
-
now resolve to a same-directory sibling file too, alongside the existing relative-dot support.
|
|
194
|
-
Found dogfooding two real local projects: neither used a single PEP 328 relative import, both
|
|
195
|
-
relied entirely on this flat-script style. Guarded by a static list of stdlib module names so a
|
|
196
|
-
bare `import os`/`import queue`/etc. is never mistaken for a same-named local file.
|
|
197
|
-
|
|
198
|
-
### Fixed
|
|
199
|
-
|
|
200
|
-
- Import graph: a Java wildcard import (`import a.b.*;`) no longer merges files from two
|
|
201
|
-
unrelated packages that happen to share a directory-name suffix (e.g. two Gradle/Maven modules
|
|
202
|
-
each with their own `.../foo`) — it now refuses to guess, same as the single-class-import case.
|
|
203
|
-
- `stale --check-contradictions`'s neighbor search now looks past same-timestamp sibling nodes
|
|
204
|
-
(e.g. the many chunks one long conversation gets split into) to reach genuinely newer content.
|
|
205
|
-
Found live-dogfooding against this repo's own database: the first real candidate's closest 15
|
|
206
|
-
neighbors were all same-conversation siblings sharing its exact timestamp, so the original
|
|
207
|
-
5-neighbor default silently found zero suggestions for every candidate, regardless of what the
|
|
208
|
-
SLM would have said.
|
|
209
|
-
|
|
210
|
-
## [0.6.0] — 2026-08-20
|
|
211
|
-
|
|
212
|
-
### Added
|
|
213
|
-
|
|
214
|
-
- Ranker: `inferred` nodes (summaries, doc snapshots) now decay twice as fast as `observed` ones
|
|
215
|
-
(git commits, diffs, shell commands) for retrieval purposes. A judgment call, not a measured
|
|
216
|
-
optimum — see `INFERRED_HALF_LIFE_RATIO` in `src/retrieval/rank.ts`.
|
|
217
|
-
- `nexusmem stale`: lists aging `inferred` nodes nothing has superseded yet, as candidates for
|
|
218
|
-
`mark-stale`. A heuristic on age and provenance, not real contradiction detection — writes nothing.
|
|
219
|
-
- `nexusmem status` now surfaces an `aging` line with the stale-candidate count when any exist.
|
|
220
|
-
- Import graph: Python relative imports (`from .foo import bar`, `from . import x`) now produce
|
|
221
|
-
file edges too, alongside the existing JS/TS support. Absolute Python imports are still skipped —
|
|
222
|
-
same "missed edge over wrong edge" reasoning as JS/TS's bare-specifier skip.
|
|
223
|
-
- Import graph: Go internal imports (resolved against `go.mod`'s module path, one edge per
|
|
224
|
-
non-test file in the imported package) and Rust `mod foo;` declarations (2018+ edition module
|
|
225
|
-
layout) now produce file edges too. External Go imports and Rust `use` paths are out of scope for
|
|
226
|
-
the same reason.
|
|
227
|
-
- Import graph: Java imports (`import a.b.C;` / `import a.b.*;`, resolved by unambiguous suffix
|
|
228
|
-
match against the tracked source tree) and `__DIR__`-anchored PHP `require`/`include` now produce
|
|
229
|
-
file edges too — all six of `nexusmem scan-structure`'s tracked languages. `import static`, PHP's
|
|
230
|
-
autoloaded `use Namespace\Class;`, and any unanchored PHP include are out of scope for the same
|
|
231
|
-
"missed edge over wrong edge" reason as everywhere else in the import graph.
|
|
232
|
-
|
|
233
|
-
## [0.5.4] — 2026-08-20
|
|
234
|
-
|
|
235
|
-
### Added
|
|
236
|
-
|
|
237
|
-
- `nexusmem status --share`: a plain-text, no-color summary (node count, failure→fix chains
|
|
238
|
-
linked, days of history) meant to be pasted somewhere, not scraped by a script.
|
|
239
|
-
|
|
240
|
-
## [0.5.3] — 2026-08-19
|
|
241
|
-
|
|
242
|
-
No functional changes to the CLI, MCP server, or published package -- a test-coverage and
|
|
243
|
-
CI-reliability pass.
|
|
244
|
-
|
|
245
|
-
### Verified
|
|
246
|
-
|
|
247
|
-
- **Closed 15 genuine test-coverage gaps**, found by grepping for real call sites in `tests/`
|
|
248
|
-
rather than guessing from filenames (an Explore agent found the first 7; the rest by hand the
|
|
249
|
-
same way). All 7 `scan-*` CLI subcommands (`scan-git`/`diff`/`shell`/`docs`/`conversation`/
|
|
250
|
-
`structure`/`session`) had zero coverage -- `tests/cli-scan.test.ts`'s name was misleading, it
|
|
251
|
-
only pinned `cli/format.ts`'s shared helpers. Also closed: `runStatus`'s report body beyond the
|
|
252
|
-
stale-identity branch, `runQuery`, `forget --import`'s dry-run preview path, `shell/detect.ts`'s
|
|
253
|
-
zsh scraping and hook-log `repoRoot` scoping, `openAllProjectSources`'s `unreadable` branch,
|
|
254
|
-
`cli/context.ts`'s `loadContext`, `store/fts.ts`'s `toStrictMatchQuery`/`significantTokens`,
|
|
255
|
-
`list_recent_memory`'s MCP `structuredContent` at the protocol level, `structure/resolve.ts`'s
|
|
256
|
-
`.mjs`/`.cjs` rewrite, `docs/read.ts`'s `include` option, and `git/log.ts`'s `buildLogArgs`. Test
|
|
257
|
-
suite: 559 → 592 tests.
|
|
258
|
-
|
|
259
|
-
### Fixed
|
|
260
|
-
|
|
261
|
-
- **A real test-isolation bug: a test could scrape a developer's actual shell history into a
|
|
262
|
-
throwaway test database.** `tests/setup.ts` isolated `NEXUSMEM_HOME` (so a test `sync()` can't
|
|
263
|
-
pollute the real `~/.nexusmem/projects.json` registry -- fixed 2026-08-15 for the same reason)
|
|
264
|
-
but never isolated `APPDATA`/`HISTFILE_BASH`/`HISTFILE`, so any test running a real `sync()`
|
|
265
|
-
scraped whatever PSReadLine/bash/zsh history actually exists on the machine running the suite.
|
|
266
|
-
Observed directly: a new test here ingested 300 real `shell_command` nodes off this machine's
|
|
267
|
-
own history before the fix. Now isolated the same way as the registry.
|
|
268
|
-
- **CI failed (deterministically, all 4 jobs) on 4 assertions that assumed adjacent words in
|
|
269
|
-
colorized CLI output stay contiguous.** GitHub Actions' runners don't have `NO_COLOR` set the
|
|
270
|
-
way this project's local dev shells apparently do, so `picocolors` emits real ANSI codes there --
|
|
271
|
-
a regex like `/git\s+last run/` breaks when an escape sequence sits between a padded label and a
|
|
272
|
-
separately-colored value. `tests/forget.test.ts` already carries a `stripAnsi()` helper for
|
|
273
|
-
exactly this trap; applied the same fix to the 4 newly-added test files that had it. Verified by
|
|
274
|
-
reproducing the CI failure locally with `FORCE_COLOR=1` before the fix, then re-running the full
|
|
275
|
-
592-test suite under `FORCE_COLOR=1` after, to check for any other latent instance.
|
|
276
|
-
|
|
277
|
-
## [0.5.2] — 2026-08-18
|
|
278
|
-
|
|
279
|
-
### Added
|
|
280
|
-
|
|
281
|
-
- **`nexusmem mark-stale <nodeId> --supersedes <newNodeId>`, `provenance`, and manual staleness
|
|
282
|
-
tracking.** A pre-Show-HN pass answering a public critique that this couldn't tell an observed
|
|
283
|
-
fact from a guess, or retire a stale conclusion. Every node now carries `provenance`
|
|
284
|
-
(`observed`/`inferred`, set per-collector: git commits, diffs, and shell commands are `observed`;
|
|
285
|
-
conversation turns, session summaries, and doc sections are `inferred`) and an optional
|
|
286
|
-
`supersedes` link. `mark-stale` writes that link; the ranker applies a flat down-weight to whatever
|
|
287
|
-
it points at, but never deletes it -- the old node stays queryable, just usually loses to its
|
|
288
|
-
replacement. `provenance` is shown as a `[observed]`/`[inferred]` tag on every query result (CLI
|
|
289
|
-
and MCP `search_memory` share the same renderer). This is deliberately not automatic: nothing
|
|
290
|
-
detects staleness on its own, a human or agent still has to notice the contradiction and run the
|
|
291
|
-
command. See the README's new "Manual staleness & provenance" section.
|
|
292
|
-
|
|
293
|
-
### Verified
|
|
294
|
-
|
|
295
|
-
- **`forget` does not leak through vector search.** Checked whether `MemoryStore.forget` excludes a
|
|
296
|
-
tombstoned node from `sqlite-vec` results only via a query-time filter (which every vector query
|
|
297
|
-
path would need to apply consistently) or by deleting the embedding row outright. It already does
|
|
298
|
-
the latter -- `dropEmbedding` runs before the node row is deleted, in the same transaction, on both
|
|
299
|
-
query paths (`runHybridQuery` and `runCrossProjectQuery`, which both call the same
|
|
300
|
-
`MemoryStore.vectorSearch`). No fix was needed; added a regression test
|
|
301
|
-
(`tests/vector.test.ts`) that forgets an embedded node and asserts both `vectorSearch` returns
|
|
302
|
-
nothing and the `nodes_vec` row count drops to zero, so this can't silently regress.
|
|
303
|
-
- **`forget` survives `sync --rebuild`.** This was the point of v0.5.0 and was already covered
|
|
304
|
-
end-to-end in `tests/forget.test.ts`. Added a lower-level `tests/store.test.ts` case pinning the
|
|
305
|
-
exact mechanism: `clearProject` (what `--rebuild` calls before re-ingesting) only touches
|
|
306
|
-
`nodes`/`nodes_vec`/`sync_state`, never `deny_list`, so a re-ingest of the same content is denied
|
|
307
|
-
again rather than resurrected.
|
|
308
|
-
|
|
309
|
-
## [0.5.1] — 2026-08-17
|
|
310
|
-
|
|
311
|
-
### Added
|
|
312
|
-
|
|
313
|
-
- **`nexusmem forget --export <path>` / `forget --import <path>`: carry a deny-list across a clone or
|
|
314
|
-
restore.** `.nexusmem/` is gitignored by design, so `deny_list` never traveled with `git clone`/
|
|
315
|
-
`git push` — confirmed live 2026-08-17 that a fresh clone of the exact same repo resurrected a value
|
|
316
|
-
already forgotten elsewhere, with zero deny-list protection, because git history (what a fresh `sync`
|
|
317
|
-
re-derives from) is fully portable while the deny-list that would have blocked it was not. `--export`
|
|
318
|
-
writes the active entries to a plaintext JSON file (loudly warned as exactly as sensitive as the
|
|
319
|
-
values it holds — never meant for git, moved through whatever secure channel the user already
|
|
320
|
-
trusts); `--import` re-applies each new entry through `forget` itself, so an imported value is
|
|
321
|
-
deleted from the new checkout's nodes too, not just blocked going forward. Same dry-run-by-default /
|
|
322
|
-
`--yes` convention as the rest of `forget`. See `docs/forget-mechanism.md`.
|
|
323
|
-
|
|
324
|
-
## [0.5.0] — 2026-08-17
|
|
325
|
-
|
|
326
|
-
### Added
|
|
327
|
-
|
|
328
|
-
- **`nexusmem forget <value>`: permanent, value-keyed deletion.** The finer-grained complement to
|
|
329
|
-
`sync --prune-source` (which only deletes at the whole-collector-source granularity): forgets one
|
|
330
|
-
exact string or `--regex` pattern, deleting every node it currently matches *and* writing a
|
|
331
|
-
standing deny-list entry so the value can never be re-ingested — closing the one gap an external
|
|
332
|
-
source-level review flagged as its most serious finding: the append-only shell-hook log (and a
|
|
333
|
-
full transcript re-read) are untouched by any prune, so `sync --rebuild` used to resurrect exactly
|
|
334
|
-
what had just been deleted. Every removal leaves a hash-only tombstone (never the forgotten
|
|
335
|
-
content itself — `body`/`title` are stored as sha256 only) and the whole operation writes one
|
|
336
|
-
`mutation_audit` row, whether or not anything matched. Dry-run by default, matching
|
|
337
|
-
`--prune-source`'s convention exactly; `--yes` confirms; `--list` shows active entries. Permanent
|
|
338
|
-
in v0 — no `--remove`, matching the existing irreversible framing of `--prune-source`/`--rebuild`.
|
|
339
|
-
See `docs/forget-mechanism.md`.
|
|
340
|
-
|
|
341
|
-
## [0.4.0] — 2026-08-16
|
|
342
|
-
|
|
343
|
-
### Added
|
|
344
|
-
|
|
345
|
-
- **`nexusmem status` now warns when prior project identities still hold nodes.** A renamed git
|
|
346
|
-
remote can leave stale data under the old identity; status reports both the identity and node
|
|
347
|
-
counts and points to `sync --prune-source <name>` instead of leaving that data discoverable only
|
|
348
|
-
through a raw SQLite query.
|
|
349
|
-
- **`docs/competitor-comparison.md`: an honest, source-level comparison against `riponcm/projectmem`** (705
|
|
350
|
-
GitHub stars vs. this project's 8, at time of writing) — a real, working competitor read in full (not
|
|
351
|
-
judged from its README), covering where the two tools' failure-tracking, import-graph coverage, and
|
|
352
|
-
token-savings claims genuinely differ. Cites a freshly re-run benchmark (`npm run bench`, not a stale
|
|
353
|
-
table) alongside projectmem's own `pjm score` constants, read directly from its source. Linked from the
|
|
354
|
-
README's `## What it costs you` section and a short website section near the existing stats.
|
|
355
|
-
- **`nexusmem hook git install|remove|status`: a real git pre-commit hook.** Installs a marked block into
|
|
356
|
-
`.git/hooks/pre-commit` that runs `nexusmem precheck` (no `--strict`, so it can never block a commit on
|
|
357
|
-
its own) before each commit — the automatic counterpart to the advisory `precheck` command. Refuses to
|
|
358
|
-
touch a pre-existing foreign hook (husky, lint-staged, lefthook, ...) unless `--force` is passed, in which
|
|
359
|
-
case it appends after the existing content rather than before, so the foreign hook still runs first and
|
|
360
|
-
keeps deciding whatever it already decided. Idempotent (`nexusmem hook git install` twice is a no-op) and
|
|
361
|
-
removable cleanly, including restoring a foreign hook to its original content if one was appended onto.
|
|
362
|
-
Live-verified against a real scratch git repo on Windows: fresh install, a real `git commit` that
|
|
363
|
-
triggered the hook and printed a correct precheck report, clean removal, and both the refuse-without-force
|
|
364
|
-
and append-with-force foreign-hook paths.
|
|
365
|
-
- **`nexusmem precheck`: proactive pre-commit warnings.** Checks staged (or `--working`, or explicit `--files`)
|
|
366
|
-
files against project memory and warns about unresolved past failures and high recent churn *before* you
|
|
367
|
-
commit — advisory by default (always exits 0; `--strict` turns an unresolved failure into a non-zero exit).
|
|
368
|
-
Matches a file's own basename tokens against still-unlinked `shell_command` failures (no
|
|
369
|
-
`resolved_by:*` link from `sync --link-failures`), reusing `filterBoilerplateTokens` from the
|
|
370
|
-
discussion-bridge heuristic — now exported and parameterized by node kind so it can be corpus-relative
|
|
371
|
-
against `shell_command` history instead of only `conversation_turn`/`session_summary`. Churn is scoped to
|
|
372
|
-
`git_commit`-kind `node_files` touches only, so it doesn't double-count the identical touches `code_diff`
|
|
373
|
-
nodes also record. Deliberately does not yet install as a real git hook (see the module comment in
|
|
374
|
-
`src/correlate/precheck.ts` for why capture-time `git status` diffing and past-commit correlation were both
|
|
375
|
-
rejected) — that's a follow-up once this signal has been dogfooded.
|
|
376
|
-
- **JS/TS import-graph edges.** `nexusmem scan-structure` previews (and `sync` now ingests) file→file
|
|
377
|
-
import relationships across a project's tracked `.ts`/`.tsx`/`.js`/`.jsx` files — a dependency-free
|
|
378
|
-
regex extractor (`src/structure/extract.ts`) resolves relative `import`/`export ... from`/`require`/
|
|
379
|
-
dynamic-`import` specifiers against the tracked-path set, correctly rewriting the common TS/ESM
|
|
380
|
-
`./foo.js` specifier back to its real `foo.ts` source. Stored in a new `file_edges` table (schema
|
|
381
|
-
v4), replaced wholesale on every sync since edges describe current tree state, not history. Surfaced
|
|
382
|
-
as a `structure` line in `nexusmem status`; not yet wired into `query` ranking or exposed as an MCP
|
|
383
|
-
tool — that's a follow-on design question, not this pass's job.
|
|
384
|
-
|
|
385
|
-
## [0.3.3] — 2026-08-16
|
|
386
|
-
|
|
387
|
-
### Added
|
|
388
|
-
|
|
389
|
-
- **`nexusmem status` now shows failure→fix chain counts** — `N/M failure(s) resolved (X retry, Y
|
|
390
|
-
discussion)`, plus a hint to run `sync --link-failures` when failures remain unresolved. The
|
|
391
|
-
chain feature is this project's most distinctive capability, but was previously invisible to
|
|
392
|
-
anyone who didn't already know to query `node_links` directly. Backed by the new
|
|
393
|
-
`getChainStats` in `src/correlate/failure-fix.ts`, which dedupes failures resolved by both
|
|
394
|
-
heuristics rather than double-counting them.
|
|
395
|
-
|
|
396
|
-
### Fixed
|
|
397
|
-
|
|
398
|
-
- **Discussion-bridge heuristic (`sync --link-failures`): corpus-relative boilerplate tokens no
|
|
399
|
-
longer produce false-positive failure→fix links.** Re-dogfooded at larger scale against a second
|
|
400
|
-
real project's history: a command made entirely of words that saturate a project's own corpus
|
|
401
|
-
(e.g. this repo's own name/verbs) could AND-match an unrelated turn that just happened to mention
|
|
402
|
-
the same words, and bm25 score alone could not separate that from a true positive (measured: the
|
|
403
|
-
false positive scored *stronger* than two real true positives). `filterBoilerplateTokens` in
|
|
404
|
-
`src/correlate/failure-fix.ts` drops any token that appears in over 20% of a project's own
|
|
405
|
-
`conversation_turn`/`session_summary` history before building the match query — measured against
|
|
406
|
-
real data, not guessed — and skips the discussion-match attempt entirely (rather than falling back
|
|
407
|
-
unfiltered) when every token turns out to be boilerplate, since a missed link is preferred over a
|
|
408
|
-
false one for this heuristic. Below 10 discussable nodes the filter is skipped, since frequency
|
|
409
|
-
isn't a meaningful signal yet on a young project.
|
|
410
|
-
|
|
411
|
-
## [0.3.2] — 2026-08-16
|
|
412
|
-
|
|
413
|
-
### Added
|
|
414
|
-
|
|
415
|
-
- **`mcpName` field in `package.json`**, required by the official MCP registry
|
|
416
|
-
(registry.modelcontextprotocol.io) to verify that whoever publishes `server.json` under
|
|
417
|
-
`io.github.yaminbkk/nexusmem` also controls the `nexusmem` npm package itself — the registry
|
|
418
|
-
rejects a publish attempt otherwise. No behavior change for CLI/MCP users; this is purely a
|
|
419
|
-
registry-ownership proof.
|
|
420
|
-
- **`list_recent_memory` MCP tool** — chronological listing of a repository's most recently
|
|
421
|
-
remembered nodes (git commits, diffs, shell commands, docs, conversation, session summaries),
|
|
422
|
-
newest first. Distinct from `search_memory`: no query, just "what has this project's memory
|
|
423
|
-
recorded lately" — built for the VS Code extension's sidebar view, which lists rather than
|
|
424
|
-
searches. Backed by `MemoryStore.listRecentNodes`, reusing the existing `idx_nodes_project_ts`
|
|
425
|
-
index.
|
|
426
|
-
- **`sync --prune-source <name>` and `sync --prune-stale-shell`** — drop one source's nodes without a
|
|
427
|
-
full `--rebuild`, which loses history that can't be re-read from disk (the shell tail window, older
|
|
428
|
-
conversation turns). `--prune-stale-shell` is a shortcut for the three dead pre-hook shell-scrape
|
|
429
|
-
sources (`shell:pwsh`, `shell:bash`, `shell:zsh`) at once. Dry-run by default — prints the matching
|
|
430
|
-
count and does nothing until `--yes` is also given, since this is an irreversible full wipe of the
|
|
431
|
-
named source(s), unlike `--rebuild`'s no-prompt full-project reset. Also sweeps any prior project
|
|
432
|
-
identity of this same repo (the id a renamed git remote leaves behind after
|
|
433
|
-
`fix(store): reconcile memory stranded by a changed git remote URL` migrates what it can) — a
|
|
434
|
-
live-id-only prune could not reach nodes reconciliation deliberately left in place. Exposed on both
|
|
435
|
-
the CLI and the MCP `sync_project` tool.
|
|
436
|
-
- **Discussion-heuristic failure→fix chains now surface**, tightened to an AND-joined significant-
|
|
437
|
-
token match instead of the original OR match. Re-verified against this repo's own real database:
|
|
438
|
-
5/5 discussion links correct (was ~half wrong when it shipped unsurfaced in 0.3.0). Chains now
|
|
439
|
-
follow across projects in `query --all-projects` too.
|
|
440
|
-
|
|
441
|
-
### Fixed
|
|
442
|
-
|
|
443
|
-
- **`sync_project`'s summary no longer contains raw ANSI color codes on Windows.** `runInit`/`runSync`
|
|
444
|
-
format their output with picocolors for terminal display, and picocolors treats `platform ===
|
|
445
|
-
'win32'` as sufficient evidence of color support on its own, without checking `isTTY` — correct for
|
|
446
|
-
a real terminal, wrong for the MCP JSON-RPC channel, which is piped on every platform. Found live: a
|
|
447
|
-
real MCP client (the VS Code extension's Output channel) rendered the raw escape codes as literal
|
|
448
|
-
text instead of color. Stripped at the MCP boundary in `syncProject`, leaving the CLI's own terminal
|
|
449
|
-
output untouched.
|
|
450
|
-
|
|
451
|
-
## [0.3.1] — 2026-08-15
|
|
452
|
-
|
|
453
|
-
### Added
|
|
454
|
-
|
|
455
|
-
- **`scripts/benchmark.ts` (`npm run bench`) — a reproducible end-to-end token-saving benchmark.**
|
|
456
|
-
Compares `packed.tokensUsed` against two baselines (full-file-read and `git log -p`) for the same
|
|
457
|
-
files a query's packed nodes touch, over a query set derived mechanically from the corpus itself
|
|
458
|
-
rather than hand-picked. Used to measure the README's `## What it costs you` numbers against both
|
|
459
|
-
this repo (62 commits) and a 9,567-commit external corpus (`vitejs/vite`) — see README for the
|
|
460
|
-
numbers and their methodology caveats.
|
|
461
|
-
|
|
462
|
-
### Fixed
|
|
463
|
-
|
|
464
|
-
- **A generic 2-3 letter local model title, "id" as a search token, and one node's chunks flooding a
|
|
465
|
-
result set — three ranking/retrieval edge cases found by dogfooding, each verified with a red test
|
|
466
|
-
before the fix.**
|
|
467
|
-
- Session-summary titles: the local model sometimes wrote a role-framing line ("Role: Lead Systems
|
|
468
|
-
Engineer...") instead of a summary, in Thai and English alike. The existing generic-title filter
|
|
469
|
-
didn't catch either language, since it only matched English words like "summary"/"update".
|
|
470
|
-
- Search: the token `id` alone was prefix-matching unrelated shell commands like
|
|
471
|
-
`winget install --id ...`, because every query token was OR-ed with no floor and no stopword
|
|
472
|
-
list, and bm25 gives a rare-but-generic token an inflated score purely from scarcity.
|
|
473
|
-
- Packing: `conversation_turn` and `doc_section` both chunk one reply or file into several nodes
|
|
474
|
-
sharing the same timestamp; up to 2 may now appear in one packed result, down from unlimited.
|
|
475
|
-
|
|
476
|
-
- **A repo's memory could silently split in two if its git remote URL ever changed** (a GitHub
|
|
477
|
-
account rename, an org transfer). Project identity is derived from the remote URL on purpose — so
|
|
478
|
-
the same repo re-cloned to a new path or machine keeps sharing memory — but a changed URL on the
|
|
479
|
-
*same* path minted a new id and stranded every node synced under the old one, invisible to
|
|
480
|
-
`status`/`query`/MCP from then on. `sync` now detects a prior id already recorded in the repo's own
|
|
481
|
-
database and reconciles it forward: recomputable node kinds (session summaries, hook-sourced shell
|
|
482
|
-
history) are migrated under their correct new id and deduplicated against anything already synced;
|
|
483
|
-
conversation turns, whose identity can't be recomputed, are reassigned in place. Git commits,
|
|
484
|
-
diffs, and doc sections are left alone — a normal sync already re-derives them completely, so
|
|
485
|
-
there is nothing to migrate.
|
|
486
|
-
|
|
487
|
-
## [0.3.0] — 2026-08-13
|
|
488
|
-
|
|
489
|
-
**Upgrade note.** The first `sync` after upgrading drops every stored embedding and rebuilds it.
|
|
490
|
-
This is not optional and it is not a bug: embeddings now come from Ollama's `/api/embed`, which
|
|
491
|
-
returns L2-normalised vectors, while the previous `/api/embeddings` did not — measured norms of 1.0
|
|
492
|
-
and 20.7 for the same input. `nodes_vec` ranks by Euclidean distance and records no per-row
|
|
493
|
-
provenance, so a corpus holding both would separate by scale rather than by meaning. `sync` says
|
|
494
|
-
what it dropped, nodes are untouched, and BM25 keeps working while the rebuild runs. On this repo
|
|
495
|
-
the rebuild was 833 nodes in one pass, inside a 9-second sync.
|
|
496
|
-
|
|
497
|
-
### Added
|
|
498
|
-
|
|
499
|
-
- **Session summaries via a local model** (`sources.session`, opt-in, off by default). Each
|
|
500
|
-
finished session becomes one distilled `session_summary` node — decisions and their reasons —
|
|
501
|
-
alongside the raw exchanges. Runs a local Ollama chat model (`qwen2.5:3b` by default); nothing is
|
|
502
|
-
downloaded automatically and no transcript leaves the machine. New `scan-session` command, with
|
|
503
|
-
`--dry-run` to print the exact prompt a session would produce without calling the model.
|
|
504
|
-
Bounded three ways: a session must be quiet for `settleMinutes` (default 30) before it is
|
|
505
|
-
eligible, the prompt is hashed so an unchanged session never reaches the model again, and
|
|
506
|
-
`maxSessions` caps how many are summarized per sync. Every exchange is redacted before the model
|
|
507
|
-
sees it, and the model's own output is redacted again before it is stored.
|
|
508
|
-
- **`sync --embed-limit <n>`** to cap the embedding pass, for when draining the whole backlog is not
|
|
509
|
-
wanted.
|
|
510
|
-
|
|
511
|
-
### Changed
|
|
512
|
-
|
|
513
|
-
- **The embedding pass drains the backlog in one `sync`** instead of stopping after 200 nodes, and
|
|
514
|
-
sends texts to Ollama in batches of 32 — measured at 4.21x the throughput of one call per node
|
|
515
|
-
(20.3ms → 4.8ms per node, over 96 real nodes from this repo). Leaving it uncapped is safe because
|
|
516
|
-
paging walks rowids monotonically, so a node the provider failed on is passed over rather than
|
|
517
|
-
retried forever, and because three consecutive dead requests end the pass: an Ollama that is not
|
|
518
|
-
running now costs three requests rather than one timeout per node.
|
|
519
|
-
- **Embeddings carry a provider identity.** Changing the embedding model, or upgrading from a
|
|
520
|
-
release that recorded no identity, drops the vectors and re-embeds rather than ranking across a
|
|
521
|
-
mixture.
|
|
522
|
-
- **Ranking priors now share one budget instead of getting one each.** `signal` and `recency` are
|
|
523
|
-
query-independent, and each was separately capped at overturning a 2× relevance gap. The score
|
|
524
|
-
multiplies them, so together they could overturn 4× — which is not a corner case but a description
|
|
525
|
-
of every commit made during an active working day, fresh and high-signal at once. A query about the
|
|
526
|
-
PowerShell hook returned two unrelated same-day `fix:` commits at ranks 3 and 4 while the section
|
|
527
|
-
that answered it sat at rank 6. The 2× budget is now the bound on the priors *jointly*, split
|
|
528
|
-
between them (`signal^0.215 × recency^0.288`, down from `^0.431` and `^0.576`), and a third prior
|
|
529
|
-
would re-divide the same budget rather than enlarge it. Measured on four real queries against this
|
|
530
|
-
repository's memory: the answering section rose in three of them — the rationale for "why BM25
|
|
531
|
-
before vector search" went from rank 4 to rank 1 — and no query's correct top hit was displaced.
|
|
532
|
-
|
|
533
|
-
### Known limitation
|
|
534
|
-
|
|
535
|
-
- Session-summary *titles* depend on the model following a fixed output format, and a 3B model often
|
|
536
|
-
does not. Measured over 14 real sessions, roughly a third came back usable; the rest were
|
|
537
|
-
conversational preambles, stray bullets, or a bare "Summary of the Session". Those are rejected
|
|
538
|
-
and the title falls back to the first line of the question that opened the session — always
|
|
539
|
-
specific, not always elegant. Compliance was worst on long sessions and on transcripts not in
|
|
540
|
-
English. `sources.session.model` takes a larger model if it matters.
|
|
541
|
-
|
|
542
|
-
## [0.2.0] — 2026-08-12
|
|
543
|
-
|
|
544
|
-
**Upgrade note.** Both new sources are on by default, so the first `sync` after upgrading an
|
|
545
|
-
existing project ingests the patches of its 200 most recent commits and starts recording the
|
|
546
|
-
repository in `~/.nexusmem/projects.json`. Set `sources.diff.enabled` to `false` in
|
|
547
|
-
`.nexusmem/config.json` if you would rather not, and `NEXUSMEM_HOME` relocates the user-scoped
|
|
548
|
-
directory. Nothing existing is rewritten or lost.
|
|
549
|
-
|
|
550
|
-
### Added
|
|
551
|
-
|
|
552
|
-
- **Diff-level nodes.** Commit patches are now indexed, one node per changed file, so a question
|
|
553
|
-
about *what the change looked like* reaches the lines themselves rather than the commit message
|
|
554
|
-
and a `+41/-6` summary. New `code_diff` kind, `diff` source, `scan-diff` preview command, and a
|
|
555
|
-
`sources.diff` config block. Read by a second `git log --patch` walk with its own cursor: folding
|
|
556
|
-
it into the existing `--numstat` walk would put patch text and numstat rows in one field, where a
|
|
557
|
-
diff line reading `-1\t2\tfoo` is indistinguishable from a real file entry.
|
|
558
|
-
Bounded on purpose — 200 commits on a first sync, 20 files per commit, no merges (their patch
|
|
559
|
-
exists only in a combined format this parser does not read), and binaries, lockfiles and build
|
|
560
|
-
output skipped. Patches are redacted with the shape-matching rules only; the key/value rule that
|
|
561
|
-
serves prose would rewrite `const apiKey = process.env.SERVICE_API_KEY` into a redaction marker.
|
|
562
|
-
- **Cross-project recall.** `query --all-projects` (and `search_memory`'s `allProjects`) searches
|
|
563
|
-
every repository NexusMem has been run in on this machine, tagging each result with the repository
|
|
564
|
-
it came from. Databases stay per-repository — a shared global store was rejected for giving up the
|
|
565
|
-
property that deleting one repo's `.nexusmem/` removes that repo's memory and nothing else — so a
|
|
566
|
-
plain index at `~/.nexusmem/projects.json`, written by `init` and refreshed by `sync`, is what
|
|
567
|
-
makes the others findable. New `projects` command lists it; `--prune` forgets entries whose
|
|
568
|
-
database is gone. A stale or corrupt registry degrades the query, never fails it.
|
|
569
|
-
Ranking fuses each project's list by rank (RRF) instead of comparing raw BM25 costs, which are
|
|
570
|
-
computed against their own corpus and are not comparable across databases. The bias this leaves —
|
|
571
|
-
every project's rank-1 hit is worth the same, so recall favours breadth — is documented rather
|
|
572
|
-
than hidden.
|
|
573
|
-
- **Query-aware diff excerpts.** A packed summary is ~320 characters and a patch is thousands, so
|
|
574
|
-
the packer now picks the hunk whose tokens match the query and starts the excerpt at the changed
|
|
575
|
-
line. Found by dogfooding: "what flags are passed to every git invocation" retrieved the right
|
|
576
|
-
file and then spent the whole summary on a class definition seventy lines above the answer.
|
|
577
|
-
Matching splits identifiers on case and underscore boundaries, because `\bretry\b` does not match
|
|
578
|
-
`RETRY_DELAYS_MS` and a natural-language question otherwise never meets the code it is about.
|
|
579
|
-
- `CHANGELOG.md` now ships inside the npm tarball. npm's always-included list covers `package.json`,
|
|
580
|
-
`README` and `LICENSE` but not the changelog, so it previously reached GitHub readers only.
|
|
581
|
-
|
|
582
|
-
### Internal
|
|
583
|
-
|
|
584
|
-
- The test suite no longer writes to the developer's real `~/.nexusmem`. `sync` records the
|
|
585
|
-
repository it ingested in the project registry, and the suite syncs temporary repositories in
|
|
586
|
-
several places, so a green run left seven dead entries behind — found by running `nexusmem
|
|
587
|
-
projects` after the fact, not by any test. `tests/setup.ts` now points `NEXUSMEM_HOME` at a
|
|
588
|
-
throwaway directory for the whole suite, and one test fails if that guard is ever removed.
|
|
589
|
-
- `npm run smoke` drives the *packaged* artifact: build, pack, install into a throwaway directory,
|
|
590
|
-
then run the installed CLI, an end-to-end ingest/query against a fixture repository, and an
|
|
591
|
-
`initialize` handshake over real stdio. It also audits the manifest `npm publish` would send,
|
|
592
|
-
which is a different artifact from the tarball. Both defects that ever reached npm users passed a
|
|
593
|
-
green unit suite first; each is now pinned by a check verified to fail when the defect is
|
|
594
|
-
reintroduced. Runs in CI on Linux and Windows as its own job.
|
|
595
|
-
|
|
596
|
-
## [0.1.2] — 2026-08-10
|
|
597
|
-
|
|
598
|
-
### Fixed
|
|
599
|
-
|
|
600
|
-
- `nexusmem --version` printed `0.1.0` on 0.1.1. The version string in `src/cli/index.ts` was a
|
|
601
|
-
literal separate from `package.json`, and the 0.1.1 bump only touched the latter. `src/mcp/server.ts`
|
|
602
|
-
had the same problem in its `McpServer` constructor, so an MCP client's `initialize` handshake
|
|
603
|
-
would have reported the same stale version. Both now read the real version through
|
|
604
|
-
`readOwnVersion()` in `src/core/version.ts`, which resolves `package.json` via `import.meta.url`.
|
|
605
|
-
Found by running the published package end to end rather than trusting `npm publish --dry-run`
|
|
606
|
-
and the registry API, neither of which executes a `--version` flag.
|
|
607
|
-
|
|
608
|
-
The ingestion and retrieval pipeline was never affected — only the two places that report a version
|
|
609
|
-
independently of running a command.
|
|
610
|
-
|
|
611
|
-
## [0.1.1] — 2026-08-10
|
|
612
|
-
|
|
613
|
-
### Changed
|
|
614
|
-
|
|
615
|
-
- README rewritten for someone deciding whether to read the source: what it does, how retrieval
|
|
616
|
-
scores, what it costs, and where it breaks. `README.md` ships inside the package, so this is a
|
|
617
|
-
real change to what npm delivers — but no code changed between 0.1.0 and 0.1.1.
|
|
618
|
-
- Documented the ranking flaw the tool found in itself, and the fact that the conversation source
|
|
619
|
-
in the sample `status` output is opt-in rather than default.
|
|
620
|
-
- Dropped `&&` from the quickstart, which Windows PowerShell 5.1 cannot parse.
|
|
621
|
-
|
|
622
|
-
## [0.1.0] — 2026-08-10
|
|
623
|
-
|
|
624
|
-
First public release.
|
|
625
|
-
|
|
626
|
-
### Added
|
|
627
|
-
|
|
628
|
-
- **Collectors.** Git history (commit metadata and diff stats, not diff bodies), shell commands with
|
|
629
|
-
exit codes via an opt-in PowerShell hook, tracked markdown docs via `git ls-files -- '*.md'`, and
|
|
630
|
-
opt-in assistant transcripts.
|
|
631
|
-
- **Hybrid retrieval.** SQLite FTS5 BM25 and `sqlite-vec` KNN over 768-dim embeddings, fused with
|
|
632
|
-
reciprocal rank fusion, then ranked by relevance against signal and recency priors and packed into
|
|
633
|
-
an explicit token budget.
|
|
634
|
-
- **MCP server** over stdio (`nexusmem mcp`) exposing `search_memory`, `sync_project` and
|
|
635
|
-
`get_status`, for Claude Desktop, Cursor, Windsurf and other MCP clients.
|
|
636
|
-
- **CLI**: `init`, `sync`, `query`, `status`, `mcp`, and `hook install|remove|status`, plus four
|
|
637
|
-
dry-run previews — `scan-git`, `scan-shell`, `scan-docs`, `scan-conversation` — that write nothing
|
|
638
|
-
and print the nodes ingestion would create with their signal scores.
|
|
639
|
-
- Content-addressed node ids (`sha256(projectId + kind + naturalKey)`), so `sync` is idempotent and
|
|
640
|
-
two clones of one repository share a memory namespace.
|
|
641
|
-
- Everything stays on the machine: one SQLite database in WAL mode under `<repo>/.nexusmem/`.
|
|
642
|
-
|
|
643
|
-
### Notes
|
|
644
|
-
|
|
645
|
-
- Requires Node **>= 22**. `better-sqlite3` 12.11.1 publishes no prebuilt binary for Node 20 — its
|
|
646
|
-
prebuilds start at ABI 127 — so a lower floor would have been a promise the package could not keep.
|
|
647
|
-
Do not lower it without checking upstream prebuilds first.
|
|
648
|
-
- Not done at this release: diff bodies are not indexed, queries are scoped to a single project,
|
|
649
|
-
there is no local-model summarization pass, and the conversation collector has never been audited
|
|
650
|
-
for the stale-node bug that was found and fixed in the docs collector.
|
|
651
|
-
|
|
652
|
-
[Unreleased]: https://github.com/yaminbkk/NexusMem/compare/v0.
|
|
653
|
-
[0.
|
|
654
|
-
[0.10.
|
|
655
|
-
[0.
|
|
656
|
-
[0.
|
|
657
|
-
[0.
|
|
658
|
-
[0.
|
|
659
|
-
[0.
|
|
660
|
-
[0.
|
|
661
|
-
[0.
|
|
662
|
-
[0.
|
|
663
|
-
[0.
|
|
664
|
-
[0.
|
|
665
|
-
[0.4
|
|
666
|
-
[0.
|
|
667
|
-
[0.
|
|
668
|
-
[0.
|
|
669
|
-
[0.
|
|
670
|
-
[0.
|
|
671
|
-
[0.
|
|
672
|
-
[0.
|
|
673
|
-
[0.1
|
|
144
|
+
|
|
145
|
+
## [0.10.3] — 2026-08-31
|
|
146
|
+
|
|
147
|
+
### Added
|
|
148
|
+
|
|
149
|
+
- `hook git-post install|remove|status`: an opt-in git post-commit hook that runs a full
|
|
150
|
+
`nexusmem sync` (including embedding) in the background after every commit. Detached (`nohup ... &`),
|
|
151
|
+
so it never makes `git commit` itself wait; output goes to `.nexusmem/post-commit-sync.log` instead
|
|
152
|
+
of the terminal. `sync` gained a matching `--auto` flag (used only by this hook) that skips instead
|
|
153
|
+
of running if another `--auto` sync already holds an advisory lock over the project's workspace dir —
|
|
154
|
+
a burst of commits (e.g. a rebase) coalesces into one sync instead of piling up. The log file is
|
|
155
|
+
reset once it passes 2000 lines so it can't grow forever, and a skipped/coalesced run is always
|
|
156
|
+
reported there, even though the hook always passes `--quiet`.
|
|
157
|
+
|
|
158
|
+
### Known limitations
|
|
159
|
+
|
|
160
|
+
- A burst of commits fast enough to launch genuinely overlapping `sync --auto` processes can
|
|
161
|
+
interleave two runs' text mid-line in `.nexusmem/post-commit-sync.log` — the advisory lock
|
|
162
|
+
serializes the actual database writes (never at risk), not who gets to write to this diagnostic
|
|
163
|
+
log file. Cosmetic only; not planned to be fixed unless it turns out to matter in practice.
|
|
164
|
+
|
|
165
|
+
## [0.10.2] — 2026-08-30
|
|
166
|
+
|
|
167
|
+
### Security
|
|
168
|
+
|
|
169
|
+
- `shell_command` nodes never ran through the same secret-redaction pass `conversation_turn` and
|
|
170
|
+
`code_diff` nodes already get: a command like `export API_KEY=...` landed verbatim in `title`/`body`,
|
|
171
|
+
which is exactly what the FTS index and vector embeddings are built from — a secret typed at a
|
|
172
|
+
prompt could resurface later through `search_memory`/`nexusmem query`. Fixed by running `redact()`
|
|
173
|
+
over the command before it reaches those fields. `meta.command` is kept raw on purpose: project-id
|
|
174
|
+
reconciliation and failure/fix correlation both hash or exact-match against the real command text,
|
|
175
|
+
and redacting that copy too would have silently orphaned nodes on a project-id migration.
|
|
176
|
+
|
|
177
|
+
## [0.10.1] — 2026-08-30
|
|
178
|
+
|
|
179
|
+
### Added
|
|
180
|
+
|
|
181
|
+
- `nodes.retrieved_count`/`nodes.last_retrieved_at` (schema V12): retrieval outcomes are now recorded
|
|
182
|
+
instead of silently discarded. Bumped once per completed `nexusmem query`/MCP `search_memory` call,
|
|
183
|
+
for every node actually packed into the returned context (both CLI and MCP share one pipeline, so
|
|
184
|
+
both are covered from a single call site). Not yet folded into `rank.ts`'s score formula — every
|
|
185
|
+
existing ranking factor was tuned against real dogfooded queries and validated with `npm run eval`
|
|
186
|
+
before being trusted, and a new factor needs the same treatment. This release is the instrumentation
|
|
187
|
+
half only; confirmed eval-neutral (MRR 0.943 / Recall@5 0.964, unchanged).
|
|
188
|
+
|
|
189
|
+
### Fixed
|
|
190
|
+
|
|
191
|
+
- `nexusmem precheck`'s "What already failed here" warning implied the failure was located in the
|
|
192
|
+
flagged file, but the match is against basename word-tokens (`tokensForFile`) — a failing `npm run
|
|
193
|
+
precheck` flagged every file whose name contained that word, not just the one actually responsible.
|
|
194
|
+
Reworded to state what's actually true ("commands naming this file failed").
|
|
195
|
+
- The high-churn warning in `nexusmem precheck` was rendered as a `WARN`, the same severity as the
|
|
196
|
+
dogfooded failure-correlation warning, despite `HIGH_CHURN_THRESHOLD` being an admitted, untuned
|
|
197
|
+
guess. Demoted to a lower-severity `note`, explicitly labeled as an untuned heuristic, so it no
|
|
198
|
+
longer reads as equally trustworthy.
|
|
199
|
+
|
|
200
|
+
## [0.10.0] — 2026-08-29
|
|
201
|
+
|
|
202
|
+
### Added
|
|
203
|
+
|
|
204
|
+
- New opt-in `github` source: `sources.github.enabled`, `nexusmem scan-github`, `sync --github`.
|
|
205
|
+
Reads issue/PR threads (title, body, comments) from this repo's github.com remote via the `gh`
|
|
206
|
+
CLI, one `github_thread` node per thread. First source with a real external dependency (needs `gh`
|
|
207
|
+
installed and authenticated) rather than reading only what's already on disk; a missing remote or
|
|
208
|
+
an unreachable `gh` is a silent no-op, matching how an unreachable Ollama degrades elsewhere.
|
|
209
|
+
|
|
210
|
+
## [0.9.1] — 2026-08-29
|
|
211
|
+
|
|
212
|
+
Three ranking/retrieval correctness fixes, found and validated against a new 28-case retrieval-quality
|
|
213
|
+
eval harness (`npm run eval`, dev-only, not part of the published package).
|
|
214
|
+
|
|
215
|
+
### Fixed
|
|
216
|
+
|
|
217
|
+
- `nodes_vec` (vector search) over-fetched `k` globally, then filtered by `project_id` after the join —
|
|
218
|
+
a heuristic, not a guarantee. A sparse project sharing `memory.db` with a much larger one could have
|
|
219
|
+
every true nearest neighbour fall outside the over-fetch window and get silently dropped (reproduced:
|
|
220
|
+
495 rows in one project + 5 in another, global `k=50` surfaced 0 of the 5). Schema V11 gives `nodes_vec`
|
|
221
|
+
a `PARTITION KEY` on `project_id` (`sqlite-vec` 0.1.9+), pushing the equality filter into the k-NN
|
|
222
|
+
search itself so cross-project exactness is now guaranteed, not probable. Migrates existing databases
|
|
223
|
+
automatically on next `sync`/`query`. `--as-of` date filtering still over-fetches — only the
|
|
224
|
+
`project_id` dimension was ever a correctness guarantee.
|
|
225
|
+
- `MAX_PRIOR_OVERTURN` (the cap on how far a recency/signal prior may overturn relevance) raised
|
|
226
|
+
`2 → 2.4`, the highest value that both improves eval MRR (0.924→0.943) and still passes every legacy
|
|
227
|
+
regression test guarding against the original same-day-fix-commits bug this constant exists to bound.
|
|
228
|
+
- A commit's `code_diff` siblings (one node per changed file, all sharing the same `ts`) could crowd a
|
|
229
|
+
packed result and bury that commit's own `git_commit` node — its answer — several ranks down. Now
|
|
230
|
+
capped per commit, matching the existing `conversation_turn`/`doc_section` family cap.
|
|
231
|
+
- `package.json`/`tsconfig.json` diffs were ranking above more relevant results when a changed file's
|
|
232
|
+
path happened to echo its commit's conventional-commit scope (title is weighted 10x body, so the scope
|
|
233
|
+
word counted twice). Down-weighted as mechanical wiring, same reasoning `TEST_PATHS` already applies to
|
|
234
|
+
tests. Only affects nodes ingested from here on — existing manifest/config diffs need `sync --rebuild`
|
|
235
|
+
to get the corrected signal retroactively.
|
|
236
|
+
|
|
237
|
+
## [0.9.0] — 2026-08-27
|
|
238
|
+
|
|
239
|
+
Closes the three remaining mechanisms from a recurring external review (Simon Strandgaard, Agent
|
|
240
|
+
Memory Atlas): trust_state, bi-temporal reads, and a dismiss verb for standing suggestions. Also
|
|
241
|
+
fixes two real Windows-specific correctness bugs found dogfooding.
|
|
242
|
+
|
|
243
|
+
### Added
|
|
244
|
+
|
|
245
|
+
- `nodes.trust_state` (`candidate` | `verified` | `rejected`, schema V10) and `nexusmem review
|
|
246
|
+
<nodeId> --verify|--reject`: a human verdict on a node, independent of the SLM contradiction
|
|
247
|
+
checker. `--reject` down-weights ranking (harsher than the existing supersede penalty, since a
|
|
248
|
+
human said no directly); `--verify` is a label only, no ranking boost. Both render as a tag in
|
|
249
|
+
packed context. Never deletes — same demote-not-delete rule as `mark-stale`.
|
|
250
|
+
- `--as-of <date>` on `nexusmem query` and the MCP `search_memory` tool: restricts both the BM25
|
|
251
|
+
and vector arms to nodes recorded at or before that instant via `created_at` (record time),
|
|
252
|
+
independent of how old the events themselves (`ts`, event time) are. Read-only — there is no
|
|
253
|
+
equivalent write, and nodes stay write-once.
|
|
254
|
+
- `nexusmem stale --dismiss`: silences a contradiction suggestion the user disagrees with, without
|
|
255
|
+
fabricating a `supersedes` relationship. Previously a wrong `--check-contradictions` verdict had
|
|
256
|
+
no way to be rejected — it re-decorated every future `stale`/`sync` run forever.
|
|
257
|
+
- `--prune-source`/`--prune-stale-shell` now record a `mutation_audit` row on their `--yes` path,
|
|
258
|
+
matching what `forget` already did. The dry-run preview stays a pure read.
|
|
259
|
+
|
|
260
|
+
### Fixed
|
|
261
|
+
|
|
262
|
+
- `git rev-parse` failures reported as "error launching git: Access is denied." (Git for Windows'
|
|
263
|
+
launcher shim failing to exec `git.exe` under handle/AV contention) were treated as git's own
|
|
264
|
+
verdict instead of a transient spawn failure, turning a one-off environment hiccup into a hard
|
|
265
|
+
failure. Now retried the same as the other known transient-spawn classes.
|
|
266
|
+
- The PowerShell prompt hook read only `$LASTEXITCODE`, which cmdlets (`Remove-Item`, `Copy-Item`,
|
|
267
|
+
…) never set — every failing cmdlet was silently logged as a success, and a stale
|
|
268
|
+
`$LASTEXITCODE` from an earlier native command could misattribute to later successful cmdlet
|
|
269
|
+
runs. Now reads `$?` alongside `$LASTEXITCODE`, before anything else can overwrite `$?`.
|
|
270
|
+
|
|
271
|
+
## [0.8.0] — 2026-08-22
|
|
272
|
+
|
|
273
|
+
### Added
|
|
274
|
+
|
|
275
|
+
- Provenance widened from 2 tiers to a 4-tier trust hierarchy: `observed` (commits, diffs, shell
|
|
276
|
+
exit codes) > `authored` (doc sections — human-written claims) > `recorded` (conversation turns —
|
|
277
|
+
verbatim discourse) > `derived` (session summaries — a model's distillation). Schema V7 backfills
|
|
278
|
+
existing databases by kind. The ranker decays each tier at its own rate (lower trust fades
|
|
279
|
+
faster); the ordering is the design claim, the exact ratios are documented judgment calls.
|
|
280
|
+
- Contradiction checking now runs automatically during `sync`: at most 3 new SLM judgments per run
|
|
281
|
+
(`contradictions.maxPerSync`; `contradictions.autoCheck: false` turns it off), gated on the
|
|
282
|
+
embedding provider already being reachable. Every judgment — either verdict — is memoized in a new
|
|
283
|
+
`contradiction_checks` table (schema V8), so a judged pair is never sent to the model again and
|
|
284
|
+
repeat syncs converge to zero model calls. Measured live against this repo's own database: 5.4s
|
|
285
|
+
for the first 10 judgments, 0.5s for the identical re-run. Still suggest-only: nothing ever writes
|
|
286
|
+
`supersedes` automatically.
|
|
287
|
+
- Standing suggestions now surface everywhere without a model in the loop: plain `nexusmem stale`
|
|
288
|
+
decorates flagged candidates from the memoized judgments (instant, offline), `nexusmem status`
|
|
289
|
+
gains a `flagged` line, and the sync summary reports new/open suggestion counts.
|
|
290
|
+
- `stale --check-contradictions` reuses each candidate's stored embedding instead of re-embedding
|
|
291
|
+
it, and stops after two consecutive null SLM replies so a down provider costs at most two timeouts
|
|
292
|
+
rather than one per candidate.
|
|
293
|
+
|
|
294
|
+
## [0.7.0] — 2026-08-21
|
|
295
|
+
|
|
296
|
+
### Added
|
|
297
|
+
|
|
298
|
+
- `nexusmem stale --check-contradictions`: for each stale candidate, finds the most similar newer
|
|
299
|
+
node (local embedding search) and asks a local SLM (Ollama, `qwen2.5:3b` by default) whether it
|
|
300
|
+
actually contradicts the older one, instead of only surfacing by age. Suggest-only — nothing is
|
|
301
|
+
written, same as plain `stale`. Live-dogfooded against this repo's own real database and Ollama
|
|
302
|
+
instance; found and fixed a real gap along the way (below) before the feature surfaced anything
|
|
303
|
+
useful.
|
|
304
|
+
- Import graph: bare Python imports with no leading dot (`import foo`, `from foo import bar`)
|
|
305
|
+
now resolve to a same-directory sibling file too, alongside the existing relative-dot support.
|
|
306
|
+
Found dogfooding two real local projects: neither used a single PEP 328 relative import, both
|
|
307
|
+
relied entirely on this flat-script style. Guarded by a static list of stdlib module names so a
|
|
308
|
+
bare `import os`/`import queue`/etc. is never mistaken for a same-named local file.
|
|
309
|
+
|
|
310
|
+
### Fixed
|
|
311
|
+
|
|
312
|
+
- Import graph: a Java wildcard import (`import a.b.*;`) no longer merges files from two
|
|
313
|
+
unrelated packages that happen to share a directory-name suffix (e.g. two Gradle/Maven modules
|
|
314
|
+
each with their own `.../foo`) — it now refuses to guess, same as the single-class-import case.
|
|
315
|
+
- `stale --check-contradictions`'s neighbor search now looks past same-timestamp sibling nodes
|
|
316
|
+
(e.g. the many chunks one long conversation gets split into) to reach genuinely newer content.
|
|
317
|
+
Found live-dogfooding against this repo's own database: the first real candidate's closest 15
|
|
318
|
+
neighbors were all same-conversation siblings sharing its exact timestamp, so the original
|
|
319
|
+
5-neighbor default silently found zero suggestions for every candidate, regardless of what the
|
|
320
|
+
SLM would have said.
|
|
321
|
+
|
|
322
|
+
## [0.6.0] — 2026-08-20
|
|
323
|
+
|
|
324
|
+
### Added
|
|
325
|
+
|
|
326
|
+
- Ranker: `inferred` nodes (summaries, doc snapshots) now decay twice as fast as `observed` ones
|
|
327
|
+
(git commits, diffs, shell commands) for retrieval purposes. A judgment call, not a measured
|
|
328
|
+
optimum — see `INFERRED_HALF_LIFE_RATIO` in `src/retrieval/rank.ts`.
|
|
329
|
+
- `nexusmem stale`: lists aging `inferred` nodes nothing has superseded yet, as candidates for
|
|
330
|
+
`mark-stale`. A heuristic on age and provenance, not real contradiction detection — writes nothing.
|
|
331
|
+
- `nexusmem status` now surfaces an `aging` line with the stale-candidate count when any exist.
|
|
332
|
+
- Import graph: Python relative imports (`from .foo import bar`, `from . import x`) now produce
|
|
333
|
+
file edges too, alongside the existing JS/TS support. Absolute Python imports are still skipped —
|
|
334
|
+
same "missed edge over wrong edge" reasoning as JS/TS's bare-specifier skip.
|
|
335
|
+
- Import graph: Go internal imports (resolved against `go.mod`'s module path, one edge per
|
|
336
|
+
non-test file in the imported package) and Rust `mod foo;` declarations (2018+ edition module
|
|
337
|
+
layout) now produce file edges too. External Go imports and Rust `use` paths are out of scope for
|
|
338
|
+
the same reason.
|
|
339
|
+
- Import graph: Java imports (`import a.b.C;` / `import a.b.*;`, resolved by unambiguous suffix
|
|
340
|
+
match against the tracked source tree) and `__DIR__`-anchored PHP `require`/`include` now produce
|
|
341
|
+
file edges too — all six of `nexusmem scan-structure`'s tracked languages. `import static`, PHP's
|
|
342
|
+
autoloaded `use Namespace\Class;`, and any unanchored PHP include are out of scope for the same
|
|
343
|
+
"missed edge over wrong edge" reason as everywhere else in the import graph.
|
|
344
|
+
|
|
345
|
+
## [0.5.4] — 2026-08-20
|
|
346
|
+
|
|
347
|
+
### Added
|
|
348
|
+
|
|
349
|
+
- `nexusmem status --share`: a plain-text, no-color summary (node count, failure→fix chains
|
|
350
|
+
linked, days of history) meant to be pasted somewhere, not scraped by a script.
|
|
351
|
+
|
|
352
|
+
## [0.5.3] — 2026-08-19
|
|
353
|
+
|
|
354
|
+
No functional changes to the CLI, MCP server, or published package -- a test-coverage and
|
|
355
|
+
CI-reliability pass.
|
|
356
|
+
|
|
357
|
+
### Verified
|
|
358
|
+
|
|
359
|
+
- **Closed 15 genuine test-coverage gaps**, found by grepping for real call sites in `tests/`
|
|
360
|
+
rather than guessing from filenames (an Explore agent found the first 7; the rest by hand the
|
|
361
|
+
same way). All 7 `scan-*` CLI subcommands (`scan-git`/`diff`/`shell`/`docs`/`conversation`/
|
|
362
|
+
`structure`/`session`) had zero coverage -- `tests/cli-scan.test.ts`'s name was misleading, it
|
|
363
|
+
only pinned `cli/format.ts`'s shared helpers. Also closed: `runStatus`'s report body beyond the
|
|
364
|
+
stale-identity branch, `runQuery`, `forget --import`'s dry-run preview path, `shell/detect.ts`'s
|
|
365
|
+
zsh scraping and hook-log `repoRoot` scoping, `openAllProjectSources`'s `unreadable` branch,
|
|
366
|
+
`cli/context.ts`'s `loadContext`, `store/fts.ts`'s `toStrictMatchQuery`/`significantTokens`,
|
|
367
|
+
`list_recent_memory`'s MCP `structuredContent` at the protocol level, `structure/resolve.ts`'s
|
|
368
|
+
`.mjs`/`.cjs` rewrite, `docs/read.ts`'s `include` option, and `git/log.ts`'s `buildLogArgs`. Test
|
|
369
|
+
suite: 559 → 592 tests.
|
|
370
|
+
|
|
371
|
+
### Fixed
|
|
372
|
+
|
|
373
|
+
- **A real test-isolation bug: a test could scrape a developer's actual shell history into a
|
|
374
|
+
throwaway test database.** `tests/setup.ts` isolated `NEXUSMEM_HOME` (so a test `sync()` can't
|
|
375
|
+
pollute the real `~/.nexusmem/projects.json` registry -- fixed 2026-08-15 for the same reason)
|
|
376
|
+
but never isolated `APPDATA`/`HISTFILE_BASH`/`HISTFILE`, so any test running a real `sync()`
|
|
377
|
+
scraped whatever PSReadLine/bash/zsh history actually exists on the machine running the suite.
|
|
378
|
+
Observed directly: a new test here ingested 300 real `shell_command` nodes off this machine's
|
|
379
|
+
own history before the fix. Now isolated the same way as the registry.
|
|
380
|
+
- **CI failed (deterministically, all 4 jobs) on 4 assertions that assumed adjacent words in
|
|
381
|
+
colorized CLI output stay contiguous.** GitHub Actions' runners don't have `NO_COLOR` set the
|
|
382
|
+
way this project's local dev shells apparently do, so `picocolors` emits real ANSI codes there --
|
|
383
|
+
a regex like `/git\s+last run/` breaks when an escape sequence sits between a padded label and a
|
|
384
|
+
separately-colored value. `tests/forget.test.ts` already carries a `stripAnsi()` helper for
|
|
385
|
+
exactly this trap; applied the same fix to the 4 newly-added test files that had it. Verified by
|
|
386
|
+
reproducing the CI failure locally with `FORCE_COLOR=1` before the fix, then re-running the full
|
|
387
|
+
592-test suite under `FORCE_COLOR=1` after, to check for any other latent instance.
|
|
388
|
+
|
|
389
|
+
## [0.5.2] — 2026-08-18
|
|
390
|
+
|
|
391
|
+
### Added
|
|
392
|
+
|
|
393
|
+
- **`nexusmem mark-stale <nodeId> --supersedes <newNodeId>`, `provenance`, and manual staleness
|
|
394
|
+
tracking.** A pre-Show-HN pass answering a public critique that this couldn't tell an observed
|
|
395
|
+
fact from a guess, or retire a stale conclusion. Every node now carries `provenance`
|
|
396
|
+
(`observed`/`inferred`, set per-collector: git commits, diffs, and shell commands are `observed`;
|
|
397
|
+
conversation turns, session summaries, and doc sections are `inferred`) and an optional
|
|
398
|
+
`supersedes` link. `mark-stale` writes that link; the ranker applies a flat down-weight to whatever
|
|
399
|
+
it points at, but never deletes it -- the old node stays queryable, just usually loses to its
|
|
400
|
+
replacement. `provenance` is shown as a `[observed]`/`[inferred]` tag on every query result (CLI
|
|
401
|
+
and MCP `search_memory` share the same renderer). This is deliberately not automatic: nothing
|
|
402
|
+
detects staleness on its own, a human or agent still has to notice the contradiction and run the
|
|
403
|
+
command. See the README's new "Manual staleness & provenance" section.
|
|
404
|
+
|
|
405
|
+
### Verified
|
|
406
|
+
|
|
407
|
+
- **`forget` does not leak through vector search.** Checked whether `MemoryStore.forget` excludes a
|
|
408
|
+
tombstoned node from `sqlite-vec` results only via a query-time filter (which every vector query
|
|
409
|
+
path would need to apply consistently) or by deleting the embedding row outright. It already does
|
|
410
|
+
the latter -- `dropEmbedding` runs before the node row is deleted, in the same transaction, on both
|
|
411
|
+
query paths (`runHybridQuery` and `runCrossProjectQuery`, which both call the same
|
|
412
|
+
`MemoryStore.vectorSearch`). No fix was needed; added a regression test
|
|
413
|
+
(`tests/vector.test.ts`) that forgets an embedded node and asserts both `vectorSearch` returns
|
|
414
|
+
nothing and the `nodes_vec` row count drops to zero, so this can't silently regress.
|
|
415
|
+
- **`forget` survives `sync --rebuild`.** This was the point of v0.5.0 and was already covered
|
|
416
|
+
end-to-end in `tests/forget.test.ts`. Added a lower-level `tests/store.test.ts` case pinning the
|
|
417
|
+
exact mechanism: `clearProject` (what `--rebuild` calls before re-ingesting) only touches
|
|
418
|
+
`nodes`/`nodes_vec`/`sync_state`, never `deny_list`, so a re-ingest of the same content is denied
|
|
419
|
+
again rather than resurrected.
|
|
420
|
+
|
|
421
|
+
## [0.5.1] — 2026-08-17
|
|
422
|
+
|
|
423
|
+
### Added
|
|
424
|
+
|
|
425
|
+
- **`nexusmem forget --export <path>` / `forget --import <path>`: carry a deny-list across a clone or
|
|
426
|
+
restore.** `.nexusmem/` is gitignored by design, so `deny_list` never traveled with `git clone`/
|
|
427
|
+
`git push` — confirmed live 2026-08-17 that a fresh clone of the exact same repo resurrected a value
|
|
428
|
+
already forgotten elsewhere, with zero deny-list protection, because git history (what a fresh `sync`
|
|
429
|
+
re-derives from) is fully portable while the deny-list that would have blocked it was not. `--export`
|
|
430
|
+
writes the active entries to a plaintext JSON file (loudly warned as exactly as sensitive as the
|
|
431
|
+
values it holds — never meant for git, moved through whatever secure channel the user already
|
|
432
|
+
trusts); `--import` re-applies each new entry through `forget` itself, so an imported value is
|
|
433
|
+
deleted from the new checkout's nodes too, not just blocked going forward. Same dry-run-by-default /
|
|
434
|
+
`--yes` convention as the rest of `forget`. See `docs/forget-mechanism.md`.
|
|
435
|
+
|
|
436
|
+
## [0.5.0] — 2026-08-17
|
|
437
|
+
|
|
438
|
+
### Added
|
|
439
|
+
|
|
440
|
+
- **`nexusmem forget <value>`: permanent, value-keyed deletion.** The finer-grained complement to
|
|
441
|
+
`sync --prune-source` (which only deletes at the whole-collector-source granularity): forgets one
|
|
442
|
+
exact string or `--regex` pattern, deleting every node it currently matches *and* writing a
|
|
443
|
+
standing deny-list entry so the value can never be re-ingested — closing the one gap an external
|
|
444
|
+
source-level review flagged as its most serious finding: the append-only shell-hook log (and a
|
|
445
|
+
full transcript re-read) are untouched by any prune, so `sync --rebuild` used to resurrect exactly
|
|
446
|
+
what had just been deleted. Every removal leaves a hash-only tombstone (never the forgotten
|
|
447
|
+
content itself — `body`/`title` are stored as sha256 only) and the whole operation writes one
|
|
448
|
+
`mutation_audit` row, whether or not anything matched. Dry-run by default, matching
|
|
449
|
+
`--prune-source`'s convention exactly; `--yes` confirms; `--list` shows active entries. Permanent
|
|
450
|
+
in v0 — no `--remove`, matching the existing irreversible framing of `--prune-source`/`--rebuild`.
|
|
451
|
+
See `docs/forget-mechanism.md`.
|
|
452
|
+
|
|
453
|
+
## [0.4.0] — 2026-08-16
|
|
454
|
+
|
|
455
|
+
### Added
|
|
456
|
+
|
|
457
|
+
- **`nexusmem status` now warns when prior project identities still hold nodes.** A renamed git
|
|
458
|
+
remote can leave stale data under the old identity; status reports both the identity and node
|
|
459
|
+
counts and points to `sync --prune-source <name>` instead of leaving that data discoverable only
|
|
460
|
+
through a raw SQLite query.
|
|
461
|
+
- **`docs/competitor-comparison.md`: an honest, source-level comparison against `riponcm/projectmem`** (705
|
|
462
|
+
GitHub stars vs. this project's 8, at time of writing) — a real, working competitor read in full (not
|
|
463
|
+
judged from its README), covering where the two tools' failure-tracking, import-graph coverage, and
|
|
464
|
+
token-savings claims genuinely differ. Cites a freshly re-run benchmark (`npm run bench`, not a stale
|
|
465
|
+
table) alongside projectmem's own `pjm score` constants, read directly from its source. Linked from the
|
|
466
|
+
README's `## What it costs you` section and a short website section near the existing stats.
|
|
467
|
+
- **`nexusmem hook git install|remove|status`: a real git pre-commit hook.** Installs a marked block into
|
|
468
|
+
`.git/hooks/pre-commit` that runs `nexusmem precheck` (no `--strict`, so it can never block a commit on
|
|
469
|
+
its own) before each commit — the automatic counterpart to the advisory `precheck` command. Refuses to
|
|
470
|
+
touch a pre-existing foreign hook (husky, lint-staged, lefthook, ...) unless `--force` is passed, in which
|
|
471
|
+
case it appends after the existing content rather than before, so the foreign hook still runs first and
|
|
472
|
+
keeps deciding whatever it already decided. Idempotent (`nexusmem hook git install` twice is a no-op) and
|
|
473
|
+
removable cleanly, including restoring a foreign hook to its original content if one was appended onto.
|
|
474
|
+
Live-verified against a real scratch git repo on Windows: fresh install, a real `git commit` that
|
|
475
|
+
triggered the hook and printed a correct precheck report, clean removal, and both the refuse-without-force
|
|
476
|
+
and append-with-force foreign-hook paths.
|
|
477
|
+
- **`nexusmem precheck`: proactive pre-commit warnings.** Checks staged (or `--working`, or explicit `--files`)
|
|
478
|
+
files against project memory and warns about unresolved past failures and high recent churn *before* you
|
|
479
|
+
commit — advisory by default (always exits 0; `--strict` turns an unresolved failure into a non-zero exit).
|
|
480
|
+
Matches a file's own basename tokens against still-unlinked `shell_command` failures (no
|
|
481
|
+
`resolved_by:*` link from `sync --link-failures`), reusing `filterBoilerplateTokens` from the
|
|
482
|
+
discussion-bridge heuristic — now exported and parameterized by node kind so it can be corpus-relative
|
|
483
|
+
against `shell_command` history instead of only `conversation_turn`/`session_summary`. Churn is scoped to
|
|
484
|
+
`git_commit`-kind `node_files` touches only, so it doesn't double-count the identical touches `code_diff`
|
|
485
|
+
nodes also record. Deliberately does not yet install as a real git hook (see the module comment in
|
|
486
|
+
`src/correlate/precheck.ts` for why capture-time `git status` diffing and past-commit correlation were both
|
|
487
|
+
rejected) — that's a follow-up once this signal has been dogfooded.
|
|
488
|
+
- **JS/TS import-graph edges.** `nexusmem scan-structure` previews (and `sync` now ingests) file→file
|
|
489
|
+
import relationships across a project's tracked `.ts`/`.tsx`/`.js`/`.jsx` files — a dependency-free
|
|
490
|
+
regex extractor (`src/structure/extract.ts`) resolves relative `import`/`export ... from`/`require`/
|
|
491
|
+
dynamic-`import` specifiers against the tracked-path set, correctly rewriting the common TS/ESM
|
|
492
|
+
`./foo.js` specifier back to its real `foo.ts` source. Stored in a new `file_edges` table (schema
|
|
493
|
+
v4), replaced wholesale on every sync since edges describe current tree state, not history. Surfaced
|
|
494
|
+
as a `structure` line in `nexusmem status`; not yet wired into `query` ranking or exposed as an MCP
|
|
495
|
+
tool — that's a follow-on design question, not this pass's job.
|
|
496
|
+
|
|
497
|
+
## [0.3.3] — 2026-08-16
|
|
498
|
+
|
|
499
|
+
### Added
|
|
500
|
+
|
|
501
|
+
- **`nexusmem status` now shows failure→fix chain counts** — `N/M failure(s) resolved (X retry, Y
|
|
502
|
+
discussion)`, plus a hint to run `sync --link-failures` when failures remain unresolved. The
|
|
503
|
+
chain feature is this project's most distinctive capability, but was previously invisible to
|
|
504
|
+
anyone who didn't already know to query `node_links` directly. Backed by the new
|
|
505
|
+
`getChainStats` in `src/correlate/failure-fix.ts`, which dedupes failures resolved by both
|
|
506
|
+
heuristics rather than double-counting them.
|
|
507
|
+
|
|
508
|
+
### Fixed
|
|
509
|
+
|
|
510
|
+
- **Discussion-bridge heuristic (`sync --link-failures`): corpus-relative boilerplate tokens no
|
|
511
|
+
longer produce false-positive failure→fix links.** Re-dogfooded at larger scale against a second
|
|
512
|
+
real project's history: a command made entirely of words that saturate a project's own corpus
|
|
513
|
+
(e.g. this repo's own name/verbs) could AND-match an unrelated turn that just happened to mention
|
|
514
|
+
the same words, and bm25 score alone could not separate that from a true positive (measured: the
|
|
515
|
+
false positive scored *stronger* than two real true positives). `filterBoilerplateTokens` in
|
|
516
|
+
`src/correlate/failure-fix.ts` drops any token that appears in over 20% of a project's own
|
|
517
|
+
`conversation_turn`/`session_summary` history before building the match query — measured against
|
|
518
|
+
real data, not guessed — and skips the discussion-match attempt entirely (rather than falling back
|
|
519
|
+
unfiltered) when every token turns out to be boilerplate, since a missed link is preferred over a
|
|
520
|
+
false one for this heuristic. Below 10 discussable nodes the filter is skipped, since frequency
|
|
521
|
+
isn't a meaningful signal yet on a young project.
|
|
522
|
+
|
|
523
|
+
## [0.3.2] — 2026-08-16
|
|
524
|
+
|
|
525
|
+
### Added
|
|
526
|
+
|
|
527
|
+
- **`mcpName` field in `package.json`**, required by the official MCP registry
|
|
528
|
+
(registry.modelcontextprotocol.io) to verify that whoever publishes `server.json` under
|
|
529
|
+
`io.github.yaminbkk/nexusmem` also controls the `nexusmem` npm package itself — the registry
|
|
530
|
+
rejects a publish attempt otherwise. No behavior change for CLI/MCP users; this is purely a
|
|
531
|
+
registry-ownership proof.
|
|
532
|
+
- **`list_recent_memory` MCP tool** — chronological listing of a repository's most recently
|
|
533
|
+
remembered nodes (git commits, diffs, shell commands, docs, conversation, session summaries),
|
|
534
|
+
newest first. Distinct from `search_memory`: no query, just "what has this project's memory
|
|
535
|
+
recorded lately" — built for the VS Code extension's sidebar view, which lists rather than
|
|
536
|
+
searches. Backed by `MemoryStore.listRecentNodes`, reusing the existing `idx_nodes_project_ts`
|
|
537
|
+
index.
|
|
538
|
+
- **`sync --prune-source <name>` and `sync --prune-stale-shell`** — drop one source's nodes without a
|
|
539
|
+
full `--rebuild`, which loses history that can't be re-read from disk (the shell tail window, older
|
|
540
|
+
conversation turns). `--prune-stale-shell` is a shortcut for the three dead pre-hook shell-scrape
|
|
541
|
+
sources (`shell:pwsh`, `shell:bash`, `shell:zsh`) at once. Dry-run by default — prints the matching
|
|
542
|
+
count and does nothing until `--yes` is also given, since this is an irreversible full wipe of the
|
|
543
|
+
named source(s), unlike `--rebuild`'s no-prompt full-project reset. Also sweeps any prior project
|
|
544
|
+
identity of this same repo (the id a renamed git remote leaves behind after
|
|
545
|
+
`fix(store): reconcile memory stranded by a changed git remote URL` migrates what it can) — a
|
|
546
|
+
live-id-only prune could not reach nodes reconciliation deliberately left in place. Exposed on both
|
|
547
|
+
the CLI and the MCP `sync_project` tool.
|
|
548
|
+
- **Discussion-heuristic failure→fix chains now surface**, tightened to an AND-joined significant-
|
|
549
|
+
token match instead of the original OR match. Re-verified against this repo's own real database:
|
|
550
|
+
5/5 discussion links correct (was ~half wrong when it shipped unsurfaced in 0.3.0). Chains now
|
|
551
|
+
follow across projects in `query --all-projects` too.
|
|
552
|
+
|
|
553
|
+
### Fixed
|
|
554
|
+
|
|
555
|
+
- **`sync_project`'s summary no longer contains raw ANSI color codes on Windows.** `runInit`/`runSync`
|
|
556
|
+
format their output with picocolors for terminal display, and picocolors treats `platform ===
|
|
557
|
+
'win32'` as sufficient evidence of color support on its own, without checking `isTTY` — correct for
|
|
558
|
+
a real terminal, wrong for the MCP JSON-RPC channel, which is piped on every platform. Found live: a
|
|
559
|
+
real MCP client (the VS Code extension's Output channel) rendered the raw escape codes as literal
|
|
560
|
+
text instead of color. Stripped at the MCP boundary in `syncProject`, leaving the CLI's own terminal
|
|
561
|
+
output untouched.
|
|
562
|
+
|
|
563
|
+
## [0.3.1] — 2026-08-15
|
|
564
|
+
|
|
565
|
+
### Added
|
|
566
|
+
|
|
567
|
+
- **`scripts/benchmark.ts` (`npm run bench`) — a reproducible end-to-end token-saving benchmark.**
|
|
568
|
+
Compares `packed.tokensUsed` against two baselines (full-file-read and `git log -p`) for the same
|
|
569
|
+
files a query's packed nodes touch, over a query set derived mechanically from the corpus itself
|
|
570
|
+
rather than hand-picked. Used to measure the README's `## What it costs you` numbers against both
|
|
571
|
+
this repo (62 commits) and a 9,567-commit external corpus (`vitejs/vite`) — see README for the
|
|
572
|
+
numbers and their methodology caveats.
|
|
573
|
+
|
|
574
|
+
### Fixed
|
|
575
|
+
|
|
576
|
+
- **A generic 2-3 letter local model title, "id" as a search token, and one node's chunks flooding a
|
|
577
|
+
result set — three ranking/retrieval edge cases found by dogfooding, each verified with a red test
|
|
578
|
+
before the fix.**
|
|
579
|
+
- Session-summary titles: the local model sometimes wrote a role-framing line ("Role: Lead Systems
|
|
580
|
+
Engineer...") instead of a summary, in Thai and English alike. The existing generic-title filter
|
|
581
|
+
didn't catch either language, since it only matched English words like "summary"/"update".
|
|
582
|
+
- Search: the token `id` alone was prefix-matching unrelated shell commands like
|
|
583
|
+
`winget install --id ...`, because every query token was OR-ed with no floor and no stopword
|
|
584
|
+
list, and bm25 gives a rare-but-generic token an inflated score purely from scarcity.
|
|
585
|
+
- Packing: `conversation_turn` and `doc_section` both chunk one reply or file into several nodes
|
|
586
|
+
sharing the same timestamp; up to 2 may now appear in one packed result, down from unlimited.
|
|
587
|
+
|
|
588
|
+
- **A repo's memory could silently split in two if its git remote URL ever changed** (a GitHub
|
|
589
|
+
account rename, an org transfer). Project identity is derived from the remote URL on purpose — so
|
|
590
|
+
the same repo re-cloned to a new path or machine keeps sharing memory — but a changed URL on the
|
|
591
|
+
*same* path minted a new id and stranded every node synced under the old one, invisible to
|
|
592
|
+
`status`/`query`/MCP from then on. `sync` now detects a prior id already recorded in the repo's own
|
|
593
|
+
database and reconciles it forward: recomputable node kinds (session summaries, hook-sourced shell
|
|
594
|
+
history) are migrated under their correct new id and deduplicated against anything already synced;
|
|
595
|
+
conversation turns, whose identity can't be recomputed, are reassigned in place. Git commits,
|
|
596
|
+
diffs, and doc sections are left alone — a normal sync already re-derives them completely, so
|
|
597
|
+
there is nothing to migrate.
|
|
598
|
+
|
|
599
|
+
## [0.3.0] — 2026-08-13
|
|
600
|
+
|
|
601
|
+
**Upgrade note.** The first `sync` after upgrading drops every stored embedding and rebuilds it.
|
|
602
|
+
This is not optional and it is not a bug: embeddings now come from Ollama's `/api/embed`, which
|
|
603
|
+
returns L2-normalised vectors, while the previous `/api/embeddings` did not — measured norms of 1.0
|
|
604
|
+
and 20.7 for the same input. `nodes_vec` ranks by Euclidean distance and records no per-row
|
|
605
|
+
provenance, so a corpus holding both would separate by scale rather than by meaning. `sync` says
|
|
606
|
+
what it dropped, nodes are untouched, and BM25 keeps working while the rebuild runs. On this repo
|
|
607
|
+
the rebuild was 833 nodes in one pass, inside a 9-second sync.
|
|
608
|
+
|
|
609
|
+
### Added
|
|
610
|
+
|
|
611
|
+
- **Session summaries via a local model** (`sources.session`, opt-in, off by default). Each
|
|
612
|
+
finished session becomes one distilled `session_summary` node — decisions and their reasons —
|
|
613
|
+
alongside the raw exchanges. Runs a local Ollama chat model (`qwen2.5:3b` by default); nothing is
|
|
614
|
+
downloaded automatically and no transcript leaves the machine. New `scan-session` command, with
|
|
615
|
+
`--dry-run` to print the exact prompt a session would produce without calling the model.
|
|
616
|
+
Bounded three ways: a session must be quiet for `settleMinutes` (default 30) before it is
|
|
617
|
+
eligible, the prompt is hashed so an unchanged session never reaches the model again, and
|
|
618
|
+
`maxSessions` caps how many are summarized per sync. Every exchange is redacted before the model
|
|
619
|
+
sees it, and the model's own output is redacted again before it is stored.
|
|
620
|
+
- **`sync --embed-limit <n>`** to cap the embedding pass, for when draining the whole backlog is not
|
|
621
|
+
wanted.
|
|
622
|
+
|
|
623
|
+
### Changed
|
|
624
|
+
|
|
625
|
+
- **The embedding pass drains the backlog in one `sync`** instead of stopping after 200 nodes, and
|
|
626
|
+
sends texts to Ollama in batches of 32 — measured at 4.21x the throughput of one call per node
|
|
627
|
+
(20.3ms → 4.8ms per node, over 96 real nodes from this repo). Leaving it uncapped is safe because
|
|
628
|
+
paging walks rowids monotonically, so a node the provider failed on is passed over rather than
|
|
629
|
+
retried forever, and because three consecutive dead requests end the pass: an Ollama that is not
|
|
630
|
+
running now costs three requests rather than one timeout per node.
|
|
631
|
+
- **Embeddings carry a provider identity.** Changing the embedding model, or upgrading from a
|
|
632
|
+
release that recorded no identity, drops the vectors and re-embeds rather than ranking across a
|
|
633
|
+
mixture.
|
|
634
|
+
- **Ranking priors now share one budget instead of getting one each.** `signal` and `recency` are
|
|
635
|
+
query-independent, and each was separately capped at overturning a 2× relevance gap. The score
|
|
636
|
+
multiplies them, so together they could overturn 4× — which is not a corner case but a description
|
|
637
|
+
of every commit made during an active working day, fresh and high-signal at once. A query about the
|
|
638
|
+
PowerShell hook returned two unrelated same-day `fix:` commits at ranks 3 and 4 while the section
|
|
639
|
+
that answered it sat at rank 6. The 2× budget is now the bound on the priors *jointly*, split
|
|
640
|
+
between them (`signal^0.215 × recency^0.288`, down from `^0.431` and `^0.576`), and a third prior
|
|
641
|
+
would re-divide the same budget rather than enlarge it. Measured on four real queries against this
|
|
642
|
+
repository's memory: the answering section rose in three of them — the rationale for "why BM25
|
|
643
|
+
before vector search" went from rank 4 to rank 1 — and no query's correct top hit was displaced.
|
|
644
|
+
|
|
645
|
+
### Known limitation
|
|
646
|
+
|
|
647
|
+
- Session-summary *titles* depend on the model following a fixed output format, and a 3B model often
|
|
648
|
+
does not. Measured over 14 real sessions, roughly a third came back usable; the rest were
|
|
649
|
+
conversational preambles, stray bullets, or a bare "Summary of the Session". Those are rejected
|
|
650
|
+
and the title falls back to the first line of the question that opened the session — always
|
|
651
|
+
specific, not always elegant. Compliance was worst on long sessions and on transcripts not in
|
|
652
|
+
English. `sources.session.model` takes a larger model if it matters.
|
|
653
|
+
|
|
654
|
+
## [0.2.0] — 2026-08-12
|
|
655
|
+
|
|
656
|
+
**Upgrade note.** Both new sources are on by default, so the first `sync` after upgrading an
|
|
657
|
+
existing project ingests the patches of its 200 most recent commits and starts recording the
|
|
658
|
+
repository in `~/.nexusmem/projects.json`. Set `sources.diff.enabled` to `false` in
|
|
659
|
+
`.nexusmem/config.json` if you would rather not, and `NEXUSMEM_HOME` relocates the user-scoped
|
|
660
|
+
directory. Nothing existing is rewritten or lost.
|
|
661
|
+
|
|
662
|
+
### Added
|
|
663
|
+
|
|
664
|
+
- **Diff-level nodes.** Commit patches are now indexed, one node per changed file, so a question
|
|
665
|
+
about *what the change looked like* reaches the lines themselves rather than the commit message
|
|
666
|
+
and a `+41/-6` summary. New `code_diff` kind, `diff` source, `scan-diff` preview command, and a
|
|
667
|
+
`sources.diff` config block. Read by a second `git log --patch` walk with its own cursor: folding
|
|
668
|
+
it into the existing `--numstat` walk would put patch text and numstat rows in one field, where a
|
|
669
|
+
diff line reading `-1\t2\tfoo` is indistinguishable from a real file entry.
|
|
670
|
+
Bounded on purpose — 200 commits on a first sync, 20 files per commit, no merges (their patch
|
|
671
|
+
exists only in a combined format this parser does not read), and binaries, lockfiles and build
|
|
672
|
+
output skipped. Patches are redacted with the shape-matching rules only; the key/value rule that
|
|
673
|
+
serves prose would rewrite `const apiKey = process.env.SERVICE_API_KEY` into a redaction marker.
|
|
674
|
+
- **Cross-project recall.** `query --all-projects` (and `search_memory`'s `allProjects`) searches
|
|
675
|
+
every repository NexusMem has been run in on this machine, tagging each result with the repository
|
|
676
|
+
it came from. Databases stay per-repository — a shared global store was rejected for giving up the
|
|
677
|
+
property that deleting one repo's `.nexusmem/` removes that repo's memory and nothing else — so a
|
|
678
|
+
plain index at `~/.nexusmem/projects.json`, written by `init` and refreshed by `sync`, is what
|
|
679
|
+
makes the others findable. New `projects` command lists it; `--prune` forgets entries whose
|
|
680
|
+
database is gone. A stale or corrupt registry degrades the query, never fails it.
|
|
681
|
+
Ranking fuses each project's list by rank (RRF) instead of comparing raw BM25 costs, which are
|
|
682
|
+
computed against their own corpus and are not comparable across databases. The bias this leaves —
|
|
683
|
+
every project's rank-1 hit is worth the same, so recall favours breadth — is documented rather
|
|
684
|
+
than hidden.
|
|
685
|
+
- **Query-aware diff excerpts.** A packed summary is ~320 characters and a patch is thousands, so
|
|
686
|
+
the packer now picks the hunk whose tokens match the query and starts the excerpt at the changed
|
|
687
|
+
line. Found by dogfooding: "what flags are passed to every git invocation" retrieved the right
|
|
688
|
+
file and then spent the whole summary on a class definition seventy lines above the answer.
|
|
689
|
+
Matching splits identifiers on case and underscore boundaries, because `\bretry\b` does not match
|
|
690
|
+
`RETRY_DELAYS_MS` and a natural-language question otherwise never meets the code it is about.
|
|
691
|
+
- `CHANGELOG.md` now ships inside the npm tarball. npm's always-included list covers `package.json`,
|
|
692
|
+
`README` and `LICENSE` but not the changelog, so it previously reached GitHub readers only.
|
|
693
|
+
|
|
694
|
+
### Internal
|
|
695
|
+
|
|
696
|
+
- The test suite no longer writes to the developer's real `~/.nexusmem`. `sync` records the
|
|
697
|
+
repository it ingested in the project registry, and the suite syncs temporary repositories in
|
|
698
|
+
several places, so a green run left seven dead entries behind — found by running `nexusmem
|
|
699
|
+
projects` after the fact, not by any test. `tests/setup.ts` now points `NEXUSMEM_HOME` at a
|
|
700
|
+
throwaway directory for the whole suite, and one test fails if that guard is ever removed.
|
|
701
|
+
- `npm run smoke` drives the *packaged* artifact: build, pack, install into a throwaway directory,
|
|
702
|
+
then run the installed CLI, an end-to-end ingest/query against a fixture repository, and an
|
|
703
|
+
`initialize` handshake over real stdio. It also audits the manifest `npm publish` would send,
|
|
704
|
+
which is a different artifact from the tarball. Both defects that ever reached npm users passed a
|
|
705
|
+
green unit suite first; each is now pinned by a check verified to fail when the defect is
|
|
706
|
+
reintroduced. Runs in CI on Linux and Windows as its own job.
|
|
707
|
+
|
|
708
|
+
## [0.1.2] — 2026-08-10
|
|
709
|
+
|
|
710
|
+
### Fixed
|
|
711
|
+
|
|
712
|
+
- `nexusmem --version` printed `0.1.0` on 0.1.1. The version string in `src/cli/index.ts` was a
|
|
713
|
+
literal separate from `package.json`, and the 0.1.1 bump only touched the latter. `src/mcp/server.ts`
|
|
714
|
+
had the same problem in its `McpServer` constructor, so an MCP client's `initialize` handshake
|
|
715
|
+
would have reported the same stale version. Both now read the real version through
|
|
716
|
+
`readOwnVersion()` in `src/core/version.ts`, which resolves `package.json` via `import.meta.url`.
|
|
717
|
+
Found by running the published package end to end rather than trusting `npm publish --dry-run`
|
|
718
|
+
and the registry API, neither of which executes a `--version` flag.
|
|
719
|
+
|
|
720
|
+
The ingestion and retrieval pipeline was never affected — only the two places that report a version
|
|
721
|
+
independently of running a command.
|
|
722
|
+
|
|
723
|
+
## [0.1.1] — 2026-08-10
|
|
724
|
+
|
|
725
|
+
### Changed
|
|
726
|
+
|
|
727
|
+
- README rewritten for someone deciding whether to read the source: what it does, how retrieval
|
|
728
|
+
scores, what it costs, and where it breaks. `README.md` ships inside the package, so this is a
|
|
729
|
+
real change to what npm delivers — but no code changed between 0.1.0 and 0.1.1.
|
|
730
|
+
- Documented the ranking flaw the tool found in itself, and the fact that the conversation source
|
|
731
|
+
in the sample `status` output is opt-in rather than default.
|
|
732
|
+
- Dropped `&&` from the quickstart, which Windows PowerShell 5.1 cannot parse.
|
|
733
|
+
|
|
734
|
+
## [0.1.0] — 2026-08-10
|
|
735
|
+
|
|
736
|
+
First public release.
|
|
737
|
+
|
|
738
|
+
### Added
|
|
739
|
+
|
|
740
|
+
- **Collectors.** Git history (commit metadata and diff stats, not diff bodies), shell commands with
|
|
741
|
+
exit codes via an opt-in PowerShell hook, tracked markdown docs via `git ls-files -- '*.md'`, and
|
|
742
|
+
opt-in assistant transcripts.
|
|
743
|
+
- **Hybrid retrieval.** SQLite FTS5 BM25 and `sqlite-vec` KNN over 768-dim embeddings, fused with
|
|
744
|
+
reciprocal rank fusion, then ranked by relevance against signal and recency priors and packed into
|
|
745
|
+
an explicit token budget.
|
|
746
|
+
- **MCP server** over stdio (`nexusmem mcp`) exposing `search_memory`, `sync_project` and
|
|
747
|
+
`get_status`, for Claude Desktop, Cursor, Windsurf and other MCP clients.
|
|
748
|
+
- **CLI**: `init`, `sync`, `query`, `status`, `mcp`, and `hook install|remove|status`, plus four
|
|
749
|
+
dry-run previews — `scan-git`, `scan-shell`, `scan-docs`, `scan-conversation` — that write nothing
|
|
750
|
+
and print the nodes ingestion would create with their signal scores.
|
|
751
|
+
- Content-addressed node ids (`sha256(projectId + kind + naturalKey)`), so `sync` is idempotent and
|
|
752
|
+
two clones of one repository share a memory namespace.
|
|
753
|
+
- Everything stays on the machine: one SQLite database in WAL mode under `<repo>/.nexusmem/`.
|
|
754
|
+
|
|
755
|
+
### Notes
|
|
756
|
+
|
|
757
|
+
- Requires Node **>= 22**. `better-sqlite3` 12.11.1 publishes no prebuilt binary for Node 20 — its
|
|
758
|
+
prebuilds start at ABI 127 — so a lower floor would have been a promise the package could not keep.
|
|
759
|
+
Do not lower it without checking upstream prebuilds first.
|
|
760
|
+
- Not done at this release: diff bodies are not indexed, queries are scoped to a single project,
|
|
761
|
+
there is no local-model summarization pass, and the conversation collector has never been audited
|
|
762
|
+
for the stale-node bug that was found and fixed in the docs collector.
|
|
763
|
+
|
|
764
|
+
[Unreleased]: https://github.com/yaminbkk/NexusMem/compare/v0.11.0...HEAD
|
|
765
|
+
[0.11.0]: https://github.com/yaminbkk/NexusMem/compare/v0.10.5...v0.11.0
|
|
766
|
+
[0.10.5]: https://github.com/yaminbkk/NexusMem/compare/v0.10.4...v0.10.5
|
|
767
|
+
[0.10.4]: https://github.com/yaminbkk/NexusMem/compare/v0.10.3...v0.10.4
|
|
768
|
+
[0.10.3]: https://github.com/yaminbkk/NexusMem/compare/v0.10.2...v0.10.3
|
|
769
|
+
[0.10.2]: https://github.com/yaminbkk/NexusMem/compare/v0.10.1...v0.10.2
|
|
770
|
+
[0.10.1]: https://github.com/yaminbkk/NexusMem/compare/v0.10.0...v0.10.1
|
|
771
|
+
[0.10.0]: https://github.com/yaminbkk/NexusMem/compare/v0.9.1...v0.10.0
|
|
772
|
+
[0.9.1]: https://github.com/yaminbkk/NexusMem/compare/v0.9.0...v0.9.1
|
|
773
|
+
[0.9.0]: https://github.com/yaminbkk/NexusMem/compare/v0.8.0...v0.9.0
|
|
774
|
+
[0.8.0]: https://github.com/yaminbkk/NexusMem/compare/v0.7.0...v0.8.0
|
|
775
|
+
[0.7.0]: https://github.com/yaminbkk/NexusMem/compare/v0.6.0...v0.7.0
|
|
776
|
+
[0.6.0]: https://github.com/yaminbkk/NexusMem/compare/v0.5.4...v0.6.0
|
|
777
|
+
[0.5.4]: https://github.com/yaminbkk/NexusMem/compare/v0.5.3...v0.5.4
|
|
778
|
+
[0.5.3]: https://github.com/yaminbkk/NexusMem/compare/v0.5.2...v0.5.3
|
|
779
|
+
[0.5.2]: https://github.com/yaminbkk/NexusMem/compare/v0.5.1...v0.5.2
|
|
780
|
+
[0.5.1]: https://github.com/yaminbkk/NexusMem/compare/v0.5.0...v0.5.1
|
|
781
|
+
[0.5.0]: https://github.com/yaminbkk/NexusMem/compare/v0.4.0...v0.5.0
|
|
782
|
+
[0.4.0]: https://github.com/yaminbkk/NexusMem/compare/v0.3.3...v0.4.0
|
|
783
|
+
[0.3.3]: https://github.com/yaminbkk/NexusMem/compare/v0.3.2...v0.3.3
|
|
784
|
+
[0.3.2]: https://github.com/yaminbkk/NexusMem/compare/v0.3.1...v0.3.2
|
|
785
|
+
[0.3.1]: https://github.com/yaminbkk/NexusMem/compare/v0.3.0...v0.3.1
|
|
786
|
+
[0.3.0]: https://github.com/yaminbkk/NexusMem/compare/v0.2.0...v0.3.0
|
|
787
|
+
[0.2.0]: https://github.com/yaminbkk/NexusMem/compare/v0.1.2...v0.2.0
|
|
788
|
+
[0.1.2]: https://github.com/yaminbkk/NexusMem/compare/v0.1.1...v0.1.2
|
|
789
|
+
[0.1.1]: https://github.com/yaminbkk/NexusMem/compare/v0.1.0...v0.1.1
|
|
790
|
+
[0.1.0]: https://github.com/yaminbkk/NexusMem/releases/tag/v0.1.0
|