markdown-memory 0.1.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,579 @@
1
+ Metadata-Version: 2.4
2
+ Name: markdown-memory
3
+ Version: 0.1.0
4
+ Summary: MCP server that hands coding agents the one Markdown section that answers the question, not the whole file. Local hybrid search: BM25 keywords + ONNX embeddings in SQLite. No API keys, no network.
5
+ Keywords: mcp,mcp-server,model-context-protocol,coding-agent,ai-agents,markdown,documentation,rag,retrieval-augmented-generation,hybrid-search,semantic-search,vector-search,full-text-search,bm25,embeddings,onnx,sqlite,sqlite-vec,fts5
6
+ Author: Hesham Karm
7
+ Author-email: Hesham Karm <24391550+hishamkaram@users.noreply.github.com>
8
+ License-Expression: MIT
9
+ License-File: LICENSE
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Intended Audience :: Developers
12
+ Classifier: Programming Language :: Python :: 3.11
13
+ Classifier: Programming Language :: Python :: 3.12
14
+ Classifier: Programming Language :: Python :: 3.13
15
+ Classifier: Programming Language :: Python :: 3.14
16
+ Classifier: Operating System :: POSIX :: Linux
17
+ Classifier: Operating System :: MacOS
18
+ Classifier: Topic :: Software Development :: Documentation
19
+ Classifier: Topic :: Text Processing :: Indexing
20
+ Classifier: Typing :: Typed
21
+ Requires-Dist: fastembed>=0.8.0
22
+ Requires-Dist: huggingface-hub>=0.26
23
+ Requires-Dist: markdown-it-py>=4.2.0
24
+ Requires-Dist: mcp[cli]>=2.2.0,<3
25
+ Requires-Dist: numpy>=1.26
26
+ Requires-Dist: onnxruntime>=1.20
27
+ Requires-Dist: pydantic>=2.0
28
+ Requires-Dist: sqlite-vec>=0.1.9
29
+ Requires-Dist: tokenizers>=0.20
30
+ Requires-Python: >=3.11
31
+ Project-URL: Homepage, https://github.com/hishamkaram/markdown-memory
32
+ Project-URL: Repository, https://github.com/hishamkaram/markdown-memory
33
+ Project-URL: Issues, https://github.com/hishamkaram/markdown-memory/issues
34
+ Description-Content-Type: text/markdown
35
+
36
+ # markdown-memory
37
+
38
+ **Your coding agent reads whole Markdown files to answer one question. This returns the
39
+ section that answers it.**
40
+
41
+ Ask an agent a question about your docs and it opens the files that might answer it, whole.
42
+ Most of what lands in its context is about something else, and the part you wanted competes
43
+ with it. markdown-memory indexes your documentation by heading, so the same question comes
44
+ back as a few sections, each addressable by its breadcrumb and quoted verbatim.
45
+
46
+ <picture>
47
+ <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/hishamkaram/markdown-memory/main/docs/assets/how-it-works-dark.svg">
48
+ <source srcset="https://raw.githubusercontent.com/hishamkaram/markdown-memory/main/docs/assets/how-it-works-light.svg">
49
+ <img src="https://raw.githubusercontent.com/hishamkaram/markdown-memory/main/docs/assets/how-it-works-light.png" width="100%"
50
+ alt="One question asked of four documentation files. Reading them whole costs 14,987
51
+ tokens. markdown-memory splits them at every heading, ranks by keywords and by
52
+ vectors, fuses the two, and returns five sections totalling 2,887 tokens - the
53
+ one that answers is 403.">
54
+ </picture>
55
+
56
+ Measured on this repository's own documentation - `README.md`, `CLAUDE.md`, `AGENTS.md` and
57
+ `docs/evaluation-protocol.md`, 14,987 tokens in all:
58
+
59
+ ```
60
+ search_docs("where does the embedding model get downloaded")
61
+
62
+ 407 tok README.md markdown-memory > The embedding model > What downloads, when, and where
63
+ 712 tok README.md markdown-memory > The embedding model > Pre-download it, or install offline
64
+ 647 tok README.md markdown-memory
65
+ 738 tok CLAUDE.md markdown-memory > Commands
66
+ 383 tok README.md markdown-memory > The embedding model > What is checked before the model is loaded
67
+ ```
68
+
69
+ **2,887 tokens instead of 14,987**, and the section that actually answers is 403 - a
70
+ thirty-fourth of what reading the files costs. Every hit carries its full text, so a good
71
+ answer usually needs no follow-up call at all.
72
+
73
+ It is a local [Model Context Protocol](https://modelcontextprotocol.io) server - MCP is the
74
+ protocol agents use to call tools - and it runs entirely on your machine: parsing with
75
+ `markdown-it-py`, embeddings with
76
+ [EmbeddingGemma-300m](https://huggingface.co/onnx-community/embeddinggemma-300m-ONNX)
77
+ (4-bit ONNX on CPU, 768 dimensions), storage in SQLite - one database per documentation
78
+ root, with a keyword index and two vector indexes over it. No API key, no network after the
79
+ first model download, nothing leaves the machine.
80
+
81
+ ## Install
82
+
83
+ ### Prerequisites
84
+
85
+ - **Python 3.11 or newer.** The project develops on 3.12 and CI runs 3.11 through 3.14 on
86
+ x86-64 Linux.
87
+ - **[uv](https://docs.astral.sh/uv/)**, which manages the interpreter and the dependencies:
88
+ ```bash
89
+ curl -LsSf https://astral.sh/uv/install.sh | sh # macOS / Linux
90
+ ```
91
+ - **Roughly 1 GB of disk**: ~218 MB for the embedding model, the rest for the index.
92
+ - Linux or macOS, x86-64 or arm64. Everything runs on CPU; there is no GPU path and no API
93
+ key. CI covers arm64 Linux on every change and Apple Silicon on `main`, because the 4-bit
94
+ graph picks its `MatMulNBits` kernel from what the CPU offers rather than from the file -
95
+ see [What downloads, when, and where](#what-downloads-when-and-where). Those kernels do
96
+ not all return the same numbers: the gate measures four CPU families and the same text
97
+ embeds up to 1.6e-3 cosine apart between x86-64 and arm64. Copying an index between two
98
+ machines therefore searches vectors from one kernel with queries from another, which
99
+ nothing detects - and which was measured, on the labelled set and through the real
100
+ ranking, to move no top result and lose no answer. It reshuffles positions two to five.
101
+
102
+ ### Get it
103
+
104
+ ```bash
105
+ uv tool install markdown-memory # or: pipx install markdown-memory
106
+ markdown-memory --download-model # optional: fetch the ~218 MB model now
107
+ ```
108
+
109
+ `uv tool install` puts the `markdown-memory` command in `~/.local/bin`; if your shell cannot
110
+ find it, `uv tool update-shell` adds that directory to your PATH. Fetching the model up
111
+ front is optional - the first search does it otherwise - but it keeps a 218 MB download
112
+ out of your first question.
113
+
114
+ ### Register it with your client
115
+
116
+ One registration serves every project. The server indexes the project the client was
117
+ started in - Claude Code tells it (`CLAUDE_PROJECT_DIR`), Codex starts it there - and each
118
+ project gets its own index.
119
+
120
+ ```bash
121
+ claude mcp add --scope user markdown-memory -- markdown-memory # Claude Code
122
+ codex mcp add markdown-memory -- markdown-memory # Codex
123
+ ```
124
+
125
+ `claude mcp list` and `codex mcp list` should show it. Upgrade with
126
+ `uv tool upgrade markdown-memory` (or `pipx upgrade markdown-memory`); remove it with
127
+ `claude mcp remove --scope user markdown-memory`, `codex mcp remove markdown-memory` and
128
+ `uv tool uninstall markdown-memory` (or `pipx uninstall markdown-memory`).
129
+
130
+ Desktop apps (Claude Desktop, Cursor) usually do not see `~/.local/bin`, so give them the
131
+ absolute path that `command -v markdown-memory` prints. Claude Desktop has no project
132
+ either, so name the documentation root:
133
+
134
+ ```json
135
+ {
136
+ "mcpServers": {
137
+ "markdown-memory": {
138
+ "command": "/home/you/.local/bin/markdown-memory",
139
+ "env": { "MARKDOWN_MEMORY_DOCS_DIR": "/absolute/path/to/your/docs" }
140
+ }
141
+ }
142
+ }
143
+ ```
144
+
145
+ `uvx markdown-memory` works in place of the installed command too, with one catch: its
146
+ first launch downloads about 250 MB of dependencies before the server can answer, longer
147
+ than Codex waits by default - set `startup_timeout_sec = 60` under
148
+ `[mcp_servers.markdown-memory]` in `~/.codex/config.toml` if you go that way.
149
+
150
+ ### From source
151
+
152
+ For working on markdown-memory itself:
153
+
154
+ ```bash
155
+ git clone https://github.com/hishamkaram/markdown-memory
156
+ cd markdown-memory
157
+ uv sync
158
+ uv run markdown-memory --version
159
+ ```
160
+
161
+ The `.mcp.json` in the clone runs that checkout (`uv run markdown-memory`) whenever Claude
162
+ Code is opened in it, with no edit.
163
+
164
+ ### Index once, then search
165
+
166
+ The server keeps its documentation root indexed by itself. The first `search_docs` (or
167
+ `list_documents` without a directory) after it starts begins a catch-up run in the
168
+ background - whatever changed while no server ran, new files included - and answers from
169
+ the index as it stood, with `index_status.indexing` set. Nothing runs before that first
170
+ call: loading the model holds the interpreter for seconds, and a client waiting on the
171
+ handshake gives up quickly. After that, each of those calls decides whether another run is
172
+ due:
173
+
174
+ - a file it sees edited, deleted or unreadable starts one, at most every 10 seconds;
175
+ - a weights mismatch a search has just recorded starts one at once;
176
+ - otherwise one walks the tree every 5 minutes while the server is in use, which is what
177
+ finds a file nobody indexed yet.
178
+
179
+ One run at a time, in one thread, stopped between two documents when the server shuts down;
180
+ a stopped run leaves what a killed run leaves, and the next one resumes. While it runs,
181
+ `index_status.indexing` is `true` and the message says so instead of asking for
182
+ `index_directory`. Nothing watches the filesystem, so a server nobody is using does not
183
+ spend anything; `read_section` and `get_document_outline` do not start a run. The first
184
+ run of a large tree embeds everything and competes with queries for CPU until it is done;
185
+ `--no-auto-index` (or `MARKDOWN_MEMORY_AUTO_INDEX=0`) turns all of this off, and then:
186
+
187
+ ```
188
+ index_directory() # scans the docs root, embeds what it finds
189
+ list_documents() # index_status.coverage should read "verified"
190
+ ```
191
+
192
+ The first run downloads the embedding model and then indexes, so it is the slow one - see
193
+ [The embedding model](#the-embedding-model). After that, indexing is incremental: re-running
194
+ it costs milliseconds when nothing changed. If `coverage` reads `"unknown"`, something was
195
+ missed and `index_status.message` says what to run.
196
+
197
+ `index_status` also carries `changed_files`: indexed documents a cheap probe could not
198
+ confirm are still what was indexed - their bytes differ, or they are gone, unreadable, or no
199
+ longer a regular file. It is best-effort in both directions. Unreadability is noticed only
200
+ where a moved timestamp made it read the file at all: a file whose permissions changed and
201
+ whose time did not is answered from the time, and it is the next `index_directory` that
202
+ records the failure and takes `coverage` to `"unknown"`. A path that cannot even be
203
+ `stat`-ed - its parent directory lost its permissions, say - has no timestamp to compare
204
+ and is counted straight away. It counts only rows the index
205
+ holds, so a file nobody has indexed yet is not among them: finding those means walking the
206
+ tree, which is the expensive half of indexing and not something a search should pay for.
207
+ And it reads bytes only where the modification time moved, so an edit that restores a file's
208
+ own timestamp is missed - indexing is not fooled by that, since it hashes every file it
209
+ walks; what is missed is only the hint that running it is worth it. A zero means nothing was
210
+ detected, not that every file was re-hashed. `coverage` stays `"verified"` while the count is
211
+ non-zero: the walk really did finish and really did read every file it found. What moved on
212
+ is the tree, and the message says so.
213
+
214
+ ## The embedding model
215
+
216
+ ### What downloads, when, and where
217
+
218
+ The first search (or `markdown-memory --download-model`) fetches three files from
219
+ [`onnx-community/embeddinggemma-300m-ONNX`](https://huggingface.co/onnx-community/embeddinggemma-300m-ONNX)
220
+ at a pinned revision: the 4-bit ONNX graph, its external weights, and the tokenizer.
221
+ About **218 MB**, once per machine, into:
222
+
223
+ ```
224
+ $XDG_CACHE_HOME/markdown-memory/models/embeddinggemma-300m-onnx-5090578d9565/ # ~/.cache/... by default
225
+ ```
226
+
227
+ The revision is part of the folder name, so moving the pin fetches the new weights instead
228
+ of serving the old ones under a name that claims to be the new ones. One revision can publish
229
+ several graphs, though, so the folder name is not the whole answer: a `.verified` stamp that
230
+ does not name exactly the files this version needs - a cache left behind by the int8 graph at
231
+ this same revision, say - is rejected and the files are fetched, rather than half-trusted.
232
+
233
+ The cache is shared by every project on purpose - the weights are identical and read-only,
234
+ so copying them per project would be pure waste. Point `MARKDOWN_MEMORY_MODEL_CACHE`
235
+ somewhere else to move it.
236
+
237
+ The graph is `onnx/model_q4.onnx`, published 4-bit: the 262144x768 vocabulary table is
238
+ quantized and gathered by a single `GatherBlockQuantized`, and the projections run as
239
+ `MatMulNBits`, so the table is never expanded to float32. Nothing is derived or rewritten on
240
+ your machine - earlier versions ran the int8 graph and patched it here to get the same
241
+ effect. Using Gemma is covered by the
242
+ [Gemma Terms of Use](https://ai.google.dev/gemma/terms); this repository distributes no
243
+ weights.
244
+
245
+ ### What is checked before the model is loaded
246
+
247
+ Every file's size and sha256 is pinned at the pinned revision. After a full check, a
248
+ `.verified` stamp records each file's size, mtime, ctime, inode and device, so an ordinary
249
+ start is a handful of `stat` calls rather than most of a second of hashing. Anything that
250
+ differs sends the files back to be hashed against the pins, and a file that does not match
251
+ is re-fetched - just that file.
252
+
253
+ **Guaranteed:** any change to a model file's contents or metadata since it was verified is
254
+ caught. `cp -p`, `tar x` and `rsync --inplace` can overwrite a file and restore its mtime,
255
+ which is why ctime is in the stamp - nothing in user space can set that back. **Not
256
+ guaranteed:** silent disk bit-rot, with no write at all.
257
+
258
+ A failure to *load* verified files is not treated as damage: it means onnxruntime,
259
+ permissions or memory, so the error is raised as it stands and nothing is downloaded.
260
+
261
+ Verification and loading hold a shared `flock`; downloading and repairing hold it
262
+ exclusively, so several servers starting at once download once between them. A model cache
263
+ on **NFS or SMB shared between machines is not supported** - `flock` can be local-only
264
+ there.
265
+
266
+ Nothing is ever deleted to reclaim space. After a successful start, one log line on stderr
267
+ names any other `embeddinggemma-300m-onnx*` folders and what they cost, and any file sitting
268
+ in the current folder that this version does not use - the int8 graph an upgrade left behind
269
+ weighs about 310 MB - and removing them is yours to do.
270
+
271
+ ### Pre-download it, or install offline
272
+
273
+ To fetch the model deliberately rather than on the first query:
274
+
275
+ ```bash
276
+ markdown-memory --download-model # the default, EmbeddingGemma
277
+ markdown-memory --download-model --embedder bge-small # or the light preset
278
+ ```
279
+
280
+ It downloads and loads the configured model, then exits; it touches no index.
281
+
282
+ For a machine with no network, copy the three files into
283
+ `$XDG_CACHE_HOME/markdown-memory/models/embeddinggemma-300m-onnx-5090578d9565/`, keeping
284
+ `onnx/model_q4.onnx`, `onnx/model_q4.onnx_data` and `tokenizer.json` where they are. The
285
+ first load hashes them once, writes the `.verified` stamp, and never touches the network.
286
+
287
+ `bge-small` is downloaded by `fastembed`, which pins no revision: if that cache is deleted,
288
+ it can come back with different weights under the same model name. EmbeddingGemma's
289
+ weights change on purpose, when a release moves its pin or graph. Either way, every document
290
+ records the weights that embedded it, and the index records the one revision that vouches
291
+ for all of its vectors:
292
+
293
+ - **Indexing repairs in place.** A document stamped by other weights is re-embedded like a
294
+ changed file, even when its bytes are the same. Before the first new vector is written the
295
+ index is marked as being re-embedded; nothing is deleted first, so keyword search answers
296
+ throughout, and a run that is killed resumes where it stopped, because each stamp is
297
+ written with its vectors. At the end of a run the revision is restored only once no
298
+ vector-bearing document in the whole database - every root's, and one indexed on its own
299
+ inside a pruned directory such as `.venv` - is stamped by anything else. Until then
300
+ `index_status.coverage` reads `"unknown"` and its message names the directories still to
301
+ be indexed.
302
+ - **Searching** compares for itself, after embedding the query, and falls back to keyword
303
+ ranking alone when the answer differs - for the old weights and the new alike while a
304
+ repair is under way. It does not wait to be told: weights can change while no Markdown
305
+ file does, and then there is no indexing run to notice. The search that notices records
306
+ it, and the next `index_directory` loads the model first and re-embeds what it finds.
307
+
308
+ A model that loads but cannot say which weights it is may not write into an index that
309
+ names its weights: its vectors could never be told apart from the ones already stored. An
310
+ index whose vectors no revision vouches for - built while the weights could not be read -
311
+ is not ranked against a query from weights that can name themselves - from the upgrade to
312
+ schema v6 on - and the first run with such weights re-embeds it. A server that loaded
313
+ `bge-small` while its revision could not be read keeps it unnamed until it restarts: the
314
+ revision is read beside the weights it loads, never after them.
315
+
316
+ ### Presets
317
+
318
+ Two presets, chosen with `MARKDOWN_MEMORY_EMBEDDER`:
319
+
320
+ | Preset | Dimensions | Download | Peak RAM | Indexing | Held-out Top-1 / Top-3 / Top-5 |
321
+ | --- | --- | --- | --- | --- | --- |
322
+ | `embeddinggemma` (default) | 768 | ~218 MB | ~0.65 GB | ~7 vectors/s | 88% / 97% / 97% |
323
+ | `bge-small` | 384 | ~65 MB | ~1.1 GB | ~12 vectors/s | 68% / 82% / 88% |
324
+
325
+ Accuracy is the frozen baseline in `scripts/eval_data/baseline.json`, recorded by
326
+ `scripts/eval_retrieval.py` over a 54-section corpus and the 34 **held-out** paraphrase
327
+ queries, which were written before any parameter was tuned. The **dev** set - the one
328
+ tuning is allowed to look at, and deliberately harder - scores 71% / 85% / 94% with
329
+ EmbeddingGemma and 53% / 68% / 79% with bge-small. Exact identifiers - flags, environment
330
+ variables, error strings - are 100% Top-1 with either preset, because FTS5 answers them.
331
+ Query latency is not in the table on purpose: it swings by 2-3x with what else the machine
332
+ is doing, so the baseline records it as informational and so should you.
333
+
334
+ Switching preset **discards the whole index**: the two produce vectors of different sizes,
335
+ which cannot be compared, so every documentation root has to be indexed again.
336
+ `index_directory` reports that when it happens.
337
+
338
+ **The first index of a large documentation set is slow.** Every paragraph, list item, table
339
+ row and code block costs one vector; a section costs none of its own, because its vector is
340
+ pooled from its passages. Measured on a 97-file set: 1,682 sections and 8,453 passages -
341
+ 10,135 vectors - indexed in **24 minutes** at 7.1 vectors per second, peaking at 667 MB,
342
+ after which a query over those 1,682 sections takes about 200 ms. Budget for it, and run it
343
+ once: indexing is incremental by SHA-256, so a
344
+ re-index that finds nothing changed takes milliseconds (8 ms for 36 sections) and only
345
+ edited files are re-embedded. `bge-small` indexes several times faster at a real cost in
346
+ accuracy.
347
+
348
+ EmbeddingGemma is distributed under the [Gemma Terms of Use](https://ai.google.dev/gemma/terms);
349
+ the revision is pinned. See [License](#license) for what that means for you.
350
+
351
+ ### When it goes wrong
352
+
353
+ - **The server starts even when the model cannot load.** Nothing loads it until the
354
+ first search, so the failure surfaces there, as
355
+ `Cannot load embedding model onnx-community/embeddinggemma-300m-ONNX: ...`. So "the server
356
+ is running" is not evidence the model is there; `markdown-memory --download-model` is.
357
+ - **The first `search_docs` can block for the length of a 218 MB download.** Pre-download it
358
+ (above) if that matters.
359
+ - **All logging goes to stderr.** stdout carries JSON-RPC frames only, so a client that
360
+ shows you "the output" may be showing you nothing. Set `MARKDOWN_MEMORY_LOG_LEVEL=DEBUG`
361
+ and read stderr.
362
+ - **Switching preset discards the index.** The two models produce vectors of different
363
+ sizes, which cannot be compared, so every root must be indexed again. `index_directory`
364
+ says so when it happens.
365
+ - **Searches come back empty or stale.** Run `index_directory` again; it is incremental, so
366
+ it is cheap. If `index_status.coverage` stays `"unknown"`, its `message` names the files
367
+ that could not be read.
368
+
369
+ ## Tools
370
+
371
+ | Tool | Purpose |
372
+ | --- | --- |
373
+ | `index_directory(directory=None)` | Scan a tree, (re)index new/changed `.md` files (SHA-256), purge deleted ones |
374
+ | `list_documents(directory="")` | `{documents, index_status}`: indexed paths, titles and section counts, and whether a full index run vouches for them |
375
+ | `get_document_outline(file_path)` | Hierarchical TOC with line ranges and token estimates |
376
+ | `read_section(file_path, heading_path, include_subsections=False)` | Verbatim text of one section |
377
+ | `search_docs(query, limit=5)` | `{results, index_status}`: BM25 + passage-level vector search fused with Reciprocal Rank Fusion (k = 60); each hit reports the `matched_passage`, and `index_status` says whether the tree searched is known to be whole |
378
+
379
+ Sections are addressed by breadcrumb: `Root > Child > Subchild`. Oversized sections
380
+ (> ~800 tokens) are stored as `Root > Child (Part 1)`, `(Part 2)`, ...; reading the base
381
+ path reassembles them byte-for-byte. `file_path` may be absolute, relative to the docs
382
+ root, or any unique path suffix. `heading_path` is matched exactly first, then ignoring
383
+ spacing around `>`, then ignoring case, then as a trailing fragment (`Child > Subchild`
384
+ or just the title); an ambiguous request lists the exact candidates.
385
+
386
+ ## Configuration
387
+
388
+ | Environment variable | CLI flag | Default |
389
+ | --- | --- | --- |
390
+ | `MARKDOWN_MEMORY_DOCS_DIR` | `--docs-dir` | `$CLAUDE_PROJECT_DIR` if the client exports it, else the working directory |
391
+ | `MARKDOWN_MEMORY_DB` | `--db` | `$XDG_DATA_HOME/markdown-memory/projects/<root>-<digest>/index.db` — one index per docs root |
392
+ | `MARKDOWN_MEMORY_MODEL_CACHE` | - | `$XDG_CACHE_HOME/markdown-memory/models` (`~/.cache/...`) |
393
+ | `MARKDOWN_MEMORY_EXCLUDE` | `--exclude` (repeatable) | nothing excluded |
394
+ | `MARKDOWN_MEMORY_LOG_LEVEL` | `--log-level` | `INFO` |
395
+ | `MARKDOWN_MEMORY_EMBEDDER` | `--embedder` | `embeddinggemma` (or `bge-small`) |
396
+ | `MARKDOWN_MEMORY_THREADS` | - | unset: onnxruntime picks. A positive integer caps the threads one embedding pass may use |
397
+ | `MARKDOWN_MEMORY_INDEX_WORKERS` | - | `2` - files read, parsed and embedded at the same time while indexing |
398
+ | `MARKDOWN_MEMORY_AUTO_INDEX` | `--no-auto-index` | on - `0`, `false`, `off` or `no` stops the server indexing its root by itself |
399
+
400
+ The default preset does not spin-wait between operators, which is what makes a query cost ~0.6 s of
401
+ CPU instead of ~5.5 s and leaves the process idle at 0 while nothing is being asked of it; the
402
+ thread *count* is left to onnxruntime, and `MARKDOWN_MEMORY_THREADS` is there for a machine that
403
+ disagrees with its choice. `bge-small` runs through `fastembed`, which exposes no such switch, so
404
+ for that preset the count is the only lever: `MARKDOWN_MEMORY_THREADS=4` took one query from 718 ms
405
+ of CPU to 95 ms.
406
+
407
+ Indexing embeds several files at once and writes them from one thread, in the order the
408
+ tree was walked. One ONNX session is shared, and its weights are mapped once however many
409
+ threads run against it, so each extra worker costs about the 150 MB of one forward pass.
410
+ Measured over 24 files of the eval corpus (879 passages): 196.3 s with one worker, 145.9 s
411
+ with two (1.35x) and 96.3 s with four (2.04x) - less than the embedding speed-up alone,
412
+ because parsing and the writes stay serial and a long file holds the head of the queue.
413
+ Two is the default because its peak measures around 0.8 GB (757-814 MB across runs),
414
+ well inside what a tool running beside an editor should take; raise `MARKDOWN_MEMORY_INDEX_WORKERS` on a machine with cores to spare.
415
+
416
+ `MARKDOWN_MEMORY_EXCLUDE` takes glob patterns separated by commas (only commas - a colon
417
+ would split a pattern that contains one). A pattern with no `/` matches that name at any
418
+ depth, the way `.gitignore` treats one: `eval_data` excludes `scripts/eval_data/corpus/`.
419
+ A pattern containing `/` is anchored at the documentation root (`tests/fixtures`,
420
+ `docs/generated/*`). Matching is case-sensitive everywhere, and a matching directory is
421
+ pruned, so its subtree costs nothing to skip. Files already indexed before an exclusion
422
+ was added are purged on the next index.
423
+
424
+ Without it, a repository that keeps fixtures, vendored documentation or a test corpus
425
+ in-tree indexes them as if they were its own docs.
426
+
427
+ Documents are stored under their absolute path, so one database *can* hold several
428
+ projects - but **search only ever answers from the root this server was started with**,
429
+ and `list_documents` shows only that root. Each project gets its own database by default,
430
+ so this matters only if you point two of them at one file with `MARKDOWN_MEMORY_DB`: the
431
+ second is then indexed, invisible, and paying for itself in disk.
432
+
433
+ ### One index per project
434
+
435
+ The user-level registration above already gives every project its own index. To set
436
+ exclusions for one repository, drop a `.mcp.json` like this into it (it takes precedence
437
+ there):
438
+ Each project owns its index without being told to: the database is keyed on the
439
+ documentation root it serves, so a project gets its own exclusions and no chance of
440
+ another project's sections - or another project's documents, which stay resolvable by
441
+ path across any database they share - appearing in its results.
442
+
443
+ ```json
444
+ {
445
+ "mcpServers": {
446
+ "markdown-memory": {
447
+ "command": "markdown-memory",
448
+ "env": {
449
+ "MARKDOWN_MEMORY_EXCLUDE": "vendor,third_party,tests/fixtures"
450
+ }
451
+ }
452
+ }
453
+ }
454
+ ```
455
+
456
+ Any path you do add is written relative, on purpose. Claude Code expands only real
457
+ environment variables in `.mcp.json`: `${workspaceFolder}` is a VS Code idea, and even
458
+ `${CLAUDE_PROJECT_DIR}` is not set at expansion time - measured on Claude Code 2.1.278,
459
+ both produce a *"Missing environment variables"* warning and are passed through as literal
460
+ text, which would make the server index a directory named `${workspaceFolder}` and report
461
+ success over zero files. The server refuses such a value outright, and resolves a relative
462
+ path against `CLAUDE_PROJECT_DIR` (which Claude Code *does* export to the spawned server),
463
+ falling back to the working directory only when that is not set. The docs root defaults to
464
+ that same project root, so it needs no entry. Each worktree of a repository is its own
465
+ directory, so each gets its own index.
466
+
467
+ By default nothing is written into the repository. (A *relative* `MARKDOWN_MEMORY_DB`
468
+ is resolved against the project root and does land inside it - `.gitignore` covers
469
+ `.markdown-memory/` for that reason, and any other relative path you choose is yours to
470
+ ignore.)
471
+ The index lives under `$XDG_DATA_HOME/markdown-memory/projects/`, in a directory named for
472
+ the documentation root and a digest of its resolved path - out of reach of `git clean
473
+ -xdf`, writable when the checkout is not, and on local disk when the checkout is on a
474
+ network share, where SQLite's write-ahead log cannot take the locks it needs. Set
475
+ `MARKDOWN_MEMORY_DB` to override it; a relative value is resolved against the project
476
+ root. The index is a cache of the Markdown files and is rebuilt from them, so deleting it
477
+ costs only the time to index again.
478
+
479
+ ## How search ranks
480
+
481
+ 1. **Keywords (FTS5, BM25).** Stopwords are dropped, identifiers are kept verbatim. A hit
482
+ only counts if it covers at least half of the query's IDF mass or matches an
483
+ identifier-like term - a stray match on "data" or "deploy" no longer outvotes the
484
+ vector index.
485
+ 2. **Vectors.** Every paragraph, list item, table row (rendered as `Header: cell; ...`) and
486
+ code block - also inside block quotes and list items - is embedded separately. The
487
+ section's own vector is the mean of those passage vectors, not a separate embedding of
488
+ the whole section: the model truncates at 512 tokens, which a long section exceeds. It
489
+ is an aggregate of the passages rather than independent evidence about the section, and
490
+ it finds no section that the passages do not. A passage longer than 600 characters is split into
491
+ consecutive windows at line, sentence or word boundaries, so a long command list or
492
+ configuration block keeps a vector for all of itself rather than for its first 600
493
+ characters. A table split across `(Part n)` sections keeps its header for every part. A
494
+ section is ranked by its closest vector, so one relevant table row is enough.
495
+ Heading-only sections have no vectors and are never returned ahead of their children.
496
+ 3. **Reciprocal Rank Fusion** of the two rankings.
497
+
498
+ Cross-encoder rerankers (MiniLM, bge-reranker-base, jina, ColBERT) were benchmarked and
499
+ rejected: every one lowered accuracy on technical documentation and cost 2-12 s a query.
500
+
501
+ All logging goes to **stderr**. stdout carries JSON-RPC frames only.
502
+
503
+ ## How malformed Markdown is handled
504
+
505
+ - **Skipped heading levels** (`#` then `####`): a heading stack pops every level `>= L`
506
+ before pushing, so breadcrumbs stay well formed.
507
+ - **Preamble**: badges/summary before the first heading become `[Overview / Preamble]`.
508
+ - **YAML front matter**: kept out of the AST (CommonMark would read it as a setext
509
+ heading) and used as a title fallback. A leading `---` rule followed by prose is not
510
+ mistaken for it. One case is inherently ambiguous - a single `key: value` line between
511
+ two `---` lines - and is read as front matter, as static-site generators do, unless the
512
+ key is an admonition word (`Note:`, `Warning:`, `TODO:` ...).
513
+ - **Unclosed code fences**: CommonMark runs them to EOF - or to the closing marker of a
514
+ *later* fence - swallowing the sections in between. The fence is closed before the next
515
+ blank-line-preceded ATX heading instead. A level-1 `# ...` line is treated as a comment
516
+ unless the fence language cannot have `#` comments (JSON, Go, ...). This is a heuristic:
517
+ a fence that merely *looks* closed is only cut on strong evidence (it contains another
518
+ opening fence with an info string, or the document ends inside a bare fence and the cut
519
+ makes the rest well formed), and never when the repair would lose a heading that was
520
+ already found. Repair work is capped per document, so a pathological file costs a
521
+ bounded number of extra parses rather than one per fence.
522
+ - **No headings / walls of text**: split on paragraph boundaries, then on lines, then on
523
+ whitespace. A fenced block is only cut when it exceeds the limit by itself.
524
+ - **Colliding breadcrumbs** get a ` [2]`, ` [3]` suffix (also against generated
525
+ `(Part n)` paths), so every stored path addresses exactly one section.
526
+ - **Headings like `Option<T>`** keep their type parameter; only lower-case formatting
527
+ tags (`<b>`, `<sub>`, `<a>`, ...) are stripped from titles.
528
+
529
+ ## Indexing rules
530
+
531
+ - `.md` / `.markdown`, regular files only, at most 10 MB; symlinked directories are not
532
+ followed. `.git`, `node_modules`, virtualenvs and tool caches are pruned - index such a
533
+ tree by passing a directory *inside* it, and it is then left alone when an ancestor is
534
+ re-indexed.
535
+ - A document is purged only when the walk could have found it and did not. Files under a
536
+ directory that cannot be listed are kept and the directory is reported as an error.
537
+
538
+ ## Development
539
+
540
+ ```bash
541
+ uv run ruff check . && uv run ruff format --check .
542
+ uv run mypy --strict src/
543
+ uv run pytest -v # unit + integration (real ONNX model for semantic tests)
544
+ uv run python scripts/live_test.py # spawns the server, drives it over stdio JSON-RPC
545
+ uv run python scripts/eval_retrieval.py --show-misses # retrieval accuracy; fails on regression
546
+ uv run python scripts/eval_retrieval.py --rebuild # ... after discarding the cached index
547
+ uv run python scripts/reindex_docs.py DIR --force # forced re-index + integrity verification
548
+ scripts/check.sh # the whole pre-commit gate, fail-fast
549
+ ```
550
+
551
+ ### For AI coding agents
552
+
553
+ `CLAUDE.md` is the developer guide (commands, layout, architecture rules, the retrieval
554
+ regression policy). `AGENTS.md` and `.cursorrules` carry the short form for Codex and
555
+ Cursor. All three share one block of rules for *using* the MCP tools - search first, then
556
+ outline, then read a single section; never dump whole Markdown files - and a test keeps the
557
+ three copies identical and in step with the code. Project skills for Claude Code live in
558
+ `.claude/skills/`: `run-eval`, `reindex-docs`, `test-regression`.
559
+
560
+ Built on the MCP Python SDK 2.x, where `FastMCP` was renamed `MCPServer`.
561
+
562
+ ## License
563
+
564
+ MIT - see [`LICENSE`](LICENSE). Two things in this repository are *not* covered by it,
565
+ because they are not ours to license:
566
+
567
+ - **The embedding models.** EmbeddingGemma-300m, the default, is distributed under the
568
+ [Gemma Terms of Use](https://ai.google.dev/gemma/terms), which are not an OSI-approved
569
+ open-source licence; the revision is pinned. The `bge-small` preset is two licences at
570
+ once: the `BAAI/bge-small-en-v1.5` weights are MIT, and `fastembed`, which loads them,
571
+ is Apache-2.0. Nothing is bundled - both are downloaded on first use - but if the Gemma
572
+ terms do not suit you, `MARKDOWN_MEMORY_EMBEDDER=bge-small` avoids them entirely.
573
+ - **The evaluation corpus.** `scripts/eval_data/corpus_v2/` is third-party documentation
574
+ vendored verbatim from five projects, pinned by commit, and used only to measure
575
+ retrieval accuracy. Each upstream keeps its own licence, and its licence and NOTICE
576
+ files travel with it in `scripts/eval_data/corpus_v2_licenses/`. The table in
577
+ [`scripts/eval_data/corpus_v2_LICENSES.md`](scripts/eval_data/corpus_v2_LICENSES.md)
578
+ says what came from where; both it and the manifest beside it are generated by
579
+ `scripts/fetch_eval_corpus.py`, so edit the script rather than the files.
@@ -0,0 +1,21 @@
1
+ markdown_memory/__init__.py,sha256=-SGV6gYrwplcR-J1maJ4_8X7td8Tq-iyASnggOOx9wU,1125
2
+ markdown_memory/autoindex.py,sha256=mZnBuSoyFhgAbGKYGihHFgB_MeEWnvt46acUja_wIEM,7233
3
+ markdown_memory/config.py,sha256=USCCJ2NEXb6kaaJgi_UULHCFLZLtN5WG9r3bWK4B4Ac,9966
4
+ markdown_memory/db.py,sha256=4JHsxBa_tIIuvL4lW0F84hrv7wYdjd7ItaYxi2ck7Ek,72227
5
+ markdown_memory/discovery.py,sha256=NTPQLVTl5jQ2pXsbas5-SEbfNFO9Ii5pxmPuFcEpHhM,8515
6
+ markdown_memory/embedders.py,sha256=PlGwtd1xH7pelbEamE7JEum4R0qREJuW7fu4escj8gc,23141
7
+ markdown_memory/exceptions.py,sha256=OsVXFN198VETNo00-Oe72SdI_wMNFG_QDEchrP0qAWM,2482
8
+ markdown_memory/freshness.py,sha256=bchaOXp1--FELa2DRywvR3aejlboFLZ7-uU_DFu_TrU,7102
9
+ markdown_memory/headings.py,sha256=QlrmdkTpKaDBgnoGMRE0LU9PlKCl3IB_zdRbbEuo078,7059
10
+ markdown_memory/indexer.py,sha256=VS_m5lVX7F1n3UyodWkTQDCu6W0UAn6dx0z5yFN6vb4,36718
11
+ markdown_memory/model_cache.py,sha256=Txv6ajf2i3ttA9L6HRfGU-cXuRS2NmZ9owW1yjEn134,11186
12
+ markdown_memory/models.py,sha256=hA2BdSVTlfwsekp7HH9RdLLRjLg_A1IfqP94cCXBr4Q,12620
13
+ markdown_memory/parser.py,sha256=qWj23NFXBIoncAPaRRQFLKOdavchYpVN9t76FlavowM,37693
14
+ markdown_memory/py.typed,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
15
+ markdown_memory/search.py,sha256=bam9luBsojtB00YhBZLu2H5J9XxCYahzufiR0AZRzf0,22265
16
+ markdown_memory/server.py,sha256=vsXRyNCXAH77I43nYiCCkUJ4wbKjuv_6ybL8vNDbipk,22709
17
+ markdown_memory-0.1.0.dist-info/licenses/LICENSE,sha256=Ta4m9NqLCdjDz0C7vcR1k0EFYoHRt9XdKyh8ho12jHk,1067
18
+ markdown_memory-0.1.0.dist-info/WHEEL,sha256=bEhYrD-rjlF0iRRHiAnfJ0mEjMsRwm29hhDD7yRgWCY,80
19
+ markdown_memory-0.1.0.dist-info/entry_points.txt,sha256=-cIIcoPPUA8jfDOjsNP8U2TraQN55LCeUGp1SHjpbsM,65
20
+ markdown_memory-0.1.0.dist-info/METADATA,sha256=kzAYuRzSIIx1xSSJT_5uW_oVeUWAsrYZJM4LdCtoKjY,33897
21
+ markdown_memory-0.1.0.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: uv 0.11.3
3
+ Root-Is-Purelib: true
4
+ Tag: py3-none-any
@@ -0,0 +1,3 @@
1
+ [console_scripts]
2
+ markdown-memory = markdown_memory.server:main
3
+
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 heshamkarm
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.