markdown-memory 0.1.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- markdown_memory/__init__.py +47 -0
- markdown_memory/autoindex.py +170 -0
- markdown_memory/config.py +235 -0
- markdown_memory/db.py +1546 -0
- markdown_memory/discovery.py +195 -0
- markdown_memory/embedders.py +513 -0
- markdown_memory/exceptions.py +71 -0
- markdown_memory/freshness.py +138 -0
- markdown_memory/headings.py +185 -0
- markdown_memory/indexer.py +725 -0
- markdown_memory/model_cache.py +272 -0
- markdown_memory/models.py +339 -0
- markdown_memory/parser.py +869 -0
- markdown_memory/py.typed +0 -0
- markdown_memory/search.py +518 -0
- markdown_memory/server.py +520 -0
- markdown_memory-0.1.0.dist-info/METADATA +579 -0
- markdown_memory-0.1.0.dist-info/RECORD +21 -0
- markdown_memory-0.1.0.dist-info/WHEEL +4 -0
- markdown_memory-0.1.0.dist-info/entry_points.txt +3 -0
- markdown_memory-0.1.0.dist-info/licenses/LICENSE +21 -0
|
@@ -0,0 +1,579 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: markdown-memory
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: MCP server that hands coding agents the one Markdown section that answers the question, not the whole file. Local hybrid search: BM25 keywords + ONNX embeddings in SQLite. No API keys, no network.
|
|
5
|
+
Keywords: mcp,mcp-server,model-context-protocol,coding-agent,ai-agents,markdown,documentation,rag,retrieval-augmented-generation,hybrid-search,semantic-search,vector-search,full-text-search,bm25,embeddings,onnx,sqlite,sqlite-vec,fts5
|
|
6
|
+
Author: Hesham Karm
|
|
7
|
+
Author-email: Hesham Karm <24391550+hishamkaram@users.noreply.github.com>
|
|
8
|
+
License-Expression: MIT
|
|
9
|
+
License-File: LICENSE
|
|
10
|
+
Classifier: Development Status :: 3 - Alpha
|
|
11
|
+
Classifier: Intended Audience :: Developers
|
|
12
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
16
|
+
Classifier: Operating System :: POSIX :: Linux
|
|
17
|
+
Classifier: Operating System :: MacOS
|
|
18
|
+
Classifier: Topic :: Software Development :: Documentation
|
|
19
|
+
Classifier: Topic :: Text Processing :: Indexing
|
|
20
|
+
Classifier: Typing :: Typed
|
|
21
|
+
Requires-Dist: fastembed>=0.8.0
|
|
22
|
+
Requires-Dist: huggingface-hub>=0.26
|
|
23
|
+
Requires-Dist: markdown-it-py>=4.2.0
|
|
24
|
+
Requires-Dist: mcp[cli]>=2.2.0,<3
|
|
25
|
+
Requires-Dist: numpy>=1.26
|
|
26
|
+
Requires-Dist: onnxruntime>=1.20
|
|
27
|
+
Requires-Dist: pydantic>=2.0
|
|
28
|
+
Requires-Dist: sqlite-vec>=0.1.9
|
|
29
|
+
Requires-Dist: tokenizers>=0.20
|
|
30
|
+
Requires-Python: >=3.11
|
|
31
|
+
Project-URL: Homepage, https://github.com/hishamkaram/markdown-memory
|
|
32
|
+
Project-URL: Repository, https://github.com/hishamkaram/markdown-memory
|
|
33
|
+
Project-URL: Issues, https://github.com/hishamkaram/markdown-memory/issues
|
|
34
|
+
Description-Content-Type: text/markdown
|
|
35
|
+
|
|
36
|
+
# markdown-memory
|
|
37
|
+
|
|
38
|
+
**Your coding agent reads whole Markdown files to answer one question. This returns the
|
|
39
|
+
section that answers it.**
|
|
40
|
+
|
|
41
|
+
Ask an agent a question about your docs and it opens the files that might answer it, whole.
|
|
42
|
+
Most of what lands in its context is about something else, and the part you wanted competes
|
|
43
|
+
with it. markdown-memory indexes your documentation by heading, so the same question comes
|
|
44
|
+
back as a few sections, each addressable by its breadcrumb and quoted verbatim.
|
|
45
|
+
|
|
46
|
+
<picture>
|
|
47
|
+
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/hishamkaram/markdown-memory/main/docs/assets/how-it-works-dark.svg">
|
|
48
|
+
<source srcset="https://raw.githubusercontent.com/hishamkaram/markdown-memory/main/docs/assets/how-it-works-light.svg">
|
|
49
|
+
<img src="https://raw.githubusercontent.com/hishamkaram/markdown-memory/main/docs/assets/how-it-works-light.png" width="100%"
|
|
50
|
+
alt="One question asked of four documentation files. Reading them whole costs 14,987
|
|
51
|
+
tokens. markdown-memory splits them at every heading, ranks by keywords and by
|
|
52
|
+
vectors, fuses the two, and returns five sections totalling 2,887 tokens - the
|
|
53
|
+
one that answers is 403.">
|
|
54
|
+
</picture>
|
|
55
|
+
|
|
56
|
+
Measured on this repository's own documentation - `README.md`, `CLAUDE.md`, `AGENTS.md` and
|
|
57
|
+
`docs/evaluation-protocol.md`, 14,987 tokens in all:
|
|
58
|
+
|
|
59
|
+
```
|
|
60
|
+
search_docs("where does the embedding model get downloaded")
|
|
61
|
+
|
|
62
|
+
407 tok README.md markdown-memory > The embedding model > What downloads, when, and where
|
|
63
|
+
712 tok README.md markdown-memory > The embedding model > Pre-download it, or install offline
|
|
64
|
+
647 tok README.md markdown-memory
|
|
65
|
+
738 tok CLAUDE.md markdown-memory > Commands
|
|
66
|
+
383 tok README.md markdown-memory > The embedding model > What is checked before the model is loaded
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
**2,887 tokens instead of 14,987**, and the section that actually answers is 403 - a
|
|
70
|
+
thirty-fourth of what reading the files costs. Every hit carries its full text, so a good
|
|
71
|
+
answer usually needs no follow-up call at all.
|
|
72
|
+
|
|
73
|
+
It is a local [Model Context Protocol](https://modelcontextprotocol.io) server - MCP is the
|
|
74
|
+
protocol agents use to call tools - and it runs entirely on your machine: parsing with
|
|
75
|
+
`markdown-it-py`, embeddings with
|
|
76
|
+
[EmbeddingGemma-300m](https://huggingface.co/onnx-community/embeddinggemma-300m-ONNX)
|
|
77
|
+
(4-bit ONNX on CPU, 768 dimensions), storage in SQLite - one database per documentation
|
|
78
|
+
root, with a keyword index and two vector indexes over it. No API key, no network after the
|
|
79
|
+
first model download, nothing leaves the machine.
|
|
80
|
+
|
|
81
|
+
## Install
|
|
82
|
+
|
|
83
|
+
### Prerequisites
|
|
84
|
+
|
|
85
|
+
- **Python 3.11 or newer.** The project develops on 3.12 and CI runs 3.11 through 3.14 on
|
|
86
|
+
x86-64 Linux.
|
|
87
|
+
- **[uv](https://docs.astral.sh/uv/)**, which manages the interpreter and the dependencies:
|
|
88
|
+
```bash
|
|
89
|
+
curl -LsSf https://astral.sh/uv/install.sh | sh # macOS / Linux
|
|
90
|
+
```
|
|
91
|
+
- **Roughly 1 GB of disk**: ~218 MB for the embedding model, the rest for the index.
|
|
92
|
+
- Linux or macOS, x86-64 or arm64. Everything runs on CPU; there is no GPU path and no API
|
|
93
|
+
key. CI covers arm64 Linux on every change and Apple Silicon on `main`, because the 4-bit
|
|
94
|
+
graph picks its `MatMulNBits` kernel from what the CPU offers rather than from the file -
|
|
95
|
+
see [What downloads, when, and where](#what-downloads-when-and-where). Those kernels do
|
|
96
|
+
not all return the same numbers: the gate measures four CPU families and the same text
|
|
97
|
+
embeds up to 1.6e-3 cosine apart between x86-64 and arm64. Copying an index between two
|
|
98
|
+
machines therefore searches vectors from one kernel with queries from another, which
|
|
99
|
+
nothing detects - and which was measured, on the labelled set and through the real
|
|
100
|
+
ranking, to move no top result and lose no answer. It reshuffles positions two to five.
|
|
101
|
+
|
|
102
|
+
### Get it
|
|
103
|
+
|
|
104
|
+
```bash
|
|
105
|
+
uv tool install markdown-memory # or: pipx install markdown-memory
|
|
106
|
+
markdown-memory --download-model # optional: fetch the ~218 MB model now
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
`uv tool install` puts the `markdown-memory` command in `~/.local/bin`; if your shell cannot
|
|
110
|
+
find it, `uv tool update-shell` adds that directory to your PATH. Fetching the model up
|
|
111
|
+
front is optional - the first search does it otherwise - but it keeps a 218 MB download
|
|
112
|
+
out of your first question.
|
|
113
|
+
|
|
114
|
+
### Register it with your client
|
|
115
|
+
|
|
116
|
+
One registration serves every project. The server indexes the project the client was
|
|
117
|
+
started in - Claude Code tells it (`CLAUDE_PROJECT_DIR`), Codex starts it there - and each
|
|
118
|
+
project gets its own index.
|
|
119
|
+
|
|
120
|
+
```bash
|
|
121
|
+
claude mcp add --scope user markdown-memory -- markdown-memory # Claude Code
|
|
122
|
+
codex mcp add markdown-memory -- markdown-memory # Codex
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
`claude mcp list` and `codex mcp list` should show it. Upgrade with
|
|
126
|
+
`uv tool upgrade markdown-memory` (or `pipx upgrade markdown-memory`); remove it with
|
|
127
|
+
`claude mcp remove --scope user markdown-memory`, `codex mcp remove markdown-memory` and
|
|
128
|
+
`uv tool uninstall markdown-memory` (or `pipx uninstall markdown-memory`).
|
|
129
|
+
|
|
130
|
+
Desktop apps (Claude Desktop, Cursor) usually do not see `~/.local/bin`, so give them the
|
|
131
|
+
absolute path that `command -v markdown-memory` prints. Claude Desktop has no project
|
|
132
|
+
either, so name the documentation root:
|
|
133
|
+
|
|
134
|
+
```json
|
|
135
|
+
{
|
|
136
|
+
"mcpServers": {
|
|
137
|
+
"markdown-memory": {
|
|
138
|
+
"command": "/home/you/.local/bin/markdown-memory",
|
|
139
|
+
"env": { "MARKDOWN_MEMORY_DOCS_DIR": "/absolute/path/to/your/docs" }
|
|
140
|
+
}
|
|
141
|
+
}
|
|
142
|
+
}
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
`uvx markdown-memory` works in place of the installed command too, with one catch: its
|
|
146
|
+
first launch downloads about 250 MB of dependencies before the server can answer, longer
|
|
147
|
+
than Codex waits by default - set `startup_timeout_sec = 60` under
|
|
148
|
+
`[mcp_servers.markdown-memory]` in `~/.codex/config.toml` if you go that way.
|
|
149
|
+
|
|
150
|
+
### From source
|
|
151
|
+
|
|
152
|
+
For working on markdown-memory itself:
|
|
153
|
+
|
|
154
|
+
```bash
|
|
155
|
+
git clone https://github.com/hishamkaram/markdown-memory
|
|
156
|
+
cd markdown-memory
|
|
157
|
+
uv sync
|
|
158
|
+
uv run markdown-memory --version
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
The `.mcp.json` in the clone runs that checkout (`uv run markdown-memory`) whenever Claude
|
|
162
|
+
Code is opened in it, with no edit.
|
|
163
|
+
|
|
164
|
+
### Index once, then search
|
|
165
|
+
|
|
166
|
+
The server keeps its documentation root indexed by itself. The first `search_docs` (or
|
|
167
|
+
`list_documents` without a directory) after it starts begins a catch-up run in the
|
|
168
|
+
background - whatever changed while no server ran, new files included - and answers from
|
|
169
|
+
the index as it stood, with `index_status.indexing` set. Nothing runs before that first
|
|
170
|
+
call: loading the model holds the interpreter for seconds, and a client waiting on the
|
|
171
|
+
handshake gives up quickly. After that, each of those calls decides whether another run is
|
|
172
|
+
due:
|
|
173
|
+
|
|
174
|
+
- a file it sees edited, deleted or unreadable starts one, at most every 10 seconds;
|
|
175
|
+
- a weights mismatch a search has just recorded starts one at once;
|
|
176
|
+
- otherwise one walks the tree every 5 minutes while the server is in use, which is what
|
|
177
|
+
finds a file nobody indexed yet.
|
|
178
|
+
|
|
179
|
+
One run at a time, in one thread, stopped between two documents when the server shuts down;
|
|
180
|
+
a stopped run leaves what a killed run leaves, and the next one resumes. While it runs,
|
|
181
|
+
`index_status.indexing` is `true` and the message says so instead of asking for
|
|
182
|
+
`index_directory`. Nothing watches the filesystem, so a server nobody is using does not
|
|
183
|
+
spend anything; `read_section` and `get_document_outline` do not start a run. The first
|
|
184
|
+
run of a large tree embeds everything and competes with queries for CPU until it is done;
|
|
185
|
+
`--no-auto-index` (or `MARKDOWN_MEMORY_AUTO_INDEX=0`) turns all of this off, and then:
|
|
186
|
+
|
|
187
|
+
```
|
|
188
|
+
index_directory() # scans the docs root, embeds what it finds
|
|
189
|
+
list_documents() # index_status.coverage should read "verified"
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
The first run downloads the embedding model and then indexes, so it is the slow one - see
|
|
193
|
+
[The embedding model](#the-embedding-model). After that, indexing is incremental: re-running
|
|
194
|
+
it costs milliseconds when nothing changed. If `coverage` reads `"unknown"`, something was
|
|
195
|
+
missed and `index_status.message` says what to run.
|
|
196
|
+
|
|
197
|
+
`index_status` also carries `changed_files`: indexed documents a cheap probe could not
|
|
198
|
+
confirm are still what was indexed - their bytes differ, or they are gone, unreadable, or no
|
|
199
|
+
longer a regular file. It is best-effort in both directions. Unreadability is noticed only
|
|
200
|
+
where a moved timestamp made it read the file at all: a file whose permissions changed and
|
|
201
|
+
whose time did not is answered from the time, and it is the next `index_directory` that
|
|
202
|
+
records the failure and takes `coverage` to `"unknown"`. A path that cannot even be
|
|
203
|
+
`stat`-ed - its parent directory lost its permissions, say - has no timestamp to compare
|
|
204
|
+
and is counted straight away. It counts only rows the index
|
|
205
|
+
holds, so a file nobody has indexed yet is not among them: finding those means walking the
|
|
206
|
+
tree, which is the expensive half of indexing and not something a search should pay for.
|
|
207
|
+
And it reads bytes only where the modification time moved, so an edit that restores a file's
|
|
208
|
+
own timestamp is missed - indexing is not fooled by that, since it hashes every file it
|
|
209
|
+
walks; what is missed is only the hint that running it is worth it. A zero means nothing was
|
|
210
|
+
detected, not that every file was re-hashed. `coverage` stays `"verified"` while the count is
|
|
211
|
+
non-zero: the walk really did finish and really did read every file it found. What moved on
|
|
212
|
+
is the tree, and the message says so.
|
|
213
|
+
|
|
214
|
+
## The embedding model
|
|
215
|
+
|
|
216
|
+
### What downloads, when, and where
|
|
217
|
+
|
|
218
|
+
The first search (or `markdown-memory --download-model`) fetches three files from
|
|
219
|
+
[`onnx-community/embeddinggemma-300m-ONNX`](https://huggingface.co/onnx-community/embeddinggemma-300m-ONNX)
|
|
220
|
+
at a pinned revision: the 4-bit ONNX graph, its external weights, and the tokenizer.
|
|
221
|
+
About **218 MB**, once per machine, into:
|
|
222
|
+
|
|
223
|
+
```
|
|
224
|
+
$XDG_CACHE_HOME/markdown-memory/models/embeddinggemma-300m-onnx-5090578d9565/ # ~/.cache/... by default
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
The revision is part of the folder name, so moving the pin fetches the new weights instead
|
|
228
|
+
of serving the old ones under a name that claims to be the new ones. One revision can publish
|
|
229
|
+
several graphs, though, so the folder name is not the whole answer: a `.verified` stamp that
|
|
230
|
+
does not name exactly the files this version needs - a cache left behind by the int8 graph at
|
|
231
|
+
this same revision, say - is rejected and the files are fetched, rather than half-trusted.
|
|
232
|
+
|
|
233
|
+
The cache is shared by every project on purpose - the weights are identical and read-only,
|
|
234
|
+
so copying them per project would be pure waste. Point `MARKDOWN_MEMORY_MODEL_CACHE`
|
|
235
|
+
somewhere else to move it.
|
|
236
|
+
|
|
237
|
+
The graph is `onnx/model_q4.onnx`, published 4-bit: the 262144x768 vocabulary table is
|
|
238
|
+
quantized and gathered by a single `GatherBlockQuantized`, and the projections run as
|
|
239
|
+
`MatMulNBits`, so the table is never expanded to float32. Nothing is derived or rewritten on
|
|
240
|
+
your machine - earlier versions ran the int8 graph and patched it here to get the same
|
|
241
|
+
effect. Using Gemma is covered by the
|
|
242
|
+
[Gemma Terms of Use](https://ai.google.dev/gemma/terms); this repository distributes no
|
|
243
|
+
weights.
|
|
244
|
+
|
|
245
|
+
### What is checked before the model is loaded
|
|
246
|
+
|
|
247
|
+
Every file's size and sha256 is pinned at the pinned revision. After a full check, a
|
|
248
|
+
`.verified` stamp records each file's size, mtime, ctime, inode and device, so an ordinary
|
|
249
|
+
start is a handful of `stat` calls rather than most of a second of hashing. Anything that
|
|
250
|
+
differs sends the files back to be hashed against the pins, and a file that does not match
|
|
251
|
+
is re-fetched - just that file.
|
|
252
|
+
|
|
253
|
+
**Guaranteed:** any change to a model file's contents or metadata since it was verified is
|
|
254
|
+
caught. `cp -p`, `tar x` and `rsync --inplace` can overwrite a file and restore its mtime,
|
|
255
|
+
which is why ctime is in the stamp - nothing in user space can set that back. **Not
|
|
256
|
+
guaranteed:** silent disk bit-rot, with no write at all.
|
|
257
|
+
|
|
258
|
+
A failure to *load* verified files is not treated as damage: it means onnxruntime,
|
|
259
|
+
permissions or memory, so the error is raised as it stands and nothing is downloaded.
|
|
260
|
+
|
|
261
|
+
Verification and loading hold a shared `flock`; downloading and repairing hold it
|
|
262
|
+
exclusively, so several servers starting at once download once between them. A model cache
|
|
263
|
+
on **NFS or SMB shared between machines is not supported** - `flock` can be local-only
|
|
264
|
+
there.
|
|
265
|
+
|
|
266
|
+
Nothing is ever deleted to reclaim space. After a successful start, one log line on stderr
|
|
267
|
+
names any other `embeddinggemma-300m-onnx*` folders and what they cost, and any file sitting
|
|
268
|
+
in the current folder that this version does not use - the int8 graph an upgrade left behind
|
|
269
|
+
weighs about 310 MB - and removing them is yours to do.
|
|
270
|
+
|
|
271
|
+
### Pre-download it, or install offline
|
|
272
|
+
|
|
273
|
+
To fetch the model deliberately rather than on the first query:
|
|
274
|
+
|
|
275
|
+
```bash
|
|
276
|
+
markdown-memory --download-model # the default, EmbeddingGemma
|
|
277
|
+
markdown-memory --download-model --embedder bge-small # or the light preset
|
|
278
|
+
```
|
|
279
|
+
|
|
280
|
+
It downloads and loads the configured model, then exits; it touches no index.
|
|
281
|
+
|
|
282
|
+
For a machine with no network, copy the three files into
|
|
283
|
+
`$XDG_CACHE_HOME/markdown-memory/models/embeddinggemma-300m-onnx-5090578d9565/`, keeping
|
|
284
|
+
`onnx/model_q4.onnx`, `onnx/model_q4.onnx_data` and `tokenizer.json` where they are. The
|
|
285
|
+
first load hashes them once, writes the `.verified` stamp, and never touches the network.
|
|
286
|
+
|
|
287
|
+
`bge-small` is downloaded by `fastembed`, which pins no revision: if that cache is deleted,
|
|
288
|
+
it can come back with different weights under the same model name. EmbeddingGemma's
|
|
289
|
+
weights change on purpose, when a release moves its pin or graph. Either way, every document
|
|
290
|
+
records the weights that embedded it, and the index records the one revision that vouches
|
|
291
|
+
for all of its vectors:
|
|
292
|
+
|
|
293
|
+
- **Indexing repairs in place.** A document stamped by other weights is re-embedded like a
|
|
294
|
+
changed file, even when its bytes are the same. Before the first new vector is written the
|
|
295
|
+
index is marked as being re-embedded; nothing is deleted first, so keyword search answers
|
|
296
|
+
throughout, and a run that is killed resumes where it stopped, because each stamp is
|
|
297
|
+
written with its vectors. At the end of a run the revision is restored only once no
|
|
298
|
+
vector-bearing document in the whole database - every root's, and one indexed on its own
|
|
299
|
+
inside a pruned directory such as `.venv` - is stamped by anything else. Until then
|
|
300
|
+
`index_status.coverage` reads `"unknown"` and its message names the directories still to
|
|
301
|
+
be indexed.
|
|
302
|
+
- **Searching** compares for itself, after embedding the query, and falls back to keyword
|
|
303
|
+
ranking alone when the answer differs - for the old weights and the new alike while a
|
|
304
|
+
repair is under way. It does not wait to be told: weights can change while no Markdown
|
|
305
|
+
file does, and then there is no indexing run to notice. The search that notices records
|
|
306
|
+
it, and the next `index_directory` loads the model first and re-embeds what it finds.
|
|
307
|
+
|
|
308
|
+
A model that loads but cannot say which weights it is may not write into an index that
|
|
309
|
+
names its weights: its vectors could never be told apart from the ones already stored. An
|
|
310
|
+
index whose vectors no revision vouches for - built while the weights could not be read -
|
|
311
|
+
is not ranked against a query from weights that can name themselves - from the upgrade to
|
|
312
|
+
schema v6 on - and the first run with such weights re-embeds it. A server that loaded
|
|
313
|
+
`bge-small` while its revision could not be read keeps it unnamed until it restarts: the
|
|
314
|
+
revision is read beside the weights it loads, never after them.
|
|
315
|
+
|
|
316
|
+
### Presets
|
|
317
|
+
|
|
318
|
+
Two presets, chosen with `MARKDOWN_MEMORY_EMBEDDER`:
|
|
319
|
+
|
|
320
|
+
| Preset | Dimensions | Download | Peak RAM | Indexing | Held-out Top-1 / Top-3 / Top-5 |
|
|
321
|
+
| --- | --- | --- | --- | --- | --- |
|
|
322
|
+
| `embeddinggemma` (default) | 768 | ~218 MB | ~0.65 GB | ~7 vectors/s | 88% / 97% / 97% |
|
|
323
|
+
| `bge-small` | 384 | ~65 MB | ~1.1 GB | ~12 vectors/s | 68% / 82% / 88% |
|
|
324
|
+
|
|
325
|
+
Accuracy is the frozen baseline in `scripts/eval_data/baseline.json`, recorded by
|
|
326
|
+
`scripts/eval_retrieval.py` over a 54-section corpus and the 34 **held-out** paraphrase
|
|
327
|
+
queries, which were written before any parameter was tuned. The **dev** set - the one
|
|
328
|
+
tuning is allowed to look at, and deliberately harder - scores 71% / 85% / 94% with
|
|
329
|
+
EmbeddingGemma and 53% / 68% / 79% with bge-small. Exact identifiers - flags, environment
|
|
330
|
+
variables, error strings - are 100% Top-1 with either preset, because FTS5 answers them.
|
|
331
|
+
Query latency is not in the table on purpose: it swings by 2-3x with what else the machine
|
|
332
|
+
is doing, so the baseline records it as informational and so should you.
|
|
333
|
+
|
|
334
|
+
Switching preset **discards the whole index**: the two produce vectors of different sizes,
|
|
335
|
+
which cannot be compared, so every documentation root has to be indexed again.
|
|
336
|
+
`index_directory` reports that when it happens.
|
|
337
|
+
|
|
338
|
+
**The first index of a large documentation set is slow.** Every paragraph, list item, table
|
|
339
|
+
row and code block costs one vector; a section costs none of its own, because its vector is
|
|
340
|
+
pooled from its passages. Measured on a 97-file set: 1,682 sections and 8,453 passages -
|
|
341
|
+
10,135 vectors - indexed in **24 minutes** at 7.1 vectors per second, peaking at 667 MB,
|
|
342
|
+
after which a query over those 1,682 sections takes about 200 ms. Budget for it, and run it
|
|
343
|
+
once: indexing is incremental by SHA-256, so a
|
|
344
|
+
re-index that finds nothing changed takes milliseconds (8 ms for 36 sections) and only
|
|
345
|
+
edited files are re-embedded. `bge-small` indexes several times faster at a real cost in
|
|
346
|
+
accuracy.
|
|
347
|
+
|
|
348
|
+
EmbeddingGemma is distributed under the [Gemma Terms of Use](https://ai.google.dev/gemma/terms);
|
|
349
|
+
the revision is pinned. See [License](#license) for what that means for you.
|
|
350
|
+
|
|
351
|
+
### When it goes wrong
|
|
352
|
+
|
|
353
|
+
- **The server starts even when the model cannot load.** Nothing loads it until the
|
|
354
|
+
first search, so the failure surfaces there, as
|
|
355
|
+
`Cannot load embedding model onnx-community/embeddinggemma-300m-ONNX: ...`. So "the server
|
|
356
|
+
is running" is not evidence the model is there; `markdown-memory --download-model` is.
|
|
357
|
+
- **The first `search_docs` can block for the length of a 218 MB download.** Pre-download it
|
|
358
|
+
(above) if that matters.
|
|
359
|
+
- **All logging goes to stderr.** stdout carries JSON-RPC frames only, so a client that
|
|
360
|
+
shows you "the output" may be showing you nothing. Set `MARKDOWN_MEMORY_LOG_LEVEL=DEBUG`
|
|
361
|
+
and read stderr.
|
|
362
|
+
- **Switching preset discards the index.** The two models produce vectors of different
|
|
363
|
+
sizes, which cannot be compared, so every root must be indexed again. `index_directory`
|
|
364
|
+
says so when it happens.
|
|
365
|
+
- **Searches come back empty or stale.** Run `index_directory` again; it is incremental, so
|
|
366
|
+
it is cheap. If `index_status.coverage` stays `"unknown"`, its `message` names the files
|
|
367
|
+
that could not be read.
|
|
368
|
+
|
|
369
|
+
## Tools
|
|
370
|
+
|
|
371
|
+
| Tool | Purpose |
|
|
372
|
+
| --- | --- |
|
|
373
|
+
| `index_directory(directory=None)` | Scan a tree, (re)index new/changed `.md` files (SHA-256), purge deleted ones |
|
|
374
|
+
| `list_documents(directory="")` | `{documents, index_status}`: indexed paths, titles and section counts, and whether a full index run vouches for them |
|
|
375
|
+
| `get_document_outline(file_path)` | Hierarchical TOC with line ranges and token estimates |
|
|
376
|
+
| `read_section(file_path, heading_path, include_subsections=False)` | Verbatim text of one section |
|
|
377
|
+
| `search_docs(query, limit=5)` | `{results, index_status}`: BM25 + passage-level vector search fused with Reciprocal Rank Fusion (k = 60); each hit reports the `matched_passage`, and `index_status` says whether the tree searched is known to be whole |
|
|
378
|
+
|
|
379
|
+
Sections are addressed by breadcrumb: `Root > Child > Subchild`. Oversized sections
|
|
380
|
+
(> ~800 tokens) are stored as `Root > Child (Part 1)`, `(Part 2)`, ...; reading the base
|
|
381
|
+
path reassembles them byte-for-byte. `file_path` may be absolute, relative to the docs
|
|
382
|
+
root, or any unique path suffix. `heading_path` is matched exactly first, then ignoring
|
|
383
|
+
spacing around `>`, then ignoring case, then as a trailing fragment (`Child > Subchild`
|
|
384
|
+
or just the title); an ambiguous request lists the exact candidates.
|
|
385
|
+
|
|
386
|
+
## Configuration
|
|
387
|
+
|
|
388
|
+
| Environment variable | CLI flag | Default |
|
|
389
|
+
| --- | --- | --- |
|
|
390
|
+
| `MARKDOWN_MEMORY_DOCS_DIR` | `--docs-dir` | `$CLAUDE_PROJECT_DIR` if the client exports it, else the working directory |
|
|
391
|
+
| `MARKDOWN_MEMORY_DB` | `--db` | `$XDG_DATA_HOME/markdown-memory/projects/<root>-<digest>/index.db` — one index per docs root |
|
|
392
|
+
| `MARKDOWN_MEMORY_MODEL_CACHE` | - | `$XDG_CACHE_HOME/markdown-memory/models` (`~/.cache/...`) |
|
|
393
|
+
| `MARKDOWN_MEMORY_EXCLUDE` | `--exclude` (repeatable) | nothing excluded |
|
|
394
|
+
| `MARKDOWN_MEMORY_LOG_LEVEL` | `--log-level` | `INFO` |
|
|
395
|
+
| `MARKDOWN_MEMORY_EMBEDDER` | `--embedder` | `embeddinggemma` (or `bge-small`) |
|
|
396
|
+
| `MARKDOWN_MEMORY_THREADS` | - | unset: onnxruntime picks. A positive integer caps the threads one embedding pass may use |
|
|
397
|
+
| `MARKDOWN_MEMORY_INDEX_WORKERS` | - | `2` - files read, parsed and embedded at the same time while indexing |
|
|
398
|
+
| `MARKDOWN_MEMORY_AUTO_INDEX` | `--no-auto-index` | on - `0`, `false`, `off` or `no` stops the server indexing its root by itself |
|
|
399
|
+
|
|
400
|
+
The default preset does not spin-wait between operators, which is what makes a query cost ~0.6 s of
|
|
401
|
+
CPU instead of ~5.5 s and leaves the process idle at 0 while nothing is being asked of it; the
|
|
402
|
+
thread *count* is left to onnxruntime, and `MARKDOWN_MEMORY_THREADS` is there for a machine that
|
|
403
|
+
disagrees with its choice. `bge-small` runs through `fastembed`, which exposes no such switch, so
|
|
404
|
+
for that preset the count is the only lever: `MARKDOWN_MEMORY_THREADS=4` took one query from 718 ms
|
|
405
|
+
of CPU to 95 ms.
|
|
406
|
+
|
|
407
|
+
Indexing embeds several files at once and writes them from one thread, in the order the
|
|
408
|
+
tree was walked. One ONNX session is shared, and its weights are mapped once however many
|
|
409
|
+
threads run against it, so each extra worker costs about the 150 MB of one forward pass.
|
|
410
|
+
Measured over 24 files of the eval corpus (879 passages): 196.3 s with one worker, 145.9 s
|
|
411
|
+
with two (1.35x) and 96.3 s with four (2.04x) - less than the embedding speed-up alone,
|
|
412
|
+
because parsing and the writes stay serial and a long file holds the head of the queue.
|
|
413
|
+
Two is the default because its peak measures around 0.8 GB (757-814 MB across runs),
|
|
414
|
+
well inside what a tool running beside an editor should take; raise `MARKDOWN_MEMORY_INDEX_WORKERS` on a machine with cores to spare.
|
|
415
|
+
|
|
416
|
+
`MARKDOWN_MEMORY_EXCLUDE` takes glob patterns separated by commas (only commas - a colon
|
|
417
|
+
would split a pattern that contains one). A pattern with no `/` matches that name at any
|
|
418
|
+
depth, the way `.gitignore` treats one: `eval_data` excludes `scripts/eval_data/corpus/`.
|
|
419
|
+
A pattern containing `/` is anchored at the documentation root (`tests/fixtures`,
|
|
420
|
+
`docs/generated/*`). Matching is case-sensitive everywhere, and a matching directory is
|
|
421
|
+
pruned, so its subtree costs nothing to skip. Files already indexed before an exclusion
|
|
422
|
+
was added are purged on the next index.
|
|
423
|
+
|
|
424
|
+
Without it, a repository that keeps fixtures, vendored documentation or a test corpus
|
|
425
|
+
in-tree indexes them as if they were its own docs.
|
|
426
|
+
|
|
427
|
+
Documents are stored under their absolute path, so one database *can* hold several
|
|
428
|
+
projects - but **search only ever answers from the root this server was started with**,
|
|
429
|
+
and `list_documents` shows only that root. Each project gets its own database by default,
|
|
430
|
+
so this matters only if you point two of them at one file with `MARKDOWN_MEMORY_DB`: the
|
|
431
|
+
second is then indexed, invisible, and paying for itself in disk.
|
|
432
|
+
|
|
433
|
+
### One index per project
|
|
434
|
+
|
|
435
|
+
The user-level registration above already gives every project its own index. To set
|
|
436
|
+
exclusions for one repository, drop a `.mcp.json` like this into it (it takes precedence
|
|
437
|
+
there):
|
|
438
|
+
Each project owns its index without being told to: the database is keyed on the
|
|
439
|
+
documentation root it serves, so a project gets its own exclusions and no chance of
|
|
440
|
+
another project's sections - or another project's documents, which stay resolvable by
|
|
441
|
+
path across any database they share - appearing in its results.
|
|
442
|
+
|
|
443
|
+
```json
|
|
444
|
+
{
|
|
445
|
+
"mcpServers": {
|
|
446
|
+
"markdown-memory": {
|
|
447
|
+
"command": "markdown-memory",
|
|
448
|
+
"env": {
|
|
449
|
+
"MARKDOWN_MEMORY_EXCLUDE": "vendor,third_party,tests/fixtures"
|
|
450
|
+
}
|
|
451
|
+
}
|
|
452
|
+
}
|
|
453
|
+
}
|
|
454
|
+
```
|
|
455
|
+
|
|
456
|
+
Any path you do add is written relative, on purpose. Claude Code expands only real
|
|
457
|
+
environment variables in `.mcp.json`: `${workspaceFolder}` is a VS Code idea, and even
|
|
458
|
+
`${CLAUDE_PROJECT_DIR}` is not set at expansion time - measured on Claude Code 2.1.278,
|
|
459
|
+
both produce a *"Missing environment variables"* warning and are passed through as literal
|
|
460
|
+
text, which would make the server index a directory named `${workspaceFolder}` and report
|
|
461
|
+
success over zero files. The server refuses such a value outright, and resolves a relative
|
|
462
|
+
path against `CLAUDE_PROJECT_DIR` (which Claude Code *does* export to the spawned server),
|
|
463
|
+
falling back to the working directory only when that is not set. The docs root defaults to
|
|
464
|
+
that same project root, so it needs no entry. Each worktree of a repository is its own
|
|
465
|
+
directory, so each gets its own index.
|
|
466
|
+
|
|
467
|
+
By default nothing is written into the repository. (A *relative* `MARKDOWN_MEMORY_DB`
|
|
468
|
+
is resolved against the project root and does land inside it - `.gitignore` covers
|
|
469
|
+
`.markdown-memory/` for that reason, and any other relative path you choose is yours to
|
|
470
|
+
ignore.)
|
|
471
|
+
The index lives under `$XDG_DATA_HOME/markdown-memory/projects/`, in a directory named for
|
|
472
|
+
the documentation root and a digest of its resolved path - out of reach of `git clean
|
|
473
|
+
-xdf`, writable when the checkout is not, and on local disk when the checkout is on a
|
|
474
|
+
network share, where SQLite's write-ahead log cannot take the locks it needs. Set
|
|
475
|
+
`MARKDOWN_MEMORY_DB` to override it; a relative value is resolved against the project
|
|
476
|
+
root. The index is a cache of the Markdown files and is rebuilt from them, so deleting it
|
|
477
|
+
costs only the time to index again.
|
|
478
|
+
|
|
479
|
+
## How search ranks
|
|
480
|
+
|
|
481
|
+
1. **Keywords (FTS5, BM25).** Stopwords are dropped, identifiers are kept verbatim. A hit
|
|
482
|
+
only counts if it covers at least half of the query's IDF mass or matches an
|
|
483
|
+
identifier-like term - a stray match on "data" or "deploy" no longer outvotes the
|
|
484
|
+
vector index.
|
|
485
|
+
2. **Vectors.** Every paragraph, list item, table row (rendered as `Header: cell; ...`) and
|
|
486
|
+
code block - also inside block quotes and list items - is embedded separately. The
|
|
487
|
+
section's own vector is the mean of those passage vectors, not a separate embedding of
|
|
488
|
+
the whole section: the model truncates at 512 tokens, which a long section exceeds. It
|
|
489
|
+
is an aggregate of the passages rather than independent evidence about the section, and
|
|
490
|
+
it finds no section that the passages do not. A passage longer than 600 characters is split into
|
|
491
|
+
consecutive windows at line, sentence or word boundaries, so a long command list or
|
|
492
|
+
configuration block keeps a vector for all of itself rather than for its first 600
|
|
493
|
+
characters. A table split across `(Part n)` sections keeps its header for every part. A
|
|
494
|
+
section is ranked by its closest vector, so one relevant table row is enough.
|
|
495
|
+
Heading-only sections have no vectors and are never returned ahead of their children.
|
|
496
|
+
3. **Reciprocal Rank Fusion** of the two rankings.
|
|
497
|
+
|
|
498
|
+
Cross-encoder rerankers (MiniLM, bge-reranker-base, jina, ColBERT) were benchmarked and
|
|
499
|
+
rejected: every one lowered accuracy on technical documentation and cost 2-12 s a query.
|
|
500
|
+
|
|
501
|
+
All logging goes to **stderr**. stdout carries JSON-RPC frames only.
|
|
502
|
+
|
|
503
|
+
## How malformed Markdown is handled
|
|
504
|
+
|
|
505
|
+
- **Skipped heading levels** (`#` then `####`): a heading stack pops every level `>= L`
|
|
506
|
+
before pushing, so breadcrumbs stay well formed.
|
|
507
|
+
- **Preamble**: badges/summary before the first heading become `[Overview / Preamble]`.
|
|
508
|
+
- **YAML front matter**: kept out of the AST (CommonMark would read it as a setext
|
|
509
|
+
heading) and used as a title fallback. A leading `---` rule followed by prose is not
|
|
510
|
+
mistaken for it. One case is inherently ambiguous - a single `key: value` line between
|
|
511
|
+
two `---` lines - and is read as front matter, as static-site generators do, unless the
|
|
512
|
+
key is an admonition word (`Note:`, `Warning:`, `TODO:` ...).
|
|
513
|
+
- **Unclosed code fences**: CommonMark runs them to EOF - or to the closing marker of a
|
|
514
|
+
*later* fence - swallowing the sections in between. The fence is closed before the next
|
|
515
|
+
blank-line-preceded ATX heading instead. A level-1 `# ...` line is treated as a comment
|
|
516
|
+
unless the fence language cannot have `#` comments (JSON, Go, ...). This is a heuristic:
|
|
517
|
+
a fence that merely *looks* closed is only cut on strong evidence (it contains another
|
|
518
|
+
opening fence with an info string, or the document ends inside a bare fence and the cut
|
|
519
|
+
makes the rest well formed), and never when the repair would lose a heading that was
|
|
520
|
+
already found. Repair work is capped per document, so a pathological file costs a
|
|
521
|
+
bounded number of extra parses rather than one per fence.
|
|
522
|
+
- **No headings / walls of text**: split on paragraph boundaries, then on lines, then on
|
|
523
|
+
whitespace. A fenced block is only cut when it exceeds the limit by itself.
|
|
524
|
+
- **Colliding breadcrumbs** get a ` [2]`, ` [3]` suffix (also against generated
|
|
525
|
+
`(Part n)` paths), so every stored path addresses exactly one section.
|
|
526
|
+
- **Headings like `Option<T>`** keep their type parameter; only lower-case formatting
|
|
527
|
+
tags (`<b>`, `<sub>`, `<a>`, ...) are stripped from titles.
|
|
528
|
+
|
|
529
|
+
## Indexing rules
|
|
530
|
+
|
|
531
|
+
- `.md` / `.markdown`, regular files only, at most 10 MB; symlinked directories are not
|
|
532
|
+
followed. `.git`, `node_modules`, virtualenvs and tool caches are pruned - index such a
|
|
533
|
+
tree by passing a directory *inside* it, and it is then left alone when an ancestor is
|
|
534
|
+
re-indexed.
|
|
535
|
+
- A document is purged only when the walk could have found it and did not. Files under a
|
|
536
|
+
directory that cannot be listed are kept and the directory is reported as an error.
|
|
537
|
+
|
|
538
|
+
## Development
|
|
539
|
+
|
|
540
|
+
```bash
|
|
541
|
+
uv run ruff check . && uv run ruff format --check .
|
|
542
|
+
uv run mypy --strict src/
|
|
543
|
+
uv run pytest -v # unit + integration (real ONNX model for semantic tests)
|
|
544
|
+
uv run python scripts/live_test.py # spawns the server, drives it over stdio JSON-RPC
|
|
545
|
+
uv run python scripts/eval_retrieval.py --show-misses # retrieval accuracy; fails on regression
|
|
546
|
+
uv run python scripts/eval_retrieval.py --rebuild # ... after discarding the cached index
|
|
547
|
+
uv run python scripts/reindex_docs.py DIR --force # forced re-index + integrity verification
|
|
548
|
+
scripts/check.sh # the whole pre-commit gate, fail-fast
|
|
549
|
+
```
|
|
550
|
+
|
|
551
|
+
### For AI coding agents
|
|
552
|
+
|
|
553
|
+
`CLAUDE.md` is the developer guide (commands, layout, architecture rules, the retrieval
|
|
554
|
+
regression policy). `AGENTS.md` and `.cursorrules` carry the short form for Codex and
|
|
555
|
+
Cursor. All three share one block of rules for *using* the MCP tools - search first, then
|
|
556
|
+
outline, then read a single section; never dump whole Markdown files - and a test keeps the
|
|
557
|
+
three copies identical and in step with the code. Project skills for Claude Code live in
|
|
558
|
+
`.claude/skills/`: `run-eval`, `reindex-docs`, `test-regression`.
|
|
559
|
+
|
|
560
|
+
Built on the MCP Python SDK 2.x, where `FastMCP` was renamed `MCPServer`.
|
|
561
|
+
|
|
562
|
+
## License
|
|
563
|
+
|
|
564
|
+
MIT - see [`LICENSE`](LICENSE). Two things in this repository are *not* covered by it,
|
|
565
|
+
because they are not ours to license:
|
|
566
|
+
|
|
567
|
+
- **The embedding models.** EmbeddingGemma-300m, the default, is distributed under the
|
|
568
|
+
[Gemma Terms of Use](https://ai.google.dev/gemma/terms), which are not an OSI-approved
|
|
569
|
+
open-source licence; the revision is pinned. The `bge-small` preset is two licences at
|
|
570
|
+
once: the `BAAI/bge-small-en-v1.5` weights are MIT, and `fastembed`, which loads them,
|
|
571
|
+
is Apache-2.0. Nothing is bundled - both are downloaded on first use - but if the Gemma
|
|
572
|
+
terms do not suit you, `MARKDOWN_MEMORY_EMBEDDER=bge-small` avoids them entirely.
|
|
573
|
+
- **The evaluation corpus.** `scripts/eval_data/corpus_v2/` is third-party documentation
|
|
574
|
+
vendored verbatim from five projects, pinned by commit, and used only to measure
|
|
575
|
+
retrieval accuracy. Each upstream keeps its own licence, and its licence and NOTICE
|
|
576
|
+
files travel with it in `scripts/eval_data/corpus_v2_licenses/`. The table in
|
|
577
|
+
[`scripts/eval_data/corpus_v2_LICENSES.md`](scripts/eval_data/corpus_v2_LICENSES.md)
|
|
578
|
+
says what came from where; both it and the manifest beside it are generated by
|
|
579
|
+
`scripts/fetch_eval_corpus.py`, so edit the script rather than the files.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
markdown_memory/__init__.py,sha256=-SGV6gYrwplcR-J1maJ4_8X7td8Tq-iyASnggOOx9wU,1125
|
|
2
|
+
markdown_memory/autoindex.py,sha256=mZnBuSoyFhgAbGKYGihHFgB_MeEWnvt46acUja_wIEM,7233
|
|
3
|
+
markdown_memory/config.py,sha256=USCCJ2NEXb6kaaJgi_UULHCFLZLtN5WG9r3bWK4B4Ac,9966
|
|
4
|
+
markdown_memory/db.py,sha256=4JHsxBa_tIIuvL4lW0F84hrv7wYdjd7ItaYxi2ck7Ek,72227
|
|
5
|
+
markdown_memory/discovery.py,sha256=NTPQLVTl5jQ2pXsbas5-SEbfNFO9Ii5pxmPuFcEpHhM,8515
|
|
6
|
+
markdown_memory/embedders.py,sha256=PlGwtd1xH7pelbEamE7JEum4R0qREJuW7fu4escj8gc,23141
|
|
7
|
+
markdown_memory/exceptions.py,sha256=OsVXFN198VETNo00-Oe72SdI_wMNFG_QDEchrP0qAWM,2482
|
|
8
|
+
markdown_memory/freshness.py,sha256=bchaOXp1--FELa2DRywvR3aejlboFLZ7-uU_DFu_TrU,7102
|
|
9
|
+
markdown_memory/headings.py,sha256=QlrmdkTpKaDBgnoGMRE0LU9PlKCl3IB_zdRbbEuo078,7059
|
|
10
|
+
markdown_memory/indexer.py,sha256=VS_m5lVX7F1n3UyodWkTQDCu6W0UAn6dx0z5yFN6vb4,36718
|
|
11
|
+
markdown_memory/model_cache.py,sha256=Txv6ajf2i3ttA9L6HRfGU-cXuRS2NmZ9owW1yjEn134,11186
|
|
12
|
+
markdown_memory/models.py,sha256=hA2BdSVTlfwsekp7HH9RdLLRjLg_A1IfqP94cCXBr4Q,12620
|
|
13
|
+
markdown_memory/parser.py,sha256=qWj23NFXBIoncAPaRRQFLKOdavchYpVN9t76FlavowM,37693
|
|
14
|
+
markdown_memory/py.typed,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
|
|
15
|
+
markdown_memory/search.py,sha256=bam9luBsojtB00YhBZLu2H5J9XxCYahzufiR0AZRzf0,22265
|
|
16
|
+
markdown_memory/server.py,sha256=vsXRyNCXAH77I43nYiCCkUJ4wbKjuv_6ybL8vNDbipk,22709
|
|
17
|
+
markdown_memory-0.1.0.dist-info/licenses/LICENSE,sha256=Ta4m9NqLCdjDz0C7vcR1k0EFYoHRt9XdKyh8ho12jHk,1067
|
|
18
|
+
markdown_memory-0.1.0.dist-info/WHEEL,sha256=bEhYrD-rjlF0iRRHiAnfJ0mEjMsRwm29hhDD7yRgWCY,80
|
|
19
|
+
markdown_memory-0.1.0.dist-info/entry_points.txt,sha256=-cIIcoPPUA8jfDOjsNP8U2TraQN55LCeUGp1SHjpbsM,65
|
|
20
|
+
markdown_memory-0.1.0.dist-info/METADATA,sha256=kzAYuRzSIIx1xSSJT_5uW_oVeUWAsrYZJM4LdCtoKjY,33897
|
|
21
|
+
markdown_memory-0.1.0.dist-info/RECORD,,
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 heshamkarm
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|