knowledge-rag 4.1.2__tar.gz → 4.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- knowledge_rag-4.1.2/README.md → knowledge_rag-4.3.0/PKG-INFO +192 -17
- knowledge_rag-4.1.2/PKG-INFO → knowledge_rag-4.3.0/README.md +144 -57
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/__init__.py +1 -1
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/server.py +366 -91
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/pyproject.toml +14 -4
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/.gitignore +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/LICENSE +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/config.example.yaml +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/config.py +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/guarded.py +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/ingestion.py +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/instance_lock.py +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/metrics.py +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/preflight.py +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/ratelimit.py +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/npm/README.md +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/presets/cybersecurity.yaml +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/presets/developer.yaml +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/presets/general.yaml +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/presets/research.yaml +0 -0
- {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/requirements.txt +0 -0
|
@@ -1,3 +1,51 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: knowledge-rag
|
|
3
|
+
Version: 4.3.0
|
|
4
|
+
Summary: Local RAG System for Claude Code — Hybrid search + Cross-encoder Reranking + 13 MCP Tools + 20 Format Parsers. Zero external servers.
|
|
5
|
+
Project-URL: Homepage, https://github.com/lyonzin/knowledge-rag
|
|
6
|
+
Project-URL: Repository, https://github.com/lyonzin/knowledge-rag
|
|
7
|
+
Project-URL: Issues, https://github.com/lyonzin/knowledge-rag/issues
|
|
8
|
+
Project-URL: Changelog, https://github.com/lyonzin/knowledge-rag/releases
|
|
9
|
+
Author-email: "Lyon." <lyonzin@users.noreply.github.com>
|
|
10
|
+
License: MIT
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Keywords: bm25,chromadb,claude-code,embeddings,fastembed,hybrid-search,knowledge-base,local-ai,mcp,rag,reranking,retrieval-augmented-generation,semantic-search
|
|
13
|
+
Classifier: Development Status :: 4 - Beta
|
|
14
|
+
Classifier: Intended Audience :: Developers
|
|
15
|
+
Classifier: Intended Audience :: Science/Research
|
|
16
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
17
|
+
Classifier: Operating System :: OS Independent
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
20
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
21
|
+
Classifier: Topic :: Text Processing :: Indexing
|
|
22
|
+
Requires-Python: >=3.11
|
|
23
|
+
Requires-Dist: beautifulsoup4>=4.12.0
|
|
24
|
+
Requires-Dist: chromadb>=1.4.0
|
|
25
|
+
Requires-Dist: fastembed[reranking]>=0.4.0
|
|
26
|
+
Requires-Dist: mcp>=1.6.0
|
|
27
|
+
Requires-Dist: numpy>=1.24.0
|
|
28
|
+
Requires-Dist: openpyxl>=3.1.0
|
|
29
|
+
Requires-Dist: pymupdf>=1.23.0
|
|
30
|
+
Requires-Dist: python-docx>=1.0.0
|
|
31
|
+
Requires-Dist: python-pptx>=1.0.0
|
|
32
|
+
Requires-Dist: pyyaml>=6.0
|
|
33
|
+
Requires-Dist: requests>=2.33.0
|
|
34
|
+
Requires-Dist: watchdog>=4.0.0
|
|
35
|
+
Provides-Extra: gpu
|
|
36
|
+
Requires-Dist: nvidia-cublas-cu12; extra == 'gpu'
|
|
37
|
+
Requires-Dist: nvidia-cuda-runtime-cu12; extra == 'gpu'
|
|
38
|
+
Requires-Dist: nvidia-cudnn-cu12; extra == 'gpu'
|
|
39
|
+
Requires-Dist: nvidia-cufft-cu12; extra == 'gpu'
|
|
40
|
+
Requires-Dist: nvidia-curand-cu12; extra == 'gpu'
|
|
41
|
+
Requires-Dist: nvidia-cusolver-cu12; extra == 'gpu'
|
|
42
|
+
Requires-Dist: nvidia-cusparse-cu12; extra == 'gpu'
|
|
43
|
+
Requires-Dist: nvidia-nvjitlink-cu12; extra == 'gpu'
|
|
44
|
+
Requires-Dist: onnxruntime-gpu>=1.14.0; extra == 'gpu'
|
|
45
|
+
Provides-Extra: server
|
|
46
|
+
Requires-Dist: uvicorn>=0.20.0; extra == 'server'
|
|
47
|
+
Description-Content-Type: text/markdown
|
|
48
|
+
|
|
1
49
|
# Knowledge RAG
|
|
2
50
|
|
|
3
51
|
<div align="center">
|
|
@@ -17,7 +65,7 @@
|
|
|
17
65
|
### Your docs, your machine, zero cloud. Claude Code searches them natively.
|
|
18
66
|
|
|
19
67
|
Drop your PDFs, markdown, code, notebooks — **1800+ files, 39K chunks, indexed in under 3 minutes.**<br/>
|
|
20
|
-
Hybrid search (BM25 + semantic vectors + cross-encoder reranking) through
|
|
68
|
+
Hybrid search (BM25 + semantic vectors + cross-encoder reranking) through 13 MCP tools.<br/>
|
|
21
69
|
Everything runs locally via ONNX. No Docker, no Ollama, no API keys, no data leaves your machine.
|
|
22
70
|
|
|
23
71
|
```
|
|
@@ -26,9 +74,9 @@ pip install knowledge-rag → restart Claude Code → search_knowledge("your que
|
|
|
26
74
|
|
|
27
75
|
---
|
|
28
76
|
|
|
29
|
-
**
|
|
77
|
+
**13 MCP Tools** | **Hybrid Search + Reranking** | **20 File Formats** | **Optional NVIDIA GPU** | **100% Local**
|
|
30
78
|
|
|
31
|
-
[What's New](#whats-new-in-
|
|
79
|
+
[What's New](#whats-new-in-v420) | [Supported Formats](#supported-formats) | [Installation](#installation) | [Configuration](#configuration) | [API Reference](#api-reference) | [Architecture](#architecture)
|
|
32
80
|
|
|
33
81
|
</div>
|
|
34
82
|
|
|
@@ -50,7 +98,17 @@ pip install knowledge-rag → restart Claude Code → search_knowledge("your que
|
|
|
50
98
|
|
|
51
99
|
---
|
|
52
100
|
|
|
53
|
-
## What's New in v4.
|
|
101
|
+
## What's New in v4.2.0
|
|
102
|
+
|
|
103
|
+
### Search Performance & Output Quality (v4.2.0)
|
|
104
|
+
|
|
105
|
+
**128× faster BM25 search** — replaced `rank-bm25` full-corpus scan with a custom **inverted-index** implementation. Only documents containing query terms are scored, using `numpy.argpartition` for O(n) top-k selection. Adjacent chunk fetching now uses a single batched ChromaDB call instead of N round-trips, and an O(1) reverse lookup (`_source_to_docid`) eliminates linear scans.
|
|
106
|
+
|
|
107
|
+
**Smarter output** — two new parameters on `search_knowledge`:
|
|
108
|
+
- **`snippet_mode`** (default: `true`) — truncates content to ~500 characters at natural break points, reducing token consumption by ~72%. Adds `content_length` field with original size; use `get_document()` for full content.
|
|
109
|
+
- **`min_score`** — filters results below a normalized relevance threshold (0.0–1.0). Eliminates low-quality noise from results. Response includes `filtered_by_score` count for transparency.
|
|
110
|
+
|
|
111
|
+
Both parameters are fully backwards-compatible (existing callers see no change in behavior).
|
|
54
112
|
|
|
55
113
|
### Enterprise Concurrent Access — SSE/HTTP Transport (v4.0.0)
|
|
56
114
|
|
|
@@ -71,7 +129,7 @@ Or via CLI: `knowledge-rag --transport sse`
|
|
|
71
129
|
- **Prometheus metrics**: `/metrics` endpoint on separate port
|
|
72
130
|
- **Bearer auth**: Token validation for SSE/HTTP connections
|
|
73
131
|
|
|
74
|
-
All
|
|
132
|
+
All 13 MCP tools are instrumented with `@rate_limited` and `@instrument` decorators — zero overhead when features are disabled. Default transport remains **stdio** for full backwards compatibility.
|
|
75
133
|
|
|
76
134
|
> **Migration**: Existing users need zero changes. SSE mode is opt-in via `server.transport: "sse"` in config.yaml. See [Configuration](#configuration) for details.
|
|
77
135
|
|
|
@@ -195,7 +253,7 @@ See [Changelog](#changelog) for full history.
|
|
|
195
253
|
| **MMR Diversification** | Maximal Marginal Relevance reduces redundant results |
|
|
196
254
|
| **Persistent Model Cache** | Embedding models cached in `models_cache/` — survives reboots |
|
|
197
255
|
| **Auto-Migration** | Detects embedding dimension mismatch and rebuilds automatically |
|
|
198
|
-
| **
|
|
256
|
+
| **13 MCP Tools** | Full CRUD + search + evaluation via Claude Code |
|
|
199
257
|
|
|
200
258
|
---
|
|
201
259
|
|
|
@@ -207,14 +265,14 @@ See [Changelog](#changelog) for full history.
|
|
|
207
265
|
flowchart TB
|
|
208
266
|
subgraph MCP["MCP SERVER (FastMCP)"]
|
|
209
267
|
direction TB
|
|
210
|
-
TOOLS["
|
|
268
|
+
TOOLS["13 MCP Tools<br/>search | get | add | update | remove<br/>reindex | reindex_status | list | stats | url | similar | evaluate"]
|
|
211
269
|
end
|
|
212
270
|
|
|
213
271
|
subgraph SEARCH["HYBRID SEARCH ENGINE"]
|
|
214
272
|
direction LR
|
|
215
273
|
ROUTER["Keyword Router<br/>(word boundaries)"]
|
|
216
274
|
SEMANTIC["Semantic Search<br/>(ChromaDB)"]
|
|
217
|
-
BM25["BM25 Keyword<br/>(
|
|
275
|
+
BM25["BM25 Keyword<br/>(inverted-index + expansion)"]
|
|
218
276
|
RRF["Reciprocal Rank<br/>Fusion (RRF)"]
|
|
219
277
|
RERANK["Cross-Encoder<br/>Reranker"]
|
|
220
278
|
|
|
@@ -278,15 +336,23 @@ flowchart TB
|
|
|
278
336
|
subgraph HYBRID["Hybrid Search"]
|
|
279
337
|
direction LR
|
|
280
338
|
SEMANTIC["Semantic Search<br/>(ChromaDB embeddings)<br/>Conceptual similarity"]
|
|
281
|
-
BM25["BM25
|
|
339
|
+
BM25["BM25 Inverted-Index<br/>(posting lists + numpy top-k)<br/>Exact term matching"]
|
|
282
340
|
end
|
|
283
341
|
|
|
284
342
|
subgraph FUSION["Result Fusion + Reranking"]
|
|
285
343
|
RRF["Reciprocal Rank Fusion<br/>score = alpha * 1/(k+rank_sem)<br/>+ (1-alpha) * 1/(k+rank_bm25)"]
|
|
286
344
|
RERANK["Cross-Encoder Reranker<br/>Re-scores top 3x candidates<br/>query+doc pair scoring"]
|
|
287
345
|
SORT["Sort by Reranker Score<br/>Normalize to 0-1"]
|
|
346
|
+
ADJ["Adjacent Chunk Expansion<br/>(batch fetch ±1 chunk)"]
|
|
288
347
|
|
|
289
|
-
RRF --> RERANK --> SORT
|
|
348
|
+
RRF --> RERANK --> SORT --> ADJ
|
|
349
|
+
end
|
|
350
|
+
|
|
351
|
+
subgraph OUTPUT["Output Processing"]
|
|
352
|
+
MINSCORE["min_score Filter<br/>(discard below threshold)"]
|
|
353
|
+
SNIPPET["snippet_mode Truncation<br/>(~500 chars at natural break)"]
|
|
354
|
+
|
|
355
|
+
MINSCORE --> SNIPPET
|
|
290
356
|
end
|
|
291
357
|
|
|
292
358
|
CATEGORY --> HYBRID
|
|
@@ -294,7 +360,8 @@ flowchart TB
|
|
|
294
360
|
SEMANTIC --> RRF
|
|
295
361
|
BM25 --> RRF
|
|
296
362
|
|
|
297
|
-
|
|
363
|
+
ADJ --> MINSCORE
|
|
364
|
+
SNIPPET --> RESULTS["Results<br/>search_method: hybrid|semantic|keyword<br/>score + filtered_by_score + content_length"]
|
|
298
365
|
```
|
|
299
366
|
|
|
300
367
|
### Document Ingestion Flow
|
|
@@ -373,7 +440,53 @@ flowchart LR
|
|
|
373
440
|
- Claude Code CLI
|
|
374
441
|
- *…or any other MCP client (Claude Desktop, Cursor, VS Code, Antigravity, opencode, Windsurf) — see [Use with other MCP clients](#use-with-other-mcp-clients)*
|
|
375
442
|
- ~200MB disk for model cache (auto-downloaded on first run)
|
|
376
|
-
- *Optional:* NVIDIA GPU + CUDA for accelerated embeddings (
|
|
443
|
+
- *Optional:* NVIDIA GPU + CUDA 12 for accelerated embeddings (see [GPU Acceleration](#gpu-acceleration) below)
|
|
444
|
+
|
|
445
|
+
### GPU Acceleration
|
|
446
|
+
|
|
447
|
+
GPU mode accelerates embedding generation during indexing and search. It requires an NVIDIA GPU with CUDA 12 support. No GPU? No problem — the server runs on CPU by default and GPU is entirely optional.
|
|
448
|
+
|
|
449
|
+
**Requirements:**
|
|
450
|
+
|
|
451
|
+
| Component | Minimum | How to check / get it |
|
|
452
|
+
|-----------|---------|----------------------|
|
|
453
|
+
| NVIDIA GPU (Turing+) | RTX 20xx / 30xx / 40xx / 50xx, or Tesla T4+ | `nvidia-smi` |
|
|
454
|
+
| NVIDIA Driver | ≥ 525 | `nvidia-smi` — [nvidia.com/drivers](https://www.nvidia.com/drivers) |
|
|
455
|
+
| CUDA 12 runtime | Provided by pip packages below | Automatic |
|
|
456
|
+
|
|
457
|
+
**Setup (2 steps):**
|
|
458
|
+
|
|
459
|
+
```bash
|
|
460
|
+
# 1. Install GPU dependencies (onnxruntime-gpu + all CUDA 12 runtime DLLs)
|
|
461
|
+
pip install knowledge-rag[gpu]
|
|
462
|
+
|
|
463
|
+
# 2. Enable in config.yaml
|
|
464
|
+
# models:
|
|
465
|
+
# embedding:
|
|
466
|
+
# gpu: true
|
|
467
|
+
```
|
|
468
|
+
|
|
469
|
+
The `[gpu]` extra installs `onnxruntime-gpu` plus 7 NVIDIA CUDA 12 packages (`cublas`, `cudnn`, `cuda-runtime`, `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`) so you don't need a full CUDA Toolkit install.
|
|
470
|
+
|
|
471
|
+
**Verify GPU is active:**
|
|
472
|
+
|
|
473
|
+
On server startup, look for the GPU status banner:
|
|
474
|
+
```
|
|
475
|
+
============================================================
|
|
476
|
+
GPU STATUS: ACTIVE
|
|
477
|
+
Provider: CUDAExecutionProvider
|
|
478
|
+
Device: NVIDIA GeForce RTX 3080 Ti
|
|
479
|
+
VRAM: 12.0 GB
|
|
480
|
+
============================================================
|
|
481
|
+
```
|
|
482
|
+
|
|
483
|
+
Or programmatically:
|
|
484
|
+
```bash
|
|
485
|
+
python -c "import onnxruntime; print(onnxruntime.get_available_providers())"
|
|
486
|
+
# Should include: 'CUDAExecutionProvider'
|
|
487
|
+
```
|
|
488
|
+
|
|
489
|
+
> **Fallback**: If CUDA is unavailable at runtime (wrong driver, missing DLLs, no GPU), the server falls back to CPU automatically with a `[WARN]` log — it never crashes. The `gpu: true` config is a preference, not a requirement.
|
|
377
490
|
|
|
378
491
|
### Install Methods
|
|
379
492
|
|
|
@@ -644,7 +757,7 @@ search_knowledge("lateral movement strategies", hybrid_alpha=1.0)
|
|
|
644
757
|
|
|
645
758
|
### Indexing
|
|
646
759
|
|
|
647
|
-
Documents are automatically indexed on first startup.
|
|
760
|
+
Documents are automatically indexed on first startup. All reindex operations run **in background** — they return immediately and you poll progress via `get_reindex_status()`:
|
|
648
761
|
|
|
649
762
|
```python
|
|
650
763
|
# Incremental: only re-index changed files (fast)
|
|
@@ -655,6 +768,10 @@ reindex_documents(force=True)
|
|
|
655
768
|
|
|
656
769
|
# Nuclear rebuild: delete everything, re-embed all (use after model change)
|
|
657
770
|
reindex_documents(full_rebuild=True)
|
|
771
|
+
|
|
772
|
+
# Poll progress (lightweight, no full stats computation)
|
|
773
|
+
get_reindex_status()
|
|
774
|
+
# → {"reindex": {"active": true, "percent": 56, "progress": "2090/3734", ...}}
|
|
658
775
|
```
|
|
659
776
|
|
|
660
777
|
### Evaluating Retrieval Quality
|
|
@@ -683,6 +800,8 @@ Hybrid search combining semantic search + BM25 keyword search with cross-encoder
|
|
|
683
800
|
| `max_results` | int | 5 | Maximum results to return (1-20) |
|
|
684
801
|
| `category` | string | null | Filter by category |
|
|
685
802
|
| `hybrid_alpha` | float | 0.3 | Balance: 0.0 = keyword only, 1.0 = semantic only |
|
|
803
|
+
| `min_score` | float | 0.0 | Minimum relevance score (0.0-1.0) to include a result. Use 0.2-0.4 to cut noise |
|
|
804
|
+
| `snippet_mode` | bool | true | Truncate content to ~500 chars at natural break points. Adds `content_length` field |
|
|
686
805
|
|
|
687
806
|
**Returns:**
|
|
688
807
|
|
|
@@ -692,6 +811,7 @@ Hybrid search combining semantic search + BM25 keyword search with cross-encoder
|
|
|
692
811
|
"query": "mimikatz credential dump",
|
|
693
812
|
"hybrid_alpha": 0.5,
|
|
694
813
|
"result_count": 3,
|
|
814
|
+
"filtered_by_score": 2,
|
|
695
815
|
"cache_hit_rate": "0.0%",
|
|
696
816
|
"results": [
|
|
697
817
|
{
|
|
@@ -733,14 +853,39 @@ Retrieve the full content of a specific document.
|
|
|
733
853
|
|
|
734
854
|
#### `reindex_documents`
|
|
735
855
|
|
|
736
|
-
Index or reindex all documents in the knowledge base.
|
|
856
|
+
Index or reindex all documents in the knowledge base. **Runs in background** — returns immediately. Poll progress via `get_reindex_status()`.
|
|
737
857
|
|
|
738
858
|
| Parameter | Type | Default | Description |
|
|
739
859
|
|-----------|------|---------|-------------|
|
|
740
860
|
| `force` | bool | false | Smart reindex: detects changes, rebuilds BM25. Fast. |
|
|
741
861
|
| `full_rebuild` | bool | false | Nuclear rebuild: deletes everything, re-embeds all documents. Use after model change. |
|
|
742
862
|
|
|
743
|
-
**Returns:**
|
|
863
|
+
**Returns:** `{"status": "started", "operation": "..."}` immediately. If already running, returns `{"status": "already_running", "progress": "1200/3734"}`.
|
|
864
|
+
|
|
865
|
+
---
|
|
866
|
+
|
|
867
|
+
#### `get_reindex_status`
|
|
868
|
+
|
|
869
|
+
Get the current status of a background reindex operation. Lightweight — does not compute full index statistics.
|
|
870
|
+
|
|
871
|
+
**Returns (active):**
|
|
872
|
+
```json
|
|
873
|
+
{
|
|
874
|
+
"status": "success",
|
|
875
|
+
"reindex": {
|
|
876
|
+
"active": true,
|
|
877
|
+
"operation": "nuclear_rebuild",
|
|
878
|
+
"progress": "1200/3734",
|
|
879
|
+
"percent": 32,
|
|
880
|
+
"indexed": 1200,
|
|
881
|
+
"skipped": 0,
|
|
882
|
+
"errors": 0,
|
|
883
|
+
"started_at": "2026-06-17T18:29:49"
|
|
884
|
+
}
|
|
885
|
+
}
|
|
886
|
+
```
|
|
887
|
+
|
|
888
|
+
**Returns (idle):** `{"status": "success", "reindex": {"active": false}}`
|
|
744
889
|
|
|
745
890
|
---
|
|
746
891
|
|
|
@@ -1049,7 +1194,7 @@ For `.md` files, chunking splits at `##` and `###` header boundaries first. Sect
|
|
|
1049
1194
|
|-------|---------|-------------|
|
|
1050
1195
|
| `models.embedding.model` | `BAAI/bge-small-en-v1.5` | Embedding model (ONNX, runs locally) |
|
|
1051
1196
|
| `models.embedding.dimensions` | 384 | Vector dimensions (must match model) |
|
|
1052
|
-
| `models.embedding.gpu` | false | Enable CUDA GPU acceleration.
|
|
1197
|
+
| `models.embedding.gpu` | false | Enable CUDA GPU acceleration. See [GPU Acceleration](#gpu-acceleration) for full setup |
|
|
1053
1198
|
| `models.reranker.enabled` | true | Enable cross-encoder reranking |
|
|
1054
1199
|
| `models.reranker.model` | `Xenova/ms-marco-MiniLM-L-6-v2` | Reranker model |
|
|
1055
1200
|
| `models.reranker.top_k_multiplier` | 3 | Fetch N*multiplier candidates for reranking |
|
|
@@ -1309,7 +1454,7 @@ Common issues:
|
|
|
1309
1454
|
- **NEW**: ChromaDB WAL mode enabled automatically in SSE/HTTP mode for concurrent read performance.
|
|
1310
1455
|
- **NEW**: Optional rate limiting — sliding-window counter, configurable RPM and burst, disabled by default.
|
|
1311
1456
|
- **NEW**: Optional Prometheus metrics endpoint — tool call counts, latency histograms, separate port, disabled by default.
|
|
1312
|
-
- **NEW**: All
|
|
1457
|
+
- **NEW**: All 13 MCP tools instrumented with `@rate_limited` and `@instrument` decorators (zero-cost when disabled).
|
|
1313
1458
|
- **NEW**: `--transport` CLI override for Docker/systemd deployments.
|
|
1314
1459
|
- **NEW**: `pip install knowledge-rag[server]` optional dependency for SSE/HTTP (uvicorn).
|
|
1315
1460
|
- **CHANGED**: SSE/HTTP mode auto-enables single-instance lock (port collision prevention).
|
|
@@ -1369,6 +1514,36 @@ Common issues:
|
|
|
1369
1514
|
|
|
1370
1515
|
### Unreleased
|
|
1371
1516
|
|
|
1517
|
+
### v4.3.0 (2026-06-17) — Async Reindex, GPU CUDA 12, 13th MCP Tool
|
|
1518
|
+
|
|
1519
|
+
- **NEW**: `get_reindex_status` MCP tool — lightweight reindex progress polling without computing full index stats. Returns active/idle status, percent, processed/total, errors, and last result.
|
|
1520
|
+
- **NEW**: `reindex_documents` now runs in background via daemon thread — returns immediately with `{"status": "started"}`. Eliminates MCP timeout on large document sets (5K+ files). Concurrent calls return `already_running` with current progress.
|
|
1521
|
+
- **NEW**: GPU acceleration with full CUDA 12 support — `onnxruntime-gpu` + 7 NVIDIA pip packages (`cublas`, `cudnn`, `cuda-runtime`, `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`). Server auto-detects GPU on startup with 4-step verification (providers, DLLs, nvidia-smi, session creation). Falls back to CPU gracefully.
|
|
1522
|
+
- **NEW**: `_setup_cuda_dll_paths()` adds NVIDIA pip package DLL directories to `PATH` automatically on Windows — onnxruntime finds CUDA 12 DLLs without a full CUDA Toolkit install.
|
|
1523
|
+
- **DEPS**: `[gpu]` extra expanded from 3 to 8 packages (added `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`).
|
|
1524
|
+
- **FIX**: GPU status reporting now uses actual ONNX session creation test instead of just checking `get_available_providers()` — prevents false "GPU ACTIVE" when CUDA DLLs are missing.
|
|
1525
|
+
- **DOCS**: GPU Acceleration section rewritten with complete requirements table, setup steps, verification instructions, and fallback behavior.
|
|
1526
|
+
- **DOCS**: Tool reference updated — `reindex_documents` async behavior documented, `get_reindex_status` reference added.
|
|
1527
|
+
- **TEST**: Backwards-compat baseline updated for 13 MCP tools.
|
|
1528
|
+
|
|
1529
|
+
### v4.2.0 (2026-06-17) — Search Performance & Output Quality
|
|
1530
|
+
|
|
1531
|
+
- **PERF**: Custom inverted-index BM25 replaces `rank-bm25` full-corpus scan — 128× faster keyword search on 50K+ chunk corpora. Only documents containing query terms are scored via posting lists.
|
|
1532
|
+
- **PERF**: `numpy.argpartition` for O(n) top-k selection instead of O(n log n) sort.
|
|
1533
|
+
- **PERF**: Batched adjacent chunk fetch — single ChromaDB `collection.get()` call replaces N round-trips per result.
|
|
1534
|
+
- **PERF**: O(1) reverse lookup via `_source_to_docid` dict eliminates linear scans of `_indexed_docs` in `search_similar`, `update_document`, `remove_document`, and `_expand_with_adjacent_chunks`.
|
|
1535
|
+
- **NEW**: `snippet_mode` parameter on `search_knowledge` (default: `true`) — truncates content to ~500 chars at natural break points with `content_length` field. Reduces token consumption by ~72%.
|
|
1536
|
+
- **NEW**: `min_score` parameter on `search_knowledge` (default: `0.0`) — filters results below a normalized relevance threshold. Response includes `filtered_by_score` count.
|
|
1537
|
+
- **NEW**: `filtered_by_score` field in search response JSON for transparency.
|
|
1538
|
+
- **DEPS**: `numpy` added as direct dependency (was transitive via fastembed); `rank-bm25` import removed from server.py.
|
|
1539
|
+
- **TEST**: 6 new tests for `min_score` filtering and `snippet_mode` truncation.
|
|
1540
|
+
- **TEST**: Updated backwards-compat baseline to include new `search_knowledge` parameters.
|
|
1541
|
+
|
|
1542
|
+
### v4.1.2 (2026-06-17)
|
|
1543
|
+
|
|
1544
|
+
- **FIX**: `_save_metadata` dict snapshot prevents concurrent modification crash during file watcher events.
|
|
1545
|
+
- **STYLE**: ruff format applied to server.py.
|
|
1546
|
+
|
|
1372
1547
|
### v4.1.1 (2026-06-17)
|
|
1373
1548
|
|
|
1374
1549
|
- **FIX**: All `_indexed_docs` iterations now use `list()` snapshot, preventing `dictionary changed size during iteration` crash when FileWatcher modifies the index concurrently with MCP tool calls (affects `search_knowledge`, `search_similar`, `update_document`, `remove_document`, `evaluate_retrieval`, `list_categories`, `list_documents`)
|