knowledge-rag 4.1.2__tar.gz → 4.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (21) hide show
  1. knowledge_rag-4.1.2/README.md → knowledge_rag-4.3.0/PKG-INFO +192 -17
  2. knowledge_rag-4.1.2/PKG-INFO → knowledge_rag-4.3.0/README.md +144 -57
  3. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/__init__.py +1 -1
  4. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/server.py +366 -91
  5. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/pyproject.toml +14 -4
  6. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/.gitignore +0 -0
  7. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/LICENSE +0 -0
  8. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/config.example.yaml +0 -0
  9. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/config.py +0 -0
  10. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/guarded.py +0 -0
  11. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/ingestion.py +0 -0
  12. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/instance_lock.py +0 -0
  13. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/metrics.py +0 -0
  14. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/preflight.py +0 -0
  15. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/mcp_server/ratelimit.py +0 -0
  16. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/npm/README.md +0 -0
  17. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/presets/cybersecurity.yaml +0 -0
  18. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/presets/developer.yaml +0 -0
  19. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/presets/general.yaml +0 -0
  20. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/presets/research.yaml +0 -0
  21. {knowledge_rag-4.1.2 → knowledge_rag-4.3.0}/requirements.txt +0 -0
@@ -1,3 +1,51 @@
1
+ Metadata-Version: 2.4
2
+ Name: knowledge-rag
3
+ Version: 4.3.0
4
+ Summary: Local RAG System for Claude Code — Hybrid search + Cross-encoder Reranking + 13 MCP Tools + 20 Format Parsers. Zero external servers.
5
+ Project-URL: Homepage, https://github.com/lyonzin/knowledge-rag
6
+ Project-URL: Repository, https://github.com/lyonzin/knowledge-rag
7
+ Project-URL: Issues, https://github.com/lyonzin/knowledge-rag/issues
8
+ Project-URL: Changelog, https://github.com/lyonzin/knowledge-rag/releases
9
+ Author-email: "Lyon." <lyonzin@users.noreply.github.com>
10
+ License: MIT
11
+ License-File: LICENSE
12
+ Keywords: bm25,chromadb,claude-code,embeddings,fastembed,hybrid-search,knowledge-base,local-ai,mcp,rag,reranking,retrieval-augmented-generation,semantic-search
13
+ Classifier: Development Status :: 4 - Beta
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: Intended Audience :: Science/Research
16
+ Classifier: License :: OSI Approved :: MIT License
17
+ Classifier: Operating System :: OS Independent
18
+ Classifier: Programming Language :: Python :: 3.11
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
21
+ Classifier: Topic :: Text Processing :: Indexing
22
+ Requires-Python: >=3.11
23
+ Requires-Dist: beautifulsoup4>=4.12.0
24
+ Requires-Dist: chromadb>=1.4.0
25
+ Requires-Dist: fastembed[reranking]>=0.4.0
26
+ Requires-Dist: mcp>=1.6.0
27
+ Requires-Dist: numpy>=1.24.0
28
+ Requires-Dist: openpyxl>=3.1.0
29
+ Requires-Dist: pymupdf>=1.23.0
30
+ Requires-Dist: python-docx>=1.0.0
31
+ Requires-Dist: python-pptx>=1.0.0
32
+ Requires-Dist: pyyaml>=6.0
33
+ Requires-Dist: requests>=2.33.0
34
+ Requires-Dist: watchdog>=4.0.0
35
+ Provides-Extra: gpu
36
+ Requires-Dist: nvidia-cublas-cu12; extra == 'gpu'
37
+ Requires-Dist: nvidia-cuda-runtime-cu12; extra == 'gpu'
38
+ Requires-Dist: nvidia-cudnn-cu12; extra == 'gpu'
39
+ Requires-Dist: nvidia-cufft-cu12; extra == 'gpu'
40
+ Requires-Dist: nvidia-curand-cu12; extra == 'gpu'
41
+ Requires-Dist: nvidia-cusolver-cu12; extra == 'gpu'
42
+ Requires-Dist: nvidia-cusparse-cu12; extra == 'gpu'
43
+ Requires-Dist: nvidia-nvjitlink-cu12; extra == 'gpu'
44
+ Requires-Dist: onnxruntime-gpu>=1.14.0; extra == 'gpu'
45
+ Provides-Extra: server
46
+ Requires-Dist: uvicorn>=0.20.0; extra == 'server'
47
+ Description-Content-Type: text/markdown
48
+
1
49
  # Knowledge RAG
2
50
 
3
51
  <div align="center">
@@ -17,7 +65,7 @@
17
65
  ### Your docs, your machine, zero cloud. Claude Code searches them natively.
18
66
 
19
67
  Drop your PDFs, markdown, code, notebooks — **1800+ files, 39K chunks, indexed in under 3 minutes.**<br/>
20
- Hybrid search (BM25 + semantic vectors + cross-encoder reranking) through 12 MCP tools.<br/>
68
+ Hybrid search (BM25 + semantic vectors + cross-encoder reranking) through 13 MCP tools.<br/>
21
69
  Everything runs locally via ONNX. No Docker, no Ollama, no API keys, no data leaves your machine.
22
70
 
23
71
  ```
@@ -26,9 +74,9 @@ pip install knowledge-rag → restart Claude Code → search_knowledge("your que
26
74
 
27
75
  ---
28
76
 
29
- **12 MCP Tools** | **Hybrid Search + Reranking** | **20 File Formats** | **Optional NVIDIA GPU** | **100% Local**
77
+ **13 MCP Tools** | **Hybrid Search + Reranking** | **20 File Formats** | **Optional NVIDIA GPU** | **100% Local**
30
78
 
31
- [What's New](#whats-new-in-v400) | [Supported Formats](#supported-formats) | [Installation](#installation) | [Configuration](#configuration) | [API Reference](#api-reference) | [Architecture](#architecture)
79
+ [What's New](#whats-new-in-v420) | [Supported Formats](#supported-formats) | [Installation](#installation) | [Configuration](#configuration) | [API Reference](#api-reference) | [Architecture](#architecture)
32
80
 
33
81
  </div>
34
82
 
@@ -50,7 +98,17 @@ pip install knowledge-rag → restart Claude Code → search_knowledge("your que
50
98
 
51
99
  ---
52
100
 
53
- ## What's New in v4.0.0
101
+ ## What's New in v4.2.0
102
+
103
+ ### Search Performance & Output Quality (v4.2.0)
104
+
105
+ **128× faster BM25 search** — replaced `rank-bm25` full-corpus scan with a custom **inverted-index** implementation. Only documents containing query terms are scored, using `numpy.argpartition` for O(n) top-k selection. Adjacent chunk fetching now uses a single batched ChromaDB call instead of N round-trips, and an O(1) reverse lookup (`_source_to_docid`) eliminates linear scans.
106
+
107
+ **Smarter output** — two new parameters on `search_knowledge`:
108
+ - **`snippet_mode`** (default: `true`) — truncates content to ~500 characters at natural break points, reducing token consumption by ~72%. Adds `content_length` field with original size; use `get_document()` for full content.
109
+ - **`min_score`** — filters results below a normalized relevance threshold (0.0–1.0). Eliminates low-quality noise from results. Response includes `filtered_by_score` count for transparency.
110
+
111
+ Both parameters are fully backwards-compatible (existing callers see no change in behavior).
54
112
 
55
113
  ### Enterprise Concurrent Access — SSE/HTTP Transport (v4.0.0)
56
114
 
@@ -71,7 +129,7 @@ Or via CLI: `knowledge-rag --transport sse`
71
129
  - **Prometheus metrics**: `/metrics` endpoint on separate port
72
130
  - **Bearer auth**: Token validation for SSE/HTTP connections
73
131
 
74
- All 12 MCP tools are instrumented with `@rate_limited` and `@instrument` decorators — zero overhead when features are disabled. Default transport remains **stdio** for full backwards compatibility.
132
+ All 13 MCP tools are instrumented with `@rate_limited` and `@instrument` decorators — zero overhead when features are disabled. Default transport remains **stdio** for full backwards compatibility.
75
133
 
76
134
  > **Migration**: Existing users need zero changes. SSE mode is opt-in via `server.transport: "sse"` in config.yaml. See [Configuration](#configuration) for details.
77
135
 
@@ -195,7 +253,7 @@ See [Changelog](#changelog) for full history.
195
253
  | **MMR Diversification** | Maximal Marginal Relevance reduces redundant results |
196
254
  | **Persistent Model Cache** | Embedding models cached in `models_cache/` — survives reboots |
197
255
  | **Auto-Migration** | Detects embedding dimension mismatch and rebuilds automatically |
198
- | **12 MCP Tools** | Full CRUD + search + evaluation via Claude Code |
256
+ | **13 MCP Tools** | Full CRUD + search + evaluation via Claude Code |
199
257
 
200
258
  ---
201
259
 
@@ -207,14 +265,14 @@ See [Changelog](#changelog) for full history.
207
265
  flowchart TB
208
266
  subgraph MCP["MCP SERVER (FastMCP)"]
209
267
  direction TB
210
- TOOLS["12 MCP Tools<br/>search | get | add | update | remove<br/>reindex | list | stats | url | similar | evaluate"]
268
+ TOOLS["13 MCP Tools<br/>search | get | add | update | remove<br/>reindex | reindex_status | list | stats | url | similar | evaluate"]
211
269
  end
212
270
 
213
271
  subgraph SEARCH["HYBRID SEARCH ENGINE"]
214
272
  direction LR
215
273
  ROUTER["Keyword Router<br/>(word boundaries)"]
216
274
  SEMANTIC["Semantic Search<br/>(ChromaDB)"]
217
- BM25["BM25 Keyword<br/>(rank-bm25 + expansion)"]
275
+ BM25["BM25 Keyword<br/>(inverted-index + expansion)"]
218
276
  RRF["Reciprocal Rank<br/>Fusion (RRF)"]
219
277
  RERANK["Cross-Encoder<br/>Reranker"]
220
278
 
@@ -278,15 +336,23 @@ flowchart TB
278
336
  subgraph HYBRID["Hybrid Search"]
279
337
  direction LR
280
338
  SEMANTIC["Semantic Search<br/>(ChromaDB embeddings)<br/>Conceptual similarity"]
281
- BM25["BM25 Search<br/>(expanded query)<br/>Exact term matching"]
339
+ BM25["BM25 Inverted-Index<br/>(posting lists + numpy top-k)<br/>Exact term matching"]
282
340
  end
283
341
 
284
342
  subgraph FUSION["Result Fusion + Reranking"]
285
343
  RRF["Reciprocal Rank Fusion<br/>score = alpha * 1/(k+rank_sem)<br/>+ (1-alpha) * 1/(k+rank_bm25)"]
286
344
  RERANK["Cross-Encoder Reranker<br/>Re-scores top 3x candidates<br/>query+doc pair scoring"]
287
345
  SORT["Sort by Reranker Score<br/>Normalize to 0-1"]
346
+ ADJ["Adjacent Chunk Expansion<br/>(batch fetch ±1 chunk)"]
288
347
 
289
- RRF --> RERANK --> SORT
348
+ RRF --> RERANK --> SORT --> ADJ
349
+ end
350
+
351
+ subgraph OUTPUT["Output Processing"]
352
+ MINSCORE["min_score Filter<br/>(discard below threshold)"]
353
+ SNIPPET["snippet_mode Truncation<br/>(~500 chars at natural break)"]
354
+
355
+ MINSCORE --> SNIPPET
290
356
  end
291
357
 
292
358
  CATEGORY --> HYBRID
@@ -294,7 +360,8 @@ flowchart TB
294
360
  SEMANTIC --> RRF
295
361
  BM25 --> RRF
296
362
 
297
- SORT --> RESULTS["Results<br/>search_method: hybrid|semantic|keyword<br/>score + reranker_score + raw_rrf_score"]
363
+ ADJ --> MINSCORE
364
+ SNIPPET --> RESULTS["Results<br/>search_method: hybrid|semantic|keyword<br/>score + filtered_by_score + content_length"]
298
365
  ```
299
366
 
300
367
  ### Document Ingestion Flow
@@ -373,7 +440,53 @@ flowchart LR
373
440
  - Claude Code CLI
374
441
  - *…or any other MCP client (Claude Desktop, Cursor, VS Code, Antigravity, opencode, Windsurf) — see [Use with other MCP clients](#use-with-other-mcp-clients)*
375
442
  - ~200MB disk for model cache (auto-downloaded on first run)
376
- - *Optional:* NVIDIA GPU + CUDA for accelerated embeddings (`pip install knowledge-rag[gpu]` + `models.embedding.gpu: true` in config)
443
+ - *Optional:* NVIDIA GPU + CUDA 12 for accelerated embeddings (see [GPU Acceleration](#gpu-acceleration) below)
444
+
445
+ ### GPU Acceleration
446
+
447
+ GPU mode accelerates embedding generation during indexing and search. It requires an NVIDIA GPU with CUDA 12 support. No GPU? No problem — the server runs on CPU by default and GPU is entirely optional.
448
+
449
+ **Requirements:**
450
+
451
+ | Component | Minimum | How to check / get it |
452
+ |-----------|---------|----------------------|
453
+ | NVIDIA GPU (Turing+) | RTX 20xx / 30xx / 40xx / 50xx, or Tesla T4+ | `nvidia-smi` |
454
+ | NVIDIA Driver | ≥ 525 | `nvidia-smi` — [nvidia.com/drivers](https://www.nvidia.com/drivers) |
455
+ | CUDA 12 runtime | Provided by pip packages below | Automatic |
456
+
457
+ **Setup (2 steps):**
458
+
459
+ ```bash
460
+ # 1. Install GPU dependencies (onnxruntime-gpu + all CUDA 12 runtime DLLs)
461
+ pip install knowledge-rag[gpu]
462
+
463
+ # 2. Enable in config.yaml
464
+ # models:
465
+ # embedding:
466
+ # gpu: true
467
+ ```
468
+
469
+ The `[gpu]` extra installs `onnxruntime-gpu` plus 7 NVIDIA CUDA 12 packages (`cublas`, `cudnn`, `cuda-runtime`, `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`) so you don't need a full CUDA Toolkit install.
470
+
471
+ **Verify GPU is active:**
472
+
473
+ On server startup, look for the GPU status banner:
474
+ ```
475
+ ============================================================
476
+ GPU STATUS: ACTIVE
477
+ Provider: CUDAExecutionProvider
478
+ Device: NVIDIA GeForce RTX 3080 Ti
479
+ VRAM: 12.0 GB
480
+ ============================================================
481
+ ```
482
+
483
+ Or programmatically:
484
+ ```bash
485
+ python -c "import onnxruntime; print(onnxruntime.get_available_providers())"
486
+ # Should include: 'CUDAExecutionProvider'
487
+ ```
488
+
489
+ > **Fallback**: If CUDA is unavailable at runtime (wrong driver, missing DLLs, no GPU), the server falls back to CPU automatically with a `[WARN]` log — it never crashes. The `gpu: true` config is a preference, not a requirement.
377
490
 
378
491
  ### Install Methods
379
492
 
@@ -644,7 +757,7 @@ search_knowledge("lateral movement strategies", hybrid_alpha=1.0)
644
757
 
645
758
  ### Indexing
646
759
 
647
- Documents are automatically indexed on first startup. To manage the index:
760
+ Documents are automatically indexed on first startup. All reindex operations run **in background** — they return immediately and you poll progress via `get_reindex_status()`:
648
761
 
649
762
  ```python
650
763
  # Incremental: only re-index changed files (fast)
@@ -655,6 +768,10 @@ reindex_documents(force=True)
655
768
 
656
769
  # Nuclear rebuild: delete everything, re-embed all (use after model change)
657
770
  reindex_documents(full_rebuild=True)
771
+
772
+ # Poll progress (lightweight, no full stats computation)
773
+ get_reindex_status()
774
+ # → {"reindex": {"active": true, "percent": 56, "progress": "2090/3734", ...}}
658
775
  ```
659
776
 
660
777
  ### Evaluating Retrieval Quality
@@ -683,6 +800,8 @@ Hybrid search combining semantic search + BM25 keyword search with cross-encoder
683
800
  | `max_results` | int | 5 | Maximum results to return (1-20) |
684
801
  | `category` | string | null | Filter by category |
685
802
  | `hybrid_alpha` | float | 0.3 | Balance: 0.0 = keyword only, 1.0 = semantic only |
803
+ | `min_score` | float | 0.0 | Minimum relevance score (0.0-1.0) to include a result. Use 0.2-0.4 to cut noise |
804
+ | `snippet_mode` | bool | true | Truncate content to ~500 chars at natural break points. Adds `content_length` field |
686
805
 
687
806
  **Returns:**
688
807
 
@@ -692,6 +811,7 @@ Hybrid search combining semantic search + BM25 keyword search with cross-encoder
692
811
  "query": "mimikatz credential dump",
693
812
  "hybrid_alpha": 0.5,
694
813
  "result_count": 3,
814
+ "filtered_by_score": 2,
695
815
  "cache_hit_rate": "0.0%",
696
816
  "results": [
697
817
  {
@@ -733,14 +853,39 @@ Retrieve the full content of a specific document.
733
853
 
734
854
  #### `reindex_documents`
735
855
 
736
- Index or reindex all documents in the knowledge base.
856
+ Index or reindex all documents in the knowledge base. **Runs in background** — returns immediately. Poll progress via `get_reindex_status()`.
737
857
 
738
858
  | Parameter | Type | Default | Description |
739
859
  |-----------|------|---------|-------------|
740
860
  | `force` | bool | false | Smart reindex: detects changes, rebuilds BM25. Fast. |
741
861
  | `full_rebuild` | bool | false | Nuclear rebuild: deletes everything, re-embeds all documents. Use after model change. |
742
862
 
743
- **Returns:** JSON with indexing statistics (indexed, updated, skipped, deleted, chunks_added, chunks_removed, dedup_skipped, elapsed_seconds).
863
+ **Returns:** `{"status": "started", "operation": "..."}` immediately. If already running, returns `{"status": "already_running", "progress": "1200/3734"}`.
864
+
865
+ ---
866
+
867
+ #### `get_reindex_status`
868
+
869
+ Get the current status of a background reindex operation. Lightweight — does not compute full index statistics.
870
+
871
+ **Returns (active):**
872
+ ```json
873
+ {
874
+ "status": "success",
875
+ "reindex": {
876
+ "active": true,
877
+ "operation": "nuclear_rebuild",
878
+ "progress": "1200/3734",
879
+ "percent": 32,
880
+ "indexed": 1200,
881
+ "skipped": 0,
882
+ "errors": 0,
883
+ "started_at": "2026-06-17T18:29:49"
884
+ }
885
+ }
886
+ ```
887
+
888
+ **Returns (idle):** `{"status": "success", "reindex": {"active": false}}`
744
889
 
745
890
  ---
746
891
 
@@ -1049,7 +1194,7 @@ For `.md` files, chunking splits at `##` and `###` header boundaries first. Sect
1049
1194
  |-------|---------|-------------|
1050
1195
  | `models.embedding.model` | `BAAI/bge-small-en-v1.5` | Embedding model (ONNX, runs locally) |
1051
1196
  | `models.embedding.dimensions` | 384 | Vector dimensions (must match model) |
1052
- | `models.embedding.gpu` | false | Enable CUDA GPU acceleration. Requires `pip install knowledge-rag[gpu]` |
1197
+ | `models.embedding.gpu` | false | Enable CUDA GPU acceleration. See [GPU Acceleration](#gpu-acceleration) for full setup |
1053
1198
  | `models.reranker.enabled` | true | Enable cross-encoder reranking |
1054
1199
  | `models.reranker.model` | `Xenova/ms-marco-MiniLM-L-6-v2` | Reranker model |
1055
1200
  | `models.reranker.top_k_multiplier` | 3 | Fetch N*multiplier candidates for reranking |
@@ -1309,7 +1454,7 @@ Common issues:
1309
1454
  - **NEW**: ChromaDB WAL mode enabled automatically in SSE/HTTP mode for concurrent read performance.
1310
1455
  - **NEW**: Optional rate limiting — sliding-window counter, configurable RPM and burst, disabled by default.
1311
1456
  - **NEW**: Optional Prometheus metrics endpoint — tool call counts, latency histograms, separate port, disabled by default.
1312
- - **NEW**: All 12 MCP tools instrumented with `@rate_limited` and `@instrument` decorators (zero-cost when disabled).
1457
+ - **NEW**: All 13 MCP tools instrumented with `@rate_limited` and `@instrument` decorators (zero-cost when disabled).
1313
1458
  - **NEW**: `--transport` CLI override for Docker/systemd deployments.
1314
1459
  - **NEW**: `pip install knowledge-rag[server]` optional dependency for SSE/HTTP (uvicorn).
1315
1460
  - **CHANGED**: SSE/HTTP mode auto-enables single-instance lock (port collision prevention).
@@ -1369,6 +1514,36 @@ Common issues:
1369
1514
 
1370
1515
  ### Unreleased
1371
1516
 
1517
+ ### v4.3.0 (2026-06-17) — Async Reindex, GPU CUDA 12, 13th MCP Tool
1518
+
1519
+ - **NEW**: `get_reindex_status` MCP tool — lightweight reindex progress polling without computing full index stats. Returns active/idle status, percent, processed/total, errors, and last result.
1520
+ - **NEW**: `reindex_documents` now runs in background via daemon thread — returns immediately with `{"status": "started"}`. Eliminates MCP timeout on large document sets (5K+ files). Concurrent calls return `already_running` with current progress.
1521
+ - **NEW**: GPU acceleration with full CUDA 12 support — `onnxruntime-gpu` + 7 NVIDIA pip packages (`cublas`, `cudnn`, `cuda-runtime`, `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`). Server auto-detects GPU on startup with 4-step verification (providers, DLLs, nvidia-smi, session creation). Falls back to CPU gracefully.
1522
+ - **NEW**: `_setup_cuda_dll_paths()` adds NVIDIA pip package DLL directories to `PATH` automatically on Windows — onnxruntime finds CUDA 12 DLLs without a full CUDA Toolkit install.
1523
+ - **DEPS**: `[gpu]` extra expanded from 3 to 8 packages (added `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`).
1524
+ - **FIX**: GPU status reporting now uses actual ONNX session creation test instead of just checking `get_available_providers()` — prevents false "GPU ACTIVE" when CUDA DLLs are missing.
1525
+ - **DOCS**: GPU Acceleration section rewritten with complete requirements table, setup steps, verification instructions, and fallback behavior.
1526
+ - **DOCS**: Tool reference updated — `reindex_documents` async behavior documented, `get_reindex_status` reference added.
1527
+ - **TEST**: Backwards-compat baseline updated for 13 MCP tools.
1528
+
1529
+ ### v4.2.0 (2026-06-17) — Search Performance & Output Quality
1530
+
1531
+ - **PERF**: Custom inverted-index BM25 replaces `rank-bm25` full-corpus scan — 128× faster keyword search on 50K+ chunk corpora. Only documents containing query terms are scored via posting lists.
1532
+ - **PERF**: `numpy.argpartition` for O(n) top-k selection instead of O(n log n) sort.
1533
+ - **PERF**: Batched adjacent chunk fetch — single ChromaDB `collection.get()` call replaces N round-trips per result.
1534
+ - **PERF**: O(1) reverse lookup via `_source_to_docid` dict eliminates linear scans of `_indexed_docs` in `search_similar`, `update_document`, `remove_document`, and `_expand_with_adjacent_chunks`.
1535
+ - **NEW**: `snippet_mode` parameter on `search_knowledge` (default: `true`) — truncates content to ~500 chars at natural break points with `content_length` field. Reduces token consumption by ~72%.
1536
+ - **NEW**: `min_score` parameter on `search_knowledge` (default: `0.0`) — filters results below a normalized relevance threshold. Response includes `filtered_by_score` count.
1537
+ - **NEW**: `filtered_by_score` field in search response JSON for transparency.
1538
+ - **DEPS**: `numpy` added as direct dependency (was transitive via fastembed); `rank-bm25` import removed from server.py.
1539
+ - **TEST**: 6 new tests for `min_score` filtering and `snippet_mode` truncation.
1540
+ - **TEST**: Updated backwards-compat baseline to include new `search_knowledge` parameters.
1541
+
1542
+ ### v4.1.2 (2026-06-17)
1543
+
1544
+ - **FIX**: `_save_metadata` dict snapshot prevents concurrent modification crash during file watcher events.
1545
+ - **STYLE**: ruff format applied to server.py.
1546
+
1372
1547
  ### v4.1.1 (2026-06-17)
1373
1548
 
1374
1549
  - **FIX**: All `_indexed_docs` iterations now use `list()` snapshot, preventing `dictionary changed size during iteration` crash when FileWatcher modifies the index concurrently with MCP tool calls (affects `search_knowledge`, `search_similar`, `update_document`, `remove_document`, `evaluate_retrieval`, `list_categories`, `list_documents`)