knowledge-rag 4.3.0__tar.gz → 4.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (21) hide show
  1. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/PKG-INFO +75 -53
  2. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/README.md +74 -52
  3. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/mcp_server/__init__.py +1 -1
  4. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/mcp_server/server.py +36 -5
  5. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/pyproject.toml +6 -2
  6. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/.gitignore +0 -0
  7. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/LICENSE +0 -0
  8. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/config.example.yaml +0 -0
  9. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/mcp_server/config.py +0 -0
  10. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/mcp_server/guarded.py +0 -0
  11. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/mcp_server/ingestion.py +0 -0
  12. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/mcp_server/instance_lock.py +0 -0
  13. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/mcp_server/metrics.py +0 -0
  14. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/mcp_server/preflight.py +0 -0
  15. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/mcp_server/ratelimit.py +0 -0
  16. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/npm/README.md +0 -0
  17. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/presets/cybersecurity.yaml +0 -0
  18. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/presets/developer.yaml +0 -0
  19. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/presets/general.yaml +0 -0
  20. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/presets/research.yaml +0 -0
  21. {knowledge_rag-4.3.0 → knowledge_rag-4.4.0}/requirements.txt +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: knowledge-rag
3
- Version: 4.3.0
3
+ Version: 4.4.0
4
4
  Summary: Local RAG System for Claude Code — Hybrid search + Cross-encoder Reranking + 13 MCP Tools + 20 Format Parsers. Zero external servers.
5
5
  Project-URL: Homepage, https://github.com/lyonzin/knowledge-rag
6
6
  Project-URL: Repository, https://github.com/lyonzin/knowledge-rag
@@ -1447,6 +1447,79 @@ Common issues:
1447
1447
 
1448
1448
  ## Changelog
1449
1449
 
1450
+ ### Unreleased
1451
+
1452
+ ### v4.4.0 (2026-07-06) — Cross-Platform Installer & Hybrid Search Category Filter
1453
+
1454
+ - **NEW**: Cross-platform, multi-LLM-client installer (`install.py`) driving both `install.sh` (Linux/macOS) and `install.ps1` (Windows) as thin wrappers. One codebase, one behavior across every OS.
1455
+ - **NEW**: Auto-detects and registers `knowledge-rag` in 8 LLM clients — Claude Code, Claude Desktop, Cursor, Windsurf, VS Code (Copilot Chat), Cline, Gemini CLI, Zed — writing to each tool's canonical config path with the correct JSON schema per client (VS Code uses `servers`, Zed uses `context_servers`, everyone else uses `mcpServers`).
1456
+ - **NEW**: `--for <clients>` / `--exclude <clients>` opt-in/opt-out selection, `--dry-run` preview, `--list-clients` registry inspection, `--pypi-version <ver>` pinning, `--skip-init` / `--skip-model` for fast reruns. See `python install.py --help`.
1457
+ - **FIX**: `install.ps1` no longer writes MCP config to `~/.claude/mcp.json` (a stale secondary path); it now targets `~/.claude.json` — the file Claude Code actually reads — via idempotent JSON merge that preserves every existing MCP server entry with an automatic `.knowledge-rag.bak` backup.
1458
+ - **FIX**: `install.ps1` gains PyPI mode (`pip install knowledge-rag`) and runs `mcp_server.server init` on install — feature parity with `install.sh`.
1459
+ - **FIX**: MCP server spec no longer uses the fragile `cmd /c cd /d ... && python ...` wrapper on Windows; it emits the standard `command` + `cwd` shape supported natively by every modern client.
1460
+ - **FIX**: `install.sh` guards against `sh install.sh` (bash-only features now emit a clear error instead of a cryptic syntax failure).
1461
+ - **FIX**: Both scripts now correctly advertise **13 MCP tools** (was outdated at 12; `get_reindex_status` shipped in v4.3.0).
1462
+ - **FIX**: Windows-side `install.ps1` prefers `winget install Python.Python.3.12 --scope user` (no admin), falls back to python.org 3.12.7 (was pinned to 3.12.0).
1463
+ - **FIX**: Hybrid search now applies the active category filter to BM25 results before RRF fusion, preventing keyword-only BM25 hits from other categories from leaking into filtered searches. Both leak paths are closed: BM25 candidates are metadata-filtered before RRF (with `top_k` widened to `max_results * 20` to compensate for post-filter drop), and the fallback fetch during fusion re-checks `category` before adding a chunk to `combined_scores`. Note: the `category_filter` guard now also applies to the keyword-routed category — `_route_by_keywords()`-inferred routing filters BM25 too, consistent with the semantic branch. Users with warm `query_cache` entries should restart the server to invalidate stale results. (#109, thanks @Hohlas)
1464
+ - **TEST**: New `tests/test_installer_no_data_loss.py` (22 tests) locks in the installer's zero-data-loss contract across all three JSON schemas (`mcpServers` / `servers` / `context_servers`): top-level keys preserved, sibling MCP servers byte-identical, `.knowledge-rag.bak` backup written before every mutation, idempotent second run, `--dry-run` writes nothing, atomic `os.replace` write. Baseline: 231 → 266.
1465
+ - **TEST**: New `TestHybridCategoryFilter::test_bm25_results_respect_category_filter` in `tests/test_search.py` — deterministic regression covering BM25-only hits with mixed categories. Baseline: 266 → 267.
1466
+
1467
+ ### v4.3.1 (2026-06-22) — Hybrid Search Fixes
1468
+
1469
+ - **FIX**: Accept `"general"` as a valid category in `search_knowledge`. The parser hardcodes `"general"` as the fallback in `_detect_category` (`ingestion.py`), but the validator only built `valid_categories` from `config.keyword_routes` + `config.category_mappings.values()` — so users who customized `config.yaml` and dropped the default `"general": "general"` mapping hit `Invalid category` even though the index contained `general` documents. Validator now always tolerates `"general"`. (#98, thanks @Hohlas)
1470
+ - **FIX**: Skip BM25-only search results when Chroma can no longer resolve the chunk ID. Stale BM25 indices (typically right after `remove_document` or in the window between async reindex and BM25 rebuild) returned hits whose `collection.get()` came back empty; the previous fallback inserted entries with `document=""` / `metadata={}` into the reranker, polluting results with empty matches. The pipeline now `continue`s past those, dropping the stale hit cleanly. (#98, thanks @Hohlas)
1471
+ - **TEST**: Added `tests/test_pr98_regression.py` (4 tests) pinning both contracts so future refactors cannot silently revert either fix. Test count baseline: 227 → 231. (#99)
1472
+ - **CI**: Bumped `[tool.mypy] python_version` from 3.11 to 3.12 to accept PEP 695 `type` statements in the numpy stub (`numpy/__init__.pyi`) which were breaking the Pillar 7 strict gate. Only affects static analysis; `requires-python = ">=3.11"` unchanged. (#100)
1473
+
1474
+ ### v4.3.0 (2026-06-17) — Async Reindex, GPU CUDA 12, 13th MCP Tool
1475
+
1476
+ - **NEW**: `get_reindex_status` MCP tool — lightweight reindex progress polling without computing full index stats. Returns active/idle status, percent, processed/total, errors, and last result.
1477
+ - **NEW**: `reindex_documents` now runs in background via daemon thread — returns immediately with `{"status": "started"}`. Eliminates MCP timeout on large document sets (5K+ files). Concurrent calls return `already_running` with current progress.
1478
+ - **NEW**: GPU acceleration with full CUDA 12 support — `onnxruntime-gpu` + 7 NVIDIA pip packages (`cublas`, `cudnn`, `cuda-runtime`, `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`). Server auto-detects GPU on startup with 4-step verification (providers, DLLs, nvidia-smi, session creation). Falls back to CPU gracefully.
1479
+ - **NEW**: `_setup_cuda_dll_paths()` adds NVIDIA pip package DLL directories to `PATH` automatically on Windows — onnxruntime finds CUDA 12 DLLs without a full CUDA Toolkit install.
1480
+ - **DEPS**: `[gpu]` extra expanded from 3 to 8 packages (added `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`).
1481
+ - **FIX**: GPU status reporting now uses actual ONNX session creation test instead of just checking `get_available_providers()` — prevents false "GPU ACTIVE" when CUDA DLLs are missing.
1482
+ - **DOCS**: GPU Acceleration section rewritten with complete requirements table, setup steps, verification instructions, and fallback behavior.
1483
+ - **DOCS**: Tool reference updated — `reindex_documents` async behavior documented, `get_reindex_status` reference added.
1484
+ - **TEST**: Backwards-compat baseline updated for 13 MCP tools.
1485
+
1486
+ ### v4.2.0 (2026-06-17) — Search Performance & Output Quality
1487
+
1488
+ - **PERF**: Custom inverted-index BM25 replaces `rank-bm25` full-corpus scan — 128× faster keyword search on 50K+ chunk corpora. Only documents containing query terms are scored via posting lists.
1489
+ - **PERF**: `numpy.argpartition` for O(n) top-k selection instead of O(n log n) sort.
1490
+ - **PERF**: Batched adjacent chunk fetch — single ChromaDB `collection.get()` call replaces N round-trips per result.
1491
+ - **PERF**: O(1) reverse lookup via `_source_to_docid` dict eliminates linear scans of `_indexed_docs` in `search_similar`, `update_document`, `remove_document`, and `_expand_with_adjacent_chunks`.
1492
+ - **NEW**: `snippet_mode` parameter on `search_knowledge` (default: `true`) — truncates content to ~500 chars at natural break points with `content_length` field. Reduces token consumption by ~72%.
1493
+ - **NEW**: `min_score` parameter on `search_knowledge` (default: `0.0`) — filters results below a normalized relevance threshold. Response includes `filtered_by_score` count.
1494
+ - **NEW**: `filtered_by_score` field in search response JSON for transparency.
1495
+ - **DEPS**: `numpy` added as direct dependency (was transitive via fastembed); `rank-bm25` import removed from server.py.
1496
+ - **TEST**: 6 new tests for `min_score` filtering and `snippet_mode` truncation.
1497
+ - **TEST**: Updated backwards-compat baseline to include new `search_knowledge` parameters.
1498
+
1499
+ ### v4.1.2 (2026-06-17)
1500
+
1501
+ - **FIX**: `_save_metadata` dict snapshot prevents concurrent modification crash during file watcher events.
1502
+ - **STYLE**: ruff format applied to server.py.
1503
+
1504
+ ### v4.1.1 (2026-06-17)
1505
+
1506
+ - **FIX**: All `_indexed_docs` iterations now use `list()` snapshot, preventing `dictionary changed size during iteration` crash when FileWatcher modifies the index concurrently with MCP tool calls (affects `search_knowledge`, `search_similar`, `update_document`, `remove_document`, `evaluate_retrieval`, `list_categories`, `list_documents`)
1507
+
1508
+ ### v4.1.0 (2026-06-17)
1509
+
1510
+ - **Added:** `query_expansion_groups` config for symmetric synonym expansion (#92)
1511
+ - **Improved:** `expand_query()` now returns deterministic expansion order (set → ordered list with dedup)
1512
+
1513
+ ### v4.0.1 (2026-06-16)
1514
+
1515
+ - **FIX**: Orphan cleanup now runs before indexing loop, preventing chunk loss when files are moved (#90).
1516
+ - **FIX**: Chunk deduplication is now per-document instead of global, preventing cross-document chunk deletion (#91).
1517
+ - **FIX**: Added `on_moved` handler to `DocumentWatcher` for proper file move detection.
1518
+ - **FIX**: Startup preflight probes ChromaDB in a child process and moves crashing persistent indexes to `data/backups/auto-repair-*` before MCP initialization.
1519
+ - **FIX**: Reranker load failures now fall back to RRF ordering instead of failing `search_knowledge` on offline machines.
1520
+ - **FIX**: Virtualenv project-root detection now handles Python symlinks that resolve to the system interpreter.
1521
+ - **NEW**: `knowledge-rag-guarded` console script kept as an explicit guarded startup alias.
1522
+
1450
1523
  ### v4.0.0 (2026-06-09) — Enterprise Concurrent Access
1451
1524
 
1452
1525
  - **NEW**: SSE and streamable-http transport modes — 1 server serves N clients (`server.transport: "sse"` in config.yaml or `--transport sse` CLI).
@@ -1482,7 +1555,7 @@ Common issues:
1482
1555
  - **NEW** Property-based fuzzing of all parsers via Hypothesis (`tests/test_ingestion_property.py`) — 200 random examples per CI run.
1483
1556
  - **NEW** Memory baseline regression tests (`tests/test_memory_baseline.py`, cross-platform via psutil) — RSS bounded under 1000 queries; nightly soak amplifies to 50K iterations.
1484
1557
  - **NEW** Property/locale/format/preset matrices (`tests/test_presets.py`, `tests/test_locale.py`, `tests/test_format_smoke.py`).
1485
- - **NEW** Backwards-compatibility regression tests (`tests/test_backwards_compat.py`) — legacy YAML configs from v3.6.0 / v3.7.0 still parse; all 12 MCP tool parameter names frozen.
1558
+ - **NEW** Backwards-compatibility regression tests (`tests/test_backwards_compat.py`) — legacy YAML configs from v3.6.0 / v3.7.0 still parse; all 13 MCP tool parameter names frozen.
1486
1559
  - **NEW** AST-based public API surface diff (`scripts/check_api_surface.py`) — any breaking change blocks merge, baseline at `.github/api-surface-baseline.json`.
1487
1560
  - **NEW** CHANGELOG enforcement (`scripts/check_changelog.py`) — user-facing PRs must add a bullet under `## Unreleased`; bypass via `skip-changelog` label.
1488
1561
  - **NEW** Test count anti-regression (`scripts/check_test_count.py`) — guards against silent test deletion.
@@ -1512,57 +1585,6 @@ Common issues:
1512
1585
  - **CHORE**: pytest `tmp_path_retention_count=1` to avoid Windows atexit cleanup race in CI.
1513
1586
  - **ROADMAP**: Tracked v4.0 shared-service architecture (one daemon, many thin MCP clients) as the long-term fix for multi-process resource duplication. (#34)
1514
1587
 
1515
- ### Unreleased
1516
-
1517
- ### v4.3.0 (2026-06-17) — Async Reindex, GPU CUDA 12, 13th MCP Tool
1518
-
1519
- - **NEW**: `get_reindex_status` MCP tool — lightweight reindex progress polling without computing full index stats. Returns active/idle status, percent, processed/total, errors, and last result.
1520
- - **NEW**: `reindex_documents` now runs in background via daemon thread — returns immediately with `{"status": "started"}`. Eliminates MCP timeout on large document sets (5K+ files). Concurrent calls return `already_running` with current progress.
1521
- - **NEW**: GPU acceleration with full CUDA 12 support — `onnxruntime-gpu` + 7 NVIDIA pip packages (`cublas`, `cudnn`, `cuda-runtime`, `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`). Server auto-detects GPU on startup with 4-step verification (providers, DLLs, nvidia-smi, session creation). Falls back to CPU gracefully.
1522
- - **NEW**: `_setup_cuda_dll_paths()` adds NVIDIA pip package DLL directories to `PATH` automatically on Windows — onnxruntime finds CUDA 12 DLLs without a full CUDA Toolkit install.
1523
- - **DEPS**: `[gpu]` extra expanded from 3 to 8 packages (added `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`).
1524
- - **FIX**: GPU status reporting now uses actual ONNX session creation test instead of just checking `get_available_providers()` — prevents false "GPU ACTIVE" when CUDA DLLs are missing.
1525
- - **DOCS**: GPU Acceleration section rewritten with complete requirements table, setup steps, verification instructions, and fallback behavior.
1526
- - **DOCS**: Tool reference updated — `reindex_documents` async behavior documented, `get_reindex_status` reference added.
1527
- - **TEST**: Backwards-compat baseline updated for 13 MCP tools.
1528
-
1529
- ### v4.2.0 (2026-06-17) — Search Performance & Output Quality
1530
-
1531
- - **PERF**: Custom inverted-index BM25 replaces `rank-bm25` full-corpus scan — 128× faster keyword search on 50K+ chunk corpora. Only documents containing query terms are scored via posting lists.
1532
- - **PERF**: `numpy.argpartition` for O(n) top-k selection instead of O(n log n) sort.
1533
- - **PERF**: Batched adjacent chunk fetch — single ChromaDB `collection.get()` call replaces N round-trips per result.
1534
- - **PERF**: O(1) reverse lookup via `_source_to_docid` dict eliminates linear scans of `_indexed_docs` in `search_similar`, `update_document`, `remove_document`, and `_expand_with_adjacent_chunks`.
1535
- - **NEW**: `snippet_mode` parameter on `search_knowledge` (default: `true`) — truncates content to ~500 chars at natural break points with `content_length` field. Reduces token consumption by ~72%.
1536
- - **NEW**: `min_score` parameter on `search_knowledge` (default: `0.0`) — filters results below a normalized relevance threshold. Response includes `filtered_by_score` count.
1537
- - **NEW**: `filtered_by_score` field in search response JSON for transparency.
1538
- - **DEPS**: `numpy` added as direct dependency (was transitive via fastembed); `rank-bm25` import removed from server.py.
1539
- - **TEST**: 6 new tests for `min_score` filtering and `snippet_mode` truncation.
1540
- - **TEST**: Updated backwards-compat baseline to include new `search_knowledge` parameters.
1541
-
1542
- ### v4.1.2 (2026-06-17)
1543
-
1544
- - **FIX**: `_save_metadata` dict snapshot prevents concurrent modification crash during file watcher events.
1545
- - **STYLE**: ruff format applied to server.py.
1546
-
1547
- ### v4.1.1 (2026-06-17)
1548
-
1549
- - **FIX**: All `_indexed_docs` iterations now use `list()` snapshot, preventing `dictionary changed size during iteration` crash when FileWatcher modifies the index concurrently with MCP tool calls (affects `search_knowledge`, `search_similar`, `update_document`, `remove_document`, `evaluate_retrieval`, `list_categories`, `list_documents`)
1550
-
1551
- ### v4.1.0 (2026-06-17)
1552
-
1553
- - **Added:** `query_expansion_groups` config for symmetric synonym expansion (#92)
1554
- - **Improved:** `expand_query()` now returns deterministic expansion order (set → ordered list with dedup)
1555
-
1556
- ### v4.0.1 (2026-06-16)
1557
-
1558
- - **FIX**: Orphan cleanup now runs before indexing loop, preventing chunk loss when files are moved (#90).
1559
- - **FIX**: Chunk deduplication is now per-document instead of global, preventing cross-document chunk deletion (#91).
1560
- - **FIX**: Added `on_moved` handler to `DocumentWatcher` for proper file move detection.
1561
- - **FIX**: Startup preflight probes ChromaDB in a child process and moves crashing persistent indexes to `data/backups/auto-repair-*` before MCP initialization.
1562
- - **FIX**: Reranker load failures now fall back to RRF ordering instead of failing `search_knowledge` on offline machines.
1563
- - **FIX**: Virtualenv project-root detection now handles Python symlinks that resolve to the system interpreter.
1564
- - **NEW**: `knowledge-rag-guarded` console script kept as an explicit guarded startup alias.
1565
-
1566
1588
  ### v3.6.2 (2026-04-23)
1567
1589
 
1568
1590
  - **INFRA**: NPM provenance attestation (SLSA supply chain security), full README on npm page
@@ -1399,6 +1399,79 @@ Common issues:
1399
1399
 
1400
1400
  ## Changelog
1401
1401
 
1402
+ ### Unreleased
1403
+
1404
+ ### v4.4.0 (2026-07-06) — Cross-Platform Installer & Hybrid Search Category Filter
1405
+
1406
+ - **NEW**: Cross-platform, multi-LLM-client installer (`install.py`) driving both `install.sh` (Linux/macOS) and `install.ps1` (Windows) as thin wrappers. One codebase, one behavior across every OS.
1407
+ - **NEW**: Auto-detects and registers `knowledge-rag` in 8 LLM clients — Claude Code, Claude Desktop, Cursor, Windsurf, VS Code (Copilot Chat), Cline, Gemini CLI, Zed — writing to each tool's canonical config path with the correct JSON schema per client (VS Code uses `servers`, Zed uses `context_servers`, everyone else uses `mcpServers`).
1408
+ - **NEW**: `--for <clients>` / `--exclude <clients>` opt-in/opt-out selection, `--dry-run` preview, `--list-clients` registry inspection, `--pypi-version <ver>` pinning, `--skip-init` / `--skip-model` for fast reruns. See `python install.py --help`.
1409
+ - **FIX**: `install.ps1` no longer writes MCP config to `~/.claude/mcp.json` (a stale secondary path); it now targets `~/.claude.json` — the file Claude Code actually reads — via idempotent JSON merge that preserves every existing MCP server entry with an automatic `.knowledge-rag.bak` backup.
1410
+ - **FIX**: `install.ps1` gains PyPI mode (`pip install knowledge-rag`) and runs `mcp_server.server init` on install — feature parity with `install.sh`.
1411
+ - **FIX**: MCP server spec no longer uses the fragile `cmd /c cd /d ... && python ...` wrapper on Windows; it emits the standard `command` + `cwd` shape supported natively by every modern client.
1412
+ - **FIX**: `install.sh` guards against `sh install.sh` (bash-only features now emit a clear error instead of a cryptic syntax failure).
1413
+ - **FIX**: Both scripts now correctly advertise **13 MCP tools** (was outdated at 12; `get_reindex_status` shipped in v4.3.0).
1414
+ - **FIX**: Windows-side `install.ps1` prefers `winget install Python.Python.3.12 --scope user` (no admin), falls back to python.org 3.12.7 (was pinned to 3.12.0).
1415
+ - **FIX**: Hybrid search now applies the active category filter to BM25 results before RRF fusion, preventing keyword-only BM25 hits from other categories from leaking into filtered searches. Both leak paths are closed: BM25 candidates are metadata-filtered before RRF (with `top_k` widened to `max_results * 20` to compensate for post-filter drop), and the fallback fetch during fusion re-checks `category` before adding a chunk to `combined_scores`. Note: the `category_filter` guard now also applies to the keyword-routed category — `_route_by_keywords()`-inferred routing filters BM25 too, consistent with the semantic branch. Users with warm `query_cache` entries should restart the server to invalidate stale results. (#109, thanks @Hohlas)
1416
+ - **TEST**: New `tests/test_installer_no_data_loss.py` (22 tests) locks in the installer's zero-data-loss contract across all three JSON schemas (`mcpServers` / `servers` / `context_servers`): top-level keys preserved, sibling MCP servers byte-identical, `.knowledge-rag.bak` backup written before every mutation, idempotent second run, `--dry-run` writes nothing, atomic `os.replace` write. Baseline: 231 → 266.
1417
+ - **TEST**: New `TestHybridCategoryFilter::test_bm25_results_respect_category_filter` in `tests/test_search.py` — deterministic regression covering BM25-only hits with mixed categories. Baseline: 266 → 267.
1418
+
1419
+ ### v4.3.1 (2026-06-22) — Hybrid Search Fixes
1420
+
1421
+ - **FIX**: Accept `"general"` as a valid category in `search_knowledge`. The parser hardcodes `"general"` as the fallback in `_detect_category` (`ingestion.py`), but the validator only built `valid_categories` from `config.keyword_routes` + `config.category_mappings.values()` — so users who customized `config.yaml` and dropped the default `"general": "general"` mapping hit `Invalid category` even though the index contained `general` documents. Validator now always tolerates `"general"`. (#98, thanks @Hohlas)
1422
+ - **FIX**: Skip BM25-only search results when Chroma can no longer resolve the chunk ID. Stale BM25 indices (typically right after `remove_document` or in the window between async reindex and BM25 rebuild) returned hits whose `collection.get()` came back empty; the previous fallback inserted entries with `document=""` / `metadata={}` into the reranker, polluting results with empty matches. The pipeline now `continue`s past those, dropping the stale hit cleanly. (#98, thanks @Hohlas)
1423
+ - **TEST**: Added `tests/test_pr98_regression.py` (4 tests) pinning both contracts so future refactors cannot silently revert either fix. Test count baseline: 227 → 231. (#99)
1424
+ - **CI**: Bumped `[tool.mypy] python_version` from 3.11 to 3.12 to accept PEP 695 `type` statements in the numpy stub (`numpy/__init__.pyi`) which were breaking the Pillar 7 strict gate. Only affects static analysis; `requires-python = ">=3.11"` unchanged. (#100)
1425
+
1426
+ ### v4.3.0 (2026-06-17) — Async Reindex, GPU CUDA 12, 13th MCP Tool
1427
+
1428
+ - **NEW**: `get_reindex_status` MCP tool — lightweight reindex progress polling without computing full index stats. Returns active/idle status, percent, processed/total, errors, and last result.
1429
+ - **NEW**: `reindex_documents` now runs in background via daemon thread — returns immediately with `{"status": "started"}`. Eliminates MCP timeout on large document sets (5K+ files). Concurrent calls return `already_running` with current progress.
1430
+ - **NEW**: GPU acceleration with full CUDA 12 support — `onnxruntime-gpu` + 7 NVIDIA pip packages (`cublas`, `cudnn`, `cuda-runtime`, `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`). Server auto-detects GPU on startup with 4-step verification (providers, DLLs, nvidia-smi, session creation). Falls back to CPU gracefully.
1431
+ - **NEW**: `_setup_cuda_dll_paths()` adds NVIDIA pip package DLL directories to `PATH` automatically on Windows — onnxruntime finds CUDA 12 DLLs without a full CUDA Toolkit install.
1432
+ - **DEPS**: `[gpu]` extra expanded from 3 to 8 packages (added `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`).
1433
+ - **FIX**: GPU status reporting now uses actual ONNX session creation test instead of just checking `get_available_providers()` — prevents false "GPU ACTIVE" when CUDA DLLs are missing.
1434
+ - **DOCS**: GPU Acceleration section rewritten with complete requirements table, setup steps, verification instructions, and fallback behavior.
1435
+ - **DOCS**: Tool reference updated — `reindex_documents` async behavior documented, `get_reindex_status` reference added.
1436
+ - **TEST**: Backwards-compat baseline updated for 13 MCP tools.
1437
+
1438
+ ### v4.2.0 (2026-06-17) — Search Performance & Output Quality
1439
+
1440
+ - **PERF**: Custom inverted-index BM25 replaces `rank-bm25` full-corpus scan — 128× faster keyword search on 50K+ chunk corpora. Only documents containing query terms are scored via posting lists.
1441
+ - **PERF**: `numpy.argpartition` for O(n) top-k selection instead of O(n log n) sort.
1442
+ - **PERF**: Batched adjacent chunk fetch — single ChromaDB `collection.get()` call replaces N round-trips per result.
1443
+ - **PERF**: O(1) reverse lookup via `_source_to_docid` dict eliminates linear scans of `_indexed_docs` in `search_similar`, `update_document`, `remove_document`, and `_expand_with_adjacent_chunks`.
1444
+ - **NEW**: `snippet_mode` parameter on `search_knowledge` (default: `true`) — truncates content to ~500 chars at natural break points with `content_length` field. Reduces token consumption by ~72%.
1445
+ - **NEW**: `min_score` parameter on `search_knowledge` (default: `0.0`) — filters results below a normalized relevance threshold. Response includes `filtered_by_score` count.
1446
+ - **NEW**: `filtered_by_score` field in search response JSON for transparency.
1447
+ - **DEPS**: `numpy` added as direct dependency (was transitive via fastembed); `rank-bm25` import removed from server.py.
1448
+ - **TEST**: 6 new tests for `min_score` filtering and `snippet_mode` truncation.
1449
+ - **TEST**: Updated backwards-compat baseline to include new `search_knowledge` parameters.
1450
+
1451
+ ### v4.1.2 (2026-06-17)
1452
+
1453
+ - **FIX**: `_save_metadata` dict snapshot prevents concurrent modification crash during file watcher events.
1454
+ - **STYLE**: ruff format applied to server.py.
1455
+
1456
+ ### v4.1.1 (2026-06-17)
1457
+
1458
+ - **FIX**: All `_indexed_docs` iterations now use `list()` snapshot, preventing `dictionary changed size during iteration` crash when FileWatcher modifies the index concurrently with MCP tool calls (affects `search_knowledge`, `search_similar`, `update_document`, `remove_document`, `evaluate_retrieval`, `list_categories`, `list_documents`)
1459
+
1460
+ ### v4.1.0 (2026-06-17)
1461
+
1462
+ - **Added:** `query_expansion_groups` config for symmetric synonym expansion (#92)
1463
+ - **Improved:** `expand_query()` now returns deterministic expansion order (set → ordered list with dedup)
1464
+
1465
+ ### v4.0.1 (2026-06-16)
1466
+
1467
+ - **FIX**: Orphan cleanup now runs before indexing loop, preventing chunk loss when files are moved (#90).
1468
+ - **FIX**: Chunk deduplication is now per-document instead of global, preventing cross-document chunk deletion (#91).
1469
+ - **FIX**: Added `on_moved` handler to `DocumentWatcher` for proper file move detection.
1470
+ - **FIX**: Startup preflight probes ChromaDB in a child process and moves crashing persistent indexes to `data/backups/auto-repair-*` before MCP initialization.
1471
+ - **FIX**: Reranker load failures now fall back to RRF ordering instead of failing `search_knowledge` on offline machines.
1472
+ - **FIX**: Virtualenv project-root detection now handles Python symlinks that resolve to the system interpreter.
1473
+ - **NEW**: `knowledge-rag-guarded` console script kept as an explicit guarded startup alias.
1474
+
1402
1475
  ### v4.0.0 (2026-06-09) — Enterprise Concurrent Access
1403
1476
 
1404
1477
  - **NEW**: SSE and streamable-http transport modes — 1 server serves N clients (`server.transport: "sse"` in config.yaml or `--transport sse` CLI).
@@ -1434,7 +1507,7 @@ Common issues:
1434
1507
  - **NEW** Property-based fuzzing of all parsers via Hypothesis (`tests/test_ingestion_property.py`) — 200 random examples per CI run.
1435
1508
  - **NEW** Memory baseline regression tests (`tests/test_memory_baseline.py`, cross-platform via psutil) — RSS bounded under 1000 queries; nightly soak amplifies to 50K iterations.
1436
1509
  - **NEW** Property/locale/format/preset matrices (`tests/test_presets.py`, `tests/test_locale.py`, `tests/test_format_smoke.py`).
1437
- - **NEW** Backwards-compatibility regression tests (`tests/test_backwards_compat.py`) — legacy YAML configs from v3.6.0 / v3.7.0 still parse; all 12 MCP tool parameter names frozen.
1510
+ - **NEW** Backwards-compatibility regression tests (`tests/test_backwards_compat.py`) — legacy YAML configs from v3.6.0 / v3.7.0 still parse; all 13 MCP tool parameter names frozen.
1438
1511
  - **NEW** AST-based public API surface diff (`scripts/check_api_surface.py`) — any breaking change blocks merge, baseline at `.github/api-surface-baseline.json`.
1439
1512
  - **NEW** CHANGELOG enforcement (`scripts/check_changelog.py`) — user-facing PRs must add a bullet under `## Unreleased`; bypass via `skip-changelog` label.
1440
1513
  - **NEW** Test count anti-regression (`scripts/check_test_count.py`) — guards against silent test deletion.
@@ -1464,57 +1537,6 @@ Common issues:
1464
1537
  - **CHORE**: pytest `tmp_path_retention_count=1` to avoid Windows atexit cleanup race in CI.
1465
1538
  - **ROADMAP**: Tracked v4.0 shared-service architecture (one daemon, many thin MCP clients) as the long-term fix for multi-process resource duplication. (#34)
1466
1539
 
1467
- ### Unreleased
1468
-
1469
- ### v4.3.0 (2026-06-17) — Async Reindex, GPU CUDA 12, 13th MCP Tool
1470
-
1471
- - **NEW**: `get_reindex_status` MCP tool — lightweight reindex progress polling without computing full index stats. Returns active/idle status, percent, processed/total, errors, and last result.
1472
- - **NEW**: `reindex_documents` now runs in background via daemon thread — returns immediately with `{"status": "started"}`. Eliminates MCP timeout on large document sets (5K+ files). Concurrent calls return `already_running` with current progress.
1473
- - **NEW**: GPU acceleration with full CUDA 12 support — `onnxruntime-gpu` + 7 NVIDIA pip packages (`cublas`, `cudnn`, `cuda-runtime`, `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`). Server auto-detects GPU on startup with 4-step verification (providers, DLLs, nvidia-smi, session creation). Falls back to CPU gracefully.
1474
- - **NEW**: `_setup_cuda_dll_paths()` adds NVIDIA pip package DLL directories to `PATH` automatically on Windows — onnxruntime finds CUDA 12 DLLs without a full CUDA Toolkit install.
1475
- - **DEPS**: `[gpu]` extra expanded from 3 to 8 packages (added `cufft`, `cusparse`, `cusolver`, `curand`, `nvjitlink`).
1476
- - **FIX**: GPU status reporting now uses actual ONNX session creation test instead of just checking `get_available_providers()` — prevents false "GPU ACTIVE" when CUDA DLLs are missing.
1477
- - **DOCS**: GPU Acceleration section rewritten with complete requirements table, setup steps, verification instructions, and fallback behavior.
1478
- - **DOCS**: Tool reference updated — `reindex_documents` async behavior documented, `get_reindex_status` reference added.
1479
- - **TEST**: Backwards-compat baseline updated for 13 MCP tools.
1480
-
1481
- ### v4.2.0 (2026-06-17) — Search Performance & Output Quality
1482
-
1483
- - **PERF**: Custom inverted-index BM25 replaces `rank-bm25` full-corpus scan — 128× faster keyword search on 50K+ chunk corpora. Only documents containing query terms are scored via posting lists.
1484
- - **PERF**: `numpy.argpartition` for O(n) top-k selection instead of O(n log n) sort.
1485
- - **PERF**: Batched adjacent chunk fetch — single ChromaDB `collection.get()` call replaces N round-trips per result.
1486
- - **PERF**: O(1) reverse lookup via `_source_to_docid` dict eliminates linear scans of `_indexed_docs` in `search_similar`, `update_document`, `remove_document`, and `_expand_with_adjacent_chunks`.
1487
- - **NEW**: `snippet_mode` parameter on `search_knowledge` (default: `true`) — truncates content to ~500 chars at natural break points with `content_length` field. Reduces token consumption by ~72%.
1488
- - **NEW**: `min_score` parameter on `search_knowledge` (default: `0.0`) — filters results below a normalized relevance threshold. Response includes `filtered_by_score` count.
1489
- - **NEW**: `filtered_by_score` field in search response JSON for transparency.
1490
- - **DEPS**: `numpy` added as direct dependency (was transitive via fastembed); `rank-bm25` import removed from server.py.
1491
- - **TEST**: 6 new tests for `min_score` filtering and `snippet_mode` truncation.
1492
- - **TEST**: Updated backwards-compat baseline to include new `search_knowledge` parameters.
1493
-
1494
- ### v4.1.2 (2026-06-17)
1495
-
1496
- - **FIX**: `_save_metadata` dict snapshot prevents concurrent modification crash during file watcher events.
1497
- - **STYLE**: ruff format applied to server.py.
1498
-
1499
- ### v4.1.1 (2026-06-17)
1500
-
1501
- - **FIX**: All `_indexed_docs` iterations now use `list()` snapshot, preventing `dictionary changed size during iteration` crash when FileWatcher modifies the index concurrently with MCP tool calls (affects `search_knowledge`, `search_similar`, `update_document`, `remove_document`, `evaluate_retrieval`, `list_categories`, `list_documents`)
1502
-
1503
- ### v4.1.0 (2026-06-17)
1504
-
1505
- - **Added:** `query_expansion_groups` config for symmetric synonym expansion (#92)
1506
- - **Improved:** `expand_query()` now returns deterministic expansion order (set → ordered list with dedup)
1507
-
1508
- ### v4.0.1 (2026-06-16)
1509
-
1510
- - **FIX**: Orphan cleanup now runs before indexing loop, preventing chunk loss when files are moved (#90).
1511
- - **FIX**: Chunk deduplication is now per-document instead of global, preventing cross-document chunk deletion (#91).
1512
- - **FIX**: Added `on_moved` handler to `DocumentWatcher` for proper file move detection.
1513
- - **FIX**: Startup preflight probes ChromaDB in a child process and moves crashing persistent indexes to `data/backups/auto-repair-*` before MCP initialization.
1514
- - **FIX**: Reranker load failures now fall back to RRF ordering instead of failing `search_knowledge` on offline machines.
1515
- - **FIX**: Virtualenv project-root detection now handles Python symlinks that resolve to the system interpreter.
1516
- - **NEW**: `knowledge-rag-guarded` console script kept as an explicit guarded startup alias.
1517
-
1518
1540
  ### v3.6.2 (2026-04-23)
1519
1541
 
1520
1542
  - **INFRA**: NPM provenance attestation (SLSA supply chain security), full README on npm page
@@ -8,7 +8,7 @@ import sys # noqa: I001
8
8
  _original_stdout = sys.stdout
9
9
  sys.stdout = sys.stderr
10
10
 
11
- __version__ = "4.3.0"
11
+ __version__ = "4.4.0"
12
12
  __author__ = "Ailton Rocha (Lyon.)"
13
13
 
14
14
  from .config import Config # noqa: E402
@@ -1474,6 +1474,12 @@ class KnowledgeOrchestrator:
1474
1474
  elif routed_category:
1475
1475
  where_filter = {"category": routed_category}
1476
1476
 
1477
+ def _matches_category(metadata: Dict[str, Any]) -> bool:
1478
+ if not where_filter:
1479
+ return True
1480
+ expected_category = where_filter.get("category")
1481
+ return not expected_category or metadata.get("category") == expected_category
1482
+
1477
1483
  # Parallel Semantic + BM25 search (threaded for latency reduction)
1478
1484
  from concurrent.futures import ThreadPoolExecutor
1479
1485
 
@@ -1507,8 +1513,23 @@ class KnowledgeOrchestrator:
1507
1513
  r = {}
1508
1514
  if hybrid_alpha < 1.0:
1509
1515
  try:
1510
- bm25_hits = self.bm25_index.search(query_text, top_k=max_results * 3)
1511
- for rank, (chunk_id, bm25_score) in enumerate(bm25_hits):
1516
+ bm25_top_k = max_results * (20 if where_filter else 3)
1517
+ bm25_hits = self.bm25_index.search(query_text, top_k=bm25_top_k)
1518
+
1519
+ if where_filter:
1520
+ chunk_ids = [chunk_id for chunk_id, _ in bm25_hits]
1521
+ metadata_by_id = {}
1522
+ if chunk_ids:
1523
+ fetched = self.collection.get(ids=chunk_ids, include=["metadatas"])
1524
+ metadata_by_id = dict(zip(fetched.get("ids", []), fetched.get("metadatas", [])))
1525
+
1526
+ bm25_hits = [
1527
+ (chunk_id, bm25_score)
1528
+ for chunk_id, bm25_score in bm25_hits
1529
+ if _matches_category(metadata_by_id.get(chunk_id, {}))
1530
+ ]
1531
+
1532
+ for rank, (chunk_id, bm25_score) in enumerate(bm25_hits[: max_results * 3]):
1512
1533
  r[chunk_id] = {"rank": rank + 1, "bm25_score": bm25_score}
1513
1534
  except Exception as e:
1514
1535
  print(f"[WARN] BM25 search failed: {e}")
@@ -1543,14 +1564,24 @@ class KnowledgeOrchestrator:
1543
1564
  else:
1544
1565
  try:
1545
1566
  fetched = self.collection.get(ids=[chunk_id], include=["documents", "metadatas"])
1567
+ if (
1568
+ not fetched["documents"]
1569
+ or not fetched["metadatas"]
1570
+ or not fetched["documents"][0]
1571
+ or not fetched["metadatas"][0]
1572
+ ):
1573
+ continue
1546
1574
  data = {
1547
- "document": fetched["documents"][0] if fetched["documents"] else "",
1548
- "metadata": fetched["metadatas"][0] if fetched["metadatas"] else {},
1575
+ "document": fetched["documents"][0],
1576
+ "metadata": fetched["metadatas"][0],
1549
1577
  "distance": 0,
1550
1578
  }
1551
1579
  except Exception:
1552
1580
  continue
1553
1581
 
1582
+ if not _matches_category(data.get("metadata", {})):
1583
+ continue
1584
+
1554
1585
  combined_scores[chunk_id] = {
1555
1586
  "rrf_score": combined_rrf,
1556
1587
  "semantic_rank": semantic_rank if chunk_id in semantic_results else None,
@@ -2273,7 +2304,7 @@ def search_knowledge(
2273
2304
  hybrid_alpha = max(0.0, min(hybrid_alpha if hybrid_alpha is not None else 0.3, 1.0))
2274
2305
  min_score = max(0.0, min(min_score if min_score is not None else 0.0, 1.0))
2275
2306
 
2276
- valid_categories = list(config.keyword_routes.keys()) + list(set(config.category_mappings.values()))
2307
+ valid_categories = list(config.keyword_routes.keys()) + list(set(config.category_mappings.values())) + ["general"]
2277
2308
  if category and category not in valid_categories:
2278
2309
  return json.dumps(
2279
2310
  {"status": "error", "message": f"Invalid category '{category}'. Valid: {', '.join(valid_categories)}"}
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
4
4
 
5
5
  [project]
6
6
  name = "knowledge-rag"
7
- version = "4.3.0"
7
+ version = "4.4.0"
8
8
  description = "Local RAG System for Claude Code — Hybrid search + Cross-encoder Reranking + 13 MCP Tools + 20 Format Parsers. Zero external servers."
9
9
  readme = "README.md"
10
10
  license = {text = "MIT"}
@@ -141,7 +141,11 @@ fail_under = 35
141
141
  # we incrementally annotate the legacy modules. The CI job runs strict on the
142
142
  # allowlist below; new modules are added as they earn full annotations.
143
143
  [tool.mypy]
144
- python_version = "3.11"
144
+ # Pinned to 3.12 to match the CI runtime (Python 3.12) and to accept PEP 695
145
+ # ``type`` statements that recent third-party stubs (e.g. numpy/__init__.pyi)
146
+ # emit. The package itself still supports Python 3.11+ at runtime — this
147
+ # setting only governs the static-analysis target.
148
+ python_version = "3.12"
145
149
  strict = true
146
150
  show_error_codes = true
147
151
  warn_unused_configs = true
File without changes
File without changes