mongo-x-ray-risk 2.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,70 @@
1
+ Metadata-Version: 2.4
2
+ Name: mongo-x-ray-risk
3
+ Version: 2.0.0
4
+ Summary: Known-risks knowledge base (ChromaDB vector search) for x-ray
5
+ Requires-Python: <4,>=3.10
6
+ Description-Content-Type: text/markdown
7
+ Requires-Dist: mongo-x-ray>=2.0.0
8
+ Requires-Dist: chromadb==1.5.9
9
+
10
+ # mongo-x-ray-risk
11
+
12
+ [![CI](https://github.com/zhangyaoxing/mongo-x-ray-risk/actions/workflows/ci.yml/badge.svg)](https://github.com/zhangyaoxing/mongo-x-ray-risk/actions/workflows/ci.yml)
13
+ [![PyPI](https://img.shields.io/pypi/v/mongo-x-ray-risk.svg)](https://pypi.org/project/mongo-x-ray-risk/)
14
+
15
+ Known-risks knowledge base for [x-ray](https://github.com/mongodb-ps/ce-mongo-x-ray): a ChromaDB-backed
16
+ vector search that matches analysis findings against known MongoDB risks.
17
+
18
+ This is an optional plugin: it ships the `ingest` command to load the risk register, the `search`
19
+ command to look risks up by name, and the analysis plugins (healthcheck, log, gmd) detect it at
20
+ runtime — when it is installed, their reports are enriched with matched risks; when it is missing,
21
+ the enrichment is silently skipped.
22
+
23
+ ## Install
24
+
25
+ ```bash
26
+ pip install mongo-x-ray mongo-x-ray-risk
27
+ ```
28
+
29
+ ## Usage
30
+
31
+ Load a risk register CSV into the ChromaDB knowledge base:
32
+
33
+ ```bash
34
+ x-ray ingest risk_register.csv
35
+ # start from an empty register, then ingest
36
+ x-ray ingest --clear risk_register.csv
37
+ # clear the register without ingesting (no CSV needed)
38
+ x-ray ingest --clear
39
+ ```
40
+
41
+ Check whether a risk is already known by searching the `Name` column
42
+ (case-insensitive substring match). Prints the `Name` and `Risk description`
43
+ of every matching risk; exits 0 when something matches, 1 when nothing does:
44
+
45
+ ```bash
46
+ x-ray search "Replication Lag"
47
+ x-ray search replication
48
+ ```
49
+
50
+ The CSV must have the columns `ID, Risk level, Impact, Name, Risk description`
51
+ (UTF-8, a BOM is tolerated; header names are matched case-insensitively). Any
52
+ other columns, such as `Other Notes`, are ignored. Rows without an ID or a
53
+ Name are skipped; entries with an existing ID are replaced. The data is
54
+ stored under `~/.x-ray/chroma`.
55
+
56
+ Once ingested, the plugin is used automatically by the other plugins — no CLI
57
+ flags needed. It also exposes a small API for tooling:
58
+
59
+ ```python
60
+ from mongo_x_ray_risk import Risk, load_risks_from_csv, ingest_risks, match_risk, enrich_test_results
61
+ ```
62
+
63
+ ## Development
64
+
65
+ Requires Python 3.10+, MongoDB 5.0 or later, and the [mongo-x-ray](https://github.com/mongodb-ps/ce-mongo-x-ray) core package.
66
+
67
+ ```bash
68
+ make unit-test # run the unit tests
69
+ make lint # ruff check + ruff format --check
70
+ ```
@@ -0,0 +1,61 @@
1
+ # mongo-x-ray-risk
2
+
3
+ [![CI](https://github.com/zhangyaoxing/mongo-x-ray-risk/actions/workflows/ci.yml/badge.svg)](https://github.com/zhangyaoxing/mongo-x-ray-risk/actions/workflows/ci.yml)
4
+ [![PyPI](https://img.shields.io/pypi/v/mongo-x-ray-risk.svg)](https://pypi.org/project/mongo-x-ray-risk/)
5
+
6
+ Known-risks knowledge base for [x-ray](https://github.com/mongodb-ps/ce-mongo-x-ray): a ChromaDB-backed
7
+ vector search that matches analysis findings against known MongoDB risks.
8
+
9
+ This is an optional plugin: it ships the `ingest` command to load the risk register, the `search`
10
+ command to look risks up by name, and the analysis plugins (healthcheck, log, gmd) detect it at
11
+ runtime — when it is installed, their reports are enriched with matched risks; when it is missing,
12
+ the enrichment is silently skipped.
13
+
14
+ ## Install
15
+
16
+ ```bash
17
+ pip install mongo-x-ray mongo-x-ray-risk
18
+ ```
19
+
20
+ ## Usage
21
+
22
+ Load a risk register CSV into the ChromaDB knowledge base:
23
+
24
+ ```bash
25
+ x-ray ingest risk_register.csv
26
+ # start from an empty register, then ingest
27
+ x-ray ingest --clear risk_register.csv
28
+ # clear the register without ingesting (no CSV needed)
29
+ x-ray ingest --clear
30
+ ```
31
+
32
+ Check whether a risk is already known by searching the `Name` column
33
+ (case-insensitive substring match). Prints the `Name` and `Risk description`
34
+ of every matching risk; exits 0 when something matches, 1 when nothing does:
35
+
36
+ ```bash
37
+ x-ray search "Replication Lag"
38
+ x-ray search replication
39
+ ```
40
+
41
+ The CSV must have the columns `ID, Risk level, Impact, Name, Risk description`
42
+ (UTF-8, a BOM is tolerated; header names are matched case-insensitively). Any
43
+ other columns, such as `Other Notes`, are ignored. Rows without an ID or a
44
+ Name are skipped; entries with an existing ID are replaced. The data is
45
+ stored under `~/.x-ray/chroma`.
46
+
47
+ Once ingested, the plugin is used automatically by the other plugins — no CLI
48
+ flags needed. It also exposes a small API for tooling:
49
+
50
+ ```python
51
+ from mongo_x_ray_risk import Risk, load_risks_from_csv, ingest_risks, match_risk, enrich_test_results
52
+ ```
53
+
54
+ ## Development
55
+
56
+ Requires Python 3.10+, MongoDB 5.0 or later, and the [mongo-x-ray](https://github.com/mongodb-ps/ce-mongo-x-ray) core package.
57
+
58
+ ```bash
59
+ make unit-test # run the unit tests
60
+ make lint # ruff check + ruff format --check
61
+ ```
@@ -0,0 +1,56 @@
1
+ [build-system]
2
+ requires = ["setuptools==83.0.0", "wheel"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "mongo-x-ray-risk"
7
+ version = "2.0.0"
8
+ description = "Known-risks knowledge base (ChromaDB vector search) for x-ray"
9
+ readme = "README.md"
10
+ requires-python = ">=3.10,<4"
11
+ dependencies = [
12
+ "mongo-x-ray>=2.0.0",
13
+ "chromadb==1.5.9",
14
+ ]
15
+
16
+ [project.entry-points."mongo_x_ray.plugins"]
17
+ ingest = "mongo_x_ray_risk.plugin:IngestPlugin"
18
+ search = "mongo_x_ray_risk.plugin:SearchPlugin"
19
+
20
+ [tool.setuptools.packages.find]
21
+ where = ["src"]
22
+ include = ["mongo_x_ray_risk*"]
23
+
24
+ [tool.pytest.ini_options]
25
+ testpaths = ["tests"]
26
+ pythonpath = ["src"]
27
+ python_files = ["test_*.py"]
28
+ python_classes = ["Test*"]
29
+ python_functions = ["test_*"]
30
+ markers = [
31
+ "slow: marks tests as slow",
32
+ "integration: marks tests as integration tests",
33
+ "unit: marks tests as unit tests",
34
+ ]
35
+
36
+ [tool.ruff]
37
+ line-length = 120
38
+ target-version = "py310"
39
+ exclude = [".venv", "build", "dist"]
40
+
41
+ [tool.ruff.lint]
42
+ select = ["E", "F", "I"]
43
+ ignore = ["E501"]
44
+
45
+ [tool.ruff.lint.isort]
46
+ known-first-party = ["mongo_x_ray", "mongo_x_ray_ftdc", "mongo_x_ray_gmd", "mongo_x_ray_hc", "mongo_x_ray_log", "mongo_x_ray_risk"]
47
+
48
+ [tool.pyright]
49
+ pythonVersion = "3.10"
50
+ venvPath = "../ce-mongo-x-ray"
51
+ venv = ".venv"
52
+ typeCheckingMode = "basic"
53
+ extraPaths = [
54
+ "src",
55
+ "../ce-mongo-x-ray/src",
56
+ ]
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+
@@ -0,0 +1,40 @@
1
+ """
2
+ Copyright (c) 2026 MongoDB Inc.
3
+
4
+ DISCLAIMER: THESE CODE SAMPLES ARE PROVIDED FOR EDUCATIONAL AND ILLUSTRATIVE PURPOSES ONLY,
5
+ TO DEMONSTRATE THE FUNCTIONALITY OF SPECIFIC MONGODB FEATURES.
6
+ THEY ARE NOT PRODUCTION-READY AND MAY LACK THE SECURITY HARDENING, ERROR HANDLING, AND TESTING REQUIRED FOR A LIVE ENVIRONMENT.
7
+ YOU ARE RESPONSIBLE FOR TESTING, VALIDATING, AND SECURING THIS CODE WITHIN YOUR OWN ENVIRONMENT BEFORE IMPLEMENTATION.
8
+ THIS MATERIAL IS PROVIDED "AS IS" WITHOUT WARRANTY OR LIABILITY.
9
+
10
+ Risk Register — ChromaDB-backed vector search for known risks.
11
+
12
+ The public API is re-exported here; consumers should import from this package
13
+ rather than from the ``db``/``shared`` submodules, which are implementation
14
+ details. The enrichment functions operate on plain values (a ``list[dict]`` of
15
+ test results, a ``str`` category), so they never depend on any analysis
16
+ module's item types.
17
+ """
18
+
19
+ from mongo_x_ray_risk.db import (
20
+ clear_risks,
21
+ enrich_test_results,
22
+ find_risks_by_name,
23
+ has_risks,
24
+ ingest_risks,
25
+ match_risk,
26
+ search_risks,
27
+ )
28
+ from mongo_x_ray_risk.shared import Risk, load_risks_from_csv
29
+
30
+ __all__ = [
31
+ "Risk",
32
+ "load_risks_from_csv",
33
+ "ingest_risks",
34
+ "search_risks",
35
+ "find_risks_by_name",
36
+ "clear_risks",
37
+ "match_risk",
38
+ "enrich_test_results",
39
+ "has_risks",
40
+ ]
@@ -0,0 +1,259 @@
1
+ """
2
+ Copyright (c) 2026 MongoDB Inc.
3
+
4
+ DISCLAIMER: THESE CODE SAMPLES ARE PROVIDED FOR EDUCATIONAL AND ILLUSTRATIVE PURPOSES ONLY,
5
+ TO DEMONSTRATE THE FUNCTIONALITY OF SPECIFIC MONGODB FEATURES.
6
+ THEY ARE NOT PRODUCTION-READY AND MAY LACK THE SECURITY HARDENING, ERROR HANDLING, AND TESTING REQUIRED FOR A LIVE ENVIRONMENT.
7
+ YOU ARE RESPONSIBLE FOR TESTING, VALIDATING, AND SECURING THIS CODE WITHIN YOUR OWN ENVIRONMENT BEFORE IMPLEMENTATION.
8
+ THIS MATERIAL IS PROVIDED "AS IS" WITHOUT WARRANTY OR LIABILITY.
9
+
10
+ ChromaDB-backed risk register with vector search.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ import logging
16
+ from typing import Any, Mapping, Optional
17
+
18
+ from mongo_x_ray_risk.shared import (
19
+ CHROMA_COLLECTION,
20
+ CHROMA_COLLECTION_DESCRIPTION,
21
+ Risk,
22
+ get_db_path,
23
+ )
24
+
25
+ # Mute chromadb telemetry errors (posthog API mismatch)
26
+ logging.getLogger("chromadb.telemetry").setLevel(logging.CRITICAL)
27
+
28
+ _logger = logging.getLogger(__name__)
29
+
30
+
31
+ def _collection(collection_name: str = CHROMA_COLLECTION):
32
+ """Return an initialized ChromaDB collection (lazy singleton)."""
33
+ # Import chromadb lazily so importing this module stays cheap and the risk
34
+ # register remains an optional best-effort enrichment. ChromaDB is not
35
+ # bundled into the frozen x-ray binary, so a missing import surfaces as a
36
+ # clear error instead of a raw traceback.
37
+ try:
38
+ import chromadb
39
+ from chromadb.config import Settings
40
+ except ImportError as exc:
41
+ raise RuntimeError(
42
+ "ChromaDB is not available in this build. Install the "
43
+ "'mongo-x-ray-risk' pip package (which depends on chromadb) "
44
+ "to use the risk register."
45
+ ) from exc
46
+
47
+ db_path = get_db_path() / "chroma"
48
+ db_path.mkdir(parents=True, exist_ok=True)
49
+ client = chromadb.PersistentClient(
50
+ path=str(db_path),
51
+ settings=Settings(anonymized_telemetry=False),
52
+ )
53
+ return client.get_or_create_collection(collection_name)
54
+
55
+
56
+ def ingest_risks(risks: list[Risk]) -> int:
57
+ """Upsert risks into ChromaDB, returning the number of documents ingested.
58
+
59
+ Each risk is embedded twice — once for the ``Name`` field and once for the
60
+ ``Risk Description`` field — stored in two separate collections so that
61
+ matching can fall back from Name to Risk Description. Risks without a Risk
62
+ Description are only embedded in the Name collection. Existing documents
63
+ with the same ID are replaced (upsert).
64
+ """
65
+ if not risks:
66
+ return 0
67
+
68
+ name_col = _collection(CHROMA_COLLECTION)
69
+ desc_col = _collection(CHROMA_COLLECTION_DESCRIPTION)
70
+
71
+ name_ids: list[str] = []
72
+ name_documents: list[str] = []
73
+ name_metadatas: list[Mapping[str, Any]] = []
74
+ desc_ids: list[str] = []
75
+ desc_documents: list[str] = []
76
+ desc_metadatas: list[Mapping[str, Any]] = []
77
+
78
+ for risk in risks:
79
+ metadata = {
80
+ "id": risk.id,
81
+ "risk_level": risk.risk_level,
82
+ "impact": risk.impact,
83
+ "name": risk.name,
84
+ "description": risk.description,
85
+ }
86
+ name_ids.append(risk.id)
87
+ name_documents.append(risk.name)
88
+ name_metadatas.append(metadata)
89
+ if risk.description.strip():
90
+ desc_ids.append(risk.id)
91
+ desc_documents.append(risk.description)
92
+ desc_metadatas.append(metadata)
93
+
94
+ name_col.upsert(ids=name_ids, documents=name_documents, metadatas=name_metadatas)
95
+ if desc_ids:
96
+ desc_col.upsert(ids=desc_ids, documents=desc_documents, metadatas=desc_metadatas)
97
+ _logger.info("Ingested %d risks into ChromaDB", len(risks))
98
+ return len(risks)
99
+
100
+
101
+ def search_risks(
102
+ query: str,
103
+ n_results: int = 3,
104
+ collection_name: str = CHROMA_COLLECTION,
105
+ ) -> list[dict]:
106
+ """Vector search for risks matching the query text.
107
+
108
+ Args:
109
+ query: The text to search for.
110
+ n_results: Maximum number of results to return.
111
+ collection_name: Which field collection to search; defaults to the
112
+ risk ``Name`` collection.
113
+
114
+ Returns:
115
+ A list of dicts with keys: id, risk_level, impact, name,
116
+ description, distance.
117
+ """
118
+ col = _collection(collection_name)
119
+ results = col.query(query_texts=[query], n_results=n_results)
120
+ entries: list[dict] = []
121
+ if not results["ids"] or not results["ids"][0]:
122
+ return entries
123
+ for i, doc_id in enumerate(results["ids"][0]):
124
+ meta = results["metadatas"][0][i] if results["metadatas"] else {}
125
+ distance = results["distances"][0][i] if results["distances"] else None
126
+ entries.append(
127
+ {
128
+ "id": meta.get("id", doc_id),
129
+ "risk_level": meta.get("risk_level", ""),
130
+ "impact": meta.get("impact", ""),
131
+ "name": meta.get("name", ""),
132
+ "description": meta.get("description", ""),
133
+ "distance": distance,
134
+ }
135
+ )
136
+ return entries
137
+
138
+
139
+ def find_risks_by_name(query: str) -> list[dict]:
140
+ """Return risks whose ``Name`` contains *query* (case-insensitive).
141
+
142
+ Unlike :func:`search_risks` (vector similarity), this is a deterministic
143
+ substring match on the risk ``Name`` field — the right tool for checking
144
+ whether a proposed new risk name already exists in the register.
145
+
146
+ Args:
147
+ query: The name fragment to search for. Leading/trailing whitespace
148
+ is ignored; an empty query matches nothing.
149
+
150
+ Returns:
151
+ A list of dicts with keys: id, risk_level, impact, name, description.
152
+ Matches are returned in insertion order; no ``distance`` is included.
153
+ """
154
+ needle = query.strip().lower()
155
+ if not needle:
156
+ return []
157
+ col = _collection(CHROMA_COLLECTION)
158
+ got = col.get(include=["metadatas"])
159
+ entries: list[dict] = []
160
+ for doc_id, meta in zip(got["ids"], got["metadatas"] or []):
161
+ meta = meta or {}
162
+ name = str(meta.get("name", "") or "")
163
+ if needle in name.lower():
164
+ entries.append(
165
+ {
166
+ "id": str(meta.get("id", doc_id)),
167
+ "risk_level": str(meta.get("risk_level", "") or ""),
168
+ "impact": str(meta.get("impact", "") or ""),
169
+ "name": name,
170
+ "description": str(meta.get("description", "") or ""),
171
+ }
172
+ )
173
+ return entries
174
+
175
+
176
+ def clear_risks() -> None:
177
+ """Delete all documents from all risk collections."""
178
+ for collection_name in (CHROMA_COLLECTION, CHROMA_COLLECTION_DESCRIPTION):
179
+ col = _collection(collection_name)
180
+ ids = col.get()["ids"]
181
+ if ids:
182
+ col.delete(ids=ids)
183
+ _logger.info("Cleared %d risks from %s", len(ids), collection_name)
184
+
185
+
186
+ def _collection_count() -> int:
187
+ """Return the number of documents in the Name collection."""
188
+ col = _collection(CHROMA_COLLECTION)
189
+ return col.count()
190
+
191
+
192
+ def has_risks() -> bool:
193
+ """Return True if the risk register contains any ingested risks.
194
+
195
+ Never raises: a missing/corrupt ChromaDB or an empty register both mean
196
+ ``False``, so callers can decide whether to show risk-related UI.
197
+ """
198
+ try:
199
+ return _collection_count() > 0
200
+ except Exception:
201
+ return False
202
+
203
+
204
+ def match_risk(category: str, max_distance: float = 0.9) -> Optional[dict]:
205
+ """Find the closest matching risk for a given issue.
206
+
207
+ Two-stage fallback search:
208
+ 1. Match ``category`` against the risk ``Name`` field.
209
+ 2. If no ``Name`` match is found, fall back to matching the same
210
+ ``category`` against the risk ``Risk Description`` field.
211
+
212
+ Args:
213
+ category: The issue category to match — the ``Alert Category``
214
+ column in the reports (each item's ``title``).
215
+ max_distance: Maximum vector distance for a match to be considered
216
+ valid. Lower values mean closer matches. Default 0.9.
217
+
218
+ Returns:
219
+ The best matching risk dict, or ``None`` if no match found.
220
+ """
221
+ for collection_name in (CHROMA_COLLECTION, CHROMA_COLLECTION_DESCRIPTION):
222
+ results = search_risks(category, n_results=1, collection_name=collection_name)
223
+ if not results:
224
+ continue
225
+ top = results[0]
226
+ if top["distance"] is not None and top["distance"] > max_distance:
227
+ continue
228
+ return top
229
+ return None
230
+
231
+
232
+ def enrich_test_results(test_results: list[dict], max_distance: float = 0.9) -> int:
233
+ """Enrich a list of test results with matched risk information.
234
+
235
+ Each result dict that has a ``title`` key (the ``Alert Category`` column
236
+ in the reports) will be matched against the risk register's ``Name``
237
+ field using the two-stage fallback search. If a match is found, a
238
+ ``matched_risk`` key is added.
239
+
240
+ Args:
241
+ test_results: List of test result dicts (each has a ``title`` key).
242
+ max_distance: Maximum vector distance for a match.
243
+
244
+ Returns:
245
+ Number of results that were successfully matched to a risk.
246
+ """
247
+ if _collection_count() == 0:
248
+ _logger.warning("\033[33mRisk register is empty — run `x-ray ingest <csv>` first\033[0m")
249
+ return 0
250
+ matched = 0
251
+ for result in test_results:
252
+ title = result.get("title", "")
253
+ if not title:
254
+ continue
255
+ risk = match_risk(title, max_distance=max_distance)
256
+ if risk:
257
+ result["matched_risk"] = risk
258
+ matched += 1
259
+ return matched
@@ -0,0 +1,144 @@
1
+ """
2
+ Copyright (c) 2026 MongoDB Inc.
3
+
4
+ DISCLAIMER: THESE CODE SAMPLES ARE PROVIDED FOR EDUCATIONAL AND ILLUSTRATIVE PURPOSES ONLY,
5
+ TO DEMONSTRATE THE FUNCTIONALITY OF SPECIFIC MONGODB FEATURES.
6
+ THEY ARE NOT PRODUCTION-READY AND MAY LACK THE SECURITY HARDENING, ERROR HANDLING, AND TESTING REQUIRED FOR A LIVE ENVIRONMENT.
7
+ YOU ARE RESPONSIBLE FOR TESTING, VALIDATING, AND SECURING THIS CODE WITHIN YOUR OWN ENVIRONMENT BEFORE IMPLEMENTATION.
8
+ THIS MATERIAL IS PROVIDED "AS IS" WITHOUT WARRANTY OR LIABILITY.
9
+ """
10
+
11
+ import logging
12
+ from pathlib import Path
13
+
14
+ from mongo_x_ray.plugin import Plugin
15
+ from mongo_x_ray_risk import (
16
+ clear_risks,
17
+ find_risks_by_name,
18
+ ingest_risks,
19
+ load_risks_from_csv,
20
+ )
21
+
22
+ logger = logging.getLogger(__name__)
23
+
24
+
25
+ class IngestPlugin(Plugin):
26
+ name = "ingest"
27
+ distribution = "mongo-x-ray-risk"
28
+ help = "Ingest a risk register CSV into the risk knowledge base"
29
+ description = """
30
+ Ingest a risk register CSV into the ChromaDB-backed risk knowledge base used
31
+ by the analysis plugins (healthcheck, log, gmd) to match findings against
32
+ known risks.
33
+
34
+ The CSV must have the following columns:
35
+ ID, Risk level, Impact, Name, Risk description
36
+
37
+ Header names are matched case-insensitively; any other columns (e.g.
38
+ Other Notes) are ignored. Rows without an ID or a Name are skipped. Existing
39
+ entries with the same ID are replaced; use --clear to start from an empty
40
+ register.
41
+
42
+ Run 'x-ray ingest --clear' alone to clear the register without a CSV.
43
+ """
44
+ epilog = """
45
+ Examples:
46
+ x-ray ingest risk_register.csv
47
+ x-ray ingest --clear risk_register.csv # clear, then ingest
48
+ x-ray ingest --clear # clear only
49
+ """
50
+
51
+ def add_arguments(self, parser):
52
+ parser.add_argument("csv", nargs="?", help="Path to the risk register CSV file.")
53
+ parser.add_argument(
54
+ "--clear",
55
+ help="Clear the existing risk register before ingesting.",
56
+ action="store_true",
57
+ default=False,
58
+ )
59
+
60
+ def run(self, args) -> int:
61
+ """Clear the risk register and/or ingest a CSV, reporting what was done."""
62
+ if not args.csv:
63
+ if args.clear:
64
+ try:
65
+ clear_risks()
66
+ except Exception as exc:
67
+ logger.error("Risk register unavailable: %s", exc)
68
+ return 1
69
+ logger.info("Cleared the existing risk register")
70
+ return 0
71
+ logger.error(
72
+ "No CSV file given. Use 'x-ray ingest <csv>' to ingest, "
73
+ "or 'x-ray ingest --clear' to clear the register."
74
+ )
75
+ return 1
76
+ csv_path = Path(args.csv)
77
+ if not csv_path.is_file():
78
+ logger.error("CSV file not found: %s", csv_path)
79
+ return 1
80
+ try:
81
+ risks = load_risks_from_csv(csv_path)
82
+ except Exception as exc:
83
+ logger.error("Failed to read CSV %s: %s", csv_path, exc)
84
+ return 1
85
+ if not risks:
86
+ logger.error(
87
+ "No valid risk rows found in %s (expected columns: ID, Risk level, Impact, Name, Risk description)",
88
+ csv_path,
89
+ )
90
+ return 1
91
+ try:
92
+ if args.clear:
93
+ clear_risks()
94
+ logger.info("Cleared the existing risk register")
95
+ count = ingest_risks(risks)
96
+ logger.info("Ingested %d risks into the risk register", count)
97
+ return 0
98
+ except Exception as exc:
99
+ logger.error("Risk register unavailable: %s", exc)
100
+ return 1
101
+
102
+
103
+ class SearchPlugin(Plugin):
104
+ name = "search"
105
+ distribution = "mongo-x-ray-risk"
106
+ help = "Search the risk register for risks by name"
107
+ description = """
108
+ Search the risk register for risks whose Name contains the given string
109
+ (case-insensitive substring match). Use it to check whether a newly
110
+ suggested risk is already known before adding it to the register.
111
+
112
+ For every matching risk the Name and Risk description are printed.
113
+
114
+ Exits 0 when at least one risk matches, 1 when nothing matches (so callers
115
+ can tell "already known" from "new risk").
116
+ """
117
+ epilog = """
118
+ Examples:
119
+ x-ray search "Replication Lag"
120
+ x-ray search replication
121
+ """
122
+
123
+ def add_arguments(self, parser):
124
+ parser.add_argument("string", help="Risk name (or part of one) to search for.")
125
+
126
+ def run(self, args) -> int:
127
+ """Search the register by risk name and print matching risks."""
128
+ query = (args.string or "").strip()
129
+ if not query:
130
+ logger.error("No search string given. Use 'x-ray search <string>'.")
131
+ return 1
132
+ try:
133
+ matches = find_risks_by_name(query)
134
+ except Exception as exc:
135
+ logger.error("Failed to search the risk register: %s", exc)
136
+ return 1
137
+ if not matches:
138
+ logger.info("No risks found matching %r", query)
139
+ return 1
140
+ for risk in matches:
141
+ print(f"Name: {risk['name']}")
142
+ print(f"Risk description: {risk['description']}")
143
+ print()
144
+ return 0
@@ -0,0 +1,78 @@
1
+ """
2
+ Copyright (c) 2026 MongoDB Inc.
3
+
4
+ DISCLAIMER: THESE CODE SAMPLES ARE PROVIDED FOR EDUCATIONAL AND ILLUSTRATIVE PURPOSES ONLY,
5
+ TO DEMONSTRATE THE FUNCTIONALITY OF SPECIFIC MONGODB FEATURES.
6
+ THEY ARE NOT PRODUCTION-READY AND MAY LACK THE SECURITY HARDENING, ERROR HANDLING, AND TESTING REQUIRED FOR A LIVE ENVIRONMENT.
7
+ YOU ARE RESPONSIBLE FOR TESTING, VALIDATING, AND SECURING THIS CODE WITHIN YOUR OWN ENVIRONMENT BEFORE IMPLEMENTATION.
8
+ THIS MATERIAL IS PROVIDED "AS IS" WITHOUT WARRANTY OR LIABILITY.
9
+
10
+ Shared constants and data model for the Risk Register module.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ import csv
16
+ from dataclasses import dataclass
17
+ from pathlib import Path
18
+
19
+ # Embed the Risk Name and Risk Description fields for vector search.
20
+ # Each field lives in its own collection so matching can fall back from
21
+ # Name to Risk Description.
22
+ CHROMA_COLLECTION = "risk_register"
23
+ CHROMA_COLLECTION_DESCRIPTION = "risk_register_description"
24
+ EMBED_FIELDS = ("Name", "Risk Description")
25
+
26
+
27
+ @dataclass
28
+ class Risk:
29
+ """A single risk entry from the CSV risk register."""
30
+
31
+ id: str
32
+ risk_level: str
33
+ impact: str
34
+ name: str
35
+ description: str
36
+
37
+
38
+ def _normalize_header(header: str) -> str:
39
+ """Normalize a CSV header for case/whitespace-insensitive matching."""
40
+ return (header or "").strip().lower()
41
+
42
+
43
+ def load_risks_from_csv(csv_path: Path) -> list[Risk]:
44
+ """Parse a CSV risk register file.
45
+
46
+ Used columns (matched case-insensitively; any other columns, e.g.
47
+ ``Other Notes``, are ignored):
48
+ ID, Risk level, Impact, Name, Risk description
49
+
50
+ Rows without an ID or a Name are skipped.
51
+ """
52
+ risks: list[Risk] = []
53
+ with open(csv_path, newline="", encoding="utf-8-sig") as fh:
54
+ reader = csv.DictReader(fh)
55
+ headers = {_normalize_header(h): h for h in (reader.fieldnames or [])}
56
+ for row in reader:
57
+ risk = Risk(
58
+ id=(row.get(headers.get("id", "")) or "").strip(),
59
+ risk_level=(row.get(headers.get("risk level", "")) or "").strip(),
60
+ impact=(row.get(headers.get("impact", "")) or "").strip(),
61
+ name=(row.get(headers.get("name", "")) or "").strip(),
62
+ description=(row.get(headers.get("risk description", "")) or "").strip(),
63
+ )
64
+ if risk.id and risk.name:
65
+ risks.append(risk)
66
+ return risks
67
+
68
+
69
+ def get_db_path() -> Path:
70
+ """Return the platform-specific database directory path."""
71
+ import platform
72
+
73
+ system = platform.system()
74
+ if system == "Windows":
75
+ base = Path.home() / "AppData" / "Roaming"
76
+ else:
77
+ base = Path.home()
78
+ return base / ".x-ray"
@@ -0,0 +1,70 @@
1
+ Metadata-Version: 2.4
2
+ Name: mongo-x-ray-risk
3
+ Version: 2.0.0
4
+ Summary: Known-risks knowledge base (ChromaDB vector search) for x-ray
5
+ Requires-Python: <4,>=3.10
6
+ Description-Content-Type: text/markdown
7
+ Requires-Dist: mongo-x-ray>=2.0.0
8
+ Requires-Dist: chromadb==1.5.9
9
+
10
+ # mongo-x-ray-risk
11
+
12
+ [![CI](https://github.com/zhangyaoxing/mongo-x-ray-risk/actions/workflows/ci.yml/badge.svg)](https://github.com/zhangyaoxing/mongo-x-ray-risk/actions/workflows/ci.yml)
13
+ [![PyPI](https://img.shields.io/pypi/v/mongo-x-ray-risk.svg)](https://pypi.org/project/mongo-x-ray-risk/)
14
+
15
+ Known-risks knowledge base for [x-ray](https://github.com/mongodb-ps/ce-mongo-x-ray): a ChromaDB-backed
16
+ vector search that matches analysis findings against known MongoDB risks.
17
+
18
+ This is an optional plugin: it ships the `ingest` command to load the risk register, the `search`
19
+ command to look risks up by name, and the analysis plugins (healthcheck, log, gmd) detect it at
20
+ runtime — when it is installed, their reports are enriched with matched risks; when it is missing,
21
+ the enrichment is silently skipped.
22
+
23
+ ## Install
24
+
25
+ ```bash
26
+ pip install mongo-x-ray mongo-x-ray-risk
27
+ ```
28
+
29
+ ## Usage
30
+
31
+ Load a risk register CSV into the ChromaDB knowledge base:
32
+
33
+ ```bash
34
+ x-ray ingest risk_register.csv
35
+ # start from an empty register, then ingest
36
+ x-ray ingest --clear risk_register.csv
37
+ # clear the register without ingesting (no CSV needed)
38
+ x-ray ingest --clear
39
+ ```
40
+
41
+ Check whether a risk is already known by searching the `Name` column
42
+ (case-insensitive substring match). Prints the `Name` and `Risk description`
43
+ of every matching risk; exits 0 when something matches, 1 when nothing does:
44
+
45
+ ```bash
46
+ x-ray search "Replication Lag"
47
+ x-ray search replication
48
+ ```
49
+
50
+ The CSV must have the columns `ID, Risk level, Impact, Name, Risk description`
51
+ (UTF-8, a BOM is tolerated; header names are matched case-insensitively). Any
52
+ other columns, such as `Other Notes`, are ignored. Rows without an ID or a
53
+ Name are skipped; entries with an existing ID are replaced. The data is
54
+ stored under `~/.x-ray/chroma`.
55
+
56
+ Once ingested, the plugin is used automatically by the other plugins — no CLI
57
+ flags needed. It also exposes a small API for tooling:
58
+
59
+ ```python
60
+ from mongo_x_ray_risk import Risk, load_risks_from_csv, ingest_risks, match_risk, enrich_test_results
61
+ ```
62
+
63
+ ## Development
64
+
65
+ Requires Python 3.10+, MongoDB 5.0 or later, and the [mongo-x-ray](https://github.com/mongodb-ps/ce-mongo-x-ray) core package.
66
+
67
+ ```bash
68
+ make unit-test # run the unit tests
69
+ make lint # ruff check + ruff format --check
70
+ ```
@@ -0,0 +1,16 @@
1
+ README.md
2
+ pyproject.toml
3
+ src/mongo_x_ray_risk/__init__.py
4
+ src/mongo_x_ray_risk/db.py
5
+ src/mongo_x_ray_risk/plugin.py
6
+ src/mongo_x_ray_risk/shared.py
7
+ src/mongo_x_ray_risk.egg-info/PKG-INFO
8
+ src/mongo_x_ray_risk.egg-info/SOURCES.txt
9
+ src/mongo_x_ray_risk.egg-info/dependency_links.txt
10
+ src/mongo_x_ray_risk.egg-info/entry_points.txt
11
+ src/mongo_x_ray_risk.egg-info/requires.txt
12
+ src/mongo_x_ray_risk.egg-info/top_level.txt
13
+ tests/test_csv_parsing.py
14
+ tests/test_ingest_plugin.py
15
+ tests/test_risk_register.py
16
+ tests/test_search_plugin.py
@@ -0,0 +1,3 @@
1
+ [mongo_x_ray.plugins]
2
+ ingest = mongo_x_ray_risk.plugin:IngestPlugin
3
+ search = mongo_x_ray_risk.plugin:SearchPlugin
@@ -0,0 +1,2 @@
1
+ mongo-x-ray>=2.0.0
2
+ chromadb==1.5.9
@@ -0,0 +1 @@
1
+ mongo_x_ray_risk
@@ -0,0 +1,76 @@
1
+ """
2
+ Copyright (c) 2026 MongoDB Inc.
3
+
4
+ DISCLAIMER: THESE CODE SAMPLES ARE PROVIDED FOR EDUCATIONAL AND ILLUSTRATIVE PURPOSES ONLY,
5
+ TO DEMONSTRATE THE FUNCTIONALITY OF SPECIFIC MONGODB FEATURES.
6
+ THEY ARE NOT PRODUCTION-READY AND MAY LACK THE SECURITY HARDENING, ERROR HANDLING, AND TESTING REQUIRED FOR A LIVE ENVIRONMENT.
7
+ YOU ARE RESPONSIBLE FOR TESTING, VALIDATING, AND SECURING THIS CODE WITHIN YOUR OWN ENVIRONMENT BEFORE IMPLEMENTATION.
8
+ THIS MATERIAL IS PROVIDED "AS IS" WITHOUT WARRANTY OR LIABILITY.
9
+
10
+ Tests for CSV risk register parsing.
11
+ """
12
+
13
+ from mongo_x_ray_risk.shared import load_risks_from_csv
14
+
15
+
16
+ def _write(tmp_path, text, name="risks.csv"):
17
+ path = tmp_path / name
18
+ path.write_text(text, encoding="utf-8")
19
+ return path
20
+
21
+
22
+ def test_parses_new_column_names_and_ignores_other_notes(tmp_path):
23
+ csv_path = _write(
24
+ tmp_path,
25
+ "ID,Risk level,Impact,Name,Risk description,Other Notes\n"
26
+ "R1,High,Medium,Replication Lag,oplog falls behind,internal note\n"
27
+ "R2,Medium,Low,Missing Index,no matching index,\n",
28
+ )
29
+ risks = load_risks_from_csv(csv_path)
30
+ assert [r.id for r in risks] == ["R1", "R2"]
31
+ assert risks[0].risk_level == "High"
32
+ assert risks[0].impact == "Medium"
33
+ assert risks[0].name == "Replication Lag"
34
+ assert risks[0].description == "oplog falls behind"
35
+ assert risks[1].description == "no matching index"
36
+
37
+
38
+ def test_accepts_old_casing_of_headers(tmp_path):
39
+ csv_path = _write(
40
+ tmp_path,
41
+ "ID,Risk Level,Impact,Name,Risk Description\nR1,High,Medium,Replication Lag,oplog falls behind\n",
42
+ )
43
+ risks = load_risks_from_csv(csv_path)
44
+ assert risks[0].risk_level == "High"
45
+ assert risks[0].description == "oplog falls behind"
46
+
47
+
48
+ def test_ignores_unknown_columns(tmp_path):
49
+ csv_path = _write(
50
+ tmp_path,
51
+ "ID,Risk level,Impact,Name,Risk description,Random field,Owner\n"
52
+ "R1,High,Medium,Replication Lag,oplog falls behind,whatever,alice\n",
53
+ )
54
+ risks = load_risks_from_csv(csv_path)
55
+ assert len(risks) == 1
56
+ assert risks[0].id == "R1"
57
+
58
+
59
+ def test_skips_rows_without_id_or_name(tmp_path):
60
+ csv_path = _write(
61
+ tmp_path,
62
+ "ID,Risk level,Impact,Name,Risk description\n"
63
+ "R1,High,Medium,Replication Lag,oplog falls behind\n"
64
+ ",High,Medium,No Id,has id? no\n"
65
+ "R3,Low,Low,,no name\n",
66
+ )
67
+ risks = load_risks_from_csv(csv_path)
68
+ assert [r.id for r in risks] == ["R1"]
69
+
70
+
71
+ def test_handles_bom_utf8(tmp_path):
72
+ csv_path = tmp_path / "bom.csv"
73
+ csv_path.write_bytes(b"\xef\xbb\xbfID,Risk level,Impact,Name,Risk description\nR1,Low,Low,Backup Failure,failing\n")
74
+ risks = load_risks_from_csv(csv_path)
75
+ assert risks[0].id == "R1"
76
+ assert risks[0].name == "Backup Failure"
@@ -0,0 +1,102 @@
1
+ """
2
+ Copyright (c) 2026 MongoDB Inc.
3
+
4
+ DISCLAIMER: THESE CODE SAMPLES ARE PROVIDED FOR EDUCATIONAL AND ILLUSTRATIVE PURPOSES ONLY,
5
+ TO DEMONSTRATE THE FUNCTIONALITY OF SPECIFIC MONGODB FEATURES.
6
+ THEY ARE NOT PRODUCTION-READY AND MAY LACK THE SECURITY HARDENING, ERROR HANDLING, AND TESTING REQUIRED FOR A LIVE ENVIRONMENT.
7
+ YOU ARE RESPONSIBLE FOR TESTING, VALIDATING, AND SECURING THIS CODE WITHIN YOUR OWN ENVIRONMENT BEFORE IMPLEMENTATION.
8
+ THIS MATERIAL IS PROVIDED "AS IS" WITHOUT WARRANTY OR LIABILITY.
9
+
10
+ Tests for the risk register ``ingest`` command plugin.
11
+ """
12
+
13
+ import argparse
14
+
15
+ from mongo_x_ray_risk.plugin import IngestPlugin
16
+
17
+
18
+ def _args(**overrides) -> argparse.Namespace:
19
+ ns = argparse.Namespace(csv=None, clear=False)
20
+ ns.__dict__.update(overrides)
21
+ return ns
22
+
23
+
24
+ def test_plugin_metadata():
25
+ assert IngestPlugin.name == "ingest"
26
+ assert IngestPlugin.distribution == "mongo-x-ray-risk"
27
+ assert "CSV" in IngestPlugin.help
28
+ assert "ID, Risk level, Impact, Name, Risk description" in IngestPlugin.description
29
+ assert "Other Notes" in IngestPlugin.description
30
+ assert "x-ray ingest risk_register.csv" in IngestPlugin.epilog
31
+
32
+
33
+ def test_run_ingests_csv_rows(monkeypatch, tmp_path):
34
+ csv_path = tmp_path / "risks.csv"
35
+ csv_path.write_text(
36
+ "ID,Risk level,Impact,Name,Risk description,Other Notes\n"
37
+ "R1,High,Medium,Replication Lag,oplog falls behind,internal note\n"
38
+ "R2,Medium,Low,Missing Index,no matching index,\n"
39
+ ",,Low,No Id Here,skipped,note\n",
40
+ encoding="utf-8",
41
+ )
42
+ seen = {}
43
+
44
+ def fake_ingest(risks):
45
+ seen["risks"] = risks
46
+ return len(risks)
47
+
48
+ monkeypatch.setattr("mongo_x_ray_risk.plugin.ingest_risks", fake_ingest)
49
+
50
+ assert IngestPlugin().run(_args(csv=str(csv_path))) == 0
51
+ assert [r.id for r in seen["risks"]] == ["R1", "R2"]
52
+ assert seen["risks"][0].name == "Replication Lag"
53
+ assert seen["risks"][0].risk_level == "High"
54
+ assert seen["risks"][0].impact == "Medium"
55
+ assert seen["risks"][0].description == "oplog falls behind"
56
+
57
+
58
+ def test_run_clear_empties_register_first(monkeypatch, tmp_path):
59
+ csv_path = tmp_path / "risks.csv"
60
+ csv_path.write_text("ID,Risk level,Impact,Name,Risk description\nR1,Low,Low,One Risk,x\n", encoding="utf-8")
61
+ cleared = []
62
+ monkeypatch.setattr("mongo_x_ray_risk.plugin.clear_risks", lambda: cleared.append(True))
63
+ monkeypatch.setattr("mongo_x_ray_risk.plugin.ingest_risks", len)
64
+
65
+ assert IngestPlugin().run(_args(csv=str(csv_path), clear=True)) == 0
66
+ assert cleared == [True]
67
+
68
+
69
+ def test_run_clear_without_csv_only_clears(monkeypatch):
70
+ cleared = []
71
+ ingest_called = []
72
+ monkeypatch.setattr("mongo_x_ray_risk.plugin.clear_risks", lambda: cleared.append(True))
73
+ monkeypatch.setattr("mongo_x_ray_risk.plugin.ingest_risks", lambda risks: ingest_called.append(risks))
74
+
75
+ assert IngestPlugin().run(_args(clear=True)) == 0
76
+ assert cleared == [True]
77
+ assert ingest_called == []
78
+
79
+
80
+ def test_run_without_csv_or_clear_returns_error():
81
+ assert IngestPlugin().run(_args()) == 1
82
+
83
+
84
+ def test_run_missing_csv_returns_error(tmp_path):
85
+ assert IngestPlugin().run(_args(csv=str(tmp_path / "nope.csv"))) == 1
86
+
87
+
88
+ def test_run_csv_without_valid_rows_returns_error(tmp_path):
89
+ csv_path = tmp_path / "empty.csv"
90
+ csv_path.write_text("ID,Risk level,Impact,Name,Risk description\n", encoding="utf-8")
91
+ assert IngestPlugin().run(_args(csv=str(csv_path))) == 1
92
+
93
+
94
+ def test_run_unreadable_csv_returns_error(monkeypatch, tmp_path):
95
+ csv_path = tmp_path / "broken.csv"
96
+ csv_path.write_text("x\n", encoding="utf-8")
97
+
98
+ def boom(_path):
99
+ raise OSError("boom")
100
+
101
+ monkeypatch.setattr("mongo_x_ray_risk.plugin.load_risks_from_csv", boom)
102
+ assert IngestPlugin().run(_args(csv=str(csv_path))) == 1
@@ -0,0 +1,245 @@
1
+ """
2
+ Copyright (c) 2026 MongoDB Inc.
3
+
4
+ DISCLAIMER: THESE CODE SAMPLES ARE PROVIDED FOR EDUCATIONAL AND ILLUSTRATIVE PURPOSES ONLY,
5
+ TO DEMONSTRATE THE FUNCTIONALITY OF SPECIFIC MONGODB FEATURES.
6
+ THEY ARE NOT PRODUCTION-READY AND MAY LACK THE SECURITY HARDENING, ERROR HANDLING, AND TESTING REQUIRED FOR A LIVE ENVIRONMENT.
7
+ YOU ARE RESPONSIBLE FOR TESTING, VALIDATING, AND SECURING THIS CODE WITHIN YOUR OWN ENVIRONMENT BEFORE IMPLEMENTATION.
8
+ THIS MATERIAL IS PROVIDED "AS IS" WITHOUT WARRANTY OR LIABILITY.
9
+ """
10
+
11
+ import pytest
12
+
13
+ from mongo_x_ray_risk import db
14
+ from mongo_x_ray_risk.shared import (
15
+ CHROMA_COLLECTION,
16
+ CHROMA_COLLECTION_DESCRIPTION,
17
+ Risk,
18
+ )
19
+
20
+ RISK_R1 = {
21
+ "id": "R1",
22
+ "risk_level": "High",
23
+ "impact": "Medium",
24
+ "name": "Replication Lag",
25
+ "description": "oplog falls behind",
26
+ "distance": 0.1,
27
+ }
28
+
29
+
30
+ def test_match_risk_uses_name_collection_first(monkeypatch):
31
+ calls = []
32
+
33
+ def fake_search(query, n_results=3, collection_name=CHROMA_COLLECTION):
34
+ calls.append((query, n_results, collection_name))
35
+ return [dict(RISK_R1)]
36
+
37
+ monkeypatch.setattr(db, "search_risks", fake_search)
38
+
39
+ risk = db.match_risk("Replication Lag")
40
+
41
+ assert risk == RISK_R1
42
+ assert calls == [("Replication Lag", 1, CHROMA_COLLECTION)]
43
+
44
+
45
+ def test_match_risk_falls_back_to_description_with_same_query(monkeypatch):
46
+ calls = []
47
+
48
+ def fake_search(query, n_results=3, collection_name=CHROMA_COLLECTION):
49
+ calls.append((query, n_results, collection_name))
50
+ if collection_name == CHROMA_COLLECTION:
51
+ return []
52
+ return [dict(RISK_R1)]
53
+
54
+ monkeypatch.setattr(db, "search_risks", fake_search)
55
+
56
+ risk = db.match_risk("Unrelated Topic")
57
+
58
+ assert risk == RISK_R1
59
+ assert calls == [
60
+ ("Unrelated Topic", 1, CHROMA_COLLECTION),
61
+ ("Unrelated Topic", 1, CHROMA_COLLECTION_DESCRIPTION),
62
+ ]
63
+
64
+
65
+ def test_match_risk_falls_back_when_name_match_is_too_far(monkeypatch):
66
+ calls = []
67
+
68
+ def fake_search(query, n_results=3, collection_name=CHROMA_COLLECTION):
69
+ calls.append((query, n_results, collection_name))
70
+ far = dict(RISK_R1)
71
+ far["distance"] = 2.0
72
+ return [far] if collection_name == CHROMA_COLLECTION else [dict(RISK_R1)]
73
+
74
+ monkeypatch.setattr(db, "search_risks", fake_search)
75
+
76
+ risk = db.match_risk("Replication Lag", max_distance=1.0)
77
+
78
+ assert risk == RISK_R1
79
+ assert len(calls) == 2
80
+
81
+
82
+ def test_match_risk_returns_none_when_both_stages_miss(monkeypatch):
83
+ def fake_search(query, n_results=3, collection_name=CHROMA_COLLECTION):
84
+ return []
85
+
86
+ monkeypatch.setattr(db, "search_risks", fake_search)
87
+
88
+ assert db.match_risk("Unrelated Topic") is None
89
+
90
+
91
+ def test_has_risks_reflects_collection_count(monkeypatch):
92
+ monkeypatch.setattr(db, "_collection_count", lambda: 3)
93
+ assert db.has_risks() is True
94
+ monkeypatch.setattr(db, "_collection_count", lambda: 0)
95
+ assert db.has_risks() is False
96
+
97
+
98
+ def test_has_risks_returns_false_when_chromadb_unavailable(monkeypatch):
99
+ def boom():
100
+ raise RuntimeError("chromadb unavailable")
101
+
102
+ monkeypatch.setattr(db, "_collection_count", boom)
103
+ assert db.has_risks() is False
104
+
105
+
106
+ def test_enrich_test_results_matches_by_title(monkeypatch):
107
+ monkeypatch.setattr(db, "_collection_count", lambda: 1)
108
+ captured = {}
109
+
110
+ def fake_match_risk(title, max_distance=1.0):
111
+ captured["title"] = title
112
+ return dict(RISK_R1)
113
+
114
+ monkeypatch.setattr(db, "match_risk", fake_match_risk)
115
+
116
+ results = [{"host": "h", "severity": "High", "title": "Replication Lag", "message": "secondary oplog behind"}]
117
+ matched = db.enrich_test_results(results)
118
+
119
+ assert matched == 1
120
+ assert results[0]["matched_risk"] == RISK_R1
121
+ assert captured == {"title": "Replication Lag"}
122
+
123
+
124
+ class _FakeNameCollection:
125
+ """Minimal stand-in for a Chroma collection of risk names."""
126
+
127
+ def __init__(self, metadatas):
128
+ self._metadatas = metadatas
129
+
130
+ def get(self, include=None):
131
+ return {"ids": [m["id"] for m in self._metadatas], "metadatas": self._metadatas}
132
+
133
+
134
+ def _install_fake_collection(monkeypatch, metadatas):
135
+ monkeypatch.setattr(db, "_collection", lambda name: _FakeNameCollection(metadatas))
136
+
137
+
138
+ def test_find_risks_by_name_matches_substring_case_insensitive(monkeypatch):
139
+ _install_fake_collection(
140
+ monkeypatch,
141
+ [
142
+ {
143
+ "id": "R1",
144
+ "risk_level": "High",
145
+ "impact": "Medium",
146
+ "name": "Replication Lag",
147
+ "description": "oplog falls behind",
148
+ },
149
+ {
150
+ "id": "R2",
151
+ "risk_level": "Medium",
152
+ "impact": "Low",
153
+ "name": "Missing Index",
154
+ "description": "no matching index",
155
+ },
156
+ ],
157
+ )
158
+ hits = db.find_risks_by_name("replication")
159
+ assert [h["id"] for h in hits] == ["R1"]
160
+ assert hits[0]["name"] == "Replication Lag"
161
+ assert hits[0]["description"] == "oplog falls behind"
162
+
163
+
164
+ def test_find_risks_by_name_returns_all_matching(monkeypatch):
165
+ _install_fake_collection(
166
+ monkeypatch,
167
+ [
168
+ {"id": "R1", "risk_level": "High", "impact": "Medium", "name": "Replication Lag", "description": "a"},
169
+ {"id": "R2", "risk_level": "Medium", "impact": "Low", "name": "Index on Replication", "description": "b"},
170
+ {"id": "R3", "risk_level": "Low", "impact": "Low", "name": "Backup Failure", "description": "c"},
171
+ ],
172
+ )
173
+ hits = db.find_risks_by_name("replication")
174
+ assert [h["id"] for h in hits] == ["R1", "R2"]
175
+
176
+
177
+ def test_find_risks_by_name_no_match_returns_empty_list(monkeypatch):
178
+ _install_fake_collection(
179
+ monkeypatch,
180
+ [
181
+ {
182
+ "id": "R1",
183
+ "risk_level": "High",
184
+ "impact": "Medium",
185
+ "name": "Replication Lag",
186
+ "description": "oplog falls behind",
187
+ }
188
+ ],
189
+ )
190
+ assert db.find_risks_by_name("Unrelated Topic XYZ") == []
191
+
192
+
193
+ def test_find_risks_by_name_empty_query_returns_empty_list(monkeypatch):
194
+ _install_fake_collection(monkeypatch, [])
195
+ assert db.find_risks_by_name(" ") == []
196
+
197
+
198
+ @pytest.mark.integration
199
+ def test_ingest_and_two_stage_search(tmp_path, monkeypatch):
200
+ monkeypatch.setattr(db, "get_db_path", lambda: tmp_path)
201
+
202
+ risks = [
203
+ Risk(
204
+ id="R1",
205
+ risk_level="High",
206
+ impact="Medium",
207
+ name="Replication Lag",
208
+ description="Replication lag occurs when the secondary's oplog application falls behind the primary.",
209
+ ),
210
+ Risk(
211
+ id="R2",
212
+ risk_level="Medium",
213
+ impact="Low",
214
+ name="Missing Index",
215
+ description="Queries scan the entire collection because no matching index exists.",
216
+ ),
217
+ Risk(id="R3", risk_level="Low", impact="Low", name="Backup Failure", description=" "),
218
+ ]
219
+
220
+ assert db.ingest_risks(risks) == 3
221
+
222
+ name_col = db._collection(CHROMA_COLLECTION)
223
+ desc_col = db._collection(CHROMA_COLLECTION_DESCRIPTION)
224
+ assert name_col.count() == 3
225
+ assert desc_col.count() == 2 # R3 has no Risk Description
226
+
227
+ # Stage 1: title matches risk Name.
228
+ title_risk = db.match_risk("Replication Lag")
229
+ assert title_risk is not None
230
+ assert title_risk["id"] == "R1"
231
+ index_risk = db.match_risk("Missing Index")
232
+ assert index_risk is not None
233
+ assert index_risk["id"] == "R2"
234
+
235
+ # Stage 2: Name misses, the same title matches Risk Description.
236
+ risk = db.match_risk("the secondary's oplog application falls behind the primary")
237
+ assert risk is not None and risk["id"] == "R1"
238
+
239
+ # No match at either stage.
240
+ assert db.match_risk("Unrelated Topic XYZ") is None
241
+
242
+ # clear_risks empties both collections.
243
+ db.clear_risks()
244
+ assert name_col.count() == 0
245
+ assert desc_col.count() == 0
@@ -0,0 +1,79 @@
1
+ """
2
+ Copyright (c) 2026 MongoDB Inc.
3
+
4
+ DISCLAIMER: THESE CODE SAMPLES ARE PROVIDED FOR EDUCATIONAL AND ILLUSTRATIVE PURPOSES ONLY,
5
+ TO DEMONSTRATE THE FUNCTIONALITY OF SPECIFIC MONGODB FEATURES.
6
+ THEY ARE NOT PRODUCTION-READY AND MAY LACK THE SECURITY HARDENING, ERROR HANDLING, AND TESTING REQUIRED FOR A LIVE ENVIRONMENT.
7
+ YOU ARE RESPONSIBLE FOR TESTING, VALIDATING, AND SECURING THIS CODE WITHIN YOUR OWN ENVIRONMENT BEFORE IMPLEMENTATION.
8
+ THIS MATERIAL IS PROVIDED "AS IS" WITHOUT WARRANTY OR LIABILITY.
9
+
10
+ Tests for the risk register ``search`` command plugin.
11
+ """
12
+
13
+ import argparse
14
+
15
+ from mongo_x_ray_risk.plugin import SearchPlugin
16
+
17
+
18
+ def _args(**overrides) -> argparse.Namespace:
19
+ ns = argparse.Namespace(string=None)
20
+ ns.__dict__.update(overrides)
21
+ return ns
22
+
23
+
24
+ def test_plugin_metadata():
25
+ assert SearchPlugin.name == "search"
26
+ assert SearchPlugin.distribution == "mongo-x-ray-risk"
27
+ assert "risk register" in SearchPlugin.help
28
+ assert "case-insensitive" in SearchPlugin.description
29
+ assert "x-ray search" in SearchPlugin.epilog
30
+
31
+
32
+ def test_run_prints_matching_risks(monkeypatch, capsys):
33
+ matches = [
34
+ {
35
+ "id": "R1",
36
+ "risk_level": "High",
37
+ "impact": "Medium",
38
+ "name": "Replication Lag",
39
+ "description": "oplog falls behind",
40
+ },
41
+ {
42
+ "id": "R2",
43
+ "risk_level": "Medium",
44
+ "impact": "Low",
45
+ "name": "Missing Index",
46
+ "description": "no matching index",
47
+ },
48
+ ]
49
+ seen = {}
50
+
51
+ def fake_find(query):
52
+ seen["query"] = query
53
+ return matches
54
+
55
+ monkeypatch.setattr("mongo_x_ray_risk.plugin.find_risks_by_name", fake_find)
56
+
57
+ assert SearchPlugin().run(_args(string="Replication")) == 0
58
+ assert seen == {"query": "Replication"}
59
+ out = capsys.readouterr().out
60
+ assert "Name: Replication Lag" in out
61
+ assert "Risk description: oplog falls behind" in out
62
+ assert "Name: Missing Index" in out
63
+
64
+
65
+ def test_run_no_match_returns_1(monkeypatch):
66
+ monkeypatch.setattr("mongo_x_ray_risk.plugin.find_risks_by_name", lambda query: [])
67
+ assert SearchPlugin().run(_args(string="Nothing Here")) == 1
68
+
69
+
70
+ def test_run_empty_string_returns_error():
71
+ assert SearchPlugin().run(_args(string=" ")) == 1
72
+
73
+
74
+ def test_run_search_failure_returns_error(monkeypatch):
75
+ def boom(_query):
76
+ raise RuntimeError("chromadb unavailable")
77
+
78
+ monkeypatch.setattr("mongo_x_ray_risk.plugin.find_risks_by_name", boom)
79
+ assert SearchPlugin().run(_args(string="Replication")) == 1