codegraph-engine 2.2.0__py3-none-any.whl → 2.2.1__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,642 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: codegraph-engine
3
- Version: 2.2.0
4
- Summary: Deep deterministic repository intelligence for AI coding agents (code, dependencies, databases, and runtime observations)
5
- Author: CodeGraph contributors
6
- License-Expression: MIT
7
- Project-URL: Homepage, https://github.com/raghurammrsd/CODE_GRAPH_MCP
8
- Project-URL: Repository, https://github.com/raghurammrsd/CODE_GRAPH_MCP
9
- Project-URL: Documentation, https://github.com/raghurammrsd/CODE_GRAPH_MCP/blob/main/docs/agent-brain.md
10
- Project-URL: Changelog, https://github.com/raghurammrsd/CODE_GRAPH_MCP/blob/main/CHANGELOG.md
11
- Project-URL: Issues, https://github.com/raghurammrsd/CODE_GRAPH_MCP/issues
12
- Classifier: Programming Language :: Python :: 3
13
- Classifier: Programming Language :: Python :: 3.12
14
- Classifier: Programming Language :: Python :: 3.13
15
- Classifier: Operating System :: OS Independent
16
- Classifier: Topic :: Software Development :: Libraries :: Python Modules
17
- Requires-Python: >=3.12
18
- Description-Content-Type: text/markdown
19
- License-File: LICENSE
20
- Requires-Dist: pydantic>=2.7
21
- Requires-Dist: typer>=0.12
22
- Requires-Dist: mcp<2,>=1.0
23
- Provides-Extra: mcp
24
- Requires-Dist: mcp<2,>=1.0; extra == "mcp"
25
- Provides-Extra: system
26
- Requires-Dist: psutil>=5.9; extra == "system"
27
- Provides-Extra: dev
28
- Requires-Dist: pytest>=8; extra == "dev"
29
- Requires-Dist: pytest-cov>=5; extra == "dev"
30
- Requires-Dist: ruff>=0.6; extra == "dev"
31
- Requires-Dist: mypy>=1.10; extra == "dev"
32
- Dynamic: license-file
33
-
34
- <p align="center">
35
- <img src="docs/assets/codegraph_logo.jpg" alt="CodeGraph MCP — Deep Deterministic Repository Intelligence for AI Coding Agents" width="500" />
36
- </p>
37
-
38
- <h1 align="center">CodeGraph-MCP Engine (v2.2.0)</h1>
39
-
40
- <p align="center">
41
- <strong>Deep deterministic repository intelligence for AI coding agents.</strong>
42
- </p>
43
-
44
- <p align="center">
45
- <a href="https://pypi.org/project/codegraph-engine/2.2.0/"><img src="https://img.shields.io/badge/pypi-codegraph--engine%20v2.2.0-blue.svg" alt="PyPI: codegraph-engine v2.2.0" /></a>
46
- <a href="pyproject.toml"><img src="https://img.shields.io/badge/python-3.12%20%7C%203.13-3776AB.svg" alt="Python 3.12 | 3.13" /></a>
47
- <a href="src/codegraph/mcp/server.py"><img src="https://img.shields.io/badge/MCP-14%20default%20%7C%2056%20full%20tools-2ea043.svg" alt="MCP Tools: 14 default | 56 full" /></a>
48
- <a href="tests/"><img src="https://img.shields.io/badge/pytest-861%20passed-brightgreen.svg" alt="Tests: 861 passed" /></a>
49
- <a href="pyproject.toml"><img src="https://img.shields.io/badge/ruff-0%20errors-success.svg" alt="Ruff: 0 errors" /></a>
50
- <a href="src/codegraph/"><img src="https://img.shields.io/badge/mypy-0%20issues%20(80%20files)-blue.svg" alt="Mypy: strict" /></a>
51
- <a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-yellow.svg" alt="License: MIT" /></a>
52
- </p>
53
-
54
- <p align="center">
55
- <a href="https://pypi.org/project/codegraph-engine/2.2.0/"><strong>PyPI (v2.2.0)</strong></a> •
56
- <a href="#2-quickstart-30-second-setup"><strong>Quickstart</strong></a> •
57
- <a href="#1-built-for-aiml-and-backend-heavy-repositories"><strong>Built For</strong></a> •
58
- <a href="#5-database-intelligence"><strong>Database Intelligence</strong></a> •
59
- <a href="#6-runtime-intelligence--static-reconciliation"><strong>Runtime Evidence</strong></a> •
60
- <a href="#8-measured-performance-v216--v217"><strong>Measured Performance</strong></a> •
61
- <a href="docs/agent-brain.md"><strong>56-Tool Reference</strong></a> •
62
- <a href="https://github.com/raghurammrsd/CODE_GRAPH_MCP"><strong>GitHub</strong></a>
63
- </p>
64
-
65
- ```text
66
- The AI reasons.
67
- CodeGraph interrogates the repository.
68
- The evidence stays traceable.
69
- ```
70
-
71
- **CodeGraph MCP gives AI coding agents an evidence-backed understanding of code, dependencies, databases, and optional runtime observations.**
72
- > **Supported Languages & Frameworks:** First-class **Python** (`FastAPI`, `Flask`, `Django`, `SQLAlchemy`, `Celery`, `pytest`) + **TypeScript / JavaScript** (`.ts`, `.tsx`, `.js`, `.jsx`) & **Express.js** support.
73
-
74
- When an AI coding agent works inside a complex Python or full-stack codebase, raw text search forces it to open dozens of files and mentally reconstruct call chains, router prefixes, dependency injection providers, and ORM table mappings inside its context window.
75
-
76
- ```text
77
- Complex repository
78
- ↓
79
- AI agent needs architectural & dataflow understanding
80
- ↓
81
- CodeGraph interrogates the local repository index
82
- ↓
83
- Compact, cited evidence (symbols, edges, tables, bounded slices)
84
- ↓
85
- AI reasons and edits with traceable citations
86
- ```
87
-
88
- ### 30-Second Install
89
-
90
- ```bash
91
- pip install "codegraph-engine[mcp]"
92
- codegraph install
93
- cd your-project
94
- codegraph init
95
- codegraph doctor
96
- ```
97
-
98
- ---
99
-
100
- ## 1. Built for AI/ML and Backend-Heavy Repositories
101
-
102
- CodeGraph is engineered for **AI/ML engineers**, **LLM application developers**, **model/inference engineers**, **backend Python & TypeScript/JS developers**, and **maintainers of large multi-package repositories** where relationships cross module, framework, and database boundaries.
103
-
104
- ### AI/ML & LLM Engineering Workloads
105
- - **Inference & Model Serving Services**: Trace HTTP/RPC routes (`FastAPI`, `Flask`) into inference handlers, request validators, preprocessing pipelines, model forward calls, postprocessing, and database/cache persistence.
106
- - **Training & Evaluation Pipelines**: Map training entrypoints (`CLI` commands, scripts) to dataset loaders, feature transforms, trainer loops, checkpoint writers, and evaluation metrics.
107
- - **LLM Applications & Tool Registries**: Resolve decorator and call-based tool/agent registries (`@register`, `register_tool`, `ROUTING_MANIFEST`), prompt/context builders, retrieval pipelines, and model provider clients.
108
- - **Experiment & Monorepo Codebases**: Distinguish active source code (`SOURCE`) from generated protobuf/OpenAPI stubs (`GENERATED`), build outputs (`BUILD_ARTIFACT`), and vendor directories (`VENDOR`).
109
-
110
- ### Backend-Heavy Python, TypeScript/JS & Service Architectures
111
- - **Web Frameworks**: **FastAPI**, **Flask**, **Django**, and **Express.js** route registration (`ROUTE_HANDLER`, `HANDLED_BY`, `ROUTES_TO`) and nested router prefix composition (`MOUNTS` via `include_router`, `register_blueprint`, `app.use`).
112
- - **Multi-Language Full-Stack Indexing**: Deep **Python** AST & dataflow analysis alongside **TypeScript** (`.ts`, `.tsx`) and **JavaScript** (`.js`, `.jsx`) symbol, import, call, and **Express.js** route extraction.
113
- - **Dependency Injection & Event Systems**: FastAPI `Depends(...)` (`INJECTS`, `PROVIDES`, `RESOLVES_DEPENDENCY`, `DI_CYCLE`), event buses (`EVENT_LISTENER`, `DISPATCHES_TO`), and **Celery** background task queues (`TASK_HANDLER`).
114
- - **Persistence & ORM Layers**: **SQLAlchemy**, **Django ORM**, **SQLModel**, **Prisma**, **Alembic**, **Django Migrations**, and raw **SQL** (`PostgreSQL`, `MySQL`, `SQLite`) table/column read-write analysis.
115
- - **Test Suites**: Link **pytest** and `unittest` test functions and fixtures directly to the symbols, routes, DI providers, and event handlers they verify (`TESTS_SYMBOL`, `TESTS_ROUTE`, `TESTS_PROVIDER`, `TESTS_EVENT_HANDLER`).
116
-
117
- ---
118
-
119
- ## 2. Quickstart (30-Second Setup)
120
-
121
- ### Step 1: Install the Package
122
-
123
- ```bash
124
- pip install "codegraph-engine[mcp]"
125
- ```
126
-
127
- ### Step 2: Configure Your AI Coding Agents (`codegraph install`)
128
-
129
- CodeGraph-MCP includes an interactive, idempotent onboarding installer ([`src/codegraph/installer.py`](src/codegraph/installer.py)) that detects installed AI coding agents (**Claude Code**, **Cursor**, **Antigravity**, **Codex CLI**, **Gemini CLI**, and **Cline**), configures `mcpServers.codegraph`, installs marker-bounded routing instructions (`<!-- CODEGRAPH:START -->` … `<!-- CODEGRAPH:END -->`), and verifies MCP server startup:
130
-
131
- ```bash
132
- # Interactive setup (detects installed agents, previews planned changes, asks confirmation)
133
- codegraph install
134
-
135
- # Non-interactive setup for all detected agents in the current project
136
- codegraph install --yes --target auto --location local
137
-
138
- # Preview exact file modifications without writing anything
139
- codegraph install --dry-run
140
- ```
141
-
142
- ### Step 3: Initialize & Verify Your Project Index
143
-
144
- ```bash
145
- cd your-project
146
- codegraph init
147
- codegraph status
148
- codegraph doctor
149
- ```
150
-
151
- > **Dynamic Index Hot-Reloading:** If your AI coding agent launches `codegraph mcp serve` before you run `codegraph init`, you do **not** need to restart your IDE or MCP session. As soon as `codegraph init` or `codegraph index` creates or updates `.codegraph.sqlite3`, the running MCP server automatically detects the updated database on the next tool call and refreshes its in-memory graph cache (`get_repository_status(reload=True, reindex=True)` can also be invoked directly by the agent).
152
-
153
- ### Managing Installation, Process & Project Index Lifecycle
154
-
155
- CodeGraph cleanly separates agent configuration, process lifecycle, and project index files:
156
-
157
- | Command | Scope | What It Does |
158
- | :--- | :--- | :--- |
159
- | `codegraph install` | Agent configuration | Configures MCP server + marker-managed rules/skills for selected AI agents (ASCII-safe on Windows `cp1252`). |
160
- | `codegraph init` | Project repository | Initializes `.codegraph.sqlite3` and indexes the current repository (hot-reloaded automatically by running MCP servers). |
161
- | `codegraph index` | Project repository | Incrementally indexes modified files (`--verbose`, `--quiet`, `--json`). |
162
- | `codegraph stop` | Process lifecycle | Safely stops running CodeGraph-owned MCP background processes (`--all`, `--repo`, `--json`) and releases file locks. |
163
- | `codegraph doctor --processes` | Process lifecycle | Lists active CodeGraph MCP processes, parent PIDs, orphan status, and cleans stale PID files. |
164
- | `codegraph uninstall` | Agent configuration | Stops active workspace MCP processes and removes CodeGraph-managed MCP entries and instruction blocks. |
165
- | `codegraph uninit` | Project repository | Removes only `.codegraph.sqlite3` and `.codegraph/` in the project (never touches source code or Git history). |
166
-
167
- ### Safe Upgrading on Windows (`WinError 32` Prevention)
168
-
169
- On Windows, running `pip install --upgrade codegraph-engine` while an IDE or agent still holds `codegraph.exe` open can trigger `WinError 32`. Before upgrading on Windows, run:
170
-
171
- ```bash
172
- codegraph stop --all
173
- pip install --upgrade "codegraph-engine[mcp]"
174
- ```
175
-
176
- ---
177
-
178
- ## 3. Core Workflow & Architecture
179
-
180
- ```mermaid
181
- flowchart TD
182
- A["AI Coding Agent"] --> B["CodeGraph MCP"]
183
- B --> C["Repository Intelligence"]
184
-
185
- C --> D["Semantic Graph"]
186
- C --> E["Text Search (search_code)"]
187
- C --> F["Source Inspection (get_file)"]
188
- C --> G["Database Intelligence"]
189
- C --> H["Runtime Observation"]
190
-
191
- D --> I["Structured Evidence + Epistemic Labels"]
192
- E --> I
193
- F --> I
194
- G --> I
195
- H --> I
196
-
197
- I --> J["Context Optimization + Secret Redaction"]
198
- J --> A
199
- ```
200
-
201
- ### Routing Each Question to the Right Primitive
202
-
203
- ```text
204
- STRUCTURAL → Graph tools (find_symbol, find_callers, find_callees, find_references, find_routes)
205
- TEXTUAL → search_code (literal/regex search across Python, HTML/Jinja, JS/TS, CSS, YAML/JSON, SQL)
206
- SOURCE → get_file (bounded start_line..end_line source inspection with truncation metadata)
207
- GRAPH → Trace tools (trace_path, trace_flow, analyze_impact, get_git_impact)
208
- DATABASE → Database tools (get_db_schema, find_db_tables, find_db_readers, find_db_writers, get_db_impact)
209
- RUNTIME → Runtime/reconciliation tools (ingest_runtime_traces, get_runtime_trace, reconcile_static_runtime)
210
- EDITING → Native agent / IDE editing tools
211
- ```
212
-
213
- ---
214
-
215
- ## 4. AI/ML & Backend Architecture Examples
216
-
217
- ### Example 1: Model Inference Service Flow
218
-
219
- ```text
220
- POST /v1/predict (FastAPI Route)
221
- ↓ HANDLED_BY (FRAMEWORK_VERIFIED)
222
- predict_endpoint (Handler)
223
- ↓ INJECTS (DATAFLOW_VERIFIED)
224
- get_inference_service (DI Provider)
225
- ↓ CALLS (AST_VERIFIED)
226
- InferenceService.run_inference
227
- ├──►CALLS (AST_VERIFIED) ──► FeaturePreprocessor.transform
228
- ├──►CALLS (AST_VERIFIED) ──► FraudClassifier.forward
229
- ├──►CALLS (AST_VERIFIED) ──► ScorePostprocessor.calibrate
230
- └──►CALLS (AST_VERIFIED) ──► PredictionRepository.log_prediction
231
- ↓ WRITES_TABLE (AST_VERIFIED)
232
- db:table:postgresql.public.prediction_logs
233
- ```
234
-
235
- **What CodeGraph-MCP structurally proves**:
236
- - `find_routes(path="/v1/predict")` resolves composed router prefixes (`MOUNTS`) to `predict_endpoint`.
237
- - `trace_path(from_symbol="predict_endpoint", to_symbol="log_prediction")` proves the multi-hop execution chain across DI injection, preprocessing, model execution, and persistence.
238
- - `get_db_impact(symbol="InferenceService.run_inference")` identifies downstream writes to `prediction_logs`.
239
-
240
- ### Example 2: Training & Evaluation Pipeline
241
-
242
- ```text
243
- train_cli (CLI Command Handler)
244
- ↓ COMMAND_HANDLER (FRAMEWORK_VERIFIED)
245
- TrainingPipeline.run
246
- ├──►CALLS (AST_VERIFIED) ──► DatasetBuilder.load_splits
247
- ├──►CALLS (AST_VERIFIED) ──► TokenizerTransform.encode_batch
248
- ├──►CALLS (AST_VERIFIED) ──► Trainer.fit_epoch
249
- ├──►CALLS (AST_VERIFIED) ──► CheckpointManager.save_weights
250
- └──►CALLS (AST_VERIFIED) ──► Evaluator.compute_metrics
251
- ```
252
-
253
- **What CodeGraph-MCP structurally proves**:
254
- - `find_callees(symbol="TrainingPipeline.run")` enumerates every stage of the pipeline with exact file and line ranges.
255
- - `find_tests(symbol="Evaluator.compute_metrics")` locates the unit and regression tests covering metric calculation.
256
- - When a transform or model class is dynamically instantiated from a YAML string (`getattr(models, cfg.arch)`), CodeGraph explicitly records `POSSIBLE_CALLS` (`POSSIBLE`) or `UNRESOLVED_REFERENCE` (`UNKNOWN`) rather than fabricating a false static call edge.
257
-
258
- ### Example 3: LLM Agent & Tool Registry
259
-
260
- ```text
261
- POST /api/chat (LLM Endpoint)
262
- ↓ HANDLED_BY (FRAMEWORK_VERIFIED)
263
- chat_handler
264
- ↓ CALLS (AST_VERIFIED)
265
- AgentRunner.execute_step
266
- ├──►REGISTERS / REGISTERED_HANDLER ──► ToolRegistry ("search_orders", "refund_order")
267
- ├──►CALLS (AST_VERIFIED) ──► ContextBuilder.compile
268
- ├──►READS_TABLE (AST_VERIFIED) ──► db:table:postgresql.public.conversations
269
- └──►CALLS (AST_VERIFIED) ──► ModelProviderClient.generate
270
- ```
271
-
272
- **What CodeGraph-MCP structurally proves**:
273
- - Tracks decorator and call-based registrations (`@tool_registry.register("search_orders")`) via `REGISTERS` and `REGISTERED_HANDLER` edges.
274
- - Tracks environment variable dependencies (`os.getenv("OPENAI_API_KEY")`) as `READS_ENV` edges with the variable name only—never indexing or exposing secret values.
275
-
276
- ### Example 4: Full-Stack Feature Investigation (Dundoo Bill Scanner)
277
-
278
- Validated end-to-end in [`src/codegraph/tool_selection_eval.py`](src/codegraph/tool_selection_eval.py) (`run_dundoo_bill_scanner_e2e_eval()`):
279
-
280
- ```text
281
- Developer Prompt: "Wire up AI bill scanning next to the Manual Entry button"
282
- │
283
- ├── 1. get_architecture() → Maps app/, templates/, static/js/, tests/
284
- ├── 2. search_code(query="Add Manual Entry") → Matches templates/bills.html:6
285
- ├── 3. get_file(path="templates/bills.html", 1..12) → Reads bounded 12-line HTML slice
286
- ├── 4. find_routes(path="/api/scan-bill") → Resolves POST /api/scan-bill → scan_bill_endpoint
287
- ├── 5. find_symbol(symbol="parse_bill") → Grounds app/bill_scanner.py::parse_bill
288
- ├── 6. get_file(path="app/bill_scanner.py", 1..20) → Reads parse_bill() and normalize_line_items()
289
- ├── 7. find_callers(symbol="parse_bill") → Confirms scan_bill_endpoint calls parse_bill
290
- └── 8. find_tests(symbol="parse_bill") → Finds test_parse_bill_calculates_total
291
- ```
292
-
293
- ---
294
-
295
- ## 5. Database Intelligence
296
-
297
- [`src/codegraph/database/`](src/codegraph/database/) provides static schema, ORM model, migration, and query extraction across **SQLAlchemy**, **Django ORM**, **SQLModel**, **Prisma** (`.prisma`), **Alembic**, **Django Migrations**, and **Raw SQL** (`PostgreSQL`, `MySQL`, `SQLite`).
298
-
299
- ```text
300
- POST /orders
301
- ↓ HANDLED_BY (FRAMEWORK_VERIFIED)
302
- OrderService.create_order
303
- ↓ CALLS (AST_VERIFIED)
304
- OrderRepository.insert_order
305
- ↓ WRITES_TABLE (AST_VERIFIED)
306
- db:table:postgresql.public.orders
307
- ↓ FOREIGN_KEY_TO (AST_VERIFIED)
308
- orders.user_id ──► users.id
309
- ```
310
-
311
- ### Capabilities Exposed by the 11 Database MCP Tools
312
- - **Table & Column Discovery** (`get_db_schema`, `get_db_table`, `find_db_tables`, `find_db_columns`): Extract tables, column types, nullability, defaults, primary keys (`HAS_PRIMARY_KEY`), indexes (`HAS_INDEX`), unique/check constraints, and foreign keys (`FOREIGN_KEY_TO`).
313
- - **ORM Mapping** (`find_db_models`): Map SQLAlchemy `__tablename__`, Django `models.Model` (`Meta.db_table`), SQLModel `table=True`, and Prisma `model` blocks via `MAPS_TO_TABLE` and `MAPS_TO_COLUMN`.
314
- - **Readers, Writers & Callers** (`find_db_readers`, `find_db_writers`, `find_db_callers`, `find_db_queries`): Identify every function or method that executes `SELECT` (`READS_TABLE`) or `INSERT` / `UPDATE` / `DELETE` / `.add()` / `.save()` (`WRITES_TABLE`) against a table.
315
- - **Migration Lineage & Schema Blast Radius** (`find_db_relationships`, `get_db_impact`): Track Alembic (`op.create_table`, `op.add_column`) and Django (`migrations.CreateModel`, `migrations.AddField`) operations (`MIGRATES_TABLE`) and compute bidirectional code $\leftrightarrow$ database impact.
316
-
317
- ---
318
-
319
- ## 6. Runtime Intelligence & Static Reconciliation
320
-
321
- Static analysis proves what **can** happen structurally; runtime telemetry records what **was observed** during a specific execution window. [`src/codegraph/runtime/`](src/codegraph/runtime/) combines both without conflating them:
322
-
323
- ```text
324
- STATIC GRAPH (AST + Framework + Dataflow + DB)
325
- +
326
- OPTIONAL RUNTIME OBSERVATION (OTel JSON / JSONL Events / SQL Logs)
327
- ↓
328
- reconcile_static_runtime()
329
- ```
330
-
331
- ### Reconciliation Outcomes ([`ReconciliationStatus`](src/codegraph/runtime/models.py))
332
-
333
- | Reconciliation Status | Static Graph | Runtime Trace | Meaning |
334
- | :--- | :---: | :---: | :--- |
335
- | **`CONFIRMED_RUNTIME_PATH`** | Present | Observed | Static relationship is structurally proven **and** observed executing in ingested traces. |
336
- | **`NOT_OBSERVED_AT_RUNTIME`** | Present | Not observed | Statically valid edge was not exercised in the ingested trace sample. |
337
- | **`RUNTIME_ONLY_OBSERVED`** | Dynamic / `UNKNOWN` | Observed | Executed at runtime (e.g., plugin hook, `getattr`, dynamic SQL) where static analysis remained `UNKNOWN`. |
338
- | **`STATIC_RUNTIME_CONFLICT`** | Target A | Target B | Runtime execution dispatched to a different target than static resolution (e.g., dependency override or subclass). |
339
-
340
- ### Epistemic Rules for Runtime Evidence
341
- 1. **Runtime Telemetry Is Opt-In & Observational**: CodeGraph never instruments or executes your code automatically. Traces are ingested only when you call `ingest_runtime_traces` on OpenTelemetry JSON, structured JSONL, or SQL log files.
342
- 2. **`NOT_OBSERVED_AT_RUNTIME` Does Not Mean Dead Code**: It only means the code path was not triggered during the recorded trace window (for example, an error handler, admin route, or periodic job).
343
- 3. **Runtime Observations Never Overwrite Static Proof**: Runtime spans are stored with `evidence_class="RUNTIME_OBSERVED"` and kept distinct from `AST_VERIFIED`, `FRAMEWORK_VERIFIED`, and `DATAFLOW_VERIFIED` static edges.
344
-
345
- ---
346
-
347
- ## 7. Epistemic Trust & Evidence Contract
348
-
349
- CodeGraph enforces a fail-closed evidence contract ([`src/codegraph/evidence_contract.py`](src/codegraph/evidence_contract.py)) across all **51 canonical relationship types** and **9 evidence classes**. CodeGraph prefers **explicit uncertainty** over **fabricated certainty**:
350
-
351
- ```text
352
- Dynamically resolved target (getattr(handler, action_name)())
353
- ↓
354
- UNRESOLVED_REFERENCE / POSSIBLE_CALLS (status = "UNKNOWN" | "POSSIBLE")
355
- (Never fabricated into a verified CALLS edge)
356
- ```
357
-
358
- ### Current Evidence Vocabulary (`src/codegraph/evidence_contract.py`)
359
-
360
- | Epistemic Status | Allowed Evidence Classes | What It Means |
361
- | :--- | :--- | :--- |
362
- | **`FACT`** | `AST_VERIFIED`, `STATIC_VERIFIED`, `FRAMEWORK_VERIFIED`, `DATAFLOW_VERIFIED` | Proven directly from syntax tree, framework decorator/router semantics, or conservative local dataflow. |
363
- | **`RUNTIME_OBSERVED`** | `RUNTIME_OBSERVED` | Observed in user-supplied OpenTelemetry, JSONL, or SQL query logs (`hit_count`, `p50_ms`, `p95_ms`). |
364
- | **`POSSIBLE`** | `POSSIBLE` | Plausible candidate relationship (`POSSIBLE_CALLS`, `POSSIBLE_TABLE`) requiring source inspection before mutation. |
365
- | **`AMBIGUOUS`** | `AMBIGUOUS` | Multiple symbols or database tables match the bare identifier across modules or dialects; returns sorted `candidates`. |
366
- | **`UNKNOWN`** | `UNKNOWN`, `RUNTIME_UNOBSERVED` | Target cannot be statically proven (dynamic reflection, external unindexed dependency, or `reason="resolution_budget_exceeded"`). |
367
- | **`CONFLICT`** | Static vs. Runtime / Multi-Source | Static analysis and runtime observation (or competing definitions) disagree. |
368
-
369
- ---
370
-
371
- ## 8. Measured Performance (`v2.1.6` → `v2.1.7`)
372
-
373
- ### Methodology
374
- All indexing measurements below were recorded using [`benchmarks/run_v217_indexing_benchmark.py`](benchmarks/run_v217_indexing_benchmark.py) on the **same machine** (macOS `arm64`, Python `3.13`), **same repository fixtures**, and **same 16-phase telemetry harness**, comparing `v2.1.6` ([`benchmarks/reports/v217_before_metrics.json`](benchmarks/reports/v217_before_metrics.json)) against `v2.1.7` ([`benchmarks/reports/v217_after_metrics.json`](benchmarks/reports/v217_after_metrics.json)).
375
-
376
- ### 4-Tier Scaling Summary (`54` → `2,504` Files)
377
-
378
- | Workload Tier | Files | Symbols | Graph Edges | `v2.1.6` Total | `v2.1.7` Total | Improvement | `v2.1.7` Peak RSS | Peak WAL (`v2.1.6` → `v2.1.7`) | Final WAL |
379
- | :--- | ---: | ---: | ---: | ---: | ---: | :--- | ---: | ---: | ---: |
380
- | **Small** | `54` | `115` | `398` | `0.527 s` | `0.325 s` | **38.3% faster (`1.62x`)** | `46.25 MB` | `0.990 MB → 1.544 MB` | `0.0 MB` |
381
- | **Medium** | `304` | `615` | `2,248` | `2.564 s` | `1.613 s` | **37.1% faster (`1.59x`)** | `62.67 MB` | `5.610 MB → 4.098 MB` | `0.0 MB` |
382
- | **Large** | `1,004` | `2,015` | `7,428` | `8.730 s` | `5.296 s` | **39.3% faster (`1.65x`)** | `103.44 MB` | `19.300 MB → 5.033 MB` | `0.0 MB` |
383
- | **Stress** | `2,504` | `5,015` | `18,528` | `21.983 s` | `13.395 s` | **39.1% faster (`1.64x`)** | `180.89 MB` | `46.980 MB → 6.628 MB` | `0.0 MB` |
384
-
385
- ### Stress Tier (`2,504` Files) Phase Breakdown
386
-
387
- | Phase / Metric | `v2.1.6` Baseline | `v2.1.7` Release | Measured Improvement |
388
- | :--- | ---: | ---: | :--- |
389
- | **Total Indexing Time** | `21.983 s` | `13.395 s` | **39.1% faster (`1.64x`)** |
390
- | **Database Intelligence Pass** | `4.818 s` | `0.832 s` | **82.7% faster (`5.79x`)** |
391
- | **Post-Processing Phase** | `13.770 s` | `7.056 s` | **48.8% faster (`1.95x`)** |
392
- | **Symbol Resolution Phase** | `3.133 s` | `1.107 s` | **64.7% faster (`2.83x`)** |
393
- | **Single-File Incremental Update** | `4.408 s` | `2.063 s` | **53.2% faster (`2.14x`)** |
394
- | **Throughput (`files/sec`)** | `113.9 files/s` | `186.9 files/s` | **`+64.1%` throughput** |
395
- | **Peak SQLite WAL Size** | `46.980 MB` | `6.628 MB` | **85.9% reduction (`7.09x` smaller)** |
396
- | **Final SQLite WAL Size** | `0.000 MB` | `0.000 MB` | **100% reclaimed (`TRUNCATE`)** |
397
- | **Indexed Symbols / Graph Edges** | `5,015` / `18,528` | `5,015` / `18,528` | **100% exact parity** |
398
-
399
- ![CodeGraph — Measured Large-Repository Scaling](docs/assets/large_repo_scaling.svg)
400
-
401
- ---
402
-
403
- ## 9. Large-Repository Stress Testing: Home Assistant Core
404
-
405
- Home Assistant Core is a large, complex public Python repository used as a real-world stress case for CodeGraph's indexing and post-processing pipeline.
406
-
407
- ### 1. External Large-Repository Stress Observation (Pre-`v2.1.7`)
408
- During external stress testing on a Home Assistant Core checkout (`~28,573` files), pre-`v2.1.7` indexing exhibited:
409
- - Sustained single-core CPU usage (`~99%`) dominated by late post-processing
410
- - Process memory peaking around `~1.1 GB` RSS before dropping
411
- - Uncheckpointed `.codegraph/index.db-wal` growth reaching `~922 MB` because indexing held a single uncommitted transaction across all files and post-processing edges
412
-
413
- ### 2. Reproducible Benchmark Fixture & Root-Cause Fixes (`v2.1.7`)
414
- To profile and verify fixes deterministically in CI, [`benchmarks/run_v217_indexing_benchmark.py`](benchmarks/run_v217_indexing_benchmark.py) provisions a 4-tier Home Assistant-architecture fixture (`homeassistant/core`, `homeassistant/helpers`, `homeassistant/components/recorder` SQLAlchemy models/queries, `500` component domains, and `pytest` fixture suites; `2,504` files, `5,015` symbols, `18,528` edges):
415
- - **Streaming & Token-Gated Database Pass**: Replaced the in-memory `file_contents` map and 8-per-file AST parses with streaming reads, fast token pre-filters (`has_potential_database_activity`, `has_potential_orm_models`), and a single shared `ast.AST` parse per candidate file (`4.818s → 0.832s`).
416
- - **Pre-Indexed Binding & Symbol Resolution**: Replaced four $O(N_{\text{bindings}} \times N_{\text{symbols}})$ linear scans in [`src/codegraph/resolver.py`](src/codegraph/resolver.py) with pre-indexed maps and `@lru_cache(maxsize=65536)` on `normalize_module` (`3.133s → 1.107s`).
417
- - **Chunked SQLite Commits & `TRUNCATE` Checkpoints**: Added composite indexes (`idx_imports_source_line`, `idx_calls_source_line`), bounded commit batches (`500` files / `10,000` edges), and `PRAGMA wal_checkpoint(TRUNCATE)` (`46.980 MB → 6.628 MB` peak WAL; `0.0 MB` final WAL).
418
- - **Safe `Ctrl+C` Cancellation & Resume**: Interrupting `codegraph index` rolls back only the active batch, preserves committed batches, marks `resolution_dirty="1"`, and resumes cleanly on the next run.
419
-
420
- ---
421
-
422
- ## 10. Context Efficiency & Internal Agent Evaluation
423
-
424
- ### 50-Task Context Compilation Benchmark ([`benchmarks/baselines/v2_0_verified.json`](benchmarks/baselines/v2_0_verified.json))
425
-
426
- Rather than claiming a single universal token reduction percentage across all possible prompts, CodeGraph records candidate-vs-selected token metrics on every `get_context` call:
427
-
428
- | Benchmark Metric | Measured Value | Source Artifact |
429
- | :--- | ---: | :--- |
430
- | **Evaluated Tasks** | `50 tasks` across `10 categories` | [`benchmarks/baselines/v2_0_verified.json`](benchmarks/baselines/v2_0_verified.json) |
431
- | **Average Selected Tokens** | `542.0 tokens` | [`benchmarks/baselines/v2_0_verified.json`](benchmarks/baselines/v2_0_verified.json) |
432
- | **Average Candidate-to-Selected Reduction Ratio** | `65.0%` (`0.65`) | [`benchmarks/baselines/v2_0_verified.json`](benchmarks/baselines/v2_0_verified.json) |
433
- | **Compression at `budget = 200 tokens`** | `85.6%` reduction | [`benchmarks/baselines/context_baseline.json`](benchmarks/baselines/context_baseline.json) |
434
- | **Compression at `budget = 600 tokens`** | `56.3%` reduction | [`benchmarks/baselines/context_baseline.json`](benchmarks/baselines/context_baseline.json) |
435
- | **Cold vs. Warm `get_context` Latency (`p50`)** | `20.0 ms` cold → `0.67 ms` warm (`30.0x`) | [`benchmarks/baselines/context_baseline.json`](benchmarks/baselines/context_baseline.json) |
436
- | **Single-Tool Query Latency (Indexed SQLite)** | `1 ms – 15 ms` | Local SQLite B-tree / FTS5 lookup |
437
- | **FACT / UNKNOWN / AMBIGUITY Correctness** | `100.0%` / `98.0%` / `100.0%` | [`benchmarks/baselines/v2_0_verified.json`](benchmarks/baselines/v2_0_verified.json) |
438
-
439
- ### Results from the 32-Task Internal Evaluation ([`src/codegraph/tool_selection_eval.py`](src/codegraph/tool_selection_eval.py))
440
-
441
- The table below reports results from the **32-task internal evaluation harness** (`run_tool_selection_ab_benchmark()`) comparing Mode A (unassisted `grep` + full-file reading without CodeGraph routing rules) against Mode B (CodeGraph MCP + agent routing rules) on the same 32 tasks:
442
-
443
- | Metric (32-Task Internal Evaluation) | Mode A (Without CodeGraph) | Mode B (With CodeGraph MCP) | Measured Improvement |
444
- | :--- | ---: | ---: | :--- |
445
- | **Total Tool Calls (32 Tasks)** | `164 calls` (`~5.1 / task`) | `71 calls` (`~2.2 / task`) | **`-93 calls (-56.7% fewer tool calls)`** |
446
- | **Direct Full-File Reads** | `161 full-file reads` | `7 bounded slice reads` | **`-154 file reads (-95.7% fewer file reads)`** |
447
- | **Context Tokens per Task** | `3,000 – 12,000+ tokens` (full files) | `~542 tokens` avg (`get_context`) | **`65.0% – 85.6% token reduction`** |
448
- | **Query Latency (`p50`)** | Hundreds of ms (multi-step `grep` + reads) | `20.0 ms` cold / `0.67 ms` warm (`30.0x`) | **Sub-20ms indexed lookup** |
449
- | **First-Tool Selection Accuracy** | `9.38%` | `100.0%` | **`+90.62%`** |
450
- | **Unsupported Claims (Hallucinated Edges)** | `11 (34.38%)` | `0 (0.00%)` | **`-11 claims (0% unsupported)`** |
451
- | **Task Accuracy** | `59.84%` | `100.0%` | **`+40.16%`** |
452
- | **12-Prompt Natural-Language Routing Eval** | — | `12 / 12 (100.0%)` | **`12/12 on this evaluation set`** |
453
-
454
- ![CodeGraph — Measured Context & Exploration Efficiency](docs/assets/performance_comparison.svg)
455
-
456
- ---
457
-
458
- ## 11. Where CodeGraph Fits
459
-
460
- CodeGraph is designed to work **alongside** your editor's language server (LSP), `ripgrep`, and structural AST tools.
461
-
462
- Legend: `✓` supported • `◐` partial / workflow-dependent • `—` not established by cited documentation
463
-
464
- | Capability | CodeGraph MCP (`v2.1.7`) | Editor LSP [1] | Structural AST (`ast-grep`) [2] | Lexical Search (`ripgrep`) [3] | Remote Code Search (`Sourcegraph MCP`) [4] |
465
- | :--- | :---: | :---: | :---: | :---: | :---: |
466
- | **Local-First & Offline Operation** | ✓ | ✓ | ✓ | ✓ | — |
467
- | **Native MCP Server for AI Agents** | ✓ (14 default / 56 full) | — | ◐ | — | ✓ |
468
- | **Semantic Symbol Graph (Callers / Callees)** | ✓ | ◐ (Position-based) | ◐ (Pattern-based) | — | ✓ (SCIP) |
469
- | **Framework Route, Mount & DI Graph** | ✓ (`FastAPI`/`Flask`/`Django`/`Express`) | — | ◐ (Custom YAML rules) | — | — |
470
- | **Database Schema, ORM, Migration & Table R/W** | ✓ (11 DB tools) | — | — | — | — |
471
- | **Opt-In Runtime Trace Ingestion & Reconciliation** | ✓ (OTel / JSONL / SQL) | — | — | — | — |
472
- | **Explicit Epistemic States (`FACT`/`POSSIBLE`/`UNKNOWN`)** | ✓ | — | — | — | — |
473
- | **Literal Text Search Across HTML/JS/CSS/Config** | ✓ (`search_code`) | — | — | ✓ | ✓ |
474
- | **Interactive Editor Hover, Completions & Diagnostics** | — | ✓ | ✓ (Lint/Rewrite) | — | — |
475
- | **Multi-Repository Enterprise Cloud Search** | — | — | — | — | ✓ |
476
-
477
- ![Repository Retrieval Approaches — Capability Comparison](docs/assets/capability_heatmap.svg)
478
-
479
- For the complete multi-tool comparison and official references ([1] [LSP Specification](https://microsoft.github.io/language-server-protocol/specifications/lsp/current/), [2] [`ast-grep`](https://ast-grep.github.io/), [3] [`ripgrep`](https://github.com/BurntSushi/ripgrep), [4] [Sourcegraph MCP](https://sourcegraph.com/docs/api/mcp), [5] [Colby McHenry CodeGraph](https://github.com/colbymchenry/codegraph), [6] [GitHub Code Navigation](https://docs.github.com/en/repositories/working-with-files/using-files/navigating-code-on-github)), see [`docs/tool-comparison.md`](docs/tool-comparison.md).
480
-
481
- ---
482
-
483
- ## 12. Security, Privacy & Redaction Boundaries
484
-
485
- CodeGraph runs **100% locally** (`stdio` MCP + local `.codegraph.sqlite3`), never executes repository code during indexing, and enforces strict file-access and redaction boundaries ([`src/codegraph/security/paths.py`](src/codegraph/security/paths.py), [`src/codegraph/security/redaction.py`](src/codegraph/security/redaction.py)).
486
-
487
- ### `BLOCKED` vs. `REDACTED` Behavior
488
-
489
- | Security Boundary | Enforcement Mode | Exact Behavior |
490
- | :--- | :---: | :--- |
491
- | **Sensitive Files** (`.env`, `.env.*`, `*.pem`, `*.key`, `*.crt`, `*.p12`, `*.pfx`, `id_rsa*`, `id_ed25519*`, `kubeconfig*`, `.npmrc`, `.pypirc`, `.netrc`, `.git/credentials`, `credentials*`, `secrets.*`, `*secret*.json/yaml`, `service-account*`, `.aws/*`, `.ssh/*`, `.gnupg/*`, `.kube/*`, `.docker/config.json`, `*.sqlite*`, `*.db`) | **`BLOCKED`** | Excluded from indexing and FTS; direct inspection via `get_file` or `read_file` is rejected with `SENSITIVE_FILE_ACCESS_DENIED`. |
492
- | **Path Traversal & Symlink Escapes** (`../`, URL-encoded `%2e%2e`, null bytes, external symlinks) | **`BLOCKED`** | Canonical path check in `resolve_within_repo()` raises `SecurityError(ErrorCode.PATH_OUTSIDE_REPOSITORY)`. |
493
- | **Binary Files** (`.pyc`, `.so`, `.dylib`, `.dll`, `.exe`, images, archives, PDFs, fonts, or NUL-byte files) | **`BLOCKED`** | Classified as `BINARY` and skipped during indexing and text search. |
494
- | **Environment Variable Reads in Code** (`os.getenv("DATABASE_URL")`, `os.environ["OPENAI_API_KEY"]`) | **Metadata Only** | Records `READS_ENV` with the **variable name only**; never reads `.env` or runtime environment values. |
495
- | **Database Connection Strings** (`postgresql://user:pass@host:5432/prod`) | **`REDACTED`** | Preserves dialect and database name while sanitizing credentials to `postgresql://[REDACTED]@[REDACTED]/prod`. |
496
- | **API Keys, Bearer Tokens, JWTs & Private Keys** (`sk-...`, `ghp_...`, `AKIA...`, `AIza...`, `xoxb-...`, `eyJ...`, `-----BEGIN ... PRIVATE KEY-----`) | **`REDACTED`** | Replaced with `[REDACTED_SECRET]`, `Bearer [REDACTED]`, or `[REDACTED_PRIVATE_KEY]` before FTS indexing, `search_code`, `get_file`, `get_context`, or MCP responses. |
497
- | **Runtime HTTP Headers & SQL Query Literals** (`authorization`, `cookie`, `set-cookie`, `x-api-key`, `WHERE password = '...'`) | **`REDACTED`** | Sensitive runtime keys and SQL literals are scrubbed (`[REDACTED]` / `?`) during `ingest_runtime_traces`. |
498
-
499
- Run `codegraph privacy .` at any time to audit the local SQLite database and verify that no sensitive files or unredacted secrets are stored.
500
-
501
- ---
502
-
503
- ## 13. MCP Tooling & Profiles (14 Default / 56 Full)
504
-
505
- By default, `create_server()` exposes the **14-tool `agent` profile** so AI coding agents receive a focused, non-overlapping tool surface. All 56 tools are available under `--profile full`.
506
-
507
- ### Default `agent` Profile (14 High-Signal Tools)
508
-
509
- | Category | Tool | Purpose |
510
- | :--- | :--- | :--- |
511
- | **Discovery (7)** | `find_symbol` | Locate a symbol definition by short, qualified, or canonical name (`symbol=...`). |
512
- | | `search_code` | Literal or regex search across Python, JS/TS, HTML/Jinja, CSS, YAML/JSON, Markdown, and SQL. |
513
- | | `find_references` | Find verified AST reference, import, and registration sites for a symbol. |
514
- | | `find_callers` | Find functions, methods, or route handlers that call the target symbol (`CALLS`, `POSSIBLE_CALLS`). |
515
- | | `find_callees` | Find functions, methods, or constructors called by the target symbol. |
516
- | | `find_tests` | Find `pytest` / `unittest` test functions covering a symbol, route, DI provider, or event handler. |
517
- | | `find_routes` | Discover FastAPI, Flask, Django, and Express HTTP routes with composed mount prefixes. |
518
- | **Details & Context (5)** | `get_symbol` | Retrieve signature, decorators, docstring, line range, and methods for a symbol. |
519
- | | `get_file` | Read bounded line ranges (`start_line`, `end_line`, `max_lines`) and AST symbol outline of a file. |
520
- | | `get_context` | Compile a token-budgeted, task-aware context packet (`query`, `intent`, `max_tokens`). |
521
- | | `get_architecture` | Summarize repository languages, layers, packages, entrypoints, routes, and database entities. |
522
- | | `get_git_impact` | Compute blast-radius impact (`changed_files`, `affected_callers`, `affected_routes`, `tests`) for a Git diff. |
523
- | **Graph Tracing (2)** | `trace_path` | Find the shortest verified execution path between `from_symbol` and `to_symbol`. |
524
- | | `trace_flow` | Trace upstream callers and downstream callees around `symbol` up to `depth`. |
525
-
526
- ### All 6 Implemented MCP Profiles ([`src/codegraph/agent_capabilities.py`](src/codegraph/agent_capabilities.py))
527
-
528
- | Profile | Tool Count | Description |
529
- | :--- | :---: | :--- |
530
- | **`agent`** *(default in `create_server`)* | **14** | High-signal discovery, bounded file inspection, context synthesis, and path tracing. |
531
- | **`core`** | **13** | Lightweight symbol lookup, callers/callees, imports/dependents, routes, and architecture. |
532
- | **`graph`** | **17** | Call-graph traversal, blast-radius impact (`analyze_impact`), and test discovery. |
533
- | **`minimal`** | **21** | Core interrogation plus `get_context`, `search_code`, `read_file`, and `verify_evidence`. |
534
- | **`developer`** | **34** | Interactive development with Git history (`get_file_history`, `get_recent_changes`) and retrieval planning. |
535
- | **`full`** | **56** | Complete capability surface including **11 Database tools** (`get_db_schema`, `get_db_table`, `find_db_tables`, `find_db_columns`, `find_db_models`, `find_db_queries`, `find_db_readers`, `find_db_writers`, `find_db_callers`, `find_db_relationships`, `get_db_impact`) and **3 Runtime tools** (`ingest_runtime_traces`, `get_runtime_trace`, `reconcile_static_runtime`). |
536
-
537
- See [`docs/agent-brain.md`](docs/agent-brain.md) and [`agent-rules/tool-capabilities-summary.md`](agent-rules/tool-capabilities-summary.md) for the complete 56-tool reference.
538
-
539
- ---
540
-
541
- ## 14. CLI Reference
542
-
543
- Every command below is verified against [`src/codegraph/cli.py`](src/codegraph/cli.py):
544
-
545
- ```bash
546
- # Agent Onboarding & Uninstall
547
- codegraph install # Interactive agent detection & setup
548
- codegraph install --yes --target auto --location local # Non-interactive local setup
549
- codegraph install --print-config claude # Print MCP JSON + rules for an agent
550
- codegraph install --dry-run # Preview planned file changes
551
- codegraph uninstall --dry-run # Preview removal of CodeGraph agent blocks
552
- codegraph uninstall --yes # Remove CodeGraph agent integrations
553
-
554
- # Repository Initialization & Indexing
555
- codegraph init . # Initialize and index repository
556
- codegraph index . --verbose # Incremental index with 16-phase telemetry
557
- codegraph index . --json # Output structured JSON phase telemetry
558
- codegraph uninit --dry-run # Preview removal of .codegraph.sqlite3
559
-
560
- # Process Lifecycle & Upgrade Lock Safety
561
- codegraph stop # Safely stop CodeGraph MCP server for current repo
562
- codegraph stop --all # Stop all CodeGraph-owned MCP processes across repos
563
- codegraph stop --repo /path/to/repo --json # Stop MCP processes for a specific repo (JSON output)
564
- codegraph doctor --processes # Inspect active MCP processes, parent PIDs, & stale locks
565
-
566
- # Health, Integrity & Privacy Diagnostics
567
- codegraph status . # Show index freshness and graph counts
568
- codegraph doctor . --database --resources # Verify SQLite integrity, FKs, FTS, and memory
569
- codegraph privacy . # Verify zero sensitive files indexed
570
- codegraph version # Print CodeGraph version (2.1.7)
571
-
572
- # Code, Graph & Context Interrogation
573
- codegraph search "authenticate" -r . # Search indexed symbols and chunks
574
- codegraph symbols src/codegraph/cli.py -r . # List extracted symbols in a file
575
- codegraph get-symbol Indexer -r . # Get AST details for a symbol
576
- codegraph resolve-symbol resolve_Repository -r . # Ground symbol or return ambiguous candidates
577
- codegraph resolve Indexer -r . # Resolve symbol with callers and callees
578
- codegraph trace Indexer -d 2 -r . # Trace callers and callees up to depth 2
579
- codegraph graph -r . # Summarize graph nodes and edges
580
- codegraph routes -r . # List discovered HTTP routes
581
- codegraph imports src/codegraph/cli.py -r . # List file/module imports
582
- codegraph dependents src/codegraph/cli.py -r . # List reverse dependents
583
- codegraph architecture -r . # Summarize repository architecture
584
- codegraph debug "trace authentication flow" -r . # Return facts and debugging hypotheses
585
- codegraph task "trace /api/v1/auth/login" # Normalize prompt into a TaskSpec
586
- codegraph plan "trace /api/v1/auth/login" -r . # Build deterministic RetrievalPlan
587
- codegraph context "trace /api/v1/auth/login" -r . # Compile token-budgeted ContextPacket
588
- codegraph explain-context "trace /api/v1/auth/login" -r . # ContextPacket with budget rejection reasons
589
- codegraph memory list -r . # Inspect repository-scoped notes
590
- codegraph benchmark -r . # Run deterministic benchmark suite
591
-
592
- # MCP Server Subcommands
593
- codegraph mcp serve . # Start stdio MCP server (with parent/EOF lifecycle guard)
594
- codegraph mcp serve . --profile full # Start stdio MCP server with all 56 tools
595
- codegraph mcp stop # Alias for codegraph stop
596
- codegraph mcp kill # Alias for codegraph stop (1s graceful timeout)
597
- codegraph mcp doctor . # End-to-end MCP startup & query check
598
- codegraph mcp config-check . # Read-only MCP config validation
599
- codegraph mcp capabilities # Print machine-readable capability manifest
600
- codegraph mcp rules --agent claude # Render agent rules for a specific agent
601
- ```
602
-
603
- ---
604
-
605
- ## 15. Honest Limitations
606
-
607
- 1. **Dynamic Metaprogramming & Reflection**: Calls constructed dynamically (`getattr(obj, dynamic_name)()`, `eval`, `exec`, `importlib.import_module(var)`, or runtime monkey-patching) cannot be proven statically. CodeGraph intentionally records these as `UNKNOWN` (`UNRESOLVED_REFERENCE` or `POSSIBLE_CALLS`) rather than inventing edges.
608
- 2. **Bounded Wildcard & Re-Export Chains**: To guarantee termination on pathological repositories with circular `from x import *` chains, resolution halts at `max_reexport_depth=16` (`max_wildcard_expansions=64`) and emits `UNKNOWN` with `reason="resolution_budget_exceeded"`.
609
- 3. **Python-First Depth vs. JS/TS Secondary Support**: Python receives deep AST, decorator, local dataflow (`LocalBindingResolver`), FastAPI/Flask/Django route, and SQLAlchemy/Django/SQLModel/Alembic analysis. JavaScript/TypeScript supports functions, classes, imports, calls, Express routes, and Prisma schemas, without full TypeScript compiler type evaluation.
610
- 4. **Opt-In Runtime Telemetry Scope**: Runtime edges (`RUNTIME_OBSERVED`) reflect only the trace files you explicitly ingest. Unobserved paths (`NOT_OBSERVED_AT_RUNTIME`) are not dead code.
611
- 5. **Very Large Pathological Repositories**: While `v2.1.7` reduces indexing time by `39.1%` and caps WAL size via chunked commits, initial cold indexing on repositories with tens of thousands of files still requires proportional CPU and disk I/O time (subsequent runs are incremental).
612
-
613
- ---
614
-
615
- ## 16. Documentation Map
616
-
617
- - **Deep Agent Brain & 56-Tool Reference**: [`docs/agent-brain.md`](docs/agent-brain.md)
618
- - **Compact Tool Capabilities Summary**: [`agent-rules/tool-capabilities-summary.md`](agent-rules/tool-capabilities-summary.md)
619
- - **Detailed Multi-Tool Capability Comparison**: [`docs/tool-comparison.md`](docs/tool-comparison.md)
620
- - **Antigravity Skill (`SKILL.md`)**: [`.agents/skills/codegraph/SKILL.md`](.agents/skills/codegraph/SKILL.md)
621
- - **Agent Rule Packs (`Claude`, `Cursor`, `Antigravity`, `Codex`, `Gemini`, `Cline`)**: [`agent-rules/README.md`](agent-rules/README.md) & [`agent-rules/AGENTS.md`](agent-rules/AGENTS.md)
622
- - **Engineering & Production Readiness**: [`docs/engineering/production-readiness.md`](docs/engineering/production-readiness.md)
623
- - **Reproducible `v2.1.7` Benchmark Script & Artifacts**: [`benchmarks/run_v217_indexing_benchmark.py`](benchmarks/run_v217_indexing_benchmark.py), [`benchmarks/reports/v217_before_metrics.json`](benchmarks/reports/v217_before_metrics.json), [`benchmarks/reports/v217_after_metrics.json`](benchmarks/reports/v217_after_metrics.json)
624
- - **Changelog**: [`CHANGELOG.md`](CHANGELOG.md)
625
-
626
- ---
627
-
628
- ## 17. Contributing & License
629
-
630
- ```bash
631
- git clone https://github.com/raghurammrsd/CODE_GRAPH_MCP.git
632
- cd CODE_GRAPH_MCP
633
- python3 -m venv .venv
634
- source .venv/bin/activate
635
- pip install -e ".[dev,mcp,system]"
636
-
637
- python3 -m ruff check .
638
- python3 -m mypy src/
639
- python3 -m pytest -q
640
- ```
641
-
642
- Licensed under the [MIT License](LICENSE).