open-context-engine 0.1.2 → 0.1.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -35,7 +35,9 @@ Measured with the optional batch rerank API. 40 source-derived tasks, each asked
35
35
 
36
36
  ## Quick start
37
37
 
38
- Requires **macOS or Linux**, Node.js 22.14+, Python 3.10+, Git, and configured embedding/reranking services. Go repositories also need Go 1.22+.
38
+ Requires **macOS, Linux, or Windows**, Node.js 22.14+, Python 3.10+, Git, and configured embedding/reranking services. Go repositories also need Go 1.22+.
39
+
40
+ Native Windows support is available in version **0.1.3** and later. Run the installation commands in PowerShell. Version **0.1.4** adds automatic worker sharing across MCP sessions to prevent index-lock conflicts.
39
41
 
40
42
  Install from npm, then expand your client's guide. Model settings are shared across clients on the same machine.
41
43
 
@@ -148,7 +150,7 @@ Your agent supplies the current project's absolute path as `directory_path`. To
148
150
 
149
151
  ## Explore
150
152
 
151
- [Benchmark report](docs/BENCHMARKS.md) · [Raw evaluations](docs/eval/results) · [Retrieval engine](src/retrieval)
153
+ [Benchmark report](https://github.com/AnnaSuSu/OpenContextEngine/blob/main/docs/BENCHMARKS.md) · [Raw evaluations](https://github.com/AnnaSuSu/OpenContextEngine/tree/main/docs/eval/results) · [Retrieval engine](https://github.com/AnnaSuSu/OpenContextEngine/tree/main/src/retrieval)
152
154
 
153
155
  ## License
154
156
 
package/README.zh-CN.md CHANGED
@@ -35,7 +35,9 @@ OpenContextEngine 是一个**可自行部署、面向 AI 编程助手的代码
35
35
 
36
36
  ## 快速开始
37
37
 
38
- 需要 **macOS 或 Linux**、Node.js 22.14+、Python 3.10+、Git,以及已配置好的向量和重排服务。分析 Go 仓库还需要 Go 1.22+。
38
+ 需要 **macOS、Linux 或 Windows**、Node.js 22.14+、Python 3.10+、Git,以及已配置好的向量和重排服务。分析 Go 仓库还需要 Go 1.22+。
39
+
40
+ **0.1.3** 起支持原生 Windows,安装命令可在 PowerShell 中运行。**0.1.4** 起支持多个 MCP 会话自动共享 worker,避免索引锁冲突。
39
41
 
40
42
  通过 npm 安装,展开你所用客户端的教程即可。同一台机器、同一用户下的多个客户端可以共用模型配置。下方链接的详细技术文档目前为英文。
41
43
 
@@ -148,7 +150,7 @@ Claude 会通过 `directory_path` 传入项目的绝对路径。首次请求会
148
150
 
149
151
  ## 进一步了解
150
152
 
151
- [评测报告](docs/BENCHMARKS.md) · [原始评测数据](docs/eval/results) · [检索引擎源码](src/retrieval)
153
+ [评测报告](https://github.com/AnnaSuSu/OpenContextEngine/blob/main/docs/BENCHMARKS.md) · [原始评测数据](https://github.com/AnnaSuSu/OpenContextEngine/tree/main/docs/eval/results) · [检索引擎源码](https://github.com/AnnaSuSu/OpenContextEngine/tree/main/src/retrieval)
152
154
 
153
155
  ## 许可证
154
156
 
@@ -10,7 +10,7 @@ import { serviceConfig } from '../src/service.mjs';
10
10
  const help = `OpenContextEngine — repository context for AI coding agents
11
11
 
12
12
  Usage:
13
- open-context-engine setup [--python /path/to/python3] [--non-interactive]
13
+ open-context-engine setup [--python <interpreter-path>] [--non-interactive]
14
14
  open-context-engine mcp [--root /project] [--state /outside/index]
15
15
  open-context-engine mcp --connect
16
16
  open-context-engine mcp-config
@@ -19,7 +19,7 @@ Usage:
19
19
 
20
20
  Setup installs isolated Python dependencies, saves shared model settings, and
21
21
  prints MCP configuration. Without --root, your agent supplies directory_path.
22
- Requires macOS/Linux, Node.js 22.14+, Python 3.10+, and Git.
22
+ Requires macOS/Linux/Windows, Node.js 22.14+, Python 3.10+, and Git.
23
23
  `;
24
24
  function mcpConfig() {
25
25
  const env = process.env.OCE_CONFIG_HOME ? {OCE_CONFIG_HOME:configDirectory()} : undefined;
@@ -1,6 +1,6 @@
1
1
  # Connect OpenContextEngine to your agent
2
2
 
3
- OpenContextEngine runs a local repository index and calls configured model services for embeddings and reranking. Its MCP interface uses stdio. Currently supported hosts: **macOS and Linux**.
3
+ OpenContextEngine runs a local repository index and calls configured model services for embeddings and reranking. Its MCP interface uses stdio. Supported hosts: **macOS, Linux, and Windows**. Native Windows support requires version **0.1.3** or later.
4
4
 
5
5
  ## 1. Install
6
6
 
@@ -13,7 +13,7 @@ npm install -g open-context-engine
13
13
  open-context-engine setup
14
14
  ```
15
15
 
16
- `setup` asks for your embedding and reranking endpoints, keys, model names, and embedding dimensions. It creates a private Python virtual environment, installs NumPy and tiktoken, and preloads tokenizer data. No model weights are installed. API key input is not echoed. Use `--python /absolute/path/to/python3` to select a base interpreter.
16
+ `setup` asks for your embedding and reranking endpoints, keys, model names, and embedding dimensions. It creates a dedicated Python virtual environment, installs NumPy and tiktoken, and preloads tokenizer data. No model weights are installed. API key input is not echoed. Use `--python /absolute/path/to/python3` to select a base interpreter.
17
17
 
18
18
  You can also run the same setup from source:
19
19
 
@@ -24,13 +24,21 @@ npm ci
24
24
  node bin/opencontextengine.mjs setup
25
25
  ```
26
26
 
27
- For internal builds, maintainers can create an archive with `npm pack` and install it with `npm install -g /path/to/open-context-engine-0.1.2.tgz`. See the [release checklist](https://github.com/AnnaSuSu/OpenContextEngine/blob/main/docs/RELEASING.md).
27
+ On Windows, run the npm installation commands in PowerShell with `node`, `npm`, `python`, and `git` on `PATH`; WSL is not required. Setup defaults to `python` and uses `Scripts/python.exe` inside its virtual environment. To select a particular interpreter, run:
28
+
29
+ ```powershell
30
+ open-context-engine setup --python "C:\Program Files\Python312\python.exe"
31
+ ```
32
+
33
+ Use native absolute project paths such as `C:\Users\you\project` in MCP calls. Generated MCP JSON escapes backslashes automatically.
34
+
35
+ For internal builds, maintainers can create an archive with `npm pack` and install it with `npm install -g /path/to/open-context-engine-0.1.4.tgz`. See the [release checklist](https://github.com/AnnaSuSu/OpenContextEngine/blob/main/docs/RELEASING.md).
28
36
 
29
37
  The CLI and package are named `open-context-engine`. The previous `opencontextengine` command remains an alias. Existing configuration and cache directories keep their paths, so saved keys and indexes are reused.
30
38
 
31
39
  ## 2. Shared model configuration
32
40
 
33
- Setup saves `~/.config/opencontextengine/config.json` with owner-only file permissions. All MCP clients running under the same user share these settings. Python environments live in the adjacent `runtimes/` directory. Use `OCE_CONFIG_HOME` to select a separate configuration directory; setup includes that override in its generated MCP configuration.
41
+ Setup saves `.config/opencontextengine/config.json` under your home directory. Files use owner-only permissions on macOS/Linux; Windows uses the directory's inherited access permissions. All MCP clients running under the same user share these settings. Python environments live in the adjacent `runtimes/` directory. Use `OCE_CONFIG_HOME` to select a separate configuration directory; setup includes that override in its generated MCP configuration.
34
42
 
35
43
  Configuration precedence is **process environment → saved user settings → source checkout `.env` defaults**. The `.env` of the project being searched is never loaded. Existing source installations using `.env` and `OCE_PYTHON` continue to work.
36
44
 
@@ -52,6 +60,20 @@ The setup prompt initially suggests `1024` dimensions; replace it with your endp
52
60
 
53
61
  The embedding service must implement `POST /v1/embeddings`. By default, the reranker uses the ordinary `/rerank` API: requests contain `model`, `query`, `documents`, and `top_n`; responses must return every requested document in `results`, with its original `index` and a finite `relevance_score` between 0 and 1. Results may arrive in relevance order. Set the reranker base URL to the part before `/rerank`: for example, `https://provider.example/v1`, `/v2`, or `https://provider.example` for an unversioned endpoint.
54
62
 
63
+ For Alibaba Cloud Model Studio, use its OpenAI-compatible embedding endpoint and select the DashScope rerank format:
64
+
65
+ ```dotenv
66
+ EMBEDDING_BASE_URL=https://YOUR_WORKSPACE.cn-beijing.maas.aliyuncs.com/compatible-mode/v1
67
+ EMBEDDING_MODEL=qwen3.7-text-embedding
68
+ OCE_EMBEDDING_DIMENSIONS=1024
69
+ OCE_EMBEDDING_BATCH_SIZE=20
70
+ RERANK_BASE_URL=https://YOUR_WORKSPACE.cn-beijing.maas.aliyuncs.com/api/v1
71
+ RERANK_MODEL=qwen3.7-text-rerank
72
+ OCE_RERANK_API=dashscope
73
+ ```
74
+
75
+ Supply your API key through `EMBEDDING_API_KEY` and `RERANK_API_KEY`, then run setup to save these settings. DashScope mode calls `/services/rerank/text-rerank/text-rerank`, nests the request under `input` and `parameters`, and reads `output.results`. The embedding model defaults to 1024 dimensions and accepts up to 20 texts per request; `OCE_EMBEDDING_BATCH_SIZE` limits index-building batches (default 64, range 1–64). Query embedding uses at most five texts per MCP search. Clear any previous SSH transport overrides when switching to direct HTTPS endpoints. Changing the embedding provider or model rebuilds the index; existing embeddings are retained for reuse with their original configuration.
76
+
55
77
  HTTPS is the default. For an explicitly trusted remote HTTP deployment, set `OCE_ALLOW_HTTP=1` before running `open-context-engine setup`; setup saves this choice in the shared configuration. HTTP transmits API keys and source text without encryption. Local model endpoints remain prohibited. Set `OCE_EMBEDDING_DIMENSIONS` to the service's actual output size (for example, `2560`); changing the provider or dimensions creates a new index generation and does not mix incompatible cached vectors.
56
78
 
57
79
  OpenContextEngine groups the needed pairs by query, reuses scores within each search, and makes at most two concurrent rerank requests by default. Optional `OCE_RERANK_CONCURRENCY` (1–8, default 2) and `OCE_RERANK_MAX_DOCUMENTS` (1–1,024, default 128) control concurrency and documents per request. It requests all scores and rejects missing, duplicate, or invalid result indices; errors are surfaced without silently switching endpoints.
@@ -98,7 +120,9 @@ If your client already has `open-context-engine` on `PATH`, this shorter equival
98
120
 
99
121
  Use your client's equivalent configuration format. Configuration and dependency errors go to stderr; MCP stdout is reserved for protocol messages.
100
122
 
101
- Without `--root`, the server uses **automatic workspace mode**. Your agent supplies the absolute project directory in `directory_path` on each tool call. No indexing starts until a project is requested; first access starts a repository worker and background indexing. Later calls reuse it. One MCP session can search multiple projects, each with an independent worker and persistent index. Workers stay active until the client disconnects, when all are stopped.
123
+ Without `--root`, the server uses **automatic workspace mode**. Your agent supplies the absolute project directory in `directory_path` on each tool call. No indexing starts until a project is requested; first access starts a repository worker and background indexing. Later calls reuse it, including calls from other MCP processes under the same local user. One MCP session can search multiple projects, each with an independent worker and persistent index.
124
+
125
+ Version **0.1.4** and later automatically share one worker per index directory across compatible clients. Closing a client releases only its lease; a worker exits after all leases expire or are released, no searches remain in flight, and it has been idle for 30 seconds. Clients renew their leases automatically, and reconnect if the worker exits.
102
126
 
103
127
  The server does not infer your editor's project from its own launch directory. Its tool instructions tell the agent to use the project path supplied by the host, or inspect the current project directory. Missing, relative, or invalid paths return an error. Symbolic links to the same directory share a worker. Supply the same project root consistently, rather than a different subdirectory on each call.
104
128
 
@@ -150,7 +174,15 @@ Saved files are checked every second by default (`OCE_POLL_SECONDS=1`), with a 3
150
174
 
151
175
  Search actively checks source hashes before retrieval and again before returning. It waits up to 30 seconds for synchronization (`freshnessWaitMs`, maximum 120 seconds). Failed updates, timeouts, or edits during retrieval produce explicit errors. Unsaved editor buffers are not indexed.
152
176
 
153
- State is stored in `~/.cache/opencontextengine/<repository-path-hash>/`. Override it with `--state /outside/repository/index`: automatic mode creates a separate path-hash subdirectory for each project; fixed `--root` mode uses that exact state directory. Existing installations automatically reuse their previous cache location. One worker may write to a state directory at a time. Stop that worker and remove the directory to delete stored source and embeddings.
177
+ Python, JavaScript, TypeScript, and Go files with syntax errors (including unfilled templates) are indexed as plain text with their original paths and line numbers. Healthy files keep their structural analysis; no structural relations are inferred for degraded files. The index remains `ready`, while `index_status` reports `generation.degradedFiles` and `generation.parseDiagnostics` (path, language, error type, line, column, and fallback mode). Search responses include a degraded-file count and MCP displays a short notice. Diagnostics persist across restarts and disappear when the file is repaired or deleted. Model, toolchain, storage, and source-integrity failures still fail explicitly; stale source is never substituted.
178
+
179
+ State is stored in `~/.cache/opencontextengine/<repository-path-hash>/`. Override it with `--state /outside/repository/index`: automatic mode creates a separate path-hash subdirectory for each project; fixed `--root` mode uses that exact state directory. Existing installations automatically reuse their previous cache location. One worker may write to a state directory at a time. Stop the worker and remove the directory to delete stored source and embeddings.
180
+
181
+ Automatic mode discovers the writer through an authenticated loopback handshake; concurrent starts keep the existing writer lock intact. Its `worker.json` connection record contains a local access token and is restricted to the current user (POSIX file permissions or a Windows file ACL). Do not share this file.
182
+
183
+ Configuration and runtime compatibility include model credentials, endpoints, language options, and worker source. Incompatible clients receive an explicit error instead of silently adopting another configuration; close the existing clients before changing settings, or select a separate `--state` directory. An older or manually started worker cannot be adopted automatically: use `--connect` or stop its owning clients before switching to automatic sharing.
184
+
185
+ Searches are serialized with a bounded queue of 16 active/waiting requests and a 30-second queue wait; saturation returns a retryable busy error.
154
186
 
155
187
  When model weights change under the same name, increment `OCE_EMBEDDING_REVISION`. A different provider, model name, or dimension count also invalidates vector reuse. Other models need separate compatibility and quality validation.
156
188
 
@@ -185,7 +217,7 @@ First-time setup requires network access to npm/PyPI and tokenizer data. If Pyth
185
217
 
186
218
  ## Installation verification
187
219
 
188
- 1. Install from npm (or an internal tarball) on macOS or Linux and run setup with your model endpoints.
220
+ 1. Install from npm (or an internal tarball) on macOS, Linux, or Windows and run setup with your model endpoints.
189
221
  2. Paste the generated MCP configuration into your client, restart it, and search a small project.
190
222
  3. Save an edit, add a file, and delete a file; verify search returns current source.
191
223
  4. Switch to another project and back; verify the results belong to the requested project.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "open-context-engine",
3
- "version": "0.1.2",
3
+ "version": "0.1.4",
4
4
  "type": "module",
5
5
  "engines": {
6
6
  "node": ">=22.14"
@@ -33,7 +33,8 @@
33
33
  },
34
34
  "os": [
35
35
  "darwin",
36
- "linux"
36
+ "linux",
37
+ "win32"
37
38
  ],
38
39
  "files": [
39
40
  "README.zh-CN.md",
@@ -42,6 +43,7 @@
42
43
  "src/environment.mjs",
43
44
  "src/client.mjs",
44
45
  "src/service.mjs",
46
+ "src/shared-service.mjs",
45
47
  "src/workspaces.mjs",
46
48
  "src/mcp.mjs",
47
49
  "src/runtime.mjs",
@@ -79,6 +81,6 @@
79
81
  "access": "public",
80
82
  "registry": "https://registry.npmjs.org/"
81
83
  },
82
- "readme": "<div align=\"center\">\n <picture>\n <source media=\"(prefers-color-scheme: dark)\" srcset=\"assets/brand/logo-lockup-dark.svg\">\n <img src=\"assets/brand/logo-lockup.svg\" alt=\"OpenContextEngine\" width=\"660\">\n </picture>\n <p><strong>Precise code context for AI coding agents.</strong></p>\n <p>Find related code across files. Follow its connections. Give your agent the evidence it needs.</p>\n <p>\n <a href=\"docs/BENCHMARKS.md#seven-method-comparison\"><img src=\"https://img.shields.io/badge/dev_evidence_coverage-94.79%25-23875b?style=flat-square\" alt=\"Development evidence coverage: 94.79%\"></a>\n <a href=\"docs/BENCHMARKS.md#seven-method-comparison\"><img src=\"https://img.shields.io/badge/median_retrieval-1.73_s-23875b?style=flat-square\" alt=\"Median retrieval: 1.73 seconds\"></a>\n <a href=\"docs/BENCHMARKS.md#engineering-validation\"><img src=\"https://img.shields.io/badge/verified_tests-107-23875b?style=flat-square\" alt=\"107 verified tests\"></a>\n <a href=\"docs/QUICKSTART.md\"><img src=\"https://img.shields.io/badge/MCP-stdio-193c34?style=flat-square\" alt=\"MCP over stdio\"></a>\n </p>\n <p><a href=\"#quick-start\">Quick start</a> · <a href=\"docs/BENCHMARKS.md\">Benchmarks</a> · <a href=\"docs/QUICKSTART.md\">MCP setup</a> · <strong>English</strong> | <a href=\"README.zh-CN.md\">简体中文</a></p>\n</div>\n\nOpenContextEngine is a **self-hostable code context engine for AI coding agents**. Connect it to your agent through MCP to help it explore an unfamiliar codebase, locate implementations, and find the related code needed for a fix or feature.\n\nIt indexes your working directory and follows saved changes. Given a natural-language task, it combines semantic and keyword search, code relationships, and reranking to return relevant source snippets with file paths and line numbers, within a fixed context budget.\n\n## Why OpenContextEngine\n\n- **Search beyond exact words.** Describe a behavior; retrieve its implementation and connected code across files.\n- **Understand code structure.** Python, TypeScript, JavaScript, and Go adapters, plus text fallback for other languages, configuration, and scripts.\n- **Stay current as you edit.** Saved changes, file deletions, and branch switches sync automatically. Unchanged embeddings are reused; incomplete updates never replace a complete index.\n- **Work across projects through MCP.** Your agent supplies the project path; indexes start on demand and are reused. `search_code` retrieves evidence; `index_status` reports synchronization.\n\n## Measured results\n\n![Required evidence coverage and observed query time](assets/benchmarks/method-comparison.svg)\n\n**94.79% required evidence coverage · 1.73 s median retrieval · 69/80 queries with complete evidence.** Seven engines, four repositories, the same 4,000-token output budget. OpenContextEngine retained the most required evidence in this internal development evaluation.\n\nMeasured with the optional batch rerank API. 40 source-derived tasks, each asked in Chinese and English. Coverage measures source evidence, not coding-agent success. Timings reflect native retrieval for open tools and SDK client calls for ACE. [Full comparison, configurations, and per-query results →](docs/eval/METHOD_COMPARISON.md)\n\n## Quick start\n\nRequires **macOS or Linux**, Node.js 22.14+, Python 3.10+, Git, and configured embedding/reranking services. Go repositories also need Go 1.22+.\n\nInstall from npm, then expand your client's guide. Model settings are shared across clients on the same machine.\n\n<details open>\n<summary><strong>Codex — install, connect, and search</strong></summary>\n\n**1. Install the CLI**\n\nWith Codex CLI already installed, run:\n\n```sh\nnpm install -g open-context-engine\n```\n\n**2. Configure your models**\n\n```sh\nopen-context-engine setup\n```\n\nEnter your embedding and reranking base URLs, API keys, model names, and embedding dimensions. Setup installs isolated Python dependencies and saves your settings. The default reranker uses the ordinary `/rerank` API. [Endpoint examples →](docs/QUICKSTART.md#2-shared-model-configuration)\n\n**3. Add the MCP server**\n\n```sh\ncodex mcp add open-context-engine -- open-context-engine mcp\ncodex mcp get open-context-engine\n```\n\nThe second command checks the saved configuration. These commands assume `open-context-engine` is on the client's `PATH`. For the desktop app or source installations, use the [absolute-path configuration](docs/QUICKSTART.md#client-setup-notes). Restart an already-running Codex client after adding the server.\n\n**4. Search your project**\n\n```sh\ncd /path/to/your-project\ncodex\n```\n\nAsk:\n\n> Use open-context-engine's search_code tool to explain this project's main functionality. Include the entry points and relevant file paths and line numbers.\n\nCodex supplies the project's absolute path as `directory_path`. The first request starts indexing; if it is still building, ask Codex to check `index_status` and retry when ready. Later searches reuse the index, and saved changes update automatically.\n\n</details>\n\n<details>\n<summary><strong>Claude Code — install, connect, and search</strong></summary>\n\n**1. Install the CLI**\n\nWith Claude Code already installed, run:\n\n```sh\nnpm install -g open-context-engine\n```\n\n**2. Configure your models**\n\n```sh\nopen-context-engine setup\n```\n\nEnter your embedding and reranking base URLs, API keys, model names, and embedding dimensions. Setup installs isolated Python dependencies and saves your settings. If you already completed setup for Codex, reuse those settings and skip this step. [Endpoint examples →](docs/QUICKSTART.md#2-shared-model-configuration)\n\n**3. Add the MCP server**\n\n```sh\nclaude mcp add --transport stdio --scope user open-context-engine -- open-context-engine mcp\n```\n\nUser scope makes the server available across your projects. For a shared project configuration, run the command from that project and replace `--scope user` with `--scope project`. These commands assume `open-context-engine` is on the client's `PATH`; see [client setup notes](docs/QUICKSTART.md#client-setup-notes) for absolute paths. Restart an already-running Claude Code session after adding the server.\n\n**4. Search your project**\n\n```sh\ncd /path/to/your-project\nclaude\n```\n\nRun `/mcp` to check the connection, then ask:\n\n> Use open-context-engine's search_code tool to explain this project's main functionality. Include the entry points and relevant file paths and line numbers.\n\nClaude supplies the project's absolute path as `directory_path`. The first request starts indexing; if it is still building, ask Claude to check `index_status` and retry when ready. Later searches reuse the index, and saved changes update automatically.\n\n</details>\n\n<details>\n<summary>Other MCP clients</summary>\n\n[Install and run setup](docs/QUICKSTART.md#1-install), then use the configuration printed by `open-context-engine mcp-config` in your client's supported format. For clients that accept `mcpServers` JSON and can find the installed command on `PATH`:\n\n```json\n{\n \"mcpServers\": {\n \"open-context-engine\": {\n \"command\": \"open-context-engine\",\n \"args\": [\"mcp\"]\n }\n }\n}\n```\n\nYour agent supplies the current project's absolute path as `directory_path`. To pin one project, add `\"--root\", \"/absolute/path/to/your-repository\"` to `args`.\n\n</details>\n\n[Model configuration, troubleshooting, and update behavior →](docs/QUICKSTART.md)\n\n## Explore\n\n[Benchmark report](docs/BENCHMARKS.md) · [Raw evaluations](docs/eval/results) · [Retrieval engine](src/retrieval)\n\n## License\n\n[MIT](LICENSE) © 2026 AnnaSuSu.\n\n## Acknowledgments\n\nThanks to the [LINUX DO](https://linux.do/) community.\n",
84
+ "readme": "<div align=\"center\">\n <picture>\n <source media=\"(prefers-color-scheme: dark)\" srcset=\"assets/brand/logo-lockup-dark.svg\">\n <img src=\"assets/brand/logo-lockup.svg\" alt=\"OpenContextEngine\" width=\"660\">\n </picture>\n <p><strong>Precise code context for AI coding agents.</strong></p>\n <p>Find related code across files. Follow its connections. Give your agent the evidence it needs.</p>\n <p>\n <a href=\"docs/BENCHMARKS.md#seven-method-comparison\"><img src=\"https://img.shields.io/badge/dev_evidence_coverage-94.79%25-23875b?style=flat-square\" alt=\"Development evidence coverage: 94.79%\"></a>\n <a href=\"docs/BENCHMARKS.md#seven-method-comparison\"><img src=\"https://img.shields.io/badge/median_retrieval-1.73_s-23875b?style=flat-square\" alt=\"Median retrieval: 1.73 seconds\"></a>\n <a href=\"docs/BENCHMARKS.md#engineering-validation\"><img src=\"https://img.shields.io/badge/verified_tests-107-23875b?style=flat-square\" alt=\"107 verified tests\"></a>\n <a href=\"docs/QUICKSTART.md\"><img src=\"https://img.shields.io/badge/MCP-stdio-193c34?style=flat-square\" alt=\"MCP over stdio\"></a>\n </p>\n <p><a href=\"#quick-start\">Quick start</a> · <a href=\"docs/BENCHMARKS.md\">Benchmarks</a> · <a href=\"docs/QUICKSTART.md\">MCP setup</a> · <strong>English</strong> | <a href=\"README.zh-CN.md\">简体中文</a></p>\n</div>\n\nOpenContextEngine is a **self-hostable code context engine for AI coding agents**. Connect it to your agent through MCP to help it explore an unfamiliar codebase, locate implementations, and find the related code needed for a fix or feature.\n\nIt indexes your working directory and follows saved changes. Given a natural-language task, it combines semantic and keyword search, code relationships, and reranking to return relevant source snippets with file paths and line numbers, within a fixed context budget.\n\n## Why OpenContextEngine\n\n- **Search beyond exact words.** Describe a behavior; retrieve its implementation and connected code across files.\n- **Understand code structure.** Python, TypeScript, JavaScript, and Go adapters, plus text fallback for other languages, configuration, and scripts.\n- **Stay current as you edit.** Saved changes, file deletions, and branch switches sync automatically. Unchanged embeddings are reused; incomplete updates never replace a complete index.\n- **Work across projects through MCP.** Your agent supplies the project path; indexes start on demand and are reused. `search_code` retrieves evidence; `index_status` reports synchronization.\n\n## Measured results\n\n![Required evidence coverage and observed query time](assets/benchmarks/method-comparison.svg)\n\n**94.79% required evidence coverage · 1.73 s median retrieval · 69/80 queries with complete evidence.** Seven engines, four repositories, the same 4,000-token output budget. OpenContextEngine retained the most required evidence in this internal development evaluation.\n\nMeasured with the optional batch rerank API. 40 source-derived tasks, each asked in Chinese and English. Coverage measures source evidence, not coding-agent success. Timings reflect native retrieval for open tools and SDK client calls for ACE. [Full comparison, configurations, and per-query results →](docs/eval/METHOD_COMPARISON.md)\n\n## Quick start\n\nRequires **macOS, Linux, or Windows**, Node.js 22.14+, Python 3.10+, Git, and configured embedding/reranking services. Go repositories also need Go 1.22+.\n\nNative Windows support is available in version **0.1.3** and later. Run the installation commands in PowerShell. Version **0.1.4** adds automatic worker sharing across MCP sessions to prevent index-lock conflicts.\n\nInstall from npm, then expand your client's guide. Model settings are shared across clients on the same machine.\n\n<details open>\n<summary><strong>Codex — install, connect, and search</strong></summary>\n\n**1. Install the CLI**\n\nWith Codex CLI already installed, run:\n\n```sh\nnpm install -g open-context-engine\n```\n\n**2. Configure your models**\n\n```sh\nopen-context-engine setup\n```\n\nEnter your embedding and reranking base URLs, API keys, model names, and embedding dimensions. Setup installs isolated Python dependencies and saves your settings. The default reranker uses the ordinary `/rerank` API. [Endpoint examples →](docs/QUICKSTART.md#2-shared-model-configuration)\n\n**3. Add the MCP server**\n\n```sh\ncodex mcp add open-context-engine -- open-context-engine mcp\ncodex mcp get open-context-engine\n```\n\nThe second command checks the saved configuration. These commands assume `open-context-engine` is on the client's `PATH`. For the desktop app or source installations, use the [absolute-path configuration](docs/QUICKSTART.md#client-setup-notes). Restart an already-running Codex client after adding the server.\n\n**4. Search your project**\n\n```sh\ncd /path/to/your-project\ncodex\n```\n\nAsk:\n\n> Use open-context-engine's search_code tool to explain this project's main functionality. Include the entry points and relevant file paths and line numbers.\n\nCodex supplies the project's absolute path as `directory_path`. The first request starts indexing; if it is still building, ask Codex to check `index_status` and retry when ready. Later searches reuse the index, and saved changes update automatically.\n\n</details>\n\n<details>\n<summary><strong>Claude Code — install, connect, and search</strong></summary>\n\n**1. Install the CLI**\n\nWith Claude Code already installed, run:\n\n```sh\nnpm install -g open-context-engine\n```\n\n**2. Configure your models**\n\n```sh\nopen-context-engine setup\n```\n\nEnter your embedding and reranking base URLs, API keys, model names, and embedding dimensions. Setup installs isolated Python dependencies and saves your settings. If you already completed setup for Codex, reuse those settings and skip this step. [Endpoint examples →](docs/QUICKSTART.md#2-shared-model-configuration)\n\n**3. Add the MCP server**\n\n```sh\nclaude mcp add --transport stdio --scope user open-context-engine -- open-context-engine mcp\n```\n\nUser scope makes the server available across your projects. For a shared project configuration, run the command from that project and replace `--scope user` with `--scope project`. These commands assume `open-context-engine` is on the client's `PATH`; see [client setup notes](docs/QUICKSTART.md#client-setup-notes) for absolute paths. Restart an already-running Claude Code session after adding the server.\n\n**4. Search your project**\n\n```sh\ncd /path/to/your-project\nclaude\n```\n\nRun `/mcp` to check the connection, then ask:\n\n> Use open-context-engine's search_code tool to explain this project's main functionality. Include the entry points and relevant file paths and line numbers.\n\nClaude supplies the project's absolute path as `directory_path`. The first request starts indexing; if it is still building, ask Claude to check `index_status` and retry when ready. Later searches reuse the index, and saved changes update automatically.\n\n</details>\n\n<details>\n<summary>Other MCP clients</summary>\n\n[Install and run setup](docs/QUICKSTART.md#1-install), then use the configuration printed by `open-context-engine mcp-config` in your client's supported format. For clients that accept `mcpServers` JSON and can find the installed command on `PATH`:\n\n```json\n{\n \"mcpServers\": {\n \"open-context-engine\": {\n \"command\": \"open-context-engine\",\n \"args\": [\"mcp\"]\n }\n }\n}\n```\n\nYour agent supplies the current project's absolute path as `directory_path`. To pin one project, add `\"--root\", \"/absolute/path/to/your-repository\"` to `args`.\n\n</details>\n\n[Model configuration, troubleshooting, and update behavior →](docs/QUICKSTART.md)\n\n## Explore\n\n[Benchmark report](https://github.com/AnnaSuSu/OpenContextEngine/blob/main/docs/BENCHMARKS.md) · [Raw evaluations](https://github.com/AnnaSuSu/OpenContextEngine/tree/main/docs/eval/results) · [Retrieval engine](https://github.com/AnnaSuSu/OpenContextEngine/tree/main/src/retrieval)\n\n## License\n\n[MIT](LICENSE) © 2026 AnnaSuSu.\n\n## Acknowledgments\n\nThanks to the [LINUX DO](https://linux.do/) community.\n",
83
85
  "readmeFilename": "README.md"
84
86
  }
@@ -1,10 +1,13 @@
1
1
  """Authenticated, persistent retrieval worker; model inference stays remote."""
2
2
  import hashlib
3
+ import errno
3
4
  from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
4
5
  import json
6
+ import os
5
7
  from pathlib import Path
6
8
  import re
7
9
  import secrets
10
+ from socketserver import TCPServer
8
11
  import sys
9
12
  import threading
10
13
  import time
@@ -13,33 +16,39 @@ import numpy as np
13
16
 
14
17
  ROOT = Path(__file__).resolve().parents[1]
15
18
  sys.path.insert(0, str(ROOT / 'src' / 'retrieval'))
16
- from routed import RoutedEngine, VERSION
19
+ from evidence import EvidenceEngine, VERSION
20
+ from planning import plan_query
17
21
  from live import LiveIndex, IndexUnavailable
22
+ from writer_lock import WriterBusy
23
+ from shared_worker import SharedWorker
18
24
 
19
25
 
20
- def plan_query(query):
21
- parts = [part.strip() for part in re.split(r'[::;;,]|,\s+(?=how|which|why|where|what)|\s+and\s+(?=how|which|why|where|what)', query, flags=re.I) if len(part.strip()) > 7]
22
- facets = parts if 1 < len(parts) <= 4 else [query]
23
- return {'intent':query,'facets':[{'question':part,'terms':re.findall(r'[A-Za-z][A-Za-z0-9_]*',part)} for part in facets]}
26
+ class LoopbackHTTPServer(ThreadingHTTPServer):
27
+ def server_bind(self):
28
+ # This worker binds a numeric loopback address; reverse DNS is unnecessary.
29
+ TCPServer.server_bind(self)
30
+ self.server_name, self.server_port = self.server_address
31
+
24
32
 
25
33
 
26
34
  def serve(config):
27
35
  initialized = time.monotonic()
28
36
  live = LiveIndex(config).start() if config.get('root') else None
37
+ shared = SharedWorker(config) if config.get('shared') and live else None
29
38
  state = Path(config['state'])
30
39
  if live:
31
40
  index, retrieval = None, None
32
41
  else:
33
42
  units = json.loads((state / 'units.json').read_text())
34
43
  index = json.loads((state / 'metadata.json').read_text())
35
- retrieval = RoutedEngine(units, np.load(state/'vectors.npy'), config['embeddingUrl'], config['reranker'], config['embeddingKey'])
44
+ retrieval = EvidenceEngine(units, np.load(state/'vectors.npy'), config['embeddingUrl'], config['reranker'], config['embeddingKey'])
36
45
  health = {'status':'ready','engine':VERSION,'index':index,'queryCache':False,
37
46
  'initializationMs':round((time.monotonic()-initialized)*1000),
38
47
  'sourceSha256':{name:hashlib.sha256((ROOT/name).read_bytes()).hexdigest() for name in
39
- ['src/retrieval/engine.py','src/retrieval/batched.py','src/retrieval/routed.py','src/retrieval/reranker.py','scripts/retrieval-server.py',
48
+ ['src/retrieval/engine.py','src/retrieval/batched.py','src/retrieval/routed.py','src/retrieval/evidence.py','src/retrieval/planning.py','src/retrieval/reranker.py','scripts/retrieval-server.py',
40
49
  'src/retrieval/languages/__init__.py','src/retrieval/languages/schema.py',
41
50
  'src/retrieval/languages/text.py','src/retrieval/languages/files.py',
42
- 'src/retrieval/languages/python.py','src/retrieval/languages/go.py','src/retrieval/languages/go_ast.go','src/retrieval/languages/go_types.go',
51
+ 'src/retrieval/languages/python.py','src/retrieval/languages/python_calls.py','src/retrieval/languages/go.py','src/retrieval/languages/go_ast.go','src/retrieval/languages/go_types.go',
43
52
  'src/retrieval/languages/typescript.py',
44
53
  'src/retrieval/languages/typescript.mjs','package.json','package-lock.json']
45
54
  if name != 'package-lock.json' or (ROOT/name).is_file()}}
@@ -48,6 +57,7 @@ def serve(config):
48
57
  with urlopen(model_base+'/healthz',timeout=10) as response:
49
58
  health['reranker'] = json.load(response)
50
59
  lock = threading.Lock()
60
+ slots = threading.BoundedSemaphore(16)
51
61
 
52
62
  class Handler(BaseHTTPRequestHandler):
53
63
  def log_message(self, *args):
@@ -71,6 +81,19 @@ def serve(config):
71
81
  self.reply(404,{'error':'Not found'})
72
82
 
73
83
  def do_POST(self):
84
+ if shared and self.path in ('/lease', '/release'):
85
+ if not secrets.compare_digest(self.headers.get('Authorization', ''), 'Bearer ' + config['serviceKey']):
86
+ return self.reply(401, {'error': 'Unauthorized'})
87
+ try:
88
+ length = int(self.headers.get('Content-Length', '0'))
89
+ if not 1 <= length <= 4096:
90
+ raise ValueError('Invalid body size')
91
+ body = json.loads(self.rfile.read(length))
92
+ if not isinstance(body, dict):
93
+ raise ValueError('Invalid lease')
94
+ except (ValueError, TypeError):
95
+ return self.reply(422, {'error': 'Invalid lease request'})
96
+ return self.reply(*shared.lease(body, release=self.path == '/release'))
74
97
  if self.path!='/search':
75
98
  return self.reply(404,{'error':'Not found'})
76
99
  if not secrets.compare_digest(self.headers.get('Authorization',''), 'Bearer '+config['serviceKey']):
@@ -91,8 +114,16 @@ def serve(config):
91
114
  except (ValueError,TypeError):
92
115
  return self.reply(422,{'error':'Expected query text and a token budget between 256 and 8000'})
93
116
  start = time.monotonic()
94
- if not lock.acquire(timeout=5):
95
- return self.reply(429,{'error':'Retrieval worker busy'})
117
+ if not slots.acquire(blocking=False):
118
+ return self.reply(429, {'error': 'Retrieval queue is full; retry later'})
119
+ if shared and not shared.enter():
120
+ slots.release()
121
+ return self.reply(503, {'error': 'Worker is stopping'})
122
+ if not lock.acquire(timeout=30):
123
+ slots.release()
124
+ if shared:
125
+ shared.leave()
126
+ return self.reply(429, {'error': 'Retrieval queue wait exceeded 30 seconds; retry later'})
96
127
  try:
97
128
  queued = round((time.monotonic()-start)*1000)
98
129
  generation = live.current(wait_ms/1000) if live else None
@@ -107,7 +138,8 @@ def serve(config):
107
138
  'retrievalMs':debug['elapsedMs'],'queueMs':queued,
108
139
  'serverElapsedMs':round((time.monotonic()-start)*1000),'queryCache':False,
109
140
  'index': {'mode':'live','identity':generation.identity,'freshness':'verified-after-search',
110
- 'completedAt':generation.info['completedAt']} if live else {'mode':'frozen'}}
141
+ 'completedAt':generation.info['completedAt'],
142
+ 'degradedFiles':generation.info.get('degradedFiles', 0)} if live else {'mode':'frozen'}}
111
143
  if body.get('trace'):
112
144
  response['diagnostics'] = debug
113
145
  self.reply(200,response)
@@ -118,18 +150,43 @@ def serve(config):
118
150
  self.reply(502,{'error':'Retrieval or model request failed'})
119
151
  finally:
120
152
  lock.release()
153
+ slots.release()
154
+ if shared:
155
+ shared.leave()
121
156
 
122
- server = ThreadingHTTPServer(('127.0.0.1',config.get('port',23505)),Handler)
157
+ server = LoopbackHTTPServer(('127.0.0.1',config.get('port',23505)),Handler)
123
158
  server.daemon_threads = True
124
- print(json.dumps({'listening':f'http://127.0.0.1:{server.server_port}',
125
- 'health':{'status':'running','mode':'live'} if live else health}),flush=True)
159
+ if shared:
160
+ shared.publish(server.server_port)
161
+ shared.monitor(server)
162
+ try:
163
+ print(json.dumps({'listening':f'http://127.0.0.1:{server.server_port}',
164
+ 'health':{'status':'running','mode':'live'} if live else health}),flush=True)
165
+ except OSError as error:
166
+ # Another client may already have attached through discovery when the
167
+ # launcher disappears. Its closed startup pipe must not kill that writer.
168
+ if not shared or error.errno not in (errno.EPIPE, errno.EINVAL):
169
+ raise
170
+ if shared:
171
+ # Shared workers outlive their launching MCP process and its stdio pipes.
172
+ sys.stdout = open(os.devnull, 'w')
173
+ sys.stderr = open(os.devnull, 'w')
126
174
  try:
127
175
  server.serve_forever()
128
176
  finally:
129
177
  server.server_close()
178
+ if shared:
179
+ shared.close()
130
180
  if live:
131
181
  live.close()
132
182
 
133
183
 
134
184
  if __name__ == '__main__':
135
- serve(json.loads(sys.stdin.readline()))
185
+ try:
186
+ serve(json.loads(sys.stdin.readline()))
187
+ except Exception as error:
188
+ code = 'INDEX_LOCKED' if isinstance(error, WriterBusy) else 'STARTUP_FAILED'
189
+ # Only known-safe diagnostics cross the startup protocol; never serialize config or keys.
190
+ message = str(error) if isinstance(error, WriterBusy) else type(error).__name__
191
+ print(json.dumps({'startupError': {'code': code, 'message': message}}), flush=True)
192
+ sys.exit(1)
package/src/client.mjs CHANGED
@@ -12,7 +12,7 @@ export function clientConfig(environment=process.env) {
12
12
 
13
13
  export async function search(query,{budget=4000,trace=false,freshnessWaitMs=30000,config=clientConfig(),signal}={}) {
14
14
  const started=performance.now();
15
- const timeout = AbortSignal.timeout(freshnessWaitMs + 30000);
15
+ const timeout = AbortSignal.timeout(freshnessWaitMs + 60000);
16
16
  const response=await fetch(`${config.baseUrl}/search`,{method:'POST',redirect:'error',signal:signal ? AbortSignal.any([signal,timeout]) : timeout,
17
17
  headers:{'content-type':'application/json',authorization:`Bearer ${config.apiKey}`},body:JSON.stringify({query,budget,trace,freshnessWaitMs})});
18
18
  if (!response.ok) {
@@ -20,6 +20,10 @@ export async function search(query,{budget=4000,trace=false,freshnessWaitMs=3000
20
20
  const body = await response.json();
21
21
  throw new Error(`Index unavailable: ${body.error || 'update pending'}`);
22
22
  }
23
+ if (response.status === 429) {
24
+ const body = await response.json();
25
+ throw new Error(`Retrieval busy: ${body.error || 'retry later'}`);
26
+ }
23
27
  throw new Error(`Retrieval HTTP ${response.status}`);
24
28
  }
25
29
  const result=await response.json();
package/src/config.mjs CHANGED
@@ -12,7 +12,7 @@ export const configKeys = new Set([
12
12
  'RERANK_BASE_URL','RERANK_API_KEY','RERANK_MODEL','RERANK_REMOTE_RUNTIME_URL','RERANK_REMOTE_HOST',
13
13
  'RERANK_SSH_TUNNEL_URL','RERANK_SSH_REMOTE',
14
14
  'OCE_ALLOW_HTTP','OCE_RERANK_API','OCE_RERANK_CONCURRENCY','OCE_RERANK_MAX_DOCUMENTS','OCE_EMBEDDING_DIMENSIONS',
15
- 'OCE_EMBEDDING_REVISION','OCE_PYTHON','OCE_GO_BINARY','OCE_LANGUAGE_OPTIONS','OCE_POLL_SECONDS',
15
+ 'OCE_EMBEDDING_REVISION','OCE_EMBEDDING_BATCH_SIZE','OCE_PYTHON','OCE_GO_BINARY','OCE_LANGUAGE_OPTIONS','OCE_POLL_SECONDS',
16
16
  'OCE_DEBOUNCE_SECONDS','OCE_API_KEY','OCE_BASE_URL','TIKTOKEN_CACHE_DIR',
17
17
  ]);
18
18
  export function configDirectory(environment = process.env) {
@@ -20,7 +20,7 @@ export function remoteRerankerConfig(env) {
20
20
  throw new Error('A user-approved remote reranker is required. No local model will be installed or started.');
21
21
  }
22
22
  const api = env.OCE_RERANK_API ?? 'rerank';
23
- if (!['rerank', 'rerank-batch'].includes(api)) throw new Error('OCE_RERANK_API must be rerank or rerank-batch');
23
+ if (!['rerank', 'rerank-batch', 'dashscope'].includes(api)) throw new Error('OCE_RERANK_API must be rerank, rerank-batch or dashscope');
24
24
  const integer = (name, fallback, maximum) => {
25
25
  const value = Number(env[name] ?? fallback);
26
26
  if (!Number.isInteger(value) || value < 1 || value > maximum) throw new Error(`${name} must be an integer from 1 to ${maximum}`);
package/src/mcp.mjs CHANGED
@@ -15,7 +15,7 @@ export function createMcpServer(config, {resolveConfig, automatic = false} = {})
15
15
  }
16
16
  const pathSchema = z.string().min(1).describe('Absolute path to the project directory. Required in automatic workspace mode.');
17
17
  const directoryPath = automatic ? pathSchema : pathSchema.optional();
18
- const server = new McpServer({name:'open-context-engine',version:'0.1.2'}, {
18
+ const server = new McpServer({name:'open-context-engine',version:'0.1.4'}, {
19
19
  instructions:workspaceInstructions + 'Search for source evidence. Results include source paths and line numbers. '
20
20
  + 'Search waits for saved file changes to be indexed. If an update is pending or fails, inspect index_status and retry after it completes. '
21
21
  + 'Read target files again before editing, because code may change after a search.',
@@ -34,7 +34,9 @@ export function createMcpServer(config, {resolveConfig, automatic = false} = {})
34
34
  const selected = await selectConfig(directory_path);
35
35
  const result = await search(query,{budget,freshnessWaitMs,config:selected,signal:extra.signal});
36
36
  if (result.index?.mode !== 'live') throw new Error('This service uses a frozen index; connect to a service started with --root');
37
- return {content:[{type:'text',text:result.context || 'No matching source context.'}],structuredContent:result};
37
+ const warning = result.index.degradedFiles
38
+ ? `Note: ${result.index.degradedFiles} file(s) indexed as plain text after syntax errors; inspect index_status for paths and locations.\n\n` : '';
39
+ return {content:[{type:'text',text:warning + (result.context || 'No matching source context.')}],structuredContent:result};
38
40
  } catch (error) {
39
41
  return {isError:true,content:[{type:'text',text:error.message}]};
40
42
  }
@@ -209,7 +209,7 @@ class BatchedEngine(Engine):
209
209
  'selectionAnchor':uid,'contextOnly':item not in retained})
210
210
  raw = '\n'.join(self.render(self.units[uid]) for uid in selected)
211
211
  end = time.monotonic()
212
- return raw,{'version':getattr(self, 'version', VERSION),'elapsedMs':round((end-start)*1000),'tokens':len(self.encoding.encode(raw)),
212
+ return raw,{'version':getattr(self, 'version', VERSION),'elapsedMs':round((end-start)*1000),'tokens':len(self.encoding.encode_ordinary(raw)),
213
213
  'candidateCount':len(candidates),'rerankedCount':len(retained),'expandedCount':len(expanded),
214
214
  'retention':retention,
215
215
  'modelRequests':{'embedding':1,'rerank':sum(w['requests'] for w in waves)},
@@ -130,7 +130,7 @@ class CascadeEngine(Engine):
130
130
  raw, selected = self.pack(retained, scores, affinity, budget)
131
131
  finished = time.monotonic()
132
132
  return raw, {'version': VERSION, 'elapsedMs': round((finished-start)*1000),
133
- 'tokens': len(self.encoding.encode(raw)), 'candidateCount': candidate_count,
133
+ 'tokens': len(self.encoding.encode_ordinary(raw)), 'candidateCount': candidate_count,
134
134
  'rerankedCount': len(retained), 'expandedCount': len(expanded),
135
135
  'modelRequests': {'embedding': 1, 'rerank': 1}, 'queryCache': False,
136
136
  'timingMs': {'embedding': round((embedded_at-start)*1000), 'recall': round((recalled_at-embedded_at)*1000),
@@ -56,7 +56,7 @@ class Engine:
56
56
  for u in units:
57
57
  for target in u['edges']:
58
58
  self.incoming[target].append(u['id'])
59
- self.costs = [len(self.encoding.encode(self.render(u))) + 2 for u in units]
59
+ self.costs = [len(self.encoding.encode_ordinary(self.render(u))) + 2 for u in units]
60
60
 
61
61
  @staticmethod
62
62
  def render(u):
@@ -153,7 +153,7 @@ class Engine:
153
153
  trace.append({'id': uid, 'overall': overall.get(uid, 0), 'facets': values.tolist(),
154
154
  'tokens': self.costs[uid], 'gain': gain, 'graphExpanded': uid in expanded})
155
155
  raw = '\n'.join(self.render(self.units[uid]) for uid in selected)
156
- return raw, {'elapsedMs': round((time.monotonic()-start)*1000), 'tokens': len(self.encoding.encode(raw)),
156
+ return raw, {'elapsedMs': round((time.monotonic()-start)*1000), 'tokens': len(self.encoding.encode_ordinary(raw)),
157
157
  'candidateCount': len(candidates), 'rerankedCount': len(retained), 'expandedCount': len(expanded),
158
158
  'selected': trace, 'plan': plan}
159
159
 
@@ -175,7 +175,7 @@ class EntityEngine(Engine):
175
175
  assembled, expanded_scores, provenance = self.complete(candidates,scores,dense)
176
176
  raw,selected = self.pack_entities(assembled,expanded_scores,dense,unit_dense,budget)
177
177
  finished = time.monotonic()
178
- return raw,{'version':VERSION,'elapsedMs':round((finished-start)*1000),'tokens':len(self.encoding.encode(raw)),
178
+ return raw,{'version':VERSION,'elapsedMs':round((finished-start)*1000),'tokens':len(self.encoding.encode_ordinary(raw)),
179
179
  'queryCache':False,'modelRequests':{'embedding':1,'rerank':1},'candidateCount':len(candidates),
180
180
  'expandedCount':len(provenance),'assembledCount':len(assembled),'policy':POLICY,'plan':plan,
181
181
  'timingMs':{'embedding':round((embedded-start)*1000),'recall':round((recalled-embedded)*1000),