xiaodcs-copilot-api 2.2.1-recovery.2 → 2.3.9-public.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,9 +1,19 @@
1
- # Copilot API Proxy
1
+ # Copilot API
2
+
3
+ <p align="center">
4
+ <img src="./docs/hero/copilot-api-hero.svg" alt="Copilot API - Universal AI Gateway" width="1600" />
5
+ </p>
6
+
7
+ <p align="center">
8
+ <strong>Universal AI Gateway</strong><br />
9
+ One Gateway. Any Client. Multiple AI Providers.<br />
10
+ Chat Completions &middot; OpenAI Responses &middot; Anthropic Messages
11
+ </p>
2
12
 
3
13
  > **XiaoDcs downstream build:** this public npm artifact is built from a
4
- > private downstream repository and adds Responses image-body budgeting and
5
- > stale compaction recovery. The original MIT-licensed project is
6
- > [caozhiyuan/copilot-api](https://github.com/caozhiyuan/copilot-api).
14
+ > private downstream repository and adds Responses image-body budgeting,
15
+ > stale compaction recovery, and Codex Fast tier routing. The original
16
+ > MIT-licensed project is [caozhiyuan/copilot-api](https://github.com/caozhiyuan/copilot-api).
7
17
 
8
18
  <p align="center">
9
19
  <a href="https://www.npmjs.com/package/xiaodcs-copilot-api"><img src="https://img.shields.io/npm/v/xiaodcs-copilot-api.svg" alt="npm version"></a>
@@ -13,80 +23,9 @@
13
23
  <a href="https://nodejs.org"><img src="https://img.shields.io/badge/Node-%3E%3D22.13.0-green.svg" alt="Node >= 22.13.0"></a>
14
24
  </p>
15
25
 
16
- English | [简体中文](./README.zh-CN.md)
17
-
18
- ## Table of Contents
19
-
20
- - [Copilot API Proxy](#copilot-api-proxy)
21
- - [Table of Contents](#table-of-contents)
22
- - [Important Notes](#important-notes)
23
- - [Project Overview](#project-overview)
24
- - [Quick Start](#quick-start)
25
- - [Features](#features)
26
- - [Prerequisites](#prerequisites)
27
- - [Installation](#installation)
28
- - [Running from Source](#running-from-source)
29
- - [Development Mode](#development-mode)
30
- - [Production Mode](#production-mode)
31
- - [Using with npx](#using-with-npx)
32
- - [Using with Docker](#using-with-docker)
33
- - [Electron Desktop App](#electron-desktop-app)
34
- - [Desktop App Screenshots](#desktop-app-screenshots)
35
- - [Using with Claude Code](#using-with-claude-code)
36
- - [Interactive Setup with `--claude-code` flag](#interactive-setup-with---claude-code-flag)
37
- - [Manual Configuration with `settings.json`](#manual-configuration-with-settingsjson)
38
- - [Using with OpenCode](#using-with-opencode)
39
- - [Minimal setup](#minimal-setup)
40
- - [Using with Codex](#using-with-codex)
41
- - [Codex `config.toml` Reference](#codex-configtoml-reference)
42
- - [GPT Tool Search](#gpt-tool-search)
43
- - [Plugin Integrations](#plugin-integrations)
44
- - [Claude Code plugin integration (marketplace-based)](#claude-code-plugin-integration-marketplace-based)
45
- - [Opencode plugin](#opencode-plugin)
46
- - [Using the Usage Viewer](#using-the-usage-viewer)
47
- - [Usage Viewer Screenshot](#usage-viewer-screenshot)
48
- - [Command Structure](#command-structure)
49
- - [Command Line Options](#command-line-options)
50
- - [Global Options](#global-options)
51
- - [Start Command Options](#start-command-options)
52
- - [Auth Command Options](#auth-command-options)
53
- - [Debug Command Options](#debug-command-options)
54
- - [Configuration (config.json)](#configuration-configjson)
55
- - [API Authentication](#api-authentication)
56
- - [API Endpoints](#api-endpoints)
57
- - [OpenAI Compatible Endpoints](#openai-compatible-endpoints)
58
- - [Codex Backend Proxy Endpoints](#codex-backend-proxy-endpoints)
59
- - [Anthropic Compatible Endpoints](#anthropic-compatible-endpoints)
60
- - [Usage Monitoring Endpoints](#usage-monitoring-endpoints)
61
- - [Admin / Configuration Endpoints](#admin--configuration-endpoints)
62
- - [Example Usage](#example-usage)
63
- - [Usage Tips](#usage-tips)
64
- - [CLAUDE.md or AGENTS.md Recommended Content](#claudemd-or-agentsmd-recommended-content)
65
-
66
- ## Important Notes
67
-
68
- > [!IMPORTANT]
69
- > **Before using, please be aware of the following:**
70
- >
71
- > 1. **Codex configuration:** When using with Codex, add the gateway provider to `~/.codex/config.toml`. See [Codex `config.toml` Reference](#codex-configtoml-reference).
72
- >
73
- > 2. **Claude Code configuration:** When using with Claude Code, please configure the model ID as `claude-opus-4-8[1m]`. Example claude `settings.json` see [Manual Configuration with `settings.json`](#manual-configuration-with-settingsjson).
74
- >
75
- > 3. **OpenCode configuration:** When using with OpenCode, configure `~/.config/opencode/opencode.json` with `@ai-sdk/anthropic`. See [Using with OpenCode](#using-with-opencode).
76
- >
77
- > 4. **Built-in `copilot`, `codex` and third-party providers:** Run `npx xiaodcs-copilot-api@latest auth` and choose `copilot`, `codex`, `deepseek`, `custom`, or other providers.
78
- >
79
- > 5. **Note:** See [GitHub Copilot Security Notice](./NOTICE.md#github-copilot-security-notice) for the warning removed from the README header.
80
-
81
- ---
82
-
83
- ## Project Overview
84
-
85
- A small AI gateway that can use GitHub Copilot, the built-in `codex` provider, or configured third-party providers such as DashScope. GitHub Copilot is optional: if no GitHub token is available, the server can still start in provider-only mode as long as at least one enabled provider is configured.
86
-
87
- The gateway exposes OpenAI- and Anthropic-compatible APIs from one local endpoint, so tools like [Claude Code](https://docs.anthropic.com/en/docs/claude-code/overview), OpenCode, Codex, and OpenAI-compatible clients can share the same local server.
88
-
89
- On the GitHub Copilot path, the gateway prefers Copilot's native Anthropic-style Messages API when available, preserving more Claude-native behavior for tool-heavy workflows.
26
+ <p align="center">
27
+ English | <a href="./README.zh-CN.md">简体中文</a>
28
+ </p>
90
29
 
91
30
  ## Quick Start
92
31
 
@@ -111,136 +50,56 @@ curl http://localhost:4141/v1/models
111
50
  > [!NOTE]
112
51
  > Token usage storage requires Node.js >= 22.13.0 or Bun. See [Using with npx](#using-with-npx) for details.
113
52
 
114
- From here, jump to the guide for your client: [Claude Code](#using-with-claude-code), [OpenCode](#using-with-opencode), [Codex](#using-with-codex), or run it with [Docker](#using-with-docker).
53
+ ### Public package traffic policy
115
54
 
116
- ## Features
55
+ The `xiaodcs-copilot-api` public package applies a fixed, process-local policy before any model upstream request begins:
117
56
 
118
- - **OpenAI and Anthropic compatibility**: Serve `/v1/responses`, `/v1/chat/completions`, `/v1/models`, `/v1/embeddings`, and `/v1/messages` from one local gateway.
119
- - **Copilot is optional**: Use GitHub Copilot when credentials are present, or run the server with only configured providers.
120
- - **One gateway for Copilot, `codex`, and external providers**: Route GitHub Copilot, the built-in `codex` provider, and configured third-party providers behind the same endpoint.
121
- - **Standalone third-party providers**: Configure providers such as DashScope, DeepSeek, OpenRouter, or a custom provider and start the gateway without a GitHub Copilot login.
122
- - **OpenAI-compatible providers on chat and Messages APIs**: `openai-compatible` providers can serve top-level `/v1/chat/completions` through `model: "provider/model"` and Anthropic-style `/v1/messages` through request/response translation.
123
- - **Agent-friendly Claude handling on Copilot**: Prefer native `/v1/messages` when available, preserve Claude-style tool flows, support Anthropic beta features, Claude WebSearch through Responses-capable models, and keep subagent/session markers intact.
124
- - **Claude Code and OpenCode integration**: Works with Claude Code and OpenCode, including direct Anthropic-compatible usage through `@ai-sdk/anthropic`.
125
- - **Flexible auth and deployment options**: Supports interactive login or direct tokens, individual/business/enterprise plans, GitHub Enterprise, opencode OAuth, and custom data directories.
126
- - **Multi-provider routing**: Expose provider-specific `/:provider/...` routes or use `model: "provider/model"` on the top-level API.
57
+ - at most 4 active upstream model requests;
58
+ - at most 1 request start per second and 20 starts per rolling minute;
59
+ - 300–900 ms of admission jitter;
60
+ - a FIFO queue of 16 requests with a 45-second wait limit;
61
+ - 5% pre-upstream load shedding, returned as HTTP `429` with `Retry-After`;
62
+ - image generation and editing endpoints are not registered. Image inputs to Chat Completions, Responses, and Messages remain supported where the model supports vision.
127
63
 
128
- ## Prerequisites
64
+ The counters are local to one running process. Separate machines or gateway processes have independent limits.
129
65
 
130
- - Bun (>= 1.2.x)
131
- - Node.js if you plan to run the published CLI with `npx`
132
- - GitHub account with Copilot subscription only if you want to use the GitHub Copilot provider
133
- - An API key or OAuth login for at least one configured provider if you want to run without GitHub Copilot
134
-
135
- ## Installation
136
-
137
- To install dependencies, run:
138
-
139
- ```sh
140
- bun install
141
- ```
142
-
143
- ## Running from Source
144
-
145
- The project can be run from source in several ways:
146
-
147
- ### Development Mode
148
-
149
- ```sh
150
- bun run dev start
151
- ```
152
-
153
- ### Production Mode
154
-
155
- ```sh
156
- bun run start start
157
- ```
158
-
159
- > The trailing `start` is the CLI subcommand passed to `src/main.ts`, not a typo: `bun run dev start` runs watch mode, `bun run start start` runs production.
160
-
161
- ## Using with npx
162
-
163
- You can run the project directly using npx:
164
-
165
- > [!IMPORTANT]
166
- > Token usage storage uses Node's built-in `node:sqlite` module when running with `npx`. It is enabled on Node.js >= 22.13.0. On Node.js < 22.13.0, the CLI still starts, but token usage storage is disabled.
167
- >
168
- > If you want token usage storage without upgrading Node.js, run the published CLI with Bun instead: `bunx --bun xiaodcs-copilot-api@latest start`.
169
-
170
- ```sh
171
- npx xiaodcs-copilot-api@latest start
172
- ```
173
-
174
- With options:
175
-
176
- ```sh
177
- npx xiaodcs-copilot-api@latest start --port 8080
178
- ```
179
-
180
- For authentication or provider configuration only:
181
-
182
- ```sh
183
- npx xiaodcs-copilot-api@latest auth
184
- ```
185
-
186
- To run without GitHub Copilot, configure at least one provider first, then start the server normally:
187
-
188
- ```sh
189
- npx xiaodcs-copilot-api@latest auth login --provider dashscope
190
- npx xiaodcs-copilot-api@latest start
191
- ```
192
-
193
- ## Using with Docker
194
-
195
- Build the image:
196
-
197
- ```sh
198
- docker build -t copilot-api .
199
- ```
200
-
201
- Run the container with a bind mount so auth data survives restarts:
202
-
203
- ```sh
204
- mkdir -p ./copilot-data
205
- docker run -p 4141:4141 -v $(pwd)/copilot-data:/root/.local/share/copilot-api copilot-api
206
- ```
207
-
208
- This stores GitHub auth data, provider config, and other gateway state in `./copilot-data` on the host, mapped to `/root/.local/share/copilot-api` in the container.
209
-
210
- Or pass a GitHub token directly:
211
-
212
- ```sh
213
- docker run -p 4141:4141 -e GH_TOKEN=your_github_token_here copilot-api
214
- ```
66
+ From here, jump to the guide for your client: [Claude Code](#using-with-claude-code), [OpenCode](#using-with-opencode), [Codex](#using-with-codex), or run it with [Docker](#using-with-docker).
215
67
 
216
- ## Electron Desktop App
68
+ ## Highlights
217
69
 
218
- If you prefer a GUI, this repository also includes an Electron desktop app in `desktop/`. It supports GitHub Copilot sign-in, OpenAI Codex OAuth, and API-key configuration for Kimi, DeepSeek, DashScope, OpenRouter, or a custom provider. After authorization or provider configuration, it can start and stop the local proxy with one click and shows the local endpoint, auth header, available models, usage, and logs in the app.
70
+ - **Unified API Gateway**: Serve OpenAI-compatible Chat Completions (`/v1/chat/completions`), the OpenAI Responses API (`/v1/responses`), and Anthropic-compatible Messages (`/v1/messages`) from one local endpoint.
71
+ - **Multi-Provider**: Route GitHub Copilot, the built-in `codex` provider, and third-party providers (Kimi, DeepSeek, DashScope, OpenRouter, OpenCode Go, or a custom provider) behind the same gateway. GitHub Copilot is optional — with at least one enabled provider, the server starts in provider-only mode without a GitHub token.
72
+ - **Coding Agent Ready**: First-class setups for Claude Code, OpenCode, and Codex, including the interactive `--claude-code` launcher and a merged model catalog for Codex.
73
+ - **Streaming & WebSocket**: SSE streaming on all three client-facing protocols. Upstream Copilot Responses traffic selects WebSocket or HTTP from each model's advertised endpoints; streamed Responses traffic for the built-in `codex` provider uses WebSocket by default and uses HTTP when `useResponsesApiWebSocket` is disabled.
74
+ - **Desktop App**: Electron GUI with GitHub Copilot sign-in, Codex OAuth, provider configuration, token usage, logs, and one-click start/stop.
219
75
 
220
- The settings screen also exposes `OAuth App`, `API Home`, `Enterprise URL`, verbose logging, and minimize-to-tray. Windows x64 (`.exe`), macOS Apple Silicon (`.dmg`), and Linux x64 (`.AppImage`) packages are published in GitHub Releases:
76
+ ## Compatibility
221
77
 
222
- https://github.com/caozhiyuan/copilot-api/releases
223
-
224
- On Linux, make the downloaded AppImage executable before launching it:
78
+ Every client talks to the same local endpoint. The gateway routes each request to GitHub Copilot, the built-in `codex` provider, or a configured third-party provider, translating between protocols when the provider speaks a different one.
225
79
 
226
- ```sh
227
- chmod +x Copilot-API-*-linux-x86_64.AppImage
228
- ./Copilot-API-*-linux-x86_64.AppImage
229
- ```
80
+ **Client / Protocol Matrix**
230
81
 
231
- Download the installer for your platform, authorize or configure a provider inside the app, choose a port, start the server, then point your client at the local endpoint shown in the app. Packaged desktop builds use the bundled Electron runtime, so normal desktop usage does not require installing Node.js separately. Token usage history is enabled when that bundled runtime supports SQLite.
82
+ | Client | Chat Completions | Responses | Anthropic Messages | Recommended |
83
+ |---|:---:|:---:|:---:|---|
84
+ | Claude Code | — | — | ✅ Native / Adapter | Anthropic Messages |
85
+ | OpenCode | ✅ Native | ✅ Native / Adapter | ✅ Native / Adapter via `@ai-sdk/anthropic` | Anthropic Messages |
86
+ | Codex | — | ✅ Native / Adapter | — | Responses |
87
+ | OpenAI-compatible clients | ✅ Native | ✅ Native / Adapter | — | Chat Completions |
88
+ | Anthropic-compatible clients | — | — | ✅ Native / Adapter | Anthropic Messages |
232
89
 
233
- The desktop app's Advanced Config page reads and writes the shared model mappings through `GET/POST /admin/config/model-mappings`. The same mappings apply across `POST /v1/messages`, `POST /v1/messages/count_tokens`, `POST /v1/responses`, and `POST /v1/chat/completions` instead of being split per interface. It uses `auth.adminApiKey` instead of the regular `auth.apiKeys`, and the app reads that key directly from `config.json` after the server has generated it on startup.
90
+ **Providers and protocols.** Protocol support is model-specific. Chat Completions requires a native endpoint, while Responses and Messages can use supported adapters. The built-in `codex` provider uses Responses natively; third-party providers can use `anthropic`, `openai-compatible`, or `openai-responses`, with per-model overrides.
234
91
 
235
- ### Desktop App Screenshots
92
+ ## Desktop App
236
93
 
237
- Main dashboard, token usage breakdown in the bundled Electron app:
94
+ Prefer a GUI? The Electron desktop app in `desktop/` covers GitHub Copilot sign-in, OpenAI Codex OAuth, and API-key configuration for Kimi, DeepSeek, DashScope, OpenRouter, or a custom provider — with one-click start/stop of the local server, and the local endpoint, auth header, available models, usage, and logs in one window.
238
95
 
239
96
  <p align="center">
240
97
  <img src="./docs/screenshots/desktop-dashboard.png" alt="Copilot API desktop app dashboard" width="49%" />
241
98
  <img src="./docs/screenshots/desktop-token-usage.png" alt="Copilot API desktop app token usage view" width="49%" />
242
99
  </p>
243
100
 
101
+ Windows x64 (`.exe`), macOS Apple Silicon (`.dmg`), and Linux x64 (`.AppImage`) packages are published in [GitHub Releases](https://github.com/caozhiyuan/copilot-api/releases). See [Electron Desktop App](#electron-desktop-app) for full setup and advanced configuration.
102
+
244
103
  ## Using with Claude Code
245
104
 
246
105
  This AI gateway can be used to power [Claude Code](https://docs.anthropic.com/en/claude-code), an experimental conversational AI assistant for developers from Anthropic.
@@ -404,6 +263,7 @@ base_url = "http://localhost:4141"
404
263
  env_key = "GITHUB_COPILOT_API_KEY"
405
264
  requires_openai_auth = true
406
265
  supports_websockets = false
266
+ supports_standalone_web_search = true
407
267
  wire_api = "responses"
408
268
  request_max_retries = 3
409
269
  stream_max_retries = 3
@@ -413,6 +273,7 @@ stream_idle_timeout_ms = 300000
413
273
  remote_compaction_v2 = true
414
274
  # optional: set false only when the model does not support tool_search
415
275
  apps = false
276
+ standalone_web_search = true
416
277
 
417
278
  [analytics]
418
279
  enabled = false
@@ -422,8 +283,62 @@ enabled = false
422
283
  > `name` must be set to `"OpenAI"`.
423
284
  >
424
285
  > For third-party models that do not support `tool_search`, we recommend disabling features.apps. Otherwise, each prompt may consume an additional 20,000 or more tokens.
286
+ >
287
+ > `supports_standalone_web_search` and `[features] standalone_web_search` must both be enabled to expose the standalone `web.run` search tool.
288
+
289
+ When Copilot exposes both a base Responses model and an exact `-fast` sibling
290
+ (for example, `gpt-5.6-sol` and `gpt-5.6-sol-fast`), the gateway presents them
291
+ to Codex as one model with a native **Fast** service tier. Selecting Fast makes
292
+ Codex send `service_tier: "priority"`; the gateway routes that request to the
293
+ paired Fast model and removes the unsupported field before forwarding it to
294
+ GitHub Copilot. The raw `/v1/models` endpoint continues to list both model IDs.
295
+ To make Fast the default for new Codex turns, add this top-level setting:
296
+
297
+ ```toml
298
+ service_tier = "fast"
299
+ ```
425
300
 
426
- When a Codex client (`User-Agent` starts with `codex`) requests the top-level `GET /v1/models`, the gateway merges native Codex models with models available through the Messages adapter. The latter advertise `use_responses_lite: true`: `/v1/responses` uses **Responses → Messages** for Anthropic providers, while OpenAI-compatible providers and Chat-only Copilot models reuse the existing Messages route for **Responses → Messages → Chat Completions**, then translate streaming or JSON results back to Responses.
301
+ ### If Codex Is Not Signed In to a GPT Account
302
+
303
+ ```toml
304
+ [model_providers.copilot_api]
305
+ name = "OpenAI"
306
+ base_url = "http://localhost:4141"
307
+ requires_openai_auth = false
308
+ supports_websockets = false
309
+ supports_standalone_web_search = true
310
+ wire_api = "responses"
311
+ request_max_retries = 3
312
+ stream_max_retries = 3
313
+ stream_idle_timeout_ms = 300000
314
+
315
+ [features]
316
+ standalone_web_search = true
317
+
318
+ [model_providers.copilot_api.auth]
319
+ command = "powershell.exe"
320
+ args = [
321
+ "-NoProfile",
322
+ "-NonInteractive",
323
+ "-Command",
324
+ "[Console]::Out.Write($env:GITHUB_COPILOT_API_KEY)"
325
+ ]
326
+ ```
327
+
328
+ macOS, replace the `auth` block with:
329
+
330
+ ```toml
331
+ [model_providers.copilot_api.auth]
332
+ command = "/bin/zsh"
333
+ args = [
334
+ "-c",
335
+ "printf '%s' \"$GITHUB_COPILOT_API_KEY\""
336
+ ]
337
+ ```
338
+
339
+ Without this configuration, Codex cannot fetch `/v1/models` while not signed in to a GPT account, so custom models are unavailable in the model picker.
340
+
341
+ When a Codex client (`User-Agent` starts with `codex`) requests the top-level `GET /v1/models`, the gateway merges native Codex models with models available through the Messages adapter. The latter advertise `use_responses_lite: true`, except DeepSeek models, which use `use_responses_lite: false` and `tool_mode: null`. For other models, `/v1/responses` uses **Responses → Messages** for Anthropic providers, while OpenAI-compatible providers and Chat-only Copilot models reuse the existing Messages route for **Responses → Messages → Chat Completions**, then translate streaming or JSON results back to Responses.
427
342
 
428
343
  The merged catalog is what Codex shows in its model picker, including the models exposed by your configured providers:
429
344
 
@@ -445,6 +360,138 @@ When Codex uses the top-level GitHub Copilot route with `approvals_reviewer = "a
445
360
 
446
361
  This mapping only applies to the top-level GitHub Copilot route. Provider-scoped routes do not use `modelMappings`, so the built-in `/codex` provider continues to handle `codex-auto-review` natively.
447
362
 
363
+ ---
364
+
365
+ ## Project Overview
366
+
367
+ A small AI gateway that can use GitHub Copilot, the built-in `codex` provider, or configured third-party providers such as DashScope. GitHub Copilot is optional: if no GitHub token is available, the server can still start in provider-only mode as long as at least one enabled provider is configured.
368
+
369
+ The gateway exposes OpenAI- and Anthropic-compatible APIs from one local endpoint, so tools like [Claude Code](https://docs.anthropic.com/en/docs/claude-code/overview), OpenCode, Codex, and OpenAI-compatible clients can share the same local server.
370
+
371
+ On the GitHub Copilot path, the gateway prefers Copilot's native Anthropic-style Messages API when available, preserving more Claude-native behavior for tool-heavy workflows.
372
+
373
+ ## Important Notes
374
+
375
+ > [!IMPORTANT]
376
+ > **Before using, please be aware of the following:**
377
+ >
378
+ > 1. **Codex configuration:** When using with Codex, add the gateway provider to `~/.codex/config.toml`. See [Codex `config.toml` Reference](#codex-configtoml-reference).
379
+ >
380
+ > 2. **Claude Code configuration:** When using with Claude Code, please configure the model ID as `claude-opus-4-8[1m]`. Example claude `settings.json` see [Manual Configuration with `settings.json`](#manual-configuration-with-settingsjson).
381
+ >
382
+ > 3. **OpenCode configuration:** When using with OpenCode, configure `~/.config/opencode/opencode.json` with `@ai-sdk/anthropic`. See [Using with OpenCode](#using-with-opencode).
383
+ >
384
+ > 4. **Built-in `copilot`, `codex` and third-party providers:** Run `npx xiaodcs-copilot-api@latest auth` and choose `copilot`, `codex`, `deepseek`, `custom`, or other providers.
385
+ >
386
+ > 5. **Note:** See [GitHub Copilot Security Notice](./NOTICE.md#github-copilot-security-notice) for the warning removed from the README header.
387
+
388
+ ## Prerequisites
389
+
390
+ - Bun (>= 1.2.x)
391
+ - Node.js if you plan to run the published CLI with `npx`
392
+ - GitHub account with Copilot subscription only if you want to use the GitHub Copilot provider
393
+ - An API key or OAuth login for at least one configured provider if you want to run without GitHub Copilot
394
+
395
+ ## Installation
396
+
397
+ To install dependencies, run:
398
+
399
+ ```sh
400
+ bun install
401
+ ```
402
+
403
+ ## Running from Source
404
+
405
+ The project can be run from source in several ways:
406
+
407
+ ### Development Mode
408
+
409
+ ```sh
410
+ bun run dev start
411
+ ```
412
+
413
+ ### Production Mode
414
+
415
+ ```sh
416
+ bun run start start
417
+ ```
418
+
419
+ > The trailing `start` is the CLI subcommand passed to `src/main.ts`, not a typo: `bun run dev start` runs watch mode, `bun run start start` runs production.
420
+
421
+ ## Using with npx
422
+
423
+ You can run the project directly using npx:
424
+
425
+ > [!IMPORTANT]
426
+ > Token usage storage uses Node's built-in `node:sqlite` module when running with `npx`. It is enabled on Node.js >= 22.13.0. On Node.js < 22.13.0, the CLI still starts, but token usage storage is disabled.
427
+ >
428
+ > If you want token usage storage without upgrading Node.js, run the published CLI with Bun instead: `bunx --bun xiaodcs-copilot-api@latest start`.
429
+
430
+ ```sh
431
+ npx xiaodcs-copilot-api@latest start
432
+ ```
433
+
434
+ With options:
435
+
436
+ ```sh
437
+ npx xiaodcs-copilot-api@latest start --port 8080
438
+ ```
439
+
440
+ For authentication or provider configuration only:
441
+
442
+ ```sh
443
+ npx xiaodcs-copilot-api@latest auth
444
+ ```
445
+
446
+ To run without GitHub Copilot, configure at least one provider first, then start the server normally:
447
+
448
+ ```sh
449
+ npx xiaodcs-copilot-api@latest auth login --provider dashscope
450
+ npx xiaodcs-copilot-api@latest start
451
+ ```
452
+
453
+ ## Using with Docker
454
+
455
+ Build the image:
456
+
457
+ ```sh
458
+ docker build -t copilot-api .
459
+ ```
460
+
461
+ Run the container with a bind mount so auth data survives restarts:
462
+
463
+ ```sh
464
+ mkdir -p ./copilot-data
465
+ docker run -p 4141:4141 -v $(pwd)/copilot-data:/root/.local/share/copilot-api copilot-api
466
+ ```
467
+
468
+ This stores GitHub auth data, provider config, and other gateway state in `./copilot-data` on the host, mapped to `/root/.local/share/copilot-api` in the container.
469
+
470
+ Or pass a GitHub token directly:
471
+
472
+ ```sh
473
+ docker run -p 4141:4141 -e GH_TOKEN=your_github_token_here copilot-api
474
+ ```
475
+
476
+ ## Electron Desktop App
477
+
478
+ If you prefer a GUI, this repository also includes an Electron desktop app in `desktop/`. It supports GitHub Copilot sign-in, OpenAI Codex OAuth, and API-key configuration for Kimi, DeepSeek, DashScope, OpenRouter, or a custom provider. After authorization or provider configuration, it can start and stop the local proxy with one click and shows the local endpoint, auth header, available models, usage, and logs in the app.
479
+
480
+ The settings screen also exposes `OAuth App`, `API Home`, `Enterprise URL`, verbose logging, and minimize-to-tray. Windows x64 (`.exe`), macOS Apple Silicon (`.dmg`), and Linux x64 (`.AppImage`) packages are published in GitHub Releases:
481
+
482
+ https://github.com/caozhiyuan/copilot-api/releases
483
+
484
+ On Linux, make the downloaded AppImage executable before launching it:
485
+
486
+ ```sh
487
+ chmod +x Copilot-API-*-linux-x86_64.AppImage
488
+ ./Copilot-API-*-linux-x86_64.AppImage
489
+ ```
490
+
491
+ Download the installer for your platform, authorize or configure a provider inside the app, choose a port, start the server, then point your client at the local endpoint shown in the app. Packaged desktop builds use the bundled Electron runtime, so normal desktop usage does not require installing Node.js separately. Token usage history is enabled when that bundled runtime supports SQLite.
492
+
493
+ The desktop app's Advanced Config page reads and writes the shared model mappings through `GET/POST /admin/config/model-mappings`. The same mappings apply across `POST /v1/messages`, `POST /v1/messages/count_tokens`, `POST /v1/responses`, and `POST /v1/chat/completions` instead of being split per interface. It uses `auth.adminApiKey` instead of the regular `auth.apiKeys`, and the app reads that key directly from `config.json` after the server has generated it on startup.
494
+
448
495
  ## GPT Tool Search
449
496
 
450
497
  For GPT Responses models such as `gpt-5.4+`, this AI gateway can expose Responses `tool_search` through a small MCP bridge. The same bridge can be used by Claude Code and opencode, as long as the client loads MCP servers and sends Anthropic Messages traffic through this gateway.
@@ -565,7 +612,7 @@ The dashboard provides a user-friendly interface to view your Copilot usage data
565
612
  > Token usage history requires Bun or Node.js >= 22.13.0. On Node.js < 22.13.0, the server runs normally but token usage storage is disabled.
566
613
 
567
614
  - **API Endpoint URL**: The dashboard is pre-configured to fetch data from your local server endpoint via a URL query parameter. You can manually switch this to any other compatible API endpoint.
568
- - **x-api-key Authentication**: If API Key authentication is enabled, you can provide the `x-api-key` request header. The key is persisted in the browser's local storage.
615
+ - **API Key Authentication**: If API Key authentication is enabled, enter a raw API key (sent as the `x-api-key` header) or `Authorization: Bearer <key>`. Credentials are remembered in the browser's local storage per endpoint origin, and switching to a different endpoint origin does not automatically send the previous credential.
569
616
  - **Period Selector**: Choose from Day, Week, or Month time ranges. The URL query parameter updates automatically when you switch, making it easy to bookmark and share.
570
617
  - **Fetch Data**: Click the "Refresh" button to load or refresh the usage data. The dashboard also fetches data automatically on page load.
571
618
  - **Copilot Quotas**: View quota usage for services such as Chat and Completions via progress bars. Hover over a card to see used/remaining details.
@@ -626,10 +673,12 @@ The following command line options are available for the `start` command:
626
673
 
627
674
  Use `copilot-api auth login --provider copilot` only when you want to enable the GitHub Copilot provider. Copilot is not required for `codex` or third-party provider-only usage.
628
675
 
629
- Use `copilot-api auth login --provider deepseek`, `--provider dashscope`, `--provider openrouter`, `--provider opencode-go`, or `--provider kimi` to add or update those common third-party providers from the CLI. DeepSeek prompts for masked `apiKey`, provider `type` (default `anthropic`), and `baseUrl` defaulting to `https://api.deepseek.com/anthropic`. DashScope prompts for masked `apiKey`, provider `type` (default `openai-compatible`), and prefilled `baseUrl`. OpenRouter prompts for masked `apiKey` and prefilled `baseUrl` only, and writes `type: "anthropic"`. OpenCode Go prompts for masked `apiKey` and prefilled `baseUrl` only, and writes `type: "openai-compatible"` (baseUrl `https://opencode.ai/zen/go`). Kimi prompts for masked `apiKey` and prefilled `baseUrl` only, and writes `type: "openai-compatible"` (baseUrl `https://api.kimi.com/coding`). OpenCode Go additionally routes built-in `qwen*` and `minimax*` models through Anthropic Messages and `gpt*`/`grok*` models through OpenAI Responses; other models keep the OpenAI-compatible default. After a provider is configured and enabled, `copilot-api start` can run without any GitHub token.
676
+ Use `copilot-api auth login --provider deepseek`, `--provider dashscope`, `--provider openrouter`, `--provider opencode-go`, or `--provider kimi` to add or update those common third-party providers from the CLI. DeepSeek prompts for masked `apiKey`, provider `type` (default `anthropic`), and `baseUrl` defaulting to `https://api.deepseek.com/anthropic`. DashScope prompts for masked `apiKey`, provider `type` (default `openai-compatible`), and prefilled `baseUrl`. OpenRouter prompts for masked `apiKey` and prefilled `baseUrl` only, and writes `type: "anthropic"`. OpenCode Go prompts for masked `apiKey` and prefilled `baseUrl` only, and writes `type: "openai-compatible"` (baseUrl `https://opencode.ai/zen/go`). Kimi prompts for masked `apiKey`, provider `type` (default `openai-compatible`), and `baseUrl` defaulting to `https://api.kimi.com/coding` (the same base URL serves both the Anthropic and OpenAI-compatible endpoints). OpenCode Go additionally routes built-in `qwen*` and `minimax*` models through Anthropic Messages and `gpt*`/`grok*`/`muse-spark*` models through OpenAI Responses; other models keep the OpenAI-compatible default. After a provider is configured and enabled, `copilot-api start` can run without any GitHub token.
630
677
 
631
678
  Use `copilot-api auth login --provider custom` to add or update another third-party provider from the CLI. The command prompts for the provider name, supported type (`anthropic`, `openai-compatible`, or `openai-responses`), `baseUrl`, masked `apiKey`, and `authType`; `authType` may be left as the type default or set to `x-api-key` / `authorization`.
632
679
 
680
+ Gateway API keys live under `auth.apiKeys` in `config.json`. Manage them with `copilot-api auth keys` (one operation per invocation): add a key with `--add <key>`, remove one with `--remove <key>`, list all with `--list`, or clear them all with `--clear`. Clients authenticate with any configured key via `x-api-key` or `Authorization: Bearer`. When no keys are configured, `copilot-api start` starts with authentication bypassed and prints a startup info message.
681
+
633
682
  ### Debug Command Options
634
683
 
635
684
  | Option | Description | Default | Alias |
@@ -667,8 +716,7 @@ Use `copilot-api auth login --provider custom` to add or update another third-pa
667
716
  "useResponsesApiCompactionRecovery": false,
668
717
  "useResponsesApiWebSocket": true,
669
718
  "responsesTransport": {
670
- "compactHeadersTimeoutMs": 180000,
671
- "headersTimeoutMs": 30000,
719
+ "headersTimeoutMsV2": 300000,
672
720
  "streamInactivityTimeoutMs": 300000,
673
721
  "websocketOpenTimeoutMs": 30000,
674
722
  "websocketPoolIdleTimeoutMs": 60000,
@@ -685,7 +733,7 @@ Use `copilot-api auth login --provider custom` to add or update another third-pa
685
733
  - **auth.adminApiKey:** Single admin key used only for `/admin/*` routes. If missing, the server generates a random key at startup and writes it back to `config.json`. Requests use the same `x-api-key` or `Authorization: Bearer` headers, but regular `auth.apiKeys` never grant access to `/admin/*`.
686
734
  - **modelMappings:** Exact `sourceModel -> targetModel` rewrites shared by top-level `POST /v1/messages`, `POST /v1/messages/count_tokens`, `POST /v1/responses`, and `POST /v1/chat/completions` requests. Omit it or leave it as `{}` to disable rewrites. Both the source and target must be non-empty strings. Targets can be regular model IDs or `provider/model` aliases such as `dashscope/qwen3.6-plus`, and the rewrite happens before provider alias parsing. These mappings are not split per interface. The admin endpoints `GET/POST /admin/config/model-mappings` read and update only this field.
687
735
  - **extraPrompts:** Map of `model -> prompt` appended to the first system prompt when translating Anthropic-style requests to Responses API. Use this to inject guardrails or guidance per model. Missing default entries are auto-added without overwriting your custom prompts. For GPT-5.3+ models (e.g. `gpt-5.3-codex`, `gpt-5.4`, `gpt-5.5`), a built-in commentary prompt is used as fallback when not explicitly configured. The built-in prompts enable phase-aware commentary, which lets the model emit a short user-facing progress update before tools or deeper reasoning.
688
- - **providers:** Global upstream provider map. Each provider key (for example `dashscope`) becomes a route prefix (`/dashscope/v1/messages`). Supports `type: "anthropic"`, `type: "openai-compatible"`, and `type: "openai-responses"`. Top-level clients can also use `model: "dashscope/model-id"` with `/v1/messages`, `/v1/messages/count_tokens`, `/v1/responses`, and `/v1/chat/completions`; the gateway strips the `dashscope/` prefix before forwarding upstream. The `/v1/responses` route for `anthropic` and `openai-compatible` providers uses the Responses Lite → Messages adapter; `openai-compatible` providers then reuse the Messages → Chat translation. Codex clients (`User-Agent` starting with `codex`) also use the adapter for non-`gpt-*` models on `openai-responses` providers. `GET /v1/models` aggregates enabled provider models with `provider/model-id` IDs, while the top-level Codex-UA catalog also merges these adaptable models as `use_responses_lite` entries. Use `GET /dashscope/v1/models` for a single provider's raw model list.
736
+ - **providers:** Global upstream provider map. Each provider key (for example `dashscope`) becomes a route prefix (`/dashscope/v1/messages`). Supports `type: "anthropic"`, `type: "openai-compatible"`, and `type: "openai-responses"`. Top-level clients can also use `model: "dashscope/model-id"` with `/v1/messages`, `/v1/messages/count_tokens`, `/v1/responses`, and `/v1/chat/completions`; the gateway strips the `dashscope/` prefix before forwarding upstream. The `/v1/responses` route for `anthropic` and `openai-compatible` providers uses the Responses Lite → Messages adapter; `openai-compatible` providers then reuse the Messages → Chat translation. Codex clients (`User-Agent` starting with `codex`) also use the adapter for non-`gpt-*` models on `openai-responses` providers. `GET /v1/models` aggregates enabled provider models with `provider/model-id` IDs, while the top-level Codex-UA catalog also merges these adaptable models as `use_responses_lite` entries (except DeepSeek models, which use `use_responses_lite: false` and `tool_mode: null`). Use `GET /dashscope/v1/models` for a single provider's raw model list.
689
737
  - `enabled` defaults to `true` if omitted.
690
738
  - `baseUrl` should be provider API base URL without the final endpoint. For Anthropic providers, omit `/v1/messages`; for OpenAI-compatible providers, omit `/v1/chat/completions`; for OpenAI Responses providers, omit `/v1/responses`.
691
739
  - `apiKey` is used as the upstream credential value and is required for regular providers.
@@ -699,7 +747,7 @@ Use `copilot-api auth login --provider custom` to add or update another third-pa
699
747
  - `pricing` (optional): Per-model token prices, in the provider `pricingCurrency`, per 1M tokens. Supported fields are `input`, `output`, `cachedInput` (implicit cache read), `explicitCachedInput` (explicit cache read), and `cacheCreationInput`. Use `tiers` with `maxInputTokens` for input-size tiered pricing.
700
748
  - `contextCache` (optional): Defaults to `true` for providers whose name is `dashscope` or whose `baseUrl` contains `aliyuncs.com`; defaults to `false` for other OpenAI-compatible providers. This enables Alibaba Cloud Model Studio/DashScope explicit context cache by injecting `cache_control: { "type": "ephemeral" }` on up to 4 content blocks using the Context Cache format. The cache breakpoint strategy matches opencode's main provider flow: the first 2 system messages plus the last 2 non-system messages. Marked string content is converted to text content part arrays for `system` / `user` / `assistant` / `tool` messages; existing array content is marked on the last part. Set this to `false` when the model already supports implicit caching, or when the upstream does not accept this explicit-cache extension field. Set this to `true` for non-DashScope providers that support the same explicit-cache extension. Applied on both `/v1/messages` and `/v1/chat/completions` routes.
701
749
  - `supportPdf` (optional): Controls whether the model supports PDF/document content. Defaults to `false`; unsupported PDFs are converted to a text notice. Set it to `true` to send PDF/document blocks as OpenAI Chat Completions file parts.
702
- - `toolContentSupportType` (optional): Tool result content capabilities for that model, as an array of `array`, `image`, and `pdf`. Provider routes default to string-only tool content when omitted. If `supportPdf` is `true` but this list does not include `pdf`, file parts in tool results are moved to user role messages. This provider default does not change the Copilot main flow, which continues to support array + image and not PDF.
750
+ - `toolContentSupportType` (optional): Tool result content capabilities for that model, as an array of `array`, `image`, and `pdf`. Provider routes default to string-only tool content when omitted. If `supportPdf` is `true` but this list does not include `pdf`, file parts in tool results are moved to user role messages. The Copilot main flow uses the same string-only default, because some Copilot models do not support array or image tool content either.
703
751
  - `type` (optional): Per-model override of the provider protocol type. Supports `anthropic`, `openai-compatible`, and `openai-responses`. When set, the provider's `/v1/messages` route uses this model's type instead of the provider-level type for request routing, auth header resolution, and upstream endpoint selection. This is useful for providers like OpenCode Go whose upstream supports both OpenAI-compatible and Anthropic Messages APIs for different models. When the type is overridden, the auth header is resolved from the overridden type's default (Anthropic defaults to `x-api-key`; OpenAI-compatible/Responses default to `authorization`).
704
752
  - `contextWindow` (optional): Context window token limit advertised when this model is merged into the Codex-UA model catalog; for example, `1000000` declares a 1M-token context window. Missing configured values use upstream metadata first, then the built-in non-GPT model catalog, then `256000`.
705
753
  - `maxOutputTokens` (optional): Maximum output token limit advertised in the Codex-UA model catalog. Missing configured values use upstream metadata first, then the built-in non-GPT model catalog, where defaults are capped at `64000`, then `32000`.
@@ -713,10 +761,10 @@ Use `copilot-api auth login --provider custom` to add or update another third-pa
713
761
  - **Priority:** request `output_config.effort` > `modelReasoningEfforts[model]` > built-in default (`xhigh` for GPT-5.3+ models, otherwise `high`).
714
762
  - **Forwarding:** the resolved value remains `output_config.effort` for the Copilot native Messages API and becomes `reasoning.effort` when translated to the Responses API.
715
763
  - **Configuration values:** `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, and `max`.
716
- - **useMessagesApi:** When `true`, Claude-family models that support Copilot's native `/v1/messages` endpoint will use the Messages API; otherwise they fall back to `/chat/completions`. Set to `false` to disable Messages API routing and always use `/chat/completions`. Defaults to `true`.
717
- - **useResponsesApiCompactionRecovery (experimental):** When `true`, successful remote Responses compactions enqueue a non-blocking, low-reasoning shadow-summary request and store only that summary under a hash of the opaque compaction item in the private `~/.local/share/copilot-api/compaction-recovery/cache.sqlite` directory. If Copilot later rejects that compaction with an encrypted-content `invalid_request_body`, or returns `input item does not belong to this connection` after the active Copilot credential changes, HTTP requests prefer the cached Codex-compatible summary. For uncached legacy checkpoints, recovery progressively removes dependent opaque reasoning and, only if the explicit opaque-item error remains, replaces the unusable checkpoint while preserving visible messages and tool state; this may require up to two retries. Subsequent requests replace cached items already marked invalid before forwarding. The cache is bounded to 128 summaries, expires entries after 7 days, and never stores the original request or encrypted content. This option defaults to `false` because each remote compaction causes one additional background model request, which is included in token-usage accounting, and every fallback is lossy. WebSocket errors are marked for recovery on the client's next retry; automatic same-request recovery is HTTP-only.
718
- - **useResponsesApiWebSocket:** When `true`, Responses API requests use Copilot's websocket transport for models that advertise `ws:/responses`; models that only advertise `/responses` continue to use HTTP. Set to `false` to disable websocket routing and use HTTP `/responses` whenever the selected model supports it. Defaults to `true`. If the Responses API WebSocket gets closed, it is usually caused by your own network. If you are using a VPN, try switching to a different node.
719
- - **responsesTransport:** Positive integer lifecycle and buffering limits for every upstream Responses transport. Invalid, zero, or negative values fall back to the defaults shown above. `headersTimeoutMs` covers connection setup through receipt of HTTP response headers for ordinary requests; it is not a total generation deadline. `compactHeadersTimeoutMs` provides a separate, longer time-to-first-header limit for remote compaction requests, which always use HTTP and can remain queued longer under load. `streamInactivityTimeoutMs` is reset by every HTTP body chunk or WebSocket message, allowing long generations to continue while they remain active. `websocketOpenTimeoutMs` limits the WebSocket handshake, while `websocketPoolIdleTimeoutMs` controls only completed, reusable pooled sockets. The byte and message limits bound queued WebSocket events; exceeding either limit fails that stream and invalidates its socket rather than dropping or reordering events.
764
+ - **useMessagesApi:** When `true`, models that advertise Copilot's native `/v1/messages` endpoint use the Messages API. If Messages is disabled or unavailable for the selected model, the gateway uses Responses when that model advertises a Responses endpoint, then falls back to Chat Completions when supported. Set this to `false` to skip native Messages routing. Defaults to `true`.
765
+ - **useResponsesApiCompactionRecovery (experimental):** When `true`, successful remote Responses compactions enqueue a non-blocking, low-reasoning shadow-summary request and store only that summary under a hash of the opaque compaction item in the private `~/.local/share/copilot-api/compaction-recovery/cache.sqlite` directory. If Copilot later rejects that compaction or connection-bound history, HTTP requests progressively rebuild the request from cached summaries and visible messages. This option defaults to `false` because shadow summaries add one background model request and recovery is lossy. WebSocket errors are recovered on the client's next retry; automatic same-request recovery is HTTP-only.
766
+ - **useResponsesApiWebSocket:** When `true`, Copilot Responses requests use WebSocket for models that advertise `ws:/responses`; models that advertise only `/responses` use HTTP. Streamed Responses requests for the built-in `codex` provider use WebSocket whenever this setting is enabled, while non-streaming Codex requests always use HTTP. Set this to `false` to make Copilot use HTTP `/responses` where the selected model advertises it and to send streamed Codex Responses requests over HTTP. WebSocket failures are not retried automatically over HTTP. Defaults to `true`. If a proxy, VPN, or network blocks or destabilizes WebSocket traffic, disable this setting or switch networks.
767
+ - **responsesTransport:** Positive integer lifecycle and buffering limits for every upstream Responses transport. Invalid, zero, or negative values fall back to the defaults shown above. `headersTimeoutMsV2` covers connection setup through receipt of HTTP response headers; it is not a total generation deadline. `streamInactivityTimeoutMs` is reset by every HTTP body chunk or WebSocket message, allowing long generations to continue while they remain active. `websocketOpenTimeoutMs` limits the WebSocket handshake, while `websocketPoolIdleTimeoutMs` controls only completed, reusable pooled sockets. The byte and message limits bound queued WebSocket events; exceeding either limit fails that stream and invalidates its socket rather than dropping or reordering events.
720
768
  - **useResponsesApiWebSearch:** When `true`, the server keeps Responses API tools with `type: "web_search"` and forwards them upstream. Set to `false` to strip those tools from `/responses` payloads. Defaults to `true`.
721
769
  - **alphaSearchCodexPriority:** Defaults to `true`. Top-level alpha-search requests prefer the Codex alpha-search endpoint because it does not consume provider quota. If Codex is unavailable, or this setting is `false`, requests with a `provider/model` alias other than `codex/model` use that provider's `/v1/responses` endpoint, and requests without a provider prefix use GitHub Copilot Responses web search. The adapter recognizes every current Codex search command; unsupported `image_query` and `screenshot` operations return successful no-retry tool output.
722
770
  - **alphaSearchModel:** Native Responses search model used when a Messages-backed Responses Lite model cannot run Responses web search directly. Defaults to `gpt-5-mini`; it may be a regular Copilot model or an `openai-responses` `provider/model` alias. Set it to an empty string to disable this redirect, in which case alpha-search requests for those models return an invalid-request error.
@@ -753,7 +801,7 @@ curl http://localhost:4141/admin/config/model-mappings \
753
801
 
754
802
  ## API Endpoints
755
803
 
756
- The server exposes several OpenAI- and Anthropic-compatible endpoints. Requests can target GitHub Copilot, the built-in `codex` provider, or configured providers depending on the selected model and `provider/model` alias.
804
+ The server exposes several OpenAI- and Anthropic-compatible endpoints. Requests can target GitHub Copilot, the built-in `codex` provider, or configured providers depending on the selected model and `provider/model` alias. Every `/v1/...` endpoint below also supports a provider-scoped path in the form `/:provider/v1/...`; those variants are omitted from the tables.
757
805
 
758
806
  ### OpenAI Compatible Endpoints
759
807
 
@@ -766,33 +814,24 @@ These endpoints mimic the OpenAI API structure.
766
814
  | `GET /v1/models` | `GET` | Lists Copilot models plus enabled provider models using `provider/model-id` IDs. Requests from Codex clients (`User-Agent` beginning with `codex`) are forwarded to the Codex Models upstream. |
767
815
  | `POST /v1/embeddings` | `POST` | Creates an embedding vector representing the input text. |
768
816
 
769
- ### Codex Backend Proxy Endpoints
817
+ ### Codex Backend Endpoints
770
818
 
771
- These endpoints require an active Codex login. Each endpoint is available both without a version prefix and under `/v1`.
819
+ These endpoints implement the supported Codex backend APIs. Alpha search can use either the Codex backend or a Responses web-search adapter. Image generation and editing are disabled in this public package.
772
820
 
773
821
  | Endpoint | Method | Description |
774
822
  | -------------------------------------------------------------- | ------ | --------------------------------------------------------------- |
775
- | `POST /alpha/search`<br>`POST /v1/alpha/search` | `POST` | Transparently forwards the JSON body and query parameters to the Codex Alpha Search upstream. |
776
- | `POST /images/generations`<br>`POST /v1/images/generations` | `POST` | Forwards a JSON image generation request to the Codex Images upstream. When the request omits `Content-Type`, the gateway defaults it to `application/json`. |
777
- | `POST /images/edits`<br>`POST /v1/images/edits` | `POST` | Forwards an image edit request to the Codex Images upstream. Send this request as `multipart/form-data` and let the HTTP client generate the `boundary`; the gateway preserves the incoming content type and streams the upload body. |
823
+ | `POST /v1/alpha/search` | `POST` | Routes Codex alpha-search requests to the Codex backend, or handles supported commands locally and through Responses web search. |
778
824
 
779
- For every endpoint above, the gateway replaces client authorization and account headers with the active Codex login, preserves query parameters and compatible request headers, and returns the upstream status, headers, and body.
825
+ For requests routed to the Codex backend, the gateway replaces client authorization and account headers with the active Codex login and preserves compatible request metadata. Responses-backed alpha search instead follows the selected Copilot or provider route.
780
826
 
781
827
  ### Anthropic Compatible Endpoints
782
828
 
783
- These endpoints are designed to be compatible with the Anthropic Messages API. Provider-scoped models, Responses, alpha-search, and images routes accept both unversioned and `/v1` paths; Messages routes remain under `/v1`.
829
+ These endpoints are designed to be compatible with the Anthropic Messages API.
784
830
 
785
831
  | Endpoint | Method | Description |
786
832
  | -------------------------------- | ------ | ------------------------------------------------------------ |
787
833
  | `POST /v1/messages` | `POST` | Creates a model response for a given conversation. Supports `provider/model` aliases for configured providers, including translation through `openai-compatible` providers. |
788
834
  | `POST /v1/messages/count_tokens` | `POST` | Calculates the number of tokens for a given set of messages. Supports `provider/model` aliases for configured providers. |
789
- | `POST /:provider/v1/messages` | `POST` | Proxies Anthropic Messages requests to the configured Anthropic provider, translates them through an OpenAI-compatible provider, or translates them through an OpenAI Responses provider. |
790
- | `GET /:provider/models`<br>`GET /:provider/v1/models` | `GET` | Proxies model listing requests to the configured provider. For `codex`, returns the built-in catalog by default; Codex clients (`User-Agent` starting with `codex`) are forwarded to the Codex Models upstream. |
791
- | `POST /:provider/v1/messages/count_tokens` | `POST` | Calculates tokens locally for provider route requests. |
792
- | `POST /:provider/responses`<br>`POST /:provider/v1/responses` | `POST` | Proxies OpenAI Responses requests to a configured `openai-responses` provider (including `codex`). |
793
- | `POST /:provider/alpha/search`<br>`POST /:provider/v1/alpha/search` | `POST` | Proxies alpha-search requests. For `codex`, forwards to the Codex Alpha Search upstream; for other providers, forwards to `{baseUrl}/v1/alpha/search`. |
794
- | `POST /:provider/images/generations`<br>`POST /:provider/v1/images/generations` | `POST` | Proxies image generation. For `codex`, uses the Codex Images upstream; for other providers, forwards to `{baseUrl}/v1/images/generations` (15-minute timeout). |
795
- | `POST /:provider/images/edits`<br>`POST /:provider/v1/images/edits` | `POST` | Proxies image edits. For `codex`, uses the Codex Images upstream; for other providers, forwards multipart/streamed bodies to `{baseUrl}/v1/images/edits` (15-minute timeout). |
796
835
 
797
836
  ### Usage Monitoring Endpoints
798
837
 
@@ -801,7 +840,6 @@ New endpoints for monitoring your Copilot usage and quotas.
801
840
  | Endpoint | Method | Description |
802
841
  | ------------ | ------ | ------------------------------------------------------------ |
803
842
  | `GET /usage` | `GET` | Get detailed Copilot usage statistics and quota information. |
804
- | `GET /token` | `GET` | Get the current Copilot token being used by the API. |
805
843
 
806
844
  ### Admin / Configuration Endpoints
807
845