@jeffreycao/copilot-api 2.6.16 → 2.6.18
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +37 -846
- package/README.zh-CN.md +39 -890
- package/dist/{auth-Bs_UWxAm.js → auth-B1ZRVwUG.js} +2 -2
- package/dist/{auth-Bs_UWxAm.js.map → auth-B1ZRVwUG.js.map} +1 -1
- package/dist/auth-YGms_d01.js +2 -0
- package/dist/main.js +2 -2
- package/dist/{models-bh4G4bGm.js → models-DH6XfUL5.js} +2 -2
- package/dist/{models-bh4G4bGm.js.map → models-DH6XfUL5.js.map} +1 -1
- package/dist/{server-DQy_3eXx.js → server-BPn42Ogo.js} +13 -11
- package/dist/server-BPn42Ogo.js.map +1 -0
- package/dist/{start-CmHlGpmG.js → start-DGRUGnBh.js} +5 -5
- package/dist/{start-CmHlGpmG.js.map → start-DGRUGnBh.js.map} +1 -1
- package/dist/{token-BUnKN_vY.js → token-BKpZGZR0.js} +3 -2
- package/dist/token-BKpZGZR0.js.map +1 -0
- package/package.json +1 -1
- package/dist/auth-ZoN2_BLX.js +0 -2
- package/dist/server-DQy_3eXx.js.map +0 -1
- package/dist/token-BUnKN_vY.js.map +0 -1
package/README.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Copilot API
|
|
2
2
|
|
|
3
3
|
<p align="center">
|
|
4
|
-
<img src="
|
|
4
|
+
<img src="docs/hero/copilot-api-hero.svg" alt="Copilot API - Universal AI Gateway" width="1600" />
|
|
5
5
|
</p>
|
|
6
6
|
|
|
7
7
|
<p align="center">
|
|
@@ -19,12 +19,22 @@
|
|
|
19
19
|
</p>
|
|
20
20
|
|
|
21
21
|
<p align="center">
|
|
22
|
-
English | <a href="
|
|
22
|
+
English | <a href="README.zh-CN.md">简体中文</a>
|
|
23
23
|
</p>
|
|
24
24
|
|
|
25
|
+
Copilot API is a local AI gateway that connects Claude Code, OpenCode, Codex, and other clients to GitHub Copilot, the built-in Codex provider, and third-party model providers through a unified API.
|
|
26
|
+
|
|
27
|
+
## Highlights
|
|
28
|
+
|
|
29
|
+
- **Unified API Gateway**: Serve OpenAI-compatible Chat Completions (`/v1/chat/completions`), the OpenAI Responses API (`/v1/responses`), and Anthropic-compatible Messages (`/v1/messages`) from one local endpoint.
|
|
30
|
+
- **Multi-Provider**: Route GitHub Copilot, the built-in `codex` provider, and third-party providers (Kimi, DeepSeek, DashScope, OpenRouter, OpenCode Go, or a custom provider) behind the same gateway. GitHub Copilot is optional — with at least one enabled provider, the server starts in provider-only mode without a GitHub token.
|
|
31
|
+
- **Coding Agent Ready**: First-class setups for Claude Code, OpenCode, and Codex, including the interactive `--claude-code` launcher and a merged model catalog for Codex.
|
|
32
|
+
- **Streaming & WebSocket**: SSE streaming on all three client-facing protocols. Upstream Copilot Responses traffic selects WebSocket or HTTP from each model's advertised endpoints; streamed Responses traffic for the built-in `codex` provider uses WebSocket by default and uses HTTP when `useResponsesApiWebSocket` is disabled.
|
|
33
|
+
- **Desktop App**: Electron GUI with GitHub Copilot sign-in, Codex OAuth, provider configuration, token usage, logs, and one-click start/stop.
|
|
34
|
+
|
|
25
35
|
## Quick Start
|
|
26
36
|
|
|
27
|
-
|
|
37
|
+
Requires **Node.js >= 22.13.0** (npx) or **Bun >= 1.2.x**. A Copilot subscription is needed only for the GitHub Copilot provider; other configured providers can run independently.
|
|
28
38
|
|
|
29
39
|
```sh
|
|
30
40
|
npx @jeffreycao/copilot-api@latest start
|
|
@@ -43,17 +53,9 @@ curl http://localhost:4141/v1/models
|
|
|
43
53
|
```
|
|
44
54
|
|
|
45
55
|
> [!NOTE]
|
|
46
|
-
> Token usage storage requires Node.js >= 22.13.0 or Bun. See [Using with npx](#using-with-npx) for details.
|
|
56
|
+
> Token usage storage requires Node.js >= 22.13.0 or Bun. See [Using with npx](docs/guides/en/getting-started.md#using-with-npx) for details.
|
|
47
57
|
|
|
48
|
-
From here, jump to the guide for your client: [Claude Code](#using-with-claude-code), [OpenCode](#using-with-opencode), [Codex](#using-with-codex), or run it with [Docker](#using-with-docker).
|
|
49
|
-
|
|
50
|
-
## Highlights
|
|
51
|
-
|
|
52
|
-
- **Unified API Gateway**: Serve OpenAI-compatible Chat Completions (`/v1/chat/completions`), the OpenAI Responses API (`/v1/responses`), and Anthropic-compatible Messages (`/v1/messages`) from one local endpoint.
|
|
53
|
-
- **Multi-Provider**: Route GitHub Copilot, the built-in `codex` provider, and third-party providers (Kimi, DeepSeek, DashScope, OpenRouter, OpenCode Go, or a custom provider) behind the same gateway. GitHub Copilot is optional — with at least one enabled provider, the server starts in provider-only mode without a GitHub token.
|
|
54
|
-
- **Coding Agent Ready**: First-class setups for Claude Code, OpenCode, and Codex, including the interactive `--claude-code` launcher and a merged model catalog for Codex.
|
|
55
|
-
- **Streaming & WebSocket**: SSE streaming on all three client-facing protocols. Upstream Copilot Responses traffic selects WebSocket or HTTP from each model's advertised endpoints; streamed Responses traffic for the built-in `codex` provider uses WebSocket by default and uses HTTP when `useResponsesApiWebSocket` is disabled.
|
|
56
|
-
- **Desktop App**: Electron GUI with GitHub Copilot sign-in, Codex OAuth, provider configuration, token usage, logs, and one-click start/stop.
|
|
58
|
+
From here, jump to the guide for your client: [Claude Code](docs/guides/en/claude-code.md#using-with-claude-code), [OpenCode](docs/guides/en/opencode.md#using-with-opencode), [Codex](docs/guides/en/codex.md#using-with-codex), or run it with [Docker](docs/guides/en/docker.md#using-with-docker).
|
|
57
59
|
|
|
58
60
|
## Compatibility
|
|
59
61
|
|
|
@@ -76,838 +78,27 @@ Every client talks to the same local endpoint. The gateway routes each request t
|
|
|
76
78
|
Prefer a GUI? The Electron desktop app in `desktop/` covers GitHub Copilot sign-in, OpenAI Codex OAuth, and API-key configuration for Kimi, DeepSeek, DashScope, OpenRouter, or a custom provider — with one-click start/stop of the local server, and the local endpoint, auth header, available models, usage, and logs in one window.
|
|
77
79
|
|
|
78
80
|
<p align="center">
|
|
79
|
-
<img src="
|
|
80
|
-
<img src="
|
|
81
|
+
<img src="docs/screenshots/desktop-dashboard.png" alt="Copilot API desktop app dashboard" width="49%" />
|
|
82
|
+
<img src="docs/screenshots/desktop-token-usage.png" alt="Copilot API desktop app token usage view" width="49%" />
|
|
81
83
|
</p>
|
|
82
84
|
|
|
83
|
-
Windows x64 (`.exe`), macOS Apple Silicon (`.dmg`), and Linux x64 (`.AppImage`) packages are published in [GitHub Releases](https://github.com/caozhiyuan/copilot-api/releases). See [Electron Desktop App](#electron-desktop-app) for full setup and advanced configuration.
|
|
84
|
-
|
|
85
|
-
##
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
### Manual Configuration with `settings.json`
|
|
104
|
-
|
|
105
|
-
Alternatively, you can configure Claude Code by creating a `.claude/settings.json` file in your project's root directory. This file should contain the environment variables needed by Claude Code. This way you don't need to run the interactive setup every time.
|
|
106
|
-
|
|
107
|
-
Here is an example `.claude/settings.json` file:
|
|
108
|
-
|
|
109
|
-
```json
|
|
110
|
-
{
|
|
111
|
-
"env": {
|
|
112
|
-
"ANTHROPIC_BASE_URL": "http://localhost:4141",
|
|
113
|
-
"ANTHROPIC_AUTH_TOKEN": "dummy",
|
|
114
|
-
"ANTHROPIC_MODEL": "gpt-5.6-sol[1m]",
|
|
115
|
-
"ANTHROPIC_DEFAULT_OPUS_MODEL": "gpt-5.6-sol[1m]",
|
|
116
|
-
"ANTHROPIC_DEFAULT_SONNET_MODEL": "gpt-5.6-sol[1m]",
|
|
117
|
-
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "gpt-5.6-luna[1m]",
|
|
118
|
-
"CLAUDE_CODE_AUTO_COMPACT_WINDOW": "272000",
|
|
119
|
-
"CLAUDE_CODE_USE_VERTEX": "0",
|
|
120
|
-
"CLAUDE_CODE_USE_BEDROCK": "0",
|
|
121
|
-
"DISABLE_NON_ESSENTIAL_MODEL_CALLS": "1",
|
|
122
|
-
"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1",
|
|
123
|
-
"CLAUDE_CODE_ATTRIBUTION_HEADER": "0",
|
|
124
|
-
"CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION": "false",
|
|
125
|
-
"CLAUDE_CODE_DISABLE_TERMINAL_TITLE": "true",
|
|
126
|
-
"CLAUDE_CODE_ENABLE_AWAY_SUMMARY": "0",
|
|
127
|
-
"CLAUDE_CODE_TOTAL_TOKENS_REMINDER": "off",
|
|
128
|
-
"CLAUDE_CODE_EFFORT_LEVEL": "max",
|
|
129
|
-
"MCP_CONNECT_TIMEOUT_MS": "20000"
|
|
130
|
-
},
|
|
131
|
-
"alwaysThinkingEnabled": true,
|
|
132
|
-
"showThinkingSummaries": true
|
|
133
|
-
}
|
|
134
|
-
```
|
|
135
|
-
|
|
136
|
-
- Replace `ANTHROPIC_MODEL`, `ANTHROPIC_DEFAULT_OPUS_MODEL`, `ANTHROPIC_DEFAULT_SONNET_MODEL`, and `ANTHROPIC_DEFAULT_HAIKU_MODEL` according to your needs. After configuration, please install the claude code plugin [Plugin Integrations](#plugin-integrations).
|
|
137
|
-
- `CLAUDE_CODE_TOTAL_TOKENS_REMINDER: "off"` disables Claude Code's total-tokens reminder, which injects a `<total_tokens>N tokens left</total_tokens>` block into the conversation to pace the model against a remaining token budget. The default budget is 15,000,000 (15M) tokens, which is not very meaningful, so it is turned off here.
|
|
138
|
-
- If you are using the codex provider, it is recommended **not** to configure the model name in the `codex/xxx` format (e.g. `codex/gpt-5.6-sol`). Claude Code treats the `codex/` prefix as a special pattern and applies degraded behavior — for example, it strips all previously returned thinking blocks on every request. Use the plain model name (e.g. `gpt-5.6-sol`) instead, and add a `modelMappings` entry in `config.json` to route it back to the codex provider:
|
|
139
|
-
```json
|
|
140
|
-
"modelMappings": {
|
|
141
|
-
"gpt-5.6-sol": "codex/gpt-5.6-sol",
|
|
142
|
-
"gpt-5.6-terra": "codex/gpt-5.6-terra",
|
|
143
|
-
"gpt-5.6-luna": "codex/gpt-5.6-luna"
|
|
144
|
-
},
|
|
145
|
-
```
|
|
146
|
-
- Setting CLAUDE_CODE_ATTRIBUTION_HEADER to 0 can prevent Claude code from adding billing and version information in system prompts, thereby avoiding prompt cache invalidation.
|
|
147
|
-
- Turning off CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION and CLAUDE_CODE_ENABLE_AWAY_SUMMARY can prevent quota from being consumed unnecessarily.
|
|
148
|
-
- Claude Code WebSearch is supported for pure search requests. For Copilot, keep the global `messageApiWebSearchModel` set to a Responses-capable GPT model or a `provider/model` alias. For provider routes, use a native Anthropic provider or an `openai-responses` provider. Add `WebSearch` to `permissions.deny` only if you want to forbid this traffic.
|
|
149
|
-
- If using a non-Claude model, do not enable ENABLE_TOOL_SEARCH. If using the Claude model, can enable ENABLE_TOOL_SEARCH. The current Claude Code uses the client tool search mode. In this mode, loading defer tools requires an additional request each time.
|
|
150
|
-
- `CLAUDE_CODE_AUTO_COMPACT_WINDOW`: Set the context capacity in tokens used for auto-compaction calculations. Defaults to the model's context window: 200K for standard models or 1M for extended context models. Use a lower value like `500000` on a 1M model (e.g., `claude-opus-4-6[1m]`) to treat the window as 500K for compaction purposes. The value is capped at the model's actual context window. `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE` is applied as a percentage of this value. Setting this variable decouples the compaction threshold from the status line's `used_percentage`, which always uses the model's full context window.
|
|
151
|
-
|
|
152
|
-
You can find more options here: [Claude Code settings](https://docs.anthropic.com/en/docs/claude-code/settings#environment-variables)
|
|
153
|
-
|
|
154
|
-
You can also read more about IDE integration here: [Add Claude Code to your IDE](https://docs.anthropic.com/en/docs/claude-code/ide-integrations)
|
|
155
|
-
|
|
156
|
-
## Using with OpenCode
|
|
157
|
-
|
|
158
|
-
OpenCode already has a direct GitHub Copilot provider. Use this section when you want OpenCode to point at this AI gateway through `@ai-sdk/anthropic` and reuse the agent behaviors described earlier in this README.
|
|
159
|
-
|
|
160
|
-
### Minimal setup
|
|
161
|
-
|
|
162
|
-
Start the AI gateway with the OpenCode OAuth app:
|
|
163
|
-
|
|
164
|
-
```sh
|
|
165
|
-
npx @jeffreycao/copilot-api@latest auth --oauth-app=opencode
|
|
166
|
-
npx @jeffreycao/copilot-api@latest start
|
|
167
|
-
```
|
|
168
|
-
|
|
169
|
-
Then point OpenCode at the gateway with `@ai-sdk/anthropic`.
|
|
170
|
-
|
|
171
|
-
Example `~/.config/opencode/opencode.json`:
|
|
172
|
-
|
|
173
|
-
```json
|
|
174
|
-
{
|
|
175
|
-
"$schema": "https://opencode.ai/config.json",
|
|
176
|
-
"provider": {
|
|
177
|
-
"local": {
|
|
178
|
-
"npm": "@ai-sdk/anthropic",
|
|
179
|
-
"name": "My Local",
|
|
180
|
-
"options": {
|
|
181
|
-
"baseURL": "http://localhost:4141/v1",
|
|
182
|
-
"apiKey": "dummy"
|
|
183
|
-
},
|
|
184
|
-
"models": {
|
|
185
|
-
"gpt-5.4": {
|
|
186
|
-
"name": "gpt-5.4",
|
|
187
|
-
"modalities": {
|
|
188
|
-
"input": ["text", "image"],
|
|
189
|
-
"output": ["text"]
|
|
190
|
-
},
|
|
191
|
-
"limit": {
|
|
192
|
-
"context": 400000,
|
|
193
|
-
"input": 272000,
|
|
194
|
-
"output": 128000
|
|
195
|
-
}
|
|
196
|
-
},
|
|
197
|
-
"claude-sonnet-4.6": {
|
|
198
|
-
"id": "claude-sonnet-4.6",
|
|
199
|
-
"name": "claude-sonnet-4.6",
|
|
200
|
-
"modalities": {
|
|
201
|
-
"input": ["text", "image"],
|
|
202
|
-
"output": ["text"]
|
|
203
|
-
},
|
|
204
|
-
"limit": {
|
|
205
|
-
"context": 200000,
|
|
206
|
-
"output": 32000
|
|
207
|
-
},
|
|
208
|
-
"options": {
|
|
209
|
-
"thinking": {
|
|
210
|
-
"type": "adaptive"
|
|
211
|
-
},
|
|
212
|
-
"effort": "max"
|
|
213
|
-
}
|
|
214
|
-
}
|
|
215
|
-
}
|
|
216
|
-
}
|
|
217
|
-
}
|
|
218
|
-
}
|
|
219
|
-
```
|
|
220
|
-
|
|
221
|
-
Why these fields matter:
|
|
222
|
-
|
|
223
|
-
- `npm: "@ai-sdk/anthropic"` is the important part. OpenCode will speak Anthropic Messages semantics to this AI gateway instead of flattening everything into OpenAI Chat Completions.
|
|
224
|
-
- `options.baseURL` should be `http://localhost:4141/v1`; the Anthropic SDK will append `/messages`, `/models`, and `/messages/count_tokens` automatically.
|
|
225
|
-
- If you enable `auth.apiKeys` in this AI gateway, replace `dummy` with a real key. Otherwise any placeholder value is fine.
|
|
226
|
-
|
|
227
|
-
## Using with Codex
|
|
228
|
-
|
|
229
|
-
This AI gateway can also power Codex.
|
|
230
|
-
|
|
231
|
-
Recommended Codex version: `0.155.1`.
|
|
232
|
-
|
|
233
|
-
### Codex `config.toml` Reference
|
|
234
|
-
|
|
235
|
-
Add the following `[model_providers.copilot_api]` section to your Codex `~/.codex/config.toml`:
|
|
236
|
-
|
|
237
|
-
```toml
|
|
238
|
-
model_provider = "copilot_api"
|
|
239
|
-
model_reasoning_summary = "auto"
|
|
240
|
-
model_context_window = 272000
|
|
241
|
-
model_auto_compact_token_limit = 244800
|
|
242
|
-
web_search = "live"
|
|
243
|
-
|
|
244
|
-
[model_providers.copilot_api]
|
|
245
|
-
name = "OpenAI"
|
|
246
|
-
base_url = "http://localhost:4141"
|
|
247
|
-
env_key = "GITHUB_COPILOT_API_KEY"
|
|
248
|
-
requires_openai_auth = true
|
|
249
|
-
supports_websockets = false
|
|
250
|
-
supports_standalone_web_search = true
|
|
251
|
-
wire_api = "responses"
|
|
252
|
-
request_max_retries = 3
|
|
253
|
-
stream_max_retries = 3
|
|
254
|
-
stream_idle_timeout_ms = 300000
|
|
255
|
-
|
|
256
|
-
[features]
|
|
257
|
-
remote_compaction_v2 = true
|
|
258
|
-
# optional: set false only when the model does not support tool_search
|
|
259
|
-
apps = false
|
|
260
|
-
standalone_web_search = true
|
|
261
|
-
|
|
262
|
-
[analytics]
|
|
263
|
-
enabled = false
|
|
264
|
-
```
|
|
265
|
-
|
|
266
|
-
> [!NOTE]
|
|
267
|
-
> `name` must be set to `"OpenAI"`.
|
|
268
|
-
>
|
|
269
|
-
> For third-party models that do not support `tool_search`, we recommend disabling features.apps. Otherwise, each prompt may consume an additional 20,000 or more tokens.
|
|
270
|
-
>
|
|
271
|
-
> `supports_standalone_web_search` and `[features] standalone_web_search` must both be enabled to expose the standalone `web.run` search tool.
|
|
272
|
-
|
|
273
|
-
### If Codex Is Not Signed In to a GPT Account
|
|
274
|
-
|
|
275
|
-
```toml
|
|
276
|
-
[model_providers.copilot_api]
|
|
277
|
-
name = "OpenAI"
|
|
278
|
-
base_url = "http://localhost:4141"
|
|
279
|
-
requires_openai_auth = false
|
|
280
|
-
supports_websockets = false
|
|
281
|
-
supports_standalone_web_search = true
|
|
282
|
-
wire_api = "responses"
|
|
283
|
-
request_max_retries = 3
|
|
284
|
-
stream_max_retries = 3
|
|
285
|
-
stream_idle_timeout_ms = 300000
|
|
286
|
-
|
|
287
|
-
[features]
|
|
288
|
-
standalone_web_search = true
|
|
289
|
-
|
|
290
|
-
[model_providers.copilot_api.auth]
|
|
291
|
-
command = "powershell.exe"
|
|
292
|
-
args = [
|
|
293
|
-
"-NoProfile",
|
|
294
|
-
"-NonInteractive",
|
|
295
|
-
"-Command",
|
|
296
|
-
"[Console]::Out.Write($env:GITHUB_COPILOT_API_KEY)"
|
|
297
|
-
]
|
|
298
|
-
```
|
|
299
|
-
|
|
300
|
-
macOS, replace the `auth` block with:
|
|
301
|
-
|
|
302
|
-
```toml
|
|
303
|
-
[model_providers.copilot_api.auth]
|
|
304
|
-
command = "/bin/zsh"
|
|
305
|
-
args = [
|
|
306
|
-
"-c",
|
|
307
|
-
"printf '%s' \"$GITHUB_COPILOT_API_KEY\""
|
|
308
|
-
]
|
|
309
|
-
```
|
|
310
|
-
|
|
311
|
-
Without this configuration, Codex cannot fetch `/v1/models` while not signed in to a GPT account, so custom models are unavailable in the model picker.
|
|
312
|
-
|
|
313
|
-
When a Codex client (`User-Agent` starts with `codex`) requests the top-level `GET /v1/models`, the gateway merges native Codex models with models available through the Messages adapter. The latter advertise `use_responses_lite: true`, except DeepSeek models, which use `use_responses_lite: false` and `tool_mode: null`. For other models, `/v1/responses` uses **Responses → Messages** for Anthropic providers, while OpenAI-compatible providers and Chat-only Copilot models reuse the existing Messages route for **Responses → Messages → Chat Completions**, then translate streaming or JSON results back to Responses.
|
|
314
|
-
|
|
315
|
-
> **Note:** DeepSeek models do not use Responses Lite (`use_responses_lite: false`, `tool_mode: null`), so the tool set they advertise to Codex differs from other models, which use `tool_mode: "code_mode_only"`. Switching between a DeepSeek model and a Responses Lite model mid-session is not compatible, because tool calls and conversation history produced under one tool set do not translate to the other. Start a new Codex session when switching between them.
|
|
316
|
-
|
|
317
|
-
The merged catalog is what Codex shows in its model picker, including the models exposed by your configured providers:
|
|
318
|
-
|
|
319
|
-
<img src="./docs/screenshots/codex-models.png" alt="Codex model picker showing models provided by the gateway" width="900" />
|
|
320
|
-
|
|
321
|
-
For Codex clients, only `gpt-*` Copilot models use the native Responses API; non-GPT Copilot models always go through the adapter, even when they advertise native `/responses` support. The same Codex rule applies on provider `/v1/responses` routes (top-level `provider/model` aliases and `/:provider/v1/responses`): for `openai-responses` providers, non-`gpt-*` models fall back to the Messages adapter, while `gpt-*` models keep native Responses forwarding.
|
|
322
|
-
|
|
323
|
-
Responses Lite tool definitions are read from `input.additional_tools`, without relying on top-level `tools`. Function, `namespace`, and custom tools are supported; clients must declare `apply_patch` as `type: "custom"`, and it is not handled as a standalone special tool type. Returned calls recover their original `name` and `namespace`. Tools are collected before old history is trimmed, so compaction requests retain them. The Messages fallback does not support Responses `tool_search` mode. Anthropic `output_config.effort` keeps the project's existing valid levels; Responses `minimal` maps to `low`, while `none` omits Anthropic effort.
|
|
324
|
-
|
|
325
|
-
When Codex uses the top-level GitHub Copilot route with `approvals_reviewer = "auto_review"`, map its internal review model to a Responses-capable Copilot model in the gateway's `config.json`:
|
|
326
|
-
|
|
327
|
-
```json
|
|
328
|
-
{
|
|
329
|
-
"modelMappings": {
|
|
330
|
-
"codex-auto-review": "gpt-5.6-luna"
|
|
331
|
-
}
|
|
332
|
-
}
|
|
333
|
-
```
|
|
334
|
-
|
|
335
|
-
This mapping only applies to the top-level GitHub Copilot route. Provider-scoped routes do not use `modelMappings`, so the built-in `/codex` provider continues to handle `codex-auto-review` natively.
|
|
336
|
-
|
|
337
|
-
---
|
|
338
|
-
|
|
339
|
-
## Project Overview
|
|
340
|
-
|
|
341
|
-
A small AI gateway that can use GitHub Copilot, the built-in `codex` provider, or configured third-party providers such as DashScope. GitHub Copilot is optional: if no GitHub token is available, the server can still start in provider-only mode as long as at least one enabled provider is configured.
|
|
342
|
-
|
|
343
|
-
The gateway exposes OpenAI- and Anthropic-compatible APIs from one local endpoint, so tools like [Claude Code](https://docs.anthropic.com/en/docs/claude-code/overview), OpenCode, Codex, and OpenAI-compatible clients can share the same local server.
|
|
344
|
-
|
|
345
|
-
On the GitHub Copilot path, the gateway prefers Copilot's native Anthropic-style Messages API when available, preserving more Claude-native behavior for tool-heavy workflows.
|
|
346
|
-
|
|
347
|
-
## Important Notes
|
|
348
|
-
|
|
349
|
-
> [!IMPORTANT]
|
|
350
|
-
> **Before using, please be aware of the following:**
|
|
351
|
-
>
|
|
352
|
-
> 1. **Codex configuration:** When using with Codex, add the gateway provider to `~/.codex/config.toml`. See [Codex `config.toml` Reference](#codex-configtoml-reference).
|
|
353
|
-
>
|
|
354
|
-
> 2. **Claude Code configuration:** When using with Claude Code, please configure the model ID as `claude-opus-4-8[1m]`. Example claude `settings.json` see [Manual Configuration with `settings.json`](#manual-configuration-with-settingsjson).
|
|
355
|
-
>
|
|
356
|
-
> 3. **OpenCode configuration:** When using with OpenCode, configure `~/.config/opencode/opencode.json` with `@ai-sdk/anthropic`. See [Using with OpenCode](#using-with-opencode).
|
|
357
|
-
>
|
|
358
|
-
> 4. **Built-in `copilot`, `codex` and third-party providers:** Run `npx @jeffreycao/copilot-api@latest auth` and choose `copilot`, `codex`, `deepseek`, `custom`, or other providers.
|
|
359
|
-
>
|
|
360
|
-
> 5. **Note:** See [GitHub Copilot Security Notice](./NOTICE.md#github-copilot-security-notice) for the warning removed from the README header.
|
|
361
|
-
|
|
362
|
-
## Prerequisites
|
|
363
|
-
|
|
364
|
-
- Bun (>= 1.2.x)
|
|
365
|
-
- Node.js >= 22.13.0 if you plan to run the published CLI with `npx`
|
|
366
|
-
- GitHub account with Copilot subscription only if you want to use the GitHub Copilot provider
|
|
367
|
-
- An API key or OAuth login for at least one configured provider if you want to run without GitHub Copilot
|
|
368
|
-
|
|
369
|
-
## Installation
|
|
370
|
-
|
|
371
|
-
To install dependencies, run:
|
|
372
|
-
|
|
373
|
-
```sh
|
|
374
|
-
bun install
|
|
375
|
-
```
|
|
376
|
-
|
|
377
|
-
## Running from Source
|
|
378
|
-
|
|
379
|
-
> [!NOTE]
|
|
380
|
-
> Building from source with `tsdown@0.23` requires Node.js `^22.18.0 || ^24.11.0 || >=26.0.0`. This is a build-time requirement only; the published CLI supports Node.js >= 22.13.0.
|
|
381
|
-
|
|
382
|
-
The project can be run from source in several ways:
|
|
383
|
-
|
|
384
|
-
### Development Mode
|
|
385
|
-
|
|
386
|
-
```sh
|
|
387
|
-
bun run dev start
|
|
388
|
-
```
|
|
389
|
-
|
|
390
|
-
### Production Mode
|
|
391
|
-
|
|
392
|
-
```sh
|
|
393
|
-
bun run start start
|
|
394
|
-
```
|
|
395
|
-
|
|
396
|
-
> The trailing `start` is the CLI subcommand passed to `src/main.ts`, not a typo: `bun run dev start` runs watch mode, `bun run start start` runs production.
|
|
397
|
-
|
|
398
|
-
## Using with npx
|
|
399
|
-
|
|
400
|
-
You can run the project directly using npx:
|
|
401
|
-
|
|
402
|
-
> [!IMPORTANT]
|
|
403
|
-
> Token usage storage uses Node's built-in `node:sqlite` module when running with `npx`. It is enabled on Node.js >= 22.13.0, the first release where `node:sqlite` works without `--experimental-sqlite`. On older Node.js versions the CLI still starts, but token usage storage is disabled.
|
|
404
|
-
>
|
|
405
|
-
> If you want token usage storage without upgrading Node.js, run the published CLI with Bun instead: `bunx --bun @jeffreycao/copilot-api@latest start`.
|
|
406
|
-
|
|
407
|
-
```sh
|
|
408
|
-
npx @jeffreycao/copilot-api@latest start
|
|
409
|
-
```
|
|
410
|
-
|
|
411
|
-
With options:
|
|
412
|
-
|
|
413
|
-
```sh
|
|
414
|
-
npx @jeffreycao/copilot-api@latest auth keys --add your-gateway-api-key
|
|
415
|
-
npx @jeffreycao/copilot-api@latest start --host 0.0.0.0 --port 8080
|
|
416
|
-
```
|
|
417
|
-
|
|
418
|
-
Binding to `0.0.0.0` exposes the gateway to the network, so the server requires at least one gateway API key and restricts CORS to same-origin requests.
|
|
419
|
-
|
|
420
|
-
For authentication or provider configuration only:
|
|
421
|
-
|
|
422
|
-
```sh
|
|
423
|
-
npx @jeffreycao/copilot-api@latest auth
|
|
424
|
-
```
|
|
425
|
-
|
|
426
|
-
To run without GitHub Copilot, configure at least one provider first, then start the server normally:
|
|
427
|
-
|
|
428
|
-
```sh
|
|
429
|
-
npx @jeffreycao/copilot-api@latest auth login --provider dashscope
|
|
430
|
-
npx @jeffreycao/copilot-api@latest start
|
|
431
|
-
```
|
|
432
|
-
|
|
433
|
-
## Using with Docker
|
|
434
|
-
|
|
435
|
-
The supplied Compose file uses the current published `ghcr.io/caozhiyuan/copilot-api:latest` image. No local image build is required. It stores gateway state in `/data` and runs the server as the non-root `bun` user.
|
|
436
|
-
|
|
437
|
-
### Quick start with Docker Compose
|
|
438
|
-
|
|
439
|
-
Run these commands from the repository root. Replace `YOUR_GATEWAY_API_KEY` with a strong key for clients connecting to this gateway:
|
|
440
|
-
|
|
441
|
-
```sh
|
|
442
|
-
mkdir -p copilot-data
|
|
443
|
-
docker compose pull
|
|
444
|
-
docker compose run --rm copilot-api --auth keys --add YOUR_GATEWAY_API_KEY
|
|
445
|
-
docker compose run --rm copilot-api --auth login
|
|
446
|
-
docker compose up -d
|
|
447
|
-
docker compose ps
|
|
448
|
-
```
|
|
449
|
-
|
|
450
|
-
If `COPILOT_API_GITHUB_TOKEN` or the legacy `GH_TOKEN` is already available in your environment or a private `.env` file, you can skip `--auth login`. A GitHub token authorizes access to GitHub Copilot; it does not replace the gateway API key configured above.
|
|
451
|
-
|
|
452
|
-
Before every server or authentication run, the one-shot `data-init` service repairs ownership of the gateway's own state in the mounted data directory: `config.json`, `github_token` (including the enterprise `ent_github_token` and OAuth app subdirectories such as `opencode/github_token`), `codex_credentials.json`, `desktop-config.json`, `copilot-api.sqlite*`, `logs/`, and `cache/`. Other files and directories in the mount are left untouched. This lets the non-root server reuse files written by an earlier root container, including mode `0600` configuration files. The host directory defaults to `./copilot-data`, matching the earlier Docker instructions; Compose mounts it at `/data` and sets `COPILOT_API_HOME` accordingly. Set `COPILOT_API_DATA_DIR` in the environment or your own `.env` to use an existing directory elsewhere.
|
|
453
|
-
|
|
454
|
-
```dotenv
|
|
455
|
-
COPILOT_API_DATA_DIR=/absolute/path/to/copilot-data
|
|
456
|
-
```
|
|
457
|
-
|
|
458
|
-
Create the host directory before the first run. Compose does not auto-create a missing host directory (`create_host_path: false`), so a typo in `COPILOT_API_DATA_DIR` fails fast instead of starting with empty state. The one-shot `data-init` service also refuses unsafe values such as `/`, and refuses to run when the mount contains system directories (`/etc`, `/usr`, and so on), which means the path resolved to a system root.
|
|
459
|
-
|
|
460
|
-
The local endpoint is `http://127.0.0.1:4141`. To publish the gateway on every host interface after configuring a gateway API key, add this to your `.env`:
|
|
461
|
-
|
|
462
|
-
```dotenv
|
|
463
|
-
COPILOT_API_BIND=0.0.0.0
|
|
464
|
-
COPILOT_API_PORT=4141
|
|
465
|
-
```
|
|
466
|
-
|
|
467
|
-
The Compose service also forwards `COPILOT_API_SQLITE_DB_PATH`, `COPILOT_API_ENTERPRISE_URL`, and `COPILOT_API_OAUTH_APP` from the environment or a private `.env` file. SQLite paths are container paths and should stay under the writable `/data` mount, for example:
|
|
468
|
-
|
|
469
|
-
```dotenv
|
|
470
|
-
COPILOT_API_SQLITE_DB_PATH=/data/copilot-api.sqlite
|
|
471
|
-
COPILOT_API_ENTERPRISE_URL=company.ghe.com
|
|
472
|
-
COPILOT_API_OAUTH_APP=opencode
|
|
473
|
-
```
|
|
474
|
-
|
|
475
|
-
Token and proxy variables can be overridden in the same file. Proxy addresses must be reachable from inside the container. Keep the internal port at `4141` so the health check remains valid.
|
|
476
|
-
|
|
477
|
-
## Electron Desktop App
|
|
478
|
-
|
|
479
|
-
If you prefer a GUI, this repository also includes an Electron desktop app in `desktop/`. It supports GitHub Copilot sign-in, OpenAI Codex OAuth with manual switching among up to 3 Codex accounts and removal of accounts that are not in use, and API-key configuration for Kimi, DeepSeek, DashScope, OpenRouter, or a custom provider. After a Codex account switch, the app prompts you to restart the service manually. After authorization or provider configuration, it can start and stop the local proxy with one click and shows the local endpoint, auth header, available models, usage, and logs in the app.
|
|
480
|
-
|
|
481
|
-
The settings screen also exposes `OAuth App`, `API Home`, `SQLite DB Path`, `Enterprise URL`, verbose logging, and minimize-to-tray. Windows x64 (`.exe`), macOS Apple Silicon (`.dmg`), and Linux x64 (`.AppImage`) packages are published in GitHub Releases:
|
|
482
|
-
|
|
483
|
-
https://github.com/caozhiyuan/copilot-api/releases
|
|
484
|
-
|
|
485
|
-
On Linux, make the downloaded AppImage executable before launching it:
|
|
486
|
-
|
|
487
|
-
```sh
|
|
488
|
-
chmod +x Copilot-API-*-linux-x86_64.AppImage
|
|
489
|
-
./Copilot-API-*-linux-x86_64.AppImage
|
|
490
|
-
```
|
|
491
|
-
|
|
492
|
-
Download the installer for your platform, authorize or configure a provider inside the app, choose a port, start the server, then point your client at the local endpoint shown in the app. Packaged desktop builds use the bundled Electron runtime, so normal desktop usage does not require installing Node.js separately. Token usage history is enabled when that bundled runtime supports SQLite.
|
|
493
|
-
|
|
494
|
-
The desktop app's Advanced Config page reads and writes the shared model mappings through `GET/POST /admin/config/model-mappings`. The same mappings apply across `POST /v1/messages`, `POST /v1/messages/count_tokens`, `POST /v1/responses`, and `POST /v1/chat/completions` instead of being split per interface. It uses `auth.adminApiKey` instead of the regular `auth.apiKeys`, and the app reads that key directly from `config.json` after the server has generated it on startup.
|
|
495
|
-
|
|
496
|
-
## GPT Tool Search
|
|
497
|
-
|
|
498
|
-
For GPT Responses models such as `gpt-5.4+`, this AI gateway can expose Responses `tool_search` through a small MCP bridge. The same bridge can be used by Claude Code and opencode, as long as the client loads MCP servers and sends Anthropic Messages traffic through this gateway.
|
|
499
|
-
|
|
500
|
-
Do not set Claude Code's native `ENABLE_TOOL_SEARCH` for GPT models. That flag enables Claude Code's own client-side tool search mode, and it may stop forwarding deferred tool definitions. This gateway needs the full tool definitions so it can keep the small always-loaded tool set eager and translate every other tool into Responses deferred namespaces.
|
|
501
|
-
|
|
502
|
-
If you install `tool-search@copilot-api-marketplace`, Claude Code receives this MCP bridge automatically and you can skip the manual Claude Code MCP setup below.
|
|
503
|
-
|
|
504
|
-
Add the tool search bridge to the MCP config used by Claude Code:
|
|
505
|
-
|
|
506
|
-
```json
|
|
507
|
-
{
|
|
508
|
-
"mcpServers": {
|
|
509
|
-
"tool_search": {
|
|
510
|
-
"type": "stdio",
|
|
511
|
-
"command": "npx",
|
|
512
|
-
"args": ["-y", "@jeffreycao/copilot-api@latest", "mcp"]
|
|
513
|
-
}
|
|
514
|
-
}
|
|
515
|
-
}
|
|
516
|
-
```
|
|
517
|
-
|
|
518
|
-
Add the tool search bridge to the MCP config used by opencode:
|
|
519
|
-
|
|
520
|
-
```json
|
|
521
|
-
{
|
|
522
|
-
"mcp": {
|
|
523
|
-
"tool_search": {
|
|
524
|
-
"type": "local",
|
|
525
|
-
"command": ["npx", "-y", "@jeffreycao/copilot-api@latest", "mcp"]
|
|
526
|
-
}
|
|
527
|
-
}
|
|
528
|
-
}
|
|
529
|
-
```
|
|
530
|
-
|
|
531
|
-
For local development, use `bun` as the command and `["run", "./src/main.ts", "mcp"]` as the args.
|
|
532
|
-
|
|
533
|
-
Internally, the gateway now configures OpenAI Responses `tool_search` in client-executed mode. Deferred tools are still exposed as searchable namespaces, but the model is explicitly asked to return the exact deferred tool names it wants to load next.
|
|
534
|
-
|
|
535
|
-
The bridge uses direct tool selection, not query search. Its tool input is `names`, a comma-separated list of exact deferred tool names, for example `TaskList,TaskGet,mcp__fetch__fetch`.
|
|
536
|
-
|
|
537
|
-
## Plugin Integrations
|
|
538
|
-
|
|
539
|
-
Plugin integrations are available for Claude Code and opencode.
|
|
540
|
-
|
|
541
|
-
### Claude Code plugin integration (marketplace-based)
|
|
542
|
-
|
|
543
|
-
The Claude Code integration is packaged as two plugins:
|
|
544
|
-
|
|
545
|
-
- `agent-inject` injects `__SUBAGENT_MARKER__...` on `SubagentStart`, so the gateway can infer `x-initiator: agent`.
|
|
546
|
-
- `tool-search` registers the `tool_search` MCP bridge used for GPT Responses deferred tool loading.
|
|
547
|
-
|
|
548
|
-
- Marketplace catalog in this repository: `.claude-plugin/marketplace.json`
|
|
549
|
-
- Plugin sources in this repository: `plugin/claude/agent-inject`, `plugin/claude/tool-search`
|
|
550
|
-
|
|
551
|
-
Add the marketplace remotely:
|
|
552
|
-
|
|
553
|
-
```sh
|
|
554
|
-
/plugin marketplace add https://github.com/caozhiyuan/copilot-api.git
|
|
555
|
-
```
|
|
556
|
-
|
|
557
|
-
Install the plugins from the marketplace:
|
|
558
|
-
|
|
559
|
-
```sh
|
|
560
|
-
/plugin install agent-inject@copilot-api-marketplace
|
|
561
|
-
/plugin install tool-search@copilot-api-marketplace
|
|
562
|
-
```
|
|
563
|
-
|
|
564
|
-
After installation, `agent-inject` injects `__SUBAGENT_MARKER__...` on `SubagentStart`, and the gateway uses it to infer `x-initiator: agent`.
|
|
565
|
-
|
|
566
|
-
The `agent-inject` plugin also registers a `UserPromptSubmit` hook that returns `{"continue": true}`, and it can inject `SessionStart` reminder rules through environment variables:
|
|
567
|
-
|
|
568
|
-
- `CLAUDE_PLUGIN_ENABLE_QUESTION_RULES=1` enables the two reminders about using the `question` tool automatically for Claude Code.
|
|
569
|
-
- `CLAUDE_PLUGIN_ENABLE_NO_BACKGROUND_AGENTS_RULE=1` enables the `run_in_background: true` avoidance reminder for agent hooks.
|
|
570
|
-
|
|
571
|
-
The `tool-search` plugin bundles the same MCP bridge described in [GPT Tool Search](#gpt-tool-search), so Claude Code users do not need to add the `tool_search` server manually when they install that plugin.
|
|
572
|
-
|
|
573
|
-
The plugin also auto-approves bridge calls through a `PermissionRequest` hook scoped exactly to `mcp__plugin_tool-search_tool_search__search`. The hook does not approve other MCP tools and does not override explicit `ask` or `deny` permission rules.
|
|
574
|
-
|
|
575
|
-
### Opencode plugin
|
|
576
|
-
|
|
577
|
-
The subagent marker producer is packaged as an opencode plugin located at `plugin/opencode/subagent-marker.js`.
|
|
578
|
-
|
|
579
|
-
**Installation:**
|
|
580
|
-
|
|
581
|
-
Copy the plugin file to your opencode plugins directory:
|
|
582
|
-
|
|
583
|
-
```sh
|
|
584
|
-
# Clone or download this repository, then copy the plugin
|
|
585
|
-
cp plugin/opencode/subagent-marker.js ~/.config/opencode/plugins/
|
|
586
|
-
```
|
|
587
|
-
|
|
588
|
-
Or manually create the file at `~/.config/opencode/plugins/subagent-marker.js` with the plugin content.
|
|
589
|
-
|
|
590
|
-
**Features:**
|
|
591
|
-
|
|
592
|
-
- Tracks sub-sessions created by subagents
|
|
593
|
-
- Automatically prepends a marker system reminder (`__SUBAGENT_MARKER__...`) to subagent chat messages
|
|
594
|
-
- Sets `x-session-id` header for session tracking
|
|
595
|
-
- Enables the gateway to infer `x-initiator: agent` for subagent-originated requests
|
|
596
|
-
|
|
597
|
-
The plugin hooks into `session.created`, `session.deleted`, `chat.message`, and `chat.headers` events to provide seamless subagent marker functionality.
|
|
598
|
-
|
|
599
|
-
## Using the Usage Viewer
|
|
600
|
-
|
|
601
|
-
After starting the server, a URL to the Copilot Usage Dashboard will be displayed in your console. This dashboard is a web interface for monitoring your API usage.
|
|
602
|
-
|
|
603
|
-
1. Start the server. For example, using npx:
|
|
604
|
-
```sh
|
|
605
|
-
npx @jeffreycao/copilot-api@latest start
|
|
606
|
-
```
|
|
607
|
-
2. The server will output a URL to the usage viewer. Copy and paste this URL into your browser. It will look something like this:
|
|
608
|
-
`http://localhost:4141/usage-viewer?endpoint=http://localhost:4141/usage`
|
|
609
|
-
- If you use the `start.bat` script on Windows, this page will open automatically.
|
|
610
|
-
|
|
611
|
-
The dashboard provides a user-friendly interface to view your Copilot usage data:
|
|
612
|
-
|
|
613
|
-
> Token usage history requires Bun or Node.js >= 22.13.0. On older Node.js versions the server runs normally but token usage storage is disabled.
|
|
614
|
-
|
|
615
|
-
- **API Endpoint URL**: The dashboard is pre-configured to fetch data from your local server endpoint via a URL query parameter. You can manually switch this to any other compatible API endpoint.
|
|
616
|
-
- **API Key Authentication**: If API Key authentication is enabled, enter a raw API key (sent as the `x-api-key` header) or `Authorization: Bearer <key>`. Credentials are remembered in the browser's local storage per endpoint origin, and switching to a different endpoint origin does not automatically send the previous credential.
|
|
617
|
-
- **Period Selector**: Choose from six time ranges: `today` (the current local calendar day so far), `this_week` (Monday at 00:00 through now), `last_7_days` (the rolling seven calendar days through now), `this_month` (the first day of the current month at 00:00 through now), `last_30_days` (the rolling 30 calendar days through now), and `lifetime` (the earliest recorded event through now). Today is selected by default, and the exact date range appears next to the selector. The URL query parameter updates automatically when you switch, making it easy to bookmark and share. The legacy values `day`, `week`, and `month` are still accepted and mapped to their new equivalents.
|
|
618
|
-
- **Fetch Data**: Click the "Refresh" button to load or refresh the usage data. The dashboard also fetches data automatically on page load.
|
|
619
|
-
- **Copilot Quotas**: View quota usage for services such as Chat and Completions via progress bars. Hover over a card to see used/remaining details.
|
|
620
|
-
- **Token Usage Metric Cards**: See a summary of Total, Input, Output, Cache Read, Cache Write, Requests, and estimated cost for the current period.
|
|
621
|
-
- **Trend Chart**: An interactive line chart with model and metric filters for the selected period. Click a data point to inspect the usage breakdown for a day; Lifetime chart data is sampled from the daily buckets and capped at 180 points for readability.
|
|
622
|
-
- **Model Breakdown Table**: A per-model summary of requests, input/output/cache tokens, and estimated cost for the selected period.
|
|
623
|
-
- **Request Events (Paginated)**: A time-sorted list of request event records with pagination support, showing timestamps, models, request IDs, and token counts.
|
|
624
|
-
- **Detailed Information**: See the full JSON response from the API for a detailed breakdown of all available usage statistics.
|
|
625
|
-
- **URL-based Configuration**: You can also specify the API endpoint and period directly via `endpoint` and `period` query parameters. For example:
|
|
626
|
-
`http://localhost:4141/usage-viewer?endpoint=http://your-api-server/usage&period=this_week`
|
|
627
|
-
|
|
628
|
-
### Usage Viewer Screenshot
|
|
629
|
-
|
|
630
|
-
<p align="center">
|
|
631
|
-
<img src="./docs/screenshots/usage-viewer.png" alt="Copilot API usage viewer" width="900" />
|
|
632
|
-
</p>
|
|
633
|
-
|
|
634
|
-
## Command Structure
|
|
635
|
-
|
|
636
|
-
Copilot API now uses a subcommand structure with these main commands:
|
|
637
|
-
|
|
638
|
-
- `start`: Start the gateway server. If a GitHub token is available, the server starts with Copilot enabled. If no GitHub token is available, it starts in provider-only mode when at least one enabled provider exists; otherwise it guides you through provider setup.
|
|
639
|
-
- `auth`: Run provider login or configuration without starting the server. Use it for GitHub Copilot login, Codex OAuth, or third-party provider API key setup.
|
|
640
|
-
- `debug`: Display diagnostic information including version, runtime details, file paths, and authentication status. Useful for troubleshooting and support.
|
|
641
|
-
|
|
642
|
-
## Command Line Options
|
|
643
|
-
|
|
644
|
-
### Global Options
|
|
645
|
-
|
|
646
|
-
The following options can be used with any subcommand. When passing them before the subcommand, use the `--key=value` form:
|
|
647
|
-
|
|
648
|
-
| Option | Description | Default | Alias |
|
|
649
|
-
| ----------------- | ------------------------------------------------------ | ------- | ----- |
|
|
650
|
-
| --api-home | Path to the API home directory (sets COPILOT_API_HOME) | none | none |
|
|
651
|
-
| --oauth-app | OAuth app identifier (sets COPILOT_API_OAUTH_APP) | none | none |
|
|
652
|
-
| --enterprise-url | Enterprise URL for GitHub (sets COPILOT_API_ENTERPRISE_URL) | none | none |
|
|
653
|
-
|
|
654
|
-
### Start Command Options
|
|
655
|
-
|
|
656
|
-
The following command line options are available for the `start` command:
|
|
657
|
-
|
|
658
|
-
| Option | Description | Default | Alias |
|
|
659
|
-
| -------------- | ----------------------------------------------------------------------------- | ---------- | ----- |
|
|
660
|
-
| --host | Host to listen on; non-loopback hosts require a configured gateway API key | 127.0.0.1 | none |
|
|
661
|
-
| --port | Port to listen on | 4141 | -p |
|
|
662
|
-
| --verbose | Enable verbose logging | false | -v |
|
|
663
|
-
| --github-token | Provide GitHub token directly (must be generated using the `auth` subcommand); prefer `COPILOT_API_GITHUB_TOKEN`, since arguments are visible in the process list | none | -g |
|
|
664
|
-
| --claude-code | Generate a command to launch Claude Code with Copilot API config | false | -c |
|
|
665
|
-
| --show-token | Show GitHub and Copilot tokens on fetch and refresh | false | none |
|
|
666
|
-
| --proxy-env | Initialize proxy from environment variables | false | none |
|
|
667
|
-
|
|
668
|
-
Passing the GitHub token on the command line exposes it to every local user through the process list, so prefer the `COPILOT_API_GITHUB_TOKEN` environment variable. The gateway resolves the token in this order: `--github-token` → `COPILOT_API_GITHUB_TOKEN` → the token file written by `auth login`.
|
|
669
|
-
|
|
670
|
-
### Auth Command Options
|
|
671
|
-
|
|
672
|
-
| Option | Description | Default | Alias |
|
|
673
|
-
| ------------ | ------------------------- | ------- | ----- |
|
|
674
|
-
| --provider | Provider to log in with or configure (`copilot`, `codex`, `opencode-go`, `kimi`, `deepseek`, `dashscope`, `openrouter`, or `custom`) | prompt | none |
|
|
675
|
-
| --alias | Optional Codex account alias; only valid with `--provider codex` | none | none |
|
|
676
|
-
| --verbose | Enable verbose logging | false | -v |
|
|
677
|
-
| --show-token | Show GitHub token on auth | false | none |
|
|
678
|
-
|
|
679
|
-
Use `copilot-api auth login --provider copilot` only when you want to enable the GitHub Copilot provider. Copilot is not required for `codex` or third-party provider-only usage.
|
|
680
|
-
|
|
681
|
-
The Codex provider stores up to 3 accounts. Add or update an account with `copilot-api auth login --provider codex --alias work`; the newly signed-in account becomes active. List accounts with `copilot-api auth codex --list`, switch with `copilot-api auth codex --use <alias-or-accountId>`, and remove an account that is not in use with `copilot-api auth codex --remove <alias-or-accountId>`. Aliases are case-insensitive and cannot duplicate another account's alias or ID; the account currently in use cannot be removed. Restart a running server after switching or removing an account so it stops using the previous account; a credential refresh never writes a removed account back. Legacy single-account `codex_credentials.json` files remain compatible and upgrade to the multi-account shape on the next credential write.
|
|
682
|
-
|
|
683
|
-
Use `copilot-api auth login --provider deepseek`, `--provider dashscope`, `--provider openrouter`, `--provider opencode-go`, or `--provider kimi` to add or update those common third-party providers from the CLI. DeepSeek prompts for masked `apiKey`, provider `type` (default `anthropic`), and `baseUrl` defaulting to `https://api.deepseek.com/anthropic`. DashScope prompts for masked `apiKey`, provider `type` (default `openai-compatible`), and prefilled `baseUrl`. OpenRouter prompts for masked `apiKey` and prefilled `baseUrl` only, and writes `type: "anthropic"`. OpenCode Go prompts for masked `apiKey` and prefilled `baseUrl` only, and writes `type: "openai-compatible"` (baseUrl `https://opencode.ai/zen/go`). Kimi prompts for masked `apiKey`, provider `type` (default `openai-compatible`), and `baseUrl` defaulting to `https://api.kimi.com/coding` (the same base URL serves both the Anthropic and OpenAI-compatible endpoints). OpenCode Go selects the protocol from the model-level models.dev `provider.npm`, then the provider-level `npm`: `@ai-sdk/anthropic` uses Anthropic Messages, `@ai-sdk/openai` uses OpenAI Responses (unless the model has `shape: completions`), and unrecognized or missing packages use OpenAI-compatible. After a provider is configured and enabled, `copilot-api start` can run without any GitHub token.
|
|
684
|
-
|
|
685
|
-
OpenCode Go model listings (`/opencode-go/v1/models` and its entries in `/v1/models`) use the `opencode-go.models` section of [models.dev/api.json](https://models.dev/api.json). At startup the server loads `models-dev-api.json` from `COPILOT_API_HOME` (or the default app data directory). If a valid cache exists, the server starts listening and refreshes the catalog in the background every five minutes. Without a valid cache, startup waits for the first download and fails if it cannot obtain a catalog. Refreshes use `ETag` or `Last-Modified` when available; both validators are saved in `models-dev-api.meta.json` with a checksum of the JSON. A `304` keeps the existing cache, and a failed background refresh keeps the last valid disk or memory catalog. Models marked `status: deprecated` are excluded.
|
|
686
|
-
|
|
687
|
-
Use `copilot-api auth login --provider custom` to add or update another third-party provider from the CLI. Choose manual entry or search the existing models.dev catalog. The catalog selection includes providers with an HTTP API URL using Anthropic Messages, OpenAI-compatible Chat Completions, or OpenAI Responses; `openrouter`, `github-copilot`, and `opencode-go` are excluded. The selected provider ID, protocol, and API URL prefill editable fields, and models.dev model prices are used for token cost estimates in USD unless a model has explicit pricing in `config.json`. The command then prompts for masked `apiKey` and `authType`; `authType` may be left as the type default or set to `x-api-key` / `authorization`. The desktop custom-provider form offers the same catalog selection and manual entry.
|
|
688
|
-
|
|
689
|
-
Gateway API keys live under `auth.apiKeys` in `config.json`. Manage them with `copilot-api auth keys` (one operation per invocation): add a key with `--add <key>`, remove one with `--remove <key>`, list all with `--list`, or clear them all with `--clear`. Clients authenticate with any configured key via `x-api-key` or `Authorization: Bearer`. Without keys, loopback listeners start with authentication bypassed and print an info message; non-loopback listeners refuse to start.
|
|
690
|
-
|
|
691
|
-
### Debug Command Options
|
|
692
|
-
|
|
693
|
-
| Option | Description | Default | Alias |
|
|
694
|
-
| ------ | ------------------------- | ------- | ----- |
|
|
695
|
-
| --json | Output debug info as JSON | false | none |
|
|
696
|
-
|
|
697
|
-
## Configuration (config.json)
|
|
698
|
-
|
|
699
|
-
- **Location:** `~/.local/share/copilot-api/config.json` (Linux/macOS) or `%USERPROFILE%\.local\share\copilot-api\config.json` (Windows).
|
|
700
|
-
- **Default shape:**
|
|
701
|
-
```json
|
|
702
|
-
{
|
|
703
|
-
"auth": {
|
|
704
|
-
"apiKeys": [],
|
|
705
|
-
"adminApiKey": "<auto-generated-on-startup>"
|
|
706
|
-
},
|
|
707
|
-
"providers": {},
|
|
708
|
-
"modelMappings": {},
|
|
709
|
-
"smallModels": {
|
|
710
|
-
"codex": "gpt-6-luna",
|
|
711
|
-
"copilot": "gpt-6-luna"
|
|
712
|
-
},
|
|
713
|
-
"contextManagement": {
|
|
714
|
-
"messages": true,
|
|
715
|
-
"responses": false
|
|
716
|
-
},
|
|
717
|
-
"modelResponsesApiCompactThresholds": {
|
|
718
|
-
"gpt-5.4": 217600,
|
|
719
|
-
"gpt-5.5": 217600
|
|
720
|
-
},
|
|
721
|
-
"useMessagesApi": true,
|
|
722
|
-
"useResponsesApiWebSocket": true,
|
|
723
|
-
"upstreamTransport": {
|
|
724
|
-
"headersTimeoutMs": 300000,
|
|
725
|
-
"streamInactivityTimeoutMs": 300000,
|
|
726
|
-
"websocketOpenTimeoutMs": 30000,
|
|
727
|
-
"websocketPoolIdleTimeoutMs": 60000,
|
|
728
|
-
"websocketMaxBufferedBytes": 8388608,
|
|
729
|
-
"websocketMaxBufferedMessages": 1024
|
|
730
|
-
},
|
|
731
|
-
"useResponsesApiWebSearch": true,
|
|
732
|
-
"alphaSearchCodexPriority": true,
|
|
733
|
-
"alphaSearchModel": "gpt-6-luna",
|
|
734
|
-
"messageApiWebSearchModel": "gpt-6-luna"
|
|
735
|
-
}
|
|
736
|
-
```
|
|
737
|
-
- **auth.apiKeys:** API keys used for request authentication on non-admin routes. Supports multiple keys for rotation. Requests can authenticate with either `x-api-key: <key>` or `Authorization: Bearer <key>`. If empty or omitted, authentication for non-admin routes is disabled only on loopback listeners; non-loopback listeners refuse to start.
|
|
738
|
-
- **auth.adminApiKey:** Single admin key used only for `/admin/*` routes. If missing, the server generates a random key at startup and writes it back to `config.json`. Requests use the same `x-api-key` or `Authorization: Bearer` headers, but regular `auth.apiKeys` never grant access to `/admin/*`.
|
|
739
|
-
- **modelMappings:** Exact `sourceModel -> targetModel` rewrites shared by top-level `POST /v1/messages`, `POST /v1/messages/count_tokens`, `POST /v1/responses`, and `POST /v1/chat/completions` requests. Omit it or leave it as `{}` to disable rewrites. Both the source and target must be non-empty strings. Targets can be regular model IDs or `provider/model` aliases such as `dashscope/qwen3.6-plus`, and the rewrite happens before provider alias parsing. These mappings are not split per interface. The admin endpoints `GET/POST /admin/config/model-mappings` read and update only this field.
|
|
740
|
-
- **extraPrompts:** Map of `model -> prompt` appended to the first system prompt when translating Anthropic-style requests to Responses API. Use this to inject guardrails or guidance per model. For GPT-5.3+ models (e.g. `gpt-5.3-codex`, `gpt-5.4`, `gpt-5.5`), a built-in commentary prompt is used as fallback when not explicitly configured. The built-in prompts enable phase-aware commentary, which lets the model emit a short user-facing progress update before tools or deeper reasoning.
|
|
741
|
-
- **providers:** Global upstream provider map. Each provider key (for example `dashscope`) becomes a route prefix (`/dashscope/v1/messages`). Supports `type: "anthropic"`, `type: "openai-compatible"`, and `type: "openai-responses"`. Top-level clients can also use `model: "dashscope/model-id"` with `/v1/messages`, `/v1/messages/count_tokens`, `/v1/responses`, and `/v1/chat/completions`; the gateway strips the `dashscope/` prefix before forwarding upstream. The `/v1/responses` route for `anthropic` and `openai-compatible` providers uses the Responses Lite → Messages adapter; `openai-compatible` providers then reuse the Messages → Chat translation. Codex clients (`User-Agent` starting with `codex`) also use the adapter for non-`gpt-*` models on `openai-responses` providers. `GET /v1/models` aggregates enabled provider models with `provider/model-id` IDs, while the top-level Codex-UA catalog also merges these adaptable models as `use_responses_lite` entries (except DeepSeek models, which use `use_responses_lite: false` and `tool_mode: null`). Use `GET /dashscope/v1/models` for a single provider's raw model list.
|
|
742
|
-
- `enabled` defaults to `true` if omitted.
|
|
743
|
-
- `baseUrl` should be provider API base URL without the final endpoint. For Anthropic providers, omit `/v1/messages`; for OpenAI-compatible providers, omit `/v1/chat/completions`; for OpenAI Responses providers, omit `/v1/responses`.
|
|
744
|
-
- `modelsDevProviderId` (optional): Written when custom authorization selects a models.dev provider. In that case, `baseUrl` is the API URL supplied by models.dev (or your edited value), and the gateway appends `/messages`, `/chat/completions`, or `/responses` directly. Model-level protocol and USD pricing metadata from the cached catalog are applied when available; model-level API URL overrides apply while `baseUrl` matches the catalog's provider URL. Explicit `models.<id>.pricing` takes precedence. Manually entered providers continue to use `baseUrl` plus `/v1/<endpoint>`.
|
|
745
|
-
- `apiKey` is used as the upstream credential value and is required unless `authType` is `azure-entra`.
|
|
746
|
-
- `authType` (optional): Controls upstream authentication. Supports `x-api-key`, `authorization`, and `azure-entra` for regular providers. Anthropic providers default to `x-api-key`; OpenAI-compatible and OpenAI Responses providers default to `authorization`. `authorization` sends `Authorization: Bearer <apiKey>`. `azure-entra` uses Azure Identity's `DefaultAzureCredential` with the `https://cognitiveservices.azure.com/.default` scope, sends the resulting bearer token, and does not require `apiKey`. For an Azure OpenAI v1 endpoint, use a provider such as `{ "type": "openai-compatible", "baseUrl": "https://<resource-name>.openai.azure.com/openai", "authType": "azure-entra" }`. Authenticate locally with `az login`, use a managed identity in Azure, or set the standard `AZURE_TENANT_ID`, `AZURE_CLIENT_ID`, and `AZURE_CLIENT_SECRET` environment variables. `oauth2` is reserved for the built-in `codex` provider and is written automatically by `auth login --provider codex`.
|
|
747
|
-
- `accountId` is used only by the built-in `codex` provider and identifies its active Codex account. CLI and desktop account switching update it automatically.
|
|
748
|
-
- `pricingCurrency` (optional): Provider-level currency used for token cost calculation, for example `USD` or `CNY`. Quick providers default to `CNY` for DashScope and DeepSeek, and `USD` for Codex, Kimi, OpenCode Go, and OpenRouter. Costs are grouped by currency and are not exchange-rate converted.
|
|
749
|
-
- `models` (optional): Per-model configuration map. Each key is a model ID (matching the model name in requests), and the value is:
|
|
750
|
-
- `temperature` (optional): Default temperature value used when the request does not specify one.
|
|
751
|
-
- `topP` (optional): Default top_p value used when the request does not specify one.
|
|
752
|
-
- `topK` (optional): Default top_k value used when the request does not specify one.
|
|
753
|
-
- `extraBody` (optional): Dynamic fields merged into the upstream request body for that model. Request body fields with the same name take precedence. OpenAI-compatible providers can use this for fields such as `enable_thinking`, `preserve_thinking`, `reasoning_effort`. `thinking_budget` is a special OpenAI-compatible provider override: when configured in `extraBody`, it is forced after Anthropic `thinking.budget_tokens` translation and overrides the request-derived budget. For providers whose name is `dashscope` or whose `baseUrl` contains `aliyuncs.com`, the request-derived `thinking_budget` (from Anthropic `thinking.budget_tokens`) is forwarded upstream; for other OpenAI-compatible providers the request-derived `thinking_budget` is stripped, while an `extraBody` `thinking_budget` is still honored. For DashScope providers, `preserve_thinking` defaults to `true` when not explicitly set in `extraBody` or the request body.
|
|
754
|
-
- `pricing` (optional): Per-model token prices, in the provider `pricingCurrency`, per 1M tokens. Supported fields are `input`, `output`, `cachedInput` (implicit cache read), `explicitCachedInput` (explicit cache read), and `cacheCreationInput`. Use `tiers` with `maxInputTokens` for input-size tiered pricing. For providers that bill separate peak and off-peak rates, add `offPeak` with the same fields as the top level and declare peak hours with `peakWindows`: each entry takes `startMinuteUtc` (inclusive) and `endMinuteUtc` (exclusive) as minutes from 00:00 UTC, plus optional `weekdays` ISO day numbers (1 = Monday, 7 = Sunday, omitted means every day). Without `peakWindows` the `offPeak` prices never apply and every request bills at the peak price. The built-in catalog uses this for DeepSeek (peak on weekdays UTC 01:00-04:00 and 06:00-10:00, off-peak otherwise including weekends) and for DashScope DeepSeek (off-peak on UTC 14:00-24:00). OpenCode Go prices use the current models.dev values.
|
|
755
|
-
- `contextCache` (optional): Defaults to `true` for providers whose name is `dashscope` or whose `baseUrl` contains `aliyuncs.com`; defaults to `false` for other OpenAI-compatible providers. This enables Alibaba Cloud Model Studio/DashScope explicit context cache by injecting `cache_control: { "type": "ephemeral" }` on up to 4 content blocks using the Context Cache format. The cache breakpoint strategy matches opencode's main provider flow: the first 2 system messages plus the last 2 non-system messages. Marked string content is converted to text content part arrays for `system` / `user` / `assistant` / `tool` messages; existing array content is marked on the last part. Set this to `false` when the model already supports implicit caching, or when the upstream does not accept this explicit-cache extension field. Set this to `true` for non-DashScope providers that support the same explicit-cache extension. Applied on both `/v1/messages` and `/v1/chat/completions` routes.
|
|
756
|
-
- `supportPdf` (optional): Controls whether the model supports PDF/document content. Defaults to `false`; unsupported PDFs are converted to a text notice. Set it to `true` to send PDF/document blocks as OpenAI Chat Completions file parts.
|
|
757
|
-
- `toolContentSupportType` (optional): Tool result content capabilities for that model, as an array of `array`, `image`, and `pdf`. Provider routes default to string-only tool content when omitted. If `supportPdf` is `true` but this list does not include `pdf`, file parts in tool results are moved to user role messages. The Copilot main flow uses the same string-only default, because some Copilot models do not support array or image tool content either.
|
|
758
|
-
- `type` (optional): Per-model override of the provider protocol type. Supports `anthropic`, `openai-compatible`, and `openai-responses`. When set, the provider's `/v1/messages` route uses this model's type instead of the provider-level type for request routing, auth header resolution, and upstream endpoint selection. This is useful for providers like OpenCode Go whose upstream supports both OpenAI-compatible and Anthropic Messages APIs for different models. When the type is overridden, the auth header is resolved from the overridden type's default (Anthropic defaults to `x-api-key`; OpenAI-compatible/Responses default to `authorization`). Providers configured with `azure-entra` keep their Entra bearer credential instead of falling back to the overridden type's default.
|
|
759
|
-
- `contextWindow` (optional): Context window token limit advertised when this model is merged into the Codex-UA model catalog; for example, `1000000` declares a 1M-token context window. Missing configured values use upstream metadata first, then the built-in non-GPT model catalog, then `256000`.
|
|
760
|
-
- `maxOutputTokens` (optional): Maximum output token limit advertised in the Codex-UA model catalog. Missing configured values use upstream metadata first, then the built-in non-GPT model catalog, where defaults are capped at `64000`, then `32000`.
|
|
761
|
-
- `inputModalities` (optional): Supported Codex input types. Use `["text", "image"]` for a model that accepts both text and images. Missing configured values use upstream metadata before the built-in non-GPT model catalog. GPT models do not receive these built-in capability defaults and continue to use the native Codex catalog or upstream metadata.
|
|
762
|
-
- `reasoningEfforts` (optional): Reasoning levels advertised for Codex. Missing configured and upstream values use the built-in non-GPT model catalog before falling back to `["high", "xhigh", "max", "ultra"]`. Provider Responses requests with an unsupported effort are normalized to a supported level when these capabilities are known.
|
|
763
|
-
- `defaultReasoningEffort` (optional): Default Codex reasoning level. Built-in model metadata may provide a known default; otherwise it defaults to `max` when available, then the first configured level. Synthetic Codex models always enable parallel tool calls.
|
|
764
|
-
- `reasoningField` (optional): Assistant thinking field sent upstream on OpenAI-compatible `/v1/messages` requests. Supports `reasoning` and `reasoning_content`; defaults to `reasoning_content`. Use `reasoning` for OpenRouter-style models; OpenCode Go Hy-family models use this field based on the models.dev family metadata.
|
|
765
|
-
- **smallModels:** Per-provider model IDs for Codex task-title `/v1/responses` requests. `codex` and `copilot` default to `gpt-6-luna`; other providers switch only when named explicitly (e.g. `smallModels.dashscope`). Use a model supported by that provider; an empty string disables switching. The request must have matching user `input_text` in one of its last two input items. `copilot` also controls tool-less Claude Code warmups for non-token-based-billing accounts. The old `smallModel` key is ignored.
|
|
766
|
-
- **contextManagement:** Controls whether the proxy adds Responses API `context_management` compaction instructions. `messages` applies when Anthropic-style `/v1/messages` requests are translated to Responses API, including `openai-responses` provider message routes, and defaults to `true`. `responses` applies to native `/v1/responses` traffic, including `provider/model` aliases and the built-in `codex` provider, and defaults to `false`. Enable `responses` only after checking that your client supports context management compaction. When enabled, the request includes `context_management` in the body and keeps only the latest compaction carrier on follow-up turns. The proxy only adds context management and compacts history for `gpt-*` models; both configuration switches have no effect on non-GPT models such as Grok. **Note:** Context management is also forcibly disabled for GPT-5.6 and above models (e.g. `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`) because enabling it breaks prompt cache hits on those models. These overrides take precedence over the `contextManagement` and `modelResponsesApiCompactThresholds` settings.
|
|
767
|
-
- **modelResponsesApiCompactThresholds:** Per-model Responses API `compact_threshold` overrides used when the proxy adds `context_management`. These values take precedence over the fallback threshold from `resolveResponsesCompactThreshold` (`max_prompt_tokens * ratio`, or the default fallback). Defaults set `gpt-5.4` and `gpt-5.5` to `217600` (`272000 * 0.8`). Models not listed continue to use the normal fallback logic.
|
|
768
|
-
- **modelReasoningEfforts:** Per-model fallback reasoning effort for `/v1/messages` requests. It is used only when the request does not provide `output_config.effort`.
|
|
769
|
-
- **Priority:** request `output_config.effort` > `modelReasoningEfforts[model]` > built-in default (`xhigh` for GPT-5.3+ models, otherwise `high`).
|
|
770
|
-
- **Forwarding:** the resolved value remains `output_config.effort` for the Copilot native Messages API and becomes `reasoning.effort` when translated to the Responses API.
|
|
771
|
-
- **Configuration values:** `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, and `max`.
|
|
772
|
-
- **useMessagesApi:** When `true`, models that advertise Copilot's native `/v1/messages` endpoint use the Messages API. If Messages is disabled or unavailable for the selected model, the gateway uses Responses when that model advertises a Responses endpoint, then falls back to Chat Completions when supported. Set this to `false` to skip native Messages routing. Defaults to `true`.
|
|
773
|
-
- **useResponsesApiWebSocket:** When `true`, Copilot Responses requests use WebSocket for models that advertise `ws:/responses`; models that advertise only `/responses` use HTTP. Streamed Responses requests for the built-in `codex` provider use WebSocket whenever this setting is enabled, while non-streaming Codex requests always use HTTP. Set this to `false` to make Copilot use HTTP `/responses` where the selected model advertises it and to send streamed Codex Responses requests over HTTP. WebSocket failures are not retried automatically over HTTP. Defaults to `true`. If a proxy, VPN, or network blocks or destabilizes WebSocket traffic, disable this setting or switch networks. Also set it to `false` when the GitHub Copilot provider fails with `Encrypted function output content could not be decrypted or decoded`; see [Troubleshooting](#troubleshooting).
|
|
774
|
-
- **upstreamTransport:** Positive integer lifecycle and buffering limits for the upstream chat completions, responses, and messages transports. Invalid, zero, or negative values fall back to the defaults shown above. `headersTimeoutMs` covers connection setup through receipt of HTTP response headers; it is not a total generation deadline. `streamInactivityTimeoutMs` is reset by every HTTP body chunk or WebSocket message, allowing long generations to continue while they remain active. `websocketOpenTimeoutMs` limits the WebSocket handshake, while `websocketPoolIdleTimeoutMs` controls only completed, reusable pooled sockets. The byte and message limits bound queued WebSocket events; exceeding either limit fails that stream and invalidates its socket rather than dropping or reordering events.
|
|
775
|
-
- **useResponsesApiWebSearch:** When `true`, the server keeps Responses API tools with `type: "web_search"` and forwards them upstream. Set to `false` to strip those tools from `/responses` payloads. Defaults to `true`.
|
|
776
|
-
- **alphaSearchCodexPriority:** Defaults to `true`. Top-level alpha-search requests prefer the Codex alpha-search endpoint because it does not consume provider quota. If Codex is unavailable, or this setting is `false`, requests with a `provider/model` alias other than `codex/model` use that provider's `/v1/responses` endpoint, and requests without a provider prefix use GitHub Copilot Responses web search. The adapter recognizes every current Codex search command; unsupported `image_query` and `screenshot` operations return successful no-retry tool output.
|
|
777
|
-
- **alphaSearchModel:** Native Responses search model used when a Messages-backed Responses Lite model cannot run Responses web search directly. Defaults to `gpt-6-luna`; it may be a regular Copilot model or an `openai-responses` `provider/model` alias. Set it to an empty string to disable this redirect, in which case alpha-search requests for those models return an invalid-request error.
|
|
778
|
-
- **messageApiWebSearchModel:** Global fallback model used when a top-level Copilot `/v1/messages` request contains only the server-side `web_search` tool. Defaults to `gpt-6-luna`. If the value is a `provider/model` alias, the request is routed into that provider's Messages API path with the provider prefix stripped. For Copilot GPT models, web search runs through `/responses`. Mixed `web_search` plus custom tools are not supported and the server-side `web_search` tool is stripped.
|
|
779
|
-
- **claudeAutoModel:** Model used for Claude Code background security-monitor requests on `/v1/messages` and provider message routes. A request is treated as a security-monitor request when it carries no tools, sets `stop_sequences` to `["</block>"]`, and contains a system text block starting with `You are a security monitor for autonomous AI coding agents.`; its model is then replaced with this value. For top-level requests, a `provider/model` alias is forwarded into that provider's Messages API; provider routes keep their current provider and use this configured value directly. Defaults to empty (disabled).
|
|
780
|
-
- **claudeTokenMultiplier:** Multiplier applied to the fallback GPT-tokenizer estimate for Claude `/v1/messages/count_tokens` requests. Defaults to `1.15`. Increase it if your client is still compacting too late. This setting is only used when the proxy is estimating Claude tokens locally; if `anthropicApiKey` is configured and Anthropic token counting succeeds, the exact Anthropic count is returned instead.
|
|
781
|
-
- **anthropicApiKey:** Anthropic API key used to forward Claude `/v1/messages/count_tokens` requests to Anthropic's real token counting endpoint, which returns exact counts instead of GPT tokenizer estimates. Can also be set via the `ANTHROPIC_API_KEY` environment variable. If not set, or if the upstream call fails, token counting falls back to local GPT tokenizer estimation controlled by `claudeTokenMultiplier`.
|
|
782
|
-
|
|
783
|
-
Edit this file to customize prompts or swap in your own fast model. Restart the server (or rerun the command) after changes so the cached config is refreshed.
|
|
784
|
-
|
|
785
|
-
## API Authentication
|
|
786
|
-
|
|
787
|
-
- **Protected non-admin routes:** All routes except `/`, `/usage-viewer`, and `/usage-viewer/` require authentication when `auth.apiKeys` is configured and non-empty. Non-loopback listeners require a non-empty `auth.apiKeys` configuration at startup and continue failing closed if the keys are later cleared.
|
|
788
|
-
- **Admin routes:** All `/admin/*` routes require `auth.adminApiKey`. If it is missing, the server generates one at startup and persists it to `config.json` before serving requests.
|
|
789
|
-
- **Allowed auth headers:**
|
|
790
|
-
- `x-api-key: <your_key>`
|
|
791
|
-
- `Authorization: Bearer <your_key>`
|
|
792
|
-
- **CORS preflight:** `OPTIONS` requests are always allowed.
|
|
793
|
-
- **When no regular keys are configured:** Non-admin routes continue to allow requests. This does not apply to `/admin/*`, which only accepts `auth.adminApiKey`.
|
|
794
|
-
|
|
795
|
-
Example request for a regular protected route:
|
|
796
|
-
|
|
797
|
-
```sh
|
|
798
|
-
curl http://localhost:4141/v1/models \
|
|
799
|
-
-H "x-api-key: your_api_key"
|
|
800
|
-
```
|
|
801
|
-
|
|
802
|
-
Example request for an admin route:
|
|
803
|
-
|
|
804
|
-
```sh
|
|
805
|
-
curl http://localhost:4141/admin/config/model-mappings \
|
|
806
|
-
-H "x-api-key: your_admin_api_key"
|
|
807
|
-
```
|
|
808
|
-
|
|
809
|
-
## API Endpoints
|
|
810
|
-
|
|
811
|
-
The server exposes several OpenAI- and Anthropic-compatible endpoints. Requests can target GitHub Copilot, the built-in `codex` provider, or configured providers depending on the selected model and `provider/model` alias. Every `/v1/...` endpoint below also supports a provider-scoped path in the form `/:provider/v1/...`; those variants are omitted from the tables.
|
|
812
|
-
|
|
813
|
-
### OpenAI Compatible Endpoints
|
|
814
|
-
|
|
815
|
-
These endpoints mimic the OpenAI API structure.
|
|
816
|
-
|
|
817
|
-
| Endpoint | Method | Description |
|
|
818
|
-
| --------------------------- | ------ | ---------------------------------------------------------------- |
|
|
819
|
-
| `POST /v1/responses` | `POST` | OpenAI Most advanced interface for generating model responses. Supports `Content-Encoding: zstd` request bodies and `provider/model` aliases for `openai-responses` providers. Zstd request decompression is limited to Responses routes, including provider-scoped aliases. |
|
|
820
|
-
| `POST /v1/chat/completions` | `POST` | Creates a model response for the given chat conversation. Supports `provider/model` aliases for `openai-compatible` providers and can be used without Copilot when the target provider is configured. |
|
|
821
|
-
| `GET /v1/models` | `GET` | Lists Copilot models plus enabled provider models using `provider/model-id` IDs. Requests from Codex clients (`User-Agent` beginning with `codex`) are forwarded to the Codex Models upstream. |
|
|
822
|
-
| `POST /v1/embeddings` | `POST` | Creates an embedding vector representing the input text. |
|
|
823
|
-
|
|
824
|
-
### Codex Backend Endpoints
|
|
825
|
-
|
|
826
|
-
These endpoints implement Codex backend APIs. Top-level image requests require an active Codex login; alpha search can use either the Codex backend or a Responses web-search adapter.
|
|
827
|
-
|
|
828
|
-
| Endpoint | Method | Description |
|
|
829
|
-
| -------------------------------------------------------------- | ------ | --------------------------------------------------------------- |
|
|
830
|
-
| `POST /v1/alpha/search` | `POST` | Routes Codex alpha-search requests to the Codex backend, or handles supported commands locally and through Responses web search. |
|
|
831
|
-
| `POST /v1/images/generations` | `POST` | Forwards a JSON image generation request to the Codex Images upstream. When the request omits `Content-Type`, the gateway defaults it to `application/json`. Configured model mappings apply to the request `model`; a mapping that resolves to a `provider/model` alias forwards the request to that provider's images endpoint when the provider is configured. |
|
|
832
|
-
| `POST /v1/images/edits` | `POST` | Forwards an image edit request to the Codex Images upstream. Send this request as `multipart/form-data` and let the HTTP client generate the `boundary`. The gateway streams uploaded files to temporary disk files while receiving them and forwards them from disk, so large uploads are not held in memory. Multipart requests are limited to 128 MiB total, 64 MiB per file, and 16 files; requests over a limit return `413`. Model mappings and `provider/model` alias routing apply to this endpoint as well. |
|
|
833
|
-
|
|
834
|
-
For requests routed to the Codex backend, the gateway replaces client authorization and account headers with the active Codex login and preserves compatible request metadata. Responses-backed alpha search instead follows the selected Copilot or provider route.
|
|
835
|
-
|
|
836
|
-
### Anthropic Compatible Endpoints
|
|
837
|
-
|
|
838
|
-
These endpoints are designed to be compatible with the Anthropic Messages API.
|
|
839
|
-
|
|
840
|
-
| Endpoint | Method | Description |
|
|
841
|
-
| -------------------------------- | ------ | ------------------------------------------------------------ |
|
|
842
|
-
| `POST /v1/messages` | `POST` | Creates a model response for a given conversation. Supports `provider/model` aliases for configured providers, including translation through `openai-compatible` providers. |
|
|
843
|
-
| `POST /v1/messages/count_tokens` | `POST` | Calculates the number of tokens for a given set of messages. Supports `provider/model` aliases for configured providers. |
|
|
844
|
-
|
|
845
|
-
### Usage Monitoring Endpoints
|
|
846
|
-
|
|
847
|
-
New endpoints for monitoring your Copilot usage and quotas.
|
|
848
|
-
|
|
849
|
-
| Endpoint | Method | Description |
|
|
850
|
-
| ------------ | ------ | ------------------------------------------------------------ |
|
|
851
|
-
| `GET /usage` | `GET` | Get detailed Copilot usage statistics and quota information. |
|
|
852
|
-
|
|
853
|
-
### Admin / Configuration Endpoints
|
|
854
|
-
|
|
855
|
-
These endpoints are reserved for local administrative actions and only accept `auth.adminApiKey`.
|
|
856
|
-
|
|
857
|
-
| Endpoint | Method | Description |
|
|
858
|
-
| ------------------------------------- | ------ | --------------------------------------------------------------------------- |
|
|
859
|
-
| `GET /admin/config/model-mappings` | `GET` | Returns the current `config.json` path and the active `modelMappings` map. |
|
|
860
|
-
| `POST /admin/config/model-mappings` | `POST` | Updates only the `modelMappings` field in `config.json` and returns it back. |
|
|
861
|
-
|
|
862
|
-
## Example Usage
|
|
863
|
-
|
|
864
|
-
Common `npx` commands:
|
|
865
|
-
|
|
866
|
-
```sh
|
|
867
|
-
# Start the gateway
|
|
868
|
-
npx @jeffreycao/copilot-api@latest start
|
|
869
|
-
|
|
870
|
-
# Start on a custom port with verbose logging
|
|
871
|
-
npx @jeffreycao/copilot-api@latest start --port 8080 --verbose
|
|
872
|
-
|
|
873
|
-
# Run the auth flow
|
|
874
|
-
npx @jeffreycao/copilot-api@latest auth login
|
|
875
|
-
|
|
876
|
-
# Configure a third-party provider, then run without GitHub Copilot
|
|
877
|
-
npx @jeffreycao/copilot-api@latest auth login --provider dashscope
|
|
878
|
-
npx @jeffreycao/copilot-api@latest start
|
|
879
|
-
|
|
880
|
-
# Print debug information as JSON
|
|
881
|
-
npx @jeffreycao/copilot-api@latest debug --json
|
|
882
|
-
|
|
883
|
-
# Run the published CLI with Bun instead of Node.js
|
|
884
|
-
bunx --bun @jeffreycao/copilot-api@latest start
|
|
885
|
-
```
|
|
886
|
-
|
|
887
|
-
OpenAI-compatible provider examples after configuring `dashscope`:
|
|
888
|
-
|
|
889
|
-
```sh
|
|
890
|
-
curl http://localhost:4141/v1/chat/completions \
|
|
891
|
-
-H "content-type: application/json" \
|
|
892
|
-
-d '{"model":"dashscope/qwen3.6-plus","messages":[{"role":"user","content":"hello"}]}'
|
|
893
|
-
|
|
894
|
-
curl http://localhost:4141/dashscope/v1/messages \
|
|
895
|
-
-H "content-type: application/json" \
|
|
896
|
-
-d '{"model":"qwen3.6-plus","max_tokens":1024,"messages":[{"role":"user","content":"hello"}]}'
|
|
897
|
-
```
|
|
898
|
-
|
|
899
|
-
<a id="troubleshooting"></a>
|
|
900
|
-
|
|
901
|
-
## Troubleshooting
|
|
902
|
-
|
|
903
|
-
**GitHub Copilot encrypted output decryption failure**
|
|
904
|
-
|
|
905
|
-
When the GitHub Copilot provider returns or logs `Encrypted function output content could not be decrypted or decoded`, it is usually an upstream issue. Set `useResponsesApiWebSocket` to `false` in `config.json` so Copilot Responses traffic goes over HTTP `/responses` instead:
|
|
906
|
-
|
|
907
|
-
```json
|
|
908
|
-
{
|
|
909
|
-
"useResponsesApiWebSocket": false
|
|
910
|
-
}
|
|
911
|
-
```
|
|
912
|
-
|
|
913
|
-
Restart the server after changing the config. See [Configuration (config.json)](#configuration-configjson) for the full option reference.
|
|
85
|
+
Windows x64 (`.exe`), macOS Apple Silicon (`.dmg`), and Linux x64 (`.AppImage`) packages are published in [GitHub Releases](https://github.com/caozhiyuan/copilot-api/releases). See [Electron Desktop App](docs/guides/en/desktop.md#electron-desktop-app) for full setup and advanced configuration.
|
|
86
|
+
|
|
87
|
+
## Documentation
|
|
88
|
+
|
|
89
|
+
| Guide | Contents |
|
|
90
|
+
| --- | --- |
|
|
91
|
+
| [Installation and Startup](docs/guides/en/getting-started.md) | Prerequisites, project overview, `npx` and source runs, provider-only mode without Copilot, and gateway API key setup |
|
|
92
|
+
| [Claude Code](docs/guides/en/claude-code.md) | The `--claude-code` interactive launcher, `.claude/settings.json` environment variables, opus / sonnet / haiku tier mapping, auto-compact window, and WebSearch behavior |
|
|
93
|
+
| [OpenCode](docs/guides/en/opencode.md) | OpenCode OAuth login, the `@ai-sdk/anthropic` provider in `opencode.json`, `baseURL` conventions, model context limits, and thinking options |
|
|
94
|
+
| [Codex](docs/guides/en/codex.md) | A full `config.toml` provider block, auto-review model mapping, generating `model_catalog.json`, and the merged model picker catalog with protocol adapters |
|
|
95
|
+
| [Docker](docs/guides/en/docker.md) | Docker Compose quick start, the `/data` persistent mount and its ownership repair, supported environment variables, and host interface binding |
|
|
96
|
+
| [Desktop App](docs/guides/en/desktop.md) | Copilot sign-in, Codex OAuth account switching, API-key providers, one-click start / stop, shared model mappings, advanced settings, and per-platform installers |
|
|
97
|
+
| [Plugins and Tool Search](docs/guides/en/integrations.md) | The Responses `tool_search` MCP bridge, Claude Code `agent-inject` and `tool-search` marketplace plugins, and the opencode subagent marker plugin |
|
|
98
|
+
| [Usage Monitoring](docs/guides/en/usage.md) | The usage viewer URL and query parameters, period selectors, Copilot quota progress, token and cost metric cards, trend charts, and paginated request events |
|
|
99
|
+
| [CLI Reference](docs/guides/en/cli.md) | Command structure, global options, and the full option sets for the `start`, `auth`, and `debug` subcommands with example usage |
|
|
100
|
+
| [Configuration Reference](docs/guides/en/configuration.md) | Every `config.json` field: gateway and admin API keys, provider definitions, model mappings, WebSocket and HTTP transport, timeouts, and context management |
|
|
101
|
+
| [API and Authentication](docs/guides/en/api.md) | Allowed auth headers and CORS rules, OpenAI, Codex backend, and Anthropic endpoints, usage monitoring routes, and admin configuration endpoints |
|
|
102
|
+
| [Troubleshooting](docs/guides/en/troubleshooting.md) | Known issues with workarounds, including the Copilot encrypted output decryption failure and the `useResponsesApiWebSocket` HTTP fallback |
|
|
103
|
+
|
|
104
|
+
Before using GitHub Copilot, read the [GitHub Copilot Security Notice](NOTICE.md#github-copilot-security-notice).
|