orchestrator-mcp-server 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,22 @@
1
+ name: tests
2
+
3
+ on:
4
+ push:
5
+ branches: [main]
6
+ pull_request:
7
+
8
+ jobs:
9
+ pytest:
10
+ runs-on: ubuntu-latest
11
+ strategy:
12
+ fail-fast: false
13
+ matrix:
14
+ python: ["3.11", "3.12", "3.13"]
15
+ steps:
16
+ - uses: actions/checkout@v4
17
+ - uses: astral-sh/setup-uv@v5
18
+ with:
19
+ python-version: ${{ matrix.python }}
20
+ - run: uv sync --all-extras --dev
21
+ # No network: every deployment is stubbed, so this needs no API keys.
22
+ - run: uv run pytest -q
@@ -0,0 +1,9 @@
1
+ .venv/
2
+ __pycache__/
3
+ *.py[cod]
4
+ .pytest_cache/
5
+ .env
6
+ dist/
7
+
8
+ # Real config carries endpoints and key references; only the example is tracked.
9
+ config.yaml
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Ayberk Karataban
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,240 @@
1
+ Metadata-Version: 2.4
2
+ Name: orchestrator-mcp-server
3
+ Version: 0.1.0
4
+ Summary: Capability-routed MCP server: route work to the LLM that is best at it.
5
+ Project-URL: Homepage, https://github.com/crAK1644/orchestrator-mcp
6
+ Project-URL: Issues, https://github.com/crAK1644/orchestrator-mcp/issues
7
+ Author: Ayberk Karataban
8
+ License-Expression: MIT
9
+ License-File: LICENSE
10
+ Keywords: claude,codex,litellm,llm,mcp,orchestrator,router
11
+ Classifier: Development Status :: 4 - Beta
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: Programming Language :: Python :: 3 :: Only
14
+ Classifier: Topic :: Software Development :: Libraries
15
+ Requires-Python: >=3.11
16
+ Requires-Dist: jsonschema>=4.23
17
+ Requires-Dist: litellm>=1.55.0
18
+ Requires-Dist: mcp<3,>=2.0
19
+ Requires-Dist: pydantic>=2.9
20
+ Requires-Dist: pyyaml>=6.0
21
+ Description-Content-Type: text/markdown
22
+
23
+ # orchestrator-mcp
24
+
25
+ An MCP server that routes each request to the model configured for that kind of work —
26
+ research to one model, coding to another — and returns every answer through a fixed,
27
+ validated envelope.
28
+
29
+ Pointing a capability at your own deployment is a YAML edit. There is no code to change.
30
+
31
+ > Published to PyPI as **`orchestrator-mcp-server`** — the shorter name is an empty
32
+ > registered project owned by someone else. The import package is `orchestrator_mcp`.
33
+
34
+ ## How it works
35
+
36
+ A **capability** is a LiteLLM `model_name` alias group. Several deployments can share
37
+ one name, and `litellm.Router` already load-balances, retries, cools down, and falls
38
+ back across them. So the routing engine is the config file:
39
+
40
+ ```yaml
41
+ model_list:
42
+ - model_name: coding # capability, not a model
43
+ litellm_params:
44
+ model: anthropic/claude-sonnet-4-5
45
+ api_key: os.environ/ANTHROPIC_API_KEY
46
+
47
+ - model_name: coding # same capability, your own box
48
+ litellm_params:
49
+ model: openai/qwen-coder
50
+ api_base: http://vllm.internal:8000/v1
51
+ api_key: os.environ/LOCAL_VLLM_KEY
52
+ ```
53
+
54
+ The calling agent states which capability it wants. There is no intent classifier —
55
+ the caller is already a language model and knows whether it is asking a coding
56
+ question; paying a second model to guess what the first one already knows buys a cost
57
+ increase and a new failure mode.
58
+
59
+ ## Quick start
60
+
61
+ Write a `config.yaml` — start from [`config.example.yaml`](config.example.yaml) — and
62
+ check that it loads:
63
+
64
+ ```bash
65
+ ORCHESTRATOR_CONFIG=config.yaml uvx --from orchestrator-mcp-server python -c "from orchestrator_mcp.server import build_server; build_server(); print('config ok')"
66
+ ```
67
+
68
+ A bad config fails here rather than at request time: every deployment must route to a
69
+ declared capability, every capability must have a deployment behind it, and every
70
+ fallback must name a real capability.
71
+
72
+ Because the file is LiteLLM's own config schema, `litellm --config config.yaml` runs on
73
+ it unchanged. Keep it out of version control — it holds your endpoints.
74
+
75
+ ### Claude Code
76
+
77
+ ```bash
78
+ claude mcp add orchestrator --env ORCHESTRATOR_CONFIG=$PWD/config.yaml -- uvx orchestrator-mcp-server
79
+ ```
80
+
81
+ ### Codex
82
+
83
+ In `~/.codex/config.toml`:
84
+
85
+ ```toml
86
+ [mcp_servers.orchestrator]
87
+ command = "uvx"
88
+ args = ["orchestrator-mcp-server"]
89
+ env = { ORCHESTRATOR_CONFIG = "/absolute/path/to/config.yaml" }
90
+ ```
91
+
92
+ Both speak stdio, and every tool result is returned as structured content *and* as
93
+ JSON text, so a client that reads only one of the two still gets the whole envelope.
94
+
95
+ **Provider keys.** `config.yaml` references them as `os.environ/NAME`, and they are
96
+ read from the environment the *server* process gets — which is the client's
97
+ environment, not your shell's. If a capability comes back `auth_failed` while the same
98
+ config works from a terminal, add the key to the client's `env` block.
99
+
100
+ ### From a checkout
101
+
102
+ ```bash
103
+ uv sync && uv run pytest -q
104
+ ```
105
+
106
+ ```bash
107
+ claude mcp add orchestrator --env ORCHESTRATOR_CONFIG=$PWD/config.yaml -- uv run --directory $PWD orchestrator-mcp-server
108
+ ```
109
+
110
+ ## Tools
111
+
112
+ ### `ask`
113
+
114
+ | Argument | Notes |
115
+ |---|---|
116
+ | `capability` | Enum, built from your config. Bad values are rejected by the protocol layer. |
117
+ | `prompt` | Required. Capped by `limits.max_prompt_chars`. |
118
+ | `context` | Source material. When set, the model is told to answer only from it and to abstain otherwise. |
119
+ | `system` | Extra instructions. Applied *before* the server's own directives, so it cannot disable them. Capped by `limits.max_system_chars`. |
120
+ | `response_schema` | JSON Schema (`"type": "object"`). Switches on structured mode. Capped by `limits.max_schema_chars` — it is inlined into the prompt verbatim. |
121
+ | `temperature` | Pinned to `0` whenever `response_schema` is set. |
122
+ | `max_output_tokens` | Capped by `limits.max_output_tokens`. |
123
+
124
+ Every call returns the same envelope:
125
+
126
+ ```json
127
+ {
128
+ "ok": true,
129
+ "content": "…",
130
+ "data": null,
131
+ "insufficient_context": false,
132
+ "capability_requested": "coding",
133
+ "model_used": "anthropic/claude-sonnet-4-5",
134
+ "fallback_used": false,
135
+ "finish_reason": "stop",
136
+ "usage": { "prompt_tokens": 10, "completion_tokens": 20, "total_tokens": 30, "cost_usd": 0.0002 },
137
+ "latency_ms": 412,
138
+ "error": null
139
+ }
140
+ ```
141
+
142
+ `content` holds the prose answer and is `null` in structured mode and on failure;
143
+ `data` holds the validated object and is set only in structured mode; `error` is
144
+ `{ code, message }` whenever `ok` is `false`.
145
+
146
+ `error.code` comes from a closed set — `invalid_request`, `no_deployment`,
147
+ `upstream_error`, `rate_limited`, `context_exceeded`, `schema_validation_failed`,
148
+ `timeout`, `content_filtered`, `auth_failed`, `output_truncated` — so callers branch
149
+ on a value instead of matching substrings. `error.message` is bounded at 500
150
+ characters and never quotes the rejected output back at you.
151
+
152
+ ### `list_capabilities`
153
+
154
+ What each capability is for, the deployments behind it, and where it falls back.
155
+
156
+ ## Guardrails, and their limits
157
+
158
+ This server sees a prompt and a completion. It has no ground truth, so it **cannot**
159
+ verify factual claims, and nothing here should be read as a hallucination detector.
160
+ What it does enforce:
161
+
162
+ - **Shape is validated, not assumed.** Structured replies are checked against your
163
+ schema locally with `jsonschema`, regardless of whether the provider claims to
164
+ enforce `response_format`. A violation is a failure, not a payload.
165
+ - **Bounded repair.** An invalid structured reply gets `limits.schema_repair_attempts`
166
+ retries carrying the validator's complaint, then fails as
167
+ `schema_validation_failed`. Never a best-effort half-parsed object.
168
+ - **An unfinished answer is a failure, not a short answer.** A completion cut off by
169
+ the token limit comes back as `output_truncated` with `content: null`, and one the
170
+ provider filtered as `content_filtered`. Neither is returned as prose, because a
171
+ half answer reads exactly like a whole one.
172
+ - **The error tells you what broke, not what the model wrote.** `error.message` gives
173
+ the failing path and constraint (`schema violation at answer/city: failed the
174
+ 'maxLength' constraint`) and is capped at 500 characters. The rejected value itself
175
+ goes only back to the model that produced it, in the repair turn.
176
+ - **`request_timeout_s` bounds the call.** Retries, cross-capability fallback, and
177
+ repair turns all spend from one budget, so `120` cannot become 360.
178
+ - **Abstention is typed.** With `context` set, the model is given an explicit way to
179
+ say the material does not support an answer; it arrives as `insufficient_context`,
180
+ not as prose you have to pattern-match.
181
+ - **The server never ghostwrites.** When `ok` is `false`, `content` and `data` are
182
+ both `null`. It will not put a "Sorry, I couldn't…" string where a model's answer
183
+ goes, because callers cannot tell those apart. Enforced by an assertion on every
184
+ response and covered by tests.
185
+ - **Degradation is visible.** `fallback_used` and `model_used` always ride along, so
186
+ an answer served by the backup after the primary died never passes as the intended
187
+ one.
188
+ - **The caller cannot smuggle a model.** There is no free-form model parameter, only
189
+ the capability enum. Routing stays operator-controlled.
190
+ - **Boundaries reject early.** Unknown capability, oversized prompt or `system`, empty
191
+ prompt, and a malformed or oversized `response_schema` all fail before a provider is
192
+ called. A nonsensical `limits:` block fails at startup instead.
193
+
194
+ Two known gaps. The MCP SDK drops unknown arguments before the handler sees them, so
195
+ an unrecognized key is ignored at the protocol layer rather than rejected — direct
196
+ calls into `Orchestrator.ask` do reject it. And a `response_schema` containing a
197
+ pathological `pattern` can burn CPU on the event loop during validation: the schema is
198
+ size-capped but not analyzed, so treat schema authorship as a trusted operation.
199
+
200
+ ## Tests
201
+
202
+ ```bash
203
+ uv run pytest -q
204
+ ```
205
+
206
+ 72 tests, no network — deployments are stubbed with LiteLLM's `mock_response`, and the
207
+ shapes it cannot express (no choices, null content, a truncated or filtered reply) are
208
+ stubbed as raw `ModelResponse` objects. Includes the rate-limit-then-fallback path and
209
+ the cooled-down-group path.
210
+
211
+ Because all of that is stubbed, it proves the orchestrator's logic and nothing about
212
+ your providers. For that:
213
+
214
+ ```bash
215
+ uv run python smoke_live.py
216
+ ```
217
+
218
+ Real calls against your `config.yaml`, roughly four short ones per capability, so it
219
+ costs a little money — run it deliberately, not in CI. It checks the things only a
220
+ live endpoint can answer: whether `response_format` survives the round trip, whether
221
+ the model honours the abstention path instead of inventing, and what the provider
222
+ actually sends as `finish_reason` when it runs out of room. Name capabilities as
223
+ arguments to check only some (`uv run python smoke_live.py fast`).
224
+
225
+ ## Contributing
226
+
227
+ Issues and pull requests are welcome. The bar for a change is a test that fails
228
+ without it — the suite runs offline, so there is no key to obtain and no cost to pay.
229
+ Keep `config.yaml` out of your commits.
230
+
231
+ If you are adding a capability to your own setup, you do not need a pull request: it
232
+ is a `model_list` entry.
233
+
234
+ ## Not included
235
+
236
+ Semantic/embedding routing and RouteLLM-style predictive routing (the caller states
237
+ its capability); Redis-backed distributed cooldown state (single process — LiteLLM
238
+ enables it via config when you need a second node); streaming (MCP tool results return
239
+ whole); `sampling/createMessage` loops; PII redaction and telemetry callbacks
240
+ (available as LiteLLM callbacks when a requirement names one).
@@ -0,0 +1,218 @@
1
+ # orchestrator-mcp
2
+
3
+ An MCP server that routes each request to the model configured for that kind of work —
4
+ research to one model, coding to another — and returns every answer through a fixed,
5
+ validated envelope.
6
+
7
+ Pointing a capability at your own deployment is a YAML edit. There is no code to change.
8
+
9
+ > Published to PyPI as **`orchestrator-mcp-server`** — the shorter name is an empty
10
+ > registered project owned by someone else. The import package is `orchestrator_mcp`.
11
+
12
+ ## How it works
13
+
14
+ A **capability** is a LiteLLM `model_name` alias group. Several deployments can share
15
+ one name, and `litellm.Router` already load-balances, retries, cools down, and falls
16
+ back across them. So the routing engine is the config file:
17
+
18
+ ```yaml
19
+ model_list:
20
+ - model_name: coding # capability, not a model
21
+ litellm_params:
22
+ model: anthropic/claude-sonnet-4-5
23
+ api_key: os.environ/ANTHROPIC_API_KEY
24
+
25
+ - model_name: coding # same capability, your own box
26
+ litellm_params:
27
+ model: openai/qwen-coder
28
+ api_base: http://vllm.internal:8000/v1
29
+ api_key: os.environ/LOCAL_VLLM_KEY
30
+ ```
31
+
32
+ The calling agent states which capability it wants. There is no intent classifier —
33
+ the caller is already a language model and knows whether it is asking a coding
34
+ question; paying a second model to guess what the first one already knows buys a cost
35
+ increase and a new failure mode.
36
+
37
+ ## Quick start
38
+
39
+ Write a `config.yaml` — start from [`config.example.yaml`](config.example.yaml) — and
40
+ check that it loads:
41
+
42
+ ```bash
43
+ ORCHESTRATOR_CONFIG=config.yaml uvx --from orchestrator-mcp-server python -c "from orchestrator_mcp.server import build_server; build_server(); print('config ok')"
44
+ ```
45
+
46
+ A bad config fails here rather than at request time: every deployment must route to a
47
+ declared capability, every capability must have a deployment behind it, and every
48
+ fallback must name a real capability.
49
+
50
+ Because the file is LiteLLM's own config schema, `litellm --config config.yaml` runs on
51
+ it unchanged. Keep it out of version control — it holds your endpoints.
52
+
53
+ ### Claude Code
54
+
55
+ ```bash
56
+ claude mcp add orchestrator --env ORCHESTRATOR_CONFIG=$PWD/config.yaml -- uvx orchestrator-mcp-server
57
+ ```
58
+
59
+ ### Codex
60
+
61
+ In `~/.codex/config.toml`:
62
+
63
+ ```toml
64
+ [mcp_servers.orchestrator]
65
+ command = "uvx"
66
+ args = ["orchestrator-mcp-server"]
67
+ env = { ORCHESTRATOR_CONFIG = "/absolute/path/to/config.yaml" }
68
+ ```
69
+
70
+ Both speak stdio, and every tool result is returned as structured content *and* as
71
+ JSON text, so a client that reads only one of the two still gets the whole envelope.
72
+
73
+ **Provider keys.** `config.yaml` references them as `os.environ/NAME`, and they are
74
+ read from the environment the *server* process gets — which is the client's
75
+ environment, not your shell's. If a capability comes back `auth_failed` while the same
76
+ config works from a terminal, add the key to the client's `env` block.
77
+
78
+ ### From a checkout
79
+
80
+ ```bash
81
+ uv sync && uv run pytest -q
82
+ ```
83
+
84
+ ```bash
85
+ claude mcp add orchestrator --env ORCHESTRATOR_CONFIG=$PWD/config.yaml -- uv run --directory $PWD orchestrator-mcp-server
86
+ ```
87
+
88
+ ## Tools
89
+
90
+ ### `ask`
91
+
92
+ | Argument | Notes |
93
+ |---|---|
94
+ | `capability` | Enum, built from your config. Bad values are rejected by the protocol layer. |
95
+ | `prompt` | Required. Capped by `limits.max_prompt_chars`. |
96
+ | `context` | Source material. When set, the model is told to answer only from it and to abstain otherwise. |
97
+ | `system` | Extra instructions. Applied *before* the server's own directives, so it cannot disable them. Capped by `limits.max_system_chars`. |
98
+ | `response_schema` | JSON Schema (`"type": "object"`). Switches on structured mode. Capped by `limits.max_schema_chars` — it is inlined into the prompt verbatim. |
99
+ | `temperature` | Pinned to `0` whenever `response_schema` is set. |
100
+ | `max_output_tokens` | Capped by `limits.max_output_tokens`. |
101
+
102
+ Every call returns the same envelope:
103
+
104
+ ```json
105
+ {
106
+ "ok": true,
107
+ "content": "…",
108
+ "data": null,
109
+ "insufficient_context": false,
110
+ "capability_requested": "coding",
111
+ "model_used": "anthropic/claude-sonnet-4-5",
112
+ "fallback_used": false,
113
+ "finish_reason": "stop",
114
+ "usage": { "prompt_tokens": 10, "completion_tokens": 20, "total_tokens": 30, "cost_usd": 0.0002 },
115
+ "latency_ms": 412,
116
+ "error": null
117
+ }
118
+ ```
119
+
120
+ `content` holds the prose answer and is `null` in structured mode and on failure;
121
+ `data` holds the validated object and is set only in structured mode; `error` is
122
+ `{ code, message }` whenever `ok` is `false`.
123
+
124
+ `error.code` comes from a closed set — `invalid_request`, `no_deployment`,
125
+ `upstream_error`, `rate_limited`, `context_exceeded`, `schema_validation_failed`,
126
+ `timeout`, `content_filtered`, `auth_failed`, `output_truncated` — so callers branch
127
+ on a value instead of matching substrings. `error.message` is bounded at 500
128
+ characters and never quotes the rejected output back at you.
129
+
130
+ ### `list_capabilities`
131
+
132
+ What each capability is for, the deployments behind it, and where it falls back.
133
+
134
+ ## Guardrails, and their limits
135
+
136
+ This server sees a prompt and a completion. It has no ground truth, so it **cannot**
137
+ verify factual claims, and nothing here should be read as a hallucination detector.
138
+ What it does enforce:
139
+
140
+ - **Shape is validated, not assumed.** Structured replies are checked against your
141
+ schema locally with `jsonschema`, regardless of whether the provider claims to
142
+ enforce `response_format`. A violation is a failure, not a payload.
143
+ - **Bounded repair.** An invalid structured reply gets `limits.schema_repair_attempts`
144
+ retries carrying the validator's complaint, then fails as
145
+ `schema_validation_failed`. Never a best-effort half-parsed object.
146
+ - **An unfinished answer is a failure, not a short answer.** A completion cut off by
147
+ the token limit comes back as `output_truncated` with `content: null`, and one the
148
+ provider filtered as `content_filtered`. Neither is returned as prose, because a
149
+ half answer reads exactly like a whole one.
150
+ - **The error tells you what broke, not what the model wrote.** `error.message` gives
151
+ the failing path and constraint (`schema violation at answer/city: failed the
152
+ 'maxLength' constraint`) and is capped at 500 characters. The rejected value itself
153
+ goes only back to the model that produced it, in the repair turn.
154
+ - **`request_timeout_s` bounds the call.** Retries, cross-capability fallback, and
155
+ repair turns all spend from one budget, so `120` cannot become 360.
156
+ - **Abstention is typed.** With `context` set, the model is given an explicit way to
157
+ say the material does not support an answer; it arrives as `insufficient_context`,
158
+ not as prose you have to pattern-match.
159
+ - **The server never ghostwrites.** When `ok` is `false`, `content` and `data` are
160
+ both `null`. It will not put a "Sorry, I couldn't…" string where a model's answer
161
+ goes, because callers cannot tell those apart. Enforced by an assertion on every
162
+ response and covered by tests.
163
+ - **Degradation is visible.** `fallback_used` and `model_used` always ride along, so
164
+ an answer served by the backup after the primary died never passes as the intended
165
+ one.
166
+ - **The caller cannot smuggle a model.** There is no free-form model parameter, only
167
+ the capability enum. Routing stays operator-controlled.
168
+ - **Boundaries reject early.** Unknown capability, oversized prompt or `system`, empty
169
+ prompt, and a malformed or oversized `response_schema` all fail before a provider is
170
+ called. A nonsensical `limits:` block fails at startup instead.
171
+
172
+ Two known gaps. The MCP SDK drops unknown arguments before the handler sees them, so
173
+ an unrecognized key is ignored at the protocol layer rather than rejected — direct
174
+ calls into `Orchestrator.ask` do reject it. And a `response_schema` containing a
175
+ pathological `pattern` can burn CPU on the event loop during validation: the schema is
176
+ size-capped but not analyzed, so treat schema authorship as a trusted operation.
177
+
178
+ ## Tests
179
+
180
+ ```bash
181
+ uv run pytest -q
182
+ ```
183
+
184
+ 72 tests, no network — deployments are stubbed with LiteLLM's `mock_response`, and the
185
+ shapes it cannot express (no choices, null content, a truncated or filtered reply) are
186
+ stubbed as raw `ModelResponse` objects. Includes the rate-limit-then-fallback path and
187
+ the cooled-down-group path.
188
+
189
+ Because all of that is stubbed, it proves the orchestrator's logic and nothing about
190
+ your providers. For that:
191
+
192
+ ```bash
193
+ uv run python smoke_live.py
194
+ ```
195
+
196
+ Real calls against your `config.yaml`, roughly four short ones per capability, so it
197
+ costs a little money — run it deliberately, not in CI. It checks the things only a
198
+ live endpoint can answer: whether `response_format` survives the round trip, whether
199
+ the model honours the abstention path instead of inventing, and what the provider
200
+ actually sends as `finish_reason` when it runs out of room. Name capabilities as
201
+ arguments to check only some (`uv run python smoke_live.py fast`).
202
+
203
+ ## Contributing
204
+
205
+ Issues and pull requests are welcome. The bar for a change is a test that fails
206
+ without it — the suite runs offline, so there is no key to obtain and no cost to pay.
207
+ Keep `config.yaml` out of your commits.
208
+
209
+ If you are adding a capability to your own setup, you do not need a pull request: it
210
+ is a `model_list` entry.
211
+
212
+ ## Not included
213
+
214
+ Semantic/embedding routing and RouteLLM-style predictive routing (the caller states
215
+ its capability); Redis-backed distributed cooldown state (single process — LiteLLM
216
+ enables it via config when you need a second node); streaming (MCP tool results return
217
+ whole); `sampling/createMessage` loops; PII redaction and telemetry callbacks
218
+ (available as LiteLLM callbacks when a requirement names one).
@@ -0,0 +1,56 @@
1
+ # Orchestrator MCP config.
2
+ #
3
+ # A capability is just a LiteLLM `model_name` alias group. Several deployments can
4
+ # share one name; the Router load-balances, retries, and cools down across them.
5
+ # Pointing a capability at your own deployment is an entry in `model_list` -- there
6
+ # is no code to change.
7
+ #
8
+ # This file is LiteLLM's own config schema, so `litellm --config config.yaml` also
9
+ # runs on it unchanged.
10
+
11
+ capabilities:
12
+ coding: "Writing, refactoring, reviewing, and debugging code."
13
+ research: "Open-ended research, synthesis, and long-context reading."
14
+ fast: "Cheap, low-latency answers. Classification, extraction, short replies."
15
+
16
+ model_list:
17
+ # --- coding -------------------------------------------------------------
18
+ - model_name: coding
19
+ litellm_params:
20
+ model: anthropic/claude-sonnet-4-5
21
+ api_key: os.environ/ANTHROPIC_API_KEY
22
+
23
+ # Bring your own: same capability, your deployment. Delete or edit freely.
24
+ # - model_name: coding
25
+ # litellm_params:
26
+ # model: openai/qwen-coder
27
+ # api_base: http://vllm.internal:8000/v1
28
+ # api_key: os.environ/LOCAL_VLLM_KEY
29
+
30
+ # --- research -----------------------------------------------------------
31
+ - model_name: research
32
+ litellm_params:
33
+ model: openai/gpt-4o
34
+ api_key: os.environ/OPENAI_API_KEY
35
+
36
+ # --- fast ---------------------------------------------------------------
37
+ - model_name: fast
38
+ litellm_params:
39
+ model: anthropic/claude-haiku-4-5-20251001
40
+ api_key: os.environ/ANTHROPIC_API_KEY
41
+
42
+ router_settings:
43
+ num_retries: 2
44
+ cooldown_time: 60
45
+ fallbacks:
46
+ - coding: [research]
47
+ - research: [coding]
48
+
49
+ limits:
50
+ max_prompt_chars: 100000
51
+ max_context_chars: 400000
52
+ max_system_chars: 10000 # caller instructions reach the prompt verbatim
53
+ max_schema_chars: 20000 # so does `response_schema`
54
+ max_output_tokens: 4096
55
+ request_timeout_s: 120
56
+ schema_repair_attempts: 1