llm-proxy-cli 0.5.2__tar.gz → 0.5.4__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,117 +1,117 @@
1
- Metadata-Version: 2.4
2
- Name: llm-proxy-cli
3
- Version: 0.5.2
4
- Summary: A lightweight CLI tool for delegating LLM tasks to expert models across multiple providers.
5
- Author-email: Kerem Barbaros Karnabat <kbarbaros@hotmail.com>
6
- Classifier: Programming Language :: Python :: 3
7
- Classifier: License :: OSI Approved :: MIT License
8
- Classifier: Operating System :: OS Independent
9
- Requires-Python: >=3.8
10
- Description-Content-Type: text/markdown
11
- License-File: LICENSE
12
- Requires-Dist: openai>=1.0.0
13
- Requires-Dist: filelock>=3.12.0
14
- Requires-Dist: anthropic>=0.30.0
15
- Dynamic: license-file
16
-
17
- # LLM Proxy CLI
18
-
19
- ![CI Status](https://img.shields.io/github/actions/workflow/status/cadakerem/llm-proxy-cli/ci.yml?branch=master&label=CI&logo=github)
20
- ![PyPI Version](https://img.shields.io/pypi/v/llm-proxy-cli?color=blue&logo=pypi)
21
- ![License](https://img.shields.io/github/license/cadakerem/llm-proxy-cli)
22
-
23
- > A lightweight, fault-tolerant CLI tool for delegating LLM tasks to expert models across multiple providers (Nvidia NIM, Groq, OpenAI, Anthropic Claude, Gemini).
24
-
25
-
26
- ## ⚡ Features
27
- - **Multi-Provider Support**: Seamlessly route requests to `nvidia`, `groq`, `openai`, `anthropic`, or `gemini`.
28
- - **Dynamic Model Discovery**: Never hardcode a model name again. Use `auto-smart` or `auto-fast` and the router will auto-select the best model.
29
- - **Active Liveness Verification**: Pings candidates with a minimal chat request to drop fake/gated models before they crash your task.
30
- - **Automatic Fallbacks**: Provide a comma-separated list of models. If one fails, it instantly falls back to the next.
31
- - **Circuit Breaker**: Built-in health tracking and cooldowns to prevent spamming dead endpoints.
32
- - **Reasoning Extraction**: Automatically extracts and formats hidden `<thought>` or `reasoning` blocks (e.g., from Nemotron).
33
- - **Streaming Native**: Built on the official OpenAI SDK for fast and reliable streaming chunks.
34
-
35
- ## 🏗️ Architecture & Under the Hood
36
- - **Language**: Python 3
37
- - **Libraries**: openai, filelock, anthropic
38
- - **Design Pattern**: Circuit Breaker, Chain of Responsibility (Fallback Routing), and Dynamic Caching.
39
-
40
- The router uses a `FileLock`-backed JSON state (`circuit_breaker.json`) to track failures across concurrent runs.
41
- If an endpoint times out or returns a 5xx error more than `MAX_FAILURES` times, the circuit trips and forces the router to skip that endpoint for the next 120 seconds, immediately trying the next fallback model.
42
-
43
-
44
- ## 📦 Installation
45
-
46
- ```bash
47
- # Install via pip
48
- pip install llm-proxy-cli
49
-
50
- # Or for local development:
51
- # git clone https://github.com/cadakerem/llm-proxy-cli.git
52
- # cd llm-proxy-cli
53
- # pip install -e .
54
- ```
55
-
56
- ## 🔑 Configuration & API Keys
57
-
58
- The router looks for API keys in your environment variables or in ~/.config/llm-proxy-cli/keys.json.
59
-
60
- Supported environment variables:
61
- - NVIDIA_API_KEY
62
- - GROQ_API_KEY
63
- - OPENAI_API_KEY
64
- - ANTHROPIC_API_KEY
65
- - GEMINI_API_KEY
66
-
67
- ## 💻 Usage
68
-
69
- The tool is designed to **automatically** discover and use the best model without you having to memorize model names (using `auto-smart` and `auto-fast`).
70
-
71
- ### 1. Automatic Model Selection (Recommended)
72
- Instead of guessing which model is currently the best or active on the API, simply use `auto-smart` (for complex coding/reasoning tasks) or `auto-fast` (for quick tasks).
73
-
74
- ```bash
75
- # Auto-select the smartest model on Nvidia (e.g., Nemotron or Llama 3.1 405B)
76
- llm-proxy-cli -m "nvidia:auto-smart" -p "Write a React button."
77
-
78
- # Auto-select the fastest model on Groq
79
- llm-proxy-cli -m "groq:auto-fast" -p "Summarize this text."
80
- ```
81
-
82
- ### 2. Chained Automatic Fallback
83
- If Nvidia goes down or hits a rate limit, you can instantly fall back to Groq's best model by separating them with a comma:
84
-
85
- ```bash
86
- llm-proxy-cli -m "nvidia:auto-smart,groq:auto-smart" -p "Refactor this python script."
87
- ```
88
-
89
- ### 3. Specific / Manual Model Selection
90
- If you have a specific model you want to use, you can still hardcode it directly:
91
-
92
- ```bash
93
- # Use Laguna, and fallback to a specific Groq model if it fails
94
- llm-proxy-cli -m "nvidia:poolside/laguna-xs-2.1,groq:groq/compound" -p "Explain quantum entanglement."
95
- ```
96
-
97
- ### ⚠️ Troubleshooting & Known Quirks: Nvidia EULA (404 Not Found)
98
- Nvidia NIM requires users to manually accept the **End User License Agreement (EULA)** for certain models on their website before using them via API. If you haven't accepted the EULA for a dynamically discovered model, Nvidia returns a cryptic `404 Not Found` error.
99
- The Smart Router intercepts this behavior automatically and will print a clear warning.
100
-
101
- **How to Fix:**
102
- 1. Log into the [Nvidia Build Portal](https://build.nvidia.com).
103
- 2. Search for the exact model name shown in the warning and click to run a quick test prompt to accept the terms.
104
- 3. Or bypass auto-discovery completely by explicitly hardcoding a model:
105
- ```bash
106
- llm-proxy-cli -m "nvidia:meta/llama-3.2-11b-vision-instruct" -p "Hello"
107
- ```
108
-
109
- ## 🧑‍💻 Developer & Contributions
110
- Developed by Kerem Barbaros Karnabat ([@cadakerem](https://github.com/cadakerem)).
111
-
112
- > **Note on Repository Structure:** The core routing logic, dynamic model discovery, and circuit breaker patterns are entirely contained within `llm_proxy_cli.py` to ensure maximum portability. Unit tests are located in the `tests/` directory, and `SKILL.md` provides instructions for integrating this tool as a native AI agent skill.
113
-
114
- Contributions, issues, and feature requests are welcome! Feel free to check the [Issues page](../../issues).
115
-
116
- ## 📜 License
117
- This project is licensed under the [MIT License](LICENSE).
1
+ Metadata-Version: 2.4
2
+ Name: llm-proxy-cli
3
+ Version: 0.5.4
4
+ Summary: A lightweight CLI tool for delegating LLM tasks to expert models across multiple providers.
5
+ Author-email: Kerem Barbaros Karnabat <kbarbaros@hotmail.com>
6
+ Classifier: Programming Language :: Python :: 3
7
+ Classifier: License :: OSI Approved :: MIT License
8
+ Classifier: Operating System :: OS Independent
9
+ Requires-Python: >=3.8
10
+ Description-Content-Type: text/markdown
11
+ License-File: LICENSE
12
+ Requires-Dist: openai>=1.0.0
13
+ Requires-Dist: filelock>=3.12.0
14
+ Requires-Dist: anthropic>=0.30.0
15
+ Dynamic: license-file
16
+
17
+ # LLM Proxy CLI
18
+
19
+ ![CI Status](https://img.shields.io/github/actions/workflow/status/cadakerem/llm-proxy-cli/ci.yml?branch=master&label=CI&logo=github)
20
+ ![PyPI Version](https://img.shields.io/pypi/v/llm-proxy-cli?color=blue&logo=pypi)
21
+ ![License](https://img.shields.io/github/license/cadakerem/llm-proxy-cli)
22
+
23
+ > A lightweight, fault-tolerant CLI tool for delegating LLM tasks to expert models across multiple providers (Nvidia NIM, Groq, OpenAI, Anthropic Claude, Gemini).
24
+
25
+
26
+ ## ⚡ Features
27
+ - **Multi-Provider Support**: Seamlessly route requests to `nvidia`, `groq`, `openai`, `anthropic`, or `gemini`.
28
+ - **Dynamic Model Discovery**: Never hardcode a model name again. Use `auto-smart` or `auto-fast` and the router will auto-select the best model.
29
+ - **Active Liveness Verification**: Pings candidates with a minimal chat request to drop fake/gated models before they crash your task.
30
+ - **Automatic Fallbacks**: Provide a comma-separated list of models. If one fails, it instantly falls back to the next.
31
+ - **Circuit Breaker**: Built-in health tracking and cooldowns to prevent spamming dead endpoints.
32
+ - **Reasoning Extraction**: Automatically extracts `reasoning_content` from natively supported models (e.g., DeepSeek-R1 or Nemotron) and outputs them to stderr.
33
+ - **Streaming Native**: Built on the official OpenAI SDK for fast and reliable streaming chunks.
34
+
35
+ ## 🏗️ Architecture & Under the Hood
36
+ - **Language**: Python 3
37
+ - **Libraries**: openai, filelock, anthropic
38
+ - **Design Pattern**: Circuit Breaker, Chain of Responsibility (Fallback Routing), and Dynamic Caching.
39
+
40
+ The router uses a `FileLock`-backed JSON state (`circuit_breaker.json`) to track failures across concurrent runs.
41
+ If an endpoint times out or returns a 5xx error more than `MAX_FAILURES` times, the circuit trips and forces the router to skip that endpoint for the next 120 seconds, immediately trying the next fallback model.
42
+
43
+
44
+ ## 📦 Installation
45
+
46
+ ```bash
47
+ # Install via pip
48
+ pip install llm-proxy-cli
49
+
50
+ # Or for local development:
51
+ # git clone https://github.com/cadakerem/llm-proxy-cli.git
52
+ # cd llm-proxy-cli
53
+ # pip install -e .
54
+ ```
55
+
56
+ ## 🔑 Configuration & API Keys
57
+
58
+ The router looks for API keys in your environment variables or in ~/.config/llm-proxy-cli/keys.json.
59
+
60
+ Supported environment variables:
61
+ - NVIDIA_API_KEY
62
+ - GROQ_API_KEY
63
+ - OPENAI_API_KEY
64
+ - ANTHROPIC_API_KEY
65
+ - GEMINI_API_KEY
66
+
67
+ ## 💻 Usage
68
+
69
+ The tool is designed to **automatically** discover and use the best model without you having to memorize model names (using `auto-smart` and `auto-fast`).
70
+
71
+ ### 1. Automatic Model Selection (Recommended)
72
+ Instead of guessing which model is currently the best or active on the API, simply use `auto-smart` (for complex coding/reasoning tasks) or `auto-fast` (for quick tasks).
73
+
74
+ ```bash
75
+ # Auto-select the smartest model on Nvidia (e.g., Nemotron or Llama 3.1 405B)
76
+ llm-proxy-cli -m "nvidia:auto-smart" -p "Write a React button."
77
+
78
+ # Auto-select the fastest model on Groq
79
+ llm-proxy-cli -m "groq:auto-fast" -p "Summarize this text."
80
+ ```
81
+
82
+ ### 2. Chained Automatic Fallback
83
+ If Nvidia goes down or hits a rate limit, you can instantly fall back to Groq's best model by separating them with a comma:
84
+
85
+ ```bash
86
+ llm-proxy-cli -m "nvidia:auto-smart,groq:auto-smart" -p "Refactor this python script."
87
+ ```
88
+
89
+ ### 3. Specific / Manual Model Selection
90
+ If you have a specific model you want to use, you can still hardcode it directly:
91
+
92
+ ```bash
93
+ # Use Laguna, and fallback to a specific Groq model if it fails
94
+ llm-proxy-cli -m "nvidia:poolside/laguna-xs-2.1,groq:groq/compound" -p "Explain quantum entanglement."
95
+ ```
96
+
97
+ ### ⚠️ Troubleshooting & Known Quirks: Nvidia EULA (404 Not Found)
98
+ Nvidia NIM requires users to manually accept the **End User License Agreement (EULA)** for certain models on their website before using them via API. If you haven't accepted the EULA for a dynamically discovered model, Nvidia returns a cryptic `404 Not Found` error.
99
+ The Smart Router intercepts this behavior automatically and will print a clear warning.
100
+
101
+ **How to Fix:**
102
+ 1. Log into the [Nvidia Build Portal](https://build.nvidia.com).
103
+ 2. Search for the exact model name shown in the warning and click to run a quick test prompt to accept the terms.
104
+ 3. Or bypass auto-discovery completely by explicitly hardcoding a model:
105
+ ```bash
106
+ llm-proxy-cli -m "nvidia:meta/llama-3.2-11b-vision-instruct" -p "Hello"
107
+ ```
108
+
109
+ ## 🧑‍💻 Developer & Contributions
110
+ Developed by Kerem Barbaros Karnabat ([@cadakerem](https://github.com/cadakerem)).
111
+
112
+ > **Note on Repository Structure:** The core routing logic, dynamic model discovery, and circuit breaker patterns are entirely contained within `llm_proxy_cli.py` to ensure maximum portability. Unit tests are located in the `tests/` directory, and `SKILL.md` provides instructions for integrating this tool as a native AI agent skill.
113
+
114
+ Contributions, issues, and feature requests are welcome! Feel free to check the [Issues page](../../issues).
115
+
116
+ ## 📜 License
117
+ This project is licensed under the [MIT License](LICENSE).
@@ -1,101 +1,101 @@
1
- # LLM Proxy CLI
2
-
3
- ![CI Status](https://img.shields.io/github/actions/workflow/status/cadakerem/llm-proxy-cli/ci.yml?branch=master&label=CI&logo=github)
4
- ![PyPI Version](https://img.shields.io/pypi/v/llm-proxy-cli?color=blue&logo=pypi)
5
- ![License](https://img.shields.io/github/license/cadakerem/llm-proxy-cli)
6
-
7
- > A lightweight, fault-tolerant CLI tool for delegating LLM tasks to expert models across multiple providers (Nvidia NIM, Groq, OpenAI, Anthropic Claude, Gemini).
8
-
9
-
10
- ## ⚡ Features
11
- - **Multi-Provider Support**: Seamlessly route requests to `nvidia`, `groq`, `openai`, `anthropic`, or `gemini`.
12
- - **Dynamic Model Discovery**: Never hardcode a model name again. Use `auto-smart` or `auto-fast` and the router will auto-select the best model.
13
- - **Active Liveness Verification**: Pings candidates with a minimal chat request to drop fake/gated models before they crash your task.
14
- - **Automatic Fallbacks**: Provide a comma-separated list of models. If one fails, it instantly falls back to the next.
15
- - **Circuit Breaker**: Built-in health tracking and cooldowns to prevent spamming dead endpoints.
16
- - **Reasoning Extraction**: Automatically extracts and formats hidden `<thought>` or `reasoning` blocks (e.g., from Nemotron).
17
- - **Streaming Native**: Built on the official OpenAI SDK for fast and reliable streaming chunks.
18
-
19
- ## 🏗️ Architecture & Under the Hood
20
- - **Language**: Python 3
21
- - **Libraries**: openai, filelock, anthropic
22
- - **Design Pattern**: Circuit Breaker, Chain of Responsibility (Fallback Routing), and Dynamic Caching.
23
-
24
- The router uses a `FileLock`-backed JSON state (`circuit_breaker.json`) to track failures across concurrent runs.
25
- If an endpoint times out or returns a 5xx error more than `MAX_FAILURES` times, the circuit trips and forces the router to skip that endpoint for the next 120 seconds, immediately trying the next fallback model.
26
-
27
-
28
- ## 📦 Installation
29
-
30
- ```bash
31
- # Install via pip
32
- pip install llm-proxy-cli
33
-
34
- # Or for local development:
35
- # git clone https://github.com/cadakerem/llm-proxy-cli.git
36
- # cd llm-proxy-cli
37
- # pip install -e .
38
- ```
39
-
40
- ## 🔑 Configuration & API Keys
41
-
42
- The router looks for API keys in your environment variables or in ~/.config/llm-proxy-cli/keys.json.
43
-
44
- Supported environment variables:
45
- - NVIDIA_API_KEY
46
- - GROQ_API_KEY
47
- - OPENAI_API_KEY
48
- - ANTHROPIC_API_KEY
49
- - GEMINI_API_KEY
50
-
51
- ## 💻 Usage
52
-
53
- The tool is designed to **automatically** discover and use the best model without you having to memorize model names (using `auto-smart` and `auto-fast`).
54
-
55
- ### 1. Automatic Model Selection (Recommended)
56
- Instead of guessing which model is currently the best or active on the API, simply use `auto-smart` (for complex coding/reasoning tasks) or `auto-fast` (for quick tasks).
57
-
58
- ```bash
59
- # Auto-select the smartest model on Nvidia (e.g., Nemotron or Llama 3.1 405B)
60
- llm-proxy-cli -m "nvidia:auto-smart" -p "Write a React button."
61
-
62
- # Auto-select the fastest model on Groq
63
- llm-proxy-cli -m "groq:auto-fast" -p "Summarize this text."
64
- ```
65
-
66
- ### 2. Chained Automatic Fallback
67
- If Nvidia goes down or hits a rate limit, you can instantly fall back to Groq's best model by separating them with a comma:
68
-
69
- ```bash
70
- llm-proxy-cli -m "nvidia:auto-smart,groq:auto-smart" -p "Refactor this python script."
71
- ```
72
-
73
- ### 3. Specific / Manual Model Selection
74
- If you have a specific model you want to use, you can still hardcode it directly:
75
-
76
- ```bash
77
- # Use Laguna, and fallback to a specific Groq model if it fails
78
- llm-proxy-cli -m "nvidia:poolside/laguna-xs-2.1,groq:groq/compound" -p "Explain quantum entanglement."
79
- ```
80
-
81
- ### ⚠️ Troubleshooting & Known Quirks: Nvidia EULA (404 Not Found)
82
- Nvidia NIM requires users to manually accept the **End User License Agreement (EULA)** for certain models on their website before using them via API. If you haven't accepted the EULA for a dynamically discovered model, Nvidia returns a cryptic `404 Not Found` error.
83
- The Smart Router intercepts this behavior automatically and will print a clear warning.
84
-
85
- **How to Fix:**
86
- 1. Log into the [Nvidia Build Portal](https://build.nvidia.com).
87
- 2. Search for the exact model name shown in the warning and click to run a quick test prompt to accept the terms.
88
- 3. Or bypass auto-discovery completely by explicitly hardcoding a model:
89
- ```bash
90
- llm-proxy-cli -m "nvidia:meta/llama-3.2-11b-vision-instruct" -p "Hello"
91
- ```
92
-
93
- ## 🧑‍💻 Developer & Contributions
94
- Developed by Kerem Barbaros Karnabat ([@cadakerem](https://github.com/cadakerem)).
95
-
96
- > **Note on Repository Structure:** The core routing logic, dynamic model discovery, and circuit breaker patterns are entirely contained within `llm_proxy_cli.py` to ensure maximum portability. Unit tests are located in the `tests/` directory, and `SKILL.md` provides instructions for integrating this tool as a native AI agent skill.
97
-
98
- Contributions, issues, and feature requests are welcome! Feel free to check the [Issues page](../../issues).
99
-
100
- ## 📜 License
101
- This project is licensed under the [MIT License](LICENSE).
1
+ # LLM Proxy CLI
2
+
3
+ ![CI Status](https://img.shields.io/github/actions/workflow/status/cadakerem/llm-proxy-cli/ci.yml?branch=master&label=CI&logo=github)
4
+ ![PyPI Version](https://img.shields.io/pypi/v/llm-proxy-cli?color=blue&logo=pypi)
5
+ ![License](https://img.shields.io/github/license/cadakerem/llm-proxy-cli)
6
+
7
+ > A lightweight, fault-tolerant CLI tool for delegating LLM tasks to expert models across multiple providers (Nvidia NIM, Groq, OpenAI, Anthropic Claude, Gemini).
8
+
9
+
10
+ ## ⚡ Features
11
+ - **Multi-Provider Support**: Seamlessly route requests to `nvidia`, `groq`, `openai`, `anthropic`, or `gemini`.
12
+ - **Dynamic Model Discovery**: Never hardcode a model name again. Use `auto-smart` or `auto-fast` and the router will auto-select the best model.
13
+ - **Active Liveness Verification**: Pings candidates with a minimal chat request to drop fake/gated models before they crash your task.
14
+ - **Automatic Fallbacks**: Provide a comma-separated list of models. If one fails, it instantly falls back to the next.
15
+ - **Circuit Breaker**: Built-in health tracking and cooldowns to prevent spamming dead endpoints.
16
+ - **Reasoning Extraction**: Automatically extracts `reasoning_content` from natively supported models (e.g., DeepSeek-R1 or Nemotron) and outputs them to stderr.
17
+ - **Streaming Native**: Built on the official OpenAI SDK for fast and reliable streaming chunks.
18
+
19
+ ## 🏗️ Architecture & Under the Hood
20
+ - **Language**: Python 3
21
+ - **Libraries**: openai, filelock, anthropic
22
+ - **Design Pattern**: Circuit Breaker, Chain of Responsibility (Fallback Routing), and Dynamic Caching.
23
+
24
+ The router uses a `FileLock`-backed JSON state (`circuit_breaker.json`) to track failures across concurrent runs.
25
+ If an endpoint times out or returns a 5xx error more than `MAX_FAILURES` times, the circuit trips and forces the router to skip that endpoint for the next 120 seconds, immediately trying the next fallback model.
26
+
27
+
28
+ ## 📦 Installation
29
+
30
+ ```bash
31
+ # Install via pip
32
+ pip install llm-proxy-cli
33
+
34
+ # Or for local development:
35
+ # git clone https://github.com/cadakerem/llm-proxy-cli.git
36
+ # cd llm-proxy-cli
37
+ # pip install -e .
38
+ ```
39
+
40
+ ## 🔑 Configuration & API Keys
41
+
42
+ The router looks for API keys in your environment variables or in ~/.config/llm-proxy-cli/keys.json.
43
+
44
+ Supported environment variables:
45
+ - NVIDIA_API_KEY
46
+ - GROQ_API_KEY
47
+ - OPENAI_API_KEY
48
+ - ANTHROPIC_API_KEY
49
+ - GEMINI_API_KEY
50
+
51
+ ## 💻 Usage
52
+
53
+ The tool is designed to **automatically** discover and use the best model without you having to memorize model names (using `auto-smart` and `auto-fast`).
54
+
55
+ ### 1. Automatic Model Selection (Recommended)
56
+ Instead of guessing which model is currently the best or active on the API, simply use `auto-smart` (for complex coding/reasoning tasks) or `auto-fast` (for quick tasks).
57
+
58
+ ```bash
59
+ # Auto-select the smartest model on Nvidia (e.g., Nemotron or Llama 3.1 405B)
60
+ llm-proxy-cli -m "nvidia:auto-smart" -p "Write a React button."
61
+
62
+ # Auto-select the fastest model on Groq
63
+ llm-proxy-cli -m "groq:auto-fast" -p "Summarize this text."
64
+ ```
65
+
66
+ ### 2. Chained Automatic Fallback
67
+ If Nvidia goes down or hits a rate limit, you can instantly fall back to Groq's best model by separating them with a comma:
68
+
69
+ ```bash
70
+ llm-proxy-cli -m "nvidia:auto-smart,groq:auto-smart" -p "Refactor this python script."
71
+ ```
72
+
73
+ ### 3. Specific / Manual Model Selection
74
+ If you have a specific model you want to use, you can still hardcode it directly:
75
+
76
+ ```bash
77
+ # Use Laguna, and fallback to a specific Groq model if it fails
78
+ llm-proxy-cli -m "nvidia:poolside/laguna-xs-2.1,groq:groq/compound" -p "Explain quantum entanglement."
79
+ ```
80
+
81
+ ### ⚠️ Troubleshooting & Known Quirks: Nvidia EULA (404 Not Found)
82
+ Nvidia NIM requires users to manually accept the **End User License Agreement (EULA)** for certain models on their website before using them via API. If you haven't accepted the EULA for a dynamically discovered model, Nvidia returns a cryptic `404 Not Found` error.
83
+ The Smart Router intercepts this behavior automatically and will print a clear warning.
84
+
85
+ **How to Fix:**
86
+ 1. Log into the [Nvidia Build Portal](https://build.nvidia.com).
87
+ 2. Search for the exact model name shown in the warning and click to run a quick test prompt to accept the terms.
88
+ 3. Or bypass auto-discovery completely by explicitly hardcoding a model:
89
+ ```bash
90
+ llm-proxy-cli -m "nvidia:meta/llama-3.2-11b-vision-instruct" -p "Hello"
91
+ ```
92
+
93
+ ## 🧑‍💻 Developer & Contributions
94
+ Developed by Kerem Barbaros Karnabat ([@cadakerem](https://github.com/cadakerem)).
95
+
96
+ > **Note on Repository Structure:** The core routing logic, dynamic model discovery, and circuit breaker patterns are entirely contained within `llm_proxy_cli.py` to ensure maximum portability. Unit tests are located in the `tests/` directory, and `SKILL.md` provides instructions for integrating this tool as a native AI agent skill.
97
+
98
+ Contributions, issues, and feature requests are welcome! Feel free to check the [Issues page](../../issues).
99
+
100
+ ## 📜 License
101
+ This project is licensed under the [MIT License](LICENSE).