ur-agent 1.82.0 → 1.82.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -74,7 +74,7 @@ UR-Nexus supports official provider access paths only:
74
74
 
75
75
  - Explicit API providers: OpenAI, Anthropic, Gemini, OpenRouter, and
76
76
  OpenAI-compatible endpoints.
77
- - Local/server providers: Ollama, LM Studio, llama.cpp, and vLLM OpenAI-compatible
77
+ - Local/server providers: Ollama, LM Studio, llama.cpp, vLLM, and Unsloth OpenAI-compatible
78
78
  server mode.
79
79
  - Subscription CLI providers: Codex CLI, Claude Code CLI, Gemini CLI, and
80
80
  Antigravity where officially supported. These dispatch turns through the
@@ -117,6 +117,7 @@ ur config set provider anthropic-api
117
117
  ur config set provider gemini-api
118
118
  ur config set provider openrouter
119
119
  ur config set provider openai-compatible
120
+ ur config set provider unsloth
120
121
  ur provider doctor agy
121
122
  ur config set provider.fallback ollama
122
123
  ur config set model <model>
@@ -130,7 +131,7 @@ recovery command; changing providers remains an explicit user action.
130
131
 
131
132
  Provider values accept canonical IDs and common aliases. Examples:
132
133
  `openai-api`, `anthropic-api`, `gemini-api`, `openrouter`, `ollama`,
133
- `lmstudio`, `LM Studio`, `llama.cpp`, `vllm`, and the subscription CLI
134
+ `lmstudio`, `LM Studio`, `llama.cpp`, `vllm`, `unsloth` (`Unsloth Studio`), and the subscription CLI
134
135
  providers `codex-cli` (`chatgpt`), `claude-code-cli` (`claude`), `gemini-cli`
135
136
  (`gemini`), and `antigravity-cli` (`agy`). Values with spaces should be quoted
136
137
  in shell commands.
@@ -143,6 +144,11 @@ incompatible saved model instead of silently carrying it across providers. The
143
144
  saved provider/model pair controls the runtime backend for the next agent
144
145
  request; Ollama is only used when `ollama` is the selected provider.
145
146
 
147
+ The configured `base_url` is provider-scoped. Setting an address while vLLM is
148
+ active does not replace the saved Ollama, llama.cpp, or Unsloth address;
149
+ returning to any provider restores its own URL. Legacy `provider.baseUrl`
150
+ settings are migrated to the old active provider when the first switch occurs.
151
+
146
152
  In the model step, Up/Down browses, Left/Right changes the focused model's
147
153
  supported effort level, Enter confirms, Ctrl+R refreshes the catalog, and Esc
148
154
  returns to providers. OpenRouter entries show pricing tier, context size,
@@ -150,6 +156,9 @@ tool/reasoning capability, compact names, and the exact ID for the focused
150
156
  entry. Its catalog is fetched fresh whenever opened; a failed refresh never
151
157
  silently displays cached entries. API-provider secret
152
158
  entry stays on one masked row and stores the value through the keychain flow.
159
+ Ollama and llama.cpp capabilities are loaded lazily for the focused model from
160
+ `/api/show` and `/props`, respectively, so the arrow selector reflects the
161
+ actual model rather than a provider-wide guess.
153
162
 
154
163
  The same provider-first picker is mandatory on the first interactive run in a
155
164
  workspace with no model in `.ur/settings.json` or `.ur/settings.local.json`.
@@ -179,8 +188,16 @@ OPENAI_COMPATIBLE_API_KEY=...
179
188
  ANTHROPIC_API_KEY=...
180
189
  GEMINI_API_KEY=...
181
190
  OPENROUTER_API_KEY=...
191
+ UNSLOTH_API_KEY=...
182
192
  ```
183
193
 
194
+ Unsloth is an inference-provider integration only. Start Unsloth Studio and
195
+ load the model outside UR, connect its generated key with `ur connect unsloth`,
196
+ then select a model discovered from `http://localhost:8888/v1` (or your
197
+ configured `base_url`). UR always sends `enable_tools: false`; standard model
198
+ function calls continue through UR's own permission, sandbox, and verifier
199
+ flow, while Unsloth's server-side tools remain disabled.
200
+
184
201
  ### OpenAI Responses transport
185
202
 
186
203
  OpenAI uses Chat Completions unless the Responses transport is selected
@@ -198,10 +198,29 @@ ur provider doctor <provider-id>
198
198
  ```sh
199
199
  curl http://localhost:11434/api/tags # Ollama
200
200
  curl http://localhost:1234/v1/models # LM Studio (llama.cpp: 8080, vLLM: 8000)
201
+ curl -H "Authorization: Bearer $UNSLOTH_API_KEY" http://localhost:8888/v1/models
201
202
  ur connect openai-api # store an API key securely
202
203
  ur provider doctor
203
204
  ```
204
205
 
206
+ ### Unsloth is selected but unavailable
207
+
208
+ - Likely cause: Studio is not running, no model is loaded, its generated API
209
+ key is not connected, or `base_url` does not end at the compatible API root.
210
+ - Fix: start/load Unsloth outside UR, connect its key, and inspect discovery.
211
+
212
+ ```sh
213
+ echo "$UNSLOTH_API_KEY" | ur connect unsloth
214
+ ur config set provider unsloth
215
+ ur config set base_url http://localhost:8888/v1
216
+ ur provider doctor unsloth
217
+ ur provider models unsloth --json
218
+ ```
219
+
220
+ UR does not start or manage Unsloth. It also disables Unsloth server-side tools
221
+ on every request; tool execution shown by UR is handled by UR's own guarded
222
+ tool loop.
223
+
205
224
  ### Local server unreachable
206
225
 
207
226
  - Likely cause: wrong `base_url` or the server is not running.
@@ -212,6 +231,10 @@ ur config set base_url http://localhost:11434
212
231
  ur provider doctor
213
232
  ```
214
233
 
234
+ Addresses are saved per provider. If the doctor probes an unexpected URL,
235
+ select that provider first and run `ur config get base_url`; changing vLLM's
236
+ address no longer overwrites Ollama, llama.cpp, or Unsloth.
237
+
215
238
  ## Sessions and workflows
216
239
 
217
240
  ### No visible progress in scripts
package/docs/USAGE.md CHANGED
@@ -162,6 +162,7 @@ ur config set provider anthropic-api
162
162
  ur config set provider gemini-api
163
163
  ur config set provider openrouter
164
164
  ur config set provider openai-compatible
165
+ ur config set provider unsloth
165
166
  ur config set model <model>
166
167
  ur config set base_url <url>
167
168
  ur config set provider.fallback ollama
@@ -206,17 +207,25 @@ kind, external CLI usage, native tool/streaming support, and the boundary text.
206
207
 
207
208
  Provider values accept canonical IDs and common aliases. For example,
208
209
  `openai-api`, `anthropic-api`, `gemini-api`, `openrouter`, `ollama`,
209
- `lmstudio`, `llama.cpp`, and `vllm` are UR-native runtime providers, and
210
+ `lmstudio`, `llama.cpp`, `vllm`, and `unsloth` are UR-native runtime providers, and
210
211
  `codex-cli` (`chatgpt`), `claude-code-cli` (`claude`), `gemini-cli` (`gemini`),
211
212
  and `antigravity-cli` (`agy`) are subscription CLI providers.
212
213
 
213
214
  API modes are explicit. Keys are read from a key stored via
214
215
  `ur connect <provider>` (OS keychain) or from the environment variables
215
216
  `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `GEMINI_API_KEY`, and
216
- `OPENROUTER_API_KEY`. Subscription CLIs are optional, never required
217
+ `OPENROUTER_API_KEY`, and `UNSLOTH_API_KEY`. Subscription CLIs are optional, never required
217
218
  dependencies, and never used as a silent fallback. UR-Nexus never scrapes
218
219
  browser sessions, extracts OAuth tokens, or bypasses provider restrictions.
219
220
  OpenAI-compatible local or cloud endpoints use `base_url` plus `model`.
221
+ Unsloth defaults to `http://localhost:8888/v1`, requires its Studio API key,
222
+ and is inference-only: UR does not manage Unsloth and disables its server-side
223
+ tools while retaining standard function calls inside UR's guarded tool loop.
224
+
225
+ UR stores `base_url` per provider. You can set different addresses for
226
+ Ollama, llama.cpp, vLLM, and Unsloth once, then switch providers without
227
+ re-entering any of them. `ur config get base_url` always reports the address
228
+ for the currently active provider.
220
229
 
221
230
  Use `/model` in an interactive session to select provider first and model
222
231
  second. OpenAI API, Claude API, Gemini API, OpenRouter, Ollama, and
@@ -19,7 +19,7 @@ You need:
19
19
 
20
20
  ```sh
21
21
  ur --version
22
- # expected for this release: "1.82.0 (UR-Nexus)"
22
+ # expected for this release: "1.82.2 (UR-Nexus)"
23
23
  ```
24
24
 
25
25
  ### 0.0 Redteam mode and Reverse Skills (1.81.0)
package/docs/providers.md CHANGED
@@ -41,6 +41,7 @@ multimodal input, external CLI boundary, and sandbox scope:
41
41
  | LM Studio | local/server | UR-native | no | yes | yes | yes | UR Bash/File sandbox | `openai-compatible:lmstudio` | local OpenAI-compatible server |
42
42
  | llama.cpp | local/server | UR-native | no | yes | yes | yes | UR Bash/File sandbox | `openai-compatible:llama.cpp` | local OpenAI-compatible server |
43
43
  | vLLM | local/server | UR-native | no | yes | yes | yes | UR Bash/File sandbox | `openai-compatible:vllm` | OpenAI-compatible server |
44
+ | Unsloth | local/server | UR-native | no | yes | yes | model-dependent | UR Bash/File sandbox | `openai-compatible:unsloth` | authenticated user-run Unsloth Studio endpoint (`UNSLOTH_API_KEY`) |
44
45
  | Codex CLI | subscription | subscription-cli | yes | no | no | no† | UR-run tools/output only† | `subscription-cli:codex` | official Codex CLI login |
45
46
  | Claude Code | subscription | subscription-cli | yes | no | no | no† | UR-run tools/output only† | `subscription-cli:claude-code` | official Claude Code CLI login |
46
47
  | Gemini CLI | subscription | subscription-cli | yes | no | no | no† | UR-run tools/output only† | `subscription-cli:gemini` | official Gemini Code Assist login |
@@ -122,6 +123,7 @@ ur config set provider anthropic-api
122
123
  ur config set provider gemini-api
123
124
  ur config set provider openrouter
124
125
  ur config set provider openai-compatible
126
+ ur config set provider unsloth
125
127
  ur config set model <model>
126
128
  ur provider select-model <provider> <model> --json
127
129
  ur config set base_url <url>
@@ -136,6 +138,11 @@ The fallback setting is a recovery hint for `ur provider doctor`; it does not
136
138
  route a failed request to another provider. Review the failure and use
137
139
  `ur config set provider <id>` to switch explicitly.
138
140
 
141
+ `base_url` is stored for the provider that is active when the command runs.
142
+ Each provider retains its own address across `/provider`, `/model`, and CLI
143
+ switches. The legacy single `provider.baseUrl` field remains readable and is
144
+ migrated to the previously active provider on the first switch.
145
+
139
146
  OpenAI API uses Chat Completions by default. `openai_transport responses` is an
140
147
  explicit opt-in to the native Responses adapter; it defaults to `store=false`
141
148
  and supports semantic streaming, background polling/cancellation, WebSocket
@@ -165,6 +172,17 @@ indicator, active-work spinner, SDK settings response, and provider request all
165
172
  use the same resolved value. If a provider advertises only boolean thinking,
166
173
  UR does not invent a graded effort selector.
167
174
 
175
+ For Ollama, UR lazily reads the focused model's `/api/show` capabilities and
176
+ sends the selected level through native `think`. Kimi K3 uses
177
+ `low|high|max`; GPT-OSS uses `low|medium|high`; other models advertising
178
+ `thinking` use Ollama's current `low|medium|high|max` contract. Direct OpenAI,
179
+ Anthropic, and Gemini models use curated model-specific ladders from their
180
+ official documentation; live discovery rows are merged with those contracts.
181
+ See [Ollama thinking](https://docs.ollama.com/capabilities/thinking),
182
+ [OpenAI model guidance](https://developers.openai.com/api/docs/guides/latest-model),
183
+ [Claude effort](https://platform.claude.com/docs/en/build-with-claude/effort),
184
+ and [Gemini thinking](https://ai.google.dev/gemini-api/docs/thinking).
185
+
168
186
  The provider-first `/model` picker supports the same control directly: use
169
187
  Left/Right to move through the effort levels advertised by the focused model,
170
188
  then Enter to apply the model and effort together. OpenRouter's live catalog
@@ -309,7 +327,7 @@ ur config set provider anthropic-api
309
327
  | --- | --- | --- |
310
328
  | API providers (openai-api, anthropic-api, gemini-api) | Live discovery from the provider's `/models` endpoint using your connected key (curated fallback until connected) | live |
311
329
  | OpenRouter | Fresh `/models` discovery every time its picker opens; no cached-list fallback in the picker | live |
312
- | Local/server providers (ollama, lmstudio, llama.cpp, vllm) | Dynamic discovery from the selected provider endpoint | live |
330
+ | Local/server providers (ollama, lmstudio, llama.cpp, vllm, unsloth) | Dynamic discovery from the selected provider endpoint | live |
313
331
  | OpenAI-compatible | Dynamic discovery from configured endpoint | live |
314
332
  | Subscription CLIs (codex-cli, claude-code-cli, gemini-cli, antigravity-cli) | Curated list (the official CLIs expose no models API); first-class in `/model`, dispatched via the official CLI. External CLI behavior depends on the vendor CLI. Log in with `ur auth <provider>` | static |
315
333
 
@@ -332,6 +350,7 @@ ur config set provider anthropic-api
332
350
  - `lmstudio` — LM Studio OpenAI-compatible server
333
351
  - `llama.cpp` — llama.cpp server mode
334
352
  - `vllm` — vLLM server
353
+ - `unsloth` — authenticated Unsloth Studio server; UR uses it for inference only
335
354
 
336
355
  **Important:**
337
356
  - A ChatGPT/Claude/Gemini subscription does NOT give API access
@@ -462,6 +481,7 @@ Provider config and doctor commands accept canonical IDs and common aliases:
462
481
  | `lmstudio` | `LM Studio`, `lm-studio` |
463
482
  | `llama.cpp` | `llama cpp`, `llamacpp`, `llama-cpp` |
464
483
  | `vllm` | `vllm server` |
484
+ | `unsloth` | `Unsloth Studio`, `unsloth server`, `unsloth local` |
465
485
 
466
486
  `ur provider doctor` checks the selected provider. It reports installed/missing
467
487
  CLIs, official login status where available, API key presence for API providers,
@@ -496,6 +516,7 @@ OPENAI_COMPATIBLE_API_KEY=...
496
516
  ANTHROPIC_API_KEY=...
497
517
  GEMINI_API_KEY=...
498
518
  OPENROUTER_API_KEY=...
519
+ UNSLOTH_API_KEY=...
499
520
  ```
500
521
 
501
522
  OpenAI-compatible endpoints can point at local or cloud endpoints:
@@ -516,6 +537,32 @@ Local/server providers use their normal endpoints:
516
537
  - LM Studio: `http://localhost:1234/v1`
517
538
  - llama.cpp server mode: `http://localhost:8080/v1`
518
539
  - vLLM server mode: `http://localhost:8000/v1`
540
+ - Unsloth Studio: `http://localhost:8888/v1`
541
+
542
+ ### Unsloth provider-only mode
543
+
544
+ UR connects to a user-run Unsloth Studio inference server through its official
545
+ OpenAI-compatible API. It does not import the Unsloth Python package, launch or
546
+ update Studio, train or convert models, manage GPUs, or load a model. Start and
547
+ load Unsloth separately, then connect the generated Studio key:
548
+
549
+ ```sh
550
+ ur config set provider unsloth
551
+ echo "$UNSLOTH_API_KEY" | ur connect unsloth
552
+ ur provider doctor unsloth
553
+ ur provider models unsloth --json
554
+ # choose one of the discovered model IDs with /model
555
+ ```
556
+
557
+ The default endpoint is `http://localhost:8888/v1`; override it with
558
+ `ur config set base_url <url>`. Authentication is mandatory. Every Unsloth
559
+ request sets `enable_tools: false`, including streaming requests. The model may
560
+ still return standard OpenAI function calls, but execution remains exclusively
561
+ inside UR's provenance gate, permissions, sandbox, and verifier. This prevents
562
+ Unsloth Studio's optional server-side web/code tools from becoming a second,
563
+ uncontrolled agent runtime. See the official
564
+ [Unsloth Studio announcement](https://github.com/unslothai/unsloth/discussions/5285)
565
+ and [Unsloth repository](https://github.com/unslothai/unsloth).
519
566
 
520
567
  Ollama allows up to 15 minutes for response headers so a cold model load or
521
568
  large prefill can begin. After headers, local and `:cloud` models use a
@@ -45,7 +45,7 @@
45
45
  <main id="content" class="content">
46
46
  <header class="topbar">
47
47
  <div>
48
- <p class="eyebrow">Version 1.82.0</p>
48
+ <p class="eyebrow">Version 1.82.2</p>
49
49
  <h1>UR-Nexus Documentation</h1>
50
50
  <p class="lead">A practical, tutorial-style reference for installing, configuring, automating, extending, and operating UR-Nexus.</p>
51
51
  </div>
@@ -7,7 +7,7 @@ plugins {
7
7
  }
8
8
 
9
9
  group = "dev.urnexus"
10
- version = "1.82.0"
10
+ version = "1.82.2"
11
11
 
12
12
  repositories {
13
13
  mavenCentral()
@@ -2,7 +2,7 @@
2
2
  "name": "ur-inline-diffs",
3
3
  "displayName": "UR Inline Diffs",
4
4
  "description": "Review, apply, and reject UR inline diff bundles from .ur/ide/diffs inside VS Code.",
5
- "version": "1.82.0",
5
+ "version": "1.82.2",
6
6
  "publisher": "ur-nexus",
7
7
  "engines": {
8
8
  "vscode": "^1.92.0"
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ur-agent",
3
- "version": "1.82.0",
3
+ "version": "1.82.2",
4
4
  "description": "UR-Nexus — autonomous engineering workflow engine (plan, execute, test, verify, document, benchmark, reproduce)",
5
5
  "type": "module",
6
6
  "packageManager": "bun@1.3.14",