ur-agent 1.83.1 → 1.84.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -314,7 +314,7 @@ ur auth antigravity
314
314
  ur config set provider ollama
315
315
  ur config set provider openai-compatible
316
316
  ur config set model qwen2.5-coder:7b
317
- ur config set base_url http://localhost:11434/v1
317
+ ur config set base_url openai-compatible http://localhost:11434/v1
318
318
  ur config set provider.fallback ollama
319
319
  ur upgrade
320
320
  ```
@@ -164,7 +164,8 @@ request; Ollama is only used when `ollama` is the selected provider.
164
164
  The configured `base_url` is provider-scoped. Setting an address while vLLM is
165
165
  active does not replace the saved Ollama, llama.cpp, or Unsloth address;
166
166
  returning to any provider restores its own URL. Legacy `provider.baseUrl`
167
- settings are migrated to the old active provider when the first switch occurs.
167
+ settings are migrated to the old active provider on the first provider switch
168
+ or scoped base-URL write.
168
169
  To configure a provider that is not active, use
169
170
  `ur config set base_url <provider> <url>`; the success message names the target
170
171
  provider. This applies to direct API providers and gateways (OpenAI, Anthropic,
@@ -185,6 +186,18 @@ Ollama and llama.cpp capabilities are loaded lazily for the focused model from
185
186
  `/api/show` and `/props`, respectively, so the arrow selector reflects the
186
187
  actual model rather than a provider-wide guess.
187
188
 
189
+ The effort row contains only capability-backed selectors UR can map to native
190
+ provider values. Ultra appears only when metadata advertises `ultra`, `max`,
191
+ `xhigh`, or an explicit alias; mappings such as `ultra→max` are shown and sent
192
+ exactly. Models that top out at `high`, boolean-thinking models, and unknown
193
+ capabilities omit Ultra. See [Reasoning effort](providers.md#reasoning-effort).
194
+
195
+ Boolean-thinking models on runtimes with a native toggle expose a two-state
196
+ control instead: Left selects off, Right selects on, and `t` toggles in
197
+ `/model`; `/thinking on|off|status` is the direct command. This updates the live
198
+ session and `alwaysThinkingEnabled`. Generic OpenAI-compatible transports do
199
+ not receive an invented boolean field.
200
+
188
201
  The same provider-first picker is mandatory on the first interactive run in a
189
202
  workspace with no model in `.ur/settings.json` or `.ur/settings.local.json`.
190
203
  The result is validated and written to the gitignored local settings file.
@@ -223,7 +236,7 @@ UNSLOTH_API_KEY=...
223
236
  Unsloth is an inference-provider integration only. Start Unsloth Studio and
224
237
  load the model outside UR, connect its generated key with `ur connect unsloth`,
225
238
  then select a model discovered from `http://localhost:8888/v1` (or your
226
- configured `base_url`). UR always sends `enable_tools: false`; standard model
239
+ configured `base_url`). UR sends `enable_tools: false` on every inference request; standard model
227
240
  function calls continue through UR's normal native tool flow, while Unsloth's
228
241
  server-side tools remain disabled so the integration stays provider-only.
229
242
 
@@ -712,6 +725,7 @@ Plugins can add commands, tools, and skills:
712
725
 
713
726
  ```sh
714
727
  ur plugin list
728
+ ur plugin marketplace add npm:@scope/catalog@latest
715
729
  ur plugin search [query] [--capability <name>] [--marketplace <name>] [--installed] [--json]
716
730
  ur plugin show <name-or-name@marketplace> [--json]
717
731
  ur plugin install <plugin>
@@ -228,8 +228,8 @@ tool loop.
228
228
  - Fix: point UR at the right endpoint.
229
229
 
230
230
  ```sh
231
- ur config set base_url http://localhost:11434
232
- ur provider doctor
231
+ ur config set base_url ollama http://localhost:11434
232
+ ur provider doctor ollama
233
233
  ```
234
234
 
235
235
  Addresses are saved per provider. If the doctor probes an unexpected URL,
package/docs/USAGE.md CHANGED
@@ -214,7 +214,7 @@ and `antigravity-cli` (`agy`) are subscription CLI providers.
214
214
 
215
215
  API modes are explicit. Keys are read from a key stored via
216
216
  `ur connect <provider>` (OS keychain) or from the environment variables
217
- `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `GEMINI_API_KEY`, and
217
+ `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `GEMINI_API_KEY`,
218
218
  `OPENROUTER_API_KEY`, and `UNSLOTH_API_KEY`. Subscription CLIs are optional, never required
219
219
  dependencies, and never used as a silent fallback. UR-Nexus never scrapes
220
220
  browser sessions, extracts OAuth tokens, or bypasses provider restrictions.
@@ -224,9 +224,10 @@ and is inference-only: UR does not manage Unsloth and disables its server-side
224
224
  tools while retaining standard function calls inside UR's guarded tool loop.
225
225
 
226
226
  UR stores `base_url` per provider. You can set different addresses for
227
- Ollama, llama.cpp, vLLM, and Unsloth once, then switch providers without
228
- re-entering any of them. `ur config get base_url` always reports the address
229
- for the currently active provider.
227
+ Ollama, LM Studio, llama.cpp, vLLM, and Unsloth once, then switch providers without
228
+ re-entering any of them. `ur config get base_url` reports the active provider's
229
+ saved scoped override when one exists; use `ur provider status` or
230
+ `ur provider doctor <provider>` to inspect the effective endpoint.
230
231
  Use `ur config set base_url <provider> <url>` to change one provider's address
231
232
  without making it active first. The `/model` picker offers the same endpoint
232
233
  entry flow for a disconnected local/server provider.
@@ -348,6 +349,10 @@ UR includes slash commands and CLI subcommands for common workflows:
348
349
  `ur plugin search [query]` for ranked cross-catalog discovery and
349
350
  `ur plugin show <name@marketplace>` to inspect provenance and capabilities
350
351
  before installation.
352
+ Add an npm-hosted catalog with
353
+ `ur plugin marketplace add npm:@scope/catalog@latest`; npm sources respect
354
+ the user's registry/authentication configuration and are refreshed with the
355
+ normal marketplace update command.
351
356
  - `ur agents` to list configured agents
352
357
  - `ur agent-trends` to inspect coverage for current agent technology trends
353
358
  - `ur a2a card` to print legacy Agent Card metadata, or
@@ -19,7 +19,7 @@ You need:
19
19
 
20
20
  ```sh
21
21
  ur --version
22
- # expected for this release: "1.83.1 (UR-Nexus)"
22
+ # expected for this release: "1.84.0 (UR-Nexus)"
23
23
  ```
24
24
 
25
25
  ### 0.0 Redteam mode and Reverse Skills (1.81.0)
@@ -72,16 +72,20 @@ changes must remain blocked. The deterministic regressions are:
72
72
  bun test test/taskListGate.test.ts test/toolExecutionFinalInput.test.ts
73
73
  ```
74
74
 
75
- ### 0.0.0 OpenRouter research routing and provider UI (1.81.3)
75
+ ### 0.0.0 OpenRouter research routing and provider UI (1.81.3; Ultra mapping updated 1.83.1)
76
76
 
77
77
  Connect OpenRouter, run `/model`, select OpenRouter, and verify that its model
78
78
  step shows catalog freshness plus pricing/context/tool/reasoning details. Focus
79
79
  a reasoning model and press Left/Right; the displayed effort must change and
80
- the row must list only that model's provider-advertised levels. Move Up/Down
81
- between high-only, xhigh, max, and native-ultra models; the level list and
82
- selected ceiling must update immediately. High-only must omit Ultra, while
80
+ the row must list only capability-backed selectors that map to that model's native levels. Move Up/Down
81
+ between models that top out at high, xhigh, max, and native-ultra models; the level list and
82
+ selected ceiling must update immediately. Models that top out at high must omit Ultra, while
83
83
  xhigh/max entries must show `ultra→xhigh` or `ultra→max`, and the
84
84
  confirmation must match `/effort status` and the request wire value. For
85
+ an Ollama model that advertises boolean thinking without a ladder, verify that
86
+ Left selects off, Right selects on, `t` toggles, and `/effort max` reports that
87
+ max was not sent while enabling `think: true`; `/thinking off` must produce
88
+ `think: false`. For
85
89
  llama.cpp, verify focus requests
86
90
  `/props?model=<focused-id>` and that a template reporting
87
91
  `supports_reasoning_effort: false` has no graded selector. Open the OpenAI API or Claude
package/docs/plugins.md CHANGED
@@ -51,6 +51,50 @@ ur --plugin-dir ./plugins/community/my-plugin
51
51
  Plugins are loaded from local UR-Nexus paths first. Network marketplace installs
52
52
  remain explicit user actions and are subject to plugin policy checks.
53
53
 
54
+ ## Marketplace sources
55
+
56
+ UR accepts GitHub shorthand, Git URLs, direct marketplace JSON URLs, npm
57
+ packages, local files/directories, and inline settings manifests. Add an npm
58
+ marketplace with the explicit `npm:` prefix:
59
+
60
+ ```sh
61
+ ur plugin marketplace add npm:acme-ur-marketplace
62
+ ur plugin marketplace add npm:@acme/ur-marketplace@latest
63
+ ur plugin marketplace add npm:@acme/ur-marketplace@^2.0.0
64
+ ur plugin marketplace update <marketplace-name>
65
+ ```
66
+
67
+ The package must ship `.ur-plugin/marketplace.json`. An omitted version follows
68
+ the registry's `latest` dist-tag; a version, semver range, or another dist-tag
69
+ can be supplied after the package name, using npm's
70
+ [package-spec syntax](https://docs.npmjs.com/cli/v11/using-npm/package-spec/).
71
+ Refreshing the marketplace re-resolves
72
+ that selector. UR uses the installed npm client, so standard `.npmrc`
73
+ authentication, scoped registries, proxies, and registry settings continue to
74
+ work. Package lifecycle scripts are disabled during marketplace download, and
75
+ only the requested package—not its staging dependency tree—is retained.
76
+
77
+ For a private registry selected in project or user settings:
78
+
79
+ ```json
80
+ {
81
+ "extraKnownMarketplaces": {
82
+ "acme": {
83
+ "source": {
84
+ "source": "npm",
85
+ "package": "@acme/ur-marketplace",
86
+ "version": "^2.0.0",
87
+ "registry": "https://registry.example.com"
88
+ }
89
+ }
90
+ }
91
+ }
92
+ ```
93
+
94
+ After an install, removal, or external registry-file change,
95
+ `/reload-plugins` clears both plugin discovery caches and the installed-plugin
96
+ snapshot before reloading.
97
+
54
98
  ## Manifest reference
55
99
 
56
100
  A plugin is a directory containing `.ur-plugin/plugin.json`. UR uses a
package/docs/providers.md CHANGED
@@ -21,8 +21,9 @@ UR-Nexus never:
21
21
  - claims provider support unless the official CLI/API path works
22
22
 
23
23
  UR-Nexus stores only safe config: provider name, model name, base URL, fallback
24
- preference, and non-secret preferences. API keys are read from environment
25
- variables only when the user explicitly selects API mode.
24
+ preference, and non-secret preferences. API keys stay out of plaintext settings
25
+ and are read from the OS keychain after `ur connect` or from environment
26
+ variables when the user explicitly selects API mode.
26
27
 
27
28
  ## Provider matrix
28
29
 
@@ -144,7 +145,8 @@ form `ur config set base_url <provider> <url>` configures a named provider
144
145
  without switching first, and the confirmation names that provider. Each
145
146
  provider retains its own address across `/provider`, `/model`, and CLI
146
147
  switches. The legacy single `provider.baseUrl` field remains readable and is
147
- migrated to the previously active provider on the first switch.
148
+ migrated to the previously active provider on the first provider switch or
149
+ scoped base-URL write.
148
150
 
149
151
  The override is not limited to local runtimes. OpenAI API, Anthropic API,
150
152
  Gemini API, and OpenRouter can each target a separate compatible gateway using
@@ -171,29 +173,46 @@ the previous provider.
171
173
 
172
174
  ### Reasoning effort
173
175
 
174
- Use `/effort minimal|low|medium|high|xhigh|max|ultra|auto` inside UR. UR normalizes
175
- the exact graded levels advertised for the active provider/model pair. `max`
176
- is provider-neutral and means the selected model's real ceiling: it therefore
177
- displays and sends `max`, `xhigh`, or `high` according to that model's contract.
176
+ Use `/effort minimal|low|medium|high|xhigh|max|ultra|auto` inside UR. UR builds a
177
+ capability-backed selector set from the native graded levels advertised for the active
178
+ provider/model pair. `max` is provider-neutral and means the selected model's
179
+ highest supported non-Ultra tier: it displays and sends the matching native
180
+ value (commonly `max`, `xhigh`, or `high`) according to that model's contract.
178
181
  For OpenRouter, UR preserves the live `/models` reasoning metadata and sends
179
182
  the unified `reasoning.effort` request. OpenAI-compatible servers receive the
180
183
  resolved value as `reasoning_effort`. The command confirmation, status
181
184
  indicator, active-work spinner, SDK settings response, and provider request all
182
- use the same resolved value. If a provider advertises only boolean thinking,
183
- UR does not invent a graded effort selector.
185
+ use the same resolved value. If a provider advertises only boolean thinking and
186
+ its runtime has a real native on/off mapping, UR does not invent a graded effort
187
+ selector. Use `/thinking on|off` directly;
188
+ in `/model`, Left selects off, Right selects on, and `t` toggles. A graded
189
+ `/effort` request on that model enables boolean thinking while clearly reporting
190
+ that the requested level was not sent.
191
+ Generic OpenAI-compatible endpoints have no universal boolean thinking field,
192
+ so metadata alone does not make this toggle appear and UR sends no invented parameter.
184
193
  `ultra` is UR's visible beyond-high ceiling selector. It is selectable only
185
194
  when the provider/model advertises `ultra`, `max`, `xhigh`, or an explicit
186
195
  provider-authored equivalent. UR shows the native mapping (for example,
187
- `ultra→max`) and sends that exact wire value; it never enables Ultra for a
188
- high-only model, boolean thinking, or unknown capability metadata. Arbitrary
196
+ `ultra→max`) and sends that exact wire value; it never enables Ultra for a model
197
+ whose graded ladder tops out at `high`, boolean thinking, or unknown capability metadata. Arbitrary
189
198
  labels such as `deep` still require an explicit provider alias because UR
190
199
  cannot infer their rank.
191
200
 
201
+ For an unknown or newly released model, UR waits for provider-authored model
202
+ metadata or a supported model-scoped probe before adding thinking parameters.
203
+ If the provider does not establish support, thinking stays off for request
204
+ shaping; UR does not optimistically send an unknown parameter and treat an API
205
+ error as capability discovery. Boolean thinking metadata enables the thinking
206
+ toggle only and never invents a graded effort ladder. On OpenRouter, UR sends
207
+ the provider-default `reasoning.enabled` control, or the exact token budget when
208
+ the model advertises `supports_max_tokens`.
209
+
192
210
  For Ollama, UR lazily reads the focused model's `/api/show` capabilities and
193
- sends the selected level through native `think`. Kimi K3 uses
194
- `low|high|max` and therefore exposes Ultra as `ultra→max`; GPT-OSS uses
195
- `low|medium|high` and does not expose Ultra; other models advertising
196
- `thinking` use Ollama's current `low|medium|high|max` contract. Direct OpenAI,
211
+ sends the resolved control through native `think`. A generic `thinking`
212
+ capability means boolean thinking only. GPT-OSS uses Ollama's documented
213
+ `low|medium|high` ladder and does not expose Ultra. Other graded ladders and
214
+ Ultra aliases are used only when the endpoint explicitly returns them in model
215
+ reasoning metadata. Direct OpenAI,
197
216
  Anthropic, and Gemini models use curated model-specific ladders from their
198
217
  official documentation; live discovery rows are merged with those contracts.
199
218
  See [Ollama thinking](https://docs.ollama.com/capabilities/thinking),
@@ -202,8 +221,10 @@ See [Ollama thinking](https://docs.ollama.com/capabilities/thinking),
202
221
  and [Gemini thinking](https://ai.google.dev/gemini-api/docs/thinking).
203
222
 
204
223
  The provider-first `/model` picker supports the same control directly: use
205
- Left/Right to move through the effort levels advertised by the focused model,
206
- then Enter to apply the model and effort together. OpenRouter's live catalog
224
+ Left/Right to move through the capability-backed selectors UR can map to a
225
+ graded model's native levels, or to choose off/on for a boolean-thinking model
226
+ when its runtime has a native two-state mapping, then Enter to apply the model
227
+ and reasoning control together. OpenRouter's live catalog
207
228
  shows pricing tier, context size, tool capability, reasoning capability, and
208
229
  the full, untruncated model ID immediately below the focused entry. Opening the
209
230
  OpenRouter catalog reuses its endpoint-scoped five-minute cache; Ctrl+R forces
@@ -215,6 +236,18 @@ provider prompt-cache markers. Explicit routing preferences and the `:nitro`,
215
236
  OpenAI, Claude, Gemini, and OpenRouter is a single aligned masked row; the key
216
237
  is stored in the OS keychain flow and is never written to settings.
217
238
 
239
+ ### Token counting
240
+
241
+ UR uses each provider's non-generating count endpoint when one covers the full
242
+ request: OpenAI Responses input tokens, Anthropic Messages token counting,
243
+ Gemini `countTokens`, llama.cpp chat input tokens, and vLLM Messages token
244
+ counting. Ollama, OpenRouter, LM Studio, Unsloth, and subscription CLIs use a
245
+ provider-wire local estimate because those runtimes do not share a dependable
246
+ preflight tokenizer for complete chat history plus tools. UR never launches a
247
+ hidden completion for token counting. If a native count call is unavailable,
248
+ file and MCP size checks retain the local estimate rather than disabling their
249
+ limits.
250
+
218
251
  For llama.cpp, `/v1/models` metadata is preserved when the server supplies it.
219
252
  Because stock llama.cpp exposes chat-template effort support per loaded model,
220
253
  UR also resolves the model currently under the Up/Down cursor through
@@ -283,12 +316,12 @@ After selecting a provider:
283
316
  - Each model shows its concise capabilities; OpenRouter includes pricing,
284
317
  context size, tool support, and reasoning support. Its list uses compact
285
318
  model names; focus an entry to see the exact provider/model ID
286
- - Model source is displayed: `live` (dynamic discovery), `cache` (fallback), or `static` (predefined)
287
- - OpenRouter is fresh-only in this picker: every opening fetches `/models`, and
288
- a failed refresh is reported instead of showing stale cached entries
319
+ - Model source is displayed: `live` (dynamic discovery), `cache` (recent endpoint result), or `static` (predefined)
320
+ - OpenRouter reuses an endpoint-scoped `/models` result for five minutes; Ctrl+R
321
+ forces a live refresh, and a failed forced refresh never substitutes stale data
289
322
  - Up/Down changes the focused model and immediately switches to that model's
290
- provider-advertised effort list; Left/Right cycles only those exact levels;
291
- Enter confirms both
323
+ capability-backed effort selectors; Left/Right cycles only values UR can map
324
+ to provider-native levels, and Enter confirms both
292
325
  - Press Esc to go back and change provider
293
326
 
294
327
  **Confirmation**
@@ -591,7 +624,7 @@ ur provider models unsloth --json
591
624
 
592
625
  The default endpoint is `http://localhost:8888/v1`; override it with
593
626
  `ur config set base_url unsloth <url>`. Authentication is mandatory. Every Unsloth
594
- request sets `enable_tools: false`, including streaming requests. The model may
627
+ inference request sets `enable_tools: false`, including streaming requests. The model may
595
628
  still return standard OpenAI function calls; UR handles those through the same
596
629
  native tool flow used by its other providers. This keeps Unsloth provider-only
597
630
  and avoids running a second tool loop inside Studio. See the official
@@ -621,6 +654,7 @@ Required variables:
621
654
  | Provider | Required env vars | Optional env vars |
622
655
  | --- | --- | --- |
623
656
  | OpenAI-compatible | `OPENAI_COMPATIBLE_BASE_URL`, `OPENAI_COMPATIBLE_MODEL` | `OPENAI_COMPATIBLE_API_KEY` |
657
+ | Unsloth | `UNSLOTH_API_KEY`, `UNSLOTH_MODEL` | `UNSLOTH_BASE_URL` (defaults to `http://localhost:8888/v1`) |
624
658
  | OpenAI | `OPENAI_API_KEY`, `OPENAI_MODEL` | `OPENAI_BASE_URL` |
625
659
  | OpenRouter | `OPENROUTER_API_KEY`, `OPENROUTER_MODEL` | `OPENROUTER_BASE_URL` |
626
660
  | Anthropic | `ANTHROPIC_API_KEY`, `ANTHROPIC_MODEL` | `ANTHROPIC_BASE_URL` |
@@ -67,9 +67,9 @@ const featureGroups = [
67
67
  },
68
68
  {
69
69
  title: 'Providers and auth',
70
- tags: ['subscription', 'API', 'local', 'status bar'],
71
- text: 'UR-native API/local/OpenAI-compatible runtimes, first-class subscription CLI providers dispatched through the official vendor CLIs, provider doctor checks, secure API-key connect, non-secret config, fallback hints, and provider-aware status-bar output.',
72
- commands: ['ur provider list', 'ur provider status', 'ur provider doctor agy', 'ur connect status', 'ur config set provider openai-api', 'ur config set provider ollama'],
70
+ tags: ['subscription', 'API', 'local', 'effort', 'status bar'],
71
+ text: 'UR-native API/local/OpenAI-compatible runtimes, provider-scoped endpoints, provider-only Unsloth inference, capability-driven reasoning effort, responsive OpenRouter routing, first-class subscription CLI providers dispatched through the official vendor CLIs, provider doctor checks, secure API-key connect, non-secret config, fallback hints, and provider-aware status-bar output.',
72
+ commands: ['ur provider list', 'ur provider status', 'ur provider doctor agy', 'ur connect status', 'ur config set provider openai-api', 'ur config set provider ollama', 'ur config set base_url llama.cpp http://localhost:9931/v1', '/effort ultra', '/thinking on'],
73
73
  },
74
74
  {
75
75
  title: 'Security and operations',
@@ -85,7 +85,7 @@ const commands = [
85
85
  category: 'Core',
86
86
  aliases: [],
87
87
  summary: 'Start an interactive session; a fresh workspace must choose and locally persist a validated provider/model pair first.',
88
- examples: ['ur', 'ur --model qwen3-coder:480b-cloud', 'ur --continue', 'ur --resume'],
88
+ examples: ['ur', 'ur --model gpt-5.6-sol --effort ultra', 'ur --continue', 'ur --resume'],
89
89
  },
90
90
  {
91
91
  name: 'ur -p',
@@ -225,7 +225,7 @@ const commands = [
225
225
  category: 'Ops',
226
226
  aliases: ['settings'],
227
227
  summary: 'Open the config panel or persist safe non-secret provider settings.',
228
- examples: ['ur config', 'ur config set provider ollama', 'ur config set provider openai-api', 'ur config set provider anthropic-api', 'ur config set model qwen3-coder:480b-cloud', 'ur config set base_url http://localhost:11434/v1', 'ur config set provider.fallback ollama'],
228
+ examples: ['ur config', 'ur config set provider ollama', 'ur config set provider openai-api', 'ur config set provider anthropic-api', 'ur config set model qwen3-coder:480b-cloud', 'ur config set base_url ollama http://localhost:11434', 'ur config set base_url llama.cpp http://localhost:9931/v1', 'ur config set provider.fallback ollama'],
229
229
  },
230
230
  {
231
231
  name: 'test-first',
@@ -378,8 +378,8 @@ const commands = [
378
378
  name: 'plugin',
379
379
  category: 'Interop',
380
380
  aliases: ['plugins'],
381
- summary: 'Manage UR plugins and marketplaces for MCP tools, skills, templates, validators, language adapters, LSP servers, agents, hooks, output styles, and commands.',
382
- examples: ['ur plugin search git', 'ur plugin search --capability skills --json', 'ur plugin show github@ur-plugins-official', 'ur plugin list', 'ur plugin install hello@ur-plugins-official', 'ur plugin update <plugin>'],
381
+ summary: 'Manage UR plugins and GitHub, Git, URL, npm, local, or settings-backed marketplaces for MCP tools, skills, templates, validators, language adapters, LSP servers, agents, hooks, output styles, and commands.',
382
+ examples: ['ur plugin marketplace add npm:@scope/catalog@latest', 'ur plugin search git', 'ur plugin search --capability skills --json', 'ur plugin show github@ur-plugins-official', 'ur plugin list', 'ur plugin install hello@ur-plugins-official', 'ur plugin update <plugin>'],
383
383
  },
384
384
  {
385
385
  name: 'provider',
@@ -547,8 +547,8 @@ const slashGroups = [
547
547
  },
548
548
  {
549
549
  title: 'Models, tools, and interop',
550
- items: ['/model', '/model-doctor', '/model-route', '/escalate', '/mcp', '/plugin', '/skills', '/skill', '/sdk', '/a2a-card'],
551
- text: 'Pick models, inspect capabilities, manage MCP/plugin extensions, browse prompt skills with /skills, run executable workflows with /skill, and expose interop surfaces.',
550
+ items: ['/model', '/provider', '/effort', '/thinking', '/fast', '/model-doctor', '/model-route', '/escalate', '/mcp', '/plugin', '/skills', '/skill', '/sdk', '/a2a-card'],
551
+ text: 'Pick providers and models, cycle only capability-backed effort selectors or provider-native boolean thinking, inspect capabilities, manage MCP/plugin extensions, browse prompt skills with /skills, run executable workflows with /skill, and expose interop surfaces.',
552
552
  },
553
553
  {
554
554
  title: 'Security operations',
@@ -45,7 +45,7 @@
45
45
  <main id="content" class="content">
46
46
  <header class="topbar">
47
47
  <div>
48
- <p class="eyebrow">Version 1.83.1</p>
48
+ <p class="eyebrow">Version 1.84.0</p>
49
49
  <h1>UR-Nexus Documentation</h1>
50
50
  <p class="lead">A practical, tutorial-style reference for installing, configuring, automating, extending, and operating UR-Nexus.</p>
51
51
  </div>
@@ -76,7 +76,7 @@
76
76
  </article>
77
77
  <article>
78
78
  <strong>Plugin marketplace</strong>
79
- <span>Plugins can add MCP tools, skills, templates, validators, language adapters, LSP servers, agents, hooks, and output styles.</span>
79
+ <span>GitHub, Git, URL, npm, local, and settings-backed catalogs can add MCP tools, skills, templates, validators, language adapters, LSP servers, agents, hooks, and output styles.</span>
80
80
  </article>
81
81
  <article>
82
82
  <strong>Legal provider routing</strong>
@@ -100,7 +100,7 @@
100
100
  </article>
101
101
  <article>
102
102
  <strong>Deterministic commands</strong>
103
- <span>The external runtime exposes 172 commands and 253 unique slash tokens; registry tests reject ambiguous names, broken loaders, and undocumented visible commands.</span>
103
+ <span>The external runtime exposes a source-derived command catalog; registry tests reject ambiguous names, broken loaders, undocumented visible aliases, and documented commands without implementations.</span>
104
104
  </article>
105
105
  <article>
106
106
  <strong>Safety and context</strong>
@@ -165,16 +165,32 @@ ur provider doctor agy</code></pre>
165
165
  <h3>API and local providers</h3>
166
166
  <pre><code>ur config set provider openai-compatible
167
167
  ur config set provider openai-api
168
- ur config set base_url http://localhost:11434/v1
168
+ ur config set base_url ollama http://localhost:11434
169
+ ur config set base_url llama.cpp http://localhost:9931/v1
170
+ ur config set provider unsloth
169
171
  ur config set model qwen3-coder:480b-cloud
170
172
  ur config set provider.fallback ollama
171
173
  ur config set openai_transport responses
172
174
  ur config set responses.store false</code></pre>
173
- <p>API providers require explicit selection and read keys from a key stored via <code>ur connect</code> (OS keychain) or from environment variables. OpenAI Responses is opt-in and privacy-first; Chat Completions remains the default. Local providers include Ollama, LM Studio, llama.cpp, and vLLM.</p>
175
+ <p>API providers require explicit selection and read keys from a key stored via <code>ur connect</code> (OS keychain) or from environment variables. Each configurable provider keeps its own <code>base_url</code>, so switching among Ollama, LM Studio, llama.cpp, vLLM, Unsloth, and API gateways restores the matching address. OpenAI Responses is opt-in and privacy-first; Chat Completions remains the default. Unsloth is an authenticated inference provider only.</p>
176
+ </article>
177
+ <article>
178
+ <h3>Capability-driven reasoning effort</h3>
179
+ <pre><code>/effort ultra
180
+ /thinking on
181
+ ur --model kimi-k3:cloud --effort high
182
+ /model # Up/Down model · Left/Right effort or boolean thinking · Enter apply</code></pre>
183
+ <p>The normalized vocabulary is <code>minimal</code>, <code>low</code>, <code>medium</code>, <code>high</code>, <code>xhigh</code>, <code>max</code>, and <code>ultra</code>; <code>/effort auto</code> clears an explicit choice. UR lists only capability-backed selectors it can map to the focused model's provider-native levels. <code>max</code> resolves to the highest supported non-Ultra tier. Ultra appears only for native <code>ultra</code>, advertised <code>max</code>/<code>xhigh</code>, or an explicit provider alias; the picker shows translations such as <code>ultra→max</code> and sends that exact provider value. Models that top out at <code>high</code>, boolean-thinking models, and unknown capabilities do not get Ultra. For boolean-thinking models on runtimes with a native toggle, Left selects off, Right selects on, and <code>t</code> toggles; <code>/thinking on|off</code> is the direct control. A graded <code>/effort</code> request on such a model enables thinking while reporting that no graded value was sent. Generic OpenAI-compatible runtimes receive no invented boolean field.</p>
184
+ </article>
185
+ <article>
186
+ <h3>OpenRouter responsive routing</h3>
187
+ <pre><code>ur config set provider openrouter
188
+ /model # cached catalog; Ctrl+R forces live refresh</code></pre>
189
+ <p>The endpoint-scoped model catalog is reused for five minutes, while forced refresh never substitutes stale data. Interactive requests prefer OpenRouter latency routing, a stable session ID, and provider-authored prompt-cache markers; explicit routing preferences and model variants still win.</p>
174
190
  </article>
175
191
  <article>
176
192
  <h3>Status bar and updates</h3>
177
- <pre><code>Ollama | llama3 | ask | main | update 1.48.0 available</code></pre>
193
+ <pre><code>Ollama | llama3 | ask | main | update available</code></pre>
178
194
  <p>The interactive status bar shows only important runtime state: provider, model, mode, branch, active tasks, checks status when known, and update availability. It is hidden in CI, dumb terminals, and print mode.</p>
179
195
  </article>
180
196
  <article>
@@ -38,7 +38,11 @@ ur config set provider codex-cli
38
38
  ```
39
39
 
40
40
  Inside a session, `/model` gives the same flow interactively: provider first,
41
- then only that provider's models. Verify the active pair any time:
41
+ then only that provider's models. Use Up/Down to focus a model and Left/Right to
42
+ cycle its capability-backed effort selectors before pressing Enter. Ultra is
43
+ shown only when UR can map it to an advertised native `ultra`, `max`, `xhigh`,
44
+ or explicit alias; the picker displays the exact mapping, such as
45
+ `ultra→max`. Verify the active pair any time:
42
46
 
43
47
  ```sh
44
48
  ur provider status
@@ -7,7 +7,7 @@ plugins {
7
7
  }
8
8
 
9
9
  group = "dev.urnexus"
10
- version = "1.83.1"
10
+ version = "1.84.0"
11
11
 
12
12
  repositories {
13
13
  mavenCentral()
@@ -2,7 +2,7 @@
2
2
  "name": "ur-inline-diffs",
3
3
  "displayName": "UR Inline Diffs",
4
4
  "description": "Review, apply, and reject UR inline diff bundles from .ur/ide/diffs inside VS Code.",
5
- "version": "1.83.1",
5
+ "version": "1.84.0",
6
6
  "publisher": "ur-nexus",
7
7
  "engines": {
8
8
  "vscode": "^1.92.0"
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ur-agent",
3
- "version": "1.83.1",
3
+ "version": "1.84.0",
4
4
  "description": "UR-Nexus — autonomous engineering workflow engine (plan, execute, test, verify, document, benchmark, reproduce)",
5
5
  "type": "module",
6
6
  "packageManager": "bun@1.3.14",