ur-agent 1.65.11 → 1.65.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -19,7 +19,7 @@ You need:
19
19
 
20
20
  ```sh
21
21
  ur --version
22
- # expected for this release: "1.65.11 (UR-Nexus)"
22
+ # expected for this release: "1.65.12 (UR-Nexus)"
23
23
  ```
24
24
 
25
25
  ## 0.1 First-workspace model selection (1.45.4)
@@ -45,7 +45,7 @@
45
45
  <main id="content" class="content">
46
46
  <header class="topbar">
47
47
  <div>
48
- <p class="eyebrow">Version 1.65.11</p>
48
+ <p class="eyebrow">Version 1.65.12</p>
49
49
  <h1>UR-Nexus Documentation</h1>
50
50
  <p class="lead">A practical, tutorial-style reference for installing, configuring, automating, extending, and operating UR-Nexus.</p>
51
51
  </div>
@@ -7,7 +7,7 @@ plugins {
7
7
  }
8
8
 
9
9
  group = "dev.urnexus"
10
- version = "1.65.11"
10
+ version = "1.65.12"
11
11
 
12
12
  repositories {
13
13
  mavenCentral()
@@ -2,7 +2,7 @@
2
2
  "name": "ur-inline-diffs",
3
3
  "displayName": "UR Inline Diffs",
4
4
  "description": "Review, apply, and reject UR inline diff bundles from .ur/ide/diffs inside VS Code.",
5
- "version": "1.65.11",
5
+ "version": "1.65.12",
6
6
  "publisher": "ur-nexus",
7
7
  "engines": {
8
8
  "vscode": "^1.92.0"
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ur-agent",
3
- "version": "1.65.11",
3
+ "version": "1.65.12",
4
4
  "description": "UR-Nexus — autonomous engineering workflow engine (plan, execute, test, verify, document, benchmark, reproduce)",
5
5
  "type": "module",
6
6
  "packageManager": "bun@1.3.14",
@@ -92,6 +92,24 @@ redirection, backgrounding, sandbox overrides, and permission-time rewrites to
92
92
  mutating commands fail closed. The preview command remains a Bash side effect
93
93
  and still follows normal permission, sandbox, and plan-worker rules.
94
94
 
95
+ Task completion also protects that lifecycle boundary. When the final
96
+ actionable `in_progress` task has a successful `Write`/`Edit`/`MultiEdit`/
97
+ `NotebookEdit` after its recorded start but no later successful inspection,
98
+ runtime, test, shell, or delegated-check result, `TaskUpdate(completed)` is
99
+ soft-deferred: it returns a non-error explanation and leaves the same task
100
+ `in_progress`. The model verifies and retries completion instead of creating a
101
+ duplicate task or discovering an all-terminal dead end on the next corrective
102
+ Edit. The guard is evidence-based and conservative: missing/compacted history,
103
+ non-file work, and intermediate tasks are not guessed into a deferred state.
104
+
105
+ Live plan mode also treats setup of the exact current session plan artifact as
106
+ planning infrastructure rather than implementation. `Write` creates the plan
107
+ file's parent automatically, but weak models may first emit `mkdir -p` for that
108
+ exact parent or the bounded `ls ... || mkdir -p ... && ls ...` check. Only those
109
+ exact-path shapes bypass the task-list requirement; Bash permission and sandbox
110
+ checks still apply, and a hook rewrite, sibling path, extra command, expansion,
111
+ background launch, or sandbox override fails closed at the final boundary.
112
+
95
113
  `AskUserQuestion` exposes a request-only model schema: one top-level
96
114
  `questions` array with 1–4 complete question objects, each containing
97
115
  `question`, a header of at most 12 characters, and 2–8 labeled choices.
@@ -101,6 +119,18 @@ question-text aliases; it does not turn arbitrary prose or flat option rows into
101
119
  invented questions. More than four blocking decisions are asked in later
102
120
  rounds.
103
121
 
122
+ One narrow end-turn recovery exists for weak models that clearly attempted this
123
+ tool but failed to emit a native call. On an interactive main-agent turn with
124
+ no existing tool use, the runtime may recover either one canonical
125
+ `questions` object at the very end of a reasoning block that explicitly says
126
+ to invoke `AskUserQuestion`, or one standalone Markdown decision menu with
127
+ exactly one bold question, 2–8 bold labeled options with descriptions, and a
128
+ terminal instruction to select an option. The recovered object must pass the
129
+ live `AskUserQuestion` schema unchanged before the normal tool executor opens
130
+ the UI. JSON repair, truncation, duplicate/ambiguous candidates, casual “A or
131
+ B?” prose, examples, incomplete menus, background workers, headless sessions,
132
+ and unavailable/disabled tools all fail closed.
133
+
104
134
  Answers and annotations are not model input fields. They are accepted only
105
135
  during post-permission validation after the interactive UI has returned one
106
136
  non-empty answer for every question; an unchanged generic approval cannot
@@ -180,10 +210,10 @@ the intended file.
180
210
 
181
211
  `Edit` remains fail-closed rather than applying a fuzzy replacement to similar
182
212
  code. When an exact contiguous `old_string` is absent, its bounded error points
183
- to a verified matching line when one exists and tells the model to re-read that
184
- region, use a smaller current 2–4-line anchor, split distant HTML/CSS/JavaScript
185
- sections, and never retry the unchanged call. One narrow idempotent case returns
186
- success without writing: a non-`replace_all` deletion-only edit whose
187
- `new_string` is already present uniquely and whose larger `old_string` is
188
- absent. General stale, fuzzy, empty-replacement, and ambiguous matches still
189
- fail.
213
+ to the most distinctive verified matching line when one exists, rather than an
214
+ unrelated generic delimiter, and tells the model to re-read that region, use a
215
+ smaller current 2–4-line anchor, split distant HTML/CSS/JavaScript sections, and
216
+ never retry the unchanged call. One narrow idempotent case returns success
217
+ without writing: a non-`replace_all` deletion-only edit whose `new_string` is
218
+ already present uniquely and whose larger `old_string` is absent. General
219
+ stale, fuzzy, empty-replacement, and ambiguous matches still fail.
@@ -111,15 +111,18 @@ ur --discover-ollama # scan the LAN for Ollama servers (ollamaDiscovery.ts)
111
111
  search, and bounded private cursor state. Compacted context persistence
112
112
  requires a 32-byte `UR_OPENAI_RESPONSES_STATE_KEY`.
113
113
  - `ollama.ts` selects timeouts from explicit request options, then
114
- `API_TIMEOUT_MS`, then runtime defaults. `:cloud` models and remote sessions
115
- use 120 seconds; local models use 300 seconds. The same model-aware value is
116
- applied while waiting for `/api/chat` response headers and as the absolute
117
- deadline in `readOllamaChunks`.
114
+ `API_TIMEOUT_MS`, then runtime defaults. Remote sessions and ordinary
115
+ `:cloud` models use 120 seconds; local models and Kimi K2.7 Cloud use 300
116
+ seconds because large Kimi coding turns can legitimately take longer than
117
+ two minutes before completing. The same model-aware value is applied while
118
+ waiting for `/api/chat` response headers and as the absolute deadline in
119
+ `readOllamaChunks`.
118
120
  - `ur.ts` identifies an Ollama Cloud runtime from both the selected provider and
119
121
  the `:cloud` suffix. It disables shared automatic request retries for that
120
- route, applies the same 120-second bound to any permitted non-streaming
122
+ route, applies the same model-aware bound to any permitted non-streaming
121
123
  fallback, and skips fallback entirely when the Ollama stream deadline itself
122
- caused the failure. Explicit `API_TIMEOUT_MS` remains authoritative.
124
+ caused the failure. Explicit `API_TIMEOUT_MS` remains authoritative, and
125
+ remote-session safety keeps its 120-second ceiling.
123
126
 
124
127
  ## Capability-aware routing
125
128
 
@@ -149,6 +149,11 @@ use legacy `TodoWrite` by default, or Task V2 when
149
149
  - Dependencies block transition or claim until prerequisites are complete.
150
150
  - Actionable `pending` or `in_progress` entries open the gate; completed and
151
151
  internal entries do not.
152
+ - The last actionable `in_progress` task cannot be terminalized immediately
153
+ after a recorded file mutation with no later successful observable check.
154
+ `TaskUpdate` soft-defers that completion and keeps the same task actionable;
155
+ it does not infer or auto-create a replacement task. A later successful
156
+ inspection/runtime/test/delegated check allows the explicit completion retry.
152
157
  - Reads remain unrestricted. Ordinary mutations have a default allowance of
153
158
  three preceding tool calls, counted by tool call rather than message.
154
159
  Delegation and child mutations always require an actionable parent task.
@@ -159,10 +164,13 @@ use legacy `TodoWrite` by default, or Task V2 when
159
164
  exempt so the agent can repair the plan.
160
165
  - Creating or updating the exact current-session plan-mode Markdown file is
161
166
  also exempt: that file is the planning artifact, not an ordinary workspace
162
- change. The exemption and live actionable-task state are re-evaluated at the
163
- final execution boundary after permission-hook input rewrites; sibling files
164
- and the plans directory are not exempt, and the filename alone is not exempt
165
- outside live plan mode.
167
+ change. A bounded Bash bootstrap may only create/check that file's exact
168
+ parent (`mkdir -p`, optionally guarded by the known `ls` pattern); this
169
+ compatibility path remains subject to Bash permission and sandbox checks.
170
+ The exemption and live actionable-task state are re-evaluated at the final
171
+ execution boundary after permission-hook input rewrites. Sibling paths,
172
+ general plans-directory commands, added shell operations, and the filename
173
+ alone outside live plan mode are not exempt.
166
174
  - Configure the behavior at
167
175
  `tasks.requireBeforeChanges.{enabled,freeReads}`.
168
176
 
@@ -1,6 +1,6 @@
1
1
  # UR-Nexus — Technical Specifications
2
2
 
3
- > Audited against the executable source and tests for `ur-agent` v1.65.11.
3
+ > Audited against the executable source and tests for `ur-agent` v1.65.12.
4
4
  > Command, tool, flag, provider, and setting claims are checked against the
5
5
  > implementation rather than copied from product prose. Release validation
6
6
  > keeps this version synchronized and packages the complete `technical/`