ur-agent 1.65.10 → 1.65.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -19,7 +19,7 @@ You need:
19
19
 
20
20
  ```sh
21
21
  ur --version
22
- # expected for this release: "1.65.10 (UR-Nexus)"
22
+ # expected for this release: "1.65.12 (UR-Nexus)"
23
23
  ```
24
24
 
25
25
  ## 0.1 First-workspace model selection (1.45.4)
@@ -45,7 +45,7 @@
45
45
  <main id="content" class="content">
46
46
  <header class="topbar">
47
47
  <div>
48
- <p class="eyebrow">Version 1.65.10</p>
48
+ <p class="eyebrow">Version 1.65.12</p>
49
49
  <h1>UR-Nexus Documentation</h1>
50
50
  <p class="lead">A practical, tutorial-style reference for installing, configuring, automating, extending, and operating UR-Nexus.</p>
51
51
  </div>
@@ -7,7 +7,7 @@ plugins {
7
7
  }
8
8
 
9
9
  group = "dev.urnexus"
10
- version = "1.65.10"
10
+ version = "1.65.12"
11
11
 
12
12
  repositories {
13
13
  mavenCentral()
@@ -2,7 +2,7 @@
2
2
  "name": "ur-inline-diffs",
3
3
  "displayName": "UR Inline Diffs",
4
4
  "description": "Review, apply, and reject UR inline diff bundles from .ur/ide/diffs inside VS Code.",
5
- "version": "1.65.10",
5
+ "version": "1.65.12",
6
6
  "publisher": "ur-nexus",
7
7
  "engines": {
8
8
  "vscode": "^1.92.0"
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ur-agent",
3
- "version": "1.65.10",
3
+ "version": "1.65.12",
4
4
  "description": "UR-Nexus — autonomous engineering workflow engine (plan, execute, test, verify, document, benchmark, reproduce)",
5
5
  "type": "module",
6
6
  "packageManager": "bun@1.3.14",
@@ -61,6 +61,10 @@ only after `EnterPlanMode` (or `/plan`) has successfully made the active mode
61
61
  `plan`. Plan approval may change that mode before permission-edited input is
62
62
  revalidated; the executor labels that second validation as post-permission so
63
63
  the already-validated exit can finish, while new out-of-mode calls still fail.
64
+ `ExitPlanMode` is exempt from the implementation task-list gate because it is
65
+ the approval/control transition that precedes implementation. Its own plan-mode
66
+ validation remains authoritative, so the exemption does not make a stale
67
+ second exit valid.
64
68
 
65
69
  For non-trivial work, the task list uses one record per cohesive outcome with
66
70
  an observable done check rather than one omnibus record. Genuine single-outcome
@@ -79,7 +83,62 @@ precision-losing numeric IDs are rejected.
79
83
  Task-gate recovery names the tracking surface that is actually present:
80
84
  interactive Task V2 sessions use `TaskCreate`, while default headless sessions
81
85
  use `TodoWrite`. It never instructs a model to recover by calling a tool absent
82
- from that runtime.
86
+ from that runtime. Runtime inspection tracks actionable and total user tasks
87
+ separately: an all-terminal list is reported truthfully and the model is told
88
+ to reopen or create the cohesive remaining task. Real Edit/Bash mutations stay
89
+ gated. One simple `open <loopback-http(s)-URL>` Bash preview is exempt only
90
+ from the task-list gate; remote/file URLs, flags, shell composition, expansion,
91
+ redirection, backgrounding, sandbox overrides, and permission-time rewrites to
92
+ mutating commands fail closed. The preview command remains a Bash side effect
93
+ and still follows normal permission, sandbox, and plan-worker rules.
94
+
95
+ Task completion also protects that lifecycle boundary. When the final
96
+ actionable `in_progress` task has a successful `Write`/`Edit`/`MultiEdit`/
97
+ `NotebookEdit` after its recorded start but no later successful inspection,
98
+ runtime, test, shell, or delegated-check result, `TaskUpdate(completed)` is
99
+ soft-deferred: it returns a non-error explanation and leaves the same task
100
+ `in_progress`. The model verifies and retries completion instead of creating a
101
+ duplicate task or discovering an all-terminal dead end on the next corrective
102
+ Edit. The guard is evidence-based and conservative: missing/compacted history,
103
+ non-file work, and intermediate tasks are not guessed into a deferred state.
104
+
105
+ Live plan mode also treats setup of the exact current session plan artifact as
106
+ planning infrastructure rather than implementation. `Write` creates the plan
107
+ file's parent automatically, but weak models may first emit `mkdir -p` for that
108
+ exact parent or the bounded `ls ... || mkdir -p ... && ls ...` check. Only those
109
+ exact-path shapes bypass the task-list requirement; Bash permission and sandbox
110
+ checks still apply, and a hook rewrite, sibling path, extra command, expansion,
111
+ background launch, or sandbox override fails closed at the final boundary.
112
+
113
+ `AskUserQuestion` exposes a request-only model schema: one top-level
114
+ `questions` array with 1–4 complete question objects, each containing
115
+ `question`, a header of at most 12 characters, and 2–8 labeled choices.
116
+ Descriptions are optional and are never fabricated from labels. The runtime
117
+ accepts only lossless compatibility forms such as string choices and recognized
118
+ question-text aliases; it does not turn arbitrary prose or flat option rows into
119
+ invented questions. More than four blocking decisions are asked in later
120
+ rounds.
121
+
122
+ One narrow end-turn recovery exists for weak models that clearly attempted this
123
+ tool but failed to emit a native call. On an interactive main-agent turn with
124
+ no existing tool use, the runtime may recover either one canonical
125
+ `questions` object at the very end of a reasoning block that explicitly says
126
+ to invoke `AskUserQuestion`, or one standalone Markdown decision menu with
127
+ exactly one bold question, 2–8 bold labeled options with descriptions, and a
128
+ terminal instruction to select an option. The recovered object must pass the
129
+ live `AskUserQuestion` schema unchanged before the normal tool executor opens
130
+ the UI. JSON repair, truncation, duplicate/ambiguous candidates, casual “A or
131
+ B?” prose, examples, incomplete menus, background workers, headless sessions,
132
+ and unavailable/disabled tools all fail closed.
133
+
134
+ Answers and annotations are not model input fields. They are accepted only
135
+ during post-permission validation after the interactive UI has returned one
136
+ non-empty answer for every question; an unchanged generic approval cannot
137
+ produce a successful “user answered” result. The UI uses prototype-safe records,
138
+ provides a real custom `Other` path for both ordinary and preview questions,
139
+ and does not count selecting `Other` itself as an answer. HTML-configured
140
+ previews are escaped into an inert preformatted-text wrapper rather than
141
+ executed as model-provided markup.
83
142
 
84
143
  ## Multi-agent tools
85
144
 
@@ -144,8 +203,17 @@ not only a modification timestamp. Full and ranged reads are compared at the
144
203
  final write boundary, preventing same-timestamp external replacements from
145
204
  being overwritten.
146
205
 
206
+ `Write` requires `file_path` and the complete literal `content` in the same
207
+ structured call. Prose outside the call is never treated as file content, and a
208
+ missing-content failure states that no file was written instead of fabricating
209
+ the intended file.
210
+
147
211
  `Edit` remains fail-closed rather than applying a fuzzy replacement to similar
148
212
  code. When an exact contiguous `old_string` is absent, its bounded error points
149
- to a verified matching line when one exists and tells the model to re-read that
150
- region, use a smaller current 2–4-line anchor, split distant HTML/CSS/JavaScript
151
- sections, and never retry the unchanged call.
213
+ to the most distinctive verified matching line when one exists, rather than an
214
+ unrelated generic delimiter, and tells the model to re-read that region, use a
215
+ smaller current 2–4-line anchor, split distant HTML/CSS/JavaScript sections, and
216
+ never retry the unchanged call. One narrow idempotent case returns success
217
+ without writing: a non-`replace_all` deletion-only edit whose `new_string` is
218
+ already present uniquely and whose larger `old_string` is absent. General
219
+ stale, fuzzy, empty-replacement, and ambiguous matches still fail.
@@ -111,15 +111,18 @@ ur --discover-ollama # scan the LAN for Ollama servers (ollamaDiscovery.ts)
111
111
  search, and bounded private cursor state. Compacted context persistence
112
112
  requires a 32-byte `UR_OPENAI_RESPONSES_STATE_KEY`.
113
113
  - `ollama.ts` selects timeouts from explicit request options, then
114
- `API_TIMEOUT_MS`, then runtime defaults. `:cloud` models and remote sessions
115
- use 120 seconds; local models use 300 seconds. The same model-aware value is
116
- applied while waiting for `/api/chat` response headers and as the absolute
117
- deadline in `readOllamaChunks`.
114
+ `API_TIMEOUT_MS`, then runtime defaults. Remote sessions and ordinary
115
+ `:cloud` models use 120 seconds; local models and Kimi K2.7 Cloud use 300
116
+ seconds because large Kimi coding turns can legitimately take longer than
117
+ two minutes before completing. The same model-aware value is applied while
118
+ waiting for `/api/chat` response headers and as the absolute deadline in
119
+ `readOllamaChunks`.
118
120
  - `ur.ts` identifies an Ollama Cloud runtime from both the selected provider and
119
121
  the `:cloud` suffix. It disables shared automatic request retries for that
120
- route, applies the same 120-second bound to any permitted non-streaming
122
+ route, applies the same model-aware bound to any permitted non-streaming
121
123
  fallback, and skips fallback entirely when the Ollama stream deadline itself
122
- caused the failure. Explicit `API_TIMEOUT_MS` remains authoritative.
124
+ caused the failure. Explicit `API_TIMEOUT_MS` remains authoritative, and
125
+ remote-session safety keeps its 120-second ceiling.
123
126
 
124
127
  ## Capability-aware routing
125
128
 
@@ -149,6 +149,11 @@ use legacy `TodoWrite` by default, or Task V2 when
149
149
  - Dependencies block transition or claim until prerequisites are complete.
150
150
  - Actionable `pending` or `in_progress` entries open the gate; completed and
151
151
  internal entries do not.
152
+ - The last actionable `in_progress` task cannot be terminalized immediately
153
+ after a recorded file mutation with no later successful observable check.
154
+ `TaskUpdate` soft-defers that completion and keeps the same task actionable;
155
+ it does not infer or auto-create a replacement task. A later successful
156
+ inspection/runtime/test/delegated check allows the explicit completion retry.
152
157
  - Reads remain unrestricted. Ordinary mutations have a default allowance of
153
158
  three preceding tool calls, counted by tool call rather than message.
154
159
  Delegation and child mutations always require an actionable parent task.
@@ -159,10 +164,13 @@ use legacy `TodoWrite` by default, or Task V2 when
159
164
  exempt so the agent can repair the plan.
160
165
  - Creating or updating the exact current-session plan-mode Markdown file is
161
166
  also exempt: that file is the planning artifact, not an ordinary workspace
162
- change. The exemption and live actionable-task state are re-evaluated at the
163
- final execution boundary after permission-hook input rewrites; sibling files
164
- and the plans directory are not exempt, and the filename alone is not exempt
165
- outside live plan mode.
167
+ change. A bounded Bash bootstrap may only create/check that file's exact
168
+ parent (`mkdir -p`, optionally guarded by the known `ls` pattern); this
169
+ compatibility path remains subject to Bash permission and sandbox checks.
170
+ The exemption and live actionable-task state are re-evaluated at the final
171
+ execution boundary after permission-hook input rewrites. Sibling paths,
172
+ general plans-directory commands, added shell operations, and the filename
173
+ alone outside live plan mode are not exempt.
166
174
  - Configure the behavior at
167
175
  `tasks.requireBeforeChanges.{enabled,freeReads}`.
168
176
 
@@ -38,6 +38,23 @@ feature. Its AI-classified `auto` mode is therefore source-only. The accepted
38
38
  `permissions.classifierPermissionsEnabled` schema field is not read by a
39
39
  shipped enforcement path and must not be treated as an active control.
40
40
 
41
+ Tools that require user interaction are semantically validated again after the
42
+ permission decision even when their input signature did not change. This keeps
43
+ an unchanged generic approval or hook allow-result from being mistaken for the
44
+ interaction result.
45
+ For `AskUserQuestion`, model requests cannot supply `answers` or `annotations`;
46
+ post-permission execution requires an exact, complete, non-empty answer map.
47
+ Question-keyed UI state uses null-prototype records and own-property checks, so
48
+ model-controlled names such as `constructor` or `__proto__` cannot become
49
+ inherited answers or mutate the record prototype. The tool rejects those
50
+ reserved question names at its outer schema as an additional boundary.
51
+
52
+ Model-provided Ask previews are never trusted as executable HTML. When an SDK
53
+ selects the HTML preview format, the runtime escapes raw input into one inert
54
+ `<pre data-ur-preview="text">…</pre>` wrapper and validates that wrapper before
55
+ interaction. Scripts, event handlers, URL attributes, style injection, and
56
+ closing-tag escapes remain text rather than active markup.
57
+
41
58
  Bash adds command parsing, command-injection checks, dangerous-pattern
42
59
  classification, project safety policy, path validation, and sandbox-aware
43
60
  permission decisions. `UR_CODE_DISABLE_COMMAND_INJECTION_CHECK` weakens one of
@@ -1,6 +1,6 @@
1
1
  # UR-Nexus — Technical Specifications
2
2
 
3
- > Audited against the executable source and tests for `ur-agent` v1.65.10.
3
+ > Audited against the executable source and tests for `ur-agent` v1.65.12.
4
4
  > Command, tool, flag, provider, and setting claims are checked against the
5
5
  > implementation rather than copied from product prose. Release validation
6
6
  > keeps this version synchronized and packages the complete `technical/`