@namzu/sdk 6.2.0 → 8.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +677 -0
- package/dist/agents/ReactiveAgent.d.ts.map +1 -1
- package/dist/agents/ReactiveAgent.js +5 -0
- package/dist/agents/ReactiveAgent.js.map +1 -1
- package/dist/agents/SupervisorAgent.d.ts.map +1 -1
- package/dist/agents/SupervisorAgent.js +172 -158
- package/dist/agents/SupervisorAgent.js.map +1 -1
- package/dist/agents/__tests__/supervisor-inbox-scope.test.d.ts +2 -0
- package/dist/agents/__tests__/supervisor-inbox-scope.test.d.ts.map +1 -0
- package/dist/agents/__tests__/supervisor-inbox-scope.test.js +125 -0
- package/dist/agents/__tests__/supervisor-inbox-scope.test.js.map +1 -0
- package/dist/agents/runAgent.d.ts +19 -1
- package/dist/agents/runAgent.d.ts.map +1 -1
- package/dist/agents/runAgent.js +2 -0
- package/dist/agents/runAgent.js.map +1 -1
- package/dist/bridge/a2a/mapper.d.ts.map +1 -1
- package/dist/bridge/a2a/mapper.js +4 -0
- package/dist/bridge/a2a/mapper.js.map +1 -1
- package/dist/bridge/sse/mapper.d.ts.map +1 -1
- package/dist/bridge/sse/mapper.js +24 -0
- package/dist/bridge/sse/mapper.js.map +1 -1
- package/dist/connector/mcp/__tests__/positional-arrays.test.d.ts +2 -0
- package/dist/connector/mcp/__tests__/positional-arrays.test.d.ts.map +1 -0
- package/dist/connector/mcp/__tests__/positional-arrays.test.js +142 -0
- package/dist/connector/mcp/__tests__/positional-arrays.test.js.map +1 -0
- package/dist/connector/mcp/adapter.d.ts.map +1 -1
- package/dist/connector/mcp/adapter.js +123 -8
- package/dist/connector/mcp/adapter.js.map +1 -1
- package/dist/constants/agent/index.d.ts +5 -0
- package/dist/constants/agent/index.d.ts.map +1 -1
- package/dist/constants/agent/index.js +5 -0
- package/dist/constants/agent/index.js.map +1 -1
- package/dist/constants/plugin/index.d.ts +15 -0
- package/dist/constants/plugin/index.d.ts.map +1 -1
- package/dist/constants/plugin/index.js +15 -0
- package/dist/constants/plugin/index.js.map +1 -1
- package/dist/contracts/api.d.ts +1 -1
- package/dist/contracts/api.d.ts.map +1 -1
- package/dist/gateway/__tests__/completion-inbox.test.js +292 -2
- package/dist/gateway/__tests__/completion-inbox.test.js.map +1 -1
- package/dist/gateway/completion-inbox.d.ts +94 -6
- package/dist/gateway/completion-inbox.d.ts.map +1 -1
- package/dist/gateway/completion-inbox.js +235 -15
- package/dist/gateway/completion-inbox.js.map +1 -1
- package/dist/gateway/local.d.ts +11 -0
- package/dist/gateway/local.d.ts.map +1 -1
- package/dist/gateway/local.js +27 -1
- package/dist/gateway/local.js.map +1 -1
- package/dist/manager/agent/lifecycle.d.ts.map +1 -1
- package/dist/manager/agent/lifecycle.js +6 -0
- package/dist/manager/agent/lifecycle.js.map +1 -1
- package/dist/manager/run/persistence.d.ts +8 -0
- package/dist/manager/run/persistence.d.ts.map +1 -1
- package/dist/manager/run/persistence.js +12 -0
- package/dist/manager/run/persistence.js.map +1 -1
- package/dist/provider/thinking-support.d.ts +2 -1
- package/dist/provider/thinking-support.d.ts.map +1 -1
- package/dist/provider/thinking-support.js +14 -0
- package/dist/provider/thinking-support.js.map +1 -1
- package/dist/public-runtime.d.ts +1 -1
- package/dist/public-runtime.d.ts.map +1 -1
- package/dist/public-runtime.js +9 -1
- package/dist/public-runtime.js.map +1 -1
- package/dist/run/reporter.d.ts.map +1 -1
- package/dist/run/reporter.js +11 -0
- package/dist/run/reporter.js.map +1 -1
- package/dist/runtime/query/__tests__/completion-does-not-erase-the-answer.test.d.ts +2 -0
- package/dist/runtime/query/__tests__/completion-does-not-erase-the-answer.test.d.ts.map +1 -0
- package/dist/runtime/query/__tests__/completion-does-not-erase-the-answer.test.js +142 -0
- package/dist/runtime/query/__tests__/completion-does-not-erase-the-answer.test.js.map +1 -0
- package/dist/runtime/query/__tests__/completion-notification.test.js +414 -32
- package/dist/runtime/query/__tests__/completion-notification.test.js.map +1 -1
- package/dist/runtime/query/__tests__/context-size-on-the-wire.test.d.ts +2 -0
- package/dist/runtime/query/__tests__/context-size-on-the-wire.test.d.ts.map +1 -0
- package/dist/runtime/query/__tests__/context-size-on-the-wire.test.js +100 -0
- package/dist/runtime/query/__tests__/context-size-on-the-wire.test.js.map +1 -0
- package/dist/runtime/query/__tests__/context.test.js +18 -0
- package/dist/runtime/query/__tests__/context.test.js.map +1 -1
- package/dist/runtime/query/__tests__/effort-reaches-the-wire.test.d.ts +2 -0
- package/dist/runtime/query/__tests__/effort-reaches-the-wire.test.d.ts.map +1 -0
- package/dist/runtime/query/__tests__/effort-reaches-the-wire.test.js +118 -0
- package/dist/runtime/query/__tests__/effort-reaches-the-wire.test.js.map +1 -0
- package/dist/runtime/query/__tests__/tool-timeout.test.js +34 -0
- package/dist/runtime/query/__tests__/tool-timeout.test.js.map +1 -1
- package/dist/runtime/query/context.d.ts.map +1 -1
- package/dist/runtime/query/context.js +16 -1
- package/dist/runtime/query/context.js.map +1 -1
- package/dist/runtime/query/executor.d.ts.map +1 -1
- package/dist/runtime/query/executor.js +11 -1
- package/dist/runtime/query/executor.js.map +1 -1
- package/dist/runtime/query/guard.d.ts +28 -0
- package/dist/runtime/query/guard.d.ts.map +1 -1
- package/dist/runtime/query/guard.js +31 -0
- package/dist/runtime/query/guard.js.map +1 -1
- package/dist/runtime/query/iteration/__tests__/settle-grace.test.d.ts +2 -0
- package/dist/runtime/query/iteration/__tests__/settle-grace.test.d.ts.map +1 -0
- package/dist/runtime/query/iteration/__tests__/settle-grace.test.js +226 -0
- package/dist/runtime/query/iteration/__tests__/settle-grace.test.js.map +1 -0
- package/dist/runtime/query/iteration/index.d.ts +92 -0
- package/dist/runtime/query/iteration/index.d.ts.map +1 -1
- package/dist/runtime/query/iteration/index.js +818 -565
- package/dist/runtime/query/iteration/index.js.map +1 -1
- package/dist/runtime/query/iteration/phases/__tests__/compaction-declined.test.d.ts +2 -0
- package/dist/runtime/query/iteration/phases/__tests__/compaction-declined.test.d.ts.map +1 -0
- package/dist/runtime/query/iteration/phases/__tests__/compaction-declined.test.js +95 -0
- package/dist/runtime/query/iteration/phases/__tests__/compaction-declined.test.js.map +1 -0
- package/dist/runtime/query/iteration/phases/compaction.d.ts +34 -0
- package/dist/runtime/query/iteration/phases/compaction.d.ts.map +1 -1
- package/dist/runtime/query/iteration/phases/compaction.js +61 -4
- package/dist/runtime/query/iteration/phases/compaction.js.map +1 -1
- package/dist/telemetry/__tests__/model-call-span.test.js +22 -4
- package/dist/telemetry/__tests__/model-call-span.test.js.map +1 -1
- package/dist/telemetry/__tests__/span-closure.test.js +12 -5
- package/dist/telemetry/__tests__/span-closure.test.js.map +1 -1
- package/dist/tools/__tests__/untrusted-envelope.test.js +16 -0
- package/dist/tools/__tests__/untrusted-envelope.test.js.map +1 -1
- package/dist/tools/coordinator/__tests__/completion-delivery.test.js +117 -0
- package/dist/tools/coordinator/__tests__/completion-delivery.test.js.map +1 -1
- package/dist/tools/coordinator/__tests__/task-list.test.js +57 -0
- package/dist/tools/coordinator/__tests__/task-list.test.js.map +1 -1
- package/dist/tools/coordinator/__tests__/wait-with-idle-bound.test.d.ts +2 -0
- package/dist/tools/coordinator/__tests__/wait-with-idle-bound.test.d.ts.map +1 -0
- package/dist/tools/coordinator/__tests__/wait-with-idle-bound.test.js +193 -0
- package/dist/tools/coordinator/__tests__/wait-with-idle-bound.test.js.map +1 -0
- package/dist/tools/coordinator/index.d.ts +19 -0
- package/dist/tools/coordinator/index.d.ts.map +1 -1
- package/dist/tools/coordinator/index.js +191 -71
- package/dist/tools/coordinator/index.js.map +1 -1
- package/dist/tools/coordinator/wait-with-idle-bound.d.ts +66 -0
- package/dist/tools/coordinator/wait-with-idle-bound.d.ts.map +1 -0
- package/dist/tools/coordinator/wait-with-idle-bound.js +78 -0
- package/dist/tools/coordinator/wait-with-idle-bound.js.map +1 -0
- package/dist/tools/untrusted-envelope.d.ts.map +1 -1
- package/dist/tools/untrusted-envelope.js +9 -1
- package/dist/tools/untrusted-envelope.js.map +1 -1
- package/dist/types/agent/base.d.ts +16 -0
- package/dist/types/agent/base.d.ts.map +1 -1
- package/dist/types/agent/gateway.d.ts +41 -0
- package/dist/types/agent/gateway.d.ts.map +1 -1
- package/dist/types/agent/lifecycle-event.d.ts +9 -1
- package/dist/types/agent/lifecycle-event.d.ts.map +1 -1
- package/dist/types/agent/task.d.ts +5 -0
- package/dist/types/agent/task.d.ts.map +1 -1
- package/dist/types/hitl/index.d.ts +10 -0
- package/dist/types/hitl/index.d.ts.map +1 -1
- package/dist/types/hitl/index.js.map +1 -1
- package/dist/types/probe/registry.d.ts +6 -0
- package/dist/types/probe/registry.d.ts.map +1 -1
- package/dist/types/provider/interface.d.ts +35 -0
- package/dist/types/provider/interface.d.ts.map +1 -1
- package/dist/types/run/config.d.ts +25 -0
- package/dist/types/run/config.d.ts.map +1 -1
- package/dist/types/run/entity.d.ts +16 -0
- package/dist/types/run/entity.d.ts.map +1 -1
- package/dist/types/run/events.d.ts +75 -0
- package/dist/types/run/events.d.ts.map +1 -1
- package/dist/types/run/events.js.map +1 -1
- package/dist/types/run/prepare-step.d.ts +17 -2
- package/dist/types/run/prepare-step.d.ts.map +1 -1
- package/dist/types/verification/index.d.ts +98 -0
- package/dist/types/verification/index.d.ts.map +1 -1
- package/dist/types/verification/index.js +10 -0
- package/dist/types/verification/index.js.map +1 -1
- package/dist/utils/__tests__/abort-reason.test.d.ts +2 -0
- package/dist/utils/__tests__/abort-reason.test.d.ts.map +1 -0
- package/dist/utils/__tests__/abort-reason.test.js +48 -0
- package/dist/utils/__tests__/abort-reason.test.js.map +1 -0
- package/dist/utils/abort.d.ts +26 -0
- package/dist/utils/abort.d.ts.map +1 -1
- package/dist/utils/abort.js +34 -0
- package/dist/utils/abort.js.map +1 -1
- package/dist/verification/__tests__/argument-pattern.test.d.ts +2 -0
- package/dist/verification/__tests__/argument-pattern.test.d.ts.map +1 -0
- package/dist/verification/__tests__/argument-pattern.test.js +122 -0
- package/dist/verification/__tests__/argument-pattern.test.js.map +1 -0
- package/dist/verification/__tests__/rule-order-and-reason.test.d.ts +2 -0
- package/dist/verification/__tests__/rule-order-and-reason.test.d.ts.map +1 -0
- package/dist/verification/__tests__/rule-order-and-reason.test.js +126 -0
- package/dist/verification/__tests__/rule-order-and-reason.test.js.map +1 -0
- package/dist/verification/gate.d.ts +17 -1
- package/dist/verification/gate.d.ts.map +1 -1
- package/dist/verification/gate.js +102 -2
- package/dist/verification/gate.js.map +1 -1
- package/dist/verification/index.d.ts +1 -1
- package/dist/verification/index.d.ts.map +1 -1
- package/dist/verification/index.js +1 -1
- package/dist/verification/index.js.map +1 -1
- package/dist/verification/rules.d.ts.map +1 -1
- package/dist/verification/rules.js +27 -0
- package/dist/verification/rules.js.map +1 -1
- package/package.json +1 -1
- package/src/agents/ReactiveAgent.ts +5 -0
- package/src/agents/SupervisorAgent.ts +175 -162
- package/src/agents/__tests__/supervisor-inbox-scope.test.ts +149 -0
- package/src/agents/runAgent.ts +22 -1
- package/src/bridge/a2a/mapper.ts +4 -0
- package/src/bridge/sse/mapper.ts +25 -0
- package/src/connector/mcp/__tests__/positional-arrays.test.ts +183 -0
- package/src/connector/mcp/adapter.ts +131 -7
- package/src/constants/agent/index.ts +5 -0
- package/src/constants/plugin/index.ts +15 -0
- package/src/contracts/api.ts +1 -0
- package/src/gateway/__tests__/completion-inbox.test.ts +348 -2
- package/src/gateway/completion-inbox.ts +248 -16
- package/src/gateway/local.ts +26 -1
- package/src/manager/agent/lifecycle.ts +6 -0
- package/src/manager/run/persistence.ts +12 -0
- package/src/provider/thinking-support.ts +19 -2
- package/src/public-runtime.ts +9 -0
- package/src/run/reporter.ts +12 -0
- package/src/runtime/query/__tests__/completion-does-not-erase-the-answer.test.ts +163 -0
- package/src/runtime/query/__tests__/completion-notification.test.ts +486 -34
- package/src/runtime/query/__tests__/context-size-on-the-wire.test.ts +122 -0
- package/src/runtime/query/__tests__/context.test.ts +24 -0
- package/src/runtime/query/__tests__/effort-reaches-the-wire.test.ts +135 -0
- package/src/runtime/query/__tests__/tool-timeout.test.ts +38 -0
- package/src/runtime/query/context.ts +16 -1
- package/src/runtime/query/executor.ts +11 -1
- package/src/runtime/query/guard.ts +32 -0
- package/src/runtime/query/iteration/__tests__/settle-grace.test.ts +265 -0
- package/src/runtime/query/iteration/index.ts +906 -635
- package/src/runtime/query/iteration/phases/__tests__/compaction-declined.test.ts +124 -0
- package/src/runtime/query/iteration/phases/compaction.ts +83 -10
- package/src/telemetry/__tests__/model-call-span.test.ts +22 -5
- package/src/telemetry/__tests__/span-closure.test.ts +12 -5
- package/src/tools/__tests__/untrusted-envelope.test.ts +23 -0
- package/src/tools/coordinator/__tests__/completion-delivery.test.ts +147 -0
- package/src/tools/coordinator/__tests__/task-list.test.ts +72 -0
- package/src/tools/coordinator/__tests__/wait-with-idle-bound.test.ts +247 -0
- package/src/tools/coordinator/index.ts +205 -78
- package/src/tools/coordinator/wait-with-idle-bound.ts +142 -0
- package/src/tools/untrusted-envelope.ts +9 -1
- package/src/types/agent/base.ts +17 -0
- package/src/types/agent/gateway.ts +42 -0
- package/src/types/agent/lifecycle-event.ts +7 -0
- package/src/types/agent/task.ts +5 -0
- package/src/types/hitl/index.ts +10 -0
- package/src/types/probe/registry.ts +6 -0
- package/src/types/provider/interface.ts +39 -0
- package/src/types/run/config.ts +26 -0
- package/src/types/run/entity.ts +17 -0
- package/src/types/run/events.ts +75 -0
- package/src/types/run/prepare-step.ts +17 -2
- package/src/types/verification/index.ts +61 -0
- package/src/utils/__tests__/abort-reason.test.ts +56 -0
- package/src/utils/abort.ts +34 -0
- package/src/verification/__tests__/argument-pattern.test.ts +158 -0
- package/src/verification/__tests__/rule-order-and-reason.test.ts +149 -0
- package/src/verification/gate.ts +106 -3
- package/src/verification/index.ts +1 -1
- package/src/verification/rules.ts +28 -0
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,682 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 8.0.0
|
|
4
|
+
|
|
5
|
+
### Major Changes
|
|
6
|
+
|
|
7
|
+
- 9ac8dd4: A delegate's output is framed as untrusted material on every path the model reads it, and it can no longer end the frame early
|
|
8
|
+
|
|
9
|
+
Blocking `create_task` and `wait_for_task` wrap a worker's text in the
|
|
10
|
+
`<namzu-untrusted>` envelope. Two other paths carried the same bytes and did
|
|
11
|
+
not: the completion notification injected into the transcript, and
|
|
12
|
+
`agent_task_list`'s rendered output. So whether a worker's words arrived as
|
|
13
|
+
material or as the parent's own reasoning depended on how the model happened to
|
|
14
|
+
fetch them — and the two unframed paths are the ones reached when a wait was
|
|
15
|
+
abandoned, which is when a run is already off its expected course.
|
|
16
|
+
|
|
17
|
+
Worse, the notification's own delimiter was forgeable. Measured: worker output
|
|
18
|
+
containing `</task-notification>` produced two closing tags in one message, with
|
|
19
|
+
attacker-controlled text sitting outside the first — reading as ordinary
|
|
20
|
+
transcript rather than as a delegate's material.
|
|
21
|
+
|
|
22
|
+
**What changed on the wire the model sees.**
|
|
23
|
+
|
|
24
|
+
- The notification now nests a `<namzu-untrusted kind="agent-result">` block
|
|
25
|
+
inside `<task-notification>`. Kernel metadata (`task_id`, `agent`, `state`,
|
|
26
|
+
`duration_ms`) stays OUTSIDE it — framing this kernel's own statements as
|
|
27
|
+
untrusted would tell the model to discount the only part of the message it
|
|
28
|
+
can rely on — and so does the truncation notice, which is an instruction
|
|
29
|
+
about how to fetch the rest.
|
|
30
|
+
- `agent_task_list` wraps each finished task's output the same way, with the
|
|
31
|
+
same `agent` and `task` attributes the blocking path uses.
|
|
32
|
+
- Both delimiters are defanged inside worker text, case-insensitively. The
|
|
33
|
+
replacements (`task_notification`, `namzu_untrusted`) share no substring with
|
|
34
|
+
the tokens they replace — a replacement that still contains the token is found
|
|
35
|
+
again by a second pass or by any looser matcher downstream.
|
|
36
|
+
- A notification is 257 characters longer than before — measured, both for a
|
|
37
|
+
five-character result and for a truncated 4 kB one, so the cost is fixed
|
|
38
|
+
rather than proportional to the output. It grows only with the length of the
|
|
39
|
+
agent id and task id, which appear in the envelope's attributes.
|
|
40
|
+
|
|
41
|
+
`data.result` on both tools is unchanged, so a host reading results
|
|
42
|
+
programmatically is unaffected. If you match on the model-facing text of either
|
|
43
|
+
tool, expect the envelope.
|
|
44
|
+
|
|
45
|
+
- 9ac8dd4: A completion inbox hears only about the tasks its own run launched, and a supervisor releases the gateway it borrowed
|
|
46
|
+
|
|
47
|
+
`TaskGateway.onTaskCompleted` is a broadcast and `TaskHandle` carries no run id,
|
|
48
|
+
so every inbox attached to a gateway was handed every completion on it.
|
|
49
|
+
Measured: two inboxes on one gateway, one run launches a task, and the OTHER
|
|
50
|
+
run drains it — it would have been told "a task you launched has finished", a
|
|
51
|
+
false statement, over another run's worker output. A shared gateway is not an
|
|
52
|
+
abuse of the API: `SupervisorAgentConfig.gateway` takes one, and a host that
|
|
53
|
+
owns a gateway reuses it.
|
|
54
|
+
|
|
55
|
+
Separately, nothing ever called `CompletionInbox.close()`. Three sequential
|
|
56
|
+
`SupervisorAgent` runs against one host gateway left three live subscriptions,
|
|
57
|
+
each still holding its run's handles, and the set only grew.
|
|
58
|
+
|
|
59
|
+
**Breaking, and what to do.**
|
|
60
|
+
|
|
61
|
+
- `CompletionInbox` now ignores a completion for a task it was not told about.
|
|
62
|
+
If you drive `buildCoordinatorTools` there is nothing to do — `create_task`
|
|
63
|
+
declares every launch, blocking and background alike. If you launch tasks
|
|
64
|
+
some other way and expect notifications, call `inbox.launched(taskId)` after
|
|
65
|
+
the launch. `inbox.expect(taskId)` already implies it.
|
|
66
|
+
- `SupervisorAgent` closes the inbox it created when the run ends, including
|
|
67
|
+
when setup throws. An inbox you construct yourself is still yours to close.
|
|
68
|
+
- `close()` now clears what the inbox owned and claimed as well as what it
|
|
69
|
+
queued, so a closed inbox cannot be re-armed through a stale reference.
|
|
70
|
+
|
|
71
|
+
The ordering that would otherwise turn this into lost results is handled in
|
|
72
|
+
two layers. `gateway.createTask` resolves one microtask before its caller can
|
|
73
|
+
say who owns the task, so a worker that finishes inside that window is
|
|
74
|
+
announced first. An unowned announcement is therefore BUFFERED rather than
|
|
75
|
+
dropped, and ownership may be claimed retroactively; the buffer is bounded at
|
|
76
|
+
32 entries so that on a shared gateway it cannot accumulate every other run's
|
|
77
|
+
worker output, and an eviction is logged at WARN so a dropped completion is
|
|
78
|
+
never inferable only from an absence. Where the buffer could not hold an entry,
|
|
79
|
+
`launched()` also asks `gateway.getTask` — an assumption that a just-settled
|
|
80
|
+
task is still findable, now stated on `TaskGateway.getTask` itself so a host
|
|
81
|
+
that cannot meet it knows it is the one paying.
|
|
82
|
+
|
|
83
|
+
- 9ac8dd4: `create_task` offers `background: true` only when there is somewhere for the result to arrive
|
|
84
|
+
|
|
85
|
+
A background launch returns a task id and tells the model its result will come
|
|
86
|
+
"later, as a task notification". The `CompletionInbox` is the only thing that
|
|
87
|
+
delivers one — it holds the run open for the outstanding worker and puts the
|
|
88
|
+
completion into the transcript. `buildCoordinatorTools` mounted the parameter
|
|
89
|
+
whether or not it was given an inbox, so a host without one had a tool
|
|
90
|
+
advertising a channel that did not exist. Nothing failed loudly, because the
|
|
91
|
+
launch itself succeeded; the result simply never arrived.
|
|
92
|
+
|
|
93
|
+
Without a `completionInbox`, `create_task` no longer declares `background` and
|
|
94
|
+
its description no longer mentions it. Everything else is unchanged: the
|
|
95
|
+
blocking path, `wait_for_task`, `cancel_task` and `agent_task_list` are all
|
|
96
|
+
still mounted. Pass a `completionInbox` — to `buildCoordinatorTools` **and** to
|
|
97
|
+
`drainQuery` — to get background launching back. `SupervisorAgent` does both
|
|
98
|
+
already, so a host using it sees no change.
|
|
99
|
+
|
|
100
|
+
A `background: true` that reaches `execute` some other way — a directly
|
|
101
|
+
constructed definition — is REFUSED, naming the missing piece, rather than
|
|
102
|
+
quietly turned into a blocking call: the caller asked for something that
|
|
103
|
+
returns immediately, and giving them a different thing is accepting work whose
|
|
104
|
+
stated terms cannot be met. The abandoned-wait messages on `create_task` and
|
|
105
|
+
`wait_for_task` no longer promise a notification either, and
|
|
106
|
+
`agent_task_list` stops telling the model to avoid the listing when the
|
|
107
|
+
listing is the only route left.
|
|
108
|
+
|
|
109
|
+
The parameter is withheld rather than refused per call, and rather than thrown
|
|
110
|
+
at construction. A parameter the model is never shown costs nothing; one it is shown and then
|
|
111
|
+
denied costs prompt-prefix tokens plus an iteration per attempt. And a throw
|
|
112
|
+
would break a caller doing something legitimate — an inbox-less coordinator
|
|
113
|
+
surface is a supported configuration. This is the same reasoning that made an
|
|
114
|
+
empty roster withhold `create_task` rather than refuse to build.
|
|
115
|
+
|
|
116
|
+
### Minor Changes
|
|
117
|
+
|
|
118
|
+
- a39c2ed: A compaction pass now reports both of its outcomes.
|
|
119
|
+
|
|
120
|
+
Two gaps, in opposite directions, in the same function.
|
|
121
|
+
|
|
122
|
+
**A compaction that sheds nothing was invisible to everyone.** All three decline
|
|
123
|
+
paths — the reducer throws, it returns no fewer messages than it was given, or
|
|
124
|
+
its result splits a `tool_use` from its `tool_result` and is refused wholesale —
|
|
125
|
+
reached a log line and stopped there. A host that silences its logger, which
|
|
126
|
+
every command-line entry point does, made a failed compaction invisible to the
|
|
127
|
+
user, to the host _and_ to the model at once. The run then continued at full
|
|
128
|
+
context toward a provider rejection several turns later that named none of this.
|
|
129
|
+
A shed that did not happen is exactly as consequential as one that did, and only
|
|
130
|
+
one of them was on the wire.
|
|
131
|
+
|
|
132
|
+
New `compaction_failed` event (wire: `compaction.failed`) carrying `cause`
|
|
133
|
+
(`reducer_threw` | `shed_nothing` | `split_tool_pair`), the unchanged message
|
|
134
|
+
count, and the reducer's error where there was one. The cause is on the event
|
|
135
|
+
because the three want different responses: one may succeed next pass, one will
|
|
136
|
+
decline identically every time, and one is a reducer bug that `findSafeTrimIndex`
|
|
137
|
+
exists to prevent.
|
|
138
|
+
|
|
139
|
+
**And a compaction that succeeded was invisible on the path most hosts take.**
|
|
140
|
+
`compaction_completed` was emitted only from the structured working-state path.
|
|
141
|
+
The reducer path — taken by any host-supplied `contextReducer` and by
|
|
142
|
+
`strategy: 'sliding-window'` — emitted nothing at all, so the event whose own
|
|
143
|
+
documentation says it exists because "a host could not show the user that context
|
|
144
|
+
was dropped" never reached the hosts most likely to need it. It is emitted from
|
|
145
|
+
both paths now.
|
|
146
|
+
|
|
147
|
+
That second one was found by a test written for the first: asserting that a
|
|
148
|
+
successful compaction does _not_ report a failure is what showed it reported
|
|
149
|
+
nothing.
|
|
150
|
+
|
|
151
|
+
**If you switch exhaustively over `RunEvent`, you need a case for
|
|
152
|
+
`compaction_failed`.** Nothing else changes: no existing event's shape moved, and
|
|
153
|
+
a host that ignores unknown events is unaffected. The A2A bridge deliberately
|
|
154
|
+
does not forward either compaction event — a peer models a task lifecycle and
|
|
155
|
+
cannot act on how this runtime manages its own context.
|
|
156
|
+
|
|
157
|
+
- f6e0594: `token_usage_updated` now carries the current context size and the window it is measured against.
|
|
158
|
+
|
|
159
|
+
A host built a context indicator, and it could not have been right. The event
|
|
160
|
+
carried `usage` — **cumulative run spend**, summed over every turn, monotonically
|
|
161
|
+
increasing and untouched by compaction — and nothing about the size of the
|
|
162
|
+
conversation being sent. So the host divided cumulative spend by a context
|
|
163
|
+
window guessed from a substring of the model name, and rendered the result as an
|
|
164
|
+
unqualified percentage, continuously.
|
|
165
|
+
|
|
166
|
+
Both terms were wrong, and the numerator was the worse of the two. A guessed
|
|
167
|
+
window is wrong by a bounded factor. Cumulative spend has three properties that
|
|
168
|
+
make it not merely imprecise but actively misleading:
|
|
169
|
+
|
|
170
|
+
- It **never decreases**, by explicit design — the accumulator is documented as
|
|
171
|
+
monotone so it can never under-report a bill. Compaction does not reduce it.
|
|
172
|
+
- It grows **superlinearly in turn count**, because every turn re-sends the whole
|
|
173
|
+
history and counts those prompt tokens again. Ten turns over a 50k context
|
|
174
|
+
accumulate roughly 500k.
|
|
175
|
+
- It measures **spend**, which is the right quantity for cost and the wrong one
|
|
176
|
+
for occupancy.
|
|
177
|
+
|
|
178
|
+
So an indicator built on it saturates at full long before the context is, and it
|
|
179
|
+
is **anti-correlated with what it claims in exactly the regime a user cares
|
|
180
|
+
about**: a long conversation reads FULL while the real context may be a fifth of
|
|
181
|
+
the window. That alarms people into compacting or restarting when they have
|
|
182
|
+
room — worse than showing nothing, because silence does not tell you something
|
|
183
|
+
false in red. In the other direction, a driver that reports no usage shows 0%
|
|
184
|
+
for a conversation that is really there.
|
|
185
|
+
|
|
186
|
+
The kernel already computed the right numbers on every iteration and kept them
|
|
187
|
+
to itself. `measureContext()` is now exported, and the event carries four new
|
|
188
|
+
optional fields: `contextTokens`, `contextMeasuredBy` (`'provider' | 'estimate'`),
|
|
189
|
+
`contextWindowTokens` and `windowSource` (`'config' | 'model-table' | 'default'`).
|
|
190
|
+
They are named apart from the cumulative figures beside them deliberately —
|
|
191
|
+
reaching for the wrong one should be a visible mistake, not a plausible guess.
|
|
192
|
+
|
|
193
|
+
**They are absent when the run has no compaction configuration**, because nothing
|
|
194
|
+
then resolves a window and inventing one would be the guess this replaces. A
|
|
195
|
+
surface should show what it can name rather than a fraction it cannot ground.
|
|
196
|
+
|
|
197
|
+
**A fraction is only as honest as the weaker of its terms.** `contextMeasuredBy`
|
|
198
|
+
and `windowSource` exist so a surface can pass that on rather than presenting an
|
|
199
|
+
estimate as a measurement. Nothing existing changes: `usage` and `cost` are
|
|
200
|
+
untouched, and the new fields are additive and optional.
|
|
201
|
+
|
|
202
|
+
- a39c2ed: A verification rule can name one tool and one argument.
|
|
203
|
+
|
|
204
|
+
Every pattern rule an operator could write was one of two wrong things.
|
|
205
|
+
|
|
206
|
+
`custom_pattern` carries no tool scope, so a rule written about `bash` decided
|
|
207
|
+
`edit` calls as well — `target: 'both'` prefixes the tool name to the subject
|
|
208
|
+
rather than requiring it, which is not a scope. And `target: 'args'` tests
|
|
209
|
+
`JSON.stringify(toolInput)`, so the subject is the JSON _text_ of the whole
|
|
210
|
+
argument object: the natural, anchored thing to write, `^git push`, is tested
|
|
211
|
+
against `{"command":"git push origin main"}` and can never match. The rule then
|
|
212
|
+
decides nothing, silently. Pinning the tool cost the anchor; anchoring cost the
|
|
213
|
+
tool scope.
|
|
214
|
+
|
|
215
|
+
New `argument_pattern` rule — `toolNames`, `argument`, `pattern`, `decision` —
|
|
216
|
+
whose subject is the named argument's own value, so an anchored pattern means
|
|
217
|
+
what it looks like it means. The refusal names the argument as well as the
|
|
218
|
+
pattern, which is what tells a model whether a different value could get through.
|
|
219
|
+
|
|
220
|
+
It deliberately decides nothing in three cases: the tool was not called, the
|
|
221
|
+
argument is absent, or the argument holds an object or an array. No string a
|
|
222
|
+
pattern could match says anything true about a structured value, and serialising
|
|
223
|
+
one to try would put this rule back where `custom_pattern` already is. To refuse
|
|
224
|
+
a tool over the _shape_ of its input, deny it by name. Numbers and booleans are
|
|
225
|
+
matched rather than skipped — they render unambiguously, and a rule about a
|
|
226
|
+
numeric argument is a reasonable thing to write.
|
|
227
|
+
|
|
228
|
+
`custom_pattern` is unchanged and not deprecated: matching anywhere in the
|
|
229
|
+
serialised input without caring where is a real use, and it is now documented as
|
|
230
|
+
being that rather than reading as something it never was. The trap was the name,
|
|
231
|
+
not the behaviour.
|
|
232
|
+
|
|
233
|
+
- 9ac8dd4: A run that ends any way other than a plain final answer no longer throws away a finished worker's output
|
|
234
|
+
|
|
235
|
+
The iteration loop consulted its completion inbox at exactly one place: the
|
|
236
|
+
branch where the model stops calling tools and answers. It leaves by eight other
|
|
237
|
+
routes, and three of them are ordinary ways for a run to END — a tool the author
|
|
238
|
+
marked `terminal`, a captured `structured_output`, and the host's `stopWhen`.
|
|
239
|
+
A background or abandoned worker that finished while any of those was deciding
|
|
240
|
+
had its result dropped: the gateway held it, the run closed, nothing read it.
|
|
241
|
+
Measured before the fix — terminal-tool exit and `stopWhen` exit both delivered
|
|
242
|
+
nothing; the final-answer exit delivered in 44 ms.
|
|
243
|
+
|
|
244
|
+
Delivery now happens in a `finally` around the loop, so it does not depend on
|
|
245
|
+
each exit remembering — including the two `return`s and a generator abandoned by
|
|
246
|
+
its consumer, which no post-loop statement reaches.
|
|
247
|
+
|
|
248
|
+
**What you may observe.** On those exits `Run.messages` can now end with a
|
|
249
|
+
`task-notification` user message after the assistant's last message. The answer
|
|
250
|
+
is on `Run.result`, as before. If you were reading the answer off the last
|
|
251
|
+
element of `Run.messages`, that assumption was already unsafe whenever a
|
|
252
|
+
notification landed mid-run; it is now unsafe in three more places.
|
|
253
|
+
|
|
254
|
+
**Which exits wait, and which only deliver.** A hold buys the model a turn in
|
|
255
|
+
which to use a result, so it is only worth paying where a turn can still
|
|
256
|
+
happen. A terminal tool and a captured `structured_output` have decided the
|
|
257
|
+
answer, so those deliver what arrived and stop. `stopWhen` is a programmable
|
|
258
|
+
halt that says nothing about whether the answer is complete, so it now HOLDS
|
|
259
|
+
like the ordinary final-answer exit — a precedence rule chosen here, not
|
|
260
|
+
something `stopWhen` implies — and costs exactly one extra turn, after which
|
|
261
|
+
the predicate fires again with nothing pending.
|
|
262
|
+
|
|
263
|
+
The stop reason survives that extra turn. `stopWhen` is consulted only after a
|
|
264
|
+
tool batch, so when the extra turn is prose the predicate is never asked again
|
|
265
|
+
and the run leaves by the ordinary route — which would have reported
|
|
266
|
+
`stopReason: 'end_turn'`, naming the shape of the last message rather than the
|
|
267
|
+
host's decision. A run that ends because a host said stop now reports
|
|
268
|
+
`'stop_condition'` whether or not a delegated result delayed it by a turn. If
|
|
269
|
+
the extra turn instead runs more tools, the predicate is asked again and
|
|
270
|
+
answers for itself.
|
|
271
|
+
|
|
272
|
+
A run that ends with a worker still running now says so on
|
|
273
|
+
`Run.abandonedTaskIds` rather than leaving the impression the result arrived.
|
|
274
|
+
|
|
275
|
+
- 9ac8dd4: A run that ends over a still-running worker says so, and the untrusted envelope's own label can no longer close it
|
|
276
|
+
|
|
277
|
+
Three things an adversarial review of the completion path found.
|
|
278
|
+
|
|
279
|
+
**`Run.abandonedTaskIds`.** A run can settle while a worker it launched is
|
|
280
|
+
still going — the model answered, a terminal tool decided the result, a
|
|
281
|
+
`stopWhen` fired. Until now nothing said so, which left the impression the
|
|
282
|
+
worker's result had been delivered. The run now names those task ids.
|
|
283
|
+
|
|
284
|
+
They are **named, not cancelled**, and that is the decision: giving up on a
|
|
285
|
+
wait is a statement about the waiter, not about the work — the rule this
|
|
286
|
+
subsystem already applies to `wait_for_task` — and "the parent answered early"
|
|
287
|
+
is a weaker warrant for killing a child than "the clock ran out", not a
|
|
288
|
+
stronger one. A worker mid-write is not the kernel's to judge. A host that
|
|
289
|
+
wants the work stopped has `cancel_task` and the run's abort controller, and
|
|
290
|
+
now has the ids to use them on.
|
|
291
|
+
|
|
292
|
+
**`wrapUntrusted` neutralises its own delimiter inside `provenance`.** The body
|
|
293
|
+
was defanged and the attributes escaped; the provenance line was interpolated
|
|
294
|
+
raw, and every caller in the SDK builds it from a value it did not author — an
|
|
295
|
+
agent id from a roster, a server name from a connector manifest. A provenance
|
|
296
|
+
carrying `</namzu-untrusted>` ended the block before the content it was
|
|
297
|
+
introducing. This affects the blocking `create_task`, `wait_for_task` and the
|
|
298
|
+
`Agent` tool as well as the two paths framed in this release.
|
|
299
|
+
|
|
300
|
+
**`background: true` with no inbox is refused, not silently made blocking**,
|
|
301
|
+
and the sentences match. The abandoned-wait messages on `create_task` and
|
|
302
|
+
`wait_for_task` promised "its result will arrive separately as a task
|
|
303
|
+
notification" unconditionally — false with no inbox, and a model told to expect
|
|
304
|
+
a message waits for it. They now say where the result actually is. The
|
|
305
|
+
`agent_task_list` description likewise stops telling the model not to use the
|
|
306
|
+
listing when, without an inbox, the listing is the only route left to an
|
|
307
|
+
abandoned launch's output.
|
|
308
|
+
|
|
309
|
+
`CompletionInbox` gains `outstandingTaskIds`, which reads the ids and cancels
|
|
310
|
+
nothing.
|
|
311
|
+
|
|
312
|
+
- 9ac8dd4: A run holding for a background worker waits a share of its own budget, not a fixed two minutes
|
|
313
|
+
|
|
314
|
+
`BACKGROUND_TASK_GRACE_MS = 120_000` was unrelated to the run it bounded, and
|
|
315
|
+
wrong in both directions at once. Measured: a run configured `timeoutMs: 20_000`
|
|
316
|
+
was held open for **120,267 ms** — six times its own budget — because the hold
|
|
317
|
+
sits inside an iteration and the run guard only checks between them, so nothing
|
|
318
|
+
could interrupt it. In the other direction, on a run with hours left the same
|
|
319
|
+
two minutes abandoned delegated workers observed at 4m21s, 5m58s and 8m04s, all
|
|
320
|
+
comfortably inside the hour `DELEGATION_TIMEOUT_MS` already declares.
|
|
321
|
+
|
|
322
|
+
The hold is now `min(remainingBeforeFinalize × 0.5, DELEGATION_TIMEOUT_MS)`,
|
|
323
|
+
where `remainingBeforeFinalize` is the time left before the run guard stops
|
|
324
|
+
asking for more work and asks for a closing summary (90% of `timeoutMs`), less
|
|
325
|
+
what the run has spent — carried across a resume, so a checkpointed run sizes
|
|
326
|
+
the hold from what is left of the RUN rather than of the process now hosting
|
|
327
|
+
it, and read when the wait starts rather than at the top of the iteration.
|
|
328
|
+
|
|
329
|
+
- **Half, not all.** The hold exists to put a worker's result where the model
|
|
330
|
+
can read it, and reading it costs a turn. Spending everything remaining would
|
|
331
|
+
deliver a notification into a run with no turn left to act on it — the same
|
|
332
|
+
failure the mechanism exists to prevent.
|
|
333
|
+
- **Bounded against the boundary that binds.** Measuring to the DEADLINE was
|
|
334
|
+
the first attempt and it looked safe: a hold cannot outlive the deadline
|
|
335
|
+
either way. But half of the time-to-deadline, started just under the warning
|
|
336
|
+
threshold, ends at 95% of the budget — so the slice the guard keeps for the
|
|
337
|
+
run to produce a closing answer is half spent waiting for the result that
|
|
338
|
+
answer was supposed to use. Against the finalize point the hold cannot reach
|
|
339
|
+
that reserve at all, which is what makes the guard's inability to interrupt
|
|
340
|
+
a hold a non-issue rather than a smaller issue.
|
|
341
|
+
- **A floor of zero, deliberately.** A run with no time left before it must
|
|
342
|
+
start finishing has no turn in which to read a notification. Nothing is
|
|
343
|
+
dropped by it: the wait returns before it looks at its timer when a
|
|
344
|
+
completion is already in hand.
|
|
345
|
+
|
|
346
|
+
**What changes for you.** A run with a short `timeoutMs` finishes when it said
|
|
347
|
+
it would instead of overrunning by minutes. A run with a long one keeps its
|
|
348
|
+
worker instead of abandoning it. If you were relying on a fixed two-minute
|
|
349
|
+
settle regardless of run configuration, set `timeoutMs` to about four and a half minutes to
|
|
350
|
+
get the same hold.
|
|
351
|
+
|
|
352
|
+
- 585a592: A caller can ask which effort levels a model accepts.
|
|
353
|
+
|
|
354
|
+
The answer existed, was modelled carefully, and was reachable only from inside
|
|
355
|
+
one driver. That matters because effort is **refused, not clamped**: a level a
|
|
356
|
+
model does not have makes the vendor reject the request, so a control offering
|
|
357
|
+
the wrong one produces a run that fails at the start rather than a quieter one.
|
|
358
|
+
|
|
359
|
+
Every option open to a caller without the answer was bad. Offering all five
|
|
360
|
+
breaks some models. Offering the intersection hides `xhigh` and `max` from every
|
|
361
|
+
model that has them, which is most of the reason to build such a control. And
|
|
362
|
+
copying the table looks fine and is worst: the ceiling has moved twice already,
|
|
363
|
+
so a copy goes stale on the next model and goes stale **silently**, surfacing as
|
|
364
|
+
a vendor rejection rather than a failing build.
|
|
365
|
+
|
|
366
|
+
**New optional `LLMProvider.effortLevelsFor(model, thinking?)`.** Three states,
|
|
367
|
+
each meaning something different: the method absent means the driver has no
|
|
368
|
+
effort concept at all and setting one will be refused; an empty array means the
|
|
369
|
+
driver implements effort and this model has none; a non-empty array is the set
|
|
370
|
+
to offer.
|
|
371
|
+
|
|
372
|
+
**`thinking` is a parameter, and that is the point.** At least one model family
|
|
373
|
+
accepts a narrower set while thinking is disabled than while it is on — so an
|
|
374
|
+
API returning two sibling arrays invites a caller to render a picker from one
|
|
375
|
+
and send the other, a combination the vendor rejects, on exactly one family.
|
|
376
|
+
Passing the configuration you will actually send makes that unspellable: there
|
|
377
|
+
is one answer and it is the one for your request.
|
|
378
|
+
|
|
379
|
+
The driver's implementation shares the same two resolution steps the request
|
|
380
|
+
path uses, so a caller's picker and the request it produces cannot disagree.
|
|
381
|
+
|
|
382
|
+
`@namzu/anthropic` also now exports `resolveThinkingCapability`,
|
|
383
|
+
`resolveThinkingBody`, `resolveEffort` and their types, for a caller that needs
|
|
384
|
+
the fuller picture — whether thinking can be switched off at all, not only which
|
|
385
|
+
effort levels apply. Prefer `effortLevelsFor` where it suffices: it is
|
|
386
|
+
provider-agnostic and cannot return the wrong one of the two sets.
|
|
387
|
+
|
|
388
|
+
Separately, the live wire-contract suite now retries a transient status rather
|
|
389
|
+
than reporting it as a contract failure. A 529 says the service is busy and
|
|
390
|
+
answers nothing about whether a schema is expressible — so a test named "every
|
|
391
|
+
shipped tool is expressible on this wire" was claiming something the run had not
|
|
392
|
+
established. That cost two manual re-runs in one day to discover the wire had no
|
|
393
|
+
opinion.
|
|
394
|
+
|
|
395
|
+
### Patch Changes
|
|
396
|
+
|
|
397
|
+
- 9ac8dd4: A background task whose completion arrived early no longer holds the run open forever
|
|
398
|
+
|
|
399
|
+
`CompletionInbox.drain()` handed the completion over and marked it claimed, but
|
|
400
|
+
left the task on the OUTSTANDING set. That set is meant to hold ids that are
|
|
401
|
+
still running, and only the gateway's completion listener takes an id off it —
|
|
402
|
+
so if the listener ran BEFORE the launching call said `expect()`, the id was
|
|
403
|
+
added to a set nothing would ever clear.
|
|
404
|
+
|
|
405
|
+
That order is reachable rather than theoretical: `expect()` runs one microtask
|
|
406
|
+
after `gateway.createTask()` resolves, and a worker that finishes fast is
|
|
407
|
+
announced in between. The result of it was `hasPendingWork === true` for the
|
|
408
|
+
rest of the run, with an empty inbox — so every attempt to settle waited out the
|
|
409
|
+
full background grace period for a result that was already in the transcript,
|
|
410
|
+
and did it again on the next turn, and the next.
|
|
411
|
+
|
|
412
|
+
Nothing to do on upgrade. If you were seeing runs pause for two minutes before
|
|
413
|
+
their final answer with no background work outstanding, this was why.
|
|
414
|
+
|
|
415
|
+
- 3d4315e: `PrepareStepResult.activeTools` documented the opposite of what it does.
|
|
416
|
+
|
|
417
|
+
Its comment promised that unregistered names are dropped so a phase list
|
|
418
|
+
outliving a tool rename would "narrow the surface, not kill the agent mid-run".
|
|
419
|
+
Since the list began bounding what may RUN rather than only what the model is
|
|
420
|
+
shown, dropping every name leaves the step able to call nothing — so the code
|
|
421
|
+
and its own documentation had said different things.
|
|
422
|
+
|
|
423
|
+
**The behaviour is right and the comment was wrong.** This list means "only
|
|
424
|
+
these": when a rename outlives it, the only set satisfying "only the tools that
|
|
425
|
+
no longer exist" is the empty one. Widening back to the run's list would grant
|
|
426
|
+
precisely the tools the caller asked to exclude, on the grounds that their own
|
|
427
|
+
list failed — a control that stops applying because it was aged out, which is
|
|
428
|
+
worse than a step that answers from what it already has. The run continues
|
|
429
|
+
either way; nothing crashes.
|
|
430
|
+
|
|
431
|
+
The warning now distinguishes the two cases, because they have different
|
|
432
|
+
consequences: some names dropped narrows the step, and all of them dropped
|
|
433
|
+
leaves it unable to call anything. "Ignoring them" was accurate for the first
|
|
434
|
+
and misleading for the second.
|
|
435
|
+
|
|
436
|
+
**Worth knowing if you rely on this:** the warning goes to the logger, so a host
|
|
437
|
+
that silences its logger sees a phase quietly stop doing anything. That is a real
|
|
438
|
+
gap and it is named here rather than papered over.
|
|
439
|
+
|
|
440
|
+
## 7.0.0
|
|
441
|
+
|
|
442
|
+
### Major Changes
|
|
443
|
+
|
|
444
|
+
- 062624c: A bridged tool's positional array is no longer flattened to "an array of
|
|
445
|
+
anything".
|
|
446
|
+
|
|
447
|
+
`mcpJsonSchemaToZod` collapsed every positional array — both the draft-07
|
|
448
|
+
spelling (`items` holding a list) and the 2020-12 one (`prefixItems`) — to
|
|
449
|
+
`z.array(z.unknown())`. The schema makes a round trip, server JSON Schema → Zod
|
|
450
|
+
→ JSON Schema on the wire, so what was dropped was dropped from what the MODEL
|
|
451
|
+
is shown: a server that spelled out `[string, number]` had the model told
|
|
452
|
+
nothing about the positions, their types, or their order.
|
|
453
|
+
|
|
454
|
+
**Why this is a major.** Where the server pinned the arity and closed the tail,
|
|
455
|
+
the converted schema is now a tuple, so input that a looser array accepted is
|
|
456
|
+
refused locally. It is only ever refused where the server itself declared it
|
|
457
|
+
invalid — the error moves from the server's response to the local validator —
|
|
458
|
+
but a host driving a bridged tool directly can see a validation failure it did
|
|
459
|
+
not see before, and code branching on the converted type (`instanceof
|
|
460
|
+
z.ZodArray`) will take a different branch. If you relied on the permissive
|
|
461
|
+
shape, the fix is to send what the server's schema declares.
|
|
462
|
+
|
|
463
|
+
**The tuple is deliberately narrow, and that is the whole design.** A rejected
|
|
464
|
+
tool schema fails the entire request rather than degrading one tool, taking down
|
|
465
|
+
every run that offered the toolset — so a faithful conversion the wire will not
|
|
466
|
+
accept is strictly worse than a lossy one it will. A tuple is therefore emitted
|
|
467
|
+
only where the server pinned the arity AND closed the tail, because that renders
|
|
468
|
+
as bounded `prefixItems`, which is the one positional shape measured as
|
|
469
|
+
accepted and the same shape a first-party builtin already ships. Every looser
|
|
470
|
+
positional array keeps today's permissive array and gains its shape in the
|
|
471
|
+
description instead, appended to whatever the server wrote rather than replacing
|
|
472
|
+
it.
|
|
473
|
+
|
|
474
|
+
The inversion worth knowing if you write these schemas: positional members do
|
|
475
|
+
not constrain LENGTH. Without `minItems` a server is permitting a shorter array,
|
|
476
|
+
which a tuple cannot express — so an absent lower bound is a reason to keep the
|
|
477
|
+
loose form, not a detail to round up.
|
|
478
|
+
|
|
479
|
+
**Also fixed, and reachable from any bridged server:** the conversion's depth
|
|
480
|
+
ceiling never fired. `MAX_CONVERSION_DEPTH` was compared against in one branch
|
|
481
|
+
that a pure array or union never reaches, and the counter was not even passed
|
|
482
|
+
down the array path — so a deeply nested schema from a remote tool listing took
|
|
483
|
+
the process down with a stack overflow instead of being left permissive as the
|
|
484
|
+
ceiling's own comment promised.
|
|
485
|
+
|
|
486
|
+
### Minor Changes
|
|
487
|
+
|
|
488
|
+
- bf0999d: a policy rule you wrote is actually consulted, and a refusal says what it said
|
|
489
|
+
|
|
490
|
+
Two defects in the verification gate, found while designing an operator-facing
|
|
491
|
+
permission surface on top of it.
|
|
492
|
+
|
|
493
|
+
**A rule could be silently unreachable.** `allowReadOnlyTools` was expanded
|
|
494
|
+
into a rule ahead of the operator's own, and the gate stops at the first match
|
|
495
|
+
— so a rule like "prompt me before every read" was never consulted while that
|
|
496
|
+
flag was on. Not rejected, not warned about, just never reached. Someone who
|
|
497
|
+
writes a control and is silently ignored gets the worst outcome available: they
|
|
498
|
+
believe it is in force and it is not.
|
|
499
|
+
|
|
500
|
+
The read-only allowance now goes LAST, which makes it what it always was in
|
|
501
|
+
substance — a default for tools nobody wrote a rule about, rather than an
|
|
502
|
+
override of the rules they did write. **The dangerous-pattern denial still goes
|
|
503
|
+
first and still outranks everything**, so an operator rule cannot open what the
|
|
504
|
+
floor closes.
|
|
505
|
+
|
|
506
|
+
**A refusal told the model nothing it could use.** The reason was built as
|
|
507
|
+
`Matched rule: ${rule.type}`, so a denial arrived as _"Blocked by the
|
|
508
|
+
verification gate: Matched rule: deny_by_name"_ — the KIND of rule and nothing
|
|
509
|
+
about it. Not which tool, not which pattern, not whether a different input
|
|
510
|
+
would fare better.
|
|
511
|
+
|
|
512
|
+
That difference is behavioural, not cosmetic. Told only that it was denied, a
|
|
513
|
+
model rewords the same call and tries again, because nothing says the retry is
|
|
514
|
+
pointless. Told that a pattern rule denies `git push*`, or that a by-name
|
|
515
|
+
denial is about the tool rather than the input, it can stop and say so. A
|
|
516
|
+
refusal that cannot be reasoned about produces thrashing; one that can produces
|
|
517
|
+
a route around it.
|
|
518
|
+
|
|
519
|
+
`describeRule` is exported, so a host rendering its own approval UI can show
|
|
520
|
+
the same sentence the model got.
|
|
521
|
+
|
|
522
|
+
- cb772c7: Export `describeRule` alongside `evaluateRule`.
|
|
523
|
+
|
|
524
|
+
`evaluateRule` has been public for some time and answers only whether a rule
|
|
525
|
+
matched. A host driving the rules directly — rather than through
|
|
526
|
+
`VerificationGate` — was left holding a verdict with no words for it, and the
|
|
527
|
+
only way to say anything about a refusal was to switch on the rule's `type`.
|
|
528
|
+
That names the KIND of rule and nothing about what it said: not which tool, not
|
|
529
|
+
which pattern, not whether a different input could ever help.
|
|
530
|
+
|
|
531
|
+
That is the same defect the gate itself carried until its `reason` stopped
|
|
532
|
+
being `Matched rule: <type>`, and it was left open one layer up for anyone
|
|
533
|
+
using the rule primitives without the gate. The two now travel together.
|
|
534
|
+
|
|
535
|
+
Nothing is removed and no behaviour changes. If you were deriving your own
|
|
536
|
+
denial text from `rule.type`, `describeRule(rule)` is the sentence the gate
|
|
537
|
+
uses, and it is worth reading before you keep your own.
|
|
538
|
+
|
|
539
|
+
- 062624c: `effort` can be set on a run — and so, for the first time, can `thinking`.
|
|
540
|
+
|
|
541
|
+
`effort` was on the provider params, exported, and read by a driver that wrote
|
|
542
|
+
it to the wire, and nothing in the kernel ever set it. Every request went out at
|
|
543
|
+
the model's default, which reads as "this model ignores effort" rather than
|
|
544
|
+
"nobody plumbed it through".
|
|
545
|
+
|
|
546
|
+
`AgentRunConfig` gains `effort`, a sibling of `thinking` rather than a field
|
|
547
|
+
inside it — on some models the two are independent controls that apply together,
|
|
548
|
+
and nesting would make that combination unsayable. It is run-level rather than
|
|
549
|
+
per-step because the provider documents that changing effort between requests
|
|
550
|
+
does not preserve a cached prefix, so a value that moves between steps buys a
|
|
551
|
+
different answer shape at the cost of the cache on every step that changes it.
|
|
552
|
+
|
|
553
|
+
**`thinking` turned out to have the same defect, and had shipped with it.** It
|
|
554
|
+
was settable only through `drainQuery`. Every ergonomic entry point — `runAgent`,
|
|
555
|
+
`ReactiveAgent`, `SupervisorAgent`, and the agent manager's bare-config branch —
|
|
556
|
+
builds its run config by hand-listing fields, so a field nobody remembered to add
|
|
557
|
+
is dropped in silence, with no cast to blame and no error to see. A caller could
|
|
558
|
+
set `thinking` on an agent config and get a run that never asked for it. Both
|
|
559
|
+
fields now live on `BaseAgentConfig` and are forwarded by all four.
|
|
560
|
+
|
|
561
|
+
This was found by watching an actual HTTP body from a real run. The unit tests
|
|
562
|
+
passed throughout, because they drive the kernel directly, and the kernel was
|
|
563
|
+
never the half that was broken.
|
|
564
|
+
|
|
565
|
+
**A driver that cannot honour `effort` now refuses rather than dropping it**,
|
|
566
|
+
the rule `thinking` already had. Effort is the worse silence of the two: a
|
|
567
|
+
dropped `thinking` leaves an empty reasoning list someone might notice, while a
|
|
568
|
+
dropped `effort` leaves a perfectly ordinary answer, so a run requested at `max`
|
|
569
|
+
is indistinguishable from one at the default — including in what it cost.
|
|
570
|
+
Nothing existing breaks, because the field could not be set until now.
|
|
571
|
+
|
|
572
|
+
Two driver-side corrections ride along, both verified against the live wire:
|
|
573
|
+
|
|
574
|
+
- The preview model's capability row claimed all five effort levels. It takes
|
|
575
|
+
`max` and not `xhigh`. That model is not reachable from the tenant the live
|
|
576
|
+
suite runs against, so the row is sourced from the reference rather than
|
|
577
|
+
measured — but the pairing itself is now measured, on a model that has it:
|
|
578
|
+
`claude-sonnet-4-6` answers `xhigh` with _"This model does not support effort
|
|
579
|
+
level 'xhigh'. Supported levels: high, low, max, medium"_ and accepts `max`.
|
|
580
|
+
Reading the levels as a ladder, where anything taking the top rung takes the
|
|
581
|
+
one below, is what produced the wrong row.
|
|
582
|
+
- `output_config` is now merged rather than assigned. It is a shared envelope on
|
|
583
|
+
that wire — a structured-output format and a task budget live in it too — so
|
|
584
|
+
assigning meant whoever wired the next one would silently delete effort, or
|
|
585
|
+
have effort delete theirs, depending only on which line ran last.
|
|
586
|
+
|
|
587
|
+
- bf0999d: a delegated worker is bounded by how long it has been quiet, not only by how long it has run
|
|
588
|
+
|
|
589
|
+
`DELEGATION_TIMEOUT_MS` gave the supervisor an hour to wait, which fixed the
|
|
590
|
+
two-minute deadline that made the blocking path structurally unreachable. An
|
|
591
|
+
hour of wall clock is still the wrong quantity to measure: it says nothing
|
|
592
|
+
about whether the worker is doing anything.
|
|
593
|
+
|
|
594
|
+
One number cannot answer both questions. It has to be generous enough for a
|
|
595
|
+
child doing real work, which is exactly what makes it useless as a stall
|
|
596
|
+
detector — so a worker wedged in its second minute held the supervisor for
|
|
597
|
+
another fifty-eight, and a worker making steady progress at minute fifty-nine
|
|
598
|
+
was cut off for being slow rather than for being stuck.
|
|
599
|
+
|
|
600
|
+
There are two clocks now:
|
|
601
|
+
|
|
602
|
+
- **the run bound**, elapsed time, never refreshed, still an hour. For a worker
|
|
603
|
+
that stays busy forever.
|
|
604
|
+
- **the idle bound**, time since the worker last did anything, reset on every
|
|
605
|
+
progress signal. Five minutes, overridable with `NAMZU_DELEGATION_IDLE_MS`.
|
|
606
|
+
For a worker that stopped.
|
|
607
|
+
|
|
608
|
+
Whichever fires first ends the wait, and **the result says which** — "it went
|
|
609
|
+
quiet" and "it ran too long" are different diagnoses that lead to different
|
|
610
|
+
next moves, and the message is what a model acts on.
|
|
611
|
+
|
|
612
|
+
Giving up on the wait does not cancel the worker. The child keeps going and its
|
|
613
|
+
completion still arrives as a task notification, because a wait that ran out is
|
|
614
|
+
a statement about the waiter, not about the work. Losing an eight-minute
|
|
615
|
+
worker's output because a clock expired is the shape of the bug this whole area
|
|
616
|
+
has been unpicking.
|
|
617
|
+
|
|
618
|
+
**`TaskGateway.onTaskProgress` is new and OPTIONAL.** The idle bound needs a
|
|
619
|
+
signal that a task did something, and only a gateway can see it. It is optional
|
|
620
|
+
because hosts implement `TaskGateway` and not all of them can observe their
|
|
621
|
+
children — a gateway without it is bounded by the wall clock alone, exactly as
|
|
622
|
+
before. That degradation is deliberately visible rather than silent: the
|
|
623
|
+
timeout result carries `idleBoundArmed`, and the message says outright that
|
|
624
|
+
this gateway cannot tell a busy worker from a stuck one.
|
|
625
|
+
|
|
626
|
+
- 69d609a: six declarations that drive nothing are marked for removal
|
|
627
|
+
|
|
628
|
+
An audit of the kernel found primitives that are declared, reachable from the
|
|
629
|
+
published typings, and read by no code at all. None is deleted yet — they are
|
|
630
|
+
on the public surface, so they get the deprecation release the repository's own
|
|
631
|
+
policy asks for, and go in the next major.
|
|
632
|
+
|
|
633
|
+
They are worth naming individually, because a dead declaration is not merely
|
|
634
|
+
untidy. Each of these tells a reader something false:
|
|
635
|
+
|
|
636
|
+
- `HOOK_MAX_CONCURRENT` reads as a concurrency cap that is in force. Hooks run
|
|
637
|
+
sequentially and always have, so a reviewer reasons about batching that does
|
|
638
|
+
not happen. Do not "fix" it by batching — ordering is the contract hooks are
|
|
639
|
+
written against.
|
|
640
|
+
- `MAX_RECENT_ACTIVITIES` — no list is trimmed to it.
|
|
641
|
+
- `AgentTask.progress` and the `progress_updated` lifecycle variant are a whole
|
|
642
|
+
reporting channel with **no producer**. A host that switches on the event has
|
|
643
|
+
written a branch that cannot run; one that waits for progress waits forever.
|
|
644
|
+
- `IterationCheckpoint.planStatus` is never set, so a host restoring a
|
|
645
|
+
checkpoint to find out whether the plan was approved gets `undefined` for
|
|
646
|
+
every run — approved or not — and cannot tell the two apart. Ask the plan
|
|
647
|
+
manager.
|
|
648
|
+
- `ProbeOptions.otel` is unimplemented: setting it changes nothing.
|
|
649
|
+
|
|
650
|
+
Each now carries `@deprecated` and a note saying which of "unused",
|
|
651
|
+
"no producer" or "unimplemented" applies, so the next reader does not have to
|
|
652
|
+
re-derive it.
|
|
653
|
+
|
|
654
|
+
### Patch Changes
|
|
655
|
+
|
|
656
|
+
- bf0999d: `continue_task` is deleted rather than left defined and unreachable
|
|
657
|
+
|
|
658
|
+
It was written, documented, and never returned from the coordinator builder —
|
|
659
|
+
so no model could call it. The question was reopened when `background: true`
|
|
660
|
+
made a live task id reachable again, since the reason it was dropped had been
|
|
661
|
+
that a blocking launch leaves every worker terminal before a later turn learns
|
|
662
|
+
its id, and the manager refuses `continue` on a terminal task.
|
|
663
|
+
|
|
664
|
+
Measured instead of assumed, and it fails on the other side. On a LIVE task the
|
|
665
|
+
manager accepts the call and pushes onto `pendingMessages` — and **nothing
|
|
666
|
+
drains that queue during a run**. The codebase already knew: `steering.ts` says
|
|
667
|
+
in as many words that `queueMessage`/`drainMessages` were never read by the
|
|
668
|
+
iteration loop, and `SteeringChannel` exists because of it, delivering guidance
|
|
669
|
+
on a tool result instead — a `tool_use` must be answered by a `tool_result`
|
|
670
|
+
with the same id, so there is no legal slot for a user message mid-batch.
|
|
671
|
+
|
|
672
|
+
So the tool had no state it worked in: terminal tasks refuse it, live tasks
|
|
673
|
+
accept it into a queue nobody reads. Registering it would have handed the model
|
|
674
|
+
a call that silently does nothing, which is worse than an unreachable
|
|
675
|
+
definition — an unreachable one at least cannot be called.
|
|
676
|
+
|
|
677
|
+
If follow-ups on a live worker are wanted, the work is a consumer for the queue
|
|
678
|
+
or a steering channel that reaches a child. Not this tool.
|
|
679
|
+
|
|
3
680
|
## 6.2.0
|
|
4
681
|
|
|
5
682
|
### Minor Changes
|