@namzu/sdk 38.2.1 → 40.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (602) hide show
  1. package/CHANGELOG.md +851 -0
  2. package/dist/advisory/executor.d.ts +10 -1
  3. package/dist/advisory/executor.d.ts.map +1 -1
  4. package/dist/advisory/executor.js +5 -26
  5. package/dist/advisory/executor.js.map +1 -1
  6. package/dist/advisory/history.d.ts +9 -0
  7. package/dist/advisory/history.d.ts.map +1 -0
  8. package/dist/advisory/history.js +120 -0
  9. package/dist/advisory/history.js.map +1 -0
  10. package/dist/advisory/index.d.ts +1 -1
  11. package/dist/advisory/index.d.ts.map +1 -1
  12. package/dist/advisory/index.js.map +1 -1
  13. package/dist/agents/ReactiveAgent.d.ts.map +1 -1
  14. package/dist/agents/ReactiveAgent.js +3 -0
  15. package/dist/agents/ReactiveAgent.js.map +1 -1
  16. package/dist/agents/runAgent.d.ts +10 -0
  17. package/dist/agents/runAgent.d.ts.map +1 -1
  18. package/dist/agents/runAgent.js +3 -0
  19. package/dist/agents/runAgent.js.map +1 -1
  20. package/dist/compaction/manual.d.ts +6 -0
  21. package/dist/compaction/manual.d.ts.map +1 -1
  22. package/dist/compaction/manual.js +21 -2
  23. package/dist/compaction/manual.js.map +1 -1
  24. package/dist/compaction/summary.d.ts.map +1 -1
  25. package/dist/compaction/summary.js +4 -1
  26. package/dist/compaction/summary.js.map +1 -1
  27. package/dist/config/runtime.js +2 -2
  28. package/dist/config/runtime.js.map +1 -1
  29. package/dist/connector/mcp/adapter.d.ts.map +1 -1
  30. package/dist/connector/mcp/adapter.js +20 -6
  31. package/dist/connector/mcp/adapter.js.map +1 -1
  32. package/dist/contracts/schemas.js +1 -1
  33. package/dist/contracts/schemas.js.map +1 -1
  34. package/dist/eval/harness-protection.d.ts +18 -0
  35. package/dist/eval/harness-protection.d.ts.map +1 -0
  36. package/dist/eval/harness-protection.js +58 -0
  37. package/dist/eval/harness-protection.js.map +1 -0
  38. package/dist/eval/harness-verification.d.ts +6 -1
  39. package/dist/eval/harness-verification.d.ts.map +1 -1
  40. package/dist/eval/harness-verification.js +19 -2
  41. package/dist/eval/harness-verification.js.map +1 -1
  42. package/dist/eval/index.d.ts +1 -0
  43. package/dist/eval/index.d.ts.map +1 -1
  44. package/dist/eval/index.js.map +1 -1
  45. package/dist/manager/resident/activity.d.ts +50 -0
  46. package/dist/manager/resident/activity.d.ts.map +1 -0
  47. package/dist/manager/resident/activity.js +125 -0
  48. package/dist/manager/resident/activity.js.map +1 -0
  49. package/dist/manager/resident/agenda.d.ts +13 -1
  50. package/dist/manager/resident/agenda.d.ts.map +1 -1
  51. package/dist/manager/resident/agenda.js +40 -5
  52. package/dist/manager/resident/agenda.js.map +1 -1
  53. package/dist/manager/resident/consumption.d.ts +104 -0
  54. package/dist/manager/resident/consumption.d.ts.map +1 -0
  55. package/dist/manager/resident/consumption.js +233 -0
  56. package/dist/manager/resident/consumption.js.map +1 -0
  57. package/dist/manager/resident/evidence-recall.d.ts +24 -0
  58. package/dist/manager/resident/evidence-recall.d.ts.map +1 -0
  59. package/dist/manager/resident/evidence-recall.js +295 -0
  60. package/dist/manager/resident/evidence-recall.js.map +1 -0
  61. package/dist/manager/resident/history-disk.d.ts +10 -0
  62. package/dist/manager/resident/history-disk.d.ts.map +1 -0
  63. package/dist/manager/resident/history-disk.js +50 -0
  64. package/dist/manager/resident/history-disk.js.map +1 -0
  65. package/dist/manager/resident/history.d.ts +79 -0
  66. package/dist/manager/resident/history.d.ts.map +1 -0
  67. package/dist/manager/resident/history.js +203 -0
  68. package/dist/manager/resident/history.js.map +1 -0
  69. package/dist/manager/resident/initiative.d.ts.map +1 -1
  70. package/dist/manager/resident/initiative.js +11 -3
  71. package/dist/manager/resident/initiative.js.map +1 -1
  72. package/dist/manager/resident/learning-cycle.d.ts +131 -0
  73. package/dist/manager/resident/learning-cycle.d.ts.map +1 -0
  74. package/dist/manager/resident/learning-cycle.js +306 -0
  75. package/dist/manager/resident/learning-cycle.js.map +1 -0
  76. package/dist/manager/resident/learning-observation.d.ts +80 -0
  77. package/dist/manager/resident/learning-observation.d.ts.map +1 -0
  78. package/dist/manager/resident/learning-observation.js +22 -0
  79. package/dist/manager/resident/learning-observation.js.map +1 -0
  80. package/dist/manager/resident/learning-store.d.ts +106 -0
  81. package/dist/manager/resident/learning-store.d.ts.map +1 -0
  82. package/dist/manager/resident/learning-store.js +598 -0
  83. package/dist/manager/resident/learning-store.js.map +1 -0
  84. package/dist/manager/resident/learning.d.ts +246 -3
  85. package/dist/manager/resident/learning.d.ts.map +1 -1
  86. package/dist/manager/resident/learning.js +96 -6
  87. package/dist/manager/resident/learning.js.map +1 -1
  88. package/dist/manager/resident/outbox.d.ts +4 -4
  89. package/dist/manager/resident/store.d.ts +37 -4
  90. package/dist/manager/resident/store.d.ts.map +1 -1
  91. package/dist/manager/resident/store.js +27 -3
  92. package/dist/manager/resident/store.js.map +1 -1
  93. package/dist/manager/resident/tool-evidence.d.ts +71 -0
  94. package/dist/manager/resident/tool-evidence.d.ts.map +1 -0
  95. package/dist/manager/resident/tool-evidence.js +285 -0
  96. package/dist/manager/resident/tool-evidence.js.map +1 -0
  97. package/dist/manager/run/persistence.d.ts +8 -0
  98. package/dist/manager/run/persistence.d.ts.map +1 -1
  99. package/dist/manager/run/persistence.js +18 -0
  100. package/dist/manager/run/persistence.js.map +1 -1
  101. package/dist/plugin/loader.d.ts.map +1 -1
  102. package/dist/plugin/loader.js +5 -3
  103. package/dist/plugin/loader.js.map +1 -1
  104. package/dist/prompt/coding-agent-doctrine.d.ts +1 -1
  105. package/dist/prompt/coding-agent-doctrine.d.ts.map +1 -1
  106. package/dist/prompt/coding-agent-doctrine.js +2 -0
  107. package/dist/prompt/coding-agent-doctrine.js.map +1 -1
  108. package/dist/prompt/index.d.ts +2 -0
  109. package/dist/prompt/index.d.ts.map +1 -1
  110. package/dist/prompt/index.js +1 -0
  111. package/dist/prompt/index.js.map +1 -1
  112. package/dist/prompt/resident-learning.d.ts +19 -0
  113. package/dist/prompt/resident-learning.d.ts.map +1 -0
  114. package/dist/prompt/resident-learning.js +125 -0
  115. package/dist/prompt/resident-learning.js.map +1 -0
  116. package/dist/prompt/resident-step.d.ts +8 -1
  117. package/dist/prompt/resident-step.d.ts.map +1 -1
  118. package/dist/prompt/resident-step.js +65 -4
  119. package/dist/prompt/resident-step.js.map +1 -1
  120. package/dist/provider/collect-chat-completion.d.ts +2 -1
  121. package/dist/provider/collect-chat-completion.d.ts.map +1 -1
  122. package/dist/provider/collect-chat-completion.js +7 -6
  123. package/dist/provider/collect-chat-completion.js.map +1 -1
  124. package/dist/provider/fallback.d.ts.map +1 -1
  125. package/dist/provider/fallback.js +2 -1
  126. package/dist/provider/fallback.js.map +1 -1
  127. package/dist/provider/stream-text.d.ts +16 -0
  128. package/dist/provider/stream-text.d.ts.map +1 -0
  129. package/dist/provider/stream-text.js +51 -0
  130. package/dist/provider/stream-text.js.map +1 -0
  131. package/dist/public-runtime.d.ts +16 -4
  132. package/dist/public-runtime.d.ts.map +1 -1
  133. package/dist/public-runtime.js +17 -4
  134. package/dist/public-runtime.js.map +1 -1
  135. package/dist/public-tools.d.ts +13 -0
  136. package/dist/public-tools.d.ts.map +1 -1
  137. package/dist/public-tools.js +16 -0
  138. package/dist/public-tools.js.map +1 -1
  139. package/dist/public-types.d.ts +15 -3
  140. package/dist/public-types.d.ts.map +1 -1
  141. package/dist/registry/tool/execute.d.ts.map +1 -1
  142. package/dist/registry/tool/execute.js +2 -3
  143. package/dist/registry/tool/execute.js.map +1 -1
  144. package/dist/registry/tool/portable.d.ts +65 -0
  145. package/dist/registry/tool/portable.d.ts.map +1 -0
  146. package/dist/registry/tool/portable.js +244 -0
  147. package/dist/registry/tool/portable.js.map +1 -0
  148. package/dist/registry/tool/schema.d.ts +32 -5
  149. package/dist/registry/tool/schema.d.ts.map +1 -1
  150. package/dist/registry/tool/schema.js +35 -9
  151. package/dist/registry/tool/schema.js.map +1 -1
  152. package/dist/registry/toolset/catalog.js +8 -8
  153. package/dist/registry/toolset/catalog.js.map +1 -1
  154. package/dist/run/LimitChecker.js +3 -3
  155. package/dist/run/LimitChecker.js.map +1 -1
  156. package/dist/run/evidence-query.d.ts +41 -0
  157. package/dist/run/evidence-query.d.ts.map +1 -0
  158. package/dist/run/evidence-query.js +270 -0
  159. package/dist/run/evidence-query.js.map +1 -0
  160. package/dist/run/evidence-recall.d.ts +99 -0
  161. package/dist/run/evidence-recall.d.ts.map +1 -0
  162. package/dist/run/evidence-recall.js +633 -0
  163. package/dist/run/evidence-recall.js.map +1 -0
  164. package/dist/run/index.d.ts +2 -0
  165. package/dist/run/index.d.ts.map +1 -1
  166. package/dist/run/index.js +1 -0
  167. package/dist/run/index.js.map +1 -1
  168. package/dist/run/json-claim-verifier.d.ts +83 -0
  169. package/dist/run/json-claim-verifier.d.ts.map +1 -0
  170. package/dist/run/json-claim-verifier.js +200 -0
  171. package/dist/run/json-claim-verifier.js.map +1 -0
  172. package/dist/run/preparation-context-error.d.ts +10 -0
  173. package/dist/run/preparation-context-error.d.ts.map +1 -0
  174. package/dist/run/preparation-context-error.js +15 -0
  175. package/dist/run/preparation-context-error.js.map +1 -0
  176. package/dist/run-query/index.d.ts +3 -1
  177. package/dist/run-query/index.d.ts.map +1 -1
  178. package/dist/runtime/jobs/awaited-jobs.d.ts +215 -0
  179. package/dist/runtime/jobs/awaited-jobs.d.ts.map +1 -0
  180. package/dist/runtime/jobs/awaited-jobs.js +259 -0
  181. package/dist/runtime/jobs/awaited-jobs.js.map +1 -0
  182. package/dist/runtime/jobs/registry.d.ts +33 -2
  183. package/dist/runtime/jobs/registry.d.ts.map +1 -1
  184. package/dist/runtime/jobs/registry.js +37 -0
  185. package/dist/runtime/jobs/registry.js.map +1 -1
  186. package/dist/runtime/query/callback-inference.d.ts +8 -0
  187. package/dist/runtime/query/callback-inference.d.ts.map +1 -0
  188. package/dist/runtime/query/callback-inference.js +89 -0
  189. package/dist/runtime/query/callback-inference.js.map +1 -0
  190. package/dist/runtime/query/checkpoint.d.ts +4 -0
  191. package/dist/runtime/query/checkpoint.d.ts.map +1 -1
  192. package/dist/runtime/query/checkpoint.js +13 -0
  193. package/dist/runtime/query/checkpoint.js.map +1 -1
  194. package/dist/runtime/query/events.d.ts +1 -0
  195. package/dist/runtime/query/events.d.ts.map +1 -1
  196. package/dist/runtime/query/events.js +18 -0
  197. package/dist/runtime/query/events.js.map +1 -1
  198. package/dist/runtime/query/executor.d.ts +35 -1
  199. package/dist/runtime/query/executor.d.ts.map +1 -1
  200. package/dist/runtime/query/executor.js +77 -8
  201. package/dist/runtime/query/executor.js.map +1 -1
  202. package/dist/runtime/query/file-evidence-context.d.ts +5 -0
  203. package/dist/runtime/query/file-evidence-context.d.ts.map +1 -0
  204. package/dist/runtime/query/file-evidence-context.js +180 -0
  205. package/dist/runtime/query/file-evidence-context.js.map +1 -0
  206. package/dist/runtime/query/file-evidence-replay.d.ts +260 -0
  207. package/dist/runtime/query/file-evidence-replay.d.ts.map +1 -0
  208. package/dist/runtime/query/file-evidence-replay.js +647 -0
  209. package/dist/runtime/query/file-evidence-replay.js.map +1 -0
  210. package/dist/runtime/query/file-evidence-seed.d.ts +50 -0
  211. package/dist/runtime/query/file-evidence-seed.d.ts.map +1 -0
  212. package/dist/runtime/query/file-evidence-seed.js +100 -0
  213. package/dist/runtime/query/file-evidence-seed.js.map +1 -0
  214. package/dist/runtime/query/guard.d.ts.map +1 -1
  215. package/dist/runtime/query/guard.js +4 -0
  216. package/dist/runtime/query/guard.js.map +1 -1
  217. package/dist/runtime/query/index.d.ts +10 -1
  218. package/dist/runtime/query/index.d.ts.map +1 -1
  219. package/dist/runtime/query/index.js +133 -26
  220. package/dist/runtime/query/index.js.map +1 -1
  221. package/dist/runtime/query/iteration/index.d.ts +92 -36
  222. package/dist/runtime/query/iteration/index.d.ts.map +1 -1
  223. package/dist/runtime/query/iteration/index.js +420 -117
  224. package/dist/runtime/query/iteration/index.js.map +1 -1
  225. package/dist/runtime/query/iteration/phases/advisory.d.ts +2 -1
  226. package/dist/runtime/query/iteration/phases/advisory.d.ts.map +1 -1
  227. package/dist/runtime/query/iteration/phases/advisory.js +2 -1
  228. package/dist/runtime/query/iteration/phases/advisory.js.map +1 -1
  229. package/dist/runtime/query/iteration/phases/context.d.ts +12 -1
  230. package/dist/runtime/query/iteration/phases/context.d.ts.map +1 -1
  231. package/dist/runtime/query/iteration/phases/context.js.map +1 -1
  232. package/dist/runtime/query/iteration/phases/tool-review.d.ts.map +1 -1
  233. package/dist/runtime/query/iteration/phases/tool-review.js +8 -1
  234. package/dist/runtime/query/iteration/phases/tool-review.js.map +1 -1
  235. package/dist/runtime/query/iteration/provider-rejected-image.d.ts +2 -1
  236. package/dist/runtime/query/iteration/provider-rejected-image.d.ts.map +1 -1
  237. package/dist/runtime/query/iteration/provider-rejected-image.js +5 -2
  238. package/dist/runtime/query/iteration/provider-rejected-image.js.map +1 -1
  239. package/dist/runtime/query/iteration/stream-turn.d.ts +4 -1
  240. package/dist/runtime/query/iteration/stream-turn.d.ts.map +1 -1
  241. package/dist/runtime/query/iteration/stream-turn.js +22 -8
  242. package/dist/runtime/query/iteration/stream-turn.js.map +1 -1
  243. package/dist/runtime/query/plugin-hooks.d.ts +14 -0
  244. package/dist/runtime/query/plugin-hooks.d.ts.map +1 -1
  245. package/dist/runtime/query/plugin-hooks.js +18 -0
  246. package/dist/runtime/query/plugin-hooks.js.map +1 -1
  247. package/dist/runtime/query/repeat-call.d.ts +17 -4
  248. package/dist/runtime/query/repeat-call.d.ts.map +1 -1
  249. package/dist/runtime/query/repeat-call.js +26 -19
  250. package/dist/runtime/query/repeat-call.js.map +1 -1
  251. package/dist/runtime/query/resume-pending.d.ts +18 -33
  252. package/dist/runtime/query/resume-pending.d.ts.map +1 -1
  253. package/dist/runtime/query/resume-pending.js +59 -42
  254. package/dist/runtime/query/resume-pending.js.map +1 -1
  255. package/dist/runtime/query/review-policy.d.ts +4 -4
  256. package/dist/runtime/query/review-policy.d.ts.map +1 -1
  257. package/dist/runtime/query/review-policy.js +10 -9
  258. package/dist/runtime/query/review-policy.js.map +1 -1
  259. package/dist/runtime/query/sandbox-lifecycle.d.ts.map +1 -1
  260. package/dist/runtime/query/sandbox-lifecycle.js +4 -0
  261. package/dist/runtime/query/sandbox-lifecycle.js.map +1 -1
  262. package/dist/runtime/query/steering.d.ts +11 -1
  263. package/dist/runtime/query/steering.d.ts.map +1 -1
  264. package/dist/runtime/query/steering.js +12 -1
  265. package/dist/runtime/query/steering.js.map +1 -1
  266. package/dist/runtime/query/tool-output-budget.d.ts +10 -6
  267. package/dist/runtime/query/tool-output-budget.d.ts.map +1 -1
  268. package/dist/runtime/query/tool-output-budget.js +54 -12
  269. package/dist/runtime/query/tool-output-budget.js.map +1 -1
  270. package/dist/runtime/query/tooling.d.ts +5 -1
  271. package/dist/runtime/query/tooling.d.ts.map +1 -1
  272. package/dist/runtime/query/tooling.js +5 -0
  273. package/dist/runtime/query/tooling.js.map +1 -1
  274. package/dist/scheduler/completion-inbox.d.ts +49 -0
  275. package/dist/scheduler/completion-inbox.d.ts.map +1 -1
  276. package/dist/scheduler/completion-inbox.js +122 -2
  277. package/dist/scheduler/completion-inbox.js.map +1 -1
  278. package/dist/store/evidence/compaction-archive.d.ts +109 -0
  279. package/dist/store/evidence/compaction-archive.d.ts.map +1 -0
  280. package/dist/store/evidence/compaction-archive.js +125 -0
  281. package/dist/store/evidence/compaction-archive.js.map +1 -0
  282. package/dist/store/evidence/compaction-provenance.d.ts +7 -0
  283. package/dist/store/evidence/compaction-provenance.d.ts.map +1 -0
  284. package/dist/store/evidence/compaction-provenance.js +49 -0
  285. package/dist/store/evidence/compaction-provenance.js.map +1 -0
  286. package/dist/store/evidence/compaction-text.d.ts +16 -0
  287. package/dist/store/evidence/compaction-text.d.ts.map +1 -0
  288. package/dist/store/evidence/compaction-text.js +51 -0
  289. package/dist/store/evidence/compaction-text.js.map +1 -0
  290. package/dist/store/evidence/disk.d.ts +11 -0
  291. package/dist/store/evidence/disk.d.ts.map +1 -0
  292. package/dist/store/evidence/disk.js +367 -0
  293. package/dist/store/evidence/disk.js.map +1 -0
  294. package/dist/store/evidence/format.d.ts +25 -0
  295. package/dist/store/evidence/format.d.ts.map +1 -0
  296. package/dist/store/evidence/format.js +115 -0
  297. package/dist/store/evidence/format.js.map +1 -0
  298. package/dist/store/evidence/index-page.d.ts +208 -0
  299. package/dist/store/evidence/index-page.d.ts.map +1 -0
  300. package/dist/store/evidence/index-page.js +262 -0
  301. package/dist/store/evidence/index-page.js.map +1 -0
  302. package/dist/store/evidence/io.d.ts +24 -0
  303. package/dist/store/evidence/io.d.ts.map +1 -0
  304. package/dist/store/evidence/io.js +69 -0
  305. package/dist/store/evidence/io.js.map +1 -0
  306. package/dist/store/evidence/linked.d.ts +9 -0
  307. package/dist/store/evidence/linked.d.ts.map +1 -0
  308. package/dist/store/evidence/linked.js +314 -0
  309. package/dist/store/evidence/linked.js.map +1 -0
  310. package/dist/store/evidence/passages.d.ts +15 -0
  311. package/dist/store/evidence/passages.d.ts.map +1 -0
  312. package/dist/store/evidence/passages.js +72 -0
  313. package/dist/store/evidence/passages.js.map +1 -0
  314. package/dist/store/evidence/record-chain.d.ts +42 -0
  315. package/dist/store/evidence/record-chain.d.ts.map +1 -0
  316. package/dist/store/evidence/record-chain.js +111 -0
  317. package/dist/store/evidence/record-chain.js.map +1 -0
  318. package/dist/store/evidence/search-input.d.ts +21 -0
  319. package/dist/store/evidence/search-input.d.ts.map +1 -0
  320. package/dist/store/evidence/search-input.js +44 -0
  321. package/dist/store/evidence/search-input.js.map +1 -0
  322. package/dist/store/evidence/selection.d.ts +9 -0
  323. package/dist/store/evidence/selection.d.ts.map +1 -0
  324. package/dist/store/evidence/selection.js +17 -0
  325. package/dist/store/evidence/selection.js.map +1 -0
  326. package/dist/store/evidence/source-kind.d.ts +11 -0
  327. package/dist/store/evidence/source-kind.d.ts.map +1 -0
  328. package/dist/store/evidence/source-kind.js +26 -0
  329. package/dist/store/evidence/source-kind.js.map +1 -0
  330. package/dist/store/evidence/source-text.d.ts +51 -0
  331. package/dist/store/evidence/source-text.d.ts.map +1 -0
  332. package/dist/store/evidence/source-text.js +213 -0
  333. package/dist/store/evidence/source-text.js.map +1 -0
  334. package/dist/store/evidence/types.d.ts +162 -0
  335. package/dist/store/evidence/types.d.ts.map +1 -0
  336. package/dist/store/evidence/types.js +2 -0
  337. package/dist/store/evidence/types.js.map +1 -0
  338. package/dist/store/memory/disk.d.ts +2 -0
  339. package/dist/store/memory/disk.d.ts.map +1 -1
  340. package/dist/store/memory/disk.js +2 -1
  341. package/dist/store/memory/disk.js.map +1 -1
  342. package/dist/store/run/disk.d.ts +10 -0
  343. package/dist/store/run/disk.d.ts.map +1 -1
  344. package/dist/store/run/disk.js +87 -7
  345. package/dist/store/run/disk.js.map +1 -1
  346. package/dist/store/run/memory.d.ts +1 -0
  347. package/dist/store/run/memory.d.ts.map +1 -1
  348. package/dist/store/run/memory.js +10 -0
  349. package/dist/store/run/memory.js.map +1 -1
  350. package/dist/store/run/tool-executions.d.ts +13 -0
  351. package/dist/store/run/tool-executions.d.ts.map +1 -0
  352. package/dist/store/run/tool-executions.js +99 -0
  353. package/dist/store/run/tool-executions.js.map +1 -0
  354. package/dist/store/session/index.d.ts +2 -0
  355. package/dist/store/session/index.d.ts.map +1 -1
  356. package/dist/store/session/index.js +1 -0
  357. package/dist/store/session/index.js.map +1 -1
  358. package/dist/store/session/sqlite.d.ts +57 -0
  359. package/dist/store/session/sqlite.d.ts.map +1 -0
  360. package/dist/store/session/sqlite.js +430 -0
  361. package/dist/store/session/sqlite.js.map +1 -0
  362. package/dist/tools/builtins/bash.d.ts.map +1 -1
  363. package/dist/tools/builtins/bash.js +4 -10
  364. package/dist/tools/builtins/bash.js.map +1 -1
  365. package/dist/tools/builtins/edit-apply.d.ts +126 -0
  366. package/dist/tools/builtins/edit-apply.d.ts.map +1 -0
  367. package/dist/tools/builtins/edit-apply.js +360 -0
  368. package/dist/tools/builtins/edit-apply.js.map +1 -0
  369. package/dist/tools/builtins/edit.d.ts +143 -1
  370. package/dist/tools/builtins/edit.d.ts.map +1 -1
  371. package/dist/tools/builtins/edit.js +37 -219
  372. package/dist/tools/builtins/edit.js.map +1 -1
  373. package/dist/tools/builtins/index.d.ts +1 -0
  374. package/dist/tools/builtins/index.d.ts.map +1 -1
  375. package/dist/tools/builtins/index.js +9 -3
  376. package/dist/tools/builtins/index.js.map +1 -1
  377. package/dist/tools/builtins/job.d.ts.map +1 -1
  378. package/dist/tools/builtins/job.js +5 -6
  379. package/dist/tools/builtins/job.js.map +1 -1
  380. package/dist/tools/builtins/read-file.d.ts +2 -2
  381. package/dist/tools/builtins/read-file.d.ts.map +1 -1
  382. package/dist/tools/builtins/read-file.js +50 -65
  383. package/dist/tools/builtins/read-file.js.map +1 -1
  384. package/dist/tools/builtins/read-render.d.ts +56 -0
  385. package/dist/tools/builtins/read-render.d.ts.map +1 -0
  386. package/dist/tools/builtins/read-render.js +73 -0
  387. package/dist/tools/builtins/read-render.js.map +1 -0
  388. package/dist/tools/builtins/wait-for-job-bounds.d.ts +67 -0
  389. package/dist/tools/builtins/wait-for-job-bounds.d.ts.map +1 -0
  390. package/dist/tools/builtins/wait-for-job-bounds.js +108 -0
  391. package/dist/tools/builtins/wait-for-job-bounds.js.map +1 -0
  392. package/dist/tools/builtins/wait-for-job.d.ts +6 -0
  393. package/dist/tools/builtins/wait-for-job.d.ts.map +1 -0
  394. package/dist/tools/builtins/wait-for-job.js +162 -0
  395. package/dist/tools/builtins/wait-for-job.js.map +1 -0
  396. package/dist/tools/builtins/write-file.js +7 -2
  397. package/dist/tools/builtins/write-file.js.map +1 -1
  398. package/dist/tools/coordinator/index.d.ts.map +1 -1
  399. package/dist/tools/coordinator/index.js +1 -7
  400. package/dist/tools/coordinator/index.js.map +1 -1
  401. package/dist/tools/defineTool.d.ts +2 -1
  402. package/dist/tools/defineTool.d.ts.map +1 -1
  403. package/dist/tools/defineTool.js +1 -1
  404. package/dist/tools/defineTool.js.map +1 -1
  405. package/dist/tools/file-read-tracker.d.ts.map +1 -1
  406. package/dist/tools/file-read-tracker.js +90 -3
  407. package/dist/tools/file-read-tracker.js.map +1 -1
  408. package/dist/tools/resident-history.d.ts +10 -0
  409. package/dist/tools/resident-history.d.ts.map +1 -0
  410. package/dist/tools/resident-history.js +80 -0
  411. package/dist/tools/resident-history.js.map +1 -0
  412. package/dist/tools/resident-tool-evidence.d.ts +5 -0
  413. package/dist/tools/resident-tool-evidence.d.ts.map +1 -0
  414. package/dist/tools/resident-tool-evidence.js +57 -0
  415. package/dist/tools/resident-tool-evidence.js.map +1 -0
  416. package/dist/types/advisory/config.d.ts +7 -0
  417. package/dist/types/advisory/config.d.ts.map +1 -1
  418. package/dist/types/agent/reactive.d.ts +1 -0
  419. package/dist/types/agent/reactive.d.ts.map +1 -1
  420. package/dist/types/authorization/index.d.ts +9 -9
  421. package/dist/types/authorization/index.d.ts.map +1 -1
  422. package/dist/types/authorization/index.js +1 -1
  423. package/dist/types/authorization/index.js.map +1 -1
  424. package/dist/types/hitl/index.d.ts +4 -0
  425. package/dist/types/hitl/index.d.ts.map +1 -1
  426. package/dist/types/hitl/index.js.map +1 -1
  427. package/dist/types/message/index.d.ts +16 -2
  428. package/dist/types/message/index.d.ts.map +1 -1
  429. package/dist/types/message/index.js +10 -1
  430. package/dist/types/message/index.js.map +1 -1
  431. package/dist/types/provider/chat.d.ts +2 -0
  432. package/dist/types/provider/chat.d.ts.map +1 -1
  433. package/dist/types/provider/stream.d.ts +4 -0
  434. package/dist/types/provider/stream.d.ts.map +1 -1
  435. package/dist/types/run/answer-review.d.ts +34 -3
  436. package/dist/types/run/answer-review.d.ts.map +1 -1
  437. package/dist/types/run/config.d.ts +3 -0
  438. package/dist/types/run/config.d.ts.map +1 -1
  439. package/dist/types/run/entity.d.ts +20 -0
  440. package/dist/types/run/entity.d.ts.map +1 -1
  441. package/dist/types/run/events.d.ts +31 -8
  442. package/dist/types/run/events.d.ts.map +1 -1
  443. package/dist/types/run/events.js.map +1 -1
  444. package/dist/types/run/prepare-step.d.ts +53 -8
  445. package/dist/types/run/prepare-step.d.ts.map +1 -1
  446. package/dist/types/run/store.d.ts +29 -4
  447. package/dist/types/run/store.d.ts.map +1 -1
  448. package/dist/types/run/store.js +0 -28
  449. package/dist/types/run/store.js.map +1 -1
  450. package/dist/types/sandbox/index.d.ts +15 -14
  451. package/dist/types/sandbox/index.d.ts.map +1 -1
  452. package/dist/types/sandbox/index.js.map +1 -1
  453. package/dist/types/tool/index.d.ts +124 -1
  454. package/dist/types/tool/index.d.ts.map +1 -1
  455. package/dist/types/tool/index.js.map +1 -1
  456. package/dist/utils/await-with-abort.d.ts +8 -0
  457. package/dist/utils/await-with-abort.d.ts.map +1 -0
  458. package/dist/utils/await-with-abort.js +28 -0
  459. package/dist/utils/await-with-abort.js.map +1 -0
  460. package/dist/utils/env.d.ts +19 -0
  461. package/dist/utils/env.d.ts.map +1 -0
  462. package/dist/utils/env.js +25 -0
  463. package/dist/utils/env.js.map +1 -0
  464. package/dist/utils/evidence-time.d.ts +3 -0
  465. package/dist/utils/evidence-time.d.ts.map +1 -0
  466. package/dist/utils/evidence-time.js +10 -0
  467. package/dist/utils/evidence-time.js.map +1 -0
  468. package/dist/utils/evidence-tokens.d.ts +13 -0
  469. package/dist/utils/evidence-tokens.d.ts.map +1 -0
  470. package/dist/utils/evidence-tokens.js +29 -0
  471. package/dist/utils/evidence-tokens.js.map +1 -0
  472. package/package.json +1 -1
  473. package/src/advisory/executor.ts +15 -30
  474. package/src/advisory/history.ts +126 -0
  475. package/src/advisory/index.ts +5 -1
  476. package/src/agents/ReactiveAgent.ts +3 -0
  477. package/src/agents/runAgent.ts +14 -1
  478. package/src/compaction/manual.ts +30 -2
  479. package/src/compaction/summary.ts +4 -1
  480. package/src/config/runtime.ts +2 -2
  481. package/src/connector/mcp/adapter.ts +20 -6
  482. package/src/contracts/schemas.ts +1 -1
  483. package/src/eval/harness-protection.ts +79 -0
  484. package/src/eval/harness-verification.ts +31 -1
  485. package/src/eval/index.ts +1 -0
  486. package/src/manager/resident/activity.ts +187 -0
  487. package/src/manager/resident/agenda.ts +59 -5
  488. package/src/manager/resident/consumption.ts +311 -0
  489. package/src/manager/resident/evidence-recall.ts +363 -0
  490. package/src/manager/resident/history-disk.ts +63 -0
  491. package/src/manager/resident/history.ts +312 -0
  492. package/src/manager/resident/initiative.ts +13 -3
  493. package/src/manager/resident/learning-cycle.ts +499 -0
  494. package/src/manager/resident/learning-observation.ts +39 -0
  495. package/src/manager/resident/learning-store.ts +813 -0
  496. package/src/manager/resident/learning.ts +132 -8
  497. package/src/manager/resident/store.ts +31 -3
  498. package/src/manager/resident/tool-evidence.ts +412 -0
  499. package/src/manager/run/persistence.ts +18 -0
  500. package/src/plugin/loader.ts +8 -3
  501. package/src/prompt/coding-agent-doctrine.ts +2 -0
  502. package/src/prompt/index.ts +2 -0
  503. package/src/prompt/resident-learning.ts +143 -0
  504. package/src/prompt/resident-step.ts +83 -3
  505. package/src/provider/collect-chat-completion.ts +7 -6
  506. package/src/provider/fallback.ts +2 -1
  507. package/src/provider/stream-text.ts +56 -0
  508. package/src/public-runtime.ts +39 -1
  509. package/src/public-tools.ts +21 -0
  510. package/src/public-types.ts +98 -0
  511. package/src/registry/tool/execute.ts +2 -4
  512. package/src/registry/tool/portable.ts +264 -0
  513. package/src/registry/tool/schema.ts +38 -8
  514. package/src/registry/toolset/catalog.ts +8 -9
  515. package/src/run/LimitChecker.ts +3 -3
  516. package/src/run/evidence-query.ts +333 -0
  517. package/src/run/evidence-recall.ts +854 -0
  518. package/src/run/index.ts +11 -0
  519. package/src/run/json-claim-verifier.ts +298 -0
  520. package/src/run/preparation-context-error.ts +16 -0
  521. package/src/run-query/index.ts +1 -1
  522. package/src/runtime/jobs/awaited-jobs.ts +271 -0
  523. package/src/runtime/jobs/registry.ts +50 -0
  524. package/src/runtime/query/callback-inference.ts +94 -0
  525. package/src/runtime/query/checkpoint.ts +14 -0
  526. package/src/runtime/query/events.ts +22 -0
  527. package/src/runtime/query/executor.ts +100 -10
  528. package/src/runtime/query/file-evidence-context.ts +209 -0
  529. package/src/runtime/query/file-evidence-replay.ts +776 -0
  530. package/src/runtime/query/file-evidence-seed.ts +126 -0
  531. package/src/runtime/query/guard.ts +2 -0
  532. package/src/runtime/query/index.ts +153 -27
  533. package/src/runtime/query/iteration/index.ts +467 -110
  534. package/src/runtime/query/iteration/phases/advisory.ts +3 -0
  535. package/src/runtime/query/iteration/phases/context.ts +12 -0
  536. package/src/runtime/query/iteration/phases/tool-review.ts +7 -0
  537. package/src/runtime/query/iteration/provider-rejected-image.ts +6 -1
  538. package/src/runtime/query/iteration/stream-turn.ts +35 -8
  539. package/src/runtime/query/plugin-hooks.ts +20 -0
  540. package/src/runtime/query/repeat-call.ts +28 -18
  541. package/src/runtime/query/resume-pending.ts +65 -40
  542. package/src/runtime/query/review-policy.ts +13 -9
  543. package/src/runtime/query/sandbox-lifecycle.ts +3 -0
  544. package/src/runtime/query/steering.ts +11 -0
  545. package/src/runtime/query/tool-output-budget.ts +65 -13
  546. package/src/runtime/query/tooling.ts +10 -1
  547. package/src/scheduler/completion-inbox.ts +124 -2
  548. package/src/store/evidence/compaction-archive.ts +139 -0
  549. package/src/store/evidence/compaction-provenance.ts +52 -0
  550. package/src/store/evidence/compaction-text.ts +61 -0
  551. package/src/store/evidence/disk.ts +461 -0
  552. package/src/store/evidence/format.ts +126 -0
  553. package/src/store/evidence/index-page.ts +292 -0
  554. package/src/store/evidence/io.ts +95 -0
  555. package/src/store/evidence/linked.ts +365 -0
  556. package/src/store/evidence/passages.ts +88 -0
  557. package/src/store/evidence/record-chain.ts +108 -0
  558. package/src/store/evidence/search-input.ts +62 -0
  559. package/src/store/evidence/selection.ts +25 -0
  560. package/src/store/evidence/source-kind.ts +36 -0
  561. package/src/store/evidence/source-text.ts +285 -0
  562. package/src/store/evidence/types.ts +177 -0
  563. package/src/store/memory/disk.ts +4 -1
  564. package/src/store/run/disk.ts +110 -8
  565. package/src/store/run/memory.ts +10 -0
  566. package/src/store/run/tool-executions.ts +112 -0
  567. package/src/store/session/index.ts +2 -0
  568. package/src/store/session/sqlite.ts +584 -0
  569. package/src/tools/builtins/bash.ts +4 -10
  570. package/src/tools/builtins/edit-apply.ts +456 -0
  571. package/src/tools/builtins/edit.ts +39 -270
  572. package/src/tools/builtins/index.ts +9 -3
  573. package/src/tools/builtins/job.ts +5 -6
  574. package/src/tools/builtins/read-file.ts +56 -77
  575. package/src/tools/builtins/read-render.ts +104 -0
  576. package/src/tools/builtins/wait-for-job-bounds.ts +179 -0
  577. package/src/tools/builtins/wait-for-job.ts +184 -0
  578. package/src/tools/builtins/write-file.ts +7 -2
  579. package/src/tools/coordinator/index.ts +1 -7
  580. package/src/tools/defineTool.ts +4 -2
  581. package/src/tools/file-read-tracker.ts +86 -2
  582. package/src/tools/resident-history.ts +86 -0
  583. package/src/tools/resident-tool-evidence.ts +68 -0
  584. package/src/types/advisory/config.ts +7 -0
  585. package/src/types/agent/reactive.ts +1 -0
  586. package/src/types/authorization/index.ts +2 -2
  587. package/src/types/hitl/index.ts +4 -0
  588. package/src/types/message/index.ts +22 -0
  589. package/src/types/provider/chat.ts +2 -0
  590. package/src/types/provider/stream.ts +4 -0
  591. package/src/types/run/answer-review.ts +34 -3
  592. package/src/types/run/config.ts +3 -0
  593. package/src/types/run/entity.ts +21 -0
  594. package/src/types/run/events.ts +31 -8
  595. package/src/types/run/prepare-step.ts +59 -8
  596. package/src/types/run/store.ts +35 -4
  597. package/src/types/sandbox/index.ts +15 -14
  598. package/src/types/tool/index.ts +123 -1
  599. package/src/utils/await-with-abort.ts +26 -0
  600. package/src/utils/env.ts +23 -0
  601. package/src/utils/evidence-time.ts +9 -0
  602. package/src/utils/evidence-tokens.ts +32 -0
package/CHANGELOG.md CHANGED
@@ -1,5 +1,856 @@
1
1
  # Changelog
2
2
 
3
+ ## 40.0.0
4
+
5
+ ### Major Changes
6
+
7
+ - 8bfe291: Tool schemas now go on the wire in a shape every provider reads the same way, instead of a draft-07 rendering each driver was expected to translate.
8
+
9
+ **What broke, and why this is the fix.** `read`'s `readRange` was a `z.tuple`, which renders as the draft-07 tuple `items: [a, b]`. A wire that validates a tool's `parameters` against the JSON Schema 2020-12 metaschema does not read that as a tuple — it reads it as "not a schema" and refuses the entire request, so one field in one tool killed every other tool in the call and the turn produced nothing. Three of the ten driver packages convert dialects at their boundary; the other seven forward the rendering verbatim, because their wires had never been measured. Converting in seven more places would need seven more measurements, including for endpoints a user configures and nobody here can probe. So the schema is fixed where it is made: the renderer now emits the intersection of draft-07 and 2020-12, which needs no conversion anywhere.
10
+
11
+ **Take this upgrade if you talk to any provider that is not Claude-backed.** Nothing you write changes; what changes is which requests come back 400.
12
+
13
+ Breaking:
14
+
15
+ - **`renderToolSchema` no longer emits a tuple.** A `z.tuple([a, b])` now renders as `{"type":"array","items":{…},"minItems":2,"maxItems":2}` — members deduplicated, or `anyOf` of them when they differ — rather than `items: [a, b]`. If you pinned the rendered bytes of a tuple-shaped tool, or built a driver that depends on receiving the draft-07 spelling in order to convert it, update the expectation. `toSchemaDialect` still exists and still converts; it is simply no longer needed for schemas this kernel rendered.
16
+ - **`ReadWindowRequest.readRange` is `readonly number[]`, not `readonly [number, number]`.** `z.infer` over `read`'s input schema changes with it. Passing a pair still type-checks; reading one out now needs an undefined check. `resolveReadWindow` already does that and ignores a range that is not two numbers, which the tool's own parser rejects before it ever gets there.
17
+ - **A bridged MCP tool's positional array reaches the model as a uniform array plus a description naming each position**, where a server that pinned the arity and closed the tail previously produced `prefixItems`. The Zod parser is unchanged and still enforces the order and the member types; only the hint the model is shown is now portable.
18
+
19
+ `read`'s parameter did NOT change for the model. It is still `readRange: [start, end]`, 1-indexed and inclusive, and the same calls parse to the same values — both members already carried the identical `integer, minimum 1` constraint, so nothing was expressible in the tuple that the array cannot say.
20
+
21
+ Added:
22
+
23
+ - `findPortableSchemaViolations(schema)` — every place a schema leaves the intersection, each with its dotted path, keyword and remedy. Use it as a gate over your own tools; the kernel sweeps all of its own with it.
24
+ - `toPortableToolSchema(schema)` — the normaliser, returning the input unchanged when there is nothing to rewrite.
25
+ - `toolWireSchema(tool)` — the schema a tool actually sends: its hand-written `modelInputSchema` when it has one, else its rendering, portable either way.
26
+ - `PortableSchemaViolation`.
27
+
28
+ ### Minor Changes
29
+
30
+ - 86a3818: A file written and then edited in the same conversation keeps its evidence: the derived step context now references the write call plus the `edit` calls applied on top of it, so the model no longer re-reads a file whose content is fully determined by history it can already see.
31
+
32
+ `FileReadTracker` gains four OPTIONAL methods: `recordEdit(key, content, callId)` and `editChain(key)` for the chain, and `recordDriftObserved(key)`/`driftObserved(key)` for the one below. Nothing is required of an existing tracker — the built-in `edit` tool falls back to today's `recordRead(key, content)` when a tracker does not implement `recordEdit`, and `writeCallId(key)` keeps its exact meaning ("the write call whose body IS the current content"), returning `undefined` as soon as an edit lands on the path. A custom tracker that implements only the old shape sees the behavior it saw before, entry for entry.
33
+
34
+ The projection admits a chain only after replaying every visible hop through the same apply core the `edit` tool runs and finding the result equal to the ledger's disk-derived fingerprint; it reads no files, never writes that fingerprint, and withholds the whole path if any hop is missing, cleared by compaction, truncated, errored, ambiguous, names another file, or no longer applies. Chains are bounded to eight edit calls, and one request may replay at most 262,144 UTF-16 code units of content — string length, not bytes on disk. A hop is replayed one operation at a time, each operation's post-image length worked out exactly from the body it is about to be applied to and checked against the request's remaining room before it is built, so the ceiling refuses work instead of measuring it and refuses none that would have fitted. The charge is the largest body the hop actually built — for a batch the largest intermediate of the fold rather than the body it ends on, since a batch that grows a file to twenty megabytes and then deletes every character has still built the twenty megabytes. An operation the ceiling turns away costs nothing, leaving the paths behind it their room.
35
+
36
+ Mutation-time drift checks still decide what reaches disk, and now also say so: `edit` on either branch, and `write`'s fresh-overwrite check, call `recordDriftObserved(key)` before returning a refusal. That refusal had already read the real file; the flag carries no body, because recording what was read there would re-baseline the comparison that refused. It leaves `fingerprint`, `hasRead`, `writeCallId` and `editChain` untouched, is cleared by the next observation of any kind, and makes the projection withhold the path in the meantime — so a reference is never offered for a body the runtime has been told is behind disk. The projection still reads no files.
37
+
38
+ - 03630cd: Admit a whole-file `read` as visible file evidence, so the derived work context
39
+ references a body the model already has in a receipt instead of leaving it to
40
+ read the file again.
41
+
42
+ `FileReadTracker` gains two optional methods. `recordFullRead(key, content,
43
+ callId, renderedFingerprint)` does everything `recordRead(key, content)` does —
44
+ always with the whole file, never the window — and additionally records that the
45
+ body is visible in that call's receipt; `readWitness(key)` reports
46
+ `{ callId, renderedFingerprint }`. Both are optional, so a custom tracker that
47
+ implements neither keeps its behavior exactly, and `createFileReadTracker()`
48
+ implements both. The built-in `read` calls `recordFullRead` only when the
49
+ window covered the whole file, and falls back to `recordRead` for a tracker
50
+ without it.
51
+
52
+ `renderedFingerprint` is of the tool's own output string, not of the file's
53
+ body: a read's body survives only as the line-numbered rendering its receipt
54
+ carries, so the projection admits the entry only while the receipt it can see
55
+ fingerprints to exactly what the tool emitted. A result the output budget elided
56
+ or spilled, one compaction cleared, or one changed in any other way withholds
57
+ the path, and nothing anywhere recovers a body by undoing the numbering. A
58
+ receipt over 32,000 UTF-16 units is not read at all; a larger file is not
59
+ admitted this way.
60
+
61
+ Such an entry carries `kind: "read"` and never `editsInCalls` — a read roots no
62
+ chain, and the first `edit` on the path withdraws it. Read-rooted entries count
63
+ against the same six paths as the write-rooted ones, which keep a path both
64
+ could claim. Existing write and chain entries are unchanged, mutation-time
65
+ drift checks are untouched, and nothing here reads the filesystem.
66
+
67
+ - f33c62b: A run now suspends for a background job the model said it was waiting on, instead of settling over it. When the model stops calling tools and a job named by `wait_for_job` is still running, the run waits — no provider request, no tokens — for the job's exit, an operator message, or the settle grace, whichever comes first. On an exit the model gets one more turn with the `[Background job update]` line in front of it; on neither, the run settles and names the job.
68
+
69
+ This is the same bounded, zero-token wait `CompletionInbox` already gave a delegated task, and it shares the delegated task's grace — half of what the run has left before it must start finishing — under a ceiling of its own: two minutes, or `NAMZU_JOB_HOLD_MAX_MS`. On a run with a `timeoutMs` the grace comes out of what is left rather than being added to it, so time a `wait_for_job` call already spent shortens the hold by the same amount. On a run WITHOUT one — no run deadline, which is what the CLI ships — there is no remainder to take a share of, and the task ceiling would be a flat hour; that hour is sound for a task, which cannot outlive it, and wrong for a job, which can run forever. The two-minute job ceiling is what bounds that case, so a `wait_for_job` that ran its own bound out is followed by two more minutes at most, not by a second hour. The iteration limit still bounds all of it, and the wait starts nothing and stops nothing.
70
+
71
+ **Wait-intent is explicit.** Only a job `wait_for_job` named is awaited, and only for the rest of the run that named it. A job nobody waited on — a dev server, a watcher — never holds a run open, and there is no opt-in flag on `bash run_in_background` that changes that.
72
+
73
+ **Why this is `minor` and not `major`.** The signal is new: no run that exists today can have an awaited job, because nothing before this could mark one. A host that never calls `wait_for_job` sees the loop it saw before, so no default changes and no existing behaviour is withdrawn.
74
+
75
+ Additive API:
76
+
77
+ - `Run.abandonedJobIds` — awaited jobs still running when the run ended, the job-side counterpart to `abandonedTaskIds`. Naming them is not stopping them: a run-owned job is still stopped by the run's own teardown, and one bound to the host's session keeps running.
78
+ - `RUNTIME_CONTEXT_MESSAGE_KINDS` gains `'job-exit'`, the provenance on the message that carries an exit delivered by the wait. Consumers that exhaustively switch on `RuntimeContextMessageKind` need a case for it.
79
+ - `BackgroundJobRegistryRef` gains an optional `markAwaited(id)`, and `bindOwner`'s options take an `onAwaited(id)` callback that backs it. Both are optional; a host that wires neither gets the previous behaviour, which is no hold.
80
+ - `NAMZU_JOB_HOLD_MAX_MS` sets the job ceiling above, in milliseconds, beside the `NAMZU_JOB_WAIT_*` knobs `wait_for_job` already reads. Unset is two minutes.
81
+
82
+ - 6ae4072: The repeat-call advisory (notices, then escalates, when a tool is called with identical arguments over and over) now reaches the model even when the repeated tool's result is structured content — an image, a document, an MCP resource block — rather than plain text. `attachRepeatNotice` previously required the trailing tool result to be a string and silently dropped the notice otherwise; it now falls back to delivering the advisory as its own runtime-context message immediately after the tool-result batch. No thresholds changed, and a repeat that keeps succeeding is still only ever noticed, never refused.
83
+
84
+ `RuntimeContextMessageKind` gains a `'repeat-call'` member for this fallback message. A consumer that exhaustively switches over the union (the CLI's transcript labeling did) needs a case for it; `@namzu/cli` adds one in this release.
85
+
86
+ - 92ab1d9: A resumed conversation keeps the file witnesses it earned. The observation ledger is process memory, and every resume path handed the run an empty one: the derived work context could admit nothing, and the first thing a resumed agent did was read back a file whose whole body was in the transcript it had just been given.
87
+
88
+ The new export `seedObservationLedger(messages, tracker, { workingDirectory, additionalDirectories, sandboxed })` rebuilds a ledger from a conversation's own history. `resumeRun` and `query`'s checkpoint resume call it for you, from the history as repaired rather than as checkpointed, so the ledger describes exactly what the model is about to be shown; the CLI calls it the first time a turn asks for a conversation's tracker, which covers `/resume`, `namzu run --resume`/`--continue`, and a forked conversation — each seeded from its own messages, once. Call it directly if you keep a tracker per conversation and restore one yourself. Nothing is persisted and no session-store schema changes; a host that does nothing sees exactly today's behaviour.
89
+
90
+ What a replay may conclude is what the projection would admit, by the same predicates and the same bounded replay. A `write` whose call and successful receipt are both intact restores its body and its witness; the `edit` calls above it are replayed hop by hop and restore the chain. A `read` never supplies a body — the line numbering is never undone to recover one — and can only confirm one already reconstructed, by rendering it forward through the read tool's own renderer and comparing the whole rendering with the receipt. A windowed read, a read that shows something else, a cleared receipt, a hop that no longer applies, a body past the bounds and a call whose arguments run past what a replay reads as evidence each withdraw whatever the pass held for that path. So do the two cases where the transcript settles no outcome: a call it never answered — the unknown-outcome result the kernel's own repair writes for one included — may have landed with the file half written, and a mutation it refused is a tool's own report about that path, a drift refusal above all, made after reading the disk. Each of those costs the path it names and no other. A path whose walk ends holding no body is entered in the ledger nowhere, and a path this conversation only ever read establishes nothing.
91
+
92
+ No file's content is read. The one thing the seed does touch the filesystem for is the key each entry is filed under: a ledger entry identifies a file rather than a spelling, so `read`, `write` and `edit` all key on the path canonicalized through its symlinks, and entries filed any other way would be entries no mutation ever checks and no drift refusal can ever withdraw. The paths named in the history are therefore resolved exactly as the tools resolve them — `additionalDirectories` included — before the walk begins. Under a sandbox the keys are the paths as written and no host path is consulted.
93
+
94
+ Only content-backed observations are restored, so a seeded ledger is never weaker than the empty one a resume starts from. A path whose body could not be reconstructed is left OUT of the ledger rather than entered without a fingerprint: `hasRead` is the read-before-overwrite refusal, and granting it with no body to compare would let a full overwrite of a file that changed while the session was closed through with nothing checked. Every path the replay does not restore therefore behaves exactly as it does today. A fingerprint it does restore is a claim derived from history and is still compared with the real file at mutation time, so a file changed while the session was closed is refused there and the refusal withdraws the path from the projection.
95
+
96
+ Three things seed nothing at all, each leaving today's empty ledger: a history naming more than 1,024 distinct path spellings — the ones only `read` names included, and two spellings of one file counting twice — which is resolved whole or not at all rather than in a prefix that cannot say what a mutation replaced; a tool call id claimed by two calls or answered by two receipts, `read` included, since the receipt that was hidden could be the observation that withdrew a claim; and a mutation no path can be recovered from, whatever came back to it — one declaring no `path`, one whose path no longer resolves inside the directories the run may reach (a refused write to a path outside them is one of these: a key is what withdrawing one path rather than the whole pass takes), or one the provider stream cut off mid-JSON, whose arguments are recorded as `{}`. A merely large call is none of these: the argument bound governs what may be believed, not what may be attributed, so an oversize `write` withdraws its own path's body and leaves every other witness standing.
97
+
98
+ `read`'s numbering and windowing move to `tools/builtins/read-render.ts` as pure functions, which is what lets the forward-render comparison run the tool's own renderer rather than a copy of it. The tool's output is unchanged, byte for byte.
99
+
100
+ - 7ca8c7d: Add a `wait_for_job` builtin tool: it blocks on a background job's exit under a run-length bound and an idle bound that resets on new output, and returns the job's accumulated output in one call — the shell-job counterpart to the existing `wait_for_task`. Neither bound stops the job; a timeout reports which clock ran out and the output read so far, with a `next_offset` to resume from. Ships by default alongside `job` and `bash`, and refuses cleanly on a host with no background job registry.
101
+
102
+ `job`'s own description no longer instructs polling with `action: "read"` in a loop; it now points at `wait_for_job` instead. `read` and `list` are unchanged.
103
+
104
+ `BackgroundJobRegistry` gains a public `waitForExit(id, { signal })`, resolving immediately for a job that has already exited and honouring an abort signal. `BackgroundJobRegistryRef` (the tool-context surface) gains an optional `waitForExit` of the same shape — additive, so an existing host implementing this interface directly keeps working without it; `wait_for_job` refuses cleanly when it is absent.
105
+
106
+ ### Patch Changes
107
+
108
+ - 68e535b: Three fixes to the bookkeeping behind the run suspend for a background job the model awaited. Nothing about when a run holds itself open changes; what changes is that the record of an exit no longer outlives the exit.
109
+
110
+ - **A job exit that has been read stops counting as pending work.** The record of an exit used to survive the notice that delivered it, gated only by whether anything at all was queued on the job-notice channel — so the next job to end, awaited or not, made that stale record read as news and bought the model a turn to re-read an exit it had already seen. The record is now dropped by the delivery that accounts for it.
111
+ - **A hold takes an exit only together with the notice that delivers it.** The delivery path took the exits first and asked for the text afterwards; on the branch that found none, the exits were already gone and nothing carried them. Neither is taken unless both are there.
112
+ - **An exit that lands while the run is settling is delivered, not lost.** An awaited job ending in the moment between the hold's grace expiring and the run finishing was delivered by nobody — the hold had already looked, `abandonedJobIds` could not honestly name a job that had finished, and the host's own between-turns announcer stays quiet while a run is in flight. It now arrives as the same `{ type: 'runtime-context', kind: 'job-exit' }` message on `Run.messages`, so the transcript has it and a continued thread opens with it.
113
+
114
+ No API is added or withdrawn. `NAMZU_JOB_HOLD_MAX_MS` is now read with the same parse the `NAMZU_JOB_WAIT_*` bounds use, which only makes an invalid value fall back to the two-minute default the way the others already did.
115
+
116
+ - a9e4b19: `bash`'s own `timeout` and `run_in_background` parameter descriptions no longer tell the model to "poll with the `job` tool" — they now point at `wait_for_job` (one call, no waiting turns) the same way `job`'s own description already does, and reserve `job` with action `"read"` for incremental output or picking up after a `wait_for_job` timeout. The tool result returned when a background job starts is worded the same way. No schema, behavior or tool-result shape changed.
117
+ - a54dc71: Internal only: `edit`'s apply core (normalizing a call's arguments and applying its replacements or insertion) now lives in a shared internal module instead of being private to the `edit` tool's file. No exported symbol, tool behavior, error message or file output changes — this is a pure refactor that lets a later projection reuse the exact same apply logic instead of a separate reimplementation.
118
+ - a8df193: `CompletionInbox.describeOwnedWork()`'s owned-work projection no longer drops a still-running task purely because more tasks were launched after it. It used to keep a single FIFO over every owned task, so a long-running task launched early fell out of the model's visibility permanently once sixteen more tasks were merely LAUNCHED — whether or not any of those newer ones had actually finished.
119
+
120
+ It now lists running tasks first, most recently launched first, and fills whatever slots are left with the most recently settled tasks — still bounded to sixteen entries. A running task that does not fit is named honestly in the preamble ("N running tasks are shown below, and N more still running.") instead of disappearing without a trace. A settled task bumped out is not individually counted; its result already reached the model once, inline or as a notification.
121
+
122
+ Model-facing text only. No exported symbol, method signature or wire shape changed.
123
+
124
+ - dd8702d: Tighten the resume seed's accounting and guards.
125
+
126
+ A `write` a `pre_tool_use` hook SKIPPED gets a non-error receipt — the hook
127
+ declined the call, nothing failed — and the ledger replay read that as a
128
+ successful write, restoring a fingerprint for a body that never reached the
129
+ disk. The next edit to that file was then refused for a drift the ledger had
130
+ invented. The skip is now recognised through the same function the executor
131
+ writes it with, and withdraws the path instead.
132
+
133
+ Three other corrections to the same pass. A read that withdraws a path no
134
+ longer counts toward the six-body bound, so a later mutation cannot evict a
135
+ path still holding a body. Attribution reads each call's `path` at most once
136
+ per seeding and not at all past about a megabyte of arguments; past that the
137
+ call is a mutation that can be placed nowhere, and the seeding establishes
138
+ nothing rather than carry a body it may have replaced. And `query`'s checkpoint
139
+ resume seeds from the repaired history plus whatever of an owned resume turn a
140
+ completed scan says already ran — a recovered `write` restores the body it put
141
+ there, an unknown outcome withdraws the path, and a call proved never started
142
+ is left out because it is about to execute. Previously the seed never saw that
143
+ turn, so an executed write inside it left the body it replaced standing as a
144
+ claim.
145
+
146
+ A seeding that throws no longer propagates out of a resume: it is logged at
147
+ debug and the run continues with the empty ledger it would otherwise have had.
148
+
149
+ No public surface changed. Hosts calling `seedObservationLedger` directly get
150
+ the corrected pass with no change to the call.
151
+
152
+ - e6d6d1e: `bash`'s and the delegation coordinator's private copies of `readPositiveIntEnv` are gone; both now import the one already shared with `wait_for_job` and the iteration runtime. Each call site still samples its environment variable at the same point it always did — module load for `NAMZU_BASH_TIMEOUT_MS`, `NAMZU_BASH_MAX_BUFFER_BYTES`, `NAMZU_BASH_MAX_TIMEOUT_MS` and `NAMZU_DELEGATION_IDLE_MS` — so no knob starts reading its variable at a different time. No behavior, schema or default changed.
153
+
154
+ ## 39.0.0
155
+
156
+ ### Major Changes
157
+
158
+ - 6663561: Prose `reviewAnswer` callbacks now fail the run when they throw or return a malformed verdict. Previously a thrown error accepted the answer without review. To keep a deliberately permissive policy, catch the error in the host callback and explicitly return `{ accept: true }`; return `{ accept: false, feedback }` only when requesting a bounded correction. Rejection feedback must be a nonempty string.
159
+
160
+ `maxAnswerReviews` now rejects negative, fractional, non-finite or unsafe values. Use a nonnegative safe integer (default three corrections). Rejection counts and feedback are saved together in checkpoints, so resuming the same checkpoint preserves the remaining allowance even after history compaction. Cancellation stops waiting for a pending reviewer; external work started by the callback must still honor its signal.
161
+
162
+ The CLI inherits these SDK semantics for host-supplied reviewers. Its command gate already converts unavailable checks to bounded rejection and keeps that behavior. Forced finalization, terminal tools and structured output retain their separate settlement paths.
163
+
164
+ - 9463b6f: Background job `read` and `list` calls now count as read-only observations by default; starting commands and `job kill` retain their existing approval requirements. SDK `defineTool` accepts an input predicate for `readOnly`.
165
+
166
+ Explicit CLI `ask` rules are now enforced rather than omitted, so they can request review ahead of a wildcard allowance or the read-only default. SDK custom-pattern rules support `decision: 'review'`, with an `authorization.explicitReview` marker on review summaries. Read-only and accept-edits exemptions honor it; explicit auto modes and prior approvals keep their meaning.
167
+
168
+ To keep reviewing every background-job operation in prompt mode, configure `permissions: { job: ask }`, or supply a matching SDK custom-pattern review rule. If an old `ask` entry was intended to inherit default behavior, remove that entry instead. Deny rules, plan-mode mutation restrictions, job ownership and sandbox boundaries remain enforced.
169
+
170
+ - b2d5b01: Support explicit unlimited run guards while retaining measured token usage.
171
+ Set `tokenBudget: 0`, `maxIterations: 0` and `timeoutMs: 0` in SDK run options,
172
+ or in the CLI's `limits` configuration, to disable those three caps. The CLI's
173
+ `--token-budget 0` and `--max-iterations 0` now override configured caps; blank,
174
+ negative and unsafe numeric values are refused. Omitted defaults are unchanged.
175
+
176
+ SDK breaking change: `maxIterations: 0` and `timeoutMs: 0` previously prevented
177
+ progress; they now disable those guards, consistently with the token limit.
178
+ Hosts that used zero to prevent a run from starting must refuse admission or
179
+ pass an already-aborted signal instead. Use positive values for finite guards.
180
+
181
+ CLI breaking change: an explicitly configured `limits.maxIterations` now applies
182
+ to built-in subagents too, instead of always giving them 40 iterations. Existing
183
+ configurations with a smaller value can stop children earlier; larger values
184
+ permit more work. Omit that setting to retain the previous child default (40)
185
+ and parent default (50), or define a specialist agent with its own iteration
186
+ configuration when the two must differ. The new `limits.timeoutMs` setting also
187
+ reaches child runs and blocking delegation tools. `0` does not bypass a finite
188
+ ancestor token cap, permissions, operator cancellation or unresolved usage.
189
+
190
+ - ebfb3b4: Preserve original messages removed by CLI `/compact`, including exact user details
191
+ absent from its summary, for conversation search/read after restart. Failed
192
+ retention keeps the existing conversation; messages whose serialized form exceeds
193
+ 3 MiB are refused before replacement.
194
+
195
+ SDK consumers handling `compaction_shed.reason` or `ShedPass.reason` exhaustively
196
+ must add the new `manual` case. Both manual compaction helpers accept optional
197
+ `onShed` to await host-owned retention before returning replacement history;
198
+ callback failure rejects the operation. Existing callers without a callback keep
199
+ their projection-only behavior.
200
+
201
+ - e63ca83: Add `PrepareStepResult.context` for current observations carried after history in a labelled runtime message for this request only. It is separate from `system` authority, counted in subsequent stages' context estimates, and never replaces operator intent or accumulates in conversation history.
202
+
203
+ The emitted `RuntimeContextMessageKind` union now includes `step-context`. Consumers with exhaustive switches or records over that exported union must handle the new kind as runtime-generated context, not operator input. This is the SDK's breaking surface; existing `prepareStep.system` callers retain their behavior.
204
+
205
+ The CLI moves its changing context inventory into this field. OpenAI and Anthropic request conversion no longer moves that inventory ahead of conversation history as system text. This preserves history placement without promising cache hits or reduced billed tokens.
206
+
207
+ Anthropic message caching now places its breakpoint before request-only step context, so the cached boundary ends on stable history rather than the inventory that changes next step. Requests without step context keep their existing breakpoint.
208
+
209
+ - 5d31eea: Resident learning hosts and direct skill promotions now require a `protection` plan with disjoint `verification` and `confirmation` task IDs chosen before candidate generation. Existing hosts without this field are refused before inference. Include at least one real preservation task per round, with two measured successful baseline trials and two successful candidate trials. Missing or uncertain controls block activation; losing one established success rejects the candidate even when aggregate scores improve.
210
+
211
+ Update `ResidentLearningCycleOptions`, discovery hosts, and `ResidentSkillEvaluation` callers to supply this plan and its actual paired evidence. Historical stored skills remain readable but do not gain protection evidence retroactively. Generic `reviewHarnessCandidate` callers can opt into the same checks with its third argument. CLI learning summaries display protected-task outcomes.
212
+
213
+ - f1e33a1: Resident wake calls now retain all accepted inputs until the next step settles, instead of replacing the previous wake reason. `ResidentState.wakeEvidence` exposes immutable reasons and receipt times; the SDK resident prompt and both CLI resident profiles include the complete batch. CLI resident status shows pending input counts.
214
+
215
+ The new default accepts at most 16 pending inputs and 16,000 total reason characters per pursuit. Overflow rejects the new wake without discarding accepted evidence. Callers that previously sent an unlimited series of replacement wakes must process each batch before sending more, or coalesce superseded inputs before calling `wake`. Custom callbacks should read `wakeEvidence` rather than only the latest `reason`.
216
+
217
+ Standalone resident records now write schema 2 and agenda records schema 6. Older processes refuse these new formats: upgrade all processes sharing the store together. Prior formats remain readable without inventing historical inputs. Crashed steps keep their pending evidence; only successful exact-claim settlement or explicit inspected reconciliation consumes it.
218
+
219
+ - 1d651d0: Keep large compacted histories searchable, including short user text attached to
220
+ large images and individual long text messages. The disk store writes a bounded
221
+ `compaction_archive` storage record and saves original messages and authenticated
222
+ text chunks under the run's `compaction-output/` directory. Full SDK event readers
223
+ restore the original `compaction_shed` event with its attachments and metadata.
224
+
225
+ Raw JSONL consumers must handle this new storage record or switch to
226
+ `RunDiskStore.readEvents()` / `readRunEventsIn()`. Preserve `compaction-output/`
227
+ with the transcript when copying a run. Upgrade SDK readers before consuming new
228
+ archives. Existing inline records remain readable; older oversized records are
229
+ not converted automatically.
230
+
231
+ CLI manual compaction now offloads messages above 3 MiB instead of refusing them.
232
+ Automatic compaction and scoped search/read use the same SDK mechanism. Archive
233
+ write failures and limits still prevent the history replacement.
234
+
235
+ ### Minor Changes
236
+
237
+ - b156888: Supply configured advisors with the successfully dispatched SDK request,
238
+ including request-only step context, followed by records appended from its
239
+ response onward. The public `AdvisoryCallContext.turn` and exported
240
+ `AdvisoryTurnContext` describe this optional trajectory. Records distinguish
241
+ request inputs from later tools and messages within one shared context window.
242
+
243
+ Snapshots are scoped to one iteration, omitted from checkpoints, and rebuilt
244
+ after resume. If the current response anchor is unavailable, the advisor sees
245
+ explicitly labelled canonical history instead. Same-batch results that have not
246
+ yet reached history are not synthesized. No additional model call is enabled.
247
+
248
+ - 7bb8163: SDK evidence sources now accept `matchMode: 'token'` for complete Unicode
249
+ letter/number/underscore terms, with the same lowercase keys used by bounded
250
+ evidence ranking. The default remains literal substring search. Token queries
251
+ must contain one token per term; use literal mode for phrases or punctuation.
252
+ Continuations retain their matching mode, and token search authenticates the
253
+ preceding chunk when checking a word boundary within the existing I/O budget.
254
+
255
+ CLI automatic evidence recall uses this mode to keep incidental substrings
256
+ such as `in` inside `Packing`, or `3` inside `13000`, from consuming its candidate
257
+ slots. Explicit conversation search still supports literal substrings.
258
+ Whole-word frequency can still limit bounded discovery; this change does not
259
+ claim complete or globally ranked archive retrieval.
260
+
261
+ - 2d26b44: Add optional `AnswerReviewContext.requestMessages` to prose and structured
262
+ review callbacks. The built-in loop supplies an isolated copy of the SDK
263
+ provider-chain request that produced the candidate, including request-only
264
+ retrieved context and the last image-recovery dispatch. `messages` continues to
265
+ mean canonical conversation history.
266
+
267
+ The snapshot is not persisted across turns or checkpoints. Runs without a
268
+ reviewer do not copy requests for review; configured reviewers incur the memory
269
+ cost of that copy. This does not install a factual judge, authenticate arbitrary
270
+ request text, or capture provider-specific wire transformations. Hosts must
271
+ still validate source scope and integrity for their task-specific checks.
272
+
273
+ - 2a1e0e5: Preserve provider-identified public assistant message items through streaming,
274
+ settlement and conversation persistence. The SDK adds optional `textParts`
275
+ snapshots, `textPart` delta metadata and `selectAssistantText`. Completed content
276
+ selects explicitly final answers instead of concatenating intermediate progress
277
+ into the answer; ordinary unphased streams retain their existing behavior.
278
+
279
+ The Codex subscription driver maps native message phases and verifies the original
280
+ public parts before native replay. The CLI exposes optional item metadata on
281
+ delta events, separates streamed item bubbles and uses the settled answer for
282
+ turn completion. Consumers that manually concatenate deltas should use completed
283
+ content when they want the final answer; deltas still contain public progress.
284
+
285
+ - ce55c21: Add experimental `createEvidenceRecallStep` and its typed host retrieval contract.
286
+ It ranks a bounded pool of authenticated historical passages and supplies exact
287
+ excerpts with source/error/preview labels in request-only context. Every request
288
+ revalidates ownership and source data; deadlines discard late reads without
289
+ accumulating overlapping retrieval or replaying actions.
290
+
291
+ Recorded CLI conversations can opt in with `compaction.recallEvidence: true`.
292
+ The default remains off. Automatic recall excludes the requesting invocation;
293
+ explicit conversation search/read still cover live evidence, more pages and
294
+ complete text. This adds historical context, not automatic verification of
295
+ current workspace state or a guarantee of exhaustive recall.
296
+
297
+ - de53442: Allow hosts to lower the read allowance for an individual resident-history or
298
+ run-evidence search/read operation with `maxReadBytes`. Resident history accepts
299
+ 1 byte through 8 MiB; run evidence accepts 1–8 MiB and cannot exceed its source's
300
+ configured ceiling. The option also reaches captured live-boundary text readers.
301
+ Existing defaults and cursor/address identity remain unchanged, so a continuation
302
+ can use another allowance while preserving its scope and source validation.
303
+
304
+ Custom source backends must honor the new option when their caller supplies it.
305
+ An insufficient allowance can stop traversal or refuse a read; it does not mean
306
+ the requested evidence is absent. These low-level controls do not yet impose a
307
+ combined budget on the resident tool-evidence wrapper.
308
+
309
+ - 6e4a820: Recover original oversized tool text in ordinary conversations after compaction or restart. `search_conversation` and `read_conversation` now use authenticated retained output for closed scoped runs while preserving assistant-message and compaction-history search. Search results can provide a UTF-8 byte position for reading near a match; returned character positions remain UTF-16. Missing or changed originals are explicitly unavailable, and partial legacy records remain previews.
310
+
311
+ The SDK adds `createDiskRunTextEvidenceSource` and its public types, a bounded text view alongside the existing tool-only evidence source, plus an optional smaller per-operation read ceiling. New spill manifests record character positions without changing the existing tool-only source interface.
312
+
313
+ Headless `run --resume`/`--continue` and persistent `run-stream --session` now receive conversation retrieval tools. Both search and read remain available with deferred tool loading. Hosts still authorize the invoking conversation; no tool is replayed to recover its result.
314
+
315
+ - bd4bd2e: Allow automatic evidence-query resolution to use one explicitly marked compaction summary as a derived lookup reference after original turns leave visible history. The existing six-excerpt, 64-message and inference limits remain; ordinary system policy and tool text are excluded. Query-resolution basis metadata may now include `source: "compaction-summary"`. This provenance marks a derived reference, not proof of the requested fact: answers still require retained originals. Recorded CLI conversations use this through their existing recall configuration.
316
+ - bb0281b: Add optional `EvidenceRecallBatch.continuations` with exported
317
+ `EvidenceRecallContinuation` hints for bounded, host-mounted read-only tools.
318
+ Incomplete recall now reports its status even when no new passage is selected,
319
+ so missing context cannot silently look like an exhaustive negative search.
320
+ Hint arguments and output are bounded within the existing context allowance.
321
+
322
+ CLI `search_conversation` accepts `cursor` alone to restore the original query,
323
+ case setting and excluded invocation. Automatic recall supplies these handles
324
+ when live or earlier-run traversal has more pages. The model can continue from
325
+ that position without replaying an action or starting the same scan again.
326
+ New searches still require a literal query. Scope, expiry, source-integrity
327
+ checks and read limits remain enforced; live handles require the same active
328
+ writer and never downgrade to another source. Automatic recall remains opt-in.
329
+
330
+ - cff2b6a: Conversation search now ignores letter case by default: searching for `destination` also finds `Destination`. Pass `caseSensitive: true` to `search_conversation` to keep the former behavior. Continue pages with the same query and case setting.
331
+
332
+ SDK run-evidence search adds optional `caseSensitive` (default `true`, unchanged). Active and closed run sources now return distinct matching passages within one text chunk, with continuation at the match limit, rather than hiding later passages in that chunk. Exact retained text, UTF-8/UTF-16 offsets, integrity verification and per-call I/O limits remain intact. Case-insensitive searches bypass exact-case filters and may read more bytes or require more pages.
333
+
334
+ - df686fc: System messages can now carry `source: { type: 'compaction-summary' }`.
335
+ Kernel-generated compaction summaries receive this marker. When retained,
336
+ their text is searchable and readable as `compaction_shed:summary`, preserving
337
+ exact text and existing part positions. Ordinary system text with the same
338
+ heading and older unmarked archives keep their previous classification.
339
+
340
+ Automatic evidence recall orders matching source records before known derived
341
+ summaries within its bounded candidate pool, using separate relevance statistics.
342
+ Summaries remain available as passages and exact read addresses; they are not
343
+ deleted. CLI evidence guidance explains that these are derived text, not
344
+ independent observations. Recall limits and opt-in settings are unchanged.
345
+
346
+ - 2869fbe: Add an optional SQLite resident learning journal with atomic event/summary updates, scoped ancestry and recorded usage, plus hash-verified immutable JSON artifacts. `runStoredResidentLearningCycle` connects existing generation and independent evaluation callbacks to the journal without adding another model loop. The store requires Node.js 22.13 or newer when used; other SDK stores retain their existing support.
347
+
348
+ Add `namzu resident learn <experiment.learning.mjs>` for explicit trusted host modules and `namzu resident learning [cycle-id]` for read-only inspection. Modules select and bound their own providers and evaluators. Interrupted work and incomplete prices remain visible; these commands do not automatically replay experiments, activate unverified guidance or start background learning. Records live in `state/learning.sqlite` and `learning/artifacts/`; accepted skills remain in the existing resident agenda.
349
+
350
+ Expose `pathBuilder`, `runStore` and `checkpointStore` on `runAgent`, forwarding the kernel's existing host storage controls. Hosts can separate generated execution evidence from a searched workspace. Omitting these options preserves the SDK's current local layout; the CLI retains its application-home layout.
351
+
352
+ - df143c8: Expose optional `excerptComplete` on retained-evidence search matches and recall candidates. Built-in sources prove whether the displayed excerpt contains a whole full-retained text part using validated UTF-8 bounds. A partial excerpt or retained preview reports false; custom sources that omit the field remain unknown.
353
+
354
+ CLI conversation search and automatic recall preserve this information and explain when reading the same unchanged part adds no text or independent evidence. The field describes one text part, not the truth of its claims or coverage of the whole conversation. Existing scope, integrity, cancellation and context limits remain enforced.
355
+
356
+ - 691342c: Automatic conversation recall now labels selected passages and visible-source
357
+ references with their producer kind. Prior assistant statements are identified
358
+ as claims rather than proof of observed file state or successful actions.
359
+
360
+ Within the existing candidate and context limits, selection keeps the best
361
+ lexical match first and then considers matching records from other producer
362
+ kinds before repeating a kind. This prevents repeated model claims from taking
363
+ every slot when a tool record is available. Derived summaries remain last.
364
+ Archive bytes, explicit search/read tools and access boundaries are unchanged;
365
+ these labels and ranking do not establish truth or independent corroboration.
366
+
367
+ - d5d2b9a: Expose optional `recordedAt` Unix milliseconds on evidence search matches, exact read pages and recall candidates. CLI conversation search/read and automatic recall preserve the stored event time, including each included occurrence of equal text. Callers can distinguish recording times without inferring them from run IDs, file times or run-start metadata.
368
+
369
+ Unknown or invalid stored timestamps stay absent; custom recall callbacks must omit unknown times and supply positive integer milliseconds within the JavaScript Date range when known. The timestamp dates recording, not fact validity; compaction copies carry their own copy time. Sequence still orders one run, and clocks across runs do not establish causal order. Existing retrieval ordering, scope, read limits and source validation remain unchanged.
370
+
371
+ - b971796: Expose `classifyEvidenceSource`, `EvidenceRecordKind` and `EVIDENCE_RECORD_GUIDANCE` for hosts presenting authenticated text evidence. Automatic recall uses the same classification. The helper interprets source tags only; it does not authenticate text or establish that its claims are true.
372
+
373
+ CLI `search_conversation` matches and located `read_conversation` pages now include `recordKind` with interpretation guidance. Exact reads also preserve recorded `toolName` and `isError`, leaving missing status unknown. Tool names exceeding 256 JSON-encoded UTF-8 bytes are omitted consistently. Original text, addresses, scope checks and pagination remain unchanged.
374
+
375
+ - e9a4192: Add optional `terms` to retained-evidence source searches. Supply 1–16 nonblank literal terms instead of `query` to discover matching passages in a shared bounded scan. Exact duplicate terms and their order do not matter; cursors bind membership and case sensitivity. Both writer-captured and closed-run sources preserve scope, integrity checks, Unicode offsets, exact reads and existing resource limits. This is candidate discovery, not automatic recall or relevance ranking. Existing literal-query callers and CLI tool schemas keep their behavior.
376
+ - 3e09024: Learning candidates and learning cycles can declare `purpose: 'exploration'` for instructions intended to improve an explorer. Their purpose is covered by the content digest, and generation cannot redirect the host-admitted purpose. A skill cannot change purpose under the same name.
377
+
378
+ `projectResidentLearning` continues to select task guidance by default. Exploration policies require an explicit matching purpose and are reported as `different-purpose` when withheld. Existing skills without a purpose retain their task behavior and hashes. CLI resident steps therefore keep exploration policies out of ordinary task context. Explicit exploration projection still requires matching source revisions and does not grant tools or start inference.
379
+
380
+ - f49a4b8: Add optional run-metered, tool-free `PrepareStepContext.generateText` and
381
+ `createEvidenceRecallStep({ resolveQuery: true })` for resolving historical
382
+ follow-ups against bounded visible conversation. Generated search terms must
383
+ occur in the question or exact cited history. SDK query resolution defaults off.
384
+
385
+ In the CLI, conversations with `compaction.recallEvidence: true` now resolve
386
+ eligible conversational queries by default. This can add one provider request
387
+ per operator input, up to 512 output tokens and ten seconds before local
388
+ retrieval. It consumes the same run token budget. Set
389
+ `compaction.resolveEvidenceQueries: false` to keep the previous literal-query,
390
+ local-only behavior. Automatic recall itself still defaults off.
391
+
392
+ - 43124f0: When automatic evidence query resolution is enabled, allow its existing bounded
393
+ planner to select grounded subject words for discovery. A named record can now
394
+ focus the search without generic field words filling context with other records.
395
+ Source spelling, quoted context and every candidate's conversation ownership
396
+ are validated before use.
397
+
398
+ Temporary context reports the selected focus, observed focus words and locally
399
+ excluded passages. An empty focused scan is explicitly not proof of archive
400
+ absence. Explicit conversation search/read tools remain available with their
401
+ existing semantics; no additional model call or retrieval budget is introduced.
402
+ SDK query resolution remains opt-in. CLI integration checks cover archived
403
+ observations after reopening a conversation.
404
+
405
+ - 0a0baf2: Add bounded resident activity inspection and a reusable consumption projection in the SDK. The CLI's new `namzu resident inspect` command reports retained admissions, settlements, archived pursuits, historical verification receipts and known versus missing usage across process restarts.
406
+
407
+ Root usage and descendant-inclusive token totals remain separate. Missing or interrupted receipts are explicitly incomplete; unpriced tokens do not imply free work. Cost reports cover the root invocation, not descendant prices or a provider bill. Inspection does not impose a new lifetime spending limit or change existing execution defaults. Use `--max-revisions` or the returned `--cursor` to inspect histories beyond the default bounded range.
408
+
409
+ - 77272e3: Read retained original observations after a process exits before recording a
410
+ terminal run status. Disk evidence factories accept `consistency: 'snapshot'`
411
+ for explicitly scoped nonterminal runs; the existing default remains `closed`.
412
+ Snapshot reads validate ownership and unchanged source bytes on every operation
413
+ without acquiring an execution lease, resuming tools or changing run metadata.
414
+
415
+ The CLI now uses this mode for recorded `idle`, `pending` and `running` runs
416
+ outside its requesting live writer. An incomplete final JSONL fragment is
417
+ excluded within the existing bounded I/O allowance without editing the source.
418
+ Search remains incomplete for nonterminal snapshots; a full read describes only
419
+ the selected retained text. File or metadata changes require a fresh search,
420
+ and missing or altered retained originals remain unavailable.
421
+
422
+ - 5996a84: Recover retained tool text while the same invocation is still running. The SDK adds optional `ToolContext.captureRunEvidence` and `RunStore.captureTextEvidence` capabilities; custom stores need not implement them. Disk events carry additive integrity links so new appends do not invalidate earlier search/read continuations. Scope changes, damaged records and modified retained outputs are refused; torn boundaries remain explicitly incomplete.
423
+
424
+ CLI conversation search and read use this capability for the requesting invocation, preserving exact output after compaction without repeating the original action. Live cursors expire when the writer is replaced; start a new search after restart. Closed-run retrieval continues to support durable run/event/part references.
425
+
426
+ Conversation search also identifies the originating tool and directs callers to read the full passage, so original observations can be distinguished from prior retrieval excerpts.
427
+
428
+ - 4828eb0: Add opt-in learning discovery from retained, host-scored failures. Hosts can record observations and authorize evaluator revisions; the SDK selects an eligible task against the installed guidance, then uses the existing generation, verification and fresh confirmation cycle. Task claims survive process restarts and prevent concurrent or accidental duplicate experiments. Provider errors, unresolved usage and obsolete observations are excluded from selection.
429
+
430
+ The CLI accepts discovery hosts in `resident learn` and adds `resident learning --observations` for paginated inspection. Compact output includes failure reasons and verification/confirmation pass counts so inspection does not require following raw artifact hashes. Existing explicit-failure hosts remain supported. Learning storage upgrades to schema 2 on its next write; older SDK builds restricted to schema 1 cannot reopen that upgraded database. Keep a database backup if a rollback to such a build is required.
431
+
432
+ - 97acc32: Resident run/start now automatically retrieve bounded original tool evidence
433
+ from earlier settled admissions before model requests, under both context
434
+ profiles. Previously these admissions exposed explicit archive tools only.
435
+ Set `compaction.recallEvidence: false` to retain that previous behavior. This
436
+ adds local archive I/O and request context; it does not add query-planning
437
+ inference, replay actions or grant ordinary chats/delegated agents access.
438
+
439
+ SDK hosts can attach `createResidentEvidenceRecallStep` to an admitted run.
440
+ It preserves historical Session/run/claim addresses and shares bounded evidence
441
+ selection with conversation recall. Resident tool sources also support bounded
442
+ token queries and exact cursor-only recovery, retaining query/filter identity
443
+ across reopening. Incomplete results do not establish absence.
444
+
445
+ Resident Sessions now leave signal handling to their enclosing host. Previously
446
+ the SDK emergency handler could exit immediately on SIGINT/SIGTERM before the
447
+ host wrote cleanup/runner receipts. Cancellation now drains through the resident
448
+ lifecycle and preserves the interrupted claim for inspected reconciliation.
449
+
450
+ - b649224: Expose optional `PrepareStepContext.captureRunEvidence(maxReadBytes?, signal?)`
451
+ for authenticated text from the current invocation's writer. It rejects local
452
+ or run cancellation and settled invocations; unsupported stores return
453
+ `undefined`. Automatic evidence recall forwards this capability with its own
454
+ deadline and revokes new captures when the recall pass ends.
455
+
456
+ With `compaction.recallEvidence: true`, recorded CLI turns now recall missing
457
+ observations from the current run, including after compaction. Up to two live
458
+ pages share the existing four-page, 8 MiB read ceiling with earlier runs;
459
+ explicit conversation tools still handle further pages and exact full text.
460
+ The default remains off. Captured observations describe the past and do not
461
+ establish current workspace contents or replay a tool action.
462
+
463
+ - 10e9984: Add optional `RunEvidenceSearchOptions.excludeSuccessfulTools` to omit successful
464
+ results from up to 16 exact tool names during bounded discovery. The default
465
+ excludes nothing. Filter membership is bound to continuations; exact reads stay
466
+ available and errors or unknown provenance remain searchable. Search results and
467
+ `EvidenceRecallBatch` can report optional `excludedToolResults`, counting skipped
468
+ visits rather than unique facts. A positive count can produce an explanatory
469
+ recall context even when no passage is selected.
470
+
471
+ Preserve tool name and explicit error status in compacted text when the same
472
+ record contains an unambiguous, correctly ordered call/result pair. Newly written
473
+ large compaction archives retain that metadata; older archives without it stay
474
+ unknown. Text addresses, original messages and copy timestamps are unchanged.
475
+
476
+ When CLI automatic evidence recall is enabled, successful `search_conversation`
477
+ and `read_conversation` results no longer occupy its initial candidate slots,
478
+ allowing original observations behind repeated archive quotes to be considered.
479
+ Automatic cursors preserve this filter. Start a new literal search without that
480
+ cursor to inspect the quoted search/read results. This fixes candidate pollution
481
+ without increasing budgets or changing the default-disabled recall option.
482
+
483
+ - b9e0f37: Add the experimental `refineEvidenceRecallTerms` SDK helper for bounded lexical
484
+ coverage checks. Hosts can use the returned strict query subset to search terms
485
+ missing from candidate excerpts without introducing another model call.
486
+
487
+ When `compaction.recallEvidence` is enabled, the CLI spends existing retrieval
488
+ pages on uncovered terms so frequent words are less likely to hide an earlier
489
+ observation. Original and focused cursors retain their own query and omission
490
+ state. The four-page, two-live-page and 8 MiB read limits remain in force;
491
+ explicit conversation searches retain literal matching. No configuration or
492
+ stored-data migration is required.
493
+
494
+ - 2bcf017: Improve automatic resident evidence selection when a long objective/summary
495
+ loses its subject or frequent matches hide a rarer requested observation.
496
+ Selection samples both ends of bounded fields and can spend existing search
497
+ pages on uncovered query words. Original and corrected observations retain
498
+ separate provenance; ambiguous references are not silently resolved.
499
+
500
+ Disk evidence sources and the resident source factory now advertise
501
+ `supportsTermRefinement`. SDK callers can supply `refineTerms` with an existing
502
+ token-search cursor to branch a strict subset at its authenticated position.
503
+ Returned cursors use the subset; the original broad cursor remains valid.
504
+ Scope, filters, read ceilings and automatic page/context limits stay enforced.
505
+ Custom sources without this capability use a fresh subset search; the resident
506
+ factory restarts within the selected invocation when its resolved backend
507
+ cannot refine a cursor.
508
+
509
+ - 6394010: Learning hosts can supply an optional `explore` callback to run environment experiments before generating guidance. The SDK retains bounded observations and their digest, then provides them to `generate` as `context.exploration`. Missing usage, cancellation, stale state or evidence-journal failure prevents continuing to synthesis or activation.
510
+
511
+ The CLI forwards the callback, shows its exploration phase and retains evidence in SQLite. Hosts that omit it keep their current behavior. Event consumers opting into this feature should handle the new `explore` stage and `exploration` event kind. Exploration needs separately authorized tools and independent evaluation; enabling it does not automatically start learning in ordinary conversations.
512
+
513
+ - f4b3ffb: Residents can retrieve earlier settled summaries and consumed wake inputs when
514
+ the latest summary omits needed evidence. The SDK adds experimental
515
+ `DiskResidentAgenda.history`, `ResidentHistorySource` and related result types,
516
+ `buildResidentHistoryTools`, and optional `ResidentStepPromptOptions.history`.
517
+ Searches are bounded and paged, tied to one pursuit and an explicit upper
518
+ revision, and report unreadable evidence without treating it as proven absence.
519
+
520
+ CLI foreground and managed resident runs mount the two read-only recall tools
521
+ in both context profiles, including deferred loading. Ordinary conversations
522
+ and delegated children do not inherit the resident's history. These tools read
523
+ existing immutable revisions; they do not restore full tool transcripts, replay
524
+ actions, or change the persisted schema.
525
+
526
+ - f92daf8: Resident `run` and `start` now default to `--learning-disclosure on-demand` in the resident context profile. Previously all accepted learned skill bodies were included automatically; now the model sees their descriptions and can read relevant guidance with `read_resident_skill`. To retain automatic inclusion, pass `--learning-disclosure eager`. The interactive context profile and ordinary chat retain their existing behavior. Stored learning is unchanged.
527
+
528
+ The SDK adds `createResidentStepContext`, which returns prompt contributions and a read-only skill tool bound by the host to one admitted run. Source dependencies are checked when instructions are read and before subsequent requests. Existing `createResidentStepContributions` callers retain eager disclosure.
529
+
530
+ - 22203b0: Add the experimental `runResidentLearningCycle` workflow for host-authorized instructional learning. It generates one candidate from recorded failure evidence, reuses paired harness verification and fresh confirmation, and activates accepted guidance through the existing resident agenda transaction.
531
+
532
+ Hosts supply generation, independent evaluation and durable event storage. The workflow preserves usage from unsuccessful work, rejects duplicate receipts, distinguishes unpriced or incomplete consumption, and reports ambiguous activation acknowledgements without replaying work. Its recorded resource allowance does not impose an in-flight spending cap; callbacks must apply their own execution limits. Existing resident defaults are unchanged.
533
+
534
+ - b888779: Separate the SDK's overflow threshold from the size of an authenticated retained-output preview. `query`/`resumeRun` and `ReactiveAgent` accept `retainedToolPreviewChars`; unset or zero preserves the existing behavior. A shorter preview is used only after full host text and its integrity manifest are saved. Storage failure keeps the ordinary text budget. Rich blocks and independently supplied model text keep their existing handling.
535
+
536
+ Recorded CLI conversations now default to at most 4,000 characters for these retained overflow previews, previously up to 40,000. The 40,000-character spill threshold and ordinary smaller results are unchanged. Set `compaction.retainedToolPreviewChars: 0` in CLI configuration to retain the previous preview size. This applies to new tool results in ordinary turns and resumed runs, without rewriting existing history. Stateless sessions and delegated workers retain their existing defaults.
537
+
538
+ - d1a6ce5: Residents can search and page original retained tool text from earlier settled invocations, even when the latest summary or compacted context omits it. The SDK adds bounded disk indexing, scoped source interfaces and `search_resident_tools` / `read_resident_tool` builders. The CLI binds them to the admitted pursuit, matching attempt receipts and invocation ownership; ordinary conversations gain no cross-session access.
539
+
540
+ Fix fresh disk-backed runs capturing their output directory before store initialization, which could leave oversized tool output as an unrecoverable preview. New spills record chunk integrity manifests, and new run metadata records its own tenant/project/Session/run scope independently of shared token accounting. The existing 40,000-character model-visible cap remains unchanged. Older unscoped runs are unavailable through this API; older truncated records without authenticated spills remain explicitly partial. Missing or modified output is never replayed or presented as an intact original.
541
+
542
+ - d81aca6: Add optional `AnswerReviewContext.latestUserMessage`, an isolated copy of the
543
+ latest accepted operator, goal-round or steering input before the candidate's
544
+ model dispatch. It uses the run's existing retained input tracking across
545
+ compaction and checkpoint resume. Later arrivals do not relabel an older
546
+ candidate, and runtime reports do not replace operator intent. This single
547
+ input is not a complete task specification or a verbatim provider-wire record.
548
+
549
+ Fix tool-mode structured settlement dropping inbound messages or steering that
550
+ arrived during the candidate's request or tool execution. While run limits
551
+ permit, the loop handles the new input before publishing a new candidate,
552
+ including when no output reviewer is installed.
553
+
554
+ - 61aab1f: Add optional `AnswerReviewContext.generateText` to prose and structured output
555
+ review. A host callback can make one bounded, tool-free inference using the
556
+ run's provider/fallback chain, selected step model and effort. It accepts the
557
+ existing `PreparationTextRequest` shape and returns `PreparationTextResult`.
558
+
559
+ Usage contributes to the owning run and token budget without changing the
560
+ candidate step's usage or provenance. Await the call before returning a verdict;
561
+ completion, error and cancellation revoke the capability. Only explicitly
562
+ supplied text is sent, and no automatic model judge is installed. Hosts must
563
+ still validate generated judgments and authenticate task-specific source data.
564
+
565
+ - ea5367d: The CLI now requires Node.js 22.13+ and stores session metadata in
566
+ `NAMZU_HOME/state/sessions.sqlite`, with artifacts directly under
567
+ `NAMZU_HOME/sessions/<sessionId>/`. It no longer creates or reads a `projects/`
568
+ runtime tree. Generated memory remains isolated under `memory/<projectId>/`,
569
+ and resident state moves to `residents/<projectId>/<agent>/`.
570
+
571
+ This changes the default persisted CLI format. Existing project trees are left
572
+ untouched and are not imported automatically. Back up the original application
573
+ home and retain the older CLI to access its conversations, generated memory
574
+ and residents. Credentials, preferences and authored configuration keep their
575
+ locations. Update custom artifact readers to the new session paths.
576
+
577
+ The SDK adds the optional `SqliteSessionStore` driver and an exact `directory`
578
+ option for `DiskMemoryStore`. Existing SDK drivers, formats and defaults remain
579
+ unchanged; SQLite is loaded only when its driver is used.
580
+
581
+ `history` now accepts real conversation UUIDs as well as host keys, and with no
582
+ key reads the most recent workspace conversation as its help documents.
583
+
584
+ - 0fa8941: Allow a host to bound a resident tool-evidence operation across history,
585
+ invocation resolution and archive reads with `maxReadBytes`. Configure the
586
+ source's `resolutionReadBytes` with a host-enforced document-read ceiling;
587
+ bounded calls refuse to proceed without that declaration. Their `chargedBytes`
588
+ includes the declared resolution allowance, and a failed source search without
589
+ a byte receipt conservatively consumes the remainder. Calls without the new
590
+ option retain their separate existing limits.
591
+
592
+ The CLI declares the existing size bounds of its two attempt receipts, making
593
+ its source usable by bounded host retrieval. This does not enable automatic
594
+ resident recall yet. Returned pages must match their resolved invocation's
595
+ tenant/project/Session/run identity, and cancelled reads cannot expose a late
596
+ backend result. Custom sources must honor the read limits they accept.
597
+
598
+ - 3c60512: Add optional writer-owned `PersistedRunEvent.previousTextRecord` links for live
599
+ text retrieval. Bounded searches can reach earlier observations without spending
600
+ their record allowance on intervening nontext lifecycle events. Operational JSONL
601
+ records and adjacent links remain intact. Missing text links retain adjacent
602
+ traversal, and malformed content or incomplete history cannot be skipped as if
603
+ the archive were complete.
604
+
605
+ The CLI's opt-in automatic evidence recall benefits from these links within its
606
+ existing page and byte limits. No extra model call or tool action replay is used.
607
+ This is selected-text integrity checking, not a full audit of skipped operational
608
+ records; the `compaction.recallEvidence` default remains off.
609
+
610
+ - 8095541: Allow resident skill candidates to declare source revision dependencies. Their approval hash includes those bindings, and context projection withholds a bound skill unless every dependency matches fresh host observations. Unbound candidates keep their existing hashes and behavior. SDK hosts can resolve revisions per model request; the CLI resident profile supports bounded workspace-file SHA-256 observations.
611
+
612
+ Correct TUI stop messages for unresolved usage and accounting failures, including unlimited runs, so they no longer claim the token allowance was exhausted.
613
+
614
+ - 6c682d8: Text evidence searches accept `excludeDerivedSummaries`, defaulting to false.
615
+ This excludes only explicitly marked compaction summaries, binds the selection
616
+ into cursors, and reports `excludedSummaries` as skipped part visits. Summary
617
+ text remains available through unfiltered searches and exact reads.
618
+
619
+ When a partial automatic CLI evidence page contains derived summaries, the host
620
+ can spend its existing refinement page on source records instead. The general
621
+ cursor and already retrieved summaries are preserved. This helps discovery reach
622
+ original observations behind repeated summaries without increasing the four-page,
623
+ 8 MiB read or context allowances. Explicit tool continuations restore the exact
624
+ filter; new literal searches remain unfiltered. An incomplete scan still cannot
625
+ establish absence.
626
+
627
+ - 656e79d: Add optional cancellation signals to `ToolContext.captureRunEvidence` and
628
+ `RunStore.captureTextEvidence`. Existing implementations that accept fewer
629
+ arguments remain compatible; custom stores should observe the supplied signal
630
+ to stop their own I/O promptly.
631
+
632
+ Tool evidence capture now observes the tool's deadline and nested dispatch
633
+ cancellation, and refuses use after the tool call settles even if its parent
634
+ run is still working. Cancelling a local read leaves other calls available.
635
+ Queued cancelled captures are skipped without releasing a writer lock early.
636
+ An uncooperative custom store can still delay later appends until its pending
637
+ operation settles, although the cancelled caller stops waiting immediately.
638
+ CLI conversation search and exact reads also forward their operation signal
639
+ when capturing live evidence.
640
+
641
+ - 3c6326f: Prevent checkpoint resume from repeating a tool that started but never recorded
642
+ its completion. Previously, resuming a partially completed batch could execute
643
+ such a call again, duplicating an external effect. The resumed conversation now
644
+ receives an explicit unknown outcome and can verify current state before further
645
+ work. Completed calls remain recovered and proven unstarted calls can continue.
646
+
647
+ Add optional `RunStore.readToolExecutions` and exported `ToolExecutionSnapshot` /
648
+ `ToolExecutionRecord` types. Disk and memory stores implement the scan; custom
649
+ stores without it use their strict `readEvents` contract. Missing or contradictory
650
+ execution evidence does not authorize replay. The disk scan has documented size
651
+ bounds; exceeded bounds produce unknown outcomes instead of automatic re-execution.
652
+
653
+ Explicitly answered durable questions may still re-enter their own asking tool,
654
+ without granting the same exception to interrupted siblings.
655
+
656
+ CLI `drain` now passes configured run limits to its resume host. Previously a
657
+ bounded run could fail with a token-budget root-limit mismatch because `drain`
658
+ silently used an unlimited limit. Keep the original token limit in configuration;
659
+ the existing ledger still enforces its spent allowance.
660
+
661
+ - 0a36260: Add `createJsonClaimVerifier` for host-configured scalar JSON claims and observation-time receipts. It rejects mismatched, incomplete, historical or foreign observations, bounds verification time and bytes, and exposes pending observation drainage. Hosts supply an authorized read adapter; this does not verify arbitrary prose or establish atomic/future source state.
662
+
663
+ Resident `run` and `start` accept `--verify <manifest>` to require configured claims before recording completion. The manifest explicitly authorizes bounded host file reads, is snapshotted per invocation, and applies to every admitted pursuit. Rejected values use the existing repair budget; only the reviewed answer can complete. Existing invocation behavior is unchanged without the flag.
664
+
665
+ - f3b377e: Project bounded file-evidence references and owned worker status into model requests. A successful write body is referenced only while its complete call input and receipt remain visible and match the conversation's observation fingerprint. Existing disk-drift checks still run before mutations. Observations without content now invalidate an earlier fingerprint instead of carrying it forward.
666
+
667
+ `FileReadTracker.recordRead` accepts an optional third argument for a successful full-body write's tool-call ID, exposed through the optional `writeCallId` method. The built-in tracker preserves this witness across identical observations and clears it on changed or unknown content. Existing custom trackers remain valid; trackers without the witness do not enable the new file reference projection.
668
+
669
+ Add `CompletionInbox.describeOwnedWork()` for a non-consuming snapshot of up to sixteen owned tasks, separating scheduler state from delivery to history. The runtime uses it to keep available results visible after operator steering; delivery does not claim that a user-facing synthesis was produced. No automatic relaunch, answer-verification inference or persisted duplicate transcript is added.
670
+
671
+ - fd0d270: Automatic evidence recall now retains bounded source references for exact text already visible in conversation history. The temporary context can contain `visibleEvidence` entries binding an exact bounded `textQuote` to an `address` for a host archive-read tool, recording time when known, source and retention/error metadata. `omittedVisibleEvidence` reports references withheld by the existing character limit. Quotes repeat at most 512 UTF-16 units to make the source association explicit; full records remain available through the read address. Visible quotes and new passages share `maxPassages`, with new text taking priority.
672
+
673
+ An otherwise complete recall pass may now return source metadata even when all matching text is already visible. Consumers should not assume every recall block contains new passage text. New-text ranking is independent of visible copies, and source ownership, revalidation, read limits and cancellation remain enforced. CLI models can use each reference's `address` with `read_conversation` to recover the exact source association without searching again or replaying an action.
674
+
675
+ ### Patch Changes
676
+
677
+ - 40651dd: Fix configured advisors losing rich tool-result text and tool-call details in
678
+ conversation context. Public records now retain roles, host provenance, call
679
+ IDs, names, arguments and explicit result status; media content and private
680
+ provider replay state are omitted.
681
+
682
+ The existing `maxContextTokens` window now counts serialized records, including
683
+ rich text, metadata and escaping, instead of array element counts. Tight windows
684
+ may retain fewer complete records; increase the configured window if needed.
685
+ The unbounded default is unchanged. Omitted records are explicitly identified.
686
+
687
+ - 28d3874: Reduce unnecessary retained-output reads during whole-token evidence search,
688
+ including the case-insensitive mode used by CLI automatic recall. New manifests
689
+ and disposable indexes carry small token-key Bloom filters. A negative skips
690
+ payload I/O; every potential match is still authenticated and matched against
691
+ original text. Literal search, scope, cancellation, query limits and exact-read
692
+ addresses retain their existing contracts.
693
+
694
+ Filters add storage and write work. Existing manifests without them or with a different runtime tag still scan
695
+ normally, and the writer omits this optional metadata when it would exceed the
696
+ existing manifest size ceiling. No data migration or configuration change is
697
+ required. The optimization can reach sparse matches within the existing I/O
698
+ budget; it does not guarantee exhaustive automatic recall.
699
+
700
+ - 985db49: Correct resident step guidance that allowed an unnamed follow-up to select one
701
+ of several plausible subjects, or treat the latest historical observation as
702
+ current. Residents are instructed to distinguish subject corrections from
703
+ changes over time, qualify alternatives, obtain fresh permitted evidence for
704
+ current-state questions, and report unavailable evidence without claiming that
705
+ the unfinished current-state task is complete.
706
+
707
+ This applies to hosts using `createResidentStepContributions`, including the
708
+ CLI's default resident context profile. It adds no inference call, retrieval
709
+ permission or stored state. Model interpretation remains fallible; the host's
710
+ answer validation contract is unchanged.
711
+
712
+ - fe6e0fb: Avoid reading and parsing a shared compaction record once for every removed
713
+ message. A search reuses one authenticated record within that operation; later
714
+ calls revalidate it. Text manifests, integrity checks, cancellation and page
715
+ limits remain in effect.
716
+
717
+ CLI manual compaction now stores one `compaction_shed` event containing all
718
+ removed messages, matching automatic compaction, instead of one event per message.
719
+ Consumers of raw manual-maintenance events must iterate the `messages` array
720
+ and use search results' `seq` and `part` addresses, rather than assuming `part: 0`
721
+ or one sequence per message. Existing archives and SDK event readers remain
722
+ supported. No config change is required for ordinary CLI use.
723
+
724
+ - e954d02: Automatic evidence recall now reports `omittedPassages` when eligible distinct
725
+ records do not fit the selected passage count or context size. Bounded
726
+ `additionalEvidence` addresses let archive tools recover withheld text;
727
+ `omittedAddresses` reports addresses which also could not fit. The original
728
+ scope checks and character ceiling remain in force.
729
+
730
+ The context can now retain an omission notice and read address even when no
731
+ whole excerpt fits. `incomplete` continues to describe source traversal, rather
732
+ than implying that every matched record was presented. In the CLI these
733
+ addresses work with the existing `read_conversation` tool. This corrects hidden
734
+ selection loss without changing the recall opt-in or adding model calls to the
735
+ retrieval hook itself.
736
+
737
+ - b1e3bc5: Resource-limit closure and recovery from an empty assistant completion now ask for a concise evidence-supported response that attributes unverified claims and identifies unresolved work. The closing request no longer demands a comprehensive answer regardless of the available evidence. This changes the kernel's model guidance, not its factual validation guarantees, tool restrictions, resource limits or stop reasons. The instruction remains local to the closing request and is not retained as an instruction for resumed turns.
738
+ - 45c8292: Fix opt-in evidence query resolution skipping follow-up questions after six or
739
+ more assistant progress messages. Within the existing 64-message scan, retain
740
+ the nearest preceding operator request and five recent updates when progress
741
+ would otherwise fill all six reference slots. The prompt, retrieval and scan
742
+ ceilings remain unchanged. A missing or compacted-away request is not invented.
743
+
744
+ Do not rewind the reference window to an older identical question when the
745
+ current retained input is outside visible history, such as steering carried on
746
+ a tool result. Use the known message object as the boundary when available;
747
+ otherwise consider bounded recent history instead of inventing a position.
748
+
749
+ Normalize grounded filenames and punctuation-separated identifiers into the
750
+ same word tokens used by evidence discovery. A valid term such as
751
+ `sevkiyatlar.txt` no longer causes the entire optional plan to fail; all expanded
752
+ tokens must remain grounded and fit the existing 16-token ceiling.
753
+
754
+ Cover this behavior through the CLI Session host, including the existing
755
+ `resolveEvidenceQueries: false` opt-out. The CLI adds no default model call.
756
+
757
+ - 6e14db9: Clarify automatic evidence-recall guidance: quoting an identifier preserves its
758
+ spelling, while an operator-requested text transformation produces a derived
759
+ value, not a replacement source fact. The guidance fits the existing character
760
+ allowance; retained text, scope checks and retrieval defaults are unchanged.
761
+
762
+ This does not add automatic answer correction or a general factual validator.
763
+
764
+ - 6a6921c: Report unavailable automatic historical evidence in temporary model context instead of silently dropping every sign of a failed query plan or read. The short status distinguishes planning failure, retrieval failure, timeout and an earlier read still pending; it does not imply that the requested history is absent.
765
+
766
+ Raw error bodies, malformed plans and rejected source data stay out of the note. Existing error diagnostics and direct callback rejections remain, parent cancellation stops work, and context/read bounds still apply. Explicit archive tools remain available, and a failed cached query plan does not trigger an extra model call each iteration. No status note is stored as operator conversation history.
767
+
768
+ - 612879e: Recover text blocks in compacted tool results that also contain images or
769
+ documents. Scoped conversation search and exact reads now include those blocks
770
+ after compaction and restart, without mixing binary bytes or inserted separators
771
+ into the text. Existing plain-text part addresses keep pointing to the same
772
+ content. Newly written large archives retain the additional text parts; old
773
+ archives are not rewritten. Unindexed legacy scans report skipped block arrays
774
+ as incomplete instead of claiming a complete search. No configuration changes
775
+ are required; automatic CLI recall remains opt-in.
776
+ - e7bc7a1: When optional conversation query planning finds competing referents, preserve
777
+ that interpretation for the main model instead of silently skipping recall.
778
+ A temporary note carries validated quotes and asks the model to clarify if
779
+ needed; it is labelled as a fallible interpretation, not historical evidence.
780
+ No subject is selected for automatic retrieval in this case. The note shares
781
+ the existing context allowance and cancellation, and new operator input clears
782
+ the cached interpretation. SDK defaults and CLI configuration keys are unchanged.
783
+ - e40044b: Correct evidence-source guidance for questions about earlier observations. The SDK coding-agent doctrine now distinguishes retained historical content from current workspace reads, preserves exact identifiers in reports and requires unavailable history to be reported honestly. Recorded CLI turns explain how to recover clipped details with their conversation search/read tools, before compaction as well as after restart. The guidance is omitted when those tools are unavailable and remains stable across a run. Search responses explicitly distinguish remaining pages from unavailable evidence so a matching announcement is not confused with the original observation. Storage, permissions, retrieval bounds and freshness checks for edits are unchanged.
784
+ - 6446182: Fix opt-in evidence query planning failing when a model rewrites a word's
785
+ spelling or inflection. The internal planner selects numbered words supplied
786
+ by the host; retrieval receives the original spellings after quote validation.
787
+ The vocabulary shares the existing 12,000-character preparation allowance,
788
+ offers at most 256 distinct spellings, and reports omissions. Each plan still
789
+ selects at most 16 words. Literal retrieval remains the SDK default, and the
790
+ CLI's existing query-resolution opt-out remains available.
791
+
792
+ Keep present-state plans from expanding with historical terms. Invalid IDs or
793
+ quotes still reject optional preparation instead of weakening source grounding.
794
+ Update the CLI Session regression for the internal selection protocol.
795
+
796
+ - c4aaf9b: Interactive CLI sessions now honor `limits.maxIterations` and `limits.tokenBudget` from user and trusted project configuration. Previously these configured limits were ignored by TUI startup, although headless commands applied them. This also applies when rebuilding a session after a model change or reopening a conversation. To keep interactive cumulative tokens unlimited, omit `limits.tokenBudget` from the effective config and use `--token-budget` for individual headless runs. Omitted defaults remain unchanged; `limits.waitForProviderMs` remains a headless policy.
797
+
798
+ SDK closing prose requested by a token, cost or time warning now preserves the triggering limit's stop reason instead of reporting `end_turn`. The partial text is still returned, including when allowance remains, but this path skips prose answer review and must not be treated as verified completion. Cancellation and validated native structured-output settlement retain their existing behavior.
799
+
800
+ - 7579aa0: Keep bounded evidence retrieval available when optional query planning fails.
801
+ Recall can search the unchanged current-query tokens and deliver validated
802
+ records together with the planning failure status, without importing terms from
803
+ an invalid plan. The existing read, candidate, context and cancellation limits
804
+ still apply. Failed planning remains diagnostic and is not retried every step;
805
+ retrieved evidence is freshly validated. Explicit archive tools remain available
806
+ when literal retrieval cannot resolve the question.
807
+ - c329408: Fix repeated historical observations filling every automatic evidence-recall
808
+ passage slot and excluding a different record such as a correction. Exact equal
809
+ text with the same producer, retention and error status now shares a passage
810
+ before bounded BM25 scoring. Copies retain their separate source addresses;
811
+ changed identifiers, previews and errors remain distinct.
812
+
813
+ Request context includes `otherOccurrences` for additional addresses and
814
+ `omittedOccurrences` when the character allowance cannot hold all addresses in
815
+ the retrieved pool. Distinct text takes priority over extra addresses. No archive
816
+ record is removed, no current-state or cross-run chronology is inferred, and
817
+ explicit search/read tools are unchanged. CLI automatic recall remains opt-in
818
+ with `compaction.recallEvidence: true`; no read or passage limits increase.
819
+
820
+ - 4801a6f: Automatic evidence recall now recognizes text already present in tool text
821
+ blocks and earlier preparation stages. Those passages receive source references
822
+ instead of occupying slots intended for missing information. Images, documents
823
+ and private reasoning are not treated as visible text, and separate blocks are
824
+ never joined to invent a matching passage. Existing scope validation, context
825
+ limits and opt-in behavior are unchanged.
826
+
827
+ CLI automatic discovery can fill a candidate page from several completely
828
+ searched runs, within the same byte, output and page limits. It no longer
829
+ spends one automatic page on every small matching run. Explicit literal
830
+ searches retain their early return; unfinished source pages still require
831
+ continuation. Serialized matches, including escaping, share the output cap.
832
+
833
+ - 830f81e: Treat uppercase and lowercase hexadecimal spellings of the same UUID as the same identity when joining and deduplicating resident consumption evidence. A copied root receipt can no longer count twice merely because its UUID spelling differs. Original identifiers remain available to the host resolver.
834
+ - 9a4877a: Keep resident selection's mean resource cost finite when valid large observations would overflow an intermediate sum. Candidate explanations retain their numeric cost and score when serialized, and very small nonzero observations are not all discarded by dividing each cost prematurely. Selection remains opt-in; policy values and resource units are unchanged.
835
+ - 9ea5074: Added integration regression coverage for resident evidence recall after
836
+ structured and sliding-window compaction. The tests verify that an exact receipt
837
+ removed from model context remains recoverable through the registered history
838
+ tools while the original pursuit boundary stays intact. Runtime behavior and
839
+ public APIs are unchanged.
840
+ - 2e93158: Preserve full permitted shell output before condensing similar lines. Previously,
841
+ condensation happened before retention, so omitted row values could be lost even
842
+ though conversation search reported the stored result as complete. Historical
843
+ search and reads can now recover those originals without repeating the command.
844
+
845
+ Compact output carries its recovery path. Authenticated retention may also write
846
+ an artifact for a condensed result below the normal size cap. If retention fails,
847
+ the ordinary bounded original is shown instead; hook-redacted text stays redacted.
848
+
849
+ - 281859f: Correct recovery guidance in shortened tool output. The kernel no longer assumes that workspace `read`/`grep` tools can open internal retained-output paths. It directs recovery through the host-authorized tools and distinguishes the saved observation from a fresh read of its source. Existing permissions, exact retention and preview limits are unchanged; previously recorded previews are not rewritten.
850
+ - bdf923d: Add `/plugins` to inspect loaded plugins and enable or disable them for an idle session. The menu reports registered tools and skills, shows plugin scope and directory, and explains how to configure loading when it is off. Changes reset on restart or model switch; configuration and plugin files are retained. Active sends, compaction and durable resumes prevent plugin changes, and session cleanup waits for a pending change to settle.
851
+
852
+ Fix discovery when project and user plugin locations resolve to the same directory under an explicit application home. Load that directory once as user scope; project-only scope still excludes it. Distinct plugin directories remain discoverable even if their authority roots match.
853
+
3
854
  ## 38.2.1
4
855
 
5
856
  ### Patch Changes