open-multi-agent-kit 0.98.5 → 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (385) hide show
  1. package/CHANGELOG.md +48 -0
  2. package/README.md +4 -3
  3. package/dist/bun/cli.d.ts.map +1 -1
  4. package/dist/bun/cli.js +1 -0
  5. package/dist/bun/cli.js.map +1 -1
  6. package/dist/bun/register-bundled-coding-agent.d.ts +10 -0
  7. package/dist/bun/register-bundled-coding-agent.d.ts.map +1 -0
  8. package/dist/bun/register-bundled-coding-agent.js +12 -0
  9. package/dist/bun/register-bundled-coding-agent.js.map +1 -0
  10. package/dist/cli/args.d.ts +1 -1
  11. package/dist/cli/args.d.ts.map +1 -1
  12. package/dist/cli/args.js +1 -1
  13. package/dist/cli/args.js.map +1 -1
  14. package/dist/cli/help.d.ts.map +1 -1
  15. package/dist/cli/help.js +2 -1
  16. package/dist/cli/help.js.map +1 -1
  17. package/dist/cli.d.ts.map +1 -1
  18. package/dist/cli.js +14 -2
  19. package/dist/cli.js.map +1 -1
  20. package/dist/commands/neo-cli.d.ts +9 -0
  21. package/dist/commands/neo-cli.d.ts.map +1 -0
  22. package/dist/commands/neo-cli.js +61 -0
  23. package/dist/commands/neo-cli.js.map +1 -0
  24. package/dist/commands/verified-run-cli.d.ts.map +1 -1
  25. package/dist/commands/verified-run-cli.js +2 -2
  26. package/dist/commands/verified-run-cli.js.map +1 -1
  27. package/dist/core/agent-session.d.ts +29 -2
  28. package/dist/core/agent-session.d.ts.map +1 -1
  29. package/dist/core/agent-session.js +126 -32
  30. package/dist/core/agent-session.js.map +1 -1
  31. package/dist/core/bundled-skills.d.ts +4 -0
  32. package/dist/core/bundled-skills.d.ts.map +1 -0
  33. package/dist/core/bundled-skills.js +32 -0
  34. package/dist/core/bundled-skills.js.map +1 -0
  35. package/dist/core/cli-diagnostics.d.ts +6 -0
  36. package/dist/core/cli-diagnostics.d.ts.map +1 -0
  37. package/dist/core/cli-diagnostics.js +20 -0
  38. package/dist/core/cli-diagnostics.js.map +1 -0
  39. package/dist/core/context-budget-headroom.d.ts +12 -0
  40. package/dist/core/context-budget-headroom.d.ts.map +1 -1
  41. package/dist/core/context-budget-headroom.js +35 -0
  42. package/dist/core/context-budget-headroom.js.map +1 -1
  43. package/dist/core/context-budget-v2-input-validation.d.ts +4 -0
  44. package/dist/core/context-budget-v2-input-validation.d.ts.map +1 -0
  45. package/dist/core/context-budget-v2-input-validation.js +79 -0
  46. package/dist/core/context-budget-v2-input-validation.js.map +1 -0
  47. package/dist/core/context-budget-v2-observability.d.ts +15 -0
  48. package/dist/core/context-budget-v2-observability.d.ts.map +1 -0
  49. package/dist/core/context-budget-v2-observability.js +35 -0
  50. package/dist/core/context-budget-v2-observability.js.map +1 -0
  51. package/dist/core/context-budget-v2-planned-items.d.ts +5 -0
  52. package/dist/core/context-budget-v2-planned-items.d.ts.map +1 -0
  53. package/dist/core/context-budget-v2-planned-items.js +40 -0
  54. package/dist/core/context-budget-v2-planned-items.js.map +1 -0
  55. package/dist/core/context-budget-v2-planner.d.ts.map +1 -1
  56. package/dist/core/context-budget-v2-planner.js +57 -38
  57. package/dist/core/context-budget-v2-planner.js.map +1 -1
  58. package/dist/core/context-budget-v2-selection.d.ts +19 -3
  59. package/dist/core/context-budget-v2-selection.d.ts.map +1 -1
  60. package/dist/core/context-budget-v2-selection.js +35 -57
  61. package/dist/core/context-budget-v2-selection.js.map +1 -1
  62. package/dist/core/context-budget-v2-tiers.d.ts.map +1 -1
  63. package/dist/core/context-budget-v2-tiers.js +9 -42
  64. package/dist/core/context-budget-v2-tiers.js.map +1 -1
  65. package/dist/core/context-budget-v2-types.d.ts +8 -1
  66. package/dist/core/context-budget-v2-types.d.ts.map +1 -1
  67. package/dist/core/context-budget-v2-types.js.map +1 -1
  68. package/dist/core/devin-harness-dispatch.d.ts +12 -0
  69. package/dist/core/devin-harness-dispatch.d.ts.map +1 -0
  70. package/dist/core/devin-harness-dispatch.js +12 -0
  71. package/dist/core/devin-harness-dispatch.js.map +1 -0
  72. package/dist/core/devin-harness.d.ts +53 -0
  73. package/dist/core/devin-harness.d.ts.map +1 -0
  74. package/dist/core/devin-harness.js +112 -0
  75. package/dist/core/devin-harness.js.map +1 -0
  76. package/dist/core/domain-dispatch.d.ts +4 -1
  77. package/dist/core/domain-dispatch.d.ts.map +1 -1
  78. package/dist/core/domain-dispatch.js +5 -0
  79. package/dist/core/domain-dispatch.js.map +1 -1
  80. package/dist/core/domain-loadouts-provider-harness.d.ts +12 -0
  81. package/dist/core/domain-loadouts-provider-harness.d.ts.map +1 -0
  82. package/dist/core/domain-loadouts-provider-harness.js +122 -0
  83. package/dist/core/domain-loadouts-provider-harness.js.map +1 -0
  84. package/dist/core/domain-loadouts.d.ts +2 -33
  85. package/dist/core/domain-loadouts.d.ts.map +1 -1
  86. package/dist/core/domain-loadouts.js +3 -56
  87. package/dist/core/domain-loadouts.js.map +1 -1
  88. package/dist/core/domain-profile.d.ts +40 -0
  89. package/dist/core/domain-profile.d.ts.map +1 -0
  90. package/dist/core/domain-profile.js +7 -0
  91. package/dist/core/domain-profile.js.map +1 -0
  92. package/dist/core/extensions/bundled-virtual-modules.d.ts +15 -0
  93. package/dist/core/extensions/bundled-virtual-modules.d.ts.map +1 -0
  94. package/dist/core/extensions/bundled-virtual-modules.js +63 -0
  95. package/dist/core/extensions/bundled-virtual-modules.js.map +1 -0
  96. package/dist/core/extensions/loader.d.ts.map +1 -1
  97. package/dist/core/extensions/loader.js +2 -42
  98. package/dist/core/extensions/loader.js.map +1 -1
  99. package/dist/core/extensions/runner.d.ts +1 -0
  100. package/dist/core/extensions/runner.d.ts.map +1 -1
  101. package/dist/core/extensions/runner.js +6 -0
  102. package/dist/core/extensions/runner.js.map +1 -1
  103. package/dist/core/extensions/types.d.ts +10 -0
  104. package/dist/core/extensions/types.d.ts.map +1 -1
  105. package/dist/core/extensions/types.js.map +1 -1
  106. package/dist/core/grok-harness-dispatch.d.ts +6 -20
  107. package/dist/core/grok-harness-dispatch.d.ts.map +1 -1
  108. package/dist/core/grok-harness-dispatch.js +6 -55
  109. package/dist/core/grok-harness-dispatch.js.map +1 -1
  110. package/dist/core/grok-harness.d.ts +7 -9
  111. package/dist/core/grok-harness.d.ts.map +1 -1
  112. package/dist/core/grok-harness.js +9 -23
  113. package/dist/core/grok-harness.js.map +1 -1
  114. package/dist/core/harness-skills.d.ts +20 -0
  115. package/dist/core/harness-skills.d.ts.map +1 -0
  116. package/dist/core/harness-skills.js +36 -0
  117. package/dist/core/harness-skills.js.map +1 -0
  118. package/dist/core/index.d.ts +2 -1
  119. package/dist/core/index.d.ts.map +1 -1
  120. package/dist/core/index.js +2 -1
  121. package/dist/core/index.js.map +1 -1
  122. package/dist/core/loadout-runtime-state.d.ts +25 -0
  123. package/dist/core/loadout-runtime-state.d.ts.map +1 -0
  124. package/dist/core/loadout-runtime-state.js +39 -0
  125. package/dist/core/loadout-runtime-state.js.map +1 -0
  126. package/dist/core/loadout-runtime.d.ts +4 -20
  127. package/dist/core/loadout-runtime.d.ts.map +1 -1
  128. package/dist/core/loadout-runtime.js +6 -13
  129. package/dist/core/loadout-runtime.js.map +1 -1
  130. package/dist/core/mcp/client.d.ts +27 -6
  131. package/dist/core/mcp/client.d.ts.map +1 -1
  132. package/dist/core/mcp/client.js +80 -20
  133. package/dist/core/mcp/client.js.map +1 -1
  134. package/dist/core/mcp/manager.d.ts +10 -1
  135. package/dist/core/mcp/manager.d.ts.map +1 -1
  136. package/dist/core/mcp/manager.js +68 -9
  137. package/dist/core/mcp/manager.js.map +1 -1
  138. package/dist/core/mcp/protocol.d.ts +9 -3
  139. package/dist/core/mcp/protocol.d.ts.map +1 -1
  140. package/dist/core/mcp/protocol.js +61 -15
  141. package/dist/core/mcp/protocol.js.map +1 -1
  142. package/dist/core/mcp/stdio-transport.d.ts +12 -1
  143. package/dist/core/mcp/stdio-transport.d.ts.map +1 -1
  144. package/dist/core/mcp/stdio-transport.js +30 -2
  145. package/dist/core/mcp/stdio-transport.js.map +1 -1
  146. package/dist/core/mcp-descriptor-injection.d.ts +26 -0
  147. package/dist/core/mcp-descriptor-injection.d.ts.map +1 -0
  148. package/dist/core/mcp-descriptor-injection.js +25 -0
  149. package/dist/core/mcp-descriptor-injection.js.map +1 -0
  150. package/dist/core/mcp-public-presets.d.ts +1 -3
  151. package/dist/core/mcp-public-presets.d.ts.map +1 -1
  152. package/dist/core/mcp-public-presets.js +3 -15
  153. package/dist/core/mcp-public-presets.js.map +1 -1
  154. package/dist/core/model-registry.d.ts.map +1 -1
  155. package/dist/core/model-registry.js +24 -3
  156. package/dist/core/model-registry.js.map +1 -1
  157. package/dist/core/model-resolver.d.ts +1 -41
  158. package/dist/core/model-resolver.d.ts.map +1 -1
  159. package/dist/core/model-resolver.js +2 -49
  160. package/dist/core/model-resolver.js.map +1 -1
  161. package/dist/core/neo/catalog.d.ts +20 -0
  162. package/dist/core/neo/catalog.d.ts.map +1 -0
  163. package/dist/core/neo/catalog.js +50 -0
  164. package/dist/core/neo/catalog.js.map +1 -0
  165. package/dist/core/neo/setup.d.ts +3 -0
  166. package/dist/core/neo/setup.d.ts.map +1 -0
  167. package/dist/core/neo/setup.js +47 -0
  168. package/dist/core/neo/setup.js.map +1 -0
  169. package/dist/core/provider-default-models.d.ts +43 -0
  170. package/dist/core/provider-default-models.d.ts.map +1 -0
  171. package/dist/core/provider-default-models.js +51 -0
  172. package/dist/core/provider-default-models.js.map +1 -0
  173. package/dist/core/provider-display-names.d.ts.map +1 -1
  174. package/dist/core/provider-display-names.js +2 -0
  175. package/dist/core/provider-display-names.js.map +1 -1
  176. package/dist/core/provider-error-classification.d.ts +57 -0
  177. package/dist/core/provider-error-classification.d.ts.map +1 -0
  178. package/dist/core/provider-error-classification.js +103 -0
  179. package/dist/core/provider-error-classification.js.map +1 -0
  180. package/dist/core/provider-harness-dispatch.d.ts +63 -0
  181. package/dist/core/provider-harness-dispatch.d.ts.map +1 -0
  182. package/dist/core/provider-harness-dispatch.js +60 -0
  183. package/dist/core/provider-harness-dispatch.js.map +1 -0
  184. package/dist/core/provider-resilience.d.ts +10 -0
  185. package/dist/core/provider-resilience.d.ts.map +1 -1
  186. package/dist/core/provider-resilience.js +36 -3
  187. package/dist/core/provider-resilience.js.map +1 -1
  188. package/dist/core/provider-usage-commandcode.d.ts +9 -0
  189. package/dist/core/provider-usage-commandcode.d.ts.map +1 -0
  190. package/dist/core/provider-usage-commandcode.js +198 -0
  191. package/dist/core/provider-usage-commandcode.js.map +1 -0
  192. package/dist/core/provider-usage-devin.d.ts +18 -0
  193. package/dist/core/provider-usage-devin.d.ts.map +1 -0
  194. package/dist/core/provider-usage-devin.js +71 -0
  195. package/dist/core/provider-usage-devin.js.map +1 -0
  196. package/dist/core/provider-usage-text.d.ts +5 -0
  197. package/dist/core/provider-usage-text.d.ts.map +1 -0
  198. package/dist/core/provider-usage-text.js +15 -0
  199. package/dist/core/provider-usage-text.js.map +1 -0
  200. package/dist/core/provider-usage-types.d.ts +2 -1
  201. package/dist/core/provider-usage-types.d.ts.map +1 -1
  202. package/dist/core/provider-usage-types.js.map +1 -1
  203. package/dist/core/provider-usage.d.ts +1 -2
  204. package/dist/core/provider-usage.d.ts.map +1 -1
  205. package/dist/core/provider-usage.js +18 -7
  206. package/dist/core/provider-usage.js.map +1 -1
  207. package/dist/core/resource-admission.d.ts +36 -0
  208. package/dist/core/resource-admission.d.ts.map +1 -1
  209. package/dist/core/resource-admission.js +59 -0
  210. package/dist/core/resource-admission.js.map +1 -1
  211. package/dist/core/resource-loader.d.ts.map +1 -1
  212. package/dist/core/resource-loader.js +2 -2
  213. package/dist/core/resource-loader.js.map +1 -1
  214. package/dist/core/run-execution-api.d.ts +1 -1
  215. package/dist/core/run-execution-api.d.ts.map +1 -1
  216. package/dist/core/run-execution-api.js.map +1 -1
  217. package/dist/core/run-usage-ledger.d.ts +47 -0
  218. package/dist/core/run-usage-ledger.d.ts.map +1 -0
  219. package/dist/core/run-usage-ledger.js +162 -0
  220. package/dist/core/run-usage-ledger.js.map +1 -0
  221. package/dist/core/run-usage-operation.d.ts +8 -0
  222. package/dist/core/run-usage-operation.d.ts.map +1 -0
  223. package/dist/core/run-usage-operation.js +12 -0
  224. package/dist/core/run-usage-operation.js.map +1 -0
  225. package/dist/core/sdk.d.ts.map +1 -1
  226. package/dist/core/sdk.js +20 -16
  227. package/dist/core/sdk.js.map +1 -1
  228. package/dist/core/session-failure-cause.d.ts.map +1 -1
  229. package/dist/core/session-failure-cause.js +14 -5
  230. package/dist/core/session-failure-cause.js.map +1 -1
  231. package/dist/core/session-prompt-lifecycle.d.ts +6 -0
  232. package/dist/core/session-prompt-lifecycle.d.ts.map +1 -1
  233. package/dist/core/session-prompt-lifecycle.js +47 -2
  234. package/dist/core/session-prompt-lifecycle.js.map +1 -1
  235. package/dist/core/session-termination.d.ts.map +1 -1
  236. package/dist/core/session-termination.js +4 -2
  237. package/dist/core/session-termination.js.map +1 -1
  238. package/dist/core/subagent-lane-authority.d.ts +28 -0
  239. package/dist/core/subagent-lane-authority.d.ts.map +1 -0
  240. package/dist/core/subagent-lane-authority.js +119 -0
  241. package/dist/core/subagent-lane-authority.js.map +1 -0
  242. package/dist/core/subagent-lane-contract.d.ts +114 -0
  243. package/dist/core/subagent-lane-contract.d.ts.map +1 -0
  244. package/dist/core/subagent-lane-contract.js +18 -0
  245. package/dist/core/subagent-lane-contract.js.map +1 -0
  246. package/dist/core/subagent-lane-launcher.d.ts +3 -6
  247. package/dist/core/subagent-lane-launcher.d.ts.map +1 -1
  248. package/dist/core/subagent-lane-launcher.js +53 -12
  249. package/dist/core/subagent-lane-launcher.js.map +1 -1
  250. package/dist/core/subagent-orchestration.d.ts +4 -17
  251. package/dist/core/subagent-orchestration.d.ts.map +1 -1
  252. package/dist/core/subagent-orchestration.js.map +1 -1
  253. package/dist/core/verified-run/broker.d.ts.map +1 -1
  254. package/dist/core/verified-run/broker.js +4 -1
  255. package/dist/core/verified-run/broker.js.map +1 -1
  256. package/dist/core/verified-run/dag-phase.d.ts +1 -1
  257. package/dist/core/verified-run/dag-phase.d.ts.map +1 -1
  258. package/dist/core/verified-run/dag-phase.js +46 -8
  259. package/dist/core/verified-run/dag-phase.js.map +1 -1
  260. package/dist/core/verified-run/dag-projection.d.ts.map +1 -1
  261. package/dist/core/verified-run/dag-projection.js +25 -12
  262. package/dist/core/verified-run/dag-projection.js.map +1 -1
  263. package/dist/core/verified-run/dag-types.d.ts +11 -0
  264. package/dist/core/verified-run/dag-types.d.ts.map +1 -1
  265. package/dist/core/verified-run/dag-types.js.map +1 -1
  266. package/dist/core/verified-run/event-parser.d.ts.map +1 -1
  267. package/dist/core/verified-run/event-parser.js +1 -0
  268. package/dist/core/verified-run/event-parser.js.map +1 -1
  269. package/dist/core/verified-run/owned-execution.d.ts +1 -0
  270. package/dist/core/verified-run/owned-execution.d.ts.map +1 -1
  271. package/dist/core/verified-run/owned-execution.js +7 -1
  272. package/dist/core/verified-run/owned-execution.js.map +1 -1
  273. package/dist/core/verified-run/process-projection.d.ts +13 -0
  274. package/dist/core/verified-run/process-projection.d.ts.map +1 -0
  275. package/dist/core/verified-run/process-projection.js +82 -0
  276. package/dist/core/verified-run/process-projection.js.map +1 -0
  277. package/dist/core/verified-run/projection.d.ts.map +1 -1
  278. package/dist/core/verified-run/projection.js +11 -55
  279. package/dist/core/verified-run/projection.js.map +1 -1
  280. package/dist/core/verified-run/run-types.d.ts +1 -0
  281. package/dist/core/verified-run/run-types.d.ts.map +1 -1
  282. package/dist/core/verified-run/run-types.js.map +1 -1
  283. package/dist/core/verified-run/task-execution.d.ts +9 -0
  284. package/dist/core/verified-run/task-execution.d.ts.map +1 -0
  285. package/dist/core/verified-run/task-execution.js +29 -0
  286. package/dist/core/verified-run/task-execution.js.map +1 -0
  287. package/dist/core/verified-run/writer-projection.d.ts.map +1 -1
  288. package/dist/core/verified-run/writer-projection.js +6 -5
  289. package/dist/core/verified-run/writer-projection.js.map +1 -1
  290. package/dist/core/workload-permit-pool.d.ts +1 -1
  291. package/dist/core/workload-permit-pool.d.ts.map +1 -1
  292. package/dist/core/workload-permit-pool.js +3 -0
  293. package/dist/core/workload-permit-pool.js.map +1 -1
  294. package/dist/guardrails/strict-evidence-approval-adapter.d.ts +34 -0
  295. package/dist/guardrails/strict-evidence-approval-adapter.d.ts.map +1 -0
  296. package/dist/guardrails/strict-evidence-approval-adapter.js +70 -0
  297. package/dist/guardrails/strict-evidence-approval-adapter.js.map +1 -0
  298. package/dist/index.d.ts +3 -0
  299. package/dist/index.d.ts.map +1 -1
  300. package/dist/index.js +1 -0
  301. package/dist/index.js.map +1 -1
  302. package/dist/main.d.ts.map +1 -1
  303. package/dist/main.js +11 -18
  304. package/dist/main.js.map +1 -1
  305. package/dist/modes/acp/acp-agent.d.ts +24 -0
  306. package/dist/modes/acp/acp-agent.d.ts.map +1 -0
  307. package/dist/modes/acp/acp-agent.js +133 -0
  308. package/dist/modes/acp/acp-agent.js.map +1 -0
  309. package/dist/modes/acp/acp-mode.d.ts +4 -0
  310. package/dist/modes/acp/acp-mode.d.ts.map +1 -0
  311. package/dist/modes/acp/acp-mode.js +34 -0
  312. package/dist/modes/acp/acp-mode.js.map +1 -0
  313. package/dist/modes/acp/acp-session.d.ts +4 -0
  314. package/dist/modes/acp/acp-session.d.ts.map +1 -0
  315. package/dist/modes/acp/acp-session.js +76 -0
  316. package/dist/modes/acp/acp-session.js.map +1 -0
  317. package/dist/modes/acp/acp-transport.d.ts +5 -0
  318. package/dist/modes/acp/acp-transport.d.ts.map +1 -0
  319. package/dist/modes/acp/acp-transport.js +90 -0
  320. package/dist/modes/acp/acp-transport.js.map +1 -0
  321. package/dist/modes/interactive/interactive-mode.d.ts.map +1 -1
  322. package/dist/modes/interactive/interactive-mode.js +3 -5
  323. package/dist/modes/interactive/interactive-mode.js.map +1 -1
  324. package/dist/modes/interactive/resource-description.d.ts +3 -0
  325. package/dist/modes/interactive/resource-description.d.ts.map +1 -0
  326. package/dist/modes/interactive/resource-description.js +15 -0
  327. package/dist/modes/interactive/resource-description.js.map +1 -0
  328. package/docs/audit-revalidation-2026-09-17.md +77 -0
  329. package/docs/devin-harness.md +150 -0
  330. package/docs/docs.json +4 -0
  331. package/docs/ecraf-normalization.md +83 -0
  332. package/docs/environment-variables.md +2 -1
  333. package/docs/index.md +1 -0
  334. package/docs/loadout-domains/README.md +2 -1
  335. package/docs/loadout-domains/devin-harness.md +72 -0
  336. package/docs/metrics.md +57 -16
  337. package/docs/model-catalog-refresh.md +51 -1
  338. package/docs/models.md +1 -1
  339. package/docs/neo.md +134 -0
  340. package/docs/providers.md +143 -1
  341. package/docs/release-audit-0.98.5.md +17 -0
  342. package/docs/release-audit-0.99.0.md +68 -0
  343. package/docs/run-protocol.md +4 -3
  344. package/docs/run-usage-ledger.md +56 -0
  345. package/docs/runtime-algorithms.md +48 -1
  346. package/docs/sdk.md +9 -5
  347. package/docs/settings.md +1 -1
  348. package/docs/skills.md +1 -1
  349. package/docs/startup-resource-labels-testing.md +47 -0
  350. package/docs/tb21-audit.md +18 -4
  351. package/docs/usage.md +7 -1
  352. package/docs/verified-run-remaining-design.md +881 -0
  353. package/docs/verified-run-testing.md +91 -0
  354. package/docs/verified-run.md +30 -9
  355. package/examples/extensions/custom-provider-anthropic/package-lock.json +2 -2
  356. package/examples/extensions/custom-provider-anthropic/package.json +1 -1
  357. package/examples/extensions/custom-provider-gitlab-duo/package.json +1 -1
  358. package/examples/extensions/gondolin/package-lock.json +2 -2
  359. package/examples/extensions/gondolin/package.json +1 -1
  360. package/examples/extensions/sandbox/package-lock.json +2 -2
  361. package/examples/extensions/sandbox/package.json +1 -1
  362. package/examples/extensions/subagent/adaptive-agent-runtime.ts +14 -1
  363. package/examples/extensions/subagent/graph-result.ts +48 -0
  364. package/examples/extensions/subagent/index.ts +246 -111
  365. package/examples/extensions/subagent/managed-process-tree.ts +42 -0
  366. package/examples/extensions/subagent/managed-process.test.ts +24 -0
  367. package/examples/extensions/subagent/managed-process.ts +93 -112
  368. package/examples/extensions/subagent/subagent-runtime-types.ts +17 -1
  369. package/examples/extensions/subagent/subagent-stream.ts +161 -0
  370. package/examples/extensions/terminal-browser/README.md +54 -0
  371. package/examples/extensions/terminal-browser/bridge-protocol.ts +29 -0
  372. package/examples/extensions/terminal-browser/bridge.ts +219 -0
  373. package/examples/extensions/terminal-browser/browser-surface.ts +288 -0
  374. package/examples/extensions/terminal-browser/index.ts +215 -0
  375. package/examples/extensions/terminal-browser/placeholders.ts +49 -0
  376. package/examples/extensions/with-deps/package-lock.json +2 -2
  377. package/examples/extensions/with-deps/package.json +1 -1
  378. package/npm-shrinkwrap.json +18 -18
  379. package/package.json +8 -7
  380. package/resources/neo/skills/omk-browser/SKILL.md +32 -0
  381. package/resources/neo/skills/omk-code-review/SKILL.md +24 -0
  382. package/resources/neo/skills/omk-computeruse/SKILL.md +34 -0
  383. package/resources/neo/skills/omk-mcp-setup/SKILL.md +48 -0
  384. package/resources/neo/skills/omk-research/SKILL.md +24 -0
  385. package/resources/neo/skills/omk-site/SKILL.md +28 -0
package/docs/models.md CHANGED
@@ -233,7 +233,7 @@ Current behavior:
233
233
 
234
234
  ### Thinking Level Map
235
235
 
236
- Use `thinkingLevelMap` on a model to describe model-specific thinking controls. `models.json` accepts `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, and `max`. Built-in metadata can also expose the `ultra` CLI tier, but the current `models.json` schema has no `ultra` key.
236
+ Use `thinkingLevelMap` on a model to describe model-specific thinking controls. `models.json` accepts `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, and `max`. Built-in metadata can also expose the `ultra` CLI tier (GPT-6 Astra maps it to documented `max`; GPT-5.6 Sol/Terra/MoA on Codex map it to backend `xhigh`), but the current `models.json` schema has no `ultra` key.
237
237
 
238
238
  | Value | Meaning |
239
239
  | --- | --- |
package/docs/neo.md ADDED
@@ -0,0 +1,134 @@
1
+ # Neo public skills and computer-use setup
2
+
3
+ Status: locally integrated and built on 2026-09-17, not a published release. Neo names the public computer-use workflow in this change; it is not a new model or an independent agent runtime.
4
+
5
+ ## What changes
6
+
7
+ The normal `DefaultResourceLoader` now reads six publicly authored skills from `resources/neo/skills` inside the installed package. No private agent-home content is copied. Explicit, project and user skills retain ownership of matching names; bundled skills fill only missing names. `noSkills` disables automatic bundled loading while preserving explicitly supplied skills, and `OMK_BUNDLED_SKILLS=0` disables just the bundled set. Skill overrides continue to run after resource loading.
8
+
9
+ Both the npm package allowlist and the two binary asset paths include the same `resources` tree. Files remain inside the installation; startup does not write skills into the user's home. A missing bundle produces a loader diagnostic, and `omk neo list` returns a nonzero exit code if one of its required skill files is missing.
10
+
11
+ | Default skill | Purpose |
12
+ | --- | --- |
13
+ | `omk-computeruse` | Neo target/capability routing and observe, act, verify workflow |
14
+ | `omk-browser` | DOM/accessibility interaction, bounded retries and side-effect approval |
15
+ | `omk-site` | Existing-project site implementation, preview and responsive validation |
16
+ | `omk-research` | Version-aware primary documentation and evidence handling |
17
+ | `omk-code-review` | Minimal changes, real regression checks and live-call-path review |
18
+ | `omk-mcp-setup` | Inventory, configuration preview, explicit setup and health verification |
19
+
20
+ These instructions are available independently of the selected model family. Tool calling must actually work. Image interpretation additionally requires compatible model and provider input support. Text-only models should use DOM/accessibility observations. This change does not establish Astra-equivalent performance, add a native desktop driver, or turn an incapable model into a computer-use model.
21
+
22
+ ## Skills and MCPs are different products
23
+
24
+ | Quantity | Meaning |
25
+ | --- | --- |
26
+ | Packaged skill | Instruction file is present in the installation |
27
+ | Loaded skill | Resource loader admitted the file after overrides and disable settings |
28
+ | Offered MCP | Public configuration preset is available |
29
+ | Configured MCP | User selected a server in a configuration file |
30
+ | Connected MCP | The running process completed its connection/handshake |
31
+ | Verified tool | A discovered tool succeeded against the intended target |
32
+
33
+ `omk neo list` reports packaged files and offered presets. It deliberately returns `connected: null` and `connectionStatus: "not_probed"`; it does not manufacture a nonzero MCP connection count. The normal session MCP status remains the connection authority.
34
+
35
+ ## Curated MCP choices
36
+
37
+ | Preset | Distribution | Activation | Scope |
38
+ | --- | --- | --- | --- |
39
+ | `playwright` | Pinned stdio recipe for `@playwright/mcp@0.0.81` | Explicit user setup | Browser actions and accessibility snapshots; compatible browser required |
40
+ | `context7` | Pinned stdio recipe for `@upstash/context7-mcp@4.1.1` | Explicit user setup | External technical documentation queries |
41
+ | Native CUA/desktop driver | Not installed or configured by this change | Separate reviewed setup | Host permissions and host identity must be verified |
42
+ | Chrome personal-profile attachment | Not enabled by default | Separate explicit session approval | May expose existing authenticated state |
43
+ | Stagehand and Browserbase cloud | Not installed or configured by this change | Separate dependency, credentials and cost approval | Optional local library or cloud browser |
44
+ | Account-connected repository/services | Not configured by this change | Per-user authentication and scope approval | No publisher credential is shared |
45
+
46
+ The recipes pin the direct package version, not an entire third-party transitive dependency graph. They do not bundle executables or browser binaries. Installation and tool behavior still need to be tested on the target platform. No remote HTTP URL is advertised as directly loadable by the current stdio-only OMK loader.
47
+
48
+ ## Commands
49
+
50
+ ```bash
51
+ omk neo list
52
+ omk neo setup playwright context7
53
+ omk neo mcp-config playwright context7
54
+ ```
55
+
56
+ All three are local and do not start a process, download a server or connect a browser. `setup` without `--apply` is a dry run. `mcp-config` emits disabled entries only. Existing configuration is never printed.
57
+
58
+ After approving the exact target and the server execution/network effects:
59
+
60
+ ```bash
61
+ # Project-local, in the intended project directory:
62
+ omk neo setup playwright context7 --apply
63
+
64
+ # Alternative: explicitly select the global configuration target:
65
+ omk neo setup playwright context7 --global --apply
66
+ ```
67
+
68
+ Project target: `<cwd>/.omk/mcp.json`. Global target: `~/.omk/mcp.json`, not `~/.omk/agent/mcp.json`. On the next ordinary OMK startup `npx` may download and execute the selected pinned servers. The command itself only creates configuration and reports `configured_not_connected`. A project entry can override a global entry with the same server name.
69
+
70
+ Setup is intentionally no-overwrite. It validates all presets before filesystem changes, writes a complete temporary file with mode `0600` and publishes it through an atomic no-replace hard link. An existing target, including malformed data or a symlink, is preserved. A symlinked `.omk` directory is rejected. Filesystems that do not support hard links fail closed. This is not a defense against a hostile process concurrently replacing an ancestor directory; use a trusted user-owned configuration location. Windows ACL and power-loss durability guarantees are outside this implementation.
71
+
72
+ For an existing configuration, generate disabled entries with `mcp-config`, then review a manual merge without exposing secrets. Do not replace the file to work around the refusal. This change does not alter legacy MCP config precedence or overwrite any account settings.
73
+
74
+ ## Browser safety and capability
75
+
76
+ The Playwright recipe requests `--headless --isolated --sandbox` and bounded action/navigation timeouts. It does not attach to a personal browser, disable TLS checks or request unrestricted file access. The upstream Linux bundled-browser default can disable its sandbox; the explicit positive flag is therefore intentional. A sandbox launch failure must be diagnosed rather than bypassed.
77
+
78
+ Profile isolation is not OS or network isolation. Origin allowlists are not complete redirect or SSRF boundaries. Operator-managed process/network isolation is required for stronger threat models. The workflow rejects instruction authority from pages, tool output and server descriptions. Skills guide the model; they do not independently enforce every permission boundary in the runtime.
79
+
80
+ After configuration, restart the intended OMK process, inspect the actual connection status and tool roster, and perform a read-only health check. Then use an observation before each action and verify the final requested state. A timeout is not permission to repeat a non-idempotent operation. Native desktop, paid cloud and private-account actions remain separate opt-ins.
81
+
82
+ ## Verification and release gates
83
+
84
+ Harness impact: `advance` for distributable capability discovery; `preserve` for existing skill precedence, disable behavior and configuration ownership. Baseline is `739bc6f3b6fe1c89bfe058aad7ec17ad252fb10e` (`0.99.0`).
85
+
86
+ Acceptance: six valid files in the npm and each platform bundle; six default skills in an empty agent/project environment; zero writes on list/dry run; byte-identical preservation of existing MCP config; no false connected count. No performance superiority claim is made.
87
+
88
+ ```bash
89
+ node --experimental-strip-types --test scripts/test/neo-distribution.test.mjs
90
+ cd packages/coding-agent
91
+ node ../../node_modules/vitest/dist/cli.js --run test/neo-distribution.test.ts
92
+ cd ../..
93
+ npm run check
94
+ npm run build
95
+ ```
96
+
97
+ The Node test covers actual catalog/CLI/setup functions, file preservation, symlinks and asset inclusion declarations. The Vitest test additionally exercises the real skill loader and `DefaultResourceLoader`; it requires the repository dependencies and a buildable base. Source inspection of binary copy commands is not a built-binary smoke test. Each final archive must be extracted and `omk neo list` checked before release.
98
+
99
+ Before publishing, run the targeted tests and complete repository gates, verify a clean npm install and real MCP/browser smoke tests, then follow the existing lockstep release procedure. A green test for a leaf module cannot override a failed monorepo build. Do not tag/publish this candidate until all release surfaces agree and blockers are resolved.
100
+
101
+ ## Local installation verification, 2026-09-17
102
+
103
+ Integrated the candidate onto `887152bc5d`. `npm run build` and `tsgo --noEmit`
104
+ passed. Standalone source tests passed 13 checks; Neo, resource-loader and SDK skill
105
+ Vitest suites passed 39 tests. The Neo fixture isolates HOME and the agent directory
106
+ so personal skills cannot contaminate its six-skill assertion.
107
+
108
+ The existing local launcher resolves this checkout's `dist/cli.js`; its `neo list`
109
+ reported all six packaged skills. The built resource loader also selected six skills
110
+ in a clean environment. `npm pack --ignore-scripts --dry-run --json` included all six.
111
+ Run `node --test scripts/test/neo-built-cli.test.mjs` after building for subprocess
112
+ checks of list, write-free preview, apply and no-replace behavior.
113
+
114
+ No live MCP configuration was created or overwritten. The existing Playwright
115
+ connection answered a tab-list request and opened/closed a task-owned blank tab;
116
+ this does not verify the candidate preset's browser or sandbox flags. Existing
117
+ Context7 returned an invalid API-key error, so documentation retrieval is blocked.
118
+ No credentials were changed and no native desktop driver was installed.
119
+
120
+ Full `npm run check` stopped at pre-existing module-size failures in `agent-session.ts`
121
+ and `model-registry.ts`. The documentation link guard also rejects the new untracked
122
+ `neo.md` until it is included in Git; no files were staged to bypass that gate.
123
+ Release-surface and shrinkwrap checks passed. Platform archives, clean dependency
124
+ installation, all-suite tests and published release verification remain unrun.
125
+ Restart OMK to load rebuilt code; changing source does not update an active process.
126
+
127
+ ## Primary source check, 2026-09-16
128
+
129
+ - Playwright MCP release `v0.0.81`: https://github.com/microsoft/playwright-mcp/releases/tag/v0.0.81
130
+ - Exact documented flags: https://github.com/microsoft/playwright-mcp/blob/v0.0.81/README.md
131
+ - Context7 release `@upstash/context7-mcp@4.1.1`: https://github.com/upstash/context7/releases/tag/%40upstash/context7-mcp%404.1.1
132
+ - Context7 stdio CLI: https://github.com/upstash/context7/blob/master/packages/mcp/src/index.ts
133
+
134
+ Upstream source/release inspection is not an installed-server integration test. Credentials, provider quality and target-host permissions are not verified by this document.
package/docs/providers.md CHANGED
@@ -19,12 +19,13 @@ Use `/login` in interactive mode, then select a provider:
19
19
  - Claude Pro/Max
20
20
  - GitHub Copilot
21
21
  - xAI Grok subscription OAuth
22
+ - Devin CLI subscription (SWE-2)
22
23
 
23
24
  Run `/login` and choose a configured subscription provider to open its account picker. Select an existing account by its ChatGPT, Claude, or Google email when available, or choose **Add another account** to sign in with a new one. OMK keeps and refreshes each account independently, pins the provider to the account you select, and does not silently fail over to another subscription. `/model` remains dedicated to model selection.
24
25
 
25
26
  Use `/logout` to clear all stored accounts for a provider. Tokens are stored in `~/.omk/agent/auth.json` and auto-refresh when expired.
26
27
 
27
- When the status sidebar is pinned, its **USAGE** section lists every configured subscription provider, with the active provider first. OMK reads quota windows from fixed provider endpoints for Codex, Claude, Kimi Code, GLM/ZAI Coding Plan, and native xAI SuperGrok, caches the result, and displays each percentage and reset countdown separately. Claude also passively merges the official `anthropic-ratelimit-unified-*` response headers used by Claude Code. If Anthropic's usage endpoint is rate limited and no complete recent snapshot exists, OMK mirrors Claude Code's own startup quota check with one fixed-endpoint Haiku request capped at one output token, no more than once per OAuth credential per hour. This fallback consumes a small amount of Claude plan quota.
28
+ When the status sidebar is pinned, its **USAGE** section lists every configured subscription provider, with the active provider first. OMK reads quota windows from fixed provider endpoints for Codex, Claude, Kimi Code, GLM/ZAI Coding Plan, native xAI SuperGrok, Devin CLI, and Command Code, caches the result, and displays each percentage and reset countdown separately. Claude also passively merges the official `anthropic-ratelimit-unified-*` response headers used by Claude Code. If Anthropic's usage endpoint is rate limited and no complete recent snapshot exists, OMK mirrors Claude Code's own startup quota check with one fixed-endpoint Haiku request capped at one output token, no more than once per OAuth credential per hour. This fallback consumes a small amount of Claude plan quota.
28
29
 
29
30
  Codex streaming passively merges `x-codex-primary-*`, `x-codex-secondary-*`, and `codex.rate_limits` signals through the non-blocking `StreamOptions.onRateLimit` observer. These signals supplement missing polling windows only when the Codex service returns them; OMK does not infer a missing 5-hour value from a 7-day value.
30
31
 
@@ -32,6 +33,8 @@ Alibaba Model Studio Token Plan is recognized as **QWEN TOKEN PLAN** and reads i
32
33
 
33
34
  With a stored native `xai` OAuth credential, OMK reads `GET https://cli-chat-proxy.grok.com/v1/billing?format=credits` and shows the weekly SuperGrok pool from `config.creditUsagePercent` plus its reset from `config.currentPeriod.end`. `XAI_API_KEY` is a separate API-billing credential and does not authorize this subscription endpoint.
34
35
 
36
+ With a configured `commandcode` API key (`models.json` or auth storage), OMK reads `GET https://api.commandcode.ai/alpha/whoami`, then `/alpha/billing/credits`, `/alpha/billing/subscriptions`, and `/alpha/usage/summary` on that same origin. The rail shows the 5-hour and weekly credit windows plus the monthly pool (`spent / remaining+spent`) and the plan name. OMK never sends the key anywhere except `api.commandcode.ai`, never reads browser cookies, and does not scrape Studio.
37
+
35
38
  ### Model Studio DeepSeek V4
36
39
 
37
40
  Model Studio uses `enable_thinking`, including for DeepSeek. OMK's Chat Completions
@@ -67,6 +70,135 @@ and [plan endpoint separation](https://www.alibabacloud.com/help/en/model-studio
67
70
  consulted 2026-09-08. The local tests exercise serialization, not account availability,
68
71
  provider compliance, billing, or benchmark performance.
69
72
 
73
+ ### Devin CLI
74
+
75
+ Run `/login devin`, select `devin/swe-2`, and use `/think medium`, `/think high`,
76
+ or `/think max`. For a new session after login:
77
+
78
+ ```bash
79
+ omk --provider devin --model swe-2 --thinking max
80
+ ```
81
+
82
+ The default is `medium`. Thinking-off and other levels are unsupported. `max`
83
+ is a reasoning level, not a requirement to buy the Devin Max subscription tier.
84
+ Your account's model access and quota still apply. For presets, effort guidance,
85
+ the 1M-token context budget, and the `devin-harness` loadout, see the
86
+ [Devin SWE-2 harness](devin-harness.md).
87
+
88
+ The `devin` catalog also carries every other lane `GetCliModelConfigs`
89
+ advertises — each wire UID is its own logical model (`claude-opus-5-high`,
90
+ `gpt-5-6-sol-xhigh`, `gemini-3-8-flash-medium`, `kimi-k3-max`, `glm-5-3-high`,
91
+ `grok-4-6-xhigh`, `deepseek-v4-pro-max`, `swe-1-7`, `inkling-max`, …). Flat
92
+ models pin their declared effort lane, so `/think` is fixed per model; entries
93
+ whose lane is a no-thinking variant report `reasoning: false`. Availability is
94
+ account- and plan-dependent: a lane absent from your catalog fails loudly
95
+ instead of being remapped. Image input is still text-only at the adapter.
96
+
97
+ ```bash
98
+ omk --provider devin --model claude-opus-5-high
99
+ ```
100
+
101
+ **Authentication.** OMK uses the CLI's PKCE flow at
102
+ `app.devin.ai/auth/cli/continue` and `api.devin.ai/auth/cli/token`. Its callback
103
+ binds only to `127.0.0.1:59653`, validates state, and closes after success,
104
+ rejection, cancellation, or a five-minute deadline. If the port is occupied,
105
+ paste the complete callback URL into OMK's local prompt. For remote sessions,
106
+ forward this loopback port to the machine running OMK.
107
+
108
+ Credentials use the existing protected `auth.json` storage and account picker.
109
+ Expired sessions require `/login devin` again: no refresh endpoint is verified,
110
+ and OMK does not renew the expiry locally. Use `/logout` to remove the stored
111
+ OMK credential. Alternatively, `DEVIN_API_KEY` accepts an already-owned CLI
112
+ session token, **not** a `cog_` REST API key. OMK does not install or invoke the
113
+ Devin agent, read browser cookies, or import another application's credentials.
114
+
115
+ **Transport and limits.** The Node-only `devin-agent` adapter uses Connect/protobuf
116
+ at the fixed HTTPS origin `https://server.codeium.com`. It exchanges the session
117
+ token for a user JWT and reads `GetCliModelConfigs` before each turn. The logical
118
+ `swe-2` model resolves its effort through the server's SWE-2 family metadata;
119
+ every other catalog model resolves by its own wire UID. OMK never invents a
120
+ `swe-2-max` ID or downgrades an unavailable route. Disabled, internal, and
121
+ ambiguous entries are excluded; fast-lane entries match only when their own UID
122
+ is selected. The family's separate 1M-context entries form
123
+ a second lane that is selected only when the model's local `contextWindow` is
124
+ 1,000,000 or more (the bundled default); a smaller budget uses the standard lane.
125
+
126
+ Text, thinking, tool calls/results, and reported usage enter the normal OMK
127
+ agent loop. Image input and enterprise-origin overrides are unsupported.
128
+ Requests reject redirects. Malformed, oversized, or unterminated streams fail;
129
+ remote error bodies are not copied into diagnostics. SDK `onPayload` receives
130
+ protobuf bytes without credential metadata.
131
+
132
+ The bundled 1,000,000-token context budget and 16,384-token output cap are
133
+ local defaults, **not published SWE-2 limits**. Output is further capped against
134
+ the authenticated catalog. If the selected lane declares a smaller context window,
135
+ the request fails and names the window; lower `contextWindow` through
136
+ `modelOverrides` in [models.json](models.md#per-model-overrides) rather than
137
+ expecting a silent downgrade. Zero catalog pricing means unpriced subscription
138
+ usage, not free inference. Quota percentages are not inferred.
139
+
140
+ **Account quota.** With a Devin credential configured, the status rail's USAGE
141
+ section calls `SeatManagementService/GetUserStatus` (session-token metadata, no
142
+ user JWT) and renders the plan's daily and weekly quota meters with reset times.
143
+ Accounts whose plan reports no quota windows show the plan name and credit
144
+ balances instead. The rail mirrors the CLI's `/usage` surface; it never sends
145
+ the session token anywhere except `server.codeium.com`.
146
+
147
+ **Verification.** Local tests exercise the public stream API, real loopback
148
+ callbacks with mocked token exchange, model selection, and protocol fixtures.
149
+ Live login, catalog compatibility, SWE-2 inference, and billing remain unverified
150
+ without a Devin account. Unauthenticated catalog probes returned HTTP 400.
151
+ Live tests require `DEVIN_API_KEY` and `LIVE_E2E=1` and consume subscription quota.
152
+
153
+ Sources consulted 2026-09-12 and 2026-09-13: the [SWE-2 announcement](https://cognition.com/blog/swe-2)
154
+ (2026-09-10; CLI availability and medium/high/max; no published context window), [CLI commands](https://docs.devin.ai/cli/reference/commands),
155
+ and the [official manifest](https://static.devin.ai/cli/current/manifest.json)
156
+ (observed identity `3000.10.21`). Protocol fields follow the third-party
157
+ [oh-my-pi snapshot](https://github.com/can1357/oh-my-pi/blob/942383f768c5f2f6a620fcab57326c0f59df623a/packages/catalog/src/discovery/devin-proto.ts),
158
+ not a stable public inference contract; attribution is in `packages/ai/DEVIN-NOTICE`.
159
+ Regenerate only the Devin models, preserving every other catalog entry, with
160
+ `npm --prefix packages/ai run generate-models -- --devin-only`.
161
+
162
+ ### Cursor
163
+
164
+ Run `/login cursor` (PKCE + polling against `cursor.com`/`api2.cursor.sh`), then
165
+ select a `cursor/*` model. `cursor/default` is the Auto router; sibling slugs
166
+ carry their effort and lane (`-low|medium|high|xhigh|max`, `-thinking-*`,
167
+ `-fast`) in the id, mirroring the CLI's flat list. `CURSOR_API_KEY` accepts an
168
+ already-owned access token instead of the OAuth flow.
169
+
170
+ ```bash
171
+ omk --provider cursor --model default
172
+ ```
173
+
174
+ **Transport and limits.** The Node-only `cursor-agent` adapter runs the
175
+ `agent.v1.AgentService/Run` bidirectional RPC over HTTP/2 at `api2.cursor.sh`
176
+ with Connect-framed protobuf. Each `stream()` call is one `Run`:
177
+ `userMessageAction` for a fresh user turn or `resumeAction` otherwise; history
178
+ is replayed through `conversationState.rootPromptMessagesJson` — SHA256-keyed
179
+ JSON message blobs the server fetches back via the kv channel — and system
180
+ prompts ride `requestContext.rules` in the exec handshake. For OpenAI-family
181
+ sibling slugs the adapter splits the effort into a `reasoning` request
182
+ parameter because the Run endpoint rejects sibling `model_id`s; other ids pass
183
+ through whole. `tokenDelta` frames accumulate output usage; plan-specific
184
+ model availability is enforced server-side (free plans serve Auto only).
185
+
186
+ Text, thinking, and reported usage enter the normal OMK agent loop; images
187
+ attach through `selectedContext.selectedImages`. Server exec frames that ask
188
+ the client to run tools are answered with the protocol's `throw` failure
189
+ channel — OMK tools are deliberately not advertised inside the provider, so a
190
+ cursor model reports the capability gap instead of bypassing tool governance.
191
+ Wiring exec frames through governed OMK tools is a coding-agent bridge
192
+ feature, not provider scope. Hosted search/fetch permission queries are
193
+ approved; interactive ones are rejected.
194
+
195
+ The bundled catalog was captured from `GetUsableModels` (2026-09-18; 223
196
+ entries). Regenerate only the Cursor section with
197
+ `npm --prefix packages/ai run generate-models -- --cursor-only`. Protocol
198
+ fields follow the same third-party snapshot attribution as the Devin adapter.
199
+ Live login, plan gating, and quota behavior remain account-dependent; the
200
+ adapter is covered by loopback HTTP/2 protocol tests.
201
+
70
202
  ### OpenAI Codex
71
203
 
72
204
  - Requires ChatGPT Plus or Pro subscription
@@ -78,6 +210,14 @@ provider compliance, billing, or benchmark performance.
78
210
  omk --provider openai-codex --model gpt-5.6-moa --thinking ultra
79
211
  ```
80
212
 
213
+ ### GPT-6 Astra
214
+
215
+ OpenAI-shaped Astra routes (`openai`, `azure-openai-responses`, `opencode`, `github-copilot`, OpenRouter) expose OMK `ultra` in the selector and send `reasoning.effort: "max"`. Official Astra effort values remain `low`, `medium`, `high`, `xhigh`, and `max`; there is no native `ultra` wire value. ChatGPT/Codex account catalogs may still reject `gpt-6-astra`.
216
+
217
+ ```bash
218
+ omk --provider openai --model gpt-6-astra --thinking ultra
219
+ ```
220
+
81
221
  ### Claude Pro/Max
82
222
 
83
223
  Anthropic subscription auth is active for Claude Pro/Max accounts. Third-party harness usage draws from [extra usage](https://claude.ai/settings/usage) and is billed per token, not against Claude plan limits.
@@ -111,6 +251,8 @@ omk
111
251
  | Azure OpenAI Responses | `AZURE_OPENAI_API_KEY` | `azure-openai-responses` |
112
252
  | OpenAI | `OPENAI_API_KEY` | `openai` |
113
253
  | DeepSeek | `DEEPSEEK_API_KEY` | `deepseek` |
254
+ | Devin CLI session token | `DEVIN_API_KEY` | `devin` |
255
+ | Cursor subscription | `CURSOR_API_KEY` | `cursor` |
114
256
  | NVIDIA NIM | `NVIDIA_API_KEY` | `nvidia` |
115
257
  | Google Gemini | `GEMINI_API_KEY` | `google` |
116
258
  | Mistral | `MISTRAL_API_KEY` | `mistral` |
@@ -76,6 +76,23 @@ An isolated HOME also hid the installed Rust toolchain: explicitly supplying its
76
76
  location restored the real cargo diagnostic check without changing the test or
77
77
  copying credentials. Neither fixture failure was treated as a product pass.
78
78
 
79
+ ## CI environment recovery
80
+
81
+ The first v0.98.5 tag run built the binaries, but its test step could not find
82
+ `/usr/bin/bwrap`. npm publication was not attempted and GitHub Release creation
83
+ was skipped. The default-branch CI workflows now install `bubblewrap` and run an
84
+ unprivileged namespace probe before the suite. Ubuntu 24.04 subsequently refused
85
+ loopback setup inside the namespace. The runtime-test jobs are pinned to Ubuntu
86
+ 22.04 LTS, with the same namespace and capability-drop checks. They do not skip
87
+ verified-run tests, disable host security controls or enable a sandbox fallback.
88
+ The older distribution `fd` lacked `--no-require-git`; CI installs the official
89
+ fd 10.4.2 static archive with SHA-256 verification before extraction and probes
90
+ that option before tests. File-search behavior and tests are not weakened.
91
+
92
+ Recovery dispatches the official workflow from `main` with both `tag` and
93
+ `source_ref` fixed to `v0.98.5`. The release tag and its source commit stay unchanged;
94
+ the existing source/tag equality checks remain mandatory.
95
+
79
96
  A source-file fingerprint is not an executed-build attestation. Linux observations
80
97
  do not establish behavior on every target platform. CI must validate the exact tag,
81
98
  build the six platform archives, run checks/tests, publish all seven npm packages
@@ -0,0 +1,68 @@
1
+ # Release audit: v0.99.0
2
+
3
+ 기준일: 2026-09-13. 사용자가 배포 준비 결과를 확인한 뒤 즉시 배포를 요청했다.
4
+ 이 기록은 후보의 검증 근거이며 GitHub/npm 게시 완료를 미리 선언하지 않는다.
5
+
6
+ ## 범위와 변경 로그
7
+
8
+ 이전 공개 릴리스 `v0.98.5`는 main의 조상이며, 준비 시점의 GitHub Release와 공개 npm
9
+ `latest` 7개도 0.98.5로 일치했다. 원격 `v0.99.0`이 없음을 확인한 뒤 준비했다.
10
+
11
+ `decee7f157`부터 `ca75f4e5cc`까지의 frontier·Devin·Codex SSE·재시도·TUI·설계 문서와,
12
+ 이번에 확정한 `3c7d3b4613`(TB 선택 v2), `b21daf1a76`(TB 감사 v2),
13
+ `003f7b081a`(문서·변경 로그·로컬 캡처 제외)을 포함한다. 마지막 세 단위는 각각
14
+ pre-commit 전체 검사를 통과했고, 선택기 40개·감사기 66개 CLI 검사를 통과했다.
15
+
16
+ TB 출력 계약 변경은 minor 증가로 처리했다. `selectionVersion: 2`의 nullable 예상
17
+ 시간·총량과 `omk-tb21-audit-report-2`의 필수 완료시각을 Breaking Changes에 기록했다.
18
+ 입력 manifest는 v1을 유지한다. 이전 버전 changelog 본문은 바꾸지 않았으며 다음
19
+ 작업용 `[Unreleased]`는 비워 두었다. 새로운 벤치마크 성능이나 SOTA 우위는 주장하지 않는다.
20
+
21
+ ## 버전·의존성
22
+
23
+ 공개 7개 패키지, root/example manifests, 내부 의존 범위, lockfiles, CLI shrinkwrap,
24
+ book compiler의 `PACKAGE_VERSION`, README 버전 링크와 릴리스 노트를 0.99.0에 맞춘다.
25
+ 외부 의존성의 버전·resolved URL·integrity 값은 변경하지 않았다. 모델 카탈로그도
26
+ 재생성하지 않았다. 운영자 설정·인증·MCP·스킬 활성 목록은 변경하지 않았다.
27
+
28
+ 첫 `version:minor` 실행은 manifest 증가 뒤 아직 이전 버전을 가리키는 내부 의존성을
29
+ npm 출시일 제한에 대조하다 중단됐다. 버전 증가를 재실행하지 않고, 기존
30
+ `sync-versions.js`와 lockfile/설치 동기화 단계만 이어갔다. 제한을 완화하거나 기존
31
+ 태그를 이동하지 않았다. 추가적인 registry 패키지 버전 변경은 없음을 대조했다.
32
+
33
+ ## 관측한 로컬 검증
34
+
35
+ - 환경: Linux, Node.js 24.19.0, npm 11.14.1.
36
+ - 새 버전의 `npm run build`가 7개 workspace 전체에서 종료 0이었다.
37
+ - `test.sh`를 인증 없는 별도 HOME·최소 환경에서 실행했다. 실제 운영자 auth 파일은
38
+ 건드리지 않았다. `taskset`으로 네 CPU에 제한했고 `LIVE_E2E=0`,
39
+ `OMK_NO_LOCAL_LLM=1`, `OMK_OFFLINE=1`을 사용했다.
40
+ - 전체 테스트: **8,279 통과, 852 환경·실계정 조건 skip, 실패 0, 종료 0**.
41
+ WPL 149, agent 870, AI 676, book compiler 22, coding-agent 5,695,
42
+ protocol 137, TUI 730개가 통과했다. 이 수치는 CI 실행 결과가 아니다.
43
+ - 준비 단계의 TB 106개와 전체 guard 378개 통과를 전체 제품 테스트 수에 다시 더하지 않는다.
44
+ - 최종 후보 확인 중 공유 트리에 Devin 요청 코드·테스트·문서와 그 변경 로그가 별도로
45
+ 바뀐 것을 발견했다. 이 4개 파일의 후속 변경은 제외하고 `003f7b081a`와 릴리스
46
+ 메타데이터만 별도 detached worktree에 옮겼다. 변경 로그의 같은 파일에 섞인 후속
47
+ 항목도 원래 커밋의 본문과 대조해 분리했다. 운영자의 변경·인증은 덮어쓰지 않는다.
48
+ - 격리 후보에서도 전체 테스트 8,279개 통과·852개 조건 skip·실패 0을 재확인했다.
49
+ 공유 트리 실행과 같은 검사를 중복 합산하지 않았다.
50
+ - 격리 후보의 `npm run check`(Node guard 378개 포함), 7개 pack dry-run·공개 entrypoint
51
+ import, 빌드 CLI의 `--version`·`--help`·`run --help`가 모두 종료 0이었다.
52
+ 모든 pack의 버전은 0.99.0이고, 필수 진입점 누락·비공개 상태 경로는 없었다.
53
+ - 후보 diff의 Gitleaks 검사는 완전 redaction과 기존 규칙으로 종료 0, 탐지 0건이었다.
54
+ - `--release`의 stale-worktree guard는 별도 최종 배포 gate다. 후보를 main에 연결하고
55
+ 이 작업의 임시 체크아웃을 제거한 뒤 기본 checkout에서 실행한다. guard나 이름을
56
+ 바꿔 검사를 피하지 않는다.
57
+
58
+ ## 배포 완료 조건
59
+
60
+ 기존 `build-binaries.yml`의 태그 기반 경로만 사용한다. 로컬 `npm publish`, 인증 변경,
61
+ 새로운 실행 권한·모델 호출은 없다. release source와 tag의 동일 SHA 검사를 유지한다.
62
+ 후보는 검토한 명시 경로만 stage하고 staged diff 전체를 확인한 뒤 commit·tag·push한다.
63
+
64
+ 공식 workflow가 6개 플랫폼 바이너리 빌드, 검사·테스트, npm 7개 패키지 게시,
65
+ GitHub Release 생성을 완료해야 한다. 최종 판정은 태그의 main 포함,
66
+ GitHub `v0.99.0` Release와 7개 npm `latest`의 일치다. 기존 token 기반 인증을 사용하며
67
+ OIDC/Sigstore provenance는 주장하지 않는다. 실패 시 원인을 확인하고 기존 배포의
68
+ 무결성 경계를 유지하며, 완료 전에는 배포 성공이라고 보고하지 않는다.
@@ -41,7 +41,8 @@ Recovery commands share command IDs and the generation cap; none resets budgets.
41
41
 
42
42
  `linux-command-dag-v1` adds `RunDagWriter` / `RunDagTask` with 1–16 nodes, explicit
43
43
  artifact dependencies, disjoint write scopes and 1–2 preapproved command attempts per
44
- node. `orderRunDag()` provides deterministic FIFO topological order and
44
+ node. Optional `maxConcurrentTasks` accepts only 1 or 2; omission preserves legacy
45
+ serialization and serial execution. `orderRunDag()` provides deterministic FIFO topological order and
45
46
  `runDagAncestors()` includes the full transitive input closure. `RunTaskRetryCommand`
46
47
  / `parseRunTaskRetryCommand()` pins a task selection to the original input and exact
47
48
  run revision/generation. These pure contracts never authorize dispatch or authenticate
@@ -53,8 +54,8 @@ then supplies authenticated checks to the existing claim-closure reducer.
53
54
  This does not replace the task/attempt/evaluation contracts or add execution to
54
55
  this package. See [Verified Run](verified-run.md) for the implemented CLI/SDK
55
56
  path, candidate recovery and checkpoint-based local writer restart, plus the
56
- serial command DAG and selective retry, plus remaining live-model, parallel frontier,
57
- plan amendment and control-surface work.
57
+ bounded command DAG, eager frontier and selective retry, plus remaining live-model,
58
+ verification-edge, plan amendment and control-surface work.
58
59
 
59
60
  ## Durable goal lifecycle
60
61
 
@@ -0,0 +1,56 @@
1
+ # Run usage ledger
2
+
3
+ Status: implemented as an in-memory accounting module and explicit operation adapter.
4
+ Not automatically connected to AgentSession, provider HTTP retries, MCP, or child CLI.
5
+ Existing `RunBudget` request/concurrency/deadline behavior is unchanged.
6
+
7
+ Source: `src/core/run-usage-ledger.ts` and `src/core/run-usage-operation.ts`.
8
+
9
+ ## Contract
10
+
11
+ - `reserve(attemptId, requestId, reservation)` admits an operation only when reported
12
+ consumption + held reservation + proposed reservation fits each supplied cap.
13
+ Retry attempts may share a logical request ID. IDs cannot be reused for attempts.
14
+ - `recordTransport(transportId, attemptId)` records a transport start observed by the
15
+ caller. An application dispatch does not automatically count as an HTTP attempt.
16
+ - `recordUsage(eventId, attemptId, usage)` accepts one cumulative report per attempt.
17
+ Identical redelivery is idempotent; changed payloads and second reports are rejected.
18
+ Reports arriving after settlement or closure still bind to the original attempt.
19
+ - `settle(attemptId)` requires observed operation settlement, not a cancellation request.
20
+ Unknown usage retains its reservation after settlement and blocks new capped admission.
21
+ - `close()` seals admission and transport starts but retains ownership and reservations.
22
+ - Snapshots distinguish the accounted subtotal from `total: null` when usage is missing.
23
+ `estimatedUsd` is an estimate in USD, never invoice-confirmed spending. Token figures
24
+ are caller-normalized reports; the ledger cannot detect provider zero-filled missing usage.
25
+ - Counts and amounts reject negative, non-finite and unsafe arithmetic. Maps are bounded
26
+ by `maxEntries` (default 10,000); there is no automatic eviction of deduplication state.
27
+
28
+ `runUsageOperation` reserves before invoking the callback and settles in `finally`.
29
+ The adapter captures the admitted attempt ID before invoking the callback, so later
30
+ input mutation cannot redirect settlement. Usage reporters must still supply the correct
31
+ attempt ID themselves. The callback promise must represent the operation lifetime. Do not pass a timeout race
32
+ that returns before the underlying work settles. No AbortSignal is treated as proof
33
+ of termination. Failed operations still count as requests/attempts.
34
+
35
+ ## Verification and limits
36
+
37
+ From `packages/coding-agent`:
38
+
39
+ ```sh
40
+ node ../../node_modules/vitest/dist/cli.js run test/run-usage-ledger.test.ts test/run-usage-ownership.test.ts test/run-budget.test.ts test/run-budget-scope.test.ts
41
+ ```
42
+
43
+ The local fixture records main work with two transport retries, continuation, summary,
44
+ and child-labelled work: four logical requests, six transports, eight input tokens.
45
+ These are explicit local adapter events, not real provider, process, or billing evidence.
46
+ Tests cover 6 consumed + 3 reserved + 2 proposed against cap 10, delayed settlement,
47
+ late usage, duplicate/conflicting events, missing usage, bounded input/state, and the
48
+ ownership boundary: denied admission never runs the operation, cancellation retains
49
+ the reservation until observed settlement, and settlement attribution survives caller
50
+ input mutation.
51
+
52
+ No durable replay/restart support, price-table revision binding, invoice corrections,
53
+ work/verification/cleanup partitioning, or automatic whole-run transport instrumentation
54
+ is provided. Actual usage can exceed its reservation: it is recorded honestly and blocks
55
+ later capped admissions rather than being clipped. This is not a hard financial limit.
56
+ R08 end-to-end product integration remains incomplete.
@@ -1,5 +1,49 @@
1
1
  # Runtime Algorithms and Direction
2
2
 
3
+ ## Current feature status (audit §19.3)
4
+
5
+ "Implemented", "wired", "enabled", "verified in CI", "released", and
6
+ "measured benefit" are different gates. A pure function passing unit tests is
7
+ not a live product path, and a live path is not a measured improvement. This
8
+ table records each mechanism's actual gate at the pinned commit; the dated
9
+ baseline below is history, not current truth.
10
+
11
+ | Mechanism | Implemented | Wired into live path | Enabled by default | Verified in CI | Released | Measured benefit |
12
+ | --- | --- | --- | --- | --- | --- | --- |
13
+ | DAG claim scheduler (`tool-dag-scheduler`) | yes | `agent-loop` frontier executor | `toolScheduler: "dag-v2"` | unit + integration tests | v0.99.x line | not measured |
14
+ | Ready-frontier admission (`runDagFrontier`) | yes | `agent-loop` | same | `tool-dag-ready-frontier`, `tool-dag-hook-replan` tests | working tree | not measured |
15
+ | Dynamic-claim conflict check (`conflictsWithUnsettledClaim`) | yes | `agent-loop` | same | `tool-dag-dependencies` + hook tests | working tree | not measured |
16
+ | ECRAF admission planner (`tool-dag-ecraf`) | yes | **no** — no call path wires it | n/a | unit tests only | unreleased | not measured |
17
+ | Context Budget V2 | yes | system-prompt assembly | opt-in policy | planner/selection/cache tests | working tree | not measured |
18
+ | Tier floor reservation | yes | context-budget-v2 planner | when V2 enabled | `context-budget-v2-tier-floor` tests | working tree | not measured |
19
+ | Reasoning router v4 | yes | `/think auto` lane | opt-in | router tests | released | classification only, not success-probability calibration |
20
+ | Workload permit pool | yes | resource admission | default | pool tests | released | not measured |
21
+ | Provider retry/failover classification | yes | `provider-retry`, `session-failure-cause` | default | resilience/classification tests | released | not measured |
22
+ | MCP descriptor injection screen | yes | `mcp/manager` import path | default | quarantine tests | released | pattern rule score, not calibrated risk |
23
+ | verified-run coordinator + evidence | yes | verified-run paths | opt-in command | coordinator/evidence tests | working tree | scope-limited binding, not general correctness |
24
+
25
+ Gates are reported per row so a green "implemented" never upgrades itself to
26
+ "released" or "measured".
27
+
28
+ ## ECRAF arithmetic boundary (2026-09-17)
29
+
30
+ The internal `planEcrafAdmissions()` planner rejects non-finite derived density
31
+ denominators, scores, and reserved usage with `RangeError`, even when each input
32
+ number is finite. Scores are computed once per candidate before sorting or
33
+ calling the conflict predicate; singleton and zero-slot batches receive the same
34
+ validation. A bounded resource overflow still defers the candidate and allows
35
+ later feasible candidates. An unbounded resource has no capacity limit, but its
36
+ usage must remain representable as a finite number. Failed passes return no
37
+ partial admission plan and do not mutate caller input.
38
+
39
+ Regression coverage is in
40
+ `packages/agent/test/tool-dag-ecraf-arithmetic.test.ts`; existing numeric and
41
+ property tests remain in `packages/agent/test/tool-dag-ecraf.test.ts`.
42
+ This is local pure-planner validation, not live frontier wiring, CI confirmation,
43
+ a release, or evidence of latency/quality improvement. Resource normalization,
44
+ slot-aware ranking, fairness, and equal-budget runtime comparisons remain separate
45
+ work.
46
+
3
47
  ## Working-tree shared run budgets
4
48
 
5
49
  The SDK `prompt(..., { runBudget })` path now shares a monotonic deadline and
@@ -154,15 +198,18 @@ Evidence:
154
198
  - `packages/coding-agent/test/context-budget-selection-policy-version.test.ts`
155
199
  - `packages/coding-agent/test/context-budget-cache-disk.test.ts`
156
200
 
157
- **Working tree:** non-queued native `xai` requests started through `AgentSession.prompt()` now derive a bounded automatic skill grant from live discovered descriptions after ordinary prompt-template expansion. The selector scores task text separately from camelCase-aware path-to-skill-name signals, excludes explicit-only skills, caps automatic matches at three, and adds `headroom` only under lexical or measured context pressure. `AgentSession.prompt()` merges the result with settings/SDK/bang selections only for that request. Queued steering/follow-up messages reuse the active run's system prompt and do not trigger another selection pass.
201
+ **Working tree:** non-queued native `xai` and `devin` requests started through `AgentSession.prompt()` now derive a bounded automatic skill grant from live discovered descriptions after ordinary prompt-template expansion. The selector scores task text separately from camelCase-aware path-to-skill-name signals, excludes explicit-only skills, caps automatic matches at three, and adds `headroom` only under lexical or measured context pressure. `AgentSession.prompt()` merges the result with settings/SDK/bang selections only for that request. Queued steering/follow-up messages reuse the active run's system prompt and do not trigger another selection pass.
158
202
 
159
203
  Evidence:
160
204
 
161
205
  - `packages/coding-agent/src/core/active-skill-state.ts`
162
206
  - `packages/coding-agent/src/core/skill-selector.ts`
207
+ - `packages/coding-agent/src/core/harness-skills.ts`
163
208
  - `packages/coding-agent/src/core/grok-harness.ts`
209
+ - `packages/coding-agent/src/core/devin-harness.ts`
164
210
  - `packages/coding-agent/src/core/agent-session.ts`
165
211
  - `packages/coding-agent/test/grok-active-skills.test.ts`
212
+ - `packages/coding-agent/test/devin-active-skills.test.ts`
166
213
  - `packages/coding-agent/test/skill-selector.property.test.ts`
167
214
 
168
215
  **Working tree:** context files now treat their global/local relevance baseline
package/docs/sdk.md CHANGED
@@ -161,7 +161,8 @@ for the JSON shape, events, hook restrictions, and uncovered paths.
161
161
  `planVerifiedRun()` and `createRunCoordinator()` provide three experimental profiles:
162
162
  `linux-command-v1` executes an approved command, `linux-scripted-agent-v1` drives
163
163
  approved steps through the real `AgentSession` and the offline Faux adapter, and
164
- `linux-command-dag-v1` executes a bounded command DAG serially.
164
+ `linux-command-dag-v1` executes a bounded command DAG, serially by default or with
165
+ an explicitly approved `writer.maxConcurrentTasks: 2`.
165
166
  Commands run in private sandboxes; native EvidenceReceipt v3 cores and a supervisor
166
167
  attestation bind the checked candidate before artifact retrieval.
167
168
 
@@ -183,13 +184,16 @@ interrupted tasks. `RunTaskRetryCommand` pins `baseDigest` and `taskIds` alongsi
183
184
  contract/revision/generation fields; an empty selection only continues pending work.
184
185
  Successful checkpoints are revalidated against their complete ancestor inputs before
185
186
  adoption. Final integration verification always uses fresh native receipts. The
186
- profile supports at most 16 tasks and two preapproved commands per task; it does not
187
+ profile supports at most 16 tasks, two preapproved commands per task and an optional
188
+ concurrency cap of 1 or 2. Omitted concurrency preserves legacy contract bytes.
189
+ The eager frontier starts a ready consumer without waiting for unrelated work;
190
+ fatal failure or cancellation drains all started tasks before returning. It does not
187
191
  synthesize a repair or accept overlapping task write scopes. `RunProjection.tasks`
188
192
  exposes task attempts and checkpoint digests; a blocked DAG returns `execution: "paused"`.
189
193
  All recovery actions share the generation cap, original budget and command-id fence.
190
194
 
191
- This does not enable live-model task generation, opaque remote replay, parallel
192
- frontier scheduling, plan amendment, or host application. See [Verified Run](verified-run.md) for contracts and trust boundaries.
195
+ This does not enable live-model task generation, opaque remote replay,
196
+ verification-conditioned edges, plan amendment, or host application. See [Verified Run](verified-run.md) for contracts and trust boundaries.
193
197
 
194
198
  ### Shared run budgets (SDK, opt-in)
195
199
 
@@ -463,7 +467,7 @@ interface PromptOptions {
463
467
  }
464
468
  ```
465
469
 
466
- `activeSkillNames` marks additional discovered skills active for this turn; `activeSkillSource` labels their provenance. They merge with global `defaultActiveSkills`, prioritize matching inventory entries, and do not expand authorization or inline full skill instructions. When the active provider is native `xai` and `OMK_GROK_HARNESS` is enabled, each non-queued `AgentSession.prompt()` request also derives up to three request-scoped matches from the live skill inventory after ordinary prompt-template expansion. Explicit-only skills are excluded from automatic selection, while explicit SDK/settings/bang selections remain authoritative additions. Queued steering and follow-up messages reuse the active run's system prompt and therefore do not perform another automatic skill-selection pass.
470
+ `activeSkillNames` marks additional discovered skills active for this turn; `activeSkillSource` labels their provenance. They merge with global `defaultActiveSkills`, prioritize matching inventory entries, and do not expand authorization or inline full skill instructions. When the active provider is native `xai` and `OMK_GROK_HARNESS` is enabled, or `devin` and `OMK_DEVIN_HARNESS` is enabled, each non-queued `AgentSession.prompt()` request also derives up to three request-scoped matches from the live skill inventory after ordinary prompt-template expansion. Explicit-only skills are excluded from automatic selection, while explicit SDK/settings/bang selections remain authoritative additions. Queued steering and follow-up messages reuse the active run's system prompt and therefore do not perform another automatic skill-selection pass.
467
471
 
468
472
  `preflightResult` is called once per `prompt()` invocation:
469
473
 
package/docs/settings.md CHANGED
@@ -272,7 +272,7 @@ Plays a short system sound only after the top-level prompt settles: retries, con
272
272
  | `retry.enabled` | boolean | `true` | Enable automatic agent-level retry on transient errors |
273
273
  | `retry.maxRetries` | number | `3` | Maximum agent-level retry attempts |
274
274
  | `retry.baseDelayMs` | number | `2000` | Base delay for agent-level exponential backoff (2s, 4s, 8s) |
275
- | `retry.provider.timeoutMs` | number | SDK default | Provider/SDK request timeout in milliseconds |
275
+ | `retry.provider.timeoutMs` | number | SDK default | Provider/SDK request timeout in milliseconds. Also bounds how long the `openai-codex` SSE transport waits for response headers (floor 10s); large contexts need more than the floor |
276
276
  | `retry.provider.maxRetries` | number | `0` | Provider/SDK retry attempts |
277
277
  | `retry.provider.maxRetryDelayMs` | number | `60000` | Max server-requested delay before failing (60s) |
278
278