open-multi-agent-kit 0.98.5 → 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (385) hide show
  1. package/CHANGELOG.md +48 -0
  2. package/README.md +4 -3
  3. package/dist/bun/cli.d.ts.map +1 -1
  4. package/dist/bun/cli.js +1 -0
  5. package/dist/bun/cli.js.map +1 -1
  6. package/dist/bun/register-bundled-coding-agent.d.ts +10 -0
  7. package/dist/bun/register-bundled-coding-agent.d.ts.map +1 -0
  8. package/dist/bun/register-bundled-coding-agent.js +12 -0
  9. package/dist/bun/register-bundled-coding-agent.js.map +1 -0
  10. package/dist/cli/args.d.ts +1 -1
  11. package/dist/cli/args.d.ts.map +1 -1
  12. package/dist/cli/args.js +1 -1
  13. package/dist/cli/args.js.map +1 -1
  14. package/dist/cli/help.d.ts.map +1 -1
  15. package/dist/cli/help.js +2 -1
  16. package/dist/cli/help.js.map +1 -1
  17. package/dist/cli.d.ts.map +1 -1
  18. package/dist/cli.js +14 -2
  19. package/dist/cli.js.map +1 -1
  20. package/dist/commands/neo-cli.d.ts +9 -0
  21. package/dist/commands/neo-cli.d.ts.map +1 -0
  22. package/dist/commands/neo-cli.js +61 -0
  23. package/dist/commands/neo-cli.js.map +1 -0
  24. package/dist/commands/verified-run-cli.d.ts.map +1 -1
  25. package/dist/commands/verified-run-cli.js +2 -2
  26. package/dist/commands/verified-run-cli.js.map +1 -1
  27. package/dist/core/agent-session.d.ts +29 -2
  28. package/dist/core/agent-session.d.ts.map +1 -1
  29. package/dist/core/agent-session.js +126 -32
  30. package/dist/core/agent-session.js.map +1 -1
  31. package/dist/core/bundled-skills.d.ts +4 -0
  32. package/dist/core/bundled-skills.d.ts.map +1 -0
  33. package/dist/core/bundled-skills.js +32 -0
  34. package/dist/core/bundled-skills.js.map +1 -0
  35. package/dist/core/cli-diagnostics.d.ts +6 -0
  36. package/dist/core/cli-diagnostics.d.ts.map +1 -0
  37. package/dist/core/cli-diagnostics.js +20 -0
  38. package/dist/core/cli-diagnostics.js.map +1 -0
  39. package/dist/core/context-budget-headroom.d.ts +12 -0
  40. package/dist/core/context-budget-headroom.d.ts.map +1 -1
  41. package/dist/core/context-budget-headroom.js +35 -0
  42. package/dist/core/context-budget-headroom.js.map +1 -1
  43. package/dist/core/context-budget-v2-input-validation.d.ts +4 -0
  44. package/dist/core/context-budget-v2-input-validation.d.ts.map +1 -0
  45. package/dist/core/context-budget-v2-input-validation.js +79 -0
  46. package/dist/core/context-budget-v2-input-validation.js.map +1 -0
  47. package/dist/core/context-budget-v2-observability.d.ts +15 -0
  48. package/dist/core/context-budget-v2-observability.d.ts.map +1 -0
  49. package/dist/core/context-budget-v2-observability.js +35 -0
  50. package/dist/core/context-budget-v2-observability.js.map +1 -0
  51. package/dist/core/context-budget-v2-planned-items.d.ts +5 -0
  52. package/dist/core/context-budget-v2-planned-items.d.ts.map +1 -0
  53. package/dist/core/context-budget-v2-planned-items.js +40 -0
  54. package/dist/core/context-budget-v2-planned-items.js.map +1 -0
  55. package/dist/core/context-budget-v2-planner.d.ts.map +1 -1
  56. package/dist/core/context-budget-v2-planner.js +57 -38
  57. package/dist/core/context-budget-v2-planner.js.map +1 -1
  58. package/dist/core/context-budget-v2-selection.d.ts +19 -3
  59. package/dist/core/context-budget-v2-selection.d.ts.map +1 -1
  60. package/dist/core/context-budget-v2-selection.js +35 -57
  61. package/dist/core/context-budget-v2-selection.js.map +1 -1
  62. package/dist/core/context-budget-v2-tiers.d.ts.map +1 -1
  63. package/dist/core/context-budget-v2-tiers.js +9 -42
  64. package/dist/core/context-budget-v2-tiers.js.map +1 -1
  65. package/dist/core/context-budget-v2-types.d.ts +8 -1
  66. package/dist/core/context-budget-v2-types.d.ts.map +1 -1
  67. package/dist/core/context-budget-v2-types.js.map +1 -1
  68. package/dist/core/devin-harness-dispatch.d.ts +12 -0
  69. package/dist/core/devin-harness-dispatch.d.ts.map +1 -0
  70. package/dist/core/devin-harness-dispatch.js +12 -0
  71. package/dist/core/devin-harness-dispatch.js.map +1 -0
  72. package/dist/core/devin-harness.d.ts +53 -0
  73. package/dist/core/devin-harness.d.ts.map +1 -0
  74. package/dist/core/devin-harness.js +112 -0
  75. package/dist/core/devin-harness.js.map +1 -0
  76. package/dist/core/domain-dispatch.d.ts +4 -1
  77. package/dist/core/domain-dispatch.d.ts.map +1 -1
  78. package/dist/core/domain-dispatch.js +5 -0
  79. package/dist/core/domain-dispatch.js.map +1 -1
  80. package/dist/core/domain-loadouts-provider-harness.d.ts +12 -0
  81. package/dist/core/domain-loadouts-provider-harness.d.ts.map +1 -0
  82. package/dist/core/domain-loadouts-provider-harness.js +122 -0
  83. package/dist/core/domain-loadouts-provider-harness.js.map +1 -0
  84. package/dist/core/domain-loadouts.d.ts +2 -33
  85. package/dist/core/domain-loadouts.d.ts.map +1 -1
  86. package/dist/core/domain-loadouts.js +3 -56
  87. package/dist/core/domain-loadouts.js.map +1 -1
  88. package/dist/core/domain-profile.d.ts +40 -0
  89. package/dist/core/domain-profile.d.ts.map +1 -0
  90. package/dist/core/domain-profile.js +7 -0
  91. package/dist/core/domain-profile.js.map +1 -0
  92. package/dist/core/extensions/bundled-virtual-modules.d.ts +15 -0
  93. package/dist/core/extensions/bundled-virtual-modules.d.ts.map +1 -0
  94. package/dist/core/extensions/bundled-virtual-modules.js +63 -0
  95. package/dist/core/extensions/bundled-virtual-modules.js.map +1 -0
  96. package/dist/core/extensions/loader.d.ts.map +1 -1
  97. package/dist/core/extensions/loader.js +2 -42
  98. package/dist/core/extensions/loader.js.map +1 -1
  99. package/dist/core/extensions/runner.d.ts +1 -0
  100. package/dist/core/extensions/runner.d.ts.map +1 -1
  101. package/dist/core/extensions/runner.js +6 -0
  102. package/dist/core/extensions/runner.js.map +1 -1
  103. package/dist/core/extensions/types.d.ts +10 -0
  104. package/dist/core/extensions/types.d.ts.map +1 -1
  105. package/dist/core/extensions/types.js.map +1 -1
  106. package/dist/core/grok-harness-dispatch.d.ts +6 -20
  107. package/dist/core/grok-harness-dispatch.d.ts.map +1 -1
  108. package/dist/core/grok-harness-dispatch.js +6 -55
  109. package/dist/core/grok-harness-dispatch.js.map +1 -1
  110. package/dist/core/grok-harness.d.ts +7 -9
  111. package/dist/core/grok-harness.d.ts.map +1 -1
  112. package/dist/core/grok-harness.js +9 -23
  113. package/dist/core/grok-harness.js.map +1 -1
  114. package/dist/core/harness-skills.d.ts +20 -0
  115. package/dist/core/harness-skills.d.ts.map +1 -0
  116. package/dist/core/harness-skills.js +36 -0
  117. package/dist/core/harness-skills.js.map +1 -0
  118. package/dist/core/index.d.ts +2 -1
  119. package/dist/core/index.d.ts.map +1 -1
  120. package/dist/core/index.js +2 -1
  121. package/dist/core/index.js.map +1 -1
  122. package/dist/core/loadout-runtime-state.d.ts +25 -0
  123. package/dist/core/loadout-runtime-state.d.ts.map +1 -0
  124. package/dist/core/loadout-runtime-state.js +39 -0
  125. package/dist/core/loadout-runtime-state.js.map +1 -0
  126. package/dist/core/loadout-runtime.d.ts +4 -20
  127. package/dist/core/loadout-runtime.d.ts.map +1 -1
  128. package/dist/core/loadout-runtime.js +6 -13
  129. package/dist/core/loadout-runtime.js.map +1 -1
  130. package/dist/core/mcp/client.d.ts +27 -6
  131. package/dist/core/mcp/client.d.ts.map +1 -1
  132. package/dist/core/mcp/client.js +80 -20
  133. package/dist/core/mcp/client.js.map +1 -1
  134. package/dist/core/mcp/manager.d.ts +10 -1
  135. package/dist/core/mcp/manager.d.ts.map +1 -1
  136. package/dist/core/mcp/manager.js +68 -9
  137. package/dist/core/mcp/manager.js.map +1 -1
  138. package/dist/core/mcp/protocol.d.ts +9 -3
  139. package/dist/core/mcp/protocol.d.ts.map +1 -1
  140. package/dist/core/mcp/protocol.js +61 -15
  141. package/dist/core/mcp/protocol.js.map +1 -1
  142. package/dist/core/mcp/stdio-transport.d.ts +12 -1
  143. package/dist/core/mcp/stdio-transport.d.ts.map +1 -1
  144. package/dist/core/mcp/stdio-transport.js +30 -2
  145. package/dist/core/mcp/stdio-transport.js.map +1 -1
  146. package/dist/core/mcp-descriptor-injection.d.ts +26 -0
  147. package/dist/core/mcp-descriptor-injection.d.ts.map +1 -0
  148. package/dist/core/mcp-descriptor-injection.js +25 -0
  149. package/dist/core/mcp-descriptor-injection.js.map +1 -0
  150. package/dist/core/mcp-public-presets.d.ts +1 -3
  151. package/dist/core/mcp-public-presets.d.ts.map +1 -1
  152. package/dist/core/mcp-public-presets.js +3 -15
  153. package/dist/core/mcp-public-presets.js.map +1 -1
  154. package/dist/core/model-registry.d.ts.map +1 -1
  155. package/dist/core/model-registry.js +24 -3
  156. package/dist/core/model-registry.js.map +1 -1
  157. package/dist/core/model-resolver.d.ts +1 -41
  158. package/dist/core/model-resolver.d.ts.map +1 -1
  159. package/dist/core/model-resolver.js +2 -49
  160. package/dist/core/model-resolver.js.map +1 -1
  161. package/dist/core/neo/catalog.d.ts +20 -0
  162. package/dist/core/neo/catalog.d.ts.map +1 -0
  163. package/dist/core/neo/catalog.js +50 -0
  164. package/dist/core/neo/catalog.js.map +1 -0
  165. package/dist/core/neo/setup.d.ts +3 -0
  166. package/dist/core/neo/setup.d.ts.map +1 -0
  167. package/dist/core/neo/setup.js +47 -0
  168. package/dist/core/neo/setup.js.map +1 -0
  169. package/dist/core/provider-default-models.d.ts +43 -0
  170. package/dist/core/provider-default-models.d.ts.map +1 -0
  171. package/dist/core/provider-default-models.js +51 -0
  172. package/dist/core/provider-default-models.js.map +1 -0
  173. package/dist/core/provider-display-names.d.ts.map +1 -1
  174. package/dist/core/provider-display-names.js +2 -0
  175. package/dist/core/provider-display-names.js.map +1 -1
  176. package/dist/core/provider-error-classification.d.ts +57 -0
  177. package/dist/core/provider-error-classification.d.ts.map +1 -0
  178. package/dist/core/provider-error-classification.js +103 -0
  179. package/dist/core/provider-error-classification.js.map +1 -0
  180. package/dist/core/provider-harness-dispatch.d.ts +63 -0
  181. package/dist/core/provider-harness-dispatch.d.ts.map +1 -0
  182. package/dist/core/provider-harness-dispatch.js +60 -0
  183. package/dist/core/provider-harness-dispatch.js.map +1 -0
  184. package/dist/core/provider-resilience.d.ts +10 -0
  185. package/dist/core/provider-resilience.d.ts.map +1 -1
  186. package/dist/core/provider-resilience.js +36 -3
  187. package/dist/core/provider-resilience.js.map +1 -1
  188. package/dist/core/provider-usage-commandcode.d.ts +9 -0
  189. package/dist/core/provider-usage-commandcode.d.ts.map +1 -0
  190. package/dist/core/provider-usage-commandcode.js +198 -0
  191. package/dist/core/provider-usage-commandcode.js.map +1 -0
  192. package/dist/core/provider-usage-devin.d.ts +18 -0
  193. package/dist/core/provider-usage-devin.d.ts.map +1 -0
  194. package/dist/core/provider-usage-devin.js +71 -0
  195. package/dist/core/provider-usage-devin.js.map +1 -0
  196. package/dist/core/provider-usage-text.d.ts +5 -0
  197. package/dist/core/provider-usage-text.d.ts.map +1 -0
  198. package/dist/core/provider-usage-text.js +15 -0
  199. package/dist/core/provider-usage-text.js.map +1 -0
  200. package/dist/core/provider-usage-types.d.ts +2 -1
  201. package/dist/core/provider-usage-types.d.ts.map +1 -1
  202. package/dist/core/provider-usage-types.js.map +1 -1
  203. package/dist/core/provider-usage.d.ts +1 -2
  204. package/dist/core/provider-usage.d.ts.map +1 -1
  205. package/dist/core/provider-usage.js +18 -7
  206. package/dist/core/provider-usage.js.map +1 -1
  207. package/dist/core/resource-admission.d.ts +36 -0
  208. package/dist/core/resource-admission.d.ts.map +1 -1
  209. package/dist/core/resource-admission.js +59 -0
  210. package/dist/core/resource-admission.js.map +1 -1
  211. package/dist/core/resource-loader.d.ts.map +1 -1
  212. package/dist/core/resource-loader.js +2 -2
  213. package/dist/core/resource-loader.js.map +1 -1
  214. package/dist/core/run-execution-api.d.ts +1 -1
  215. package/dist/core/run-execution-api.d.ts.map +1 -1
  216. package/dist/core/run-execution-api.js.map +1 -1
  217. package/dist/core/run-usage-ledger.d.ts +47 -0
  218. package/dist/core/run-usage-ledger.d.ts.map +1 -0
  219. package/dist/core/run-usage-ledger.js +162 -0
  220. package/dist/core/run-usage-ledger.js.map +1 -0
  221. package/dist/core/run-usage-operation.d.ts +8 -0
  222. package/dist/core/run-usage-operation.d.ts.map +1 -0
  223. package/dist/core/run-usage-operation.js +12 -0
  224. package/dist/core/run-usage-operation.js.map +1 -0
  225. package/dist/core/sdk.d.ts.map +1 -1
  226. package/dist/core/sdk.js +20 -16
  227. package/dist/core/sdk.js.map +1 -1
  228. package/dist/core/session-failure-cause.d.ts.map +1 -1
  229. package/dist/core/session-failure-cause.js +14 -5
  230. package/dist/core/session-failure-cause.js.map +1 -1
  231. package/dist/core/session-prompt-lifecycle.d.ts +6 -0
  232. package/dist/core/session-prompt-lifecycle.d.ts.map +1 -1
  233. package/dist/core/session-prompt-lifecycle.js +47 -2
  234. package/dist/core/session-prompt-lifecycle.js.map +1 -1
  235. package/dist/core/session-termination.d.ts.map +1 -1
  236. package/dist/core/session-termination.js +4 -2
  237. package/dist/core/session-termination.js.map +1 -1
  238. package/dist/core/subagent-lane-authority.d.ts +28 -0
  239. package/dist/core/subagent-lane-authority.d.ts.map +1 -0
  240. package/dist/core/subagent-lane-authority.js +119 -0
  241. package/dist/core/subagent-lane-authority.js.map +1 -0
  242. package/dist/core/subagent-lane-contract.d.ts +114 -0
  243. package/dist/core/subagent-lane-contract.d.ts.map +1 -0
  244. package/dist/core/subagent-lane-contract.js +18 -0
  245. package/dist/core/subagent-lane-contract.js.map +1 -0
  246. package/dist/core/subagent-lane-launcher.d.ts +3 -6
  247. package/dist/core/subagent-lane-launcher.d.ts.map +1 -1
  248. package/dist/core/subagent-lane-launcher.js +53 -12
  249. package/dist/core/subagent-lane-launcher.js.map +1 -1
  250. package/dist/core/subagent-orchestration.d.ts +4 -17
  251. package/dist/core/subagent-orchestration.d.ts.map +1 -1
  252. package/dist/core/subagent-orchestration.js.map +1 -1
  253. package/dist/core/verified-run/broker.d.ts.map +1 -1
  254. package/dist/core/verified-run/broker.js +4 -1
  255. package/dist/core/verified-run/broker.js.map +1 -1
  256. package/dist/core/verified-run/dag-phase.d.ts +1 -1
  257. package/dist/core/verified-run/dag-phase.d.ts.map +1 -1
  258. package/dist/core/verified-run/dag-phase.js +46 -8
  259. package/dist/core/verified-run/dag-phase.js.map +1 -1
  260. package/dist/core/verified-run/dag-projection.d.ts.map +1 -1
  261. package/dist/core/verified-run/dag-projection.js +25 -12
  262. package/dist/core/verified-run/dag-projection.js.map +1 -1
  263. package/dist/core/verified-run/dag-types.d.ts +11 -0
  264. package/dist/core/verified-run/dag-types.d.ts.map +1 -1
  265. package/dist/core/verified-run/dag-types.js.map +1 -1
  266. package/dist/core/verified-run/event-parser.d.ts.map +1 -1
  267. package/dist/core/verified-run/event-parser.js +1 -0
  268. package/dist/core/verified-run/event-parser.js.map +1 -1
  269. package/dist/core/verified-run/owned-execution.d.ts +1 -0
  270. package/dist/core/verified-run/owned-execution.d.ts.map +1 -1
  271. package/dist/core/verified-run/owned-execution.js +7 -1
  272. package/dist/core/verified-run/owned-execution.js.map +1 -1
  273. package/dist/core/verified-run/process-projection.d.ts +13 -0
  274. package/dist/core/verified-run/process-projection.d.ts.map +1 -0
  275. package/dist/core/verified-run/process-projection.js +82 -0
  276. package/dist/core/verified-run/process-projection.js.map +1 -0
  277. package/dist/core/verified-run/projection.d.ts.map +1 -1
  278. package/dist/core/verified-run/projection.js +11 -55
  279. package/dist/core/verified-run/projection.js.map +1 -1
  280. package/dist/core/verified-run/run-types.d.ts +1 -0
  281. package/dist/core/verified-run/run-types.d.ts.map +1 -1
  282. package/dist/core/verified-run/run-types.js.map +1 -1
  283. package/dist/core/verified-run/task-execution.d.ts +9 -0
  284. package/dist/core/verified-run/task-execution.d.ts.map +1 -0
  285. package/dist/core/verified-run/task-execution.js +29 -0
  286. package/dist/core/verified-run/task-execution.js.map +1 -0
  287. package/dist/core/verified-run/writer-projection.d.ts.map +1 -1
  288. package/dist/core/verified-run/writer-projection.js +6 -5
  289. package/dist/core/verified-run/writer-projection.js.map +1 -1
  290. package/dist/core/workload-permit-pool.d.ts +1 -1
  291. package/dist/core/workload-permit-pool.d.ts.map +1 -1
  292. package/dist/core/workload-permit-pool.js +3 -0
  293. package/dist/core/workload-permit-pool.js.map +1 -1
  294. package/dist/guardrails/strict-evidence-approval-adapter.d.ts +34 -0
  295. package/dist/guardrails/strict-evidence-approval-adapter.d.ts.map +1 -0
  296. package/dist/guardrails/strict-evidence-approval-adapter.js +70 -0
  297. package/dist/guardrails/strict-evidence-approval-adapter.js.map +1 -0
  298. package/dist/index.d.ts +3 -0
  299. package/dist/index.d.ts.map +1 -1
  300. package/dist/index.js +1 -0
  301. package/dist/index.js.map +1 -1
  302. package/dist/main.d.ts.map +1 -1
  303. package/dist/main.js +11 -18
  304. package/dist/main.js.map +1 -1
  305. package/dist/modes/acp/acp-agent.d.ts +24 -0
  306. package/dist/modes/acp/acp-agent.d.ts.map +1 -0
  307. package/dist/modes/acp/acp-agent.js +133 -0
  308. package/dist/modes/acp/acp-agent.js.map +1 -0
  309. package/dist/modes/acp/acp-mode.d.ts +4 -0
  310. package/dist/modes/acp/acp-mode.d.ts.map +1 -0
  311. package/dist/modes/acp/acp-mode.js +34 -0
  312. package/dist/modes/acp/acp-mode.js.map +1 -0
  313. package/dist/modes/acp/acp-session.d.ts +4 -0
  314. package/dist/modes/acp/acp-session.d.ts.map +1 -0
  315. package/dist/modes/acp/acp-session.js +76 -0
  316. package/dist/modes/acp/acp-session.js.map +1 -0
  317. package/dist/modes/acp/acp-transport.d.ts +5 -0
  318. package/dist/modes/acp/acp-transport.d.ts.map +1 -0
  319. package/dist/modes/acp/acp-transport.js +90 -0
  320. package/dist/modes/acp/acp-transport.js.map +1 -0
  321. package/dist/modes/interactive/interactive-mode.d.ts.map +1 -1
  322. package/dist/modes/interactive/interactive-mode.js +3 -5
  323. package/dist/modes/interactive/interactive-mode.js.map +1 -1
  324. package/dist/modes/interactive/resource-description.d.ts +3 -0
  325. package/dist/modes/interactive/resource-description.d.ts.map +1 -0
  326. package/dist/modes/interactive/resource-description.js +15 -0
  327. package/dist/modes/interactive/resource-description.js.map +1 -0
  328. package/docs/audit-revalidation-2026-09-17.md +77 -0
  329. package/docs/devin-harness.md +150 -0
  330. package/docs/docs.json +4 -0
  331. package/docs/ecraf-normalization.md +83 -0
  332. package/docs/environment-variables.md +2 -1
  333. package/docs/index.md +1 -0
  334. package/docs/loadout-domains/README.md +2 -1
  335. package/docs/loadout-domains/devin-harness.md +72 -0
  336. package/docs/metrics.md +57 -16
  337. package/docs/model-catalog-refresh.md +51 -1
  338. package/docs/models.md +1 -1
  339. package/docs/neo.md +134 -0
  340. package/docs/providers.md +143 -1
  341. package/docs/release-audit-0.98.5.md +17 -0
  342. package/docs/release-audit-0.99.0.md +68 -0
  343. package/docs/run-protocol.md +4 -3
  344. package/docs/run-usage-ledger.md +56 -0
  345. package/docs/runtime-algorithms.md +48 -1
  346. package/docs/sdk.md +9 -5
  347. package/docs/settings.md +1 -1
  348. package/docs/skills.md +1 -1
  349. package/docs/startup-resource-labels-testing.md +47 -0
  350. package/docs/tb21-audit.md +18 -4
  351. package/docs/usage.md +7 -1
  352. package/docs/verified-run-remaining-design.md +881 -0
  353. package/docs/verified-run-testing.md +91 -0
  354. package/docs/verified-run.md +30 -9
  355. package/examples/extensions/custom-provider-anthropic/package-lock.json +2 -2
  356. package/examples/extensions/custom-provider-anthropic/package.json +1 -1
  357. package/examples/extensions/custom-provider-gitlab-duo/package.json +1 -1
  358. package/examples/extensions/gondolin/package-lock.json +2 -2
  359. package/examples/extensions/gondolin/package.json +1 -1
  360. package/examples/extensions/sandbox/package-lock.json +2 -2
  361. package/examples/extensions/sandbox/package.json +1 -1
  362. package/examples/extensions/subagent/adaptive-agent-runtime.ts +14 -1
  363. package/examples/extensions/subagent/graph-result.ts +48 -0
  364. package/examples/extensions/subagent/index.ts +246 -111
  365. package/examples/extensions/subagent/managed-process-tree.ts +42 -0
  366. package/examples/extensions/subagent/managed-process.test.ts +24 -0
  367. package/examples/extensions/subagent/managed-process.ts +93 -112
  368. package/examples/extensions/subagent/subagent-runtime-types.ts +17 -1
  369. package/examples/extensions/subagent/subagent-stream.ts +161 -0
  370. package/examples/extensions/terminal-browser/README.md +54 -0
  371. package/examples/extensions/terminal-browser/bridge-protocol.ts +29 -0
  372. package/examples/extensions/terminal-browser/bridge.ts +219 -0
  373. package/examples/extensions/terminal-browser/browser-surface.ts +288 -0
  374. package/examples/extensions/terminal-browser/index.ts +215 -0
  375. package/examples/extensions/terminal-browser/placeholders.ts +49 -0
  376. package/examples/extensions/with-deps/package-lock.json +2 -2
  377. package/examples/extensions/with-deps/package.json +1 -1
  378. package/npm-shrinkwrap.json +18 -18
  379. package/package.json +8 -7
  380. package/resources/neo/skills/omk-browser/SKILL.md +32 -0
  381. package/resources/neo/skills/omk-code-review/SKILL.md +24 -0
  382. package/resources/neo/skills/omk-computeruse/SKILL.md +34 -0
  383. package/resources/neo/skills/omk-mcp-setup/SKILL.md +48 -0
  384. package/resources/neo/skills/omk-research/SKILL.md +24 -0
  385. package/resources/neo/skills/omk-site/SKILL.md +28 -0
@@ -0,0 +1,3 @@
1
+ /** Keep decorative imported harness markers out of OMK display; source metadata stays unchanged. */
2
+ export declare function formatResourceDescription(description: string | undefined, sourceTag?: string): string | undefined;
3
+ //# sourceMappingURL=resource-description.d.ts.map
@@ -0,0 +1 @@
1
+ {"version":3,"file":"resource-description.d.ts","sourceRoot":"","sources":["../../../src/modes/interactive/resource-description.ts"],"names":[],"mappings":"AAAA,oGAAoG;AACpG,wBAAgB,yBAAyB,CAAC,WAAW,EAAE,MAAM,GAAG,SAAS,EAAE,SAAS,CAAC,EAAE,MAAM,GAAG,MAAM,GAAG,SAAS,CAWjH","sourcesContent":["/** Keep decorative imported harness markers out of OMK display; source metadata stays unchanged. */\nexport function formatResourceDescription(description: string | undefined, sourceTag?: string): string | undefined {\n\tlet text = description;\n\tif (description) {\n\t\tconst marker = /^\\[(OMX|OMO)\\]\\s*/.exec(description);\n\t\tif (marker) {\n\t\t\tconst body = description.slice(marker[0].length);\n\t\t\ttext = body || \"OMK resource\";\n\t\t}\n\t}\n\tif (!sourceTag) return text;\n\treturn text ? `[${sourceTag}] ${text}` : `[${sourceTag}]`;\n}\n"]}
@@ -0,0 +1,15 @@
1
+ /** Keep decorative imported harness markers out of OMK display; source metadata stays unchanged. */
2
+ export function formatResourceDescription(description, sourceTag) {
3
+ let text = description;
4
+ if (description) {
5
+ const marker = /^\[(OMX|OMO)\]\s*/.exec(description);
6
+ if (marker) {
7
+ const body = description.slice(marker[0].length);
8
+ text = body || "OMK resource";
9
+ }
10
+ }
11
+ if (!sourceTag)
12
+ return text;
13
+ return text ? `[${sourceTag}] ${text}` : `[${sourceTag}]`;
14
+ }
15
+ //# sourceMappingURL=resource-description.js.map
@@ -0,0 +1 @@
1
+ {"version":3,"file":"resource-description.js","sourceRoot":"","sources":["../../../src/modes/interactive/resource-description.ts"],"names":[],"mappings":"AAAA,oGAAoG;AACpG,MAAM,UAAU,yBAAyB,CAAC,WAA+B,EAAE,SAAkB,EAAsB;IAClH,IAAI,IAAI,GAAG,WAAW,CAAC;IACvB,IAAI,WAAW,EAAE,CAAC;QACjB,MAAM,MAAM,GAAG,mBAAmB,CAAC,IAAI,CAAC,WAAW,CAAC,CAAC;QACrD,IAAI,MAAM,EAAE,CAAC;YACZ,MAAM,IAAI,GAAG,WAAW,CAAC,KAAK,CAAC,MAAM,CAAC,CAAC,CAAC,CAAC,MAAM,CAAC,CAAC;YACjD,IAAI,GAAG,IAAI,IAAI,cAAc,CAAC;QAC/B,CAAC;IACF,CAAC;IACD,IAAI,CAAC,SAAS;QAAE,OAAO,IAAI,CAAC;IAC5B,OAAO,IAAI,CAAC,CAAC,CAAC,IAAI,SAAS,KAAK,IAAI,EAAE,CAAC,CAAC,CAAC,IAAI,SAAS,GAAG,CAAC;AAAA,CAC1D","sourcesContent":["/** Keep decorative imported harness markers out of OMK display; source metadata stays unchanged. */\nexport function formatResourceDescription(description: string | undefined, sourceTag?: string): string | undefined {\n\tlet text = description;\n\tif (description) {\n\t\tconst marker = /^\\[(OMX|OMO)\\]\\s*/.exec(description);\n\t\tif (marker) {\n\t\t\tconst body = description.slice(marker[0].length);\n\t\t\ttext = body || \"OMK resource\";\n\t\t}\n\t}\n\tif (!sourceTag) return text;\n\treturn text ? `[${sourceTag}] ${text}` : `[${sourceTag}]`;\n}\n"]}
@@ -0,0 +1,77 @@
1
+ # Audit revalidation — 2026-09-17
2
+
3
+ Baseline: `1e58611d0c28775c8a066b6e913ea6710ccea396`, with unrelated local
4
+ provider/model changes present. This is a scoped local revalidation, not a release
5
+ approval, full security audit, or competitor benchmark.
6
+
7
+ ## Inputs and scope
8
+
9
+ The local September 14 reproduction ZIP targets
10
+ `69e0637d8fddd26eb40a7ec3f8ebfb65b500de5b`. All 25 entries listed in its
11
+ SHA256SUMS matched. Its recorded six expected failures describe that historical
12
+ snapshot, not the current checkout. The archived runner was not executed or
13
+ substituted for current repository tests.
14
+
15
+ The September 15 strategy report combines historical findings with proposed
16
+ resource-aware scheduling, recovery, and comparative evaluation. These proposals
17
+ are not measured product benefits. The September 16 GitHub audit targets
18
+ `739bc6f3b6fe1c89bfe058aad7ec17ad252fb10e`; its missing DAG contract is no longer
19
+ present in this baseline. The documents were used to select the checks below;
20
+ not every proposed acceptance criterion was executed.
21
+
22
+ ## Executed regression groups
23
+
24
+ Commands use `node ../../node_modules/vitest/dist/cli.js --run` from the named
25
+ workspace. No external provider requests were needed.
26
+
27
+ | Workspace | Test files | Result |
28
+ | --- | --- | --- |
29
+ | agent | `test/tool-dag-*.test.ts` | 83 passed, including 10,000 seeded bounded schedules |
30
+ | coding-agent | `test/mcp/{protocol,protocol-required,manager-lifecycle,manager-health,client,tools}.test.ts` | 61 passed; client tests include a local stdio fixture |
31
+ | protocol | `test/{protocol,evaluation-candidate-binding,validation,run-dag-contract,run-dag-properties}.test.ts` | 44 passed |
32
+ | coding-agent | `test/{context-budget-v2-validation-cache,context-budget-cache,context-budget-cache-policy,context-budget-v2-tier-floor,context-budget-v2-eligibility,context-budget-v2-nonfinite-tokens,context-budget-governor-v2}.test.ts` | 46 passed after the fix below |
33
+
34
+ These are 234 distinct passing test cases, not 234 independent production tasks.
35
+ An initial command included nonexistent `test/mcp/manager.test.ts`; Vitest ran
36
+ other matching files successfully. Manager coverage above comes from the actual
37
+ `manager-lifecycle`, `manager-health`, and `client` files, not that missing path.
38
+
39
+ ## Fixed: input diagnostics leaked across plan-cache reuse
40
+
41
+ `planPromptContextBudgetV2()` originally decided cache eligibility before
42
+ `validateBudgetItems()`. Duplicate IDs are dropped and invalid token estimates are
43
+ recomputed. Their sanitized items can produce exactly the same plan key as valid
44
+ input, but their input diagnostics are different.
45
+
46
+ Three regressions failed before the fix:
47
+
48
+ 1. A cached valid plan hid a duplicate-ID diagnostic on a later call.
49
+ 2. A plan produced from duplicate IDs carried its diagnostic into a later valid call.
50
+ 3. A cached plan hid the diagnostic for a recomputed non-finite token estimate.
51
+
52
+ Cache eligibility is now decided after input validation. Calls with input or
53
+ budget diagnostics neither read nor write the plan cache. Representation-cache
54
+ validation and normal valid-input plan reuse remain unchanged. The regression
55
+ uses a required item to isolate plan reuse from the existing
56
+ `cache_dependency_unsafe` rejection for plans containing representation-cache hits.
57
+
58
+ Changed implementation: `src/core/context-budget-v2-planner.ts`.
59
+ Regression: `test/context-budget-v2-validation-cache.test.ts` (3 passed).
60
+ This preserves diagnostics; it does not redesign duplicate-item handling or tier
61
+ allocation policy.
62
+
63
+ ## Remaining boundaries
64
+
65
+ - Full `npm run check` has been blocked by unrelated local module-size growth in
66
+ `agent-session.ts` and `model-registry.ts`; those files and the ratchet baseline
67
+ are not changed by this fix. A passing focused compiler/test run does not close
68
+ the full repository gate.
69
+ - General observation evaluation still has existential semantics. Passing
70
+ candidate-binding tests does not establish latest-result or complete-coverage
71
+ semantics, or independently authenticate a verifier.
72
+ - ECRAF remains an internal planner. Resource normalization, fairness, live
73
+ admission integration, and equal-budget performance comparisons were not added.
74
+ - This pass does not validate the complete timeout/ownership fault matrix,
75
+ verified-run crash recovery, all MCP authorization boundaries, remote CI,
76
+ branch protection, published artifacts, or router calibration.
77
+ - No release, commit, push, paid benchmark, or new runtime default is implied.
@@ -0,0 +1,150 @@
1
+ # Devin SWE-2 harness
2
+
3
+ This page is the canonical operator guide for the built-in `devin` provider, centered on the logical `swe-2` model. Use `/login devin` for the Devin CLI subscription PKCE flow or `DEVIN_API_KEY` for an already-owned CLI session token. A user-local `~/.omk/agent/devin.md` may add operator notes, but it is not the portable product contract.
4
+
5
+ Authentication, transport, and verification limits are owned by [Providers](providers.md#devin-cli); this page covers how OMK drives SWE-2 as a harness.
6
+
7
+ ## Presets
8
+
9
+ Project presets live in `.omk/presets.json` (or `~/.omk/agent/presets.json`) and are consumed by the preset extension from `packages/coding-agent/examples/extensions/preset.ts`. The shared SWE-2 presets intentionally omit the `tools` key so role/domain lane grants keep control of the active tools.
10
+
11
+ | Preset | Provider | Model | Thinking | Use |
12
+ | --- | --- | --- | --- | --- |
13
+ | `swe2-verified` | `devin` | `swe-2` | `high` | Default SWE-2 coding baseline: multi-file edits with tests. |
14
+ | `swe2-max` | `devin` | `swe-2` | `max` | Long-horizon, uncertain, or repository-wide work that should use the 1M budget. |
15
+ | `swe2-fast-edit` | `devin` | `swe-2` | `medium` | Small, well-specified edits where medium acts sooner and costs less. |
16
+
17
+ ```json
18
+ {
19
+ "swe2-verified": { "provider": "devin", "model": "swe-2", "thinkingLevel": "high" },
20
+ "swe2-max": { "provider": "devin", "model": "swe-2", "thinkingLevel": "max" },
21
+ "swe2-fast-edit": { "provider": "devin", "model": "swe-2", "thinkingLevel": "medium" }
22
+ }
23
+ ```
24
+
25
+ For a new session without presets:
26
+
27
+ ```bash
28
+ omk --provider devin --model swe-2 --thinking max
29
+ ```
30
+
31
+ ## Effort tiers
32
+
33
+ SWE-2 exposes exactly three server-declared efforts. OMK maps its thinking tiers as follows and rejects anything else before sending credentials; `max` is a reasoning level, not the Devin Max subscription tier.
34
+
35
+ | OMK tier | `swe-2` |
36
+ | --- | --- |
37
+ | `off`, `minimal`, `low` | unavailable |
38
+ | `medium` | `medium` |
39
+ | `high` | `high` |
40
+ | `xhigh`, `ultra` | unavailable |
41
+ | `max` | `max` |
42
+
43
+ Guidance from the [SWE-2 announcement](https://cognition.com/blog/swe-2): `medium` makes its first real edit sooner and is the cost-efficient choice for simple and intermediate tasks; `high` and `max` plan more, explore more of the codebase, and verify more on complex tasks. `recommendedDevinEffortForIntent()` encodes the same split (`code` → `medium`, `test` → `high`, `debug`/`repo` → `max`).
44
+
45
+ ## Context budget: 1,000,000 tokens
46
+
47
+ `devin/swe-2` ships with `contextWindow: 1000000` and `maxTokens: 16384`. These are local budgets that drive OMK's context budgeting and compaction, **not published SWE-2 limits**; Cognition has not published a context window for SWE-2. Chat completion settings follow the captured native Devin CLI 3000.6.2 defaults (`maxNewlines` 400, empty stop list). `topP=0.95` is protobuf field 8; field 6 is `firstTemperature` and is omitted. This removes the old 200-newline cap and synthetic stops, but does not prove that every interrupted reply had that cause. The budget also selects the catalog lane:
48
+
49
+ 1. Before each turn OMK reads `GetCliModelConfigs`. SWE-2 family entries may carry a `1M Context` axis (order `1`) beside the effort axis. A local budget of 1,000,000 or more asks for that 1M-context lane; below it, the standard lane is used and 1M entries are ignored.
50
+ 2. A catalog with no 1M-context lane keeps the standard lane for the selected effort.
51
+ 3. If the chosen lane declares a context window smaller than the local budget, the request fails with `... declares a N-token context window; lower the models.json contextWindow before retrying`. OMK never shrinks the budget silently, never invents a wire UID, and never downgrades to another effort.
52
+ 4. Fast-lane (`Fast Mode`) entries are excluded from effort routing and are reachable only through their own UID models (ids ending in `-fast` or `-priority`). Output is capped against the authenticated catalog's declared maximum.
53
+
54
+ To lower the budget (for example if your account only serves the standard lane at 262,144 tokens), override the built-in model in `~/.omk/agent/models.json`:
55
+
56
+ ```json
57
+ {
58
+ "providers": {
59
+ "devin": {
60
+ "modelOverrides": {
61
+ "swe-2": { "contextWindow": 262144 }
62
+ }
63
+ }
64
+ }
65
+ }
66
+ ```
67
+
68
+ Recommended compaction settings for 1M sessions keep the defaults but raise the recent-token window so summaries do not discard the working set:
69
+
70
+ ```json
71
+ {
72
+ "compaction": { "enabled": true, "reserveTokens": 16384, "keepRecentTokens": 60000, "maxUsageRatio": 0.85 }
73
+ }
74
+ ```
75
+
76
+ Context discipline still applies: the budget is room for the repository, not an invitation to dump it. Prefer targeted reads and searches, keep tool output bounded, and let `precompact-checkpoint` snapshot state before compaction rather than restarting sessions. Quota is per account; a larger window consumes more of it per turn.
77
+
78
+ With a Devin credential configured, the status rail's USAGE section shows the account's daily and weekly quota meters from `GetUserStatus` (or the plan name and credit balances when the plan reports no quota windows), matching the CLI's `/usage` surface.
79
+
80
+ ## Domain routing
81
+
82
+ Selecting the `devin` provider auto-applies the `devin-harness` loadout by default; this does not require `OMK_DOMAIN_ROUTING=1`. Set `OMK_DEVIN_HARNESS=0` to disable that provider-specific dispatch. The flag is independent from `OMK_GROK_HARNESS`.
83
+
84
+ Loadout policies reject extension tools that shadow builtins fail-closed. A host that intentionally replaces the builtin `bash` (for example a Landstrip shell provider) should keep `OMK_DEVIN_HARNESS=0` in its launcher until the replacement is removed or bridged; the `devin.md` overlay and the model budget still apply in that case.
85
+
86
+ General prompt-based domain routing is separate and opt-in through `OMK_DOMAIN_ROUTING=1`. It selects one of the profiles under [`loadout-domains/`](loadout-domains/README.md) and composes it with the active role loadout. SWE-2 presets only set provider, model, thinking level, and instruction pointers.
87
+
88
+ ## Model selection
89
+
90
+ The `devin` catalog leads with the logical `swe-2` model; the server's SWE-2 family metadata supplies each effort's wire UID at request time. Every other lane the account catalog advertises is its own logical model whose id is the wire UID — for example `devin/claude-opus-5-high`, `devin/gpt-5-6-sol-xhigh`, `devin/gemini-3-8-flash-medium`, `devin/kimi-k3-max`, `devin/glm-5-3-high`, `devin/grok-4-6-xhigh`, `devin/deepseek-v4-pro-max`, `devin/swe-1-7`, or `devin/inkling-max`. Flat models pin their declared effort lane, so `/think` levels are fixed per model and `No Thinking`/`None` lanes report `reasoning: false`. Availability is account- and plan-dependent: a lane absent from your catalog fails loudly instead of being remapped. Use `/model` or `omk --list-models devin` for the current list. Image input is unsupported; provide text.
91
+
92
+ ## Skill and MCP matrix summary
93
+
94
+ Use the normal OMK lane grant model: grant the smallest skill and MCP surface that matches the task.
95
+
96
+ For each non-queued `devin` request started through `AgentSession.prompt()`, OMK calls `selectDevinHarnessSkills()` against the live discovered skill descriptions after ordinary prompt-template expansion. It merges up to three matches with explicit/settings selections and rebuilds that turn's `<active_skills source="devin-harness">` marker. The scorer is the same `selectSkills()` used by the Grok harness: weak 0.35 / strong 0.7 thresholds, deterministic input-order ties, first-name-wins deduplication, explicit-only skills never auto-selected, and `headroom` only for lexical pressure cues or the session's measured context-pressure bucket. A task with no signals yields an empty automatic grant rather than the full allowlist. Queued `steer`, `followUp`, or `prompt(..., { streamingBehavior })` messages retain the active run's system prompt.
97
+
98
+ | Task class | Skills | MCP |
99
+ | --- | --- | --- |
100
+ | Multi-package or repo-context work | `packages`; add `headroom` only under context pressure | none by default |
101
+ | Repo graph or broad comprehension | `understand-anything`; optionally `packages` | `understand-anything` |
102
+ | TypeScript/Rust/Python/Go edits | `programming`; add `lsp` or `ast-grep` only for symbol/structural work | none by default |
103
+ | New behavior or bug fix with regression test | `tdd-workflow`, `programming` | none by default |
104
+ | Runtime failures or broken behavior | `debugging` | task-specific only |
105
+ | Library API lookup | task skill as needed | `context7` |
106
+ | Current public URL or docs lookup | task skill as needed | `fetch` |
107
+ | UI/TUI verification | task skill as needed | `playwright` only when browser/UI evidence is required |
108
+
109
+ Relevant evidence hooks for SWE-2 lanes are `pre-shell-guard`, `protect-secrets`, `typecheck-after-edit`, `stop-verify`, `session-context`, and `precompact-checkpoint`. Hook output is incremental evidence; code changes still need the project's required final verification command before claiming type/lint cleanliness.
110
+
111
+ ## Suggested TUI flow
112
+
113
+ 1. Run `/login devin` once; the account picker stores the CLI session token in `auth.json`.
114
+ 2. Select `/preset swe2-verified` for normal coding work, `/preset swe2-max` for long-horizon or repository-wide tasks, and `/preset swe2-fast-edit` for small edits.
115
+ 3. Use `/think medium`, `/think high`, or `/think max` to change effort mid-session; other levels are rejected.
116
+ 4. If a turn fails with an "unavailable or ambiguous" route or a smaller declared context window, treat it as a configuration signal: check `devin models list`, or lower `contextWindow` as shown above. Do not retry with a guessed wire UID.
117
+ 5. Keep credentials out of preset JSON, prompts, and logs: the session token, the exchanged user JWT, and `auth.json` contents are secrets under `protect-secrets`.
118
+
119
+ ## Troubleshooting `does not provide an export named`
120
+
121
+ If Devin fails with `The requested module './devin-connect.js' does not provide an export named 'MAX_FRAME_BYTES'`, the loaded stream module is mixed with an older unary module. That is a client load error, not an orphan tool call or a remote protocol trailer.
122
+
123
+ - Current OMK keeps the 16 MiB Connect frame cap inside `devin-connect-stream.ts`, so a stale `devin-connect.js` cannot fail that named import.
124
+ - `/new` does not reload already-imported provider modules. Quit and restart OMK after a rebuild.
125
+ - The failure banner should say to restart OMK, not to sanitize a sticky transcript.
126
+
127
+ ## Troubleshooting `invalid_argument`
128
+
129
+ A Connect `invalid_argument` trailer is a provider/request failure, not evidence of
130
+ an orphan tool call or a safety refusal. Keep any reported trace ID for support;
131
+ OMK includes only a bounded hexadecimal trace ID, never the remote error body.
132
+
133
+ - The SWE-2 adapter rejects zero, negative, and non-finite `temperature` values
134
+ before sending credentials. A controlled high-effort probe returned
135
+ `invalid_argument` at `temperature: 0`, while the default temperature completed
136
+ the same prompt. Omit the option to use the native default of `1`; OMK does not
137
+ silently replace an explicit zero.
138
+ - Check model, effort, context budget, and request settings with `/debug`. Repair
139
+ tool history only when a tool/message mismatch is actually identified.
140
+ - After updating or rebuilding the adapter, quit and restart the OMK process.
141
+ `/new` replaces conversation state; it does not reload an imported provider
142
+ module. This is an update-application step, not a guaranteed fix for every
143
+ provider error.
144
+ - If a minimal tool-free request still fails at the default temperature, preserve
145
+ the trace ID and investigate provider compatibility or availability rather than
146
+ repeatedly sanitizing the same transcript.
147
+
148
+ ## Local overlay
149
+
150
+ When the `devin` provider is active, OMK appends `~/.omk/agent/devin.md` (capped at 24,000 characters) to the system prompt, mirroring the Grok `grok.md` overlay. Treat that file as optional host configuration for effort defaults, compaction notes, or team conventions; this page and the current provider documentation remain authoritative and it cannot override higher-priority instructions.
package/docs/docs.json CHANGED
@@ -27,6 +27,10 @@
27
27
  "title": "Native xAI Grok",
28
28
  "path": "grok-harness.md"
29
29
  },
30
+ {
31
+ "title": "Devin SWE-2",
32
+ "path": "devin-harness.md"
33
+ },
30
34
  {
31
35
  "title": "Containerization",
32
36
  "path": "containerization.md"
@@ -0,0 +1,83 @@
1
+ # ECRAF dimensionless normalization (R11)
2
+
3
+ Status: **pure planner, opt-in, not wired** — the same gate vocabulary as
4
+ [runtime-algorithms](./runtime-algorithms.md) applies. This page documents what is
5
+ implemented and verified, and states explicitly what remains unproven.
6
+
7
+ ## Versions
8
+
9
+ `planEcrafAdmissions()` accepts an explicit `algorithmVersion`:
10
+
11
+ | Version | Density formula | When selected |
12
+ | --- | --- | --- |
13
+ | `legacy-v1` (default when omitted) | `P_i / (epsilon + Σ_r weight_r · a_ir)` | No normalization options supplied. |
14
+ | `normalized-v2` | `P_i / (epsilon + slotCost + Σ_r weight_r · a_ir / s_r)` | `algorithmVersion: "normalized-v2"` **or** `referenceScales` supplied without a version (compatibility with the earlier scales-only opt-in). |
15
+
16
+ Compatibility policy:
17
+
18
+ - Omitting `algorithmVersion` and every normalization option reproduces the
19
+ byte-for-byte legacy plan. Legacy regression and property tests are unchanged.
20
+ - Supplying `referenceScales` (or `slotCost`) with `legacy-v1` is rejected with
21
+ `RangeError`; the scales-only shape silently choosing a different algorithm was
22
+ the ambiguity this version policy closes.
23
+ - Unknown versions and unknown future options are rejected, not coerced.
24
+
25
+ ## normalized-v2 semantics (spec §13.2)
26
+
27
+ - Every resource with a nonzero demand needs a positive finite reference scale
28
+ `s_r`. In `normalized-v2` omitted entries default to the resource's **positive
29
+ total capacity** — never remaining headroom — so the caller cannot implicitly
30
+ re-scale ranking by scheduling pressure. Unbounded resources (capacity omitted)
31
+ cannot fall back to a scale and require an explicit one.
32
+ - Zero capacity is a **feasibility gate**, not a scale: a positive demand on a
33
+ zero-capacity resource is deferred before scoring; a zero demand still uses the
34
+ resource in the legacy sense (admission fails on held usage), so the documented
35
+ "missing capacity = unbounded" and `capacity: 0` meanings are preserved.
36
+ - `slotCost` (λ_slot) is finite and **positive**; the slot term charges every
37
+ running candidate one execution slot even when its resource vector is empty or
38
+ all-zero, so an empty node cannot dominate on `epsilon` alone.
39
+ - Infinite capacities are rejected in v2 (`RangeError`); the legacy finite-input
40
+ contract is unchanged. Finite inputs can still overflow the denominator,
41
+ density, or reserved usage, and those derived values are rejected with
42
+ `RangeError` as before.
43
+
44
+ ## Verified properties (unit invariance, spec §13.3)
45
+
46
+ `a'_ir = c_r·a_ir` with `s'_r = c_r·s_r` (`c_r > 0`) produces the identical plan.
47
+ The seeded property test (`tool-dag-ecraf-normalization.test.ts`) replays the same
48
+ batch under independent power-of-two memory/CPU rescaling, both bounded and
49
+ unbounded, with conflicts, held usage, and slot caps, and additionally asserts
50
+ partition completeness, slot bounds, feasibility, conflict invariants, and input
51
+ immutability. The spec's worked example holds: memory scale 4 GiB / CPU scale 8
52
+ gives A = 0.625 < B = 0.75 in both GiB and byte units.
53
+
54
+ Unit normalization is **not** a fairness guarantee or an optimality proof, and it
55
+ is not equivalent to DRF.
56
+
57
+ ## What this is not
58
+
59
+ - **No live path calls the planner.** `runDagFrontier()` in
60
+ `packages/agent/src/agent-loop.ts` admits ready calls by source order and
61
+ settled-claim conflicts; it does not import `tool-dag-ecraf`. There is no
62
+ shadow recording, feature flag, or default change, and **no measured benefit**
63
+ is claimed anywhere.
64
+ - **The conflict predicate covers only this pass.** A live caller must re-check
65
+ unsettled claims, re-validate post-hook arguments, and refuse stale plans
66
+ before granting (spec §13.4). None of that wiring exists.
67
+ - Starvation/fairness accounting (spec §13.5) and same-budget comparisons
68
+ (§13.6 steps 4–6) are separate, unstarted work.
69
+
70
+ ## Tests
71
+
72
+ - `packages/agent/test/tool-dag-ecraf-normalization.test.ts` — version policy,
73
+ scale derivation, zero-capacity gating, slot cost, overflow/validation
74
+ boundaries, unit-invariance property test (seed 110917, 300 runs).
75
+ - `packages/agent/test/tool-dag-ecraf.test.ts` — legacy behavior, input
76
+ rejection, admission invariants, and the §13.3 GiB/bytes ranking-flip example.
77
+ - `packages/agent/test/tool-dag-ecraf-arithmetic.test.ts` — finite-arithmetic
78
+ overflow rejection, unchanged.
79
+
80
+ Mutation/negative-control evidence: removing the demand/scale division, weakening
81
+ slot-cost positivity to `< 0`, removing the zero-capacity gate, or removing the
82
+ version whitelist each makes the assertion harness fail (run from a /tmp mutant
83
+ copy; the working tree was not modified during the check).
@@ -100,7 +100,8 @@ These variables are read by OMK itself. The four built-in harness flags below ar
100
100
  | `OMK_YOLO`, `OMK_COMMAND_SAFETY`, `OMK_DISABLE_COMMAND_SAFETY` | Disable the command-safety gate entirely (YOLO mode). `OMK_YOLO` and `OMK_DISABLE_COMMAND_SAFETY` accept `1`, `true`, `yes`, or `on`; `OMK_COMMAND_SAFETY` accepts `0`, `false`, `off`, `disable`, or `disabled`. Every verdict, including block-tier and privilege commands, is skipped in interactive and headless runs. Use only when a verified outer sandbox owns the boundary |
101
101
  | `OMK_COMMAND_SAFETY_ASSUME_YES` | `1` or `true` auto-accepts non-privilege confirm-tier commands in interactive **and headless** runs. Privilege confirmation and block-tier commands remain denied. Use only under a trusted outer sandbox when headless auto-accept is intended |
102
102
  | `OMK_GROK_HARNESS` | Default-on native `xai` provider dispatch to the `grok-harness` loadout. `0`, `false`, `off`, or `no` disables it |
103
- | `OMK_DOMAIN_ROUTING` | Set to `1` to enable general prompt-based domain routing. Native xAI harness dispatch does not require it |
103
+ | `OMK_DEVIN_HARNESS` | Default-on `devin` provider dispatch to the `devin-harness` loadout. `0`, `false`, `off`, or `no` disables it; independent from `OMK_GROK_HARNESS` |
104
+ | `OMK_DOMAIN_ROUTING` | Set to `1` to enable general prompt-based domain routing. Native xAI and Devin harness dispatch do not require it |
104
105
  | `VISUAL`, `EDITOR` | External editor fallback when `externalEditor` is unset |
105
106
  | `HTTP_PROXY`, `HTTPS_PROXY` | Proxy outbound HTTP requests |
106
107
  | `OMK_RESOURCE_GOVERNOR` | Resource-governor mode: `off`, `observe` (default), `adaptive`, or `strict`. Feeds `/resource [probe\|policy]` and `omk doctor resources [--json]`; see the resource governor section in [Settings](settings.md) |
package/docs/index.md CHANGED
@@ -39,6 +39,7 @@ For the full first-run flow, see [Quickstart](quickstart.md).
39
39
  - [Providers](providers.md) - subscription and API-key setup for built-in providers.
40
40
  - [Provider Resilience](provider-resilience.md) - retry, failover, quota, and safety-stop recovery.
41
41
  - [Native xAI Grok](grok-harness.md) - authentication, weekly SuperGrok usage, presets, and thinking tiers.
42
+ - [Devin SWE-2](devin-harness.md) - presets, effort tiers, the 1M-token context budget and lane rule, and the `devin-harness` loadout.
42
43
  - [Containerization](containerization.md) - sandbox omk with OpenShell, Gondolin, or Docker.
43
44
  - [Settings](settings.md) - global and project settings.
44
45
  - [Environment Variables](environment-variables.md) - process configuration, harness opt-outs, and bash-tool session environment.
@@ -21,7 +21,7 @@ OMK routes incoming tasks to a **domain capability profile** ("inherited documen
21
21
 
22
22
  Thresholds: `STRONG_THRESHOLD = 8`, `WEAK_THRESHOLD = 4`, `AMBIGUITY_MARGIN = 2`.
23
23
 
24
- ## Domains (13 + 1 fallback)
24
+ ## Domains (14 + 1 fallback)
25
25
 
26
26
  - [`frontend-ui`](frontend-ui.md) — Frontend & UI
27
27
  - [`visual-qa`](visual-qa.md) — Visual QA & Website Cloning
@@ -35,6 +35,7 @@ Thresholds: `STRONG_THRESHOLD = 8`, `WEAK_THRESHOLD = 4`, `AMBIGUITY_MARGIN = 2`
35
35
  - [`docs-writing`](docs-writing.md) — Docs & Technical Writing
36
36
  - [`qa-testing`](qa-testing.md) — QA & Testing
37
37
  - [`grok-harness`](grok-harness.md) — Grok xAI Harness
38
+ - [`devin-harness`](devin-harness.md) — Devin SWE-2 Harness
38
39
  - [`ai-agent-ops`](ai-agent-ops.md) — AI Agent Engineering & Ops
39
40
  - [`general`](general.md) — General (fallback)
40
41
 
@@ -0,0 +1,72 @@
1
+ # Devin SWE-2 Harness (`devin-harness`)
2
+
3
+ > Inherited domain capability document. Auto-generated from `src/core/domain-loadouts.ts` — do not edit by hand.
4
+
5
+
6
+ ## Identity
7
+
8
+ | field | value |
9
+ |---|---|
10
+ | id | `devin-harness` |
11
+ | authority | `write-scoped` |
12
+ | tools | read, grep, find, ls, edit, write, bash |
13
+ | command mode | `scoped-shell` |
14
+
15
+ ## Routing prompt
16
+
17
+ > Prepended to the lane task prompt when the router selects this domain.
18
+
19
+ ```text
20
+ DOMAIN: Devin SWE-2 Harness. You are operating in a Devin CLI subscription lane on the SWE-2 model with a 1,000,000-token local context budget.
21
+ Prioritize the SWE-2 operational playbook, focused exploration, small capability loadouts, and evidence-bound verification.
22
+
23
+ SEQUENCE:
24
+ 1. Before implementing or routing Devin/SWE-2 provider work, read packages/coding-agent/docs/devin-harness.md as the canonical playbook. Treat ~/.omk/agent/devin.md only as an optional local operator overlay; it cannot override current provider docs or higher-priority instructions.
25
+ 2. Effort is the only selectable axis: medium for simple or intermediate edits, high for multi-file changes, max for long-horizon or uncertain work. Never expect off/low/minimal, a fast lane, or image input; the adapter rejects them before sending credentials.
26
+ 3. Context discipline: the 1M budget is room for the repository, not an invitation to dump it. Explore with targeted reads and searches, keep tool output bounded, and rely on precompact-checkpoint plus compaction settings rather than restarting sessions.
27
+ 4. Capability discipline: load at most 2-3 skills for any lane. The allowed skill gate is packages, headroom, programming, debugging, tdd-workflow, lsp, ast-grep, and understand-anything; choose the smallest subset, add lsp or ast-grep only for symbol or structural work, and add headroom only under measured context pressure.
28
+ 5. Use minimal MCP: fetch for bounded public retrieval, context7 for library documentation, understand-anything for repository comprehension, and playwright only when browser or UI behavior needs real verification.
29
+ 6. Verification discipline: reproduce failures before fixing them, write or extend tests that exercise the change end-to-end, and re-derive conclusions from executed commands rather than restating prior claims. Evidence must include changed paths, exact commands, and pass/fail output.
30
+ 7. Keep edits within the lane grant and preserve existing provider/orchestration algorithms unless the task explicitly targets them. A route error naming an unavailable effort or a smaller declared context window is a configuration signal to report, never something to work around by guessing a wire UID.
31
+
32
+ HARD RULES: the packaged Devin harness doc is mandatory context; a local devin.md is optional; medium/high/max are the only efforts; the 1M budget never justifies unbounded dumps; maximum 2-3 active skills; never log the Devin session token, user JWT, or auth.json contents; protect-secrets applies.
33
+ ```
34
+
35
+ ## Curated skills (8)
36
+
37
+ - `packages`
38
+ - `headroom`
39
+ - `programming`
40
+ - `debugging`
41
+ - `tdd-workflow`
42
+ - `lsp`
43
+ - `ast-grep`
44
+ - `understand-anything`
45
+
46
+ ## Curated MCP servers (4)
47
+
48
+ - `fetch`
49
+ - `context7`
50
+ - `understand-anything`
51
+ - `playwright`
52
+
53
+ ## Curated hooks (6)
54
+
55
+ - `pre-shell-guard`
56
+ - `protect-secrets`
57
+ - `typecheck-after-edit`
58
+ - `stop-verify`
59
+ - `session-context`
60
+ - `precompact-checkpoint`
61
+
62
+ ## Routing triggers (7)
63
+
64
+ | kind | pattern | weight |
65
+ |---|---|---|
66
+ | keyword | `devin` | 8 |
67
+ | keyword | `swe-2` | 8 |
68
+ | keyword | `swe2` | 8 |
69
+ | keyword | `cognition` | 6 |
70
+ | keyword | `devin cli` | 8 |
71
+ | keyword | `1m context` | 5 |
72
+ | regex | `\b(devin|swe[- ]?2|cognition)\b` | 7 |
package/docs/metrics.md CHANGED
@@ -160,7 +160,7 @@ node scripts/tb-mini-suite.mjs --json # feed a runner
160
160
  node scripts/tb-mini-suite.mjs --seed 7 # a different fixed subset
161
161
  ```
162
162
 
163
- With identical task metadata, seed, size, and collation, selection is repeatable.
163
+ With identical task metadata, selection version, seed, size, and collation, selection is repeatable.
164
164
  The default 15-task subset oversamples easy tasks and prioritizes shorter expert
165
165
  time estimates; it is a regression signal, not a population-representative score
166
166
  or an agent runtime bound. Selection alone is not a capability result. Scoring
@@ -171,12 +171,35 @@ population. Missing difficulty quotas are filled from unselected tasks using the
171
171
  same ordering, so a valid request returns exactly that many distinct tasks.
172
172
  `--seed` accepts integers from `0` through `4294967295`. Invalid or missing option
173
173
  values and oversized requests exit with code `2`; absent, empty, or non-directory
174
- task paths exit with code `1`. The existing default subset remains unchanged.
174
+ task paths exit with code `1`.
175
+
176
+ The JSON output now declares `selectionVersion: 2`. Missing, empty, nonfinite, or
177
+ negative expert-time estimates are `null`, not zero. Within each difficulty band
178
+ and in quota refill, known estimates sort before unknown estimates. A genuine zero
179
+ or fractional estimate remains valid. Difficulty quotas still take precedence, so
180
+ unknown-estimate tasks can be selected to fill a band.
181
+
182
+ `knownExpertMinutes` sums known estimates; `unknownExpertEstimates` counts selected
183
+ tasks with unknown estimates. `totalExpertMinutes` is `null` when any selected
184
+ estimate is unknown. A nonfinite sum is refused with exit `1`, not serialized as
185
+ an apparently missing total. Expert estimates are not agent timeout limits.
186
+
187
+ For sizes 1–2, available slots go first to the highest-weight difficulty bands
188
+ (medium, then hard), with normal refill if those bands are unavailable. Task names
189
+ no longer decide which excess band quota is discarded. The normal-size quota rule
190
+ is preserved. Human-readable counts use a `Map`, including for labels such as
191
+ `__proto__` that overlap JavaScript object properties.
192
+
193
+ This is a selection and output-contract change: default membership can change when
194
+ metadata is incomplete, and the prior default JSON digest is historical only.
195
+ Freeze new task manifests before comparison; do not combine version 1 and 2 runs
196
+ as if their selection policy were identical. The limited flat TOML field reader
197
+ is unchanged; this does not add general TOML syntax support.
175
198
 
176
199
  Run the offline CLI regression tests without downloading tasks or calling models:
177
200
 
178
201
  ```bash
179
- node --test scripts/test/tb-mini-suite.test.mjs
202
+ node --test --test-concurrency=1 scripts/test/tb-mini-suite.test.mjs scripts/test/tb-mini-suite-ranking.test.mjs
180
203
  ```
181
204
 
182
205
  See [the harness roadmap](../../../ROADMAP.md) for the dated OMK versus Terminus-2
@@ -187,8 +210,8 @@ criteria. Planned runtime improvements are not measured benchmark gains.
187
210
 
188
211
  The checkout-only `scripts/tb21-audit.mjs` audits explicitly selected Harbor jobs
189
212
  against a caller-pinned manifest digest. It rejects duplicate tasks/trials, missing
190
- results or costs, mismatched task checksums or configured model labels, and
191
- contradictory success records. It never starts a model, picks the latest job, joins
213
+ results or costs, mismatched task checksums or configured model labels,
214
+ missing/invalid completion times, and contradictory success records. It never starts a model, picks the latest job, joins
192
215
  requests by timestamp, or rewrites evidence. See [TB 2.1 offline audit](tb21-audit.md)
193
216
  for the schema, invocation, error codes, and limitations.
194
217
 
@@ -196,14 +219,32 @@ A complete audit means recorded outcomes passed these checks, not that every
196
219
  provider request obeyed a single-model contract. Wire provenance, actual billing,
197
220
  repeated-trial analysis, and statistical superiority need separate evidence.
198
221
 
199
- ### Explicit output-limit validation
200
-
201
- For callers that supply `AgentLoopConfig.modelContract`, both the contract's
202
- `maxOutputTokens` and an explicitly supplied request `maxTokens` must be positive
203
- safe integers. Invalid explicit values are refused before `provider_request` and
204
- before calling the provider stream function; they are not treated as absent.
205
-
206
- An omitted request limit still leaves provider defaults unchecked by this
207
- predicate. This change does not activate a contract in the CLI or impose an
208
- effective cap on compaction and other provider paths. Full run-wide enforcement
209
- remains a roadmap item.
222
+ ### Output-limit validation: availability history
223
+
224
+ **2026-09-13 source snapshot (`ca75f4e5cc`):** the logical contract, CLI/SDK
225
+ wiring, and final Chat Completions model/output-limit checks are in the committed
226
+ source. [Model dispatch contracts](model-contract.md) defines their coverage;
227
+ [Verified Run](verified-run.md) describes the separate protected execution path.
228
+ Neither is a universal provider billing cap or a new controlled benchmark result.
229
+ See [the roadmap, section 16](../../../ROADMAP.md) for current local release-preparation checks.
230
+
231
+ The following paragraphs describe the earlier checkout only. Its missing modules
232
+ and failing collection are historical observations, not active release blockers.
233
+
234
+ **2026-09-08 follow-up:** the current worktree restores the logical contract and
235
+ connects it through the CLI/SDK, including SDK-stream summaries. See
236
+ [Model dispatch contracts](model-contract.md) and ROADMAP §13 for fresh evidence
237
+ and the remaining final-wire/accounting gaps. The warning below records the
238
+ preceding checkout, not the current availability of the restored module.
239
+
240
+ The previous worktree checkpoint tested positive-safe-integer validation of
241
+ `modelContract.maxOutputTokens` and explicit request `maxTokens`. During the
242
+ 2026-09-08 re-verification, the checkout changed: `run-model-contract.ts` and the
243
+ corresponding `AgentLoopConfig.modelContract` surface were absent. The remaining
244
+ `model-contract-output-limit.test.ts` fails collection against that checkout.
245
+
246
+ Do not treat the historical passing tests as proof that this guard is currently
247
+ available. Restoring or porting the runtime contract requires an explicit source
248
+ baseline decision and fresh send-boundary tests. No missing code was silently
249
+ recreated and no failing test was deleted. Full run-wide enforcement, including
250
+ omitted limits and compaction, remains unverified; see ROADMAP sections 11–12.
@@ -3,6 +3,56 @@
3
3
  확인일: 2026-09-09. 생성기와 공급자 어댑터를 수정한 뒤 `npm run models:refresh`로
4
4
  두 카탈로그를 재생성했다. 생성 파일을 손으로 수정하지 않았다.
5
5
 
6
+ ## 2026-09-17 후속: OpenCode Go DeepSeek V4.1 ID 변경
7
+
8
+ [OpenCode Go 공식 endpoint 목록](https://opencode.ai/docs/go/)의 현재 ID는
9
+ `deepseek-v4.1-flash`다. 직접 DeepSeek의 `deepseek-flash`와 구분한다.
10
+ 9월 17일 카탈로그 갱신은 새 ID를 반영했지만, 생성기의 V4.1 메타데이터 보정은
11
+ 이전 ID만 인식해 `low/max`와 `max_tokens` 설정이 누락됐다.
12
+
13
+ 생성기 조건에 Go의 새 ID를 추가하고 기존 Go ID의 보정도 유지했다. 회귀 테스트는
14
+ 공급자별 실제 요청 ID를 사용하며 `off/low/high/max`, 이미지 입력, 출력 상한,
15
+ 전송 직전 `thinking`과 `reasoning_effort` 검사를 그대로 유지한다.
16
+ ID만 교체한 상태에서도 5개 실패가 재현됐으며 메타데이터 수정 후 통과했다.
17
+
18
+ `node packages/ai/scripts/generate-models.ts`로 생성한 결과 중 해당 Go 모델의
19
+ 메타데이터 변경만 포함했다. 실시간 목록에서 함께 발생한 다른 모델의 추가/삭제,
20
+ 가격/상한 변경은 이번 수정에 포함하지 않았다. 생성 파일 값을 수작업으로 만들지 않았다.
21
+ 공급자 추론이나 계정별 사용 가능 여부는 검증하지 않았다. 아래 9월 10일 표는 당시 기록이다.
22
+
23
+ ## 2026-09-17 갱신: OpenRouter Union Alpha (stealth)
24
+
25
+ `stealth/union-alpha` 추가 요청으로 `npm run models:refresh`를 종료0으로 재생성했다.
26
+ 생성 파일은 손대지 않았다. 전 소스가 응답했고 `--allow-partial`은 쓰지 않았다.
27
+
28
+ | 항목 | 값 |
29
+ | --- | --- |
30
+ | 요청 ID | `stealth/union-alpha` (OpenRouter) |
31
+ | context / maxTokens | 262,144 / 131,072 |
32
+ | 입력 | text, image |
33
+ | tool 지원 | `tools`, `tool_choice` |
34
+ | 가격 | prompt/completion 모두 `0` |
35
+ | **thinking** | **없음 — route가 `reasoning`을 선언하지 않음** |
36
+
37
+ 목록과 `/api/v1/models/stealth/union-alpha/endpoints` 모두
38
+ `supported_parameters`가 `max_tokens, temperature, top_p, tools, tool_choice,
39
+ response_format`이다. `reasoning`도 `include_reasoning`도 없다. 선언이 없으므로
40
+ 수준을 지어내지 않고 `reasoning: false`로 들어갔으며 `thinkingLevelMap`도 없다.
41
+ 이름이 frontier 계열을 연상시킨다는 이유로 effort를 이식하지 않는다.
42
+
43
+ `created`는 2026-09-16으로 직전 갱신(09-09) 이후에 생긴 항목이다. 한 모델만
44
+ 집어넣는 경로가 없어 카탈로그 전체가 8일치 드리프트를 함께 반영한다.
45
+ OpenRouter 371 → 376, 전체 추가 57 · 제거 28(고유 4)이다. 제거는
46
+ `deepseek.r1-v1:0`과 mistral `devstral-small-2` · `mistral-medium` · `pixtral-12b`로,
47
+ 갱신 소스에서 더 이상 선정되지 않았다는 뜻이며 공급자의 폐기 공지나 모든 계정의
48
+ 사용 불가를 뜻하지 않는다.
49
+
50
+ 가격 `0`은 stealth 공개 기간의 목록 값이다. 무상 사용을 보장하지 않으며 stealth
51
+ 해제 시 달라질 수 있다. 실제 청구는 계정에서 따로 확인한다.
52
+
53
+ 표적 검사 `latest-model-thinking` · `catalog-thinking` · `latest-thinking-payload`
54
+ 51개가 통과했다. 공급자 추론은 호출하지 않았다.
55
+
6
56
  ## 2026-09-10 재검증: DeepSeek V4.1 Flash 제공 경로
7
57
 
8
58
  공식 문서·공개 API에서 확인한 기존 OMK 공급자 4곳을 반영했다.
@@ -114,7 +164,7 @@ OpenRouter의 `reasoning.supported_efforts`로 표시할 수준을 만들고,
114
164
 
115
165
  | 모델·route | 이번 정합성 규칙 |
116
166
  | --- | --- |
117
- | GPT-6 Astra, OpenAI/Responses 및 OpenRouter | low·medium·high·xhigh·max. off/minimal 미노출 |
167
+ | GPT-6 Astra, OpenAI/Responses 및 OpenRouter | low·medium·high·xhigh·max. OMK `ultra`는 공식 천장 `max`의 선택기 별칭이며 와이어에 `ultra`를 보내지 않음. off/minimal 미노출 |
118
168
  | Claude Opus 5, Anthropic Messages | adaptive thinking, low·medium·high·xhigh·max. off는 high 이하에서 가능 |
119
169
  | Claude Opus 5, Bedrock | legacy budget 대신 adaptive 및 xhigh 사용. application profile의 표시명 매칭 보존 |
120
170
  | Gemini 3.7/3.8 Flash, Google/Vertex | low·medium·high. minimal은 API 오류. SDK의 off 요청도 LOW로 처리하며 완전 비활성화로 주장하지 않음 |