@armadra/agent 0.5.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (471) hide show
  1. package/CHANGELOG.md +177 -294
  2. package/CHANGELOG.zh-CN.md +442 -0
  3. package/README.md +408 -364
  4. package/README.zh-CN.md +705 -0
  5. package/dist/agent/session-plan.js +2 -2
  6. package/dist/agent/session-state.d.ts +3 -0
  7. package/dist/agent/session-state.js +20 -0
  8. package/dist/agent/session-subagent.d.ts +15 -1
  9. package/dist/agent/session-subagent.js +26 -3
  10. package/dist/agent/session-telemetry.d.ts +16 -2
  11. package/dist/agent/session-telemetry.js +29 -7
  12. package/dist/agent/session-tools.js +5 -2
  13. package/dist/agent/session-trace-writer.d.ts +27 -0
  14. package/dist/agent/session-trace-writer.js +253 -0
  15. package/dist/agent/session.d.ts +2 -0
  16. package/dist/agent/session.js +12 -0
  17. package/dist/agent/subagent-direct.d.ts +20 -0
  18. package/dist/agent/subagent-direct.js +76 -0
  19. package/dist/agent/subagent-registry.d.ts +8 -1
  20. package/dist/agent/subagent-registry.js +28 -1
  21. package/dist/agent/system-prompt.d.ts +3 -1
  22. package/dist/agent/system-prompt.js +3 -0
  23. package/dist/agent/tool-runner.d.ts +2 -1
  24. package/dist/agent/tool-runner.js +5 -4
  25. package/dist/agent/types-w6.d.ts +39 -0
  26. package/dist/agent/types-w6.js +4 -0
  27. package/dist/agent/types.d.ts +5 -2
  28. package/dist/agents/external.js +1 -1
  29. package/dist/agents/task-record.d.ts +74 -1
  30. package/dist/agents/task-record.js +100 -0
  31. package/dist/ai/apis/cache-params.js +2 -1
  32. package/dist/ai/apis/chatgpt-backend.d.ts +52 -0
  33. package/dist/ai/apis/chatgpt-backend.js +224 -0
  34. package/dist/ai/apis/chatgpt-rate-limits.d.ts +26 -0
  35. package/dist/ai/apis/chatgpt-rate-limits.js +91 -0
  36. package/dist/ai/apis/openai-responses-request.d.ts +2 -1
  37. package/dist/ai/apis/openai-responses-request.js +8 -2
  38. package/dist/ai/apis/openai-responses.d.ts +2 -0
  39. package/dist/ai/apis/openai-responses.js +51 -16
  40. package/dist/ai/overflow.d.ts +1 -1
  41. package/dist/ai/overflow.js +10 -7
  42. package/dist/ai/providers/auth.d.ts +15 -2
  43. package/dist/ai/providers/auth.js +69 -2
  44. package/dist/ai/providers/builtin.js +27 -0
  45. package/dist/ai/providers/catalog-data.js +1 -0
  46. package/dist/ai/providers/enrich.js +2 -7
  47. package/dist/ai/providers/models-dev-cache.js +15 -19
  48. package/dist/ai/providers/models-dev-snapshot.js +8 -7
  49. package/dist/ai/providers/models-dev.js +5 -15
  50. package/dist/ai/providers/suggest.js +2 -16
  51. package/dist/ai/types.d.ts +27 -1
  52. package/dist/auth/chatgpt/backend-client.d.ts +26 -0
  53. package/dist/auth/chatgpt/backend-client.js +74 -0
  54. package/dist/auth/chatgpt/claims.d.ts +17 -0
  55. package/dist/auth/chatgpt/claims.js +46 -0
  56. package/dist/auth/chatgpt/cli.d.ts +33 -0
  57. package/dist/auth/chatgpt/cli.js +252 -0
  58. package/dist/auth/chatgpt/doctor.d.ts +17 -0
  59. package/dist/auth/chatgpt/doctor.js +36 -0
  60. package/dist/auth/chatgpt/host-id.d.ts +7 -0
  61. package/dist/auth/chatgpt/host-id.js +26 -0
  62. package/dist/auth/chatgpt/login.d.ts +30 -0
  63. package/dist/auth/chatgpt/login.js +130 -0
  64. package/dist/auth/chatgpt/presets.d.ts +66 -0
  65. package/dist/auth/chatgpt/presets.js +120 -0
  66. package/dist/auth/chatgpt/quota-text.d.ts +13 -0
  67. package/dist/auth/chatgpt/quota-text.js +36 -0
  68. package/dist/auth/oauth/browser.d.ts +12 -0
  69. package/dist/auth/oauth/browser.js +27 -0
  70. package/dist/auth/oauth/callback-server.d.ts +47 -0
  71. package/dist/auth/oauth/callback-server.js +147 -0
  72. package/dist/auth/oauth/flows.d.ts +62 -0
  73. package/dist/auth/oauth/flows.js +113 -0
  74. package/dist/auth/oauth/jwt.d.ts +24 -0
  75. package/dist/auth/oauth/jwt.js +81 -0
  76. package/dist/auth/oauth/live.d.ts +21 -0
  77. package/dist/auth/oauth/live.js +25 -0
  78. package/dist/auth/oauth/oidc.d.ts +26 -0
  79. package/dist/auth/oauth/oidc.js +72 -0
  80. package/dist/auth/oauth/pkce.d.ts +15 -0
  81. package/dist/auth/oauth/pkce.js +20 -0
  82. package/dist/auth/oauth/refresh.d.ts +35 -0
  83. package/dist/auth/oauth/refresh.js +129 -0
  84. package/dist/auth/oauth/token-client.d.ts +33 -0
  85. package/dist/auth/oauth/token-client.js +121 -0
  86. package/dist/auth/oauth/token-store.d.ts +23 -0
  87. package/dist/auth/oauth/token-store.js +125 -0
  88. package/dist/auth/testing/fake-oauth.d.ts +73 -0
  89. package/dist/auth/testing/fake-oauth.js +271 -0
  90. package/dist/auth/testing/refresh-child.d.ts +5 -0
  91. package/dist/auth/testing/refresh-child.js +16 -0
  92. package/dist/bundle/ama.cjs +28307 -13241
  93. package/dist/checkpoints/backend.js +6 -5
  94. package/dist/checkpoints/restore.js +3 -2
  95. package/dist/checkpoints/settings.js +2 -1
  96. package/dist/checkpoints/shadow-git.js +7 -6
  97. package/dist/checkpoints/shadow-restore.js +3 -2
  98. package/dist/checkpoints/tracker.js +4 -3
  99. package/dist/cli/args.d.ts +9 -2
  100. package/dist/cli/args.js +74 -27
  101. package/dist/cli/bootstrap.js +53 -31
  102. package/dist/cli/choice-prompt.d.ts +58 -0
  103. package/dist/cli/choice-prompt.js +142 -0
  104. package/dist/cli/codemode-notice.js +2 -2
  105. package/dist/cli/compose-agents.js +2 -1
  106. package/dist/cli/compose-extensions.js +2 -0
  107. package/dist/cli/compose-memory.d.ts +53 -0
  108. package/dist/cli/compose-memory.js +127 -0
  109. package/dist/cli/compose-providers.d.ts +5 -0
  110. package/dist/cli/compose-providers.js +32 -2
  111. package/dist/cli/compose-session.d.ts +5 -0
  112. package/dist/cli/compose-session.js +27 -13
  113. package/dist/cli/compose-store.d.ts +2 -0
  114. package/dist/cli/compose-store.js +9 -5
  115. package/dist/cli/compose.d.ts +2 -0
  116. package/dist/cli/compose.js +13 -5
  117. package/dist/cli/default-model.d.ts +0 -2
  118. package/dist/cli/default-model.js +5 -8
  119. package/dist/cli/deps.d.ts +5 -1
  120. package/dist/cli/exit-codes.d.ts +1 -1
  121. package/dist/cli/exit-codes.js +7 -16
  122. package/dist/cli/from-prompt.js +3 -2
  123. package/dist/cli/help-text.d.ts +2 -1
  124. package/dist/cli/help-text.js +5 -99
  125. package/dist/cli/main.d.ts +5 -0
  126. package/dist/cli/main.js +21 -2
  127. package/dist/cli/proxy.js +14 -9
  128. package/dist/cli/runtime.d.ts +5 -0
  129. package/dist/cli/startup-screen.js +20 -27
  130. package/dist/cli/startup-steps.js +6 -8
  131. package/dist/cli/subcommands/auth.d.ts +6 -3
  132. package/dist/cli/subcommands/auth.js +52 -22
  133. package/dist/cli/subcommands/config-set.d.ts +21 -0
  134. package/dist/cli/subcommands/config-set.js +137 -0
  135. package/dist/cli/subcommands/config.d.ts +2 -1
  136. package/dist/cli/subcommands/config.js +51 -37
  137. package/dist/cli/subcommands/doctor.d.ts +2 -1
  138. package/dist/cli/subcommands/doctor.js +97 -58
  139. package/dist/cli/subcommands/init.d.ts +1 -1
  140. package/dist/cli/subcommands/init.js +8 -7
  141. package/dist/cli/subcommands/memory.d.ts +18 -0
  142. package/dist/cli/subcommands/memory.js +189 -0
  143. package/dist/cli/subcommands/model-meta.js +11 -9
  144. package/dist/cli/subcommands/models-cache-probe.d.ts +1 -1
  145. package/dist/cli/subcommands/models-cache-probe.js +26 -32
  146. package/dist/cli/subcommands/models-discover.d.ts +3 -0
  147. package/dist/cli/subcommands/models-discover.js +44 -21
  148. package/dist/cli/subcommands/models.d.ts +1 -1
  149. package/dist/cli/subcommands/models.js +29 -22
  150. package/dist/cli/subcommands/probe-runner.js +8 -7
  151. package/dist/cli/subcommands/providers-list.js +32 -21
  152. package/dist/cli/subcommands/providers-plan.js +24 -25
  153. package/dist/cli/subcommands/providers-probe.js +5 -6
  154. package/dist/cli/subcommands/providers.d.ts +2 -2
  155. package/dist/cli/subcommands/providers.js +48 -59
  156. package/dist/cli/subcommands/sessions-export.d.ts +1 -1
  157. package/dist/cli/subcommands/sessions-export.js +10 -9
  158. package/dist/cli/subcommands/sessions-search.d.ts +1 -1
  159. package/dist/cli/subcommands/sessions-search.js +11 -10
  160. package/dist/cli/subcommands/sessions-trace.d.ts +30 -0
  161. package/dist/cli/subcommands/sessions-trace.js +156 -0
  162. package/dist/cli/subcommands/sessions.d.ts +4 -3
  163. package/dist/cli/subcommands/sessions.js +40 -37
  164. package/dist/cli/subcommands/stats.d.ts +1 -1
  165. package/dist/cli/subcommands/stats.js +49 -30
  166. package/dist/cli/system-prompt-arg.js +3 -2
  167. package/dist/codemode/capability.js +3 -2
  168. package/dist/compaction/post-compact.d.ts +1 -0
  169. package/dist/compaction/post-compact.js +15 -1
  170. package/dist/compaction/prune-tier.d.ts +1 -1
  171. package/dist/compaction/prune-tier.js +3 -3
  172. package/dist/compaction/serialize.js +2 -1
  173. package/dist/config/auth-file.d.ts +12 -2
  174. package/dist/config/auth-file.js +32 -10
  175. package/dist/config/checker.js +12 -11
  176. package/dist/config/context-files.js +3 -2
  177. package/dist/config/edit.d.ts +106 -0
  178. package/dist/config/edit.js +350 -0
  179. package/dist/config/init.d.ts +4 -3
  180. package/dist/config/init.js +10 -18
  181. package/dist/config/json-schema.d.ts +2 -1
  182. package/dist/config/json-schema.js +101 -68
  183. package/dist/config/key-docs.d.ts +18 -7
  184. package/dist/config/key-docs.js +43 -128
  185. package/dist/config/load.js +7 -6
  186. package/dist/config/merge.d.ts +4 -1
  187. package/dist/config/merge.js +42 -22
  188. package/dist/config/paths.js +6 -5
  189. package/dist/config/profile.d.ts +5 -1
  190. package/dist/config/profile.js +7 -2
  191. package/dist/config/schema-w5.js +3 -2
  192. package/dist/config/schema-w6.d.ts +19 -0
  193. package/dist/config/schema-w6.js +115 -0
  194. package/dist/config/schema.js +34 -20
  195. package/dist/config/settings-registry.d.ts +78 -0
  196. package/dist/config/settings-registry.js +184 -0
  197. package/dist/config/types-w6.d.ts +109 -0
  198. package/dist/config/types-w6.js +19 -0
  199. package/dist/config/types.d.ts +11 -8
  200. package/dist/config/types.js +1 -0
  201. package/dist/drivers/acp/client.js +1 -1
  202. package/dist/drivers/acp/driver.js +13 -8
  203. package/dist/drivers/agents.js +4 -3
  204. package/dist/drivers/base.js +3 -2
  205. package/dist/drivers/host-runners.js +2 -1
  206. package/dist/drivers/native/claude-stream.js +12 -11
  207. package/dist/drivers/native/codex-app-server.js +17 -12
  208. package/dist/drivers/native/oneshot.js +8 -7
  209. package/dist/drivers/pool.js +1 -1
  210. package/dist/drivers/runner.d.ts +3 -1
  211. package/dist/drivers/runner.js +68 -20
  212. package/dist/drivers/turn.js +1 -1
  213. package/dist/hooks/config.js +3 -2
  214. package/dist/hooks/protocol.js +18 -12
  215. package/dist/host/api-impl.js +11 -10
  216. package/dist/host/loader.js +9 -8
  217. package/dist/host/types.d.ts +3 -0
  218. package/dist/i18n/catalog.d.ts +2169 -0
  219. package/dist/i18n/catalog.js +72 -0
  220. package/dist/i18n/format.d.ts +21 -0
  221. package/dist/i18n/format.js +58 -0
  222. package/dist/i18n/index.d.ts +58 -0
  223. package/dist/i18n/index.js +72 -0
  224. package/dist/i18n/messages/agents.d.ts +84 -0
  225. package/dist/i18n/messages/agents.js +85 -0
  226. package/dist/i18n/messages/approval.d.ts +99 -0
  227. package/dist/i18n/messages/approval.js +100 -0
  228. package/dist/i18n/messages/auth.d.ts +187 -0
  229. package/dist/i18n/messages/auth.js +189 -0
  230. package/dist/i18n/messages/cli-args.d.ts +46 -0
  231. package/dist/i18n/messages/cli-args.js +46 -0
  232. package/dist/i18n/messages/cli-help.d.ts +12 -0
  233. package/dist/i18n/messages/cli-help.js +244 -0
  234. package/dist/i18n/messages/cli.d.ts +336 -0
  235. package/dist/i18n/messages/cli.js +311 -0
  236. package/dist/i18n/messages/config-keys.d.ts +289 -0
  237. package/dist/i18n/messages/config-keys.js +292 -0
  238. package/dist/i18n/messages/config.d.ts +502 -0
  239. package/dist/i18n/messages/config.js +239 -0
  240. package/dist/i18n/messages/drivers.d.ts +123 -0
  241. package/dist/i18n/messages/drivers.js +124 -0
  242. package/dist/i18n/messages/errors.d.ts +83 -0
  243. package/dist/i18n/messages/errors.js +171 -0
  244. package/dist/i18n/messages/interactive-line.d.ts +61 -0
  245. package/dist/i18n/messages/interactive-line.js +69 -0
  246. package/dist/i18n/messages/interactive-startup.d.ts +103 -0
  247. package/dist/i18n/messages/interactive-startup.js +104 -0
  248. package/dist/i18n/messages/interactive-view.d.ts +140 -0
  249. package/dist/i18n/messages/interactive-view.js +141 -0
  250. package/dist/i18n/messages/interactive.d.ts +493 -0
  251. package/dist/i18n/messages/interactive.js +248 -0
  252. package/dist/i18n/messages/memory.d.ts +123 -0
  253. package/dist/i18n/messages/memory.js +128 -0
  254. package/dist/i18n/messages/panels.d.ts +190 -0
  255. package/dist/i18n/messages/panels.js +189 -0
  256. package/dist/i18n/messages/permissions.d.ts +126 -0
  257. package/dist/i18n/messages/permissions.js +151 -0
  258. package/dist/i18n/messages/plan.d.ts +123 -0
  259. package/dist/i18n/messages/plan.js +124 -0
  260. package/dist/i18n/messages/print.d.ts +86 -0
  261. package/dist/i18n/messages/print.js +91 -0
  262. package/dist/i18n/messages/report.d.ts +391 -0
  263. package/dist/i18n/messages/report.js +490 -0
  264. package/dist/i18n/messages/rewind.d.ts +198 -0
  265. package/dist/i18n/messages/rewind.js +217 -0
  266. package/dist/i18n/messages/session.d.ts +230 -0
  267. package/dist/i18n/messages/session.js +258 -0
  268. package/dist/i18n/messages/settings.d.ts +348 -0
  269. package/dist/i18n/messages/settings.js +336 -0
  270. package/dist/i18n/messages/subcommands-config.d.ts +111 -0
  271. package/dist/i18n/messages/subcommands-config.js +127 -0
  272. package/dist/i18n/messages/subcommands.d.ts +548 -0
  273. package/dist/i18n/messages/subcommands.js +474 -0
  274. package/dist/i18n/messages/trace.d.ts +243 -0
  275. package/dist/i18n/messages/trace.js +238 -0
  276. package/dist/i18n/types.d.ts +15 -0
  277. package/dist/i18n/types.js +7 -0
  278. package/dist/index.d.ts +6 -0
  279. package/dist/index.js +2 -0
  280. package/dist/memory/edit.d.ts +24 -0
  281. package/dist/memory/edit.js +63 -0
  282. package/dist/memory/frontmatter.d.ts +37 -0
  283. package/dist/memory/frontmatter.js +79 -0
  284. package/dist/memory/index.d.ts +28 -0
  285. package/dist/memory/index.js +126 -0
  286. package/dist/memory/lock.d.ts +17 -0
  287. package/dist/memory/lock.js +62 -0
  288. package/dist/memory/paths.d.ts +57 -0
  289. package/dist/memory/paths.js +175 -0
  290. package/dist/memory/report.d.ts +36 -0
  291. package/dist/memory/report.js +86 -0
  292. package/dist/memory/runtime.d.ts +42 -0
  293. package/dist/memory/runtime.js +36 -0
  294. package/dist/memory/secrets.d.ts +8 -0
  295. package/dist/memory/secrets.js +27 -0
  296. package/dist/memory/section.d.ts +28 -0
  297. package/dist/memory/section.js +59 -0
  298. package/dist/memory/store.d.ts +70 -0
  299. package/dist/memory/store.js +284 -0
  300. package/dist/memory/tool.d.ts +31 -0
  301. package/dist/memory/tool.js +94 -0
  302. package/dist/modes/acp/acp-events.js +2 -1
  303. package/dist/modes/acp/acp-server.js +13 -12
  304. package/dist/modes/commands-core.d.ts +17 -0
  305. package/dist/modes/commands-core.js +101 -61
  306. package/dist/modes/image-input.js +20 -6
  307. package/dist/modes/interactive/agent-bar.d.ts +95 -0
  308. package/dist/modes/interactive/agent-bar.js +248 -0
  309. package/dist/modes/interactive/agent-panels.js +27 -26
  310. package/dist/modes/interactive/agent-transcript.d.ts +47 -0
  311. package/dist/modes/interactive/agent-transcript.js +191 -0
  312. package/dist/modes/interactive/agent-ui.d.ts +36 -1
  313. package/dist/modes/interactive/agent-ui.js +180 -14
  314. package/dist/modes/interactive/agent-view.d.ts +72 -0
  315. package/dist/modes/interactive/agent-view.js +281 -0
  316. package/dist/modes/interactive/approval-dialog.js +46 -34
  317. package/dist/modes/interactive/clipboard-paste.d.ts +2 -2
  318. package/dist/modes/interactive/clipboard-paste.js +9 -4
  319. package/dist/modes/interactive/commands.d.ts +22 -1
  320. package/dist/modes/interactive/commands.js +93 -38
  321. package/dist/modes/interactive/completion.js +2 -1
  322. package/dist/modes/interactive/config-panel.d.ts +70 -0
  323. package/dist/modes/interactive/config-panel.js +319 -0
  324. package/dist/modes/interactive/config-ui.d.ts +96 -0
  325. package/dist/modes/interactive/config-ui.js +397 -0
  326. package/dist/modes/interactive/confirm-dialog.d.ts +49 -0
  327. package/dist/modes/interactive/confirm-dialog.js +120 -0
  328. package/dist/modes/interactive/event-notices.js +11 -9
  329. package/dist/modes/interactive/interactive-mode.d.ts +1 -1
  330. package/dist/modes/interactive/interactive-mode.js +25 -22
  331. package/dist/modes/interactive/key-dispatch.d.ts +18 -0
  332. package/dist/modes/interactive/key-dispatch.js +39 -13
  333. package/dist/modes/interactive/line/line-editor.js +2 -1
  334. package/dist/modes/interactive/line/line-mode.d.ts +2 -0
  335. package/dist/modes/interactive/line/line-mode.js +24 -4
  336. package/dist/modes/interactive/line/line-render.js +25 -24
  337. package/dist/modes/interactive/memory-panel.d.ts +64 -0
  338. package/dist/modes/interactive/memory-panel.js +221 -0
  339. package/dist/modes/interactive/message-view.d.ts +4 -2
  340. package/dist/modes/interactive/message-view.js +42 -33
  341. package/dist/modes/interactive/panels.js +39 -31
  342. package/dist/modes/interactive/pickers.d.ts +2 -1
  343. package/dist/modes/interactive/pickers.js +10 -6
  344. package/dist/modes/interactive/plan-command.d.ts +1 -1
  345. package/dist/modes/interactive/plan-command.js +24 -27
  346. package/dist/modes/interactive/plan-dialog.js +18 -17
  347. package/dist/modes/interactive/plan-flow.js +6 -5
  348. package/dist/modes/interactive/rewind-command.d.ts +1 -1
  349. package/dist/modes/interactive/rewind-command.js +18 -16
  350. package/dist/modes/interactive/rewind-flow.d.ts +0 -1
  351. package/dist/modes/interactive/rewind-flow.js +13 -14
  352. package/dist/modes/interactive/rewind-list.js +7 -6
  353. package/dist/modes/interactive/rewind-panel.js +40 -40
  354. package/dist/modes/interactive/rewind-text.d.ts +5 -3
  355. package/dist/modes/interactive/rewind-text.js +48 -38
  356. package/dist/modes/interactive/run-indicator.js +15 -12
  357. package/dist/modes/interactive/session-events.js +3 -2
  358. package/dist/modes/interactive/startup-header.js +22 -23
  359. package/dist/modes/interactive/startup-ui.js +41 -29
  360. package/dist/modes/interactive/status-area.d.ts +2 -0
  361. package/dist/modes/interactive/status-area.js +8 -2
  362. package/dist/modes/interactive/status-bar.js +4 -3
  363. package/dist/modes/interactive/subagent-view.js +9 -7
  364. package/dist/modes/interactive/tasks-report.d.ts +6 -0
  365. package/dist/modes/interactive/tasks-report.js +28 -31
  366. package/dist/modes/interactive/terminal-setup.d.ts +9 -0
  367. package/dist/modes/interactive/terminal-setup.js +23 -0
  368. package/dist/modes/interactive/tool-summary.js +19 -17
  369. package/dist/modes/interactive/tool-view.js +7 -4
  370. package/dist/modes/interactive/trace-view.d.ts +89 -0
  371. package/dist/modes/interactive/trace-view.js +332 -0
  372. package/dist/modes/print/print-mode.js +14 -15
  373. package/dist/modes/rpc/commands.js +7 -3
  374. package/dist/modes/rpc/rpc-mode.js +9 -3
  375. package/dist/modes/session-report.d.ts +2 -0
  376. package/dist/modes/session-report.js +101 -105
  377. package/dist/modes/startup-ui-text.js +11 -6
  378. package/dist/permissions/bypass.d.ts +30 -0
  379. package/dist/permissions/bypass.js +45 -0
  380. package/dist/permissions/memory-class.d.ts +23 -0
  381. package/dist/permissions/memory-class.js +54 -0
  382. package/dist/permissions/modes.d.ts +3 -3
  383. package/dist/permissions/modes.js +21 -16
  384. package/dist/permissions/pipeline.d.ts +1 -1
  385. package/dist/permissions/pipeline.js +9 -2
  386. package/dist/permissions/preview.js +49 -42
  387. package/dist/permissions/rules.js +5 -2
  388. package/dist/permissions/types.d.ts +4 -0
  389. package/dist/plan/store.js +2 -1
  390. package/dist/rpc.d.ts +39 -1
  391. package/dist/rpc.js +3 -0
  392. package/dist/sandbox/bash.js +11 -8
  393. package/dist/sandbox/detect.js +14 -11
  394. package/dist/sdk.d.ts +17 -1
  395. package/dist/sdk.js +29 -7
  396. package/dist/session/export.js +31 -30
  397. package/dist/session/manager.d.ts +1 -1
  398. package/dist/session/manager.js +5 -3
  399. package/dist/session/reuse.js +5 -4
  400. package/dist/session/scan.js +3 -2
  401. package/dist/session/stats-aggregate.d.ts +3 -1
  402. package/dist/session/stats-aggregate.js +5 -1
  403. package/dist/session/stats-index.js +2 -1
  404. package/dist/session/stats-scan.d.ts +4 -1
  405. package/dist/session/stats-scan.js +6 -2
  406. package/dist/session/store.d.ts +11 -0
  407. package/dist/session/store.js +48 -1
  408. package/dist/session/types.d.ts +2 -0
  409. package/dist/skills/builtin.js +2 -1
  410. package/dist/tools/image-file.d.ts +26 -1
  411. package/dist/tools/image-file.js +55 -13
  412. package/dist/tools/presets.d.ts +14 -1
  413. package/dist/tools/presets.js +3 -3
  414. package/dist/tools/registry.js +1 -1
  415. package/dist/tools/truncate.d.ts +1 -1
  416. package/dist/tools/truncate.js +2 -2
  417. package/dist/tools/types.d.ts +17 -1
  418. package/dist/trace/build-index.d.ts +91 -0
  419. package/dist/trace/build-index.js +127 -0
  420. package/dist/trace/build-nodes.d.ts +24 -0
  421. package/dist/trace/build-nodes.js +279 -0
  422. package/dist/trace/build-util.d.ts +43 -0
  423. package/dist/trace/build-util.js +120 -0
  424. package/dist/trace/build.d.ts +82 -0
  425. package/dist/trace/build.js +382 -0
  426. package/dist/trace/detail.d.ts +35 -0
  427. package/dist/trace/detail.js +205 -0
  428. package/dist/trace/flatten.d.ts +69 -0
  429. package/dist/trace/flatten.js +173 -0
  430. package/dist/trace/format.d.ts +47 -0
  431. package/dist/trace/format.js +329 -0
  432. package/dist/trace/html-template.d.ts +19 -0
  433. package/dist/trace/html-template.js +150 -0
  434. package/dist/trace/html.d.ts +104 -0
  435. package/dist/trace/html.js +292 -0
  436. package/dist/trace/preview.d.ts +29 -0
  437. package/dist/trace/preview.js +59 -0
  438. package/dist/trace/query-session.d.ts +14 -0
  439. package/dist/trace/query-session.js +38 -0
  440. package/dist/trace/query.d.ts +55 -0
  441. package/dist/trace/query.js +188 -0
  442. package/dist/trace/session.d.ts +40 -0
  443. package/dist/trace/session.js +168 -0
  444. package/dist/trace/types.d.ts +243 -0
  445. package/dist/trace/types.js +13 -0
  446. package/dist/tui/components/editor-paste.d.ts +2 -0
  447. package/dist/tui/components/editor-paste.js +6 -3
  448. package/dist/tui/components/editor.js +2 -2
  449. package/dist/tui/components/loader.d.ts +3 -0
  450. package/dist/tui/components/loader.js +14 -0
  451. package/dist/tui/components/settings-list.d.ts +61 -0
  452. package/dist/tui/components/settings-list.js +135 -0
  453. package/dist/tui/keybindings.d.ts +6 -1
  454. package/dist/tui/keybindings.js +18 -6
  455. package/dist/tui.d.ts +1 -0
  456. package/dist/tui.js +2 -0
  457. package/docs/en/host-api.md +167 -0
  458. package/docs/en/permissions.md +214 -0
  459. package/docs/en/providers.md +625 -0
  460. package/docs/en/rpc.md +361 -0
  461. package/docs/en/sessions.md +219 -0
  462. package/docs/en/tui.md +527 -0
  463. package/docs/host-api.md +18 -17
  464. package/docs/memory.md +172 -0
  465. package/docs/permissions.md +3 -2
  466. package/docs/providers.md +67 -1
  467. package/docs/rpc.md +31 -3
  468. package/docs/session-format.md +8 -2
  469. package/docs/sessions.md +29 -1
  470. package/docs/tui.md +189 -21
  471. package/package.json +8 -3
@@ -0,0 +1,625 @@
1
+ # Providers and models
2
+
3
+ English · [简体中文](../providers.md)
4
+
5
+ > Translated from the Chinese [docs/providers.md](../providers.md) as of commit `e3bde2e`. When the two differ, the
6
+ > Chinese version is authoritative.
7
+
8
+ Built-in providers, model references, API keys, custom providers and relays, the compat switches of each protocol, and caching. The design rationale is in [design.md](../design.md) §3 and §9.1 (Chinese).
9
+
10
+ ## Config directory
11
+
12
+ Default `~/.config/ama/` (Windows `%APPDATA%\ama\`; `AMA_CONFIG_DIR` wins, then `XDG_CONFIG_HOME/ama`). Data (sessions, the models.dev refresh override) lives in another directory: `~/.local/share/ama/` (`AMA_DATA_DIR` / `XDG_DATA_HOME`).
13
+
14
+ ```
15
+ ~/.config/ama/ 0700
16
+ ├── config.json user-level config: providers, channels, models, permissions, tools, cache… (edit directly)
17
+ ├── config.schema.json JSON Schema of config.json (editor completion and validation; generated by ama, rewritten)
18
+ ├── auth.json API keys (0600; created only by ama auth set / ama providers add)
19
+ ├── config.json.bak backup made before ama rewrites config.json
20
+ ├── hooks.json / trust.json / keybindings.json (as needed)
21
+ └── skills/ prompts/ (as needed)
22
+ ~/.local/share/ama/
23
+ ├── sessions/ sessions
24
+ └── models-dev.json models.dev refresh override (written only by ama models refresh)
25
+ ```
26
+
27
+ - `ama init`: creates the directory (0700) and fills in a missing `config.json` and `config.schema.json`, printing "created" or "already exists, left unchanged" for each. An existing `config.json` is never overwritten (not even with `--force`); `config.schema.json` is not a user file and every `init` rewrites it for the current version (and the current interface language); an empty `auth.json` is never created.
28
+ - **Automatic initialization on first run**: commands that enter a conversation (interactive, `-p`, `--mode rpc`) and `ama providers add` silently create the directory with a minimal `config.json` and schema when the config directory does not exist (`AMA_NO_INIT=1` disables this; the SDK and tests never trigger it). Read-only subcommands (`config show` / `path`, `doctor`, `models list`, `providers list`, `auth list`, `sessions` …) never create or rewrite the config directory.
29
+ - The minimal `config.json`:
30
+
31
+ ```json
32
+ {
33
+ "$schema": "./config.schema.json",
34
+ "version": 1,
35
+ "providers": {}
36
+ }
37
+ ```
38
+
39
+ No defaults are written (so default changes apply to old configs too, and `ama config show` shows the source as default), and no `defaultModel` either (see "Default model" below). `ama init` ends by printing next steps (`ama auth set` / `ama providers add` / `ama doctor`). Every key in `config.schema.json` carries a description and default (keys decided at runtime only describe the rule), visible on hover in editors.
40
+
41
+ - `ama config path`: prints the config directory, the data directory and each file path (marking whether it exists); `ama config edit`: opens `config.json` with `$VISUAL` / `$EDITOR` (running `init` first if it does not exist), printing the path when no editor is available.
42
+
43
+ Example: a three-channel relay + an image model + an override of a built-in provider.
44
+
45
+ ```json
46
+ {
47
+ "$schema": "./config.schema.json",
48
+ "version": 1,
49
+ "defaultModel": "packy/kimi-k2.5",
50
+ "providers": {
51
+ "packy": {
52
+ "apiKey": "$PACKY_API_KEY",
53
+ "channels": {
54
+ "chat": { "api": "openai-completions", "baseUrl": "https://www.packyapi.com/v1" },
55
+ "responses": { "api": "openai-responses", "baseUrl": "https://www.packyapi.com/v1" },
56
+ "messages": { "api": "anthropic-messages", "baseUrl": "https://www.packyapi.com" }
57
+ },
58
+ "models": [
59
+ { "id": "kimi-k2.5", "channels": ["chat", "messages"] },
60
+ { "id": "grok-4.7", "channels": ["responses"] },
61
+ {
62
+ "id": "qwen3-vl-flash",
63
+ "input": ["text", "image"],
64
+ "modelsDev": "llmgateway/qwen3-vl-flash"
65
+ }
66
+ ]
67
+ },
68
+ "deepseek": { "modelOverrides": [{ "id": "deepseek-flash", "contextWindow": 131072 }] }
69
+ }
70
+ }
71
+ ```
72
+
73
+ ## Built-in providers
74
+
75
+ | id | Channels (**bold** is the default) | Default address | API key environment variables (in order) |
76
+ | ------------ | ------------------------------------------------------------------------------------- | -------------------------------------------------- | ------------------------------------------------------------- |
77
+ | `anthropic` | single channel messages | `https://api.anthropic.com` | `ANTHROPIC_API_KEY`, `AMA_API_KEY_ANTHROPIC` |
78
+ | `openai` | **responses**, chat | `https://api.openai.com/v1` | `OPENAI_API_KEY`, `AMA_API_KEY_OPENAI` |
79
+ | `google` | single channel gemini | `https://generativelanguage.googleapis.com/v1beta` | `GEMINI_API_KEY`, `GOOGLE_API_KEY`, `AMA_API_KEY_GOOGLE` |
80
+ | `deepseek` | **chat**, messages (`/anthropic`) | `https://api.deepseek.com` | `DEEPSEEK_API_KEY`, `AMA_API_KEY_DEEPSEEK` |
81
+ | `moonshot` | **chat**, messages, responses, chat-intl, messages-intl (`.ai`) | `https://api.moonshot.cn/v1` | `MOONSHOT_API_KEY`, `KIMI_API_KEY`, `AMA_API_KEY_MOONSHOT` |
82
+ | `zhipu` | **chat**, messages (`/api/anthropic`), chat-intl, messages-intl (`api.z.ai`) | `https://open.bigmodel.cn/api/paas/v4` | `ZHIPU_API_KEY`, `ZAI_API_KEY`, `AMA_API_KEY_ZHIPU` |
83
+ | `dashscope` | **messages** (`/apps/anthropic`), responses, chat, messages-intl, chat-intl | `https://dashscope.aliyuncs.com/apps/anthropic` | `DASHSCOPE_API_KEY`, `QWEN_API_KEY`, `AMA_API_KEY_DASHSCOPE` |
84
+ | `openrouter` | single channel chat | `https://openrouter.ai/api/v1` | `OPENROUTER_API_KEY`, `AMA_API_KEY_OPENROUTER` |
85
+ | `groq` | single channel chat | `https://api.groq.com/openai/v1` | `GROQ_API_KEY`, `AMA_API_KEY_GROQ` |
86
+ | `xai` | **responses**, chat | `https://api.x.ai/v1` | `XAI_API_KEY`, `AMA_API_KEY_XAI` |
87
+ | `mistral` | single channel chat | `https://api.mistral.ai/v1` | `MISTRAL_API_KEY`, `AMA_API_KEY_MISTRAL` |
88
+ | `minimax` | **messages**, responses (M3 only), chat, messages-intl, chat-intl (`api.minimax.io`) | `https://api.minimax.cn/anthropic` | `MINIMAX_API_KEY`, `AMA_API_KEY_MINIMAX` |
89
+ | `stepfun` | **messages**, chat, responses (step-5-preview only), messages-intl, chat-intl (`.ai`) | `https://api.stepfun.com` | `STEPFUN_API_KEY`, `STEP_API_KEY`, `AMA_API_KEY_STEPFUN` |
90
+ | `volcengine` | **responses**, chat | `https://ark.cn-beijing.volces.com/api/v3` | `ARK_API_KEY`, `VOLCENGINE_API_KEY`, `AMA_API_KEY_VOLCENGINE` |
91
+ | `tencent` | **messages**, chat (`/v1`) | `https://tokenhub.tencentmaas.com` | `TOKENHUB_API_KEY`, `HUNYUAN_API_KEY`, `AMA_API_KEY_TENCENT` |
92
+ | `chatgpt` | **siwc**, codex (default follows the signed-in flavor, see "ChatGPT login") | `https://api.openai.com/v1` | none (`ama auth login chatgpt`) |
93
+ | `ollama` | single channel chat | `http://127.0.0.1:11434/v1` | optional (`OLLAMA_API_KEY`) |
94
+ | `lmstudio` | single channel chat | `http://127.0.0.1:1234/v1` | optional |
95
+
96
+ There is also the test provider `fake` (models `fake/echo`, `fake/reasoning`), see below.
97
+
98
+ **Default protocol**: Messages / Responses first, Chat as fallback. Messages endpoints that honor explicit caching (Qwen, MiniMax M2.x, Tencent) and StepFun, which officially recommends Messages, default to messages; OpenAI, xAI and Volcengine Ark default to responses; DeepSeek, Zhipu and Kimi only have implicit caching and **stay on chat by default until their direct official endpoints pass the measurement gate** (see "Channel measurements"); switching is one `defaultChannel` line in `builtin.ts`.
99
+
100
+ **Built-in channels** (the same mechanism as user [channels](#channels-one-provider-several-endpoints)):
101
+
102
+ - `provider/model@channel` picks a channel, e.g. `deepseek/deepseek-v4-pro@messages`, `dashscope/qwen3.8-max@chat-intl`; catalog models mount every channel by default, and the catalog can restrict per model (MiniMax M2.x and StepFun step-3.x have no responses).
103
+ - Changing the default: `"providers": { "deepseek": { "defaultChannel": "messages" } }`; or just write the protocol `"api": "anthropic-messages"` to pick the built-in channel with that protocol.
104
+ - Changing one channel: `channels.<built-in channel name>` with only the fields to change (`headers` / `compat` merge one level deep, the rest overrides); channels with new names are appended after them, and catalog models can use them too.
105
+ - **When the provider-level `baseUrl` is changed** (config, auth.json or `OPENAI_BASE_URL` / `ANTHROPIC_BASE_URL`) without writing your own `channels`, the built-in channels are dropped and it is handled as a single channel of "provider-level `api` + the new `baseUrl`", as before built-in channels existed: the fallback protocol is Chat (`api` can change it); **catalog models** of OpenAI / xAI / Volcengine Ark still use Responses on a relay, while ids outside the catalog (local services, the relay's own models) use Chat. A baseUrl on any built-in channel host (such as the international site) does not count as a relay.
106
+ - Auth headers per channel: Anthropic-compatible endpoints default to `x-api-key`; the Messages channels of Kimi, MiniMax and StepFun use `Authorization: Bearer` per their docs.
107
+
108
+ ### Coding Plan style subscription endpoints
109
+
110
+ The Coding Plan / Token Plan quotas of Volcengine Ark, Alibaba Bailian, Zhipu, Kimi, MiniMax, StepFun and others are restricted to coding tools; `ama -p` and RPC embedding are automated calls and risk being judged violations, so **they are not built-in channels**. Once you have confirmed the terms allow it, add a channel yourself (with the key on the channel):
111
+
112
+ ```json
113
+ {
114
+ "providers": {
115
+ "volcengine": {
116
+ "channels": {
117
+ "coding": {
118
+ "api": "anthropic-messages",
119
+ "baseUrl": "https://ark.cn-beijing.volces.com/api/coding",
120
+ "apiKey": "$ARK_CODING_PLAN_KEY"
121
+ }
122
+ }
123
+ },
124
+ "zhipu": {
125
+ "channels": {
126
+ "coding": {
127
+ "api": "openai-completions",
128
+ "baseUrl": "https://open.bigmodel.cn/api/coding/paas/v4",
129
+ "apiKey": "$ZHIPU_CODING_PLAN_KEY"
130
+ }
131
+ }
132
+ },
133
+ "stepfun": {
134
+ "channels": {
135
+ "plan": {
136
+ "api": "anthropic-messages",
137
+ "baseUrl": "https://api.stepfun.com/step_plan",
138
+ "apiKey": "$STEP_PLAN_KEY",
139
+ "authHeader": "authorization-bearer"
140
+ }
141
+ }
142
+ }
143
+ }
144
+ }
145
+ ```
146
+
147
+ Then `--model volcengine/<model>@coding` (model ids of subscription endpoints follow each vendor's docs; add ids outside the catalog with `models[]` and `"channels": ["coding"]`).
148
+
149
+ ## ChatGPT login
150
+
151
+ Drive ama with your own ChatGPT Plus / Pro subscription (built-in provider `chatgpt`, protocol openai-responses). **For your own personal use only**: do not let one login serve several people; embedding hosts (Armadra) must not turn it into a multi-user or hosted service either.
152
+
153
+ ```bash
154
+ ama auth login chatgpt # default: official Sign in with ChatGPT (siwc), browser authorization
155
+ ama auth login chatgpt --paste # no browser (SSH / host): open the printed URL and paste the callback URL back
156
+ ama auth status # flavor, plan, masked email, token lifetime; codex also shows quota
157
+ ama --model chatgpt/<model> # models available to the account: ama models discover chatgpt
158
+ ama auth logout chatgpt # siwc revokes the refresh token first, then deletes it locally
159
+ ```
160
+
161
+ **Two paths** (`--flavor` or the user-level config `auth.chatgpt.flavor`; siwc by default):
162
+
163
+ | | `siwc` (default, official) | `codex` (opt-in fallback) |
164
+ | ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
165
+ | Login | OpenAI's official dynamic registration: registers as `dynamic_agent_client` the first time and stores the issued client id in the entry; `agent_name_hint=ama`; the install id is kept in `<dataDir>/chatgpt-host.json` | Borrows the Codex CLI's public client; the first time it asks in the terminal to confirm "unofficial use, personal only, may change at any time" (non-TTY needs `--yes`) |
166
+ | Verification | id_token signature checked against JWKS plus iss / aud / nonce / exp; the granted scope must include `chatgpt.tokens.use.direct` | Only decodes the id_token for the account id and plan |
167
+ | Callback | `http://127.0.0.1:1455/auth/callback`, any free port when 1455 is taken | 1455 → 1457, never takes over a busy port (suggests `--paste` / `--device`) |
168
+ | Device code | none (use `--paste`) | `--device` (may need enabling under ChatGPT → Settings → Security) |
169
+ | Inference endpoint | channel `siwc`: `https://api.openai.com/v1` | channel `codex`: `https://chatgpt.com/backend-api/codex`, also sending `ChatGPT-Account-ID` and `originator` (default `codex_cli_rs`, changeable via `auth.chatgpt.originator`) |
170
+ | Quota | only known when exceeded (429); set a weekly cap for ama under ChatGPT → Settings → Usage → App limits | response headers, `codex.rate_limits` events, `ama auth status` queries `wham/usage` |
171
+ | Logout | calls `revocation_endpoint`, then deletes locally | deletes locally only |
172
+
173
+ The default channel of `chatgpt` is chosen at assembly time from the flavor of the auth.json entry; `provider/model@siwc|codex` can name it explicitly but must match the signed-in flavor (otherwise `chatgpt_flavor_mismatch`).
174
+
175
+ **Credentials**: a `{ "type": "oauth", … }` entry in auth.json (file mode 0600); `ama auth list` only shows `oauth · <flavor> · <plan>`. The access token is refreshed automatically when less than 5 minutes remain or a request returns 401; several ama processes (several nodes on the canvas) share one auth.json, and refreshes are serialized through `auth.json.lock` (re-read after taking the lock; if another process already refreshed, its token is used), so refresh-token rotation never knocks another process out. When refreshing fails permanently the entry is marked `needsLogin` (tokens are not deleted), requests report `auth_expired`, and you sign in again with `ama auth login chatgpt` as prompted. Raw tokens, codes and id_tokens never reach logs, sessions, events or errors. Logins of other applications (Codex CLI and others) are never read or imported.
176
+
177
+ **Request body**: both backends force `store:false`, `stream:true` and an `input` array, and drop unsupported fields (siwc: 15 including `max_output_tokens`, `temperature`, `top_p`, `metadata`, `user`, `truncation`, `prompt_cache_retention`; codex: `max_output_tokens`, `temperature`, `top_p`, `prompt_cache_retention`, `prompt_cache_options`). Both send `prompt_cache_key = session id`, and `cacheRetention: "long"` is lowered to short automatically. If siwc rejects a field with `subscription_sharing_unsupported_capability`, it is dropped and the request retried once (remembered for the process); if codex reports `Instructions are not valid`, this session moves the system prompt into a leading developer message (the prefix stays stable). With compat `toolsInNamespace: true` tools go into an `additional_tools` input item (shape still to be verified with a real account).
178
+
179
+ **Error codes** (error messages start with the code; hosts decide by code):
180
+
181
+ | Code | Source | Handling |
182
+ | ---------------- | ------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------- |
183
+ | `quota_exceeded` | siwc 429 `subscription_sharing_usage_limit_exceeded`; codex 429 `usage_limit_reached` / `usage_not_included` | no retry; carries the reset time and emits `quota_update` |
184
+ | `auth_expired` | a 401 still failing after one refresh, a permanently failed refresh, an entry marked `needsLogin` | no retry; run `ama auth login chatgpt` again |
185
+ | `not_eligible` | siwc 403 `subscription_sharing_user_not_eligible` | no retry, no re-login |
186
+ | (as is) | 503 and similar | the session layer's existing backoff retries |
187
+
188
+ **Usage**: subscription requests record `usage.cost = 0` with `billing: "subscription"`; `/session` lists "subscription usage" separately (requests, tokens, cache hit rate, no USD conversion) together with the latest quota; the `quota_update` event is forwarded as is over RPC and has the same name among host events.
189
+
190
+ **Overrides** (for tests or a future own client): `auth.chatgpt.clientId` / `issuer` / `originator` / `redirectPorts` (user level and profile only), environment variables `AMA_CHATGPT_CLIENT_ID`, `AMA_CHATGPT_ISSUER`, `AMA_CHATGPT_BASE_URL` (changes the address of the current flavor's channel).
191
+
192
+ **Embedding hosts**: with a profile ama never starts an interactive login; when `chatgpt` is used and the login has expired the request reports `auth_expired`, and the host guides the user to run `ama auth login chatgpt --paste` in a terminal. Hosts never read, store or forward tokens; they only consume `quota_update` and `auth_expired` / `quota_exceeded`.
193
+
194
+ **Real-account check** (not run in CI): sign in first, then `AMA_E2E_CHATGPT=1 pnpm vitest run src/auth/chatgpt/chatgpt.e2e.test.ts` (optionally `AMA_E2E_CHATGPT_MODEL=<slug>`); it prints the flavor, the model list, the stop reason and the quota, never tokens.
195
+
196
+ ## Model references
197
+
198
+ `provider/model-id`, e.g. `deepseek/deepseek-v4-pro`, `openrouter/anthropic/claude-sonnet-5.5`. Without a provider prefix, the id must match uniquely across all catalogs; when several providers share the name, only providers with a configured key are considered, and if it is still ambiguous an error lists the candidates. Providers with an empty model table (ollama, lmstudio, custom providers without `models`) and built-in providers whose baseUrl points to a non-official host (see "Connecting relays") accept any model id.
199
+
200
+ ## API key lookup order
201
+
202
+ 1. `--api-key` (only for the provider given by `--model`)
203
+ 2. `--auth-file` / the profile's `authFile`
204
+ 3. `~/.config/ama/auth.json` (a warning when its permission is not 0600, but it is still used)
205
+ 4. `providers.<id>.apiKey` in `config.json`: supports `$ENV`, `${ENV}`, `!command`; `$$` means a literal `$`
206
+ 5. Environment variables (in the order of the table above; skipped with the profile's `authEnv: false`)
207
+ 6. Local services (`requiresApiKey: false`) work without a key
208
+
209
+ An `apiKey` in `auth.json` starting with `!` runs a command for the value (10-second timeout; empty output or a non-zero exit counts as not configured).
210
+
211
+ ## Custom providers
212
+
213
+ ```json
214
+ {
215
+ "providers": {
216
+ "my-proxy": {
217
+ "baseUrl": "https://proxy.example/v1",
218
+ "apiKey": "$MY_PROXY_KEY",
219
+ "models": [{ "id": "gpt-x", "contextWindow": 128000, "maxTokens": 16384, "reasoning": true }],
220
+ "compat": { "maxTokensField": "max_tokens" }
221
+ },
222
+ "deepseek": { "modelOverrides": [{ "id": "deepseek-flash", "contextWindow": 131072 }] }
223
+ }
224
+ }
225
+ ```
226
+
227
+ - `api` defaults to `openai-completions`; custom models default to `maxTokens: 8192`, `reasoning: false`, `input: ["text"]`; `contextWindow` is not guessed (automatic compaction is off by default).
228
+ - `models[]` entries with the same id replace the whole model, new ids are appended; `modelOverrides[]` only changes metadata of existing models.
229
+
230
+ ## Connecting relays
231
+
232
+ Under one relay, different models often support different protocols (some all three, some only Chat and Messages, some only Responses). One provider is enough: put the protocol on the model.
233
+
234
+ ```sh
235
+ export PACKY_API_KEY=sk-...
236
+ ```
237
+
238
+ ```json
239
+ {
240
+ "providers": {
241
+ "packy": {
242
+ "baseUrl": "https://proxy.example/v1",
243
+ "apiKey": "$PACKY_API_KEY",
244
+ "models": [
245
+ { "id": "deepseek-v4-flash" },
246
+ { "id": "grok-4.7", "api": "openai-responses" },
247
+ { "id": "MiniMax-M2.7", "api": "anthropic-messages" }
248
+ ]
249
+ }
250
+ }
251
+ }
252
+ ```
253
+
254
+ - A model's `api` defaults to the provider's (here `openai-completions`); `modelOverrides[]` can change `api` too.
255
+ - The three protocols share one `baseUrl`: Completions / Responses append `/chat/completions`, `/responses`; Messages appends `/messages` when the baseUrl ends with `/v1`, otherwise `/v1/messages`.
256
+ - Don't want to write `models` by hand: `ama models discover packy` lists the relay's models (`GET {baseUrl}/models`); `--probe` tries minimal requests per model with the provider's protocol, completions, responses and messages in turn and records the first that succeeds (at most 3 per model; `--limit` caps the number of probed models, default 30; an estimate is printed first; models are probed concurrently with the same rules as `providers add --probe` below); `--write` merges the result into the user-level `config.json` (existing ids are not overwritten; only `id` and an `api` differing from the provider's are written; the original is backed up as `config.json.bak`). Context window and other metadata are filled from the models.dev snapshot at runtime (see "Model metadata" below); entries without a match have no `contextWindow`, so automatic compaction is off for them; add it by hand when needed.
257
+
258
+ ```sh
259
+ ama models discover packy --probe --write --limit 8
260
+ ama -p "hi" --model packy/grok-4.7
261
+ ```
262
+
263
+ Zero config: the built-in `openai` / `anthropic` recognize `OPENAI_BASE_URL` / `ANTHROPIC_BASE_URL` (the common convention of the OpenAI SDK and Claude Code), at lower priority than `baseUrl` in config and auth.json, and not read with the profile's `authEnv: false`. When the baseUrl is not an official host, model ids outside the catalog are accepted too, and compat uses conservative defaults (no `prompt_cache_key`). The "Providers" section of `ama config show` and `ama doctor` mark which variable the baseUrl came from; the default model picked by zero config comes from the official catalog and may not exist on the relay, so specify one with `--model` or `defaultModel`.
264
+
265
+ ### Default model
266
+
267
+ Without `--model`, a resumed session's model or `defaultModel`, ama takes the first provider in provider order (built-ins first, then custom providers from config) that has a key (or a reachable local service), then picks among its models:
268
+
269
+ - **Built-in providers**: the first catalog entry (catalogs are ordered by recommendation);
270
+ - **Custom providers** (relays, whose model table follows the upstream `/models` order): among models with a models.dev price (input price > 0), tool calling and a context window ≥ 64k, the one with the **lowest input price**; ties go to the larger context, then to the earlier position in the list; only when none qualifies does it fall back to the first entry.
271
+
272
+ When there is no `defaultModel` yet, `ama providers add` picks one with the same rules and writes it to `defaultModel` (after probing, only among models that passed), stating in the summary which one and why; an existing `defaultModel` is left alone. The "Model" line of `ama config show` and `ama doctor` explains the reason the same way. When no model is usable at all, startup lists the key environment variable names, `ama auth set` and `ama providers add`.
273
+
274
+ ```sh
275
+ OPENAI_BASE_URL=https://proxy.example/v1 OPENAI_API_KEY=$PACKY_API_KEY ama -p "hi" --model openai/qwen3.8-flash
276
+ ```
277
+
278
+ ### Channels: one provider, several endpoints
279
+
280
+ A relay often offers Chat Completions (`/v1/chat/completions`), Responses (`/v1/responses`) and Anthropic Messages (`/v1/messages`) at the same time, with each model available on only some of them. A **channel** is "protocol + address (+ optional key / headers / compat)"; a provider can have several channels, and each model declares which channels it is mounted on:
281
+
282
+ ```json
283
+ {
284
+ "providers": {
285
+ "packy": {
286
+ "name": "Packy",
287
+ "apiKey": "$PACKY_API_KEY",
288
+ "channels": {
289
+ "chat": { "api": "openai-completions", "baseUrl": "https://www.packyapi.com/v1" },
290
+ "responses": { "api": "openai-responses", "baseUrl": "https://www.packyapi.com/v1" },
291
+ "messages": { "api": "anthropic-messages", "baseUrl": "https://www.packyapi.com" }
292
+ },
293
+ "defaultChannel": "chat",
294
+ "models": [
295
+ { "id": "kimi-k2.5", "channels": ["chat", "messages"] },
296
+ { "id": "grok-4.7", "channels": ["responses"] },
297
+ { "id": "deepseek-v4-flash", "channels": ["chat", "responses", "messages"] },
298
+ { "id": "glm-5" }
299
+ ]
300
+ }
301
+ }
302
+ }
303
+ ```
304
+
305
+ - **Channel fields**: `api` and `baseUrl` are required; `apiKey` (same syntax as at provider level, `$ENV` / `!command` / literal; `"<provider>@<channel>"` entries in auth.json work too), `headers`, `compat` and `authHeader` are optional and inherit from the provider by default. Channel names match `[A-Za-z0-9][A-Za-z0-9_-]*`, without `/` or `@`.
306
+ - **Model mounting**: `models[].channels` lists the usable channels, the first being preferred; without it → `defaultChannel` (by default the first key of `channels`). Referencing a channel that does not exist → a config validation error with the path.
307
+ - **Model references**: `provider/model` uses the preferred channel; `provider/model@channel` picks one explicitly (the same in `--model`, `defaultModel`, `/model`, the SDK and RPC `set_model`). A channel not in the model's `channels` → an error listing the available `provider/model@channel`. When the part after `@` is not a channel name of that provider, the whole string is still treated as the model id (compatible with models whose id contains `@`).
308
+ - **Backward compatibility**: providers without `channels` (custom providers and single-protocol built-in providers; built-in channels of multi-protocol providers are described in "Built-in providers") are handled as a single channel: the provider-level `api` + `baseUrl` is the implicit `default` channel; model-level `api` / `baseUrl` still work and override the selected channel (equivalent to an anonymous channel). Existing configs need no changes. Once `channels` is written, the provider-level `api` / `baseUrl` no longer form a channel by themselves.
309
+ - **Runtime**: the selected model carries the channel's protocol, address, key, headers and compat; session records (`model_change`) and `ModelRef` carry the `channel`; the cache endpoint key (three states, misses, granularity inference) is `provider|host|model@channel`, so different channels of the same model are tracked separately. Prices and models.dev metadata are shared per model.
310
+
311
+ ### One-step setup: `ama providers`
312
+
313
+ With just a baseUrl and a key, one command creates the provider, lists the models and fills in metadata:
314
+
315
+ ```sh
316
+ ama providers add packy --base-url https://www.packyapi.com/v1 --key-env PACKY_API_KEY --probe --limit 8 --yes
317
+ ```
318
+
319
+ ```
320
+ ama providers add <id> --base-url <url> [--channel <name>=<api>@<baseUrl> …] [--api <api>|auto]
321
+ [--key-env <VAR>] [--probe] [--limit N] [--probe-models a,b,…]
322
+ [--max-requests N] [--concurrency N] [--probe-timeout ms]
323
+ [--prefer chat,responses,messages] [--include-no-tools] [--yes]
324
+ ama providers list
325
+ ama providers channels <id>
326
+ ama providers remove <id>
327
+ ama providers refresh <id> [--probe …]
328
+ ```
329
+
330
+ - **Key**: read from stdin by default (not echoed in a terminal, never on the command line or in shell history) and stored in `auth.json` (0600); with `--key-env VAR` stdin is not read and `config.json` gets `"apiKey": "$VAR"`. If the provider already has a key, nothing is asked.
331
+ - **Candidate channels**: with `--channel` (repeatable) only those are used; otherwise three are derived from `--base-url`: `chat` (openai-completions, baseUrl as is), `responses` (openai-responses, same), `messages` (anthropic-messages, the host root without a trailing `/v1`); `--api <api>` keeps only the matching one.
332
+ - **Model list**: `GET {baseUrl}/models` (the address of the first OpenAI-family channel; with only a Messages channel it is `/v1/models`, tried with `Authorization: Bearer` first and `x-api-key` on 401 / 403, since relays were measured to accept only the former). new-api style relays give `supported_endpoint_types` (`openai` / `openai-response` / `anthropic`) on entries, used to mount models on the matching channels; without hints a model is mounted on the first candidate channel; with hints but no supporting candidate channel (such as grok when only a messages channel is configured) → not written.
333
+ - **`--probe`**: for the selected models (those listed by `--probe-models`, by default the first `--limit` ids in case-insensitive alphabetical order, default 30) one minimal request is sent per channel (only the hinted channels when hints exist), and **every channel that passes is written into the model's `channels`**, ordered by `--prefer` (default chat, responses, messages); models where everything fails are not written; unprobed models are written per hints and marked "not probed" in the table. The request count and an estimated duration are printed first, and beyond `--max-requests` (default 60) the model count is cut.
334
+ - **Verdict**: a successful HTTP response whose stream yields a first content event (text / thinking / tool call) counts as usable and is disconnected immediately, without waiting for a reasoning model to finish thinking; when only the start of the stream arrived without content, it waits 1 more second and counts as usable if no error appears (so relays that answer 200 first and report errors inside the stream are not misjudged). HTTP errors, stream errors and timeouts (`--probe-timeout`, default 15000 ms) count as unusable, with the reason recorded.
335
+ - **Concurrency**: at most `--concurrency` "model × channel" requests in flight (1–16, default 6); results are printed in model and channel order, a terminal shows a single refreshing progress line, and the total time is printed at the end.
336
+ - **Rate limits**: 401 / 403 stop immediately; on 429 concurrency is halved and the request retried once after 2 s, stopping if it is still 429.
337
+ - **Channel pruning**: before writing, candidate channels with no mounted model are removed; `defaultChannel` is the first remaining one (by `--prefer`).
338
+ - **Writing**: `providers.<id>` in the user-level `config.json` (backed up as `config.json.bak` first). Model entries only get `id` and `channels`; context, output, images, reasoning and prices are **not written to the config** but filled from the models.dev snapshot at runtime (next section), so `refresh` never overwrites hand-edited fields and hand-written values always win. Models that models.dev marks as lacking tool calling are not written by default (an agent cannot work without tool calls); `--include-no-tools` writes them anyway. Running `add` again on an existing provider only appends new channels and new models; existing channel definitions and model entries are left byte-for-byte unchanged.
339
+ - **Confirmation**: a summary is printed before writing the config, with one y/N question in a terminal; non-TTY use requires `--yes` (otherwise exit 2).
340
+ - The printed table: id, channels, context, output, images, reasoning, tool calling, price (the vendor's price from models.dev, $/M input / output; the relay's actual price may differ), match method.
341
+ - `list`: every provider (those in config.json and built-ins with a key) → channels (protocol, address, key source: auth.json / `$VAR` / literal / none; the key is never shown) → model count. `channels <id>`: each channel's protocol, address and mounted model count. `remove`: deletes `providers.<id>` (with a backup) and the provider's entries in auth.json. `refresh`: re-fetches `/models` (models.dev from the local snapshot) and only appends new models (probing like add with `--probe`), leaving existing entries alone; ids removed upstream are only reported, never deleted.
342
+
343
+ Measured (2026-10-02, a relay offering all three endpoint types, 22 models, 29 requests in total):
344
+
345
+ - `providers add --probe --probe-models kimi-k2.5,grok-4.7,deepseek-v4-flash` took 8 requests (1 list + 7 probes): kimi-k2.5 → chat, messages (Responses failed); grok-4.7 → responses; deepseek-v4-flash returned `incomplete_details.reason: "length"` on Responses (the relay omits `max_output_tokens` when forwarding to DeepSeek), which is handled as output truncation, so all three endpoints work.
346
+ - models.dev matched 22 / 22: 19 vendor entries (including `qwen3.8-max-0902`, matched to `alibaba/qwen3.8-max` after removing the date suffix), 2 by majority (kimi-k2.5: the vendor entry `moonshotai/kimi-k2.5` is not in the database, so the majority of 11 entries with the same canonical id gives 262k / images; qwen3-coder-next), 1 unique entry (qwen3-vl-flash).
347
+ - `-p` worked through both the chat and `@messages` channels; with `--image`, a four-color square picture was answered correctly as red / green / blue / yellow on kimi-k2.5 (chat, messages) and qwen3-vl-flash; for glm-5 (text-only per models.dev) it exited with 2 without sending a request.
348
+ - A second provider with only a messages channel: 19 models mounted, the 3 Responses-only grok models not written; using a shared model name without a prefix reports the ambiguity and lists both providers.
349
+
350
+ ### Model metadata: models.dev
351
+
352
+ [models.dev](https://models.dev) aggregates model parameters from over two hundred providers (`https://models.dev/api.json`, about 5 MB, MIT licensed; the notice is in `THIRD_PARTY_NOTICES.md` at the repository root). ama **ships a trimmed snapshot** for the major vendors and uses it to fill in numeric facts such as context window, output limit, input modalities, reasoning and prices for built-in catalogs and custom models. **No network at startup or runtime.**
353
+
354
+ - **Snapshot**: `src/ai/providers/models-dev/<provider>.json`, one file per vendor, with the list in `_providers.json` in the same directory (22: anthropic, openai, google, xai, deepseek, moonshotai / moonshotai-cn, zhipuai / zai, alibaba / alibaba-cn, minimax / minimax-cn, stepfun / stepfun-ai, volcengine, tencent-tokenhub, mistral, meta, xiaomi, groq, plus openrouter models whose vendor prefix is on an allowlist). `scripts/update-models-dev.mjs` (zero dependencies) fetches, filters (dropping `deprecated` entries, entries whose output lacks text, a context of 0 and `tool_call: false`), trims fields (name, family, knowledge, release_date, reasoning, modalities.input, limit.context / input / output, cost including `context_over_200k` and per-context `tiers`, interleaved, beta status, canonical_model_id) and writes them sorted by key, then generates `models-dev-data.ts`, which is inlined into the bundle; unchanged content stays byte-identical (`fetchedAt` in `_meta.json` only changes with the sha256). To control size (inlined data ≤ 200 KB, guarded by a test), `last_updated` and `reasoning_options` are not included.
355
+ - **Weekly refresh**: `.github/workflows/models-dev.yml` runs the script every Monday at 03:17 UTC (and on manual `workflow_dispatch`), automatically removes catalog overrides identical to the new snapshot, and only opens / updates a PR on the fixed branch `chore/models-dev-refresh` when something changed (the summary lists additions, removals and context / output / price changes; `needs-review` is added when more than 30% is removed). The PR is created with the repository secret `MODELS_DEV_PR_TOKEN` (a GitHub App token or a fine-grained PAT with `contents` + `pull-requests` write access) so that CI triggers; without that secret the workflow runs `pnpm run ci` itself first, then opens the PR with the default token and states the result in the description. A person merges it.
356
+ - **`ama models refresh [--provider <id>[,<id>…]]`**: explicitly fetches the latest `api.json` over the network, trims it with the same list, writes `models-dev.json` into the data directory (default `~/.local/share/ama/`) and prints additions / upstream removals / price changes. From then on the index = the built-in snapshot ⊕ this override: the override applies only when its `fetchedAt` is later than the snapshot's (so an old override expires automatically once an ama upgrade brings a newer snapshot); for the same `provider/model` the override wins, and entries removed upstream keep the snapshot's values. `refresh-catalog` is the old name. `AMA_MODELS_DEV_URL` changes the refresh source. The full cache written by older ama versions (version 1) is no longer read. `ama providers add|refresh` and `ama models discover` only use the local index and no longer fetch models.dev (their `/models` requests are unchanged).
357
+ - **Built-in catalogs (`catalog/*.json`) only hold overrides**: a file-level `"modelsDev": "<snapshot provider id>"` (local services write `false`); entries inherit name, reasoning, contextWindow, maxTokens, input, cost, family, knowledge, releaseDate, inputLimit and status from the snapshot via `<modelsDev>/<id>`, and the catalog only writes ama-specific fields (`api`, `thinkingLevelMap`, `promptCache`, `compat`) plus values deliberately different from the snapshot (such as openai's 272k window or deepseek's prices); `cost` may list only the keys to change. When the id differs from the snapshot the entry writes `"modelsDev": "provider/model"`, and `false` to inherit nothing. Models missing from the snapshot (deprecated etc.) get complete entries. `catalog.test.ts` reports fields equal to the snapshot's value as redundant; `UPDATE_CATALOG=1 pnpm vitest run src/ai/providers/catalog.test.ts` removes them automatically and regenerates `catalog-data.ts` (run prettier afterwards); to pin a value equal to the snapshot on purpose, write `"_reason": "…"` on the entry (a comment ignored at runtime that exempts the entry from the redundancy check). The inherited mapping for catalogs differs slightly from custom models: `maxTokens` is not capped (it is `min(limit.output, contextWindow)`), and missing cache prices count as 0.
358
+ - **Priority**: user config (fields written in `models[]` / `modelOverrides[]`) > built-in catalog (snapshot ⊕ catalog overrides) > the models.dev index > custom defaults (`maxTokens: 8192`, `input: ["text"]`, `reasoning: false`, no guessed `contextWindow`). `ama models list` and `ama config show` mark where each field comes from (`config` / catalog / `models.dev` / default; fields a built-in catalog inherits from the snapshot are marked `models.dev`).
359
+ - **Field mapping** (custom models): `contextWindow = limit.context`; `maxTokens = min(limit.output, 65536, contextWindow)`. `maxTokens` is sent as `max_tokens` on every request, and models.dev gives the vendor's limit (quite a few models list 1M, equal to their context); relays that switched upstreams often reject very large values, the thinking budget of the Anthropic protocol is also derived from it, and 64k is plenty for a coding agent's single-turn output; write it in the config when you need more. `input` becomes `["text","image"]` or `["text"]` depending on whether `modalities.input` contains `image`; `reasoning`; `cost` takes `input` / `output` / `cache_read` / `cache_write` ($/M), using the input price when cache prices are missing (no discount assumed, so warming economics err on the conservative side); `cost.tiers` takes models.dev's per-context tiers, folding to a single 200k tier when only `context_over_200k` exists; plus `family`, `knowledge`, `releaseDate`, `inputLimit` (when it differs from the context) and `status: "beta"`.
360
+ - **Matching rules** (the same id often appears under several resellers with differing values):
361
+ 1. `"modelsDev": "provider/model"` on the model → use that entry directly (`false` disables enrichment);
362
+ 2. an id of the form `vendor/model` that exists exactly as a `provider/model` in the index → use it;
363
+ 3. find all entries with the same id case-insensitively; among those with a `canonical_model_id`, take the one pointing to a model named like the id (otherwise the most voted one); if that can be found under the vendor's own provider → use the vendor entry;
364
+ 4. otherwise prefer vendor providers among the entries (with the same canonical id): anthropic, openai, google, deepseek, moonshotai(-cn), zhipuai, zai, alibaba(-cn), xai, mistral, minimax(-cn), meta, llama, cohere, xiaomi, stepfun(-ai), volcengine, tencent-tokenhub and so on (models.dev has no provider id like `qwen`; Qwen lives under `alibaba`);
365
+ 5. still several entries → take the majority by (context, output, images), with a warning when values disagree; a single entry is used as is;
366
+ 6. when no entry has the same name, try normalized ids in turn: strip the `vendor/` prefix, strip suffixes like `:free`, strip `-latest`, strip date suffixes (`-0902`, `-20250514`, `-2025-05-14`);
367
+ 7. nothing found → "unmatched", keeping the custom defaults (no guessed `contextWindow`, automatic compaction off). The snapshot only covers major vendors, so reseller-specific ids may not match; write `modelsDev` on the model to point at an included entry.
368
+
369
+ ### Image input
370
+
371
+ - All four protocols put images into user messages: Chat Completions `image_url` (data URL), Responses `input_image`, Anthropic `image` (base64 source), Gemini `inlineData`; images in tool results are mapped the same way.
372
+ - Entry points: `ama -p "describe this picture" --image a.png --image b.jpg`; in the interactive and line UIs write `@image-path`, or paste / drop an image file path (words in the input ending in `.png` / `.jpg` / `.jpeg` / `.gif` / `.webp` whose file exists); in the interactive UI `Ctrl+V` or `/paste` saves the clipboard image to a file and inserts `@<path>` at the cursor (see "Clipboard images" below and [tui.md](tui.md) "Clipboard images"). Images from `--image`, `@image` and pasting are all resized per `images.resize`.
373
+ - MIME detection is shared with the `read` tool (PNG / JPEG / GIF / WebP recognized by file header; when the extension disagrees, the header wins), as is the size limit. The limit is measured **after base64** (`ceil(bytes/3)*4`) and tiered by the current model's endpoint: official Anthropic (`api.anthropic.com`) 10 MB, official Gemini and OpenAI 20 MB, relays (built-in providers with a changed baseUrl) and other hosts 5 MB; either side over 8000 px is refused.
374
+ - Resizing (`images.resize`, default `auto`): when over the limit, `sips` (macOS) and then `magick` / `convert` (ImageMagick) are tried to shrink it within the limit before attaching; without a tool, or with `off`, it is refused per the rules above with a hint. Zero dependencies, no built-in image decoding.
375
+ - The per-request image budget depends on the protocol: `anthropic-messages` 32 MB, others 20 MB (after base64). Earlier images are resent every turn; over budget, images are replaced with the placeholder text `[earlier image omitted to fit request size]` starting from the oldest, in one go down to below 60% of the budget; old images above the current endpoint's per-image limit (after switching model / channel) are replaced directly; with more than 20 images in one request and some longer side > 2000 px, the oldest keep being dropped until at most 20 remain. The newest message with images is never degraded. Degradation is written into the session as `context_edit{reason:"image_budget"}`, so the prefix is stable afterwards, and cache stats treat that request as a reset point.
376
+ - Clipboard images (`Ctrl+V` / `/paste` in the interactive UI): `pasteClipboardImage` tries `osascript` / `pngpaste` (macOS), `wl-paste` (Wayland) / `xclip` (X11) and PowerShell `Get-Clipboard -Format Image` (Windows), writing to `<data dir>/clipboard/<timestamp>.png`; `ama sessions prune` cleans files there older than 7 days.
377
+ - When the model's `input` lacks `image`, the request is refused up front with a hint to switch models (`-p` exits with 2, the UI shows an error, nothing is sent); when the `read` tool reads an image it only returns the path, dimensions and a note that the current model does not accept images.
378
+
379
+ Once connected: `ama models check packy/<id>` sends one minimal request to confirm connectivity; `ama models cache-probe packy/<id>` shows whether the endpoint reports cache usage (see "Caching" below); for models on a relay that do not report it, set `compat.cacheReporting: "silent"` as suggested, and the status bar shows "not reported" instead of 0%.
380
+
381
+ ## compat for the OpenAI-compatible line
382
+
383
+ Inference order: conservative defaults ← inference table (provider id, then baseUrl substring) ← `provider.compat` ← `model.compat`.
384
+
385
+ | Switch | Effect |
386
+ | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
387
+ | `maxTokensField` | `max_tokens` or `max_completion_tokens` |
388
+ | `supportsDeveloperRole` | Use the `developer` role for the system prompt of reasoning models |
389
+ | `supportsUsageInStreaming` | Send `stream_options.include_usage` |
390
+ | `supportsFinishReason` | When false, ignore finish_reason and infer the stop reason from content |
391
+ | `supportsReasoningEffort` | Send `reasoning_effort` |
392
+ | `thinkingFormat` | `openai` / `openrouter` (`reasoning.effort`) / `deepseek`, `zai` (`thinking.type`) / `qwen` (`enable_thinking`) / `none` |
393
+ | `thinkingTokenBudgetField` | Name of the budget field (DashScope: `thinking_budget`) |
394
+ | `requiresReasoningContentOnAssistantMessages` | Earlier assistant messages of reasoning models carry `reasoning_content` (DeepSeek) |
395
+ | `requiresToolResultName` | Tool result messages carry `name` (Mistral) |
396
+ | `requiresAssistantAfterToolResult` | Insert an assistant message when a user message directly follows a tool result |
397
+ | `supportsMidConvoSystemMessages` | Later system prompt patches are inserted back as system messages at their position |
398
+ | `cacheControlFormat` | `anthropic`: put `cache_control` on the system prompt, the last tool and the last user/tool message |
399
+ | `supportsStrictTools` | Send `strict: true` for strictly compatible tool schemas |
400
+ | `supportsStore` | Send `store: false` |
401
+
402
+ compat only records **verified** differences; please attach a docs link or a real sample when adding entries.
403
+
404
+ ## compat for Anthropic Messages
405
+
406
+ Defaults are inferred from the request host (`src/ai/apis/anthropic-compat.ts`); `providers.<id>.compat` and channel- or model-level `compat` override field by field:
407
+
408
+ | Switch | Effect | Default |
409
+ | ----------------------------- | ------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
410
+ | `sendInterleavedThinkingBeta` | Send the `interleaved-thinking` beta header for budgeted thinking + tools | On for `api.anthropic.com`; off for DeepSeek, Zhipu, Kimi, Qwen, MiniMax, StepFun, Tencent, Volcengine; other hosts only for `claude*` models |
411
+ | `sendCacheControl` | Place `cache_control` breakpoints | On; off for DeepSeek (documented as ignored) |
412
+ | `adaptiveThinking` | `{type:"adaptive"}` + `output_config.effort` | Off; enabled per entry in the catalog for new official models |
413
+ | `supportsCacheControlOnTools` | Breakpoint on the last tool definition | On |
414
+ | `maxCacheBreakpoints` | Maximum number of breakpoints | 4 |
415
+
416
+ OpenRouter's Messages endpoint only gives cache usage in `message_delta`; parsing already merges non-empty fields, so no switch is needed.
417
+
418
+ ## compat for Responses and Gemini
419
+
420
+ | Protocol | Switch | Effect | Default |
421
+ | ---------------------- | ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------- |
422
+ | `openai-responses` | `supportsReasoningSummary` | `reasoning.summary: "auto"` (thinking blocks get text) | On for official OpenAI; off for xAI and others |
423
+ | `openai-responses` | `supportsStore` | Send `store: false`; reasoning models also need `include: ["reasoning.encrypted_content"]` to replay encrypted reasoning items across turns | On for official OpenAI and xAI; off otherwise |
424
+ | `google-generative-ai` | `supportsThoughtSignature` | Replay `thoughtSignature` for the same model (thinking, text, functionCall parts) | On |
425
+ | `google-generative-ai` | `supportsFunctionResponseParts` | Put images of tool results into `functionResponse.parts`; when off, a separate user turn is started | On from Gemini 3; off for Gemini 2.x and non-Gemini names |
426
+
427
+ - Responses: the system prompt goes into `instructions`; cache fields are in "Caching" below; thinking off sends `reasoning.effort` only when the mapping table gives a string for off (such as `"none"`), otherwise the server default applies.
428
+ - Gemini: a mapped string value (or the Gemini 3 family without a mapping) → discrete `thinkingLevel` (`LOW` / `HIGH`…); a numeric mapping or other models → `thinkingBudget` (`-1` dynamic); off → `thinkingBudget: 0`. Implicit caching works automatically, and `cachedContentTokenCount` counts towards `cacheRead`.
429
+
430
+ ## Caching
431
+
432
+ Most usage in long tasks is cache reads: once the prefix changes, every later request re-reads it at full price. ama handles caching in three layers: the **protocol layer** places breakpoints, sends cache keys and retention tiers the way each provider expects, and marks whether a response had cache fields; the **session layer** records each request's prefix fingerprint, detects misses, decides whether the endpoint reports cache usage and warms the cache during long tool runs; the **display layer** is the status bar, `/session`, RPC stats and `ama models cache-probe` (how to read the interface is in [tui.md](tui.md) "Cache and context").
433
+
434
+ Prefix stability is guaranteed by assembly: the system prompt sections have a fixed order without timestamps, tools are sorted by name, and mid-session changes are only appended at the end as system patches ([session-format.md](../session-format.md), Chinese).
435
+
436
+ ### Request fields
437
+
438
+ | Protocol | Field | Condition |
439
+ | ----------------------------------------- | ------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
440
+ | `anthropic-messages` | Three `cache_control` breakpoints (last user message, end of system, last tool) | `cacheRetention` is not `none` |
441
+ | `anthropic-messages` | `ttl: "1h"` | `long` and `supportsLongCacheRetention`; otherwise 5m |
442
+ | `openai-completions` | `prompt_cache_key = sessionId` (cut to 64 characters) | `sendPromptCacheKey` and not `none` |
443
+ | `openai-completions` | `prompt_cache_retention: "24h"` | `long` and `supportsLongCacheRetention` |
444
+ | `openai-completions` (`anthropic/*` etc.) | `cache_control` (`cacheControlFormat: "anthropic"`), with `ttl: "1h"` for `long` | Same as Anthropic |
445
+ | `openai-responses` | `prompt_cache_key` | Same as Completions |
446
+ | `openai-responses` | `prompt_cache_options: { ttl: "30m" }`, otherwise `prompt_cache_retention: "24h"` | `long` and `supportsExplicitPromptCacheMode`; otherwise `long` and `supportsLongCacheRetention` |
447
+ | Both OpenAI lines | Affinity header `x-session-affinity` + a per-request `x-client-request-id` (OpenRouter: `x-session-id`) | `sendSessionAffinityHeaders` and a sessionId |
448
+ | `google-generative-ai` | None (implicit caching) | — |
449
+ | All | `toolChoice: "none"` → each provider's way of forbidding tool calls | When the request has tools (summary continuation does **not** use it, see "Compaction summary continuation") |
450
+
451
+ Retention tier: `StreamOptions.cacheRetention` wins; otherwise `AMA_CACHE_RETENTION=none|short|long` is read; with neither it is `short`. The Anthropic request body gets a final TTL order check (if a 1h appears after a 5m across tools → system → messages, everything falls back to 5m). When the Anthropic `baseUrl` ends with `/v1`, requests go to `{baseUrl}/messages`, never `/v1/v1/messages`.
452
+
453
+ ### Compat switches (`providers.<id>.compat` or model-level `compat`)
454
+
455
+ | Switch | Effect | Default |
456
+ | --------------------------------- | -------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
457
+ | `sendPromptCacheKey` | Send `prompt_cache_key` | On for the official hosts of OpenAI, xAI, Mistral, Kimi (`.cn` / `.ai`) and Tencent TokenHub; off otherwise |
458
+ | `sendSessionAffinityHeaders` | Send affinity headers | Off (including OpenRouter, not measured) |
459
+ | `supportsLongCacheRetention` | `long` available (Anthropic 1h, OpenAI 24h); otherwise falls back to short | On for `api.openai.com` / `api.anthropic.com` / `tokenhub.tencentmaas.com` |
460
+ | `supportsExplicitPromptCacheMode` | Responses' `prompt_cache_options` (30m) | Off |
461
+ | `cacheReporting` | `auto` / `silent` / `reported`: force the "reports cache" state | `auto` |
462
+
463
+ Inference only looks at the final request's host name (`HOST_CACHE_CAPABILITIES` in `cache-params.ts`, listing only hosts whose official docs state support), never at the provider id: pointing `openai` at a relay with `OPENAI_BASE_URL` or a custom `baseUrl` makes it a relay. Defaults are on only for official endpoints because relays were measured to "accept without visible benefit" or "accept but not apply" (table below).
464
+
465
+ **Automatic stripping on 400**: when an endpoint rejects with 400 and the error body names `prompt_cache_key` / `prompt_cache_retention` / `prompt_cache_options` / `cache_control`, ama records `provider/model` in an in-process strip table, resends once without those fields (still with a single terminal event), and suggests once which switch to write; later requests in the same process omit them directly.
466
+
467
+ ### usage and `cacheReported`
468
+
469
+ If any cache field appears in the raw usage (even 0) → `Usage.cacheReported = true`; none → `false`. Completions recognizes `prompt_tokens_details.cached_tokens` / `cache_write_tokens`, `prompt_cache_hit_tokens` / `prompt_cache_miss_tokens` and top-level `cached_tokens`; Responses recognizes `input_tokens_details.cached_tokens`; Anthropic recognizes `cache_read_input_tokens` / `cache_creation_input_tokens`; Google recognizes `cachedContentTokenCount` (often missing on implicit cache misses). Endpoints whose fields exist but stay 0 are judged `silent` by the session layer after 3 in a row.
470
+
471
+ ### Model catalog `promptCache`
472
+
473
+ Only publicly documented values are written (seconds / tokens): all Anthropic models `short 300 / long 3600`, with `minTokens` 512–4096 by model; all OpenAI models `short 300 / long 86400 / minTokens 1024`; all Kimi models `short 300`. DeepSeek, Zhipu, Qwen, Groq, xAI, Mistral, OpenRouter and Google promise no TTL, so it stays empty (no warming; attribution assumes 10 minutes for implicit caching). You can fill it in yourself in `models[]` / `modelOverrides[]`.
474
+
475
+ ### Session layer: fingerprints, misses and three states
476
+
477
+ Every real request records an in-memory entry: the prefix fingerprint (the first 16 hex characters of sha256 for the system prompt and the tool table each, plus `provider/model`), `promptTokens` (input + cacheRead + cacheWrite), usage and the send time. The next request is compared with the previous entry:
478
+
479
+ - **Miss**: `missed = min(previous prefix, current prefix) − current cacheRead`, ignored below the noise floor (`max(1024, promptCache.minTokens, endpoint cache granularity)`). Some endpoints report cache reads in blocks (DeepSeek through a relay uses blocks of 2048); the granularity is the GCD of the non-zero cacheRead values observed for the same endpoint (provider + host + model), trusted only with at least 2 samples and within 128–8192, kept in memory only and reused across sessions in the process. It counts as one miss when the relative share exceeds a scale-adaptive threshold (about `0.10 × √(100k / prefix)`, clamped to 2%–30%) or the absolute value is ≥ 20 000. The re-billing amount is estimated from the difference between the price actually paid and the read price; without model prices there are only tokens.
480
+ - **Cause** (decided in order): the system / tool table fingerprint changed → `prefix_changed` (`detail` says which part; usually the host registering tools mid-session or hook context changes); the model changed → `model_changed`; the gap exceeds the TTL → `idle` (implicit caches without a catalog TTL assume 10 minutes); a `task` sub-task took more than 80% of the gap between the two requests → `subtask`; otherwise → `evicted` (server eviction).
481
+ - **Not a miss**: the first request after compaction, a branch summary or tier-one pruning (the context changed legitimately); prefixes below the minimum cacheable length. A model switch is **not** exempt.
482
+ - **Three states**: maintained in-process per `(provider, baseUrl host name, model)`. `unknown`: no long enough comparable request yet; `reported`: a cacheRead or cacheWrite > 0 has been seen; `silent`: reads and writes were 0 for 3 comparable requests in a row (prefix ≥ minTokens, fingerprint unchanged, gap < TTL), or `compat.cacheReporting: "silent"`. Only `reported` shows a hit rate, detects misses and warms; `unknown` / `silent` requests stay out of the hit-rate denominator, and the interface shows `—` / "not reported" instead of 0%.
483
+
484
+ `task` sub-sessions have their own record chain and stats, summarized in the "Sub-tasks" line of `/session`; forked sessions keep using the root session id as `prompt_cache_key` (just a routing hint).
485
+
486
+ ### Warming
487
+
488
+ While a tool runs for a long time (long tests, `task` sub-tasks, codemode scripts), the prefix may expire before the next request. Warming replays the previous real request before the TTL expires (same model, same context, `maxTokens: 1`), paying only one read to keep the cache alive.
489
+
490
+ | Item | Rule |
491
+ | ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
492
+ | Modes | `off`; `streaming` (default, only while running, i.e. during tool execution); `idle` (also in the idle period after a run, for expensive models) |
493
+ | Preconditions | The endpoint is `reported`; the model catalog has `promptCache.short` (TTL > 10 seconds); the request's `cacheRetention` is not `none`; the request body was not replaced by `onPayload`; no warming for Anthropic with thinking whose budget is derived from `max_tokens` |
494
+ | Timing | Counted from when the previous request was **sent**, fired after `max(1s, min(0.9·TTL, TTL − 10s))`; if the timer fires later than the deadline (sleep, a blocked event loop) it stops |
495
+ | Economics | Sent only when `p · missCost − warmCost ≥ cache.minSavingsUsd` (default $0.05), with `p` 1 for streaming and 0.15 for idle; not sent without prices |
496
+ | Limits | 60 minutes for streaming, 30 minutes for idle; stops after 2 warmings in a row with zero hits |
497
+ | Cancellation | Cancelled on model switch, thinking level change, compaction, `/tree` and exit; restarts with the next real request |
498
+ | Accounting | A successful warming appends a `usage{kind:"cache_warm"}` entry (never in the context), counted in `/session` cost and RPC stats; events `cache_warm{scheduled|sent|stopped}` |
499
+
500
+ Hosts can veto or force every warming through `api.cache.onWarmingDecision` ([host-api.md](host-api.md) "Cache warming"). Sub-sessions do not warm by default (`cache.warmSubagents: true` turns it on). By catalog prices, warming during long tool runs almost always pays off; idle warming only pays off for expensive models with long prefixes.
501
+
502
+ ### Compaction summary continuation
503
+
504
+ The summary request of tier-two compaction no longer starts a separate conversation; instead it appends a summary instruction after a prefix byte-identical to the last real request (`cacheRetention: "short"`), so the whole history is billed at the read price. The continuation request **does not send `tool_choice`**: relays and Kimi were measured to render the prompt without tool definitions under `tool_choice: "none"`, breaking the prefix at the tools section so the cache is not read; Anthropic also documents that changing tool_choice invalidates the message cache. The tool table is sent as usual, and "do not call tools, only output the summary" is in the final instruction. When the response is empty, truncated, contains tool calls or the request fails, it falls back to a standalone summary request (`cacheRetention: "none"`) with a warning.
505
+
506
+ ### Configuration
507
+
508
+ ```json
509
+ {
510
+ "cache": {
511
+ "warming": "streaming",
512
+ "retention": "short",
513
+ "minSavingsUsd": 0.05,
514
+ "missNotices": true,
515
+ "warmSubagents": false
516
+ }
517
+ }
518
+ ```
519
+
520
+ | Key | Default | Description |
521
+ | --------------- | ----------- | --------------------------------------------------------------------------------------------------------------------------------- |
522
+ | `warming` | `streaming` | `off` / `streaming` / `idle`; the environment variable `AMA_CACHE_WARMING` overrides; `/cache warm …` switches it for the session |
523
+ | `retention` | `short` | `none` / `short` / `long`; the environment variable `AMA_CACHE_RETENTION` overrides |
524
+ | `minSavingsUsd` | `0.05` | Minimum expected savings for warming (USD) |
525
+ | `missNotices` | `true` | Miss and remaining-context notices in the message area (stats are unaffected) |
526
+ | `warmSubagents` | `false` | Also warm `task` sub-sessions |
527
+
528
+ The whole section is only honored in user-level and profile `config.json`; at project level it is ignored with a warning. Provider-level switches are in `providers.<id>.compat` ("Compat switches" above), and the TTL in the model's `promptCache`.
529
+
530
+ ### `ama models cache-probe`
531
+
532
+ ```sh
533
+ ama models cache-probe <provider/id> [--tokens 2048] [--gap-ms 3000] [--json] [--yes]
534
+ ```
535
+
536
+ It sends a deterministic fixed prefix (about `--tokens` tokens) + `Reply with: ok`, `maxTokens: 16`, twice `--gap-ms` apart, and decides:
537
+
538
+ - `reported`: the second cacheRead ≥ 50% of the prefix; when the catalog has no `promptCache`, it suggests filling in `promptCache.short` to enable warming;
539
+ - `silent`: reads and writes were 0 both times; it suggests setting `compat.cacheReporting: "silent"`. When the response has cache fields that stay 0, it also hints at a possible write delay (kimi-k2.5 on the same relay read 0 twice 3 seconds apart, but the full prefix on the second request 8 seconds apart), so retry with a larger `--gap-ms`;
540
+ - `inconclusive`: a small read or only writes, most likely a cache granularity or TTL issue.
541
+
542
+ It prints input / cacheRead / cacheWrite of both requests and the usage field names that protocol reads. This is a billed action: an estimate is printed first (`$?` without prices), an interactive terminal asks y/N once, and non-interactive use requires `--yes` (otherwise exit 2); with `--json` the estimate goes to stderr and stdout carries only the result object.
543
+
544
+ ### Relay measurements (2026-10-02)
545
+
546
+ A test relay offering all three of Chat / Responses / Messages, fixed prefixes of about 8.6k–10.9k tokens, the same prefix sent twice 2–3 seconds apart (38 requests in total).
547
+
548
+ | Endpoint / model | Cache parameters sent | Result | Second cacheRead / prefix | Cache fields in usage (raw shape) |
549
+ | ------------------------------------------------- | ------------------------------------------------- | ----------------------- | -------------------------- | ----------------------------------------------------------------------------------------------------- |
550
+ | Chat · kimi-k2.5 | none | 200 | 8576 / 8597 | `prompt_tokens_details.cached_tokens` (0 on the first request) |
551
+ | Chat · kimi-k2.5 | `prompt_cache_key`; then plus affinity headers | 200, accepted | 8576 / 8597 (no gain) | Same as above |
552
+ | Chat · deepseek-v4-flash | `prompt_cache_key` + `x-session-affinity` | 200, accepted | 9472 / 9767 | `prompt_tokens_details.cached_tokens` |
553
+ | Chat · qwen3.8-flash | Same as above | 200, accepted | 10240 / 10400 | Same as above |
554
+ | Chat · glm-5 | Same as above | 200, accepted | 8704 / 9040 | Same as above |
555
+ | Chat · MiniMax-M2.7 | Same as above | 200, accepted | 9389 / 9700 | Same as above |
556
+ | Responses · qwen3.8-flash | `prompt_cache_key` + `prompt_cache_retention` | 200, accepted | 10240 / 10432 | `input_tokens_details.cached_tokens` |
557
+ | Responses · qwen3.8-flash | `prompt_cache_options: {ttl:"30m"}` | 200, accepted | Same as above | Same as above |
558
+ | Responses · grok-4.7 | Both sets above | 200, accepted | 1152 / 10929 | Same as above |
559
+ | Responses · deepseek-v4-flash | `prompt_cache_retention` / `prompt_cache_options` | **400** `unknown field` | — | 200 once removed; `input_tokens_details.cached_tokens` |
560
+ | Messages · kimi-k2.5, qwen3.8-flash | `cache_control` + `ttl:"1h"` | 200, accepted | 9464 / 9476, 10385 / 10400 | `cache_creation.ephemeral_5m_input_tokens` has values, no 1h field: **1h treated as 5m** |
561
+ | Messages · MiniMax-M2.7, deepseek-v4-flash, glm-5 | Same as above | 200, accepted | 9403, 8192, 8704 | `cache_creation_input_tokens` always 0, reads hit as usual (implicit caching managed by the endpoint) |
562
+
563
+ Conclusions and defaults:
564
+
565
+ - `prompt_cache_key` and affinity headers are accepted on the relay but show no hit improvement → `sendPromptCacheKey` and `sendSessionAffinityHeaders` default to off for non-official endpoints; turn them on yourself when needed (400 stripping is the safety net).
566
+ - 1h retention is accepted on the relay's Messages endpoint but written as 5m, and the long retention fields of Responses return 400 on some upstreams → `supportsLongCacheRetention` defaults to on only for official endpoints, and `supportsExplicitPromptCacheMode` defaults to off.
567
+ - All five measured vendors report cache fields on the Chat endpoint (the zero reports from DeepSeek / GLM in an earlier study did not reproduce this time); the `cached_tokens: 0` on their first request is exactly "the field exists but is 0", so `cacheReported` is true.
568
+ - The `/v1` deduplication and `toolChoice: "none"` were verified by real requests through ama's protocol layer: Messages requests land on `/v1/messages`, and all three endpoints accept `tool_choice: none` without producing tool calls.
569
+
570
+ ## Channel measurements
571
+
572
+ `scripts/channel-probe.mjs` (run `pnpm build:lib` first) runs the measurement gate on `provider/model@channel`, with ≤ 8 requests per model:
573
+
574
+ | Item | Method | Pass |
575
+ | --------------------- | -------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
576
+ | ① check | One minimal call | No error and some text |
577
+ | ② tool round trip | One parallel `read` of two temp files, answering with a passphrase | Tool calls made and the answer contains the passphrase |
578
+ | ③ thinking, two turns | In the same conversation, read another file with thinking=medium (thinking blocks replayed with history) | No error on the second request; no thinking blocks, or thinking blocks without signatures on Messages, are marked ⚠ (signature replay unverified) |
579
+ | ④ cache | A fixed prefix of about 3k tokens sent twice `--gap-ms` apart | Second cacheRead > 0 |
580
+ | tool_use.id | All tool call ids in the same conversation of ②③ | All distinct |
581
+
582
+ All four items passing with distinct ids = passing the gate; once the official Messages channels of DeepSeek, Zhipu and Kimi pass, their default switches to messages.
583
+
584
+ ### Relay comparison (2026-10-02, messages vs chat)
585
+
586
+ The same test relay (Chat / Messages endpoints), `--gap-ms 8000`, 84 requests (including one debugging round), conservatively priced at about $0.11:
587
+
588
+ | Model@channel | Protocol · host | ① check | ② tool round trip | ③ thinking, two turns | ④ cache (read / prefix) | ids distinct | Requests | Verdict |
589
+ | ---------------------------------- | ------------------------------------- | ------- | ----------------- | ------------------------------ | ----------------------- | ------------ | -------- | -------- |
590
+ | `packy/kimi-k2.5@messages` | anthropic-messages · www.packyapi.com | ✓ | ✓ | ⚠ 2/2 thinking blocks unsigned | ✓ 2734 / 2738 | ✓ | 7 | fail |
591
+ | `packy/kimi-k2.5@chat` | openai-completions · www.packyapi.com | ✓ | ✓ | ✓ | ✓ 2688 / 2740 | ✓ | 7 | **pass** |
592
+ | `packy/deepseek-v4-flash@messages` | anthropic-messages · www.packyapi.com | ✓ | ✓ | ⚠ no thinking blocks | ✓ 2048 / 2732 | ✓ | 7 | fail |
593
+ | `packy/deepseek-v4-flash@chat` | openai-completions · www.packyapi.com | ✓ | ✓ | ✓ | ✓ 2048 / 2732 | ✓ | 7 | **pass** |
594
+ | `packy/glm-5@messages` | anthropic-messages · www.packyapi.com | ✓ | ✓ | ⚠ 2/2 thinking blocks unsigned | ✓ 2560 / 2733 | ✓ | 7 | fail |
595
+ | `packy/glm-5@chat` | openai-completions · www.packyapi.com | ✓ | ✓ | ✓ | ✓ 2560 / 2733 | ✓ | 7 | **pass** |
596
+ | `packy/qwen3.8-flash@messages` | anthropic-messages · www.packyapi.com | ✓ | ✓ | ⚠ 2/2 thinking blocks unsigned | ✓ 3053 / 3061 | ✓ | 7 | fail |
597
+ | `packy/qwen3.8-flash@chat` | openai-completions · www.packyapi.com | ✓ | ✓ | ✓ | ✓ 2048 / 3063 | ✓ | 7 | **pass** |
598
+ | `packy/MiniMax-M2.7@messages` | anthropic-messages · www.packyapi.com | ✓ | ✓ | ⚠ 2/2 thinking blocks unsigned | ✓ 2638 / 2743 | ✓ | 7 | fail |
599
+ | `packy/MiniMax-M2.7@chat` | openai-completions · www.packyapi.com | ✓ | ✓ | ✓ | ✓ 2638 / 2744 | ✓ | 7 | **pass** |
600
+
601
+ - All five vendors pass ①②④ on both relay endpoints, and tool_use.id never repeats within one conversation (Kimi's ids look like `functions.read:<n>` and increase per conversation; different conversations restart from 0, so ids can only be compared within one conversation).
602
+ - Thinking blocks on the Messages endpoint are **all unsigned** (DeepSeek returned no thinking blocks): the relay converts upstream Chat into Messages without signatures, ama degrades to text replay by rule, and round trips succeed, but signature replay was not verified, so **relay results cannot replace the measurement gate on direct official endpoints**.
603
+ - Cache: Qwen on Messages writes 3053 tokens on the very first request (honoring `cache_control`) and reads the full prefix the second time; the other four read hits as usual with `cache_creation_input_tokens` at 0 (implicit caching managed by the endpoint), consistent with "Caching · relay measurements".
604
+
605
+ ### Direct official endpoints (to be run by the user with their own keys)
606
+
607
+ ≤ 8 requests per vendor; Kimi's cache writes are delayed, so use `--gap-ms 8000`:
608
+
609
+ ```sh
610
+ pnpm build:lib
611
+ node scripts/channel-probe.mjs --model deepseek/deepseek-v4-pro@messages,deepseek/deepseek-v4-pro@chat --json /tmp/probe-deepseek.json
612
+ node scripts/channel-probe.mjs --model zhipu/glm-5.3@messages,zhipu/glm-5.3@chat --json /tmp/probe-zhipu.json
613
+ node scripts/channel-probe.mjs --model moonshot/kimi-k3@messages,moonshot/kimi-k3@chat --gap-ms 8000 --json /tmp/probe-kimi.json
614
+ node scripts/channel-probe.mjs --model dashscope/qwen3.8-max@messages --json /tmp/probe-qwen.json
615
+ node scripts/channel-probe.mjs --model minimax/MiniMax-M3@messages,minimax/MiniMax-M2.7@messages --json /tmp/probe-minimax.json
616
+ node scripts/channel-probe.mjs --model stepfun/step-5-preview@messages --json /tmp/probe-stepfun.json
617
+ node scripts/channel-probe.mjs --model tencent/hy3@messages --json /tmp/probe-tencent.json
618
+ node scripts/channel-probe.mjs --model volcengine/doubao-seed-2-1-pro-260628@responses --json /tmp/probe-volcengine.json
619
+ ```
620
+
621
+ Keys come from each vendor's standard environment variables (table above); a vendor without a key is recorded as "no key" and no request is sent. Results are re-rendered with `--render` and added to this section.
622
+
623
+ ## The fake provider for tests
624
+
625
+ `--model fake/echo`: echoes the last user message. The zero-config model picker, `ama doctor`, `ama models list`, `ama providers list` and `ama config show` hide fake by default; it is listed with `AMA_SHOW_FAKE=1` or when `AMA_FAKE_SCRIPT` is set, and an explicit `--model fake/…` works at any time. With `AMA_FAKE_SCRIPT=<file.json>` the n-th call produces text, thinking, tool calls, 429s, overflow, dropped streams or delays per the script; the script format is in `src/ai/fake/fake-script.ts`, with examples in `test/fixtures/scripts/`.