pi-usereq 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (361) hide show
  1. package/.g.conf +20 -0
  2. package/.github/workflows/release-npm.yml +131 -0
  3. package/.gitignore +270 -0
  4. package/CHANGELOG.md +221 -0
  5. package/LICENSE +674 -0
  6. package/README.md +260 -0
  7. package/TODO.md +20 -0
  8. package/docs/pi.dev/agent-document-manifest.json +6399 -0
  9. package/docs/pi.dev/coding-agent-docs/compaction.md +394 -0
  10. package/docs/pi.dev/coding-agent-docs/custom-provider.md +596 -0
  11. package/docs/pi.dev/coding-agent-docs/development.md +71 -0
  12. package/docs/pi.dev/coding-agent-docs/extensions.md +2262 -0
  13. package/docs/pi.dev/coding-agent-docs/images/doom-extension.png +0 -0
  14. package/docs/pi.dev/coding-agent-docs/images/exy.png +0 -0
  15. package/docs/pi.dev/coding-agent-docs/images/interactive-mode.png +0 -0
  16. package/docs/pi.dev/coding-agent-docs/images/tree-view.png +0 -0
  17. package/docs/pi.dev/coding-agent-docs/json.md +82 -0
  18. package/docs/pi.dev/coding-agent-docs/keybindings.md +175 -0
  19. package/docs/pi.dev/coding-agent-docs/models.md +392 -0
  20. package/docs/pi.dev/coding-agent-docs/packages.md +218 -0
  21. package/docs/pi.dev/coding-agent-docs/prompt-templates.md +67 -0
  22. package/docs/pi.dev/coding-agent-docs/providers.md +195 -0
  23. package/docs/pi.dev/coding-agent-docs/rpc.md +1377 -0
  24. package/docs/pi.dev/coding-agent-docs/sdk.md +1124 -0
  25. package/docs/pi.dev/coding-agent-docs/session.md +412 -0
  26. package/docs/pi.dev/coding-agent-docs/settings.md +247 -0
  27. package/docs/pi.dev/coding-agent-docs/shell-aliases.md +13 -0
  28. package/docs/pi.dev/coding-agent-docs/skills.md +232 -0
  29. package/docs/pi.dev/coding-agent-docs/terminal-setup.md +106 -0
  30. package/docs/pi.dev/coding-agent-docs/termux.md +127 -0
  31. package/docs/pi.dev/coding-agent-docs/themes.md +295 -0
  32. package/docs/pi.dev/coding-agent-docs/tmux.md +61 -0
  33. package/docs/pi.dev/coding-agent-docs/tree.md +231 -0
  34. package/docs/pi.dev/coding-agent-docs/tui.md +887 -0
  35. package/docs/pi.dev/coding-agent-docs/windows.md +17 -0
  36. package/docs/pi.dev/mom-docs/artifacts-server.md +475 -0
  37. package/docs/pi.dev/mom-docs/events.md +307 -0
  38. package/docs/pi.dev/mom-docs/new.md +970 -0
  39. package/docs/pi.dev/mom-docs/sandbox.md +153 -0
  40. package/docs/pi.dev/mom-docs/slack-bot-minimal-guide.md +399 -0
  41. package/docs/pi.dev/mom-docs/v86.md +319 -0
  42. package/docs/pi.dev/pods-docs/gml-4.5.md +189 -0
  43. package/docs/pi.dev/pods-docs/gpt-oss.md +233 -0
  44. package/docs/pi.dev/pods-docs/implementation-plan.md +183 -0
  45. package/docs/pi.dev/pods-docs/kimi-k2.md +197 -0
  46. package/docs/pi.dev/pods-docs/models.md +116 -0
  47. package/docs/pi.dev/pods-docs/plan.md +166 -0
  48. package/docs/pi.dev/pods-docs/qwen3-coder.md +132 -0
  49. package/images/flowchart-bw.png +0 -0
  50. package/images/flowchart-bw.svg +102 -0
  51. package/images/flowchart.md +100 -0
  52. package/images/flowchart.png +0 -0
  53. package/images/flowchart.svg +3 -0
  54. package/package.json +46 -0
  55. package/req/docs/REFERENCES.md +4554 -0
  56. package/req/docs/REQUIREMENTS.md +475 -0
  57. package/req/docs/WORKFLOW.md +1059 -0
  58. package/scripts/debug-extension.ts +497 -0
  59. package/scripts/lib/extension-debug-harness.ts +450 -0
  60. package/scripts/lib/recording-extension-api.ts +786 -0
  61. package/scripts/lib/sdk-smoke.ts +503 -0
  62. package/scripts/pi-usereq-debug.sh +330 -0
  63. package/scripts/tool-args-to-params.ts +208 -0
  64. package/src/cli.ts +349 -0
  65. package/src/core/agent-tool-json.ts +621 -0
  66. package/src/core/compress-files.ts +72 -0
  67. package/src/core/compress-payload.ts +648 -0
  68. package/src/core/compress.ts +464 -0
  69. package/src/core/config.ts +272 -0
  70. package/src/core/doxygen-parser.ts +318 -0
  71. package/src/core/errors.ts +31 -0
  72. package/src/core/extension-status.ts +660 -0
  73. package/src/core/find-constructs.ts +319 -0
  74. package/src/core/find-payload.ts +915 -0
  75. package/src/core/generate-markdown.ts +120 -0
  76. package/src/core/path-context.ts +196 -0
  77. package/src/core/pi-notify.ts +430 -0
  78. package/src/core/pi-usereq-tools.ts +140 -0
  79. package/src/core/prompts.ts +184 -0
  80. package/src/core/reference-payload.ts +818 -0
  81. package/src/core/resources.ts +63 -0
  82. package/src/core/runtime-project-paths.ts +99 -0
  83. package/src/core/settings-menu.ts +233 -0
  84. package/src/core/source-analyzer.ts +1721 -0
  85. package/src/core/static-check.ts +674 -0
  86. package/src/core/token-counter.ts +729 -0
  87. package/src/core/tool-runner.ts +717 -0
  88. package/src/core/utils.ts +185 -0
  89. package/src/index.ts +2209 -0
  90. package/src/resources/guidelines/Google_C++_Style_Guide.md +3711 -0
  91. package/src/resources/guidelines/Google_Python_Style_Guide.md +3709 -0
  92. package/src/resources/prompts/analyze.md +130 -0
  93. package/src/resources/prompts/change.md +227 -0
  94. package/src/resources/prompts/check.md +139 -0
  95. package/src/resources/prompts/cover.md +219 -0
  96. package/src/resources/prompts/create.md +104 -0
  97. package/src/resources/prompts/fix.md +221 -0
  98. package/src/resources/prompts/flowchart.md +220 -0
  99. package/src/resources/prompts/implement.md +163 -0
  100. package/src/resources/prompts/new.md +226 -0
  101. package/src/resources/prompts/readme.md +182 -0
  102. package/src/resources/prompts/recreate.md +213 -0
  103. package/src/resources/prompts/refactor.md +213 -0
  104. package/src/resources/prompts/references.md +100 -0
  105. package/src/resources/prompts/renumber.md +119 -0
  106. package/src/resources/prompts/workflow.md +202 -0
  107. package/src/resources/prompts/write.md +99 -0
  108. package/src/resources/sounds/Machine-alert-beep-sound-effect.mp3 +0 -0
  109. package/src/resources/sounds/Soft-high-tech-notification-sound-effect.mp3 +0 -0
  110. package/src/resources/templates/Document_Source_Code_in_Doxygen_Style.md +130 -0
  111. package/src/resources/templates/HDT_Test_Authoring_Guide.md +318 -0
  112. package/src/resources/templates/Requirements_Template.md +78 -0
  113. package/tests/attended-results-scenarios.ts +758 -0
  114. package/tests/attended-results.test.ts +39 -0
  115. package/tests/cli-command-option-parity.test.ts +815 -0
  116. package/tests/debug-extension-harness.test.ts +463 -0
  117. package/tests/extension-registration.test.ts +2006 -0
  118. package/tests/fixtures/fixture_c.c +361 -0
  119. package/tests/fixtures/fixture_cpp.cpp +407 -0
  120. package/tests/fixtures/fixture_csharp.cs +411 -0
  121. package/tests/fixtures/fixture_elixir.ex +409 -0
  122. package/tests/fixtures/fixture_go.go +341 -0
  123. package/tests/fixtures/fixture_haskell.hs +250 -0
  124. package/tests/fixtures/fixture_java.java +430 -0
  125. package/tests/fixtures/fixture_javascript.js +383 -0
  126. package/tests/fixtures/fixture_kotlin.kt +451 -0
  127. package/tests/fixtures/fixture_lua.lua +276 -0
  128. package/tests/fixtures/fixture_perl.pl +310 -0
  129. package/tests/fixtures/fixture_php.php +433 -0
  130. package/tests/fixtures/fixture_python.py +502 -0
  131. package/tests/fixtures/fixture_ruby.rb +345 -0
  132. package/tests/fixtures/fixture_rust.rs +380 -0
  133. package/tests/fixtures/fixture_scala.scala +398 -0
  134. package/tests/fixtures/fixture_shell.sh +276 -0
  135. package/tests/fixtures/fixture_swift.swift +397 -0
  136. package/tests/fixtures/fixture_typescript.ts +434 -0
  137. package/tests/fixtures/fixture_zig.zig +295 -0
  138. package/tests/fixtures_attended_results/project/compress-line-numbers.json +5 -0
  139. package/tests/fixtures_attended_results/project/compress.json +5 -0
  140. package/tests/fixtures_attended_results/project/enable-static-check-invalid-command.json +5 -0
  141. package/tests/fixtures_attended_results/project/enable-static-check-valid.json +5 -0
  142. package/tests/fixtures_attended_results/project/files-static-check.json +5 -0
  143. package/tests/fixtures_attended_results/project/find-line-numbers.json +5 -0
  144. package/tests/fixtures_attended_results/project/find.json +5 -0
  145. package/tests/fixtures_attended_results/project/get-base-path.json +5 -0
  146. package/tests/fixtures_attended_results/project/git-check-clean.json +5 -0
  147. package/tests/fixtures_attended_results/project/git-check-dirty.json +5 -0
  148. package/tests/fixtures_attended_results/project/git-path.json +5 -0
  149. package/tests/fixtures_attended_results/project/git-wt-create-invalid.json +5 -0
  150. package/tests/fixtures_attended_results/project/git-wt-create-valid.json +5 -0
  151. package/tests/fixtures_attended_results/project/git-wt-delete-nonexistent.json +5 -0
  152. package/tests/fixtures_attended_results/project/git-wt-delete-valid.json +5 -0
  153. package/tests/fixtures_attended_results/project/git-wt-name.json +5 -0
  154. package/tests/fixtures_attended_results/project/references.json +5 -0
  155. package/tests/fixtures_attended_results/project/static-check.json +5 -0
  156. package/tests/fixtures_attended_results/project/tokens.json +5 -0
  157. package/tests/fixtures_attended_results/standalone/files-compress/fixture_c.c.json +5 -0
  158. package/tests/fixtures_attended_results/standalone/files-compress/fixture_cpp.cpp.json +5 -0
  159. package/tests/fixtures_attended_results/standalone/files-compress/fixture_csharp.cs.json +5 -0
  160. package/tests/fixtures_attended_results/standalone/files-compress/fixture_elixir.ex.json +5 -0
  161. package/tests/fixtures_attended_results/standalone/files-compress/fixture_go.go.json +5 -0
  162. package/tests/fixtures_attended_results/standalone/files-compress/fixture_haskell.hs.json +5 -0
  163. package/tests/fixtures_attended_results/standalone/files-compress/fixture_java.java.json +5 -0
  164. package/tests/fixtures_attended_results/standalone/files-compress/fixture_javascript.js.json +5 -0
  165. package/tests/fixtures_attended_results/standalone/files-compress/fixture_kotlin.kt.json +5 -0
  166. package/tests/fixtures_attended_results/standalone/files-compress/fixture_lua.lua.json +5 -0
  167. package/tests/fixtures_attended_results/standalone/files-compress/fixture_perl.pl.json +5 -0
  168. package/tests/fixtures_attended_results/standalone/files-compress/fixture_php.php.json +5 -0
  169. package/tests/fixtures_attended_results/standalone/files-compress/fixture_python.py.json +5 -0
  170. package/tests/fixtures_attended_results/standalone/files-compress/fixture_ruby.rb.json +5 -0
  171. package/tests/fixtures_attended_results/standalone/files-compress/fixture_rust.rs.json +5 -0
  172. package/tests/fixtures_attended_results/standalone/files-compress/fixture_scala.scala.json +5 -0
  173. package/tests/fixtures_attended_results/standalone/files-compress/fixture_shell.sh.json +5 -0
  174. package/tests/fixtures_attended_results/standalone/files-compress/fixture_swift.swift.json +5 -0
  175. package/tests/fixtures_attended_results/standalone/files-compress/fixture_typescript.ts.json +5 -0
  176. package/tests/fixtures_attended_results/standalone/files-compress/fixture_zig.zig.json +5 -0
  177. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_c.c.json +5 -0
  178. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_cpp.cpp.json +5 -0
  179. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_csharp.cs.json +5 -0
  180. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_elixir.ex.json +5 -0
  181. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_go.go.json +5 -0
  182. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_haskell.hs.json +5 -0
  183. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_java.java.json +5 -0
  184. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_javascript.js.json +5 -0
  185. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_kotlin.kt.json +5 -0
  186. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_lua.lua.json +5 -0
  187. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_perl.pl.json +5 -0
  188. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_php.php.json +5 -0
  189. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_python.py.json +5 -0
  190. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_ruby.rb.json +5 -0
  191. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_rust.rs.json +5 -0
  192. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_scala.scala.json +5 -0
  193. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_shell.sh.json +5 -0
  194. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_swift.swift.json +5 -0
  195. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_typescript.ts.json +5 -0
  196. package/tests/fixtures_attended_results/standalone/files-compress-line-numbers/fixture_zig.zig.json +5 -0
  197. package/tests/fixtures_attended_results/standalone/files-find/fixture_c.c.json +5 -0
  198. package/tests/fixtures_attended_results/standalone/files-find/fixture_cpp.cpp.json +5 -0
  199. package/tests/fixtures_attended_results/standalone/files-find/fixture_csharp.cs.json +5 -0
  200. package/tests/fixtures_attended_results/standalone/files-find/fixture_elixir.ex.json +5 -0
  201. package/tests/fixtures_attended_results/standalone/files-find/fixture_go.go.json +5 -0
  202. package/tests/fixtures_attended_results/standalone/files-find/fixture_haskell.hs.json +5 -0
  203. package/tests/fixtures_attended_results/standalone/files-find/fixture_java.java.json +5 -0
  204. package/tests/fixtures_attended_results/standalone/files-find/fixture_javascript.js.json +5 -0
  205. package/tests/fixtures_attended_results/standalone/files-find/fixture_kotlin.kt.json +5 -0
  206. package/tests/fixtures_attended_results/standalone/files-find/fixture_lua.lua.json +5 -0
  207. package/tests/fixtures_attended_results/standalone/files-find/fixture_perl.pl.json +5 -0
  208. package/tests/fixtures_attended_results/standalone/files-find/fixture_php.php.json +5 -0
  209. package/tests/fixtures_attended_results/standalone/files-find/fixture_python.py.json +5 -0
  210. package/tests/fixtures_attended_results/standalone/files-find/fixture_ruby.rb.json +5 -0
  211. package/tests/fixtures_attended_results/standalone/files-find/fixture_rust.rs.json +5 -0
  212. package/tests/fixtures_attended_results/standalone/files-find/fixture_scala.scala.json +5 -0
  213. package/tests/fixtures_attended_results/standalone/files-find/fixture_shell.sh.json +5 -0
  214. package/tests/fixtures_attended_results/standalone/files-find/fixture_swift.swift.json +5 -0
  215. package/tests/fixtures_attended_results/standalone/files-find/fixture_typescript.ts.json +5 -0
  216. package/tests/fixtures_attended_results/standalone/files-find/fixture_zig.zig.json +5 -0
  217. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_c.c.json +5 -0
  218. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_cpp.cpp.json +5 -0
  219. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_csharp.cs.json +5 -0
  220. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_elixir.ex.json +5 -0
  221. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_go.go.json +5 -0
  222. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_haskell.hs.json +5 -0
  223. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_java.java.json +5 -0
  224. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_javascript.js.json +5 -0
  225. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_kotlin.kt.json +5 -0
  226. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_lua.lua.json +5 -0
  227. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_perl.pl.json +5 -0
  228. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_php.php.json +5 -0
  229. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_python.py.json +5 -0
  230. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_ruby.rb.json +5 -0
  231. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_rust.rs.json +5 -0
  232. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_scala.scala.json +5 -0
  233. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_shell.sh.json +5 -0
  234. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_swift.swift.json +5 -0
  235. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_typescript.ts.json +5 -0
  236. package/tests/fixtures_attended_results/standalone/files-find-line-numbers/fixture_zig.zig.json +5 -0
  237. package/tests/fixtures_attended_results/standalone/files-references/fixture_c.c.json +5 -0
  238. package/tests/fixtures_attended_results/standalone/files-references/fixture_cpp.cpp.json +5 -0
  239. package/tests/fixtures_attended_results/standalone/files-references/fixture_csharp.cs.json +5 -0
  240. package/tests/fixtures_attended_results/standalone/files-references/fixture_elixir.ex.json +5 -0
  241. package/tests/fixtures_attended_results/standalone/files-references/fixture_go.go.json +5 -0
  242. package/tests/fixtures_attended_results/standalone/files-references/fixture_haskell.hs.json +5 -0
  243. package/tests/fixtures_attended_results/standalone/files-references/fixture_java.java.json +5 -0
  244. package/tests/fixtures_attended_results/standalone/files-references/fixture_javascript.js.json +5 -0
  245. package/tests/fixtures_attended_results/standalone/files-references/fixture_kotlin.kt.json +5 -0
  246. package/tests/fixtures_attended_results/standalone/files-references/fixture_lua.lua.json +5 -0
  247. package/tests/fixtures_attended_results/standalone/files-references/fixture_perl.pl.json +5 -0
  248. package/tests/fixtures_attended_results/standalone/files-references/fixture_php.php.json +5 -0
  249. package/tests/fixtures_attended_results/standalone/files-references/fixture_python.py.json +5 -0
  250. package/tests/fixtures_attended_results/standalone/files-references/fixture_ruby.rb.json +5 -0
  251. package/tests/fixtures_attended_results/standalone/files-references/fixture_rust.rs.json +5 -0
  252. package/tests/fixtures_attended_results/standalone/files-references/fixture_scala.scala.json +5 -0
  253. package/tests/fixtures_attended_results/standalone/files-references/fixture_shell.sh.json +5 -0
  254. package/tests/fixtures_attended_results/standalone/files-references/fixture_swift.swift.json +5 -0
  255. package/tests/fixtures_attended_results/standalone/files-references/fixture_typescript.ts.json +5 -0
  256. package/tests/fixtures_attended_results/standalone/files-references/fixture_zig.zig.json +5 -0
  257. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_c.c.json +5 -0
  258. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_cpp.cpp.json +5 -0
  259. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_csharp.cs.json +5 -0
  260. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_elixir.ex.json +5 -0
  261. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_go.go.json +5 -0
  262. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_haskell.hs.json +5 -0
  263. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_java.java.json +5 -0
  264. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_javascript.js.json +5 -0
  265. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_kotlin.kt.json +5 -0
  266. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_lua.lua.json +5 -0
  267. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_perl.pl.json +5 -0
  268. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_php.php.json +5 -0
  269. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_python.py.json +5 -0
  270. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_ruby.rb.json +5 -0
  271. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_rust.rs.json +5 -0
  272. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_scala.scala.json +5 -0
  273. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_shell.sh.json +5 -0
  274. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_swift.swift.json +5 -0
  275. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_typescript.ts.json +5 -0
  276. package/tests/fixtures_attended_results/standalone/files-tokens/fixture_zig.zig.json +5 -0
  277. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_c.c.json +5 -0
  278. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_cpp.cpp.json +5 -0
  279. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_csharp.cs.json +5 -0
  280. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_elixir.ex.json +5 -0
  281. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_go.go.json +5 -0
  282. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_haskell.hs.json +5 -0
  283. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_java.java.json +5 -0
  284. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_javascript.js.json +5 -0
  285. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_kotlin.kt.json +5 -0
  286. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_lua.lua.json +5 -0
  287. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_perl.pl.json +5 -0
  288. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_php.php.json +5 -0
  289. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_python.py.json +5 -0
  290. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_ruby.rb.json +5 -0
  291. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_rust.rs.json +5 -0
  292. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_scala.scala.json +5 -0
  293. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_shell.sh.json +5 -0
  294. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_swift.swift.json +5 -0
  295. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_typescript.ts.json +5 -0
  296. package/tests/fixtures_attended_results/standalone/test-static-check-command/fixture_zig.zig.json +5 -0
  297. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_c.c.json +5 -0
  298. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_cpp.cpp.json +5 -0
  299. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_csharp.cs.json +5 -0
  300. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_elixir.ex.json +5 -0
  301. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_go.go.json +5 -0
  302. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_haskell.hs.json +5 -0
  303. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_java.java.json +5 -0
  304. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_javascript.js.json +5 -0
  305. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_kotlin.kt.json +5 -0
  306. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_lua.lua.json +5 -0
  307. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_perl.pl.json +5 -0
  308. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_php.php.json +5 -0
  309. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_python.py.json +5 -0
  310. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_ruby.rb.json +5 -0
  311. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_rust.rs.json +5 -0
  312. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_scala.scala.json +5 -0
  313. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_shell.sh.json +5 -0
  314. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_swift.swift.json +5 -0
  315. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_typescript.ts.json +5 -0
  316. package/tests/fixtures_attended_results/standalone/test-static-check-dummy/fixture_zig.zig.json +5 -0
  317. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_c.c.json +5 -0
  318. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_cpp.cpp.json +5 -0
  319. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_csharp.cs.json +5 -0
  320. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_elixir.ex.json +5 -0
  321. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_go.go.json +5 -0
  322. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_haskell.hs.json +5 -0
  323. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_java.java.json +5 -0
  324. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_javascript.js.json +5 -0
  325. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_kotlin.kt.json +5 -0
  326. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_lua.lua.json +5 -0
  327. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_perl.pl.json +5 -0
  328. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_php.php.json +5 -0
  329. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_python.py.json +5 -0
  330. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_ruby.rb.json +5 -0
  331. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_rust.rs.json +5 -0
  332. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_scala.scala.json +5 -0
  333. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_shell.sh.json +5 -0
  334. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_swift.swift.json +5 -0
  335. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_typescript.ts.json +5 -0
  336. package/tests/fixtures_attended_results/standalone/test-static-check-pylance/fixture_zig.zig.json +5 -0
  337. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_c.c.json +5 -0
  338. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_cpp.cpp.json +5 -0
  339. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_csharp.cs.json +5 -0
  340. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_elixir.ex.json +5 -0
  341. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_go.go.json +5 -0
  342. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_haskell.hs.json +5 -0
  343. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_java.java.json +5 -0
  344. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_javascript.js.json +5 -0
  345. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_kotlin.kt.json +5 -0
  346. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_lua.lua.json +5 -0
  347. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_perl.pl.json +5 -0
  348. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_php.php.json +5 -0
  349. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_python.py.json +5 -0
  350. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_ruby.rb.json +5 -0
  351. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_rust.rs.json +5 -0
  352. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_scala.scala.json +5 -0
  353. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_shell.sh.json +5 -0
  354. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_swift.swift.json +5 -0
  355. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_typescript.ts.json +5 -0
  356. package/tests/fixtures_attended_results/standalone/test-static-check-ruff/fixture_zig.zig.json +5 -0
  357. package/tests/helpers.ts +204 -0
  358. package/tests/oracle-project.test.ts +63 -0
  359. package/tests/oracle-standalone.test.ts +48 -0
  360. package/tests/prompt-rendering.test.ts +66 -0
  361. package/tests/release-workflow.test.ts +133 -0
@@ -0,0 +1,116 @@
1
+ ### Qwen-Coder
2
+ - [ ] Qwen2.5-Coder-32B-Instruct
3
+ - HF: Qwen/Qwen2.5-Coder-32B-Instruct
4
+ - Hardware:
5
+ - 1x H100/H200
6
+ - --tool-call-parser hermes --enable-auto-tool-choice
7
+ - 2x H100/H200
8
+ - --tensor-parallel-size 2 --tool-call-parser hermes --enable-auto-tool-choice
9
+ - Notes: Good balance of size and performance. Single GPU capable.
10
+ - [ ] Qwen3-Coder-480B-A35B-Instruct (BF16)
11
+ - HF: Qwen/Qwen3-Coder-480B-A35B-Instruct
12
+ - Hardware:
13
+ - 8x H200/H20
14
+ - --tensor-parallel-size 8 --max-model-len 32000 --enable-auto-tool-choice --tool-call-parser qwen3_coder
15
+ - Notes: Cannot serve full 262K context on single node. Reduce max-model-len or increase gpu-memory-utilization.
16
+ - [ ] Qwen3-Coder-480B-A35B-Instruct-FP8
17
+ - HF: Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
18
+ - Hardware:
19
+ - 8x H200/H20
20
+ - --max-model-len 131072 --enable-expert-parallel --data-parallel-size 8 --enable-auto-tool-choice --tool-call-parser qwen3_coder
21
+ - Env: VLLM_USE_DEEP_GEMM=1
22
+ - Notes: Use data-parallel mode (not tensor-parallel) to avoid weight quantization errors. DeepGEMM recommended.
23
+ - [ ] Qwen3-Coder-30B-A3B-Instruct (BF16)
24
+ - HF: Qwen/Qwen3-Coder-30B-A3B-Instruct
25
+ - Hardware:
26
+ - 1x H100/H200
27
+ - --enable-auto-tool-choice --tool-call-parser qwen3_coder
28
+ - Notes: Fits comfortably on single GPU. ~60GB model weight.
29
+ - 2x H100/H200
30
+ - --tensor-parallel-size 2 --enable-auto-tool-choice --tool-call-parser qwen3_coder
31
+ - Notes: For higher throughput/longer context.
32
+ - [ ] Qwen3-Coder-30B-A3B-Instruct-FP8
33
+ - HF: Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8
34
+ - Hardware:
35
+ - 1x H100/H200
36
+ - --enable-auto-tool-choice --tool-call-parser qwen3_coder
37
+ - Env: VLLM_USE_DEEP_GEMM=1
38
+ - Notes: FP8 quantized, ~30GB model weight. Excellent for single GPU deployment.
39
+
40
+ ### GPT-OSS
41
+ - Notes: Requires vLLM 0.10.1+gptoss. Built-in tools via /v1/responses endpoint (browsing, Python). Function calling not yet supported. --async-scheduling recommended for higher perf (not compatible with structured output).
42
+ - [ ] GPT-OSS-20B
43
+ - HF: openai/gpt-oss-20b
44
+ - Hardware:
45
+ - 1x H100/H200
46
+ - --async-scheduling
47
+ - 1x B200
48
+ - --async-scheduling
49
+ - Env: VLLM_USE_TRTLLM_ATTENTION=1 VLLM_USE_TRTLLM_DECODE_ATTENTION=1 VLLM_USE_TRTLLM_CONTEXT_ATTENTION=1 VLLM_USE_FLASHINFER_MXFP4_MOE=1
50
+ - [ ] GPT-OSS-120B
51
+ - HF: openai/gpt-oss-120b
52
+ - Hardware:
53
+ - 1x H100/H200
54
+ - --async-scheduling
55
+ - Notes: Needs --gpu-memory-utilization 0.95 --max-num-batched-tokens 1024 to avoid OOM
56
+ - 2x H100/H200
57
+ - --tensor-parallel-size 2 --async-scheduling
58
+ - Notes: Set --gpu-memory-utilization <0.95 to avoid OOM
59
+ - 4x H100/H200
60
+ - --tensor-parallel-size 4 --async-scheduling
61
+ - 8x H100/H200
62
+ - --tensor-parallel-size 8 --async-scheduling --max-model-len 131072 --max-num-batched-tokens 10240 --max-num-seqs 128 --gpu-memory-utilization 0.85 --no-enable-prefix-caching
63
+ - 1x B200
64
+ - --async-scheduling
65
+ - Env: VLLM_USE_TRTLLM_ATTENTION=1 VLLM_USE_TRTLLM_DECODE_ATTENTION=1 VLLM_USE_TRTLLM_CONTEXT_ATTENTION=1 VLLM_USE_FLASHINFER_MXFP4_MOE=1
66
+ - 2x B200
67
+ - --tensor-parallel-size 2 --async-scheduling
68
+ - Env: VLLM_USE_TRTLLM_ATTENTION=1 VLLM_USE_TRTLLM_DECODE_ATTENTION=1 VLLM_USE_TRTLLM_CONTEXT_ATTENTION=1 VLLM_USE_FLASHINFER_MXFP4_MOE=1
69
+
70
+ ### GLM-4.5
71
+ - Notes: Listed configs support reduced context. For full 128K context, double the GPU count. Models default to thinking mode (disable with API param).
72
+ - [ ] GLM-4.5 (BF16)
73
+ - HF: zai-org/GLM-4.5
74
+ - Hardware:
75
+ - 16x H100
76
+ - --tensor-parallel-size 16 --tool-call-parser glm45 --reasoning-parser glm45 --enable-auto-tool-choice
77
+ - 8x H200
78
+ - --tensor-parallel-size 8 --tool-call-parser glm45 --reasoning-parser glm45 --enable-auto-tool-choice
79
+ - Notes: On 8x H100, may need --cpu-offload-gb 16 to avoid OOM. For full 128K: needs 32x H100 or 16x H200.
80
+ - [ ] GLM-4.5-FP8
81
+ - HF: zai-org/GLM-4.5-FP8
82
+ - Hardware:
83
+ - 8x H100
84
+ - --tensor-parallel-size 8 --tool-call-parser glm45 --reasoning-parser glm45 --enable-auto-tool-choice
85
+ - 4x H200
86
+ - --tensor-parallel-size 4 --tool-call-parser glm45 --reasoning-parser glm45 --enable-auto-tool-choice
87
+ - Notes: For full 128K context: needs 16x H100 or 8x H200.
88
+ - [ ] GLM-4.5-Air (BF16)
89
+ - HF: zai-org/GLM-4.5-Air
90
+ - Hardware:
91
+ - 4x H100
92
+ - --tensor-parallel-size 4 --tool-call-parser glm45 --reasoning-parser glm45 --enable-auto-tool-choice
93
+ - 2x H200
94
+ - --tensor-parallel-size 2 --tool-call-parser glm45 --reasoning-parser glm45 --enable-auto-tool-choice
95
+ - Notes: For full 128K context: needs 8x H100 or 4x H200.
96
+ - [ ] GLM-4.5-Air-FP8
97
+ - HF: zai-org/GLM-4.5-Air-FP8
98
+ - Hardware:
99
+ - 2x H100
100
+ - --tensor-parallel-size 2 --tool-call-parser glm45 --reasoning-parser glm45 --enable-auto-tool-choice
101
+ - 1x H200
102
+ - --tensor-parallel-size 1 --tool-call-parser glm45 --reasoning-parser glm45 --enable-auto-tool-choice
103
+ - Notes: For full 128K context: needs 4x H100 or 2x H200.
104
+
105
+ ### Kimi
106
+ - Notes: Requires vLLM v0.10.0rc1+. Minimum 16 GPUs for FP8 with 128k context. Reuses DeepSeekV3 architecture with model_type="kimi_k2".
107
+ - [ ] Kimi-K2-Instruct
108
+ - HF: moonshotai/Kimi-K2-Instruct
109
+ - Hardware:
110
+ - 16x H200/H20
111
+ - --tensor-parallel-size 16 --trust-remote-code --enable-auto-tool-choice --tool-call-parser kimi_k2
112
+ - Notes: Pure TP mode. For >16 GPUs, combine with pipeline-parallelism.
113
+ - 16x H200/H20 (DP+EP mode)
114
+ - --data-parallel-size 16 --data-parallel-size-local 8 --enable-expert-parallel --max-num-batched-tokens 8192 --max-num-seqs 256 --gpu-memory-utilization 0.85 --trust-remote-code --enable-auto-tool-choice --tool-call-parser kimi_k2
115
+ - Notes: Data parallel + expert parallel mode for higher throughput. Requires multi-node setup with proper networking.
116
+
@@ -0,0 +1,166 @@
1
+ ## Pi
2
+
3
+ Pi automates vLLM deployment on GPU pods from DataCrunch, Vast.ai, Prime Intellect, RunPod (or any Ubuntu machine with NVIDIA GPUs). It manages multiple concurrent model deployments via separate vLLM instances, each accessible through the OpenAI API protocol with API key authentication.
4
+
5
+ Pods are treated as ephemeral - spin up when needed, tear down when done. To avoid re-downloading models (30+ minutes for 100GB+ models), pi uses persistent network volumes for model storage that can be shared across pods on the same provider. This minimizes both cost (only pay for active compute) and setup time (models already cached).
6
+
7
+ ## Usage
8
+
9
+ ### Pods
10
+ ```bash
11
+ pi pods setup dc1 "ssh root@1.2.3.4" --mount "mount -t nfs..." # Setup pod (requires HF_TOKEN, PI_API_KEY env vars)
12
+ pi pods # List all pods (* = active)
13
+ pi pods active dc2 # Switch active pod
14
+ pi pods remove dc1 # Remove pod
15
+ ```
16
+
17
+ ### Models
18
+ ```bash
19
+ pi start Qwen/Qwen2.5-72B-Instruct --name qwen72b # Known model - pi handles vLLM args
20
+ pi start some/unknown-model --name mymodel --vllm --tensor-parallel-size 4 --max-model-len 32768 # Custom vLLM args
21
+ pi list # List running models with ports
22
+ pi stop qwen72b # Stop model
23
+ pi logs qwen72b # View model logs
24
+ ```
25
+
26
+ For known models, pi automatically configures appropriate vLLM arguments from model documentation based on the hardware of the pod. For unknown models or custom configurations, pass vLLM args after `--vllm`.
27
+
28
+ ## Pod management
29
+
30
+ Pi manages GPU pods from various providers (DataCrunch, Vast.ai, Prime Intellect, RunPod) as ephemeral compute resources. Users manually create pods via provider dashboards, then register them with pi for automated setup and management.
31
+
32
+ Key capabilities:
33
+ - **Pod setup**: Transform bare Ubuntu/Debian machines into vLLM-ready environments in ~2 minutes
34
+ - **Model caching**: Optional persistent storage shared by pods to avoid re-downloading 100GB+ models
35
+ - **Multi-pod management**: Register multiple pods, switch between them, maintain different environments
36
+
37
+ ### Pod setup
38
+
39
+ When a user creates a fresh pod on a provider, they register it with pi using the SSH command from the provider:
40
+
41
+ ```bash
42
+ pi pods setup dc1 "ssh root@1.2.3.4" --mount "mount -t nfs..."
43
+ ```
44
+
45
+ This copies and executes `pod_setup.sh` which:
46
+ 1. Detects GPUs via `nvidia-smi` and stores count/memory in local config
47
+ 2. Installs CUDA toolkit matching the driver version
48
+ 3. Creates Python environment
49
+ - Installs uv and Python 3.12
50
+ - Creates venv at ~/venv with PyTorch (--torch-backend=auto)
51
+ - Installs vLLM (model-specific versions when needed)
52
+ - Installs FlashInfer (builds from source if required)
53
+ - Installs huggingface-hub (for model downloads)
54
+ - Installs hf-transfer (for accelerated downloads)
55
+ 4. Mounts persistent storage if provided
56
+ - Symlinks to ~/.cache/huggingface for model caching
57
+ 5. Configures environment variables persistently
58
+
59
+ Required environment variables:
60
+ - `HF_TOKEN`: HuggingFace token for model downloads
61
+ - `PI_API_KEY`: API key for securing vLLM endpoints
62
+
63
+ ### Model caching
64
+
65
+ Models can be 100GB+ and take 30+ minutes to download. The `--mount` flag enables persistent model caching:
66
+
67
+ - **DataCrunch**: NFS shared filesystems, mountable across multiple running pods in same region
68
+ - **RunPod**: Network volumes persist independently but cannot be shared between running pods
69
+ - **Vast.ai**: Volumes locked to specific machine - no sharing
70
+ - **Prime Intellect**: No persistent storage documented
71
+
72
+ Without `--mount`, models download to pod-local storage and are lost on termination.
73
+
74
+ ### Multi-pod management
75
+
76
+ Users can register multiple pods and switch between them:
77
+
78
+ ```bash
79
+ pi pods # List all pods (* = active)
80
+ pi pods active dc2 # Switch active pod
81
+ pi pods remove dc1 # Remove pod from local config but doesn't destroy pod remotely.
82
+ ```
83
+
84
+ All model commands (`pi start`, `pi stop`, etc.) target the active pod, unless `--pod <podname>` is given, which overrides the active pod for that command.
85
+
86
+ ## Model deployment
87
+
88
+ Pi uses direct SSH commands to manage vLLM instances on pods. No remote manager component is needed - everything is controlled from the local pi CLI.
89
+
90
+ ### Architecture
91
+ The pi CLI maintains all state locally in `~/.pi/pods.json`:
92
+ ```json
93
+ {
94
+ "pods": {
95
+ "dc1": {
96
+ "ssh": "ssh root@1.2.3.4",
97
+ "gpus": [
98
+ {"id": 0, "name": "H100", "memory": "80GB"},
99
+ {"id": 1, "name": "H100", "memory": "80GB"}
100
+ ],
101
+ "models": {
102
+ "qwen": {
103
+ "model": "Qwen/Qwen2.5-72B",
104
+ "port": 8001,
105
+ "gpu": "0",
106
+ "pid": 12345
107
+ }
108
+ }
109
+ }
110
+ },
111
+ "active": "dc1"
112
+ }
113
+ ```
114
+
115
+ The location of the pi config dir can also be specified via the `PI_CONFIG_DIR` env var, e.g. for testing.
116
+
117
+ Pods are assumed to be fully managed by pi - no other processes compete for ports or GPUs.
118
+
119
+ ### Starting models
120
+ When user runs `pi start Qwen/Qwen2.5-72B --name qwen`:
121
+ 1. CLI determines next available port (starting from 8001)
122
+ 2. Selects GPU (round-robin based on stored GPU info)
123
+ 3. Downloads model if not cached:
124
+ - Sets `HF_HUB_ENABLE_HF_TRANSFER=1` for fast downloads
125
+ - Runs via SSH with output piped to local terminal
126
+ - Ctrl+C cancels download and returns control
127
+ 4. Builds vLLM command with appropriate args and PI_API_KEY
128
+ 5. Executes via SSH: `ssh pod "nohup vllm serve ... > ~/.vllm_logs/qwen.log 2>&1 & echo $!"`
129
+ 6. Waits for vLLM to be ready (checks health endpoint)
130
+ 7. On success: stores port, GPU, PID in local state
131
+ 8. On failure: shows exact error from vLLM logs, doesn't save to config
132
+
133
+ ### Managing models
134
+ - **List**: Show models from local state, optionally verify PIDs still running
135
+ - **Stop**: SSH to kill process by PID
136
+ - **Logs**: SSH to tail -f log files (Ctrl+C stops tailing, doesn't kill vLLM)
137
+
138
+ ### Error handling
139
+ - **SSH failures**: Prompt user to check connection or remove pod from config
140
+ - **Stale state**: Commands that fail with "process not found" auto-clean local state
141
+ - **Setup failures**: Ctrl+C during setup kills remote script and exits cleanly
142
+
143
+ ### Testing models
144
+ The `pi prompt` command provides a quick way to test deployed models:
145
+ ```bash
146
+ pi prompt qwen "What is 2+2?" # Simple prompt
147
+ pi prompt qwen "Read file.txt and summarize" # Uses built-in tools
148
+ ```
149
+
150
+ Built-in tools for agentic testing:
151
+ - `ls(path, ignore?)`: List files and directories at path, with optional ignore patterns
152
+ - `read(file_path, offset?, limit?)`: Read file contents with optional line offset/limit
153
+ - `glob(pattern, path?)`: Find files matching glob pattern (e.g., "**/*.py", "src/**/*.ts")
154
+ - `rg(args)`: Run ripgrep with any arguments (e.g., "pattern -t py -C 3", "TODO --type-not test")
155
+
156
+ The provided prompt will be augmented with info on the current local working directory. File tools expect absolute paths.
157
+
158
+ This allows testing basic agent capabilities without external tool configuration.
159
+
160
+ `prompt` is implemented using the latest OpenAI SDK for NodeJS. It outputs thinking content, tool calls and results, and normal assistant messages.
161
+
162
+ ## Models
163
+ We want to support these models specifically, with alternative models being marked as "possibly works". This list will be updated with new models regularly. A checked
164
+ box means "supported".
165
+
166
+ See [models.md](./models.md) for a list of models, their HW reqs, vLLM args and notes, we want to support out of the box with a simple `pi start <model-name> --name <local-name>`
@@ -0,0 +1,132 @@
1
+ # Qwen3-Coder Usage Guide
2
+
3
+ [Qwen3-Coder](https://github.com/QwenLM/Qwen3-Coder) is an advanced large language model created by the Qwen team from Alibaba Cloud. vLLM already supports Qwen3-Coder, and `tool-call` functionality will be available in vLLM v0.10.0 and higher You can install vLLM with `tool-call` support using the following method:
4
+
5
+ ## Installing vLLM
6
+
7
+ ```bash
8
+ uv venv
9
+ source .venv/bin/activate
10
+ uv pip install -U vllm --torch-backend auto
11
+ ```
12
+
13
+ ## Launching Qwen3-Coder with vLLM
14
+
15
+ ### Serving on 8xH200 (or H20) GPUs (141GB × 8)
16
+
17
+ **BF16 Model**
18
+
19
+ ```bash
20
+ vllm serve Qwen/Qwen3-Coder-480B-A35B-Instruct \
21
+ --tensor-parallel-size 8 \
22
+ --max-model-len 32000 \
23
+ --enable-auto-tool-choice \
24
+ --tool-call-parser qwen3_coder
25
+ ```
26
+
27
+ **FP8 Model**
28
+
29
+ ```bash
30
+ VLLM_USE_DEEP_GEMM=1 vllm serve Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 \
31
+ --max-model-len 131072 \
32
+ --enable-expert-parallel \
33
+ --data-parallel-size 8 \
34
+ --enable-auto-tool-choice \
35
+ --tool-call-parser qwen3_coder
36
+ ```
37
+
38
+ ## Performance Metrics
39
+
40
+ ### Evaluation
41
+ We launched `Qwen3-Coder-480B-A35B-Instruct-FP8` using vLLM and evaluated its performance using [EvalPlus](https://github.com/evalplus/evalplus). The results are displayed below:
42
+
43
+ | Dataset | Test Type | Pass@1 Score |
44
+ |-----------|-----------|--------------|
45
+ | HumanEval | Base tests | 0.939 |
46
+ | HumanEval+ | Base + extra tests | 0.902 |
47
+ | MBPP | Base tests | 0.918 |
48
+ | MBPP+ | Base + extra tests | 0.794 |
49
+
50
+ ### Benchmarking
51
+ We used the following script to benchmark `Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8`
52
+
53
+ ```bash
54
+ vllm bench serve \
55
+ --backend vllm \
56
+ --model Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 \
57
+ --endpoint /v1/completions \
58
+ --dataset-name random \
59
+ --random-input 2048 \
60
+ --random-output 1024 \
61
+ --max-concurrency 10 \
62
+ --num-prompt 100 \
63
+ ```
64
+ If successful, you will see the following output.
65
+
66
+ ```shell
67
+ ============ Serving Benchmark Result ============
68
+ Successful requests: 100
69
+ Benchmark duration (s): 776.49
70
+ Total input tokens: 204169
71
+ Total generated tokens: 102400
72
+ Request throughput (req/s): 0.13
73
+ Output token throughput (tok/s): 131.88
74
+ Total Token throughput (tok/s): 394.81
75
+ ---------------Time to First Token----------------
76
+ Mean TTFT (ms): 7639.31
77
+ Median TTFT (ms): 6935.71
78
+ P99 TTFT (ms): 13766.68
79
+ -----Time per Output Token (excl. 1st token)------
80
+ Mean TPOT (ms): 68.43
81
+ Median TPOT (ms): 67.23
82
+ P99 TPOT (ms): 72.14
83
+ ---------------Inter-token Latency----------------
84
+ Mean ITL (ms): 68.43
85
+ Median ITL (ms): 66.34
86
+ P99 ITL (ms): 69.38
87
+ ==================================================
88
+
89
+ ```
90
+
91
+
92
+ ## Using Tips
93
+
94
+ ### BF16 Models
95
+ - **Context Length Limitation**: A single H20 node cannot serve the original context length (262144). You can reduce the `max-model-len` or increase `gpu-memory-utilization` to work within memory constraints.
96
+
97
+ ### FP8 Models
98
+ - **Context Length Limitation**: A single H20 node cannot serve the original context length (262144). You can reduce the `max-model-len` or increase `gpu-memory-utilization` to work within memory constraints.
99
+ - **DeepGEMM Usage**: To use [DeepGEMM](https://github.com/deepseek-ai/DeepGEMM), set `VLLM_USE_DEEP_GEMM=1`. Follow the [setup instructions](https://github.com/vllm-project/vllm/blob/main/benchmarks/kernels/deepgemm/README.md#setup) to install it.
100
+ - **Tensor Parallelism Issue**: When using `tensor-parallel-size 8`, the following failures are expected. Switch to data-parallel mode using `--data-parallel-size`.
101
+ - **Additional Resources**: Refer to the [Data Parallel Deployment documentation](https://docs.vllm.ai/en/latest/serving/data_parallel_deployment.html) for more parallelism groups.
102
+
103
+ ```shell
104
+ ERROR [multiproc_executor.py:511] File "/vllm/vllm/model_executor/models/qwen3_moe.py", line 336, in <lambda>
105
+ ERROR [multiproc_executor.py:511] lambda prefix: Qwen3MoeDecoderLayer(config=config,
106
+ ERROR [multiproc_executor.py:511] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
107
+ ERROR [multiproc_executor.py:511] File "/vllm/vllm/model_executor/models/qwen3_moe.py", line 278, in __init__
108
+ ERROR [multiproc_executor.py:511] self.mlp = Qwen3MoeSparseMoeBlock(config=config,
109
+ ERROR [multiproc_executor.py:511] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
110
+ ERROR [multiproc_executor.py:511] File "/vllm/vllm/model_executor/models/qwen3_moe.py", line 113, in __init__
111
+ ERROR [multiproc_executor.py:511] self.experts = FusedMoE(num_experts=config.num_experts,
112
+ ERROR [multiproc_executor.py:511] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
113
+ ERROR [multiproc_executor.py:511] File "/vllm/vllm/model_executor/layers/fused_moe/layer.py", line 773, in __init__
114
+ ERROR [multiproc_executor.py:511] self.quant_method.create_weights(layer=self, **moe_quant_params)
115
+ ERROR [multiproc_executor.py:511] File "/vllm/vllm/model_executor/layers/quantization/fp8.py", line 573, in create_weights
116
+ ERROR [multiproc_executor.py:511] raise ValueError(
117
+ ERROR [multiproc_executor.py:511] ValueError: The output_size of gate's and up's weight = 320 is not divisible by weight quantization block_n = 128.
118
+ ```
119
+
120
+ ### Tool Calling
121
+ - **Enable Tool Calls**: Add `--tool-call-parser qwen3_coder` to enable tool call parsing functionality, please refer to: [tool_calling](https://docs.vllm.ai/en/latest/features/tool_calling.html)
122
+
123
+ ## Roadmap
124
+
125
+ - [x] Add benchmark results
126
+
127
+
128
+ ## Additional Resources
129
+
130
+ - [EvalPlus](https://github.com/evalplus/evalplus)
131
+ - [Qwen3-Coder](https://github.com/QwenLM/Qwen3-Coder)
132
+ - [vLLM Documentation](https://docs.vllm.ai/)
Binary file