hipengine 0.1.0__tar.gz → 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (665) hide show
  1. hipengine-0.2.0/.gitattributes +10 -0
  2. hipengine-0.2.0/.github/workflows/publish.yml +103 -0
  3. {hipengine-0.1.0 → hipengine-0.2.0}/AGENTS.md +2 -0
  4. hipengine-0.2.0/CHANGELOG.md +112 -0
  5. {hipengine-0.1.0 → hipengine-0.2.0}/PKG-INFO +139 -46
  6. {hipengine-0.1.0 → hipengine-0.2.0}/README.md +138 -45
  7. {hipengine-0.1.0 → hipengine-0.2.0}/WORKLOG.md +12684 -2807
  8. hipengine-0.2.0/benchmarks/7900XTX.md +166 -0
  9. hipengine-0.2.0/benchmarks/CHANGELOG.md +203 -0
  10. hipengine-0.2.0/benchmarks/MTP.md +115 -0
  11. hipengine-0.2.0/benchmarks/README.md +415 -0
  12. hipengine-0.2.0/benchmarks/W7900.md +117 -0
  13. hipengine-0.2.0/benchmarks/configs/llamacpp-mtp-qwen36-27b.json +53 -0
  14. hipengine-0.2.0/benchmarks/prompts/mtpbench-code-general-ja.jsonl +10 -0
  15. hipengine-0.2.0/benchmarks/results/2026-05-16-hipengine-gguf-q4k-pack8-bf16out-diagnostic.json +102 -0
  16. hipengine-0.2.0/benchmarks/results/2026-05-16-hipengine-gguf-q4k-pack8-gemv-diagnostic.json +72 -0
  17. hipengine-0.2.0/benchmarks/results/2026-05-16-hipengine-gguf-qwen35-e2e-correctness-diagnostic.json +71 -0
  18. hipengine-0.2.0/benchmarks/results/2026-05-16-hipengine-gguf-vs-paro-diagnostic.json +295 -0
  19. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-aotriton-v3-prefill-diagnostic.json +101 -0
  20. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-bulk-prefill-q4km-accepted.json +258 -0
  21. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-decode-graph-replay-diagnostic.json +100 -0
  22. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-full-attn-gpu-prelude-diagnostic.json +122 -0
  23. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-local-quant-coverage-diagnostic.json +145 -0
  24. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-prefill-projection-diagnostic.json +157 -0
  25. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-q4km-parity-benchmark-diagnostic.json +814 -0
  26. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-resident-session-diagnostic.json +106 -0
  27. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-512-128-paro-blocker-profile.json +685 -0
  28. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-bench-paro-comparison-diagnostic.json +370 -0
  29. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-bulk-moe-prefill-diagnostic.json +482 -0
  30. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-bulk-parity-diagnostic.json +261 -0
  31. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-decode-pack8-raw-partial.json +546 -0
  32. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-decode-profile-diagnostic.json +2013 -0
  33. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-expert-pack8-sidecar-diagnostic.json +240 -0
  34. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-fast-bulk-default-promoted-diagnostic.json +169 -0
  35. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-full-attn-parity-fixed-diagnostic.json +118 -0
  36. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-intake-diagnostic.json +388 -0
  37. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-linear-recurrent-parity-fixed-diagnostic.json +149 -0
  38. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-native-attention-bulk-moe-diagnostic.json +383 -0
  39. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-public-generate-smoke.json +135 -0
  40. hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-selected-device-experts-diagnostic.json +457 -0
  41. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-gfx1100-qwen36-27b-paro-diagnostic.json +554 -0
  42. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-p8_2-dense-q4k-wmma-prefill-diagnostic.json +713 -0
  43. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen35-int8-kv-128k-quality-perf-diagnostic.json +1062 -0
  44. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen35-int8-kv-256k-capacity-blocked.json +334 -0
  45. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen35-int8-kv-256k-single-buffer-capacity-diagnostic.json +327 -0
  46. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen35-int8-kv-aotriton-query-reuse-diagnostic.json +443 -0
  47. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen35-int8-kv-scratch-release-diagnostic.json +219 -0
  48. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p8-compact-moe-wmma-accepted.json +1640 -0
  49. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_a3-gdn-k2-chain-accepted.json +1986 -0
  50. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c1-wmma-tile-sweep-blocked.json +1239 -0
  51. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c10-combined-gap-analysis.json +309 -0
  52. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c11-hot-expert-final-blocked.json +334 -0
  53. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c3-selected-moe-profile.json +421 -0
  54. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c4-q4-hot-fulltile-v1-rejected.json +53 -0
  55. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c5-q4-sidemeta-v1-rejected.json +32 -0
  56. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c7-q5-opt-v1-rejected.json +40 -0
  57. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c8-q6-retain-legacy.json +56 -0
  58. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c9-tail-no-padding-not-retained.json +67 -0
  59. hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-prefill-q8-wmma-p8_1.json +371 -0
  60. hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_b7-decode-gemv-rejected.json +880 -0
  61. hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_c12-q4t16-repack-design.json +102 -0
  62. hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_c13-q4t16-materializer.json +48 -0
  63. hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_e2-e2e-correctness-rejected.json +303 -0
  64. hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_h1-fastpath-safety-correctness-accepted.json +333 -0
  65. hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_h2-decode-repack-design.json +50 -0
  66. hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_h3-t16-512x128-bench.json +1279 -0
  67. hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_h3-t16-e2e-correctness-accepted.json +334 -0
  68. hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_h3-t16-rocprof-512x16-summary.json +1699 -0
  69. hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-packed-int8-kv-readme-memory-diagnostic.json +225 -0
  70. hipengine-0.2.0/benchmarks/results/2026-05-19-llamacpp-mtp-qwen36-27b-diagnostic.json +537 -0
  71. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p10-b5-p9-e2e-gate-blocked.json +351 -0
  72. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p10-b6-acceptance-blocked.json +98 -0
  73. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p10-wave1-blocked.json +121 -0
  74. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_c14-q4t16-selected-wmma-prototype.json +60 -0
  75. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_c15-q4t16-replay-rejected.json +100 -0
  76. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_c16-selected-moe-alternatives.json +135 -0
  77. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_c17-no-q4-redesign-blocked.json +74 -0
  78. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d1-router-split-coop.json +82 -0
  79. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d10-q8t16-dual-split-64.json +110 -0
  80. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d11-rejected-q8t16-shared-silu.json +109 -0
  81. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d15-dense-dual-alpha-beta.json +115 -0
  82. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d16-q8t16-f32-ssm-out.json +121 -0
  83. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d17-key-bf16-rope.json +127 -0
  84. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d18-splitk-gqa-gate.json +130 -0
  85. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d2-bf16-key-rope-rejected.json +90 -0
  86. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d4-q4t16-silu-decode.json +94 -0
  87. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d6-q8t16-pair-dispatch.json +82 -0
  88. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d7-q8t16-qkv-gate-pair.json +74 -0
  89. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d8-q8t16-dcache-rejected.json +62 -0
  90. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d9-q8t16-triple-qkv.json +80 -0
  91. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_g1-final-acceptance-blocked.json +183 -0
  92. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_h3-rejected-attn-gate-fusion.json +51 -0
  93. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_h3-rejected-attn128.json +57 -0
  94. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_h3-rejected-q4t16-silu256.json +50 -0
  95. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_h3-rejected-q5q6-direct-probes.json +78 -0
  96. hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_h3-rejected-q6dense-dpreload.json +50 -0
  97. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-aotriton-v2-v3-diagnostic.json +155 -0
  98. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-d10-splitk-rocprof.json +472 -0
  99. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-d11-comparison-review.json +191 -0
  100. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-d8-splitk-decode.json +62 -0
  101. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-d9-splitk-sweep.json +102 -0
  102. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-decode-repack-residency-audit.json +96 -0
  103. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-long-context-chunked-smoke.json +111 -0
  104. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-long-context-preflight-blocked.json +65 -0
  105. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-memory-decode-pass-review.json +94 -0
  106. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-no-prefill-scratch-kv.json +44 -0
  107. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-q8-t16-scale-broadcast-rejected.json +55 -0
  108. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-r0-rocprof-baseline.json +1013 -0
  109. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-r1-post-x1-rocprof.json +509 -0
  110. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-retained-safe-mode.json +62 -0
  111. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-selected-moe-t16-launchbounds-rejected.json +79 -0
  112. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-task39-32k-smoke.json +109 -0
  113. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-task40-prefill-vs-paro-diagnosis.json +233 -0
  114. hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-x1-correctness-plus-x2-wmma-blocker.json +176 -0
  115. hipengine-0.2.0/benchmarks/results/2026-05-21-local-rx7900xtx-gguf-vs-paro-memory-comparison.json +403 -0
  116. hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-after-memory-decode-pass-review.json +453 -0
  117. hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-direct-selected-moe-c1-4k128-accepted.json +69 -0
  118. hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-final-gate-4k128-accepted.json +134 -0
  119. hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-q8-t16-decode-probes-rejected.json +65 -0
  120. hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-q8-t16-second-pass-rejected.json +68 -0
  121. hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-router256-4k128-accepted.json +69 -0
  122. hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-selected-moe-down64-rejected.json +70 -0
  123. hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-selected-moe-t16-qk256-rejected.json +37 -0
  124. hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-small-kernel-second-pass-rejected.json +110 -0
  125. hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-small-kernel-third-pass-rejected.json +102 -0
  126. hipengine-0.2.0/benchmarks/results/2026-05-23-hipengine-qwen36-35b-a3b-q4ks-persistent-session-w7900-gap-review.json +525 -0
  127. hipengine-0.2.0/benchmarks/results/2026-05-23-hipengine-qwen36-35b-a3b-q4ks-w7900-cold-start-diagnostic.json +144 -0
  128. hipengine-0.2.0/benchmarks/results/2026-05-23-hipengine-qwen36-35b-a3b-q4ks-w7900-readme-sweep-accepted.json +200 -0
  129. hipengine-0.2.0/benchmarks/results/2026-05-23-hipengine-qwen36-35b-a3b-q4ks-w7900-therock713-diagnostic.json +136 -0
  130. hipengine-0.2.0/benchmarks/results/2026-05-23-hipengine-w7900-therock713-gguf-paro-512-4k-spot-diagnostic.json +186 -0
  131. hipengine-0.2.0/benchmarks/results/2026-05-23-paro-512-prefill-workspace-overlap-rootcause-diagnostic.json +244 -0
  132. hipengine-0.2.0/benchmarks/results/2026-05-23-paro-prefill-workspace-overlap-threshold-diagnostic.json +271 -0
  133. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-1024-128.json +495 -0
  134. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-131072-128-blocked.json +22 -0
  135. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-32768-128.json +495 -0
  136. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-4096-128.json +495 -0
  137. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-512-128-rerun.json +479 -0
  138. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-65536-128-blocked.json +22 -0
  139. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4ks-4096-128.json +495 -0
  140. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-1024-128.json +6976 -0
  141. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-131072-128.json +6984 -0
  142. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-32768-128.json +6984 -0
  143. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-4096-128.json +6984 -0
  144. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-512-128.json +6976 -0
  145. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-65536-128.json +6984 -0
  146. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-int8kv-131072-128.json +11784 -0
  147. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-int8kv-65536-128.json +11784 -0
  148. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-llamacpp-hip-q4km-f16kv-sweep.json +1395 -0
  149. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-llamacpp-hip-q4km-q8kv-maxctx.json +499 -0
  150. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-llamacpp-vulkan-q4km-f16kv-sweep.json +1395 -0
  151. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-llamacpp-vulkan-q4km-q8kv-maxctx.json +499 -0
  152. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-paro-v011-current-regression-check.json +62 -0
  153. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-rocm7130423-current-head-paro-512-4k-check.json +176 -0
  154. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-rocm7130423-pure-current-head-paro-512-4k-check.json +196 -0
  155. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-rocm7130423-pure-packed-qwen36-paro-512-4k-check.json +395 -0
  156. hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-rocm714-current-head-paro-512-4k-check.json +169 -0
  157. hipengine-0.2.0/benchmarks/results/2026-05-23-w7900-hipengine-therock713-paro-gguf-sweep-diagnostic.json +352 -0
  158. hipengine-0.2.0/benchmarks/results/2026-05-25-w7900-hipengine-gguf-q4ks-readme-persistent-5run.json +2768 -0
  159. hipengine-0.2.0/benchmarks/results/2026-05-25-w7900-hipengine-paro-readme-persistent-5run.json +2947 -0
  160. hipengine-0.2.0/benchmarks/results/2026-05-25-w7900-hipengine-readme-persistent-5run-diagnostic.json +313 -0
  161. {hipengine-0.1.0 → hipengine-0.2.0}/docs/BENCHMARK.md +61 -0
  162. hipengine-0.2.0/docs/ENVS.md +156 -0
  163. hipengine-0.2.0/docs/GGUF.md +3023 -0
  164. hipengine-0.2.0/docs/GGUF_DECODE_REPACK.md +333 -0
  165. {hipengine-0.1.0 → hipengine-0.2.0}/docs/KERNELS.md +121 -10
  166. hipengine-0.2.0/docs/KVCACHE.md +448 -0
  167. hipengine-0.2.0/docs/LESSONS-LEARNED.md +1068 -0
  168. hipengine-0.2.0/docs/OPTIMIZE-DENSE.md +423 -0
  169. {hipengine-0.1.0 → hipengine-0.2.0}/docs/PUBLISH.md +6 -1
  170. {hipengine-0.1.0 → hipengine-0.2.0}/docs/README.md +5 -2
  171. hipengine-0.2.0/docs/RELAXED.md +651 -0
  172. hipengine-0.2.0/docs/ROOFLINE-gfx1151.md +805 -0
  173. hipengine-0.2.0/docs/SPECULATIVE-DECODE.md +1465 -0
  174. {hipengine-0.1.0 → hipengine-0.2.0}/docs/TESTING.md +55 -1
  175. {hipengine-0.1.0 → hipengine-0.2.0}/docs/source_lineage.json +7 -7
  176. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/dtype.py +2 -0
  177. hipengine-0.2.0/hipengine/dispatch/__init__.py +38 -0
  178. hipengine-0.2.0/hipengine/dispatch/kv.py +259 -0
  179. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/generation/__init__.py +1 -0
  180. hipengine-0.2.0/hipengine/generation/qwen35_gguf.py +101 -0
  181. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/generation/qwen35_paro.py +24 -3
  182. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/generation/registry.py +3 -0
  183. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/cpu_reference/__init__.py +24 -0
  184. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/cpu_reference/ops.py +383 -1
  185. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/__init__.py +16 -0
  186. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_wrap.py +5 -0
  187. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/paged_attn_decode.hip +696 -20
  188. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/paged_attn_decode.py +579 -0
  189. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/paged_kv_write.hip +374 -1
  190. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/attention/paged_kv_write.py +888 -0
  191. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/fused/__init__.py +24 -0
  192. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/fused/gguf_ops.hip +621 -0
  193. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/fused/gguf_ops.py +683 -0
  194. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear/dense_gemv.hip +84 -0
  195. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear/dense_gemv.py +48 -0
  196. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear_attn/conv.hip +12 -3
  197. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear_attn/gdn.hip +259 -6
  198. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear_attn/gdn.py +93 -0
  199. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/__init__.py +4 -0
  200. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/router.hip +155 -22
  201. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/router.py +141 -0
  202. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/__init__.py +68 -0
  203. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_expert_pack8_gemv.hip +517 -0
  204. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_expert_pack8_gemv.py +274 -0
  205. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_gemv.hip +767 -0
  206. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_gemv.py +458 -0
  207. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_selected_pack8_gemv.hip +524 -0
  208. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_selected_pack8_gemv.py +344 -0
  209. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_selected_prefill.hip +800 -0
  210. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_selected_prefill.py +419 -0
  211. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_t16_selected_prefill.hip +498 -0
  212. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_t16_selected_prefill.py +289 -0
  213. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_gemv.hip +1078 -0
  214. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_gemv.py +806 -0
  215. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_pack8_gemv.hip +336 -0
  216. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_pack8_gemv.py +182 -0
  217. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_prefill.hip +649 -0
  218. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_prefill.py +319 -0
  219. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_selected_pack8_gemv.hip +423 -0
  220. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_selected_pack8_gemv.py +289 -0
  221. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_selected_prefill.hip +1400 -0
  222. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_selected_prefill.py +785 -0
  223. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_t16_selected_prefill.hip +419 -0
  224. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_t16_selected_prefill.py +314 -0
  225. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_embedding.hip +190 -0
  226. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_embedding.py +198 -0
  227. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_pack8_gemv.hip +349 -0
  228. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_pack8_gemv.py +170 -0
  229. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_t16_gemv.hip +185 -0
  230. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_t16_gemv.py +190 -0
  231. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_pack8_gemv.hip +514 -0
  232. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_pack8_gemv.py +352 -0
  233. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_prefill.hip +642 -0
  234. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_prefill.py +363 -0
  235. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_t16_gemv.hip +559 -0
  236. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_t16_gemv.py +652 -0
  237. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_t16_prefill.hip +333 -0
  238. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_t16_prefill.py +274 -0
  239. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_t16_selected_gemv.hip +919 -0
  240. hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_t16_selected_gemv.py +1068 -0
  241. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/runtime/state.py +5 -0
  242. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/registry.py +13 -0
  243. hipengine-0.2.0/hipengine/kvcache/__init__.py +30 -0
  244. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kvcache/policy.py +128 -3
  245. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kvcache/spans.py +43 -1
  246. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/llm.py +27 -6
  247. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/loading/__init__.py +90 -0
  248. hipengine-0.2.0/hipengine/loading/gguf.py +327 -0
  249. hipengine-0.2.0/hipengine/loading/qwen35_gguf.py +443 -0
  250. hipengine-0.2.0/hipengine/loading/qwen35_gguf_expert_sidecar.py +577 -0
  251. hipengine-0.2.0/hipengine/loading/qwen35_gguf_materialize.py +514 -0
  252. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/models/__init__.py +12 -1
  253. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/models/qwen35.py +58 -0
  254. hipengine-0.2.0/hipengine/quant/__init__.py +164 -0
  255. hipengine-0.2.0/hipengine/quant/gguf.py +678 -0
  256. hipengine-0.2.0/hipengine/quant/gguf_k.py +93 -0
  257. hipengine-0.2.0/hipengine/quant/gguf_q4_k.py +343 -0
  258. hipengine-0.2.0/hipengine/quant/gguf_t16.py +507 -0
  259. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/runtime/__init__.py +22 -0
  260. hipengine-0.2.0/hipengine/runtime/gguf_embedding.py +145 -0
  261. hipengine-0.2.0/hipengine/runtime/gguf_linear.py +1040 -0
  262. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/runtime/prefill.py +12 -0
  263. hipengine-0.2.0/hipengine/runtime/qwen35_gguf_runner.py +5718 -0
  264. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/runtime/qwen35_paro.py +316 -44
  265. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/runtime/qwen35_paro_runner.py +590 -50
  266. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/server/api.py +19 -0
  267. hipengine-0.2.0/hipengine/tokenization/__init__.py +5 -0
  268. hipengine-0.2.0/hipengine/tokenization/gguf.py +166 -0
  269. {hipengine-0.1.0 → hipengine-0.2.0}/pyproject.toml +1 -1
  270. hipengine-0.2.0/scripts/gguf_k_gemv_smoke.py +304 -0
  271. hipengine-0.2.0/scripts/gguf_prefill_projection_smoke.py +308 -0
  272. hipengine-0.2.0/scripts/gguf_q6_k_embedding_smoke.py +151 -0
  273. hipengine-0.2.0/scripts/inspect_gguf.py +188 -0
  274. hipengine-0.2.0/scripts/llamacpp_mtp_bench.py +518 -0
  275. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_batch_packed_prefill_correctness.py +17 -0
  276. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_batch_serial_bench.py +14 -2
  277. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_batch_serial_correctness.py +22 -0
  278. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_decode_graph_fixture_gate.py +13 -0
  279. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_e2e_correctness.py +16 -1
  280. hipengine-0.2.0/scripts/qwen35_gguf_aotriton_prefill_sweep.py +112 -0
  281. hipengine-0.2.0/scripts/qwen35_gguf_bench.py +920 -0
  282. hipengine-0.2.0/scripts/qwen35_gguf_build_expert_sidecar.py +111 -0
  283. hipengine-0.2.0/scripts/qwen35_gguf_bulk_parity.py +357 -0
  284. hipengine-0.2.0/scripts/qwen35_gguf_decode_graph_smoke.py +377 -0
  285. hipengine-0.2.0/scripts/qwen35_gguf_e2e_correctness.py +357 -0
  286. hipengine-0.2.0/scripts/qwen35_gguf_expert_pack8_smoke.py +315 -0
  287. hipengine-0.2.0/scripts/qwen35_gguf_moe_replay.py +801 -0
  288. hipengine-0.2.0/scripts/qwen35_gguf_p9_e2e_correctness.py +442 -0
  289. hipengine-0.2.0/scripts/qwen35_gguf_rocprof_summary.py +648 -0
  290. hipengine-0.2.0/scripts/qwen35_kv_e2e_fixture_gate.py +459 -0
  291. hipengine-0.2.0/scripts/qwen35_kv_int8_accuracy.py +870 -0
  292. hipengine-0.2.0/scripts/qwen35_kv_policy_args.py +80 -0
  293. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_correctness.py +39 -4
  294. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_fixture_gate.py +55 -4
  295. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_fullattn_stage_probe.py +3 -5
  296. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_stage_probe.py +2 -7
  297. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_paro_bench.py +38 -1
  298. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_paro_packed_bench.py +6 -0
  299. hipengine-0.2.0/scripts/qwen35_readme_sweep.py +574 -0
  300. hipengine-0.2.0/scripts/resolve_worklog_conflict.py +186 -0
  301. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/smoke.py +971 -2
  302. hipengine-0.2.0/tests/_gguf_synthetic_weights.py +187 -0
  303. hipengine-0.2.0/tests/conftest.py +39 -0
  304. hipengine-0.2.0/tests/fixtures/cpu_reference/kv_int8_dequant_per_token_head.json +33 -0
  305. hipengine-0.2.0/tests/fixtures/cpu_reference/paged_attn_decode_int8_per_token_head.json +48 -0
  306. hipengine-0.2.0/tests/fixtures/gguf/qwen35_0_8b_q4_1_e2e.json +45 -0
  307. hipengine-0.2.0/tests/fixtures/gguf/qwen35_0_8b_q4_k_m_e2e.json +46 -0
  308. hipengine-0.2.0/tests/fixtures/gguf/qwen35_0_8b_q8_0_e2e.json +41 -0
  309. hipengine-0.2.0/tests/fixtures/gguf/qwen35_0_8b_ud_q4_k_xl_e2e.json +47 -0
  310. hipengine-0.2.0/tests/fixtures/gguf/qwen36_35b_a3b_q4km_p9_e2e.json +54 -0
  311. hipengine-0.2.0/tests/fixtures/gguf/qwen36_35b_a3b_q4km_smoke.json +39 -0
  312. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_aotriton_discovery.py +9 -0
  313. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_cpu_reference.py +93 -0
  314. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_generation_qwen35_paro.py +51 -2
  315. hipengine-0.2.0/tests/test_gguf_e2e_acceptance.py +131 -0
  316. hipengine-0.2.0/tests/test_gguf_embedding_dispatch.py +66 -0
  317. hipengine-0.2.0/tests/test_gguf_expert_pack8_gemv.py +261 -0
  318. hipengine-0.2.0/tests/test_gguf_gemv_decode_dispatch.py +344 -0
  319. hipengine-0.2.0/tests/test_gguf_k_gemv.py +238 -0
  320. hipengine-0.2.0/tests/test_gguf_k_selected_pack8_gemv_decode.py +272 -0
  321. hipengine-0.2.0/tests/test_gguf_k_selected_wmma_prefill.py +469 -0
  322. hipengine-0.2.0/tests/test_gguf_k_t16_selected_wmma_prefill.py +448 -0
  323. hipengine-0.2.0/tests/test_gguf_linear_dispatch.py +781 -0
  324. hipengine-0.2.0/tests/test_gguf_ops.py +260 -0
  325. hipengine-0.2.0/tests/test_gguf_q4_k_gemv.py +193 -0
  326. hipengine-0.2.0/tests/test_gguf_q4_k_pack8_gemv_decode.py +316 -0
  327. hipengine-0.2.0/tests/test_gguf_q4_k_selected_dual_pack8_gemv_decode.py +287 -0
  328. hipengine-0.2.0/tests/test_gguf_q4_k_selected_wmma_prefill.py +603 -0
  329. hipengine-0.2.0/tests/test_gguf_q4_k_t16_selected_wmma_prefill.py +235 -0
  330. hipengine-0.2.0/tests/test_gguf_q4_k_tile16_repack.py +74 -0
  331. hipengine-0.2.0/tests/test_gguf_q4_k_wmma_prefill.py +528 -0
  332. hipengine-0.2.0/tests/test_gguf_q6_k_embedding.py +176 -0
  333. hipengine-0.2.0/tests/test_gguf_q6_k_pack8_gemv_decode.py +204 -0
  334. hipengine-0.2.0/tests/test_gguf_q6_k_t16_gemv_decode.py +125 -0
  335. hipengine-0.2.0/tests/test_gguf_q8_0_pack8_gemv_decode.py +271 -0
  336. hipengine-0.2.0/tests/test_gguf_q8_0_t16_gemv_decode.py +590 -0
  337. hipengine-0.2.0/tests/test_gguf_q8_0_t16_wmma_prefill.py +315 -0
  338. hipengine-0.2.0/tests/test_gguf_q8_0_wmma_prefill.py +533 -0
  339. hipengine-0.2.0/tests/test_gguf_q8_0_wmma_prefill_dual.py +232 -0
  340. hipengine-0.2.0/tests/test_gguf_quant_layout.py +96 -0
  341. hipengine-0.2.0/tests/test_gguf_reader.py +161 -0
  342. hipengine-0.2.0/tests/test_gguf_t16_repack.py +179 -0
  343. hipengine-0.2.0/tests/test_gguf_t16_selected_gemv_decode.py +715 -0
  344. hipengine-0.2.0/tests/test_kv_dispatch.py +163 -0
  345. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_kvcache_policy.py +66 -1
  346. hipengine-0.2.0/tests/test_kvcache_spans.py +130 -0
  347. hipengine-0.2.0/tests/test_llm_gguf_generate_path.py +118 -0
  348. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_model_quant_and_imports.py +30 -0
  349. hipengine-0.2.0/tests/test_qwen35_bench_memory_audit.py +37 -0
  350. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_decode_state.py +222 -6
  351. hipengine-0.2.0/tests/test_qwen35_gguf_chunked_prefill.py +72 -0
  352. hipengine-0.2.0/tests/test_qwen35_gguf_compact_moe_gemv_routing.py +468 -0
  353. hipengine-0.2.0/tests/test_qwen35_gguf_compact_moe_wmma_resolver.py +140 -0
  354. hipengine-0.2.0/tests/test_qwen35_gguf_compact_moe_wmma_routing.py +263 -0
  355. hipengine-0.2.0/tests/test_qwen35_gguf_decode_graph_policy.py +173 -0
  356. hipengine-0.2.0/tests/test_qwen35_gguf_decode_repack_dispatch.py +221 -0
  357. hipengine-0.2.0/tests/test_qwen35_gguf_decode_repack_semantics.py +36 -0
  358. hipengine-0.2.0/tests/test_qwen35_gguf_expert_sidecar.py +121 -0
  359. hipengine-0.2.0/tests/test_qwen35_gguf_fastpath_safety.py +150 -0
  360. hipengine-0.2.0/tests/test_qwen35_gguf_full_attention_gpu.py +407 -0
  361. hipengine-0.2.0/tests/test_qwen35_gguf_gdn_prefill_correctness.py +695 -0
  362. hipengine-0.2.0/tests/test_qwen35_gguf_gdn_prefill_routing.py +290 -0
  363. hipengine-0.2.0/tests/test_qwen35_gguf_mapping.py +138 -0
  364. hipengine-0.2.0/tests/test_qwen35_gguf_materialize.py +273 -0
  365. hipengine-0.2.0/tests/test_qwen35_gguf_moe_replay.py +57 -0
  366. hipengine-0.2.0/tests/test_qwen35_gguf_p10_x2_layer_correctness.py +163 -0
  367. hipengine-0.2.0/tests/test_qwen35_gguf_p9_e2e_correctness.py +145 -0
  368. hipengine-0.2.0/tests/test_qwen35_gguf_rocprof_summary.py +388 -0
  369. hipengine-0.2.0/tests/test_qwen35_gguf_runner.py +154 -0
  370. hipengine-0.2.0/tests/test_qwen35_gguf_tokenizer.py +68 -0
  371. hipengine-0.2.0/tests/test_qwen35_kv_e2e_fixture_gate.py +207 -0
  372. hipengine-0.2.0/tests/test_qwen35_kv_int8_accuracy.py +82 -0
  373. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_linear_attn_gdn_plan.py +53 -0
  374. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_paged_attn_decode_plan.py +112 -1
  375. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_paged_kv_write_plan.py +127 -1
  376. hipengine-0.2.0/tests/test_qwen35_prefill_workspace_policy.py +33 -0
  377. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_resident_batch_layout.py +371 -0
  378. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_router_plan.py +124 -0
  379. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_server_api.py +10 -1
  380. hipengine-0.1.0/.gitattributes +0 -3
  381. hipengine-0.1.0/CHANGELOG.md +0 -38
  382. hipengine-0.1.0/benchmarks/CHANGELOG.md +0 -112
  383. hipengine-0.1.0/benchmarks/README.md +0 -175
  384. hipengine-0.1.0/docs/GGUF.md +0 -576
  385. hipengine-0.1.0/docs/KVCACHE.md +0 -307
  386. hipengine-0.1.0/docs/LESSONS-LEARNED.md +0 -99
  387. hipengine-0.1.0/hipengine/dispatch/__init__.py +0 -18
  388. hipengine-0.1.0/hipengine/kernels/hip_gfx1100/attention/paged_kv_write.py +0 -453
  389. hipengine-0.1.0/hipengine/kvcache/__init__.py +0 -6
  390. hipengine-0.1.0/hipengine/quant/__init__.py +0 -28
  391. hipengine-0.1.0/tests/test_kvcache_spans.py +0 -62
  392. hipengine-0.1.0/uv.lock +0 -1602
  393. {hipengine-0.1.0 → hipengine-0.2.0}/.gitignore +0 -0
  394. {hipengine-0.1.0 → hipengine-0.2.0}/CLAUDE.md +0 -0
  395. {hipengine-0.1.0 → hipengine-0.2.0}/LICENSE +0 -0
  396. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/.gitkeep +0 -0
  397. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-13-hipengine-qwen35-paro-optimal-blocked.json +0 -0
  398. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-13-source-lineage-qwen35-paro-optimal-4k-128.json +0 -0
  399. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-13-source-lineage-qwen35-paro-optimal-512-128.json +0 -0
  400. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-c1-parent-fixture-blocked.json +0 -0
  401. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-cn-correctness-blocked.json +0 -0
  402. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-ab-fused-lmhead128-graph-diagnostic.json +0 -0
  403. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-c1-diagnostic.json +0 -0
  404. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-graph-diagnostic.json +0 -0
  405. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-linear-qkv-z-full-qk-fused-graph-diagnostic.json +0 -0
  406. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-linear-qkv-z-fused-graph-diagnostic.json +0 -0
  407. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-lmhead128-qk-qkvz-fused-graph-diagnostic.json +0 -0
  408. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-tokenizer-cache-diagnostic.json +0 -0
  409. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-0p8b-paro-512-128-blocked.json +0 -0
  410. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c1-parent-fixture-accepted.json +0 -0
  411. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c1-parent-mixed-blocked.json +0 -0
  412. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c1-router-qnorm-blocked.json +0 -0
  413. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c1-scheduler-serial-bench-blocked.json +0 -0
  414. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c2-native-compact-prefill-correctness-accepted.json +0 -0
  415. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c2-scheduler-serial-bench-blocked.json +0 -0
  416. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c2-scheduler-serial-runner-accepted.json +0 -0
  417. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c2-serial-slot-runner-accepted.json +0 -0
  418. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c4-native-compact-prefill-correctness-accepted.json +0 -0
  419. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c4-scheduler-serial-bench-blocked.json +0 -0
  420. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c4-scheduler-serial-runner-accepted.json +0 -0
  421. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c8-native-compact-prefill-correctness-accepted.json +0 -0
  422. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c8-scheduler-serial-bench-blocked.json +0 -0
  423. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c8-scheduler-serial-runner-accepted.json +0 -0
  424. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-cn-generated-equality-accepted.json +0 -0
  425. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-dflash-ddtree-blocked.json +0 -0
  426. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-linear-attn-segment-prefill-accepted.json +0 -0
  427. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefill-compact-c8-blocked.json +0 -0
  428. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefill-full-attn-boundary-blocked.json +0 -0
  429. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefill-full-single-request-accepted.json +0 -0
  430. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefill-multiloop-512-4k-diagnostic.json +0 -0
  431. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefill-plan-blocked.json +0 -0
  432. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer0-attention-accepted.json +0 -0
  433. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer0-attn-rejected.json +0 -0
  434. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer0-decode-rejected.json +0 -0
  435. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer0-gated-recurrent-rejected.json +0 -0
  436. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer0-stage-bisect-rejected.json +0 -0
  437. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer3-fullattn-stage-accepted.json +0 -0
  438. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-prefill-rejected.json +0 -0
  439. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-scratch-restore-sweep.json +0 -0
  440. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-serial-fullattn-layer4-accepted.json +0 -0
  441. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-serial-suffix-full40-accepted.json +0 -0
  442. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-sweep-rejected.json +0 -0
  443. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-prefix-bisect-blocked.json +0 -0
  444. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-varlen-full-attn-prefill-accepted.json +0 -0
  445. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-aotriton-cast-glue-diagnostic.json +0 -0
  446. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-aotriton-gate-rotate-diagnostic.json +0 -0
  447. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-aotriton-threshold-sweep-diagnostic.json +0 -0
  448. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-aotriton-v3-memory-diagnostic.json +0 -0
  449. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-aotriton-v3-prefill-diagnostic.json +0 -0
  450. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-comparison-tables-diagnostic.json +0 -0
  451. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-decode-graph-replay-diagnostic.json +0 -0
  452. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-long-checkpoint-diagnostic.json +0 -0
  453. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-prefill-chunking-diagnostic.json +0 -0
  454. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-gfx1151-shisa-qwen36-packed-canonical-sweep-diagnostic.json +0 -0
  455. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-gfx1151-shisa-qwen36-packed-chunk256-sweep-diagnostic.json +0 -0
  456. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-08b-gfx1151-dense-diagnostic.json +0 -0
  457. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-35b-qwen36-27b-gfx1151-diagnostic.json +0 -0
  458. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d11-rotate-dual-pack8-fusion-rejected.json +0 -0
  459. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d12-rmsnorm-producer-fusion-deferred.json +0 -0
  460. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d13-same-input-projection-fusions-rejected.json +0 -0
  461. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d14-selected-moe-postop-fold-rejected.json +0 -0
  462. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d15-router-coop-fold-rejected.json +0 -0
  463. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d16-kv-pack8-fusion-rejected.json +0 -0
  464. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d21-marlin-k-qweight-neutral-diagnostic.json +0 -0
  465. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d31-d33-grouped-gqa-long-context-diagnostic.json +0 -0
  466. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d42-dispatch-cap-rejected.json +0 -0
  467. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d44-launch-bounds-deferred.json +0 -0
  468. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d51-gdn-decode-audit.json +0 -0
  469. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d52-w8a16-decode-audit.json +0 -0
  470. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-p33-moe-metadata-fanout-deferred.json +0 -0
  471. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-p52-prefill-chunk-autotune-accepted.json +0 -0
  472. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p11-rocblas-ab-rejected.json +0 -0
  473. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p12-shared-gate-up-token-tile-diagnostic.json +0 -0
  474. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p13-shared-down-combine-token-tile-diagnostic.json +0 -0
  475. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p14-moe-wmma-threshold-diagnostic.json +0 -0
  476. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p16-prefill-mcumode-rejected.json +0 -0
  477. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p31-gdn-rotate-rejected.json +0 -0
  478. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p32-router-sigmoid-rejected.json +0 -0
  479. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-paro-dual-format-diagnostic.json +0 -0
  480. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-w1-unroll600-ablation-diagnostic.json +0 -0
  481. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-rocprof-amdahl-diagnostic.json +0 -0
  482. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen36-packed-shared-decode-fusion-diagnostic.json +0 -0
  483. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen36-shisa-force-legacy-diagnostic.json +0 -0
  484. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen36-shisa-packed-vs-legacy-refresh-diagnostic.json +0 -0
  485. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-llamacpp-hip-qwen36-peak.json +0 -0
  486. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-llamacpp-upstream-gfx1151-qwen36-gguf-rerun-diagnostic.json +0 -0
  487. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-llamacpp-vulkan-qwen36-peak.json +0 -0
  488. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-18-hipengine-gfx1100-shisa-qwen36-packed-gt1k-default-diagnostic.json +0 -0
  489. {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-18-hipengine-qwen35-gt1k-prefill-chunk-policy-diagnostic.json +0 -0
  490. {hipengine-0.1.0 → hipengine-0.2.0}/docs/API.md +0 -0
  491. {hipengine-0.1.0 → hipengine-0.2.0}/docs/DFLASH.md +0 -0
  492. {hipengine-0.1.0 → hipengine-0.2.0}/docs/IMPLEMENTATION.md +0 -0
  493. {hipengine-0.1.0 → hipengine-0.2.0}/docs/MARLIN.md +0 -0
  494. {hipengine-0.1.0 → hipengine-0.2.0}/docs/MTP.md +0 -0
  495. {hipengine-0.1.0 → hipengine-0.2.0}/docs/OPTIMIZE.md +0 -0
  496. {hipengine-0.1.0 → hipengine-0.2.0}/docs/PLAN.md +0 -0
  497. {hipengine-0.1.0 → hipengine-0.2.0}/docs/PREFILL.md +0 -0
  498. {hipengine-0.1.0 → hipengine-0.2.0}/docs/ROOFLINE.md +0 -0
  499. {hipengine-0.1.0 → hipengine-0.2.0}/fixtures/qwen35_paro/parent_512_32_seed1234.json +0 -0
  500. {hipengine-0.1.0 → hipengine-0.2.0}/hatch_build.py +0 -0
  501. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/__init__.py +0 -0
  502. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/benchmark/__init__.py +0 -0
  503. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/benchmark/correctness.py +0 -0
  504. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/__init__.py +0 -0
  505. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/build.py +0 -0
  506. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/device.py +0 -0
  507. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/hip.py +0 -0
  508. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/memory.py +0 -0
  509. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/rocblas.py +0 -0
  510. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/tensor.py +0 -0
  511. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/dispatch/batch.py +0 -0
  512. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/dispatch/fusion.py +0 -0
  513. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/distributed/__init__.py +0 -0
  514. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/generation/batch_scheduler.py +0 -0
  515. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/__init__.py +0 -0
  516. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/backends.py +0 -0
  517. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/cpu_reference/fixtures.py +0 -0
  518. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/cuda_sm86/__init__.py +0 -0
  519. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/__init__.py +0 -0
  520. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton.py +0 -0
  521. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_release.toml +0 -0
  522. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/MANIFEST.vendor.json +0 -0
  523. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/aiter_hip_common.h +0 -0
  524. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/flash/aiter.h +0 -0
  525. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/kernel_cluster.h +0 -0
  526. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/lazy_tensor_internal.h +0 -0
  527. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/packed_kernel.h +0 -0
  528. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/triton_kernel.h +0 -0
  529. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/util.h +0 -0
  530. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/config.h +0 -0
  531. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/cpp_tune.h +0 -0
  532. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/dtypes.h +0 -0
  533. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/flash.h +0 -0
  534. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/runtime.h +0 -0
  535. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/util.h +0 -0
  536. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/v2/flash.h +0 -0
  537. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_F_0_0___gfx11xx.aks2" +0 -0
  538. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_F_0_1___gfx11xx.aks2" +0 -0
  539. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_F_3_0___gfx11xx.aks2" +0 -0
  540. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_T_0_0___gfx11xx.aks2" +0 -0
  541. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_T_0_1___gfx11xx.aks2" +0 -0
  542. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_T_3_0___gfx11xx.aks2" +0 -0
  543. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_F_0_0___gfx11xx.aks2" +0 -0
  544. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_F_0_1___gfx11xx.aks2" +0 -0
  545. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_F_3_0___gfx11xx.aks2" +0 -0
  546. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_T_0_0___gfx11xx.aks2" +0 -0
  547. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_T_0_1___gfx11xx.aks2" +0 -0
  548. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_T_3_0___gfx11xx.aks2" +0 -0
  549. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/libaotriton_v2.so +0 -0
  550. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/libaotriton_v2.so.0.11.2 +0 -0
  551. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_wrap.cc +0 -0
  552. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/common/__init__.py +0 -0
  553. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/convert/__init__.py +0 -0
  554. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/convert/cast.hip +0 -0
  555. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/convert/cast.py +0 -0
  556. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/fused/paro_combine.hip +0 -0
  557. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/fused/paro_combine.py +0 -0
  558. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/fused/paro_silu.hip +0 -0
  559. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/fused/paro_silu.py +0 -0
  560. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear/__init__.py +0 -0
  561. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear/lm_head.hip +0 -0
  562. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear/lm_head.py +0 -0
  563. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear_attn/__init__.py +0 -0
  564. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear_attn/conv.py +0 -0
  565. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/group_scatter.hip +0 -0
  566. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/group_scatter.py +0 -0
  567. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/prefill.py +0 -0
  568. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/norm/__init__.py +0 -0
  569. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/norm/rmsnorm.hip +0 -0
  570. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/norm/rmsnorm.py +0 -0
  571. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/paro_awq_gemv.hip +0 -0
  572. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/paro_awq_gemv.py +0 -0
  573. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/paro_marlin_k.hip +0 -0
  574. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/paro_marlin_k.py +0 -0
  575. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/w8a16_linear.hip +0 -0
  576. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/w8a16_linear.py +0 -0
  577. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/rotary/__init__.py +0 -0
  578. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/rotary/paro_rotate.hip +0 -0
  579. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/rotary/paro_rotate.py +0 -0
  580. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/rotary/qwen35_rotary.hip +0 -0
  581. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/rotary/qwen35_rotary.py +0 -0
  582. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/runtime/__init__.py +0 -0
  583. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/runtime/state.hip +0 -0
  584. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/smoke/__init__.py +0 -0
  585. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/smoke/smoke_add.hip +0 -0
  586. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/smoke/smoke_add.py +0 -0
  587. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/wmma/__init__.py +0 -0
  588. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/wmma/paro_awq_wmma.hip +0 -0
  589. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/wmma/paro_awq_wmma.py +0 -0
  590. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1151/__init__.py +0 -0
  591. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/layers/__init__.py +0 -0
  592. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/layers/base.py +0 -0
  593. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/loading/materialize.py +0 -0
  594. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/loading/qwen35_paro.py +0 -0
  595. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/loading/safetensors.py +0 -0
  596. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/models/base.py +0 -0
  597. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/models/registry.py +0 -0
  598. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/models/toy.py +0 -0
  599. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/quant/base.py +0 -0
  600. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/quant/bf16.py +0 -0
  601. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/quant/fp16.py +0 -0
  602. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/quant/registry.py +0 -0
  603. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/quant/w4_paro.py +0 -0
  604. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/runtime/workspace.py +0 -0
  605. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/server/__init__.py +0 -0
  606. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/server/__main__.py +0 -0
  607. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/speculative/__init__.py +0 -0
  608. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/speculative/interfaces.py +0 -0
  609. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/util/__init__.py +0 -0
  610. {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/util/amdgpu_vram.py +0 -0
  611. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/__init__.py +0 -0
  612. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/check_fixtures.py +0 -0
  613. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/check_lineage.py +0 -0
  614. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/fetch_aotriton.sh +0 -0
  615. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/gdn_decode_probe.py +0 -0
  616. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/llamacpp_bench_with_peak.py +0 -0
  617. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_batch_correctness.py +0 -0
  618. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_compare_tables.py +0 -0
  619. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_dflash_ddtree_blocker.py +0 -0
  620. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_compact_prefill_plan.py +0 -0
  621. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_boundary.py +0 -0
  622. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_plan.py +0 -0
  623. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_paro_next_token.py +0 -0
  624. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_rocprof_audit.py +0 -0
  625. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/strip_paro_safetensors.py +0 -0
  626. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/vendor_aotriton.sh +0 -0
  627. {hipengine-0.1.0 → hipengine-0.2.0}/scripts/w8a16_decode_probe.py +0 -0
  628. {hipengine-0.1.0 → hipengine-0.2.0}/tests/fixtures/cpu_reference/attention_decode_masked.json +0 -0
  629. {hipengine-0.1.0 → hipengine-0.2.0}/tests/fixtures/cpu_reference/full_attn_prefill_causal_gqa_gate.json +0 -0
  630. {hipengine-0.1.0 → hipengine-0.2.0}/tests/fixtures/cpu_reference/linear_basic.json +0 -0
  631. {hipengine-0.1.0 → hipengine-0.2.0}/tests/fixtures/cpu_reference/rmsnorm_basic.json +0 -0
  632. {hipengine-0.1.0 → hipengine-0.2.0}/tests/fixtures/cpu_reference/rotate_split_half.json +0 -0
  633. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_build.py +0 -0
  634. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_cast_plan.py +0 -0
  635. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_check_lineage.py +0 -0
  636. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_dense_gemv_plan.py +0 -0
  637. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_dispatch_batch.py +0 -0
  638. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_fusion_spike.py +0 -0
  639. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_generation_batch_scheduler.py +0 -0
  640. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_gfx1151_backend.py +0 -0
  641. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_hip_runtime.py +0 -0
  642. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_kernel_registry.py +0 -0
  643. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_llm_generate.py +0 -0
  644. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_lm_head_plan.py +0 -0
  645. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_loading_materialize.py +0 -0
  646. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_loading_safetensors.py +0 -0
  647. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_memory_stats.py +0 -0
  648. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_paro_awq_gemv_plan.py +0 -0
  649. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_paro_awq_wmma_plan.py +0 -0
  650. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_paro_combine_plan.py +0 -0
  651. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_paro_rotate_plan.py +0 -0
  652. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_paro_silu_plan.py +0 -0
  653. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_linear_attn_conv_plan.py +0 -0
  654. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_moe_group_scatter_plan.py +0 -0
  655. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_native_prefill_boundary.py +0 -0
  656. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_native_prefill_fullattn_stage_probe.py +0 -0
  657. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_paro_layout.py +0 -0
  658. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_paro_marlin_k.py +0 -0
  659. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_rmsnorm_plan.py +0 -0
  660. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_rotary_plan.py +0 -0
  661. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_runtime_state_plan.py +0 -0
  662. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_runtime_workspace.py +0 -0
  663. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_smoke_add_plan.py +0 -0
  664. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_speculative_interfaces.py +0 -0
  665. {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_w8a16_linear_plan.py +0 -0
@@ -0,0 +1,10 @@
1
+ # WORKLOG.md is append-only: when two branches both add new entries, git's
2
+ # built-in union driver concatenates both sides instead of emitting conflict
3
+ # markers. For an append-only file this is the correct semantics (common
4
+ # prefix + ours-tail + theirs-tail). See scripts/resolve_worklog_conflict.py
5
+ # for the recovery path when conflict markers already exist in the file.
6
+ WORKLOG.md merge=union
7
+
8
+ hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/** -whitespace
9
+ hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/**/*.aks2 filter=lfs diff=lfs merge=lfs -text
10
+ hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/**/*.so* filter=lfs diff=lfs merge=lfs -text
@@ -0,0 +1,103 @@
1
+ name: Publish to PyPI
2
+
3
+ on:
4
+ push:
5
+ tags:
6
+ - "v*"
7
+
8
+ permissions: {}
9
+
10
+ jobs:
11
+ build:
12
+ runs-on: ubuntu-24.04
13
+ permissions:
14
+ contents: read
15
+ outputs:
16
+ version: ${{ steps.meta.outputs.version }}
17
+
18
+ steps:
19
+ - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
20
+ with:
21
+ persist-credentials: false
22
+ lfs: true
23
+
24
+ - name: Install uv
25
+ uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57 # v8.0.0
26
+
27
+ - name: Set up Python 3.12
28
+ run: |
29
+ uv python install 3.12
30
+ uv venv --python 3.12
31
+
32
+ - name: Extract version metadata
33
+ id: meta
34
+ run: |
35
+ version="$(uv run --no-sync python - <<'PY'
36
+ from pathlib import Path
37
+ import tomllib
38
+
39
+ data = tomllib.loads(Path("pyproject.toml").read_text())
40
+ print(data["project"]["version"])
41
+ PY
42
+ )"
43
+ echo "version=${version}" >> "$GITHUB_OUTPUT"
44
+
45
+ - name: Verify tag matches package version
46
+ env:
47
+ RELEASE_TAG: ${{ github.ref_name }}
48
+ RELEASE_VERSION: ${{ steps.meta.outputs.version }}
49
+ run: |
50
+ expected="v${RELEASE_VERSION}"
51
+ if [ "${RELEASE_TAG}" != "${expected}" ]; then
52
+ echo "::error::Tag '${RELEASE_TAG}' does not match package version '${expected}'"
53
+ exit 1
54
+ fi
55
+
56
+ - name: Install dependencies
57
+ run: uv pip install -e '.[dev]' build twine
58
+
59
+ - name: Run release validation
60
+ run: uv run --no-sync pytest -q
61
+
62
+ - name: Remove stale build artifacts
63
+ run: rm -rf dist/
64
+
65
+ - name: Build
66
+ run: uv run --no-sync python -m build
67
+
68
+ - name: Verify package metadata
69
+ run: uv run --no-sync python -m twine check dist/*
70
+
71
+ - name: Upload build artifacts
72
+ uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
73
+ with:
74
+ name: dist
75
+ path: |
76
+ dist/hipengine-*.whl
77
+ dist/hipengine-*.tar.gz
78
+
79
+ publish:
80
+ needs: build
81
+ runs-on: ubuntu-24.04
82
+ environment: pypi-publish
83
+ permissions:
84
+ id-token: write
85
+ contents: read
86
+ attestations: write
87
+
88
+ steps:
89
+ - name: Download build artifacts
90
+ uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
91
+ with:
92
+ name: dist
93
+ path: dist/
94
+
95
+ - name: Generate artifact attestations
96
+ uses: actions/attest-build-provenance@a2bbfa25375fe432b6a289bc6b6cd05ecd0c4c32 # v4.1.0
97
+ with:
98
+ subject-path: |
99
+ dist/hipengine-*.whl
100
+ dist/hipengine-*.tar.gz
101
+
102
+ - name: Publish to PyPI (trusted publishing)
103
+ uses: pypa/gh-action-pypi-publish@cef221092ed1bacb1cc03d23a2d87d1d172e277b # v1.14.0
@@ -64,6 +64,7 @@ Do not drift these casually. They define what hipEngine is.
64
64
 
65
65
  - Keep changes scoped to one logical unit (one kernel family, one plugin, one doc, one phase milestone).
66
66
  - Write or update the targeted test/fixture before implementation when behavior or math changes. If RED-first is impractical, record why in `WORKLOG.md`.
67
+ - When adding tests that call HIP/ROCm runtime, `hipcc`, or GPU kernels, add an explicit HIP-availability guard (for example `ctypes.CDLL("libamdhip64.so")` + `pytest.skip`) so no-ROCm CI/publish runners skip them instead of failing release validation.
67
68
  - Log non-trivial decisions, measurements, and dependency additions in `WORKLOG.md` as they happen.
68
69
  - When profiling Python/ctypes JIT-built kernels with `rocprofv3`, prebuild the `.so` outside the profiler and run the profiled command with a precomputed compiler-version file plus `require_cached`; do not let the profiled process spawn `hipcc`/clang.
69
70
  - Do not silently add `import torch`, `flash_attn`, or other CUDA-only deps to hot-path modules.
@@ -144,6 +145,7 @@ Working tree is shared state. Other agents or the human may be editing concurren
144
145
  - **High-conflict files:** `AGENTS.md`, `CLAUDE.md`, `docs/PLAN.md`, `docs/BENCHMARK.md`, `docs/TESTING.md`, `docs/KERNELS.md`, `docs/IMPLEMENTATION.md`, `WORKLOG.md`, `pyproject.toml`, `hipengine/kernels/registry.py`, `hipengine/quant/registry.py`, `hipengine/models/registry.py`, `hipengine/dispatch/fusion.py`, `hipengine/core/*`.
145
146
  - Same-file contention: stop and coordinate. The designated agent stages and commits their scoped hunks first to unblock others.
146
147
  - `WORKLOG.md` appends are expected and not a conflict unless there are actual conflict markers or interleaved garbled lines. Re-read the live tail, append after it, commit with your logical unit.
148
+ - `WORKLOG.md` is configured with git's built-in `merge=union` driver (see `.gitattributes`), so concurrent appends auto-resolve as `common prefix + ours-tail + theirs-tail` with no conflict markers. If markers do appear (e.g. a stash or a rebase started before this was configured), run `python3 scripts/resolve_worklog_conflict.py WORKLOG.md` (add `--sort-by-date` to re-order `## YYYY-MM-DD` sections in each resolved block; `--check` for a pre-commit gate). The script only touches conflict blocks; content outside markers is left exactly as-is.
147
149
  - Do not clean up another agent's benchmark outputs, staged files, or local artifacts unless the task explicitly asks for that cleanup.
148
150
 
149
151
  ## External Reference Repos
@@ -0,0 +1,112 @@
1
+ # Changelog
2
+
3
+ All notable user-facing changes for hipEngine releases are documented here.
4
+
5
+ This changelog is for package/API releases. Performance rollup history remains in
6
+ [`benchmarks/CHANGELOG.md`](benchmarks/CHANGELOG.md), with detailed benchmark
7
+ evidence under [`benchmarks/results/`](benchmarks/results/).
8
+
9
+ ## v0.2.0 - 2026-05-25
10
+
11
+ Minor release for the GGUF runtime path and W7900 benchmark refresh. GGUF is a
12
+ meaningful new model-loading surface rather than a patch-level fix, so this
13
+ supersedes the previously planned v0.1.2 patch.
14
+
15
+ ### Added
16
+
17
+ - Added Qwen3.6 35B MoE GGUF support for `Q4_K_M` and `Q4_K_S` model files,
18
+ including resident GGUF loading, bulk prefill, graph-replay decode,
19
+ decode-repacked T16 layouts, and WMMA/GEMV fast-path controls used by the
20
+ W7900 benchmark profile.
21
+ - Added `docs/ENVS.md` as the canonical environment-variable reference, including
22
+ TheRock ROCm process setup, cached-build profiling guidance, and safe GGUF
23
+ benchmark profiles.
24
+ - Added a persistent README sweep harness that loads each hipEngine model once
25
+ and runs repeated in-session workload measurements, matching llama-bench-style
26
+ repetition without multiplying model load/decode-repack time by every shape.
27
+
28
+ ### Changed
29
+
30
+ - Refreshed W7900 README performance tables with 5-run persistent-session medians
31
+ for packed PARO and GGUF Q4_K_S while keeping the existing llama.cpp HIP/Vulkan
32
+ comparison rows unchanged.
33
+ - Documented the current GGUF tradeoffs: higher one-time load cost and resident
34
+ memory from decode-repack, Q4_K_S preferred for tighter VRAM budgets, and
35
+ performance still behind PARO on some shapes while already competitive in the
36
+ broader W7900 comparison.
37
+
38
+ ### Fixed
39
+
40
+ - Fixed the PARO resident prefill workspace-overlap regression that shipped in
41
+ v0.1.1: short and mid prompts now keep prefill workspaces resident through
42
+ 32K tokens, restoring 512/128-class prefill throughput while retaining the
43
+ long-context memory-saving path for prompts above 32K when active chunking
44
+ splits the prompt.
45
+ - Fixed GGUF non-split full-attention decode in max-context persistent sessions
46
+ by launching the context kernel with the active decode context instead of the
47
+ session's maximum allocation length.
48
+
49
+ ### Known limitations
50
+
51
+ - GGUF support remains alpha: production correctness and performance coverage is
52
+ strongest for the documented Qwen3.6 35B MoE Q4_K_M/Q4_K_S paths on gfx1100,
53
+ and other GGUF quants/models require local validation.
54
+ - GGUF model load is slower than packed PARO on the same host because current
55
+ decode-repack happens on load and is not yet cached on disk.
56
+
57
+ ## v0.1.1 - 2026-05-19
58
+
59
+ Patch release focused on long-context memory documentation and the INT8 KV cache
60
+ bring-up that landed after v0.1.0.
61
+
62
+ ### Added
63
+
64
+ - INT8 KV cache policy controls and dispatch coverage for Qwen/PARO resident
65
+ inference paths, including CPU/layer/E2E correctness gates and memory audits.
66
+ - Documented Qwen3.6 packed PARO memory rows for 128K BF16 KV, 128K INT8 KV, and
67
+ 256K INT8 KV on W7900/gfx1100, with retained-KV and loaded-weight VRAM notes.
68
+
69
+ ### Changed
70
+
71
+ - Reduced the 256K INT8 KV tracked allocator high-water mark below the 24 GiB
72
+ class target by releasing/reusing prefill scratch and AOTriton query buffers.
73
+ - Clarified that packed vs unstripped PARO checkpoint size does not translate to
74
+ meaningfully different resident model-weight VRAM for the current text runtime.
75
+
76
+ ### Known limitations
77
+
78
+ - INT8 KV correctness is gated by deterministic fixtures and layer probes; it is
79
+ not yet a long-rollout perplexity or compounding-error study.
80
+ - Qwen3.6 packed throughput rows remain diagnostic pending a promoted public
81
+ `LLM.generate()` correctness/repetition gate.
82
+
83
+ ## v0.1.0 - 2026-05-18
84
+
85
+ Initial public alpha release.
86
+
87
+ ### Added
88
+
89
+ - Torch-free Python runtime hot path for local ROCm inference bring-up.
90
+ - Plugin registries keyed by model/backend/quant/layer variants.
91
+ - HIP backends for `gfx1100` and `gfx1151`, plus `backend="auto"` detection with
92
+ `HIPENGINE_BACKEND` force override guidance for nearby targets.
93
+ - Qwen3.5/Qwen3.6 PARO W4 runtime path, JIT HIP build/cache plumbing, AOTriton
94
+ prefill runtime packaging, and OpenAI-compatible server entry point.
95
+ - CPU reference kernels and focused correctness/performance documentation.
96
+
97
+ ### Packaging
98
+
99
+ - PyPI project name: `hipengine`.
100
+ - Python import package: `hipengine`.
101
+ - Canonical repository/wordmark: `hipEngine`.
102
+ - Release wheels are Linux x86-64 `manylinux_2_39` platform wheels because the
103
+ package bundles a ROCm/AOTriton shared-library runtime; ROCm runtime libraries
104
+ remain external system dependencies.
105
+
106
+ ### Known limitations
107
+
108
+ - Alpha-quality API and model coverage; expect sharp edges outside the documented
109
+ Qwen/PARO paths.
110
+ - Default supported GPU targets are `gfx1100` and `gfx1151`; other AMD targets
111
+ require explicit backend forcing and local validation.
112
+ - Model weights are not distributed with the package.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: hipengine
3
- Version: 0.1.0
3
+ Version: 0.2.0
4
4
  Summary: ROCm-native local LLM inference engine with a torch-free runtime hot path
5
5
  Project-URL: Homepage, https://github.com/shisa-ai/hipEngine
6
6
  Project-URL: Repository, https://github.com/shisa-ai/hipEngine
@@ -75,52 +75,143 @@ supported GPUs and models.
75
75
 
76
76
  ## Status
77
77
 
78
- **v0.1.x.** The runtime hot path is torch-free by construction; kernel
79
- families and registry plumbing are landing under
80
- [`hipengine/kernels/hip_gfx1100/`](hipengine/kernels/hip_gfx1100/). Current
81
- single-model tuning targets
78
+ **v0.2.0 alpha.** The runtime hot path is torch-free by construction, and the
79
+ first two 35B-class model-loading surfaces are now available on gfx1100:
82
80
  [shisa-ai/Qwen3.6-35B-A3B-PARO-full4096-e5-packed](https://huggingface.co/shisa-ai/Qwen3.6-35B-A3B-PARO-full4096-e5-packed)
83
81
  (19.07 GiB, 4.68 bpw) in packed
84
- [ParoQuant](https://github.com/shisa-ai/paroquant) format.
82
+ [ParoQuant](https://github.com/shisa-ai/paroquant) format, plus Qwen3.6 GGUF
83
+ `Q4_K_M` / `Q4_K_S` files through the new resident GGUF path.
85
84
 
86
- ## gfx1100 (Radeon RX 7900 XTX / Radeon Pro W7900)
85
+ - INT8 KV cache support has been added for PARO. Qwen 3 MoE's full 256K context window can fit in <24GB tracked memory; see [Memory Usage](#memory-usage).
86
+ - Qwen 3.6 [Q4_K_M](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF?show_file_info=Qwen3.6-35B-A3B-UD-Q4_K_M.gguf) and [Q4_K_S](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF?show_file_info=Qwen3.6-35B-A3B-UD-IQ4_XS.gguf) GGUF support has landed (W7900 Q4_K_S sweep is in [Performance](#performance) alongside packed PARO and llama.cpp Q4_K_M HIP/Vulkan baselines). GGUF uses a substantial GGUF-specific runtime path with bulk prefill, graph decode, and on-load decode-repack into T16 tile layouts. Q4_K_S is recommended on 24 GiB cards because Q4_K_M is bigger; on the 48 GiB W7900 Q4_K_S fits all the way to 128K context, while on 24 GiB cards expect roughly 64K. GGUF also has a higher per-session load cost (~60 s vs ~24 s for PARO packed on the same hardware) for the same decode-repack reason.
87
+ - Current gfx1100 performance snapshots are summarized in [Performance](#performance) and compared against recent llama.cpp Q4_K_M baselines.
87
88
 
88
- While we are far from [gfx1100 roofline](https://github.com/shisa-ai/hipEngine/blob/main/docs/ROOFLINE.md), the current gfx1100 implementation does well compared to Q4_K_M quants of recent llama.cpp builds (`b9042`) on the same model family. The latest W7900 packed rows use the default prefill policy: 512-token prompts stay unchunked and prompts above 1K use `1024/1024/4096/1024/1024` chunks.
89
89
 
90
- ### Prefill tok/s
90
+ ## Hardware targets
91
+
92
+ | Backend | Hardware | Status |
93
+ | --- | --- | --- |
94
+ | `cpu_reference` | Any CPU, numpy | Correctness oracle; CI without GPU |
95
+ | `hip_gfx1100` | AMD Radeon Pro W7900 / RX 7900 XTX (RDNA3) | Active backend |
96
+ | `hip_gfx1151` | AMD Ryzen AI MAX+ 395 / Radeon 8060S (Strix Halo, RDNA3.5) | Active backend |
97
+ | `cuda_sm86` | NVIDIA Ampere consumer (3090-class) | Planned peer backend |
91
98
 
92
- | Workload | hipEngine shisa Qwen3.6 packed PARO | llama.cpp HIP | llama.cpp Vulkan |
99
+ `backend="auto"` is the public API/server default. It maps exact `gfx1100` and
100
+ `gfx1151` detections to the matching HIP backend; unknown ROCm targets warn and
101
+ select `cpu_reference` where a CPU implementation exists. Users on nearby targets
102
+ such as `gfx1101`/`gfx1102` can force a backend with `backend="hip_gfx1100"`,
103
+ `--backend hip_gfx1100`, or `HIPENGINE_BACKEND=hip_gfx1100` after validating
104
+ correctness/performance.
105
+
106
+ Wave32 is the default for `hip_gfx1100` device code; wave64 is treated as an
107
+ isolated experiment with its own gates (see
108
+ [`docs/PLAN.md`](docs/PLAN.md#rdna3-wavefront-and-scheduling-caveat)).
109
+
110
+ ## Memory Usage
111
+
112
+ With BF16 KV cache, hipEngine running the packed Qwen 3.6 PARO model fits a
113
+ 128K context window in a 24GB-class memory budget. The INT8 KV cache option
114
+ (with FP16 per-token/per-head scales) uses the
115
+ `--kv-storage int8_per_token_head` flag and lets the **full 256K context** fit
116
+ under 24 GiB tracked allocator peak.
117
+
118
+ The numbers below are for
119
+ `shisa-ai/Qwen3.6-35B-A3B-PARO-full4096-e5-packed` on W7900/gfx1100 with q3072
120
+ full-attention prefill chunks:
121
+
122
+ | Model | Context | KV cache | Sampled peak | Allocator peak | Retained KV | Prefill | Decode |
123
+ | -------------------- | ------: | -------- | -----------: | -------------: | ----------: | -----------: | ---------: |
124
+ | Qwen3.6 35B-A3B PARO | 128K | BF16 | 21.04 GiB | 21.88 GiB | 2.69 GiB | 1091.9 tok/s | 62.2 tok/s |
125
+ | Qwen3.6 35B-A3B PARO | 128K | INT8 | 19.80 GiB | 20.89 GiB | 1.36 GiB | 1076.5 tok/s | 60.0 tok/s |
126
+ | Qwen3.6 35B-A3B PARO | 256K | INT8 | 21.96 GiB | 23.71 GiB | 2.71 GiB | 670.2 tok/s | 40.3 tok/s |
127
+
128
+ Regardless of the difference in PARO weight storage (legacy or packed),
129
+ loaded-weight memory is about the same — approximately 16.4 GiB in VRAM.
130
+
131
+ The INT8 KV correctness gate is currently the deterministic Qwen3.5 PARO
132
+ fixture `fixtures/qwen35_paro/parent_512_32_seed1234.json` (512-token prompt,
133
+ 32 greedy decode tokens): `max_kl=0.015328`, `mean_kl=0.001639`, top-1 agreement
134
+ 100%, and generated IDs match BF16 KV exactly. Layer attention probes at context
135
+ 64 and 520 also had top-1 agreement 100% with max quantized-vs-BF16 KL
136
+ `2.34e-7`. This is a fixture/regression gate, not a long-rollout perplexity
137
+ study, so long context generations may have unmeasured compounding errors.
138
+
139
+ The same 128K/128 Qwen3.5 BF16-vs-INT8 run measured -0.99% prefill tok/s and
140
+ -3.20% decode tok/s for INT8 KV, so speed loss is also very small.
141
+
142
+ See
143
+ [`benchmarks/results/2026-05-19-hipengine-qwen36-packed-int8-kv-readme-memory-diagnostic.json`](benchmarks/results/2026-05-19-hipengine-qwen36-packed-int8-kv-readme-memory-diagnostic.json),
144
+ [`benchmarks/README.md`](benchmarks/README.md#blocked--diagnostic-benchmark-attempts),
145
+ and [`docs/KVCACHE.md`](docs/KVCACHE.md) for commands, artifacts, and the full
146
+ no-shadow memory audit.
147
+
148
+ ### llama.cpp
149
+
150
+ When run with `q8_0` kvcache, llama.cpp can also fit in 24GB:
151
+
152
+ ```bash
153
+ --flash-attn on -ctk q8_0 -ctv q8_0 -c 262144 -b 128 -ub 128
154
+ ```
155
+
156
+ Results:
157
+
158
+ | Model | llama.cpp model buffer | KV cache | Compute buffer | rocm-smi VRAM used | Free VRAM |
159
+ | --- | ---: | ---: | ---: | ---: | ---: |
160
+ | Q4_K_M | 20583 MiB | 2720 MiB | 203 MiB | 24017 MiB / 23.45 GiB | ~543 MiB |
161
+ | Q4_K_S | 19399 MiB | 2720 MiB | 203 MiB | 22832 MiB / 22.30 GiB | ~1728 MiB |
162
+
163
+ With `-ub 512`:
164
+
165
+ | Model | Compute buffer | rocm-smi VRAM used | Free VRAM |
93
166
  | --- | ---: | ---: | ---: |
94
- | 512/128 | **2500.565** | 2436.049 | 1816.927 |
95
- | 4K/128 | **2899.685** | 2176.905 | 1705.093 |
96
- | 32K/128 | **2115.050** | 1496.409 | 1128.554 |
97
- | 128K/128 | **1054.291** | 710.213 | 480.539 |
167
+ | Q4_K_M | 812 MiB | 24540 MiB | ~20 MiB |
168
+ | Q4_K_S | 812 MiB | 23443 MiB | ~1117 MiB |
169
+
170
+ - Note Q4_K_M is incredibly tight with only 20 MiB of headroom and you may either need to resize down or set `-b 512 -ub 128`.
171
+ - Q4_K_S does not need small `-b`/`-ub`; `-ub 512` fits fine, and can even increase to `-b 2048` (but `-ub` is the more important VRAM knob that controls the physical microbatch / compute buffer size for llama.cpp).
172
+
173
+ ## Performance
174
+
175
+ ### gfx1100 (Radeon RX 7900 XTX / Radeon Pro W7900)
176
+
177
+ While we are far from [gfx1100 roofline](https://github.com/shisa-ai/hipEngine/blob/main/docs/ROOFLINE.md), the current gfx1100 implementation does well compared to Q4_K_M quants of recent llama.cpp builds (`b9042`) on the same model family. The latest W7900 hipEngine rows use TheRock ROCm 7.13 and load each resident model once for 1 warmup + 5 measured in-session repetitions per shape. PARO uses the default prefill policy: 512-token prompts stay unchunked and prompts above 1K use `1024/1024/4096/1024/1024` chunks. The `hipEngine GGUF Q4_K_S` column uses the same chunked-prefill policy plus the WMMA prefill + GEMV decode fast paths and the persistent on-load decode-repack into T16 tile layouts.
178
+
179
+ ### Prefill tok/s
180
+
181
+ | Workload | hipEngine PARO | hipEngine GGUF Q4_K_S | llama.cpp HIP | llama.cpp Vulkan |
182
+ | --- | ---: | ---: | ---: | ---: |
183
+ | 512/128 | **2718.497** | 2258.847 | 2436.049 | 1816.927 |
184
+ | 4K/128 | **2838.773** | 2576.673 | 2176.905 | 1705.093 |
185
+ | 32K/128 | **2074.699** | 1893.967 | 1496.409 | 1128.554 |
186
+ | 128K/128 | **1055.454** | 998.143 | 710.213 | 480.539 |
98
187
 
99
188
  ### Decode tok/s
100
189
 
101
- | Workload | hipEngine shisa Qwen3.6 packed PARO | llama.cpp HIP | llama.cpp Vulkan |
102
- | --- | ---: | ---: | ---: |
103
- | 512/128 | 111.516 | 85.487 | **127.515** |
104
- | 4K/128 | 113.094 | 87.375 | **120.163** |
105
- | 32K/128 | 97.594 | 76.994 | **98.073** |
106
- | 128K/128 | 62.027 | 57.341 | **64.478** |
190
+ | Workload | hipEngine PARO | hipEngine GGUF Q4_K_S | llama.cpp HIP | llama.cpp Vulkan |
191
+ | --- | ---: | ---: | ---: | ---: |
192
+ | 512/128 | 103.460 | 109.152 | 85.487 | **127.515** |
193
+ | 4K/128 | 101.964 | 100.048 | 87.375 | **120.163** |
194
+ | 32K/128 | 90.438 | 86.774 | 76.994 | **98.073** |
195
+ | 128K/128 | 59.598 | 57.954 | 57.341 | **64.478** |
107
196
 
108
197
  ### Peak GiB
109
198
 
110
- | Workload | hipEngine shisa Qwen3.6 packed PARO | llama.cpp HIP | llama.cpp Vulkan |
111
- | --- | ---: | ---: | ---: |
112
- | 512/128 | **18.123** | 21.125 | 20.844 |
113
- | 4K/128 | **19.455** | 21.197 | 20.969 |
114
- | 32K/128 | **20.267** | 21.738 | 21.533 |
115
- | 128K/128 | **23.235** | 23.605 | 23.596 |
199
+ | Workload | hipEngine PARO | hipEngine GGUF Q4_K_S | llama.cpp HIP | llama.cpp Vulkan |
200
+ | --- | ---: | ---: | ---: | ---: |
201
+ | 512/128 | 20.962 | 25.108 | 21.125 | **20.844** |
202
+ | 4K/128 | 21.906 | 25.108 | 21.197 | **20.969** |
203
+ | 32K/128 | 22.016 | 25.108 | 21.738 | **21.533** |
204
+ | 128K/128 | **22.122** | 25.108 | 23.605 | 23.596 |
205
+
206
+ hipEngine W7900 row source: [`benchmarks/results/2026-05-25-w7900-hipengine-readme-persistent-5run-diagnostic.json`](benchmarks/results/2026-05-25-w7900-hipengine-readme-persistent-5run-diagnostic.json). Both hipEngine columns are 5-run medians from one resident session allocated for the maximum requested context (`128K/128`), so the peak-memory column is a max-context persistent-session high-water mark rather than each shape's minimum allocation. Existing W7900 llama.cpp HIP/Vulkan Q4_K_M rows are reused unchanged. The hipEngine GGUF Q4_K_S column is compared against the existing llama.cpp Q4_K_M baselines because that is the lineage of measured baselines we have on this host; cross-quant comparisons should be read as approximate.
116
207
 
117
- ## gfx1151 (AMD Ryzen AI MAX+ 395 / Radeon 8060S)
208
+ ### gfx1151 (AMD Ryzen AI MAX+ 395 / Radeon 8060S)
118
209
 
119
210
  The gfx1151 backend is a native `--offload-arch=gfx1151` peer backend using the same registry-keyed kernel surface. The Strix Halo snapshot below uses 256-row prefill chunks, which removed the 4K prefill gap without hurting long-context decode.
120
211
 
121
212
  ### Prefill tok/s
122
213
 
123
- | Workload | hipEngine shisa Qwen3.6 packed PARO | llama.cpp HIP | llama.cpp Vulkan |
214
+ | Workload | hipEngine PARO | llama.cpp HIP | llama.cpp Vulkan |
124
215
  | --- | ---: | ---: | ---: |
125
216
  | 512/128 | 983.206 | **1058.738** | 638.008 |
126
217
  | 4K/128 | **1029.402** | 1004.220 | 595.400 |
@@ -129,7 +220,7 @@ The gfx1151 backend is a native `--offload-arch=gfx1151` peer backend using the
129
220
 
130
221
  ### Decode tok/s
131
222
 
132
- | Workload | hipEngine shisa Qwen3.6 packed PARO | llama.cpp HIP | llama.cpp Vulkan |
223
+ | Workload | hipEngine PARO | llama.cpp HIP | llama.cpp Vulkan |
133
224
  | --- | ---: | ---: | ---: |
134
225
  | 512/128 | **62.060** | 50.537 | 57.615 |
135
226
  | 4K/128 | **63.605** | 49.379 | 55.027 |
@@ -141,25 +232,26 @@ On Strix Halo, `rocm-smi` / sysfs expose only a 512 MiB VRAM aperture, so cross-
141
232
  See [`benchmarks/README.md`](benchmarks/README.md) for full protocol details,
142
233
  correctness status, source-lineage targets, and external comparison baselines.
143
234
 
144
- ## Hardware targets
235
+ ## GGUF Support
145
236
 
146
- | Backend | Hardware | Status |
147
- | --- | --- | --- |
148
- | `hip_gfx1100` | AMD Radeon Pro W7900 / RX 7900 XTX (RDNA3) | Primary, in active bring-up |
149
- | `hip_gfx1151` | AMD Ryzen AI MAX+ 395 / Radeon 8060S (Strix Halo, RDNA3.5) | Active backend |
150
- | `cuda_sm86` | NVIDIA Ampere consumer (3090-class) | Planned peer backend |
151
- | `cpu_reference` | Any CPU, numpy | Correctness oracle; CI without GPU |
237
+ As of v0.2.0, hipEngine includes resident Qwen3.6 GGUF support for `Q4_K_M` and
238
+ `Q4_K_S` model files (with more formats planned). This is a major runtime path,
239
+ not just a loader shim: GGUF has its own quant readers, bulk-prefill path,
240
+ decode-repacked T16 layouts, and fast-path controls.
152
241
 
153
- `backend="auto"` is the public API/server default. It maps exact `gfx1100` and
154
- `gfx1151` detections to the matching HIP backend; unknown ROCm targets warn and
155
- select `cpu_reference` where a CPU implementation exists. Users on nearby targets
156
- such as `gfx1101`/`gfx1102` can force a backend with `backend="hip_gfx1100"`,
157
- `--backend hip_gfx1100`, or `HIPENGINE_BACKEND=hip_gfx1100` after validating
158
- correctness/performance.
242
+ Current caveats:
243
+
244
+ - PARO models take ~24s to load on the W7900 test host; GGUF currently takes
245
+ about 60s because decode-repack happens on load. On-disk caching could reduce
246
+ startup time later, but would require additional storage for repacked layouts.
247
+ - GGUF has higher resident memory than packed PARO. In the current W7900 README
248
+ sweep, the max-context Q4_K_S session peaks at ~25.1 GiB tracked, so 128K is
249
+ W7900/48 GiB territory; on 24 GiB cards, expect roughly 64K context with
250
+ Q4_K_S.
251
+ - GGUF is close enough to PARO to share some high-level scheduling ideas, but in
252
+ practice it needs substantial GGUF-only kernels and dispatch. The goal for
253
+ future releases is to keep closing the remaining PARO/GGUF speed gap.
159
254
 
160
- Wave32 is the default for `hip_gfx1100` device code; wave64 is treated as an
161
- isolated experiment with its own gates (see
162
- [`docs/PLAN.md`](docs/PLAN.md#rdna3-wavefront-and-scheduling-caveat)).
163
255
 
164
256
  ## Architecture at a glance
165
257
 
@@ -185,7 +277,7 @@ isolated experiment with its own gates (see
185
277
  │ KERNELS (backend-keyed, 120 __global__ in the Qwen/PARO port) │
186
278
  │ kernels/hip_gfx1100/ attention / linear_attn / moe / quant │
187
279
  │ wmma / norm / rotary / fused │
188
- │ kernels/hip_gfx1151/ native target-arch peer backend │
280
+ │ kernels/hip_gfx1151/ native target-arch peer backend │
189
281
  │ kernels/cuda_sm86/ (future) │
190
282
  │ kernels/cpu_reference/ correctness oracle, no GPU required │
191
283
  └─────────────────────────────────────────────────────────────────┘
@@ -262,6 +354,7 @@ current limitations.
262
354
  | [`docs/BENCHMARK.md`](docs/BENCHMARK.md) | Benchmark protocols, baselines, correctness gate, artifact format |
263
355
  | [`docs/TESTING.md`](docs/TESTING.md) | RED/GREEN workflow, correctness oracles, fixture policy |
264
356
  | [`docs/KERNELS.md`](docs/KERNELS.md) | Kernel catalog, source-lineage drift workflow, JIT cache gotchas, build profiles |
357
+ | [`docs/ENVS.md`](docs/ENVS.md) | Environment variables, TheRock setup, benchmark/profiling profiles |
265
358
  | [`docs/ROOFLINE.md`](docs/ROOFLINE.md) | RDNA3 / W7900 performance model and decision tree |
266
359
  | [`docs/IMPLEMENTATION.md`](docs/IMPLEMENTATION.md) | Implementation status and concrete milestones |
267
360
  | [`docs/API.md`](docs/API.md) | OpenAI-compatible server usage and endpoint support |