hipengine 0.1.0__tar.gz → 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- hipengine-0.2.0/.gitattributes +10 -0
- hipengine-0.2.0/.github/workflows/publish.yml +103 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/AGENTS.md +2 -0
- hipengine-0.2.0/CHANGELOG.md +112 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/PKG-INFO +139 -46
- {hipengine-0.1.0 → hipengine-0.2.0}/README.md +138 -45
- {hipengine-0.1.0 → hipengine-0.2.0}/WORKLOG.md +12684 -2807
- hipengine-0.2.0/benchmarks/7900XTX.md +166 -0
- hipengine-0.2.0/benchmarks/CHANGELOG.md +203 -0
- hipengine-0.2.0/benchmarks/MTP.md +115 -0
- hipengine-0.2.0/benchmarks/README.md +415 -0
- hipengine-0.2.0/benchmarks/W7900.md +117 -0
- hipengine-0.2.0/benchmarks/configs/llamacpp-mtp-qwen36-27b.json +53 -0
- hipengine-0.2.0/benchmarks/prompts/mtpbench-code-general-ja.jsonl +10 -0
- hipengine-0.2.0/benchmarks/results/2026-05-16-hipengine-gguf-q4k-pack8-bf16out-diagnostic.json +102 -0
- hipengine-0.2.0/benchmarks/results/2026-05-16-hipengine-gguf-q4k-pack8-gemv-diagnostic.json +72 -0
- hipengine-0.2.0/benchmarks/results/2026-05-16-hipengine-gguf-qwen35-e2e-correctness-diagnostic.json +71 -0
- hipengine-0.2.0/benchmarks/results/2026-05-16-hipengine-gguf-vs-paro-diagnostic.json +295 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-aotriton-v3-prefill-diagnostic.json +101 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-bulk-prefill-q4km-accepted.json +258 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-decode-graph-replay-diagnostic.json +100 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-full-attn-gpu-prelude-diagnostic.json +122 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-local-quant-coverage-diagnostic.json +145 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-prefill-projection-diagnostic.json +157 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-q4km-parity-benchmark-diagnostic.json +814 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-gguf-resident-session-diagnostic.json +106 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-512-128-paro-blocker-profile.json +685 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-bench-paro-comparison-diagnostic.json +370 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-bulk-moe-prefill-diagnostic.json +482 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-bulk-parity-diagnostic.json +261 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-decode-pack8-raw-partial.json +546 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-decode-profile-diagnostic.json +2013 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-expert-pack8-sidecar-diagnostic.json +240 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-fast-bulk-default-promoted-diagnostic.json +169 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-full-attn-parity-fixed-diagnostic.json +118 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-intake-diagnostic.json +388 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-linear-recurrent-parity-fixed-diagnostic.json +149 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-native-attention-bulk-moe-diagnostic.json +383 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-public-generate-smoke.json +135 -0
- hipengine-0.2.0/benchmarks/results/2026-05-17-hipengine-qwen36-35b-a3b-q4km-selected-device-experts-diagnostic.json +457 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-gfx1100-qwen36-27b-paro-diagnostic.json +554 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-p8_2-dense-q4k-wmma-prefill-diagnostic.json +713 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen35-int8-kv-128k-quality-perf-diagnostic.json +1062 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen35-int8-kv-256k-capacity-blocked.json +334 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen35-int8-kv-256k-single-buffer-capacity-diagnostic.json +327 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen35-int8-kv-aotriton-query-reuse-diagnostic.json +443 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen35-int8-kv-scratch-release-diagnostic.json +219 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p8-compact-moe-wmma-accepted.json +1640 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_a3-gdn-k2-chain-accepted.json +1986 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c1-wmma-tile-sweep-blocked.json +1239 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c10-combined-gap-analysis.json +309 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c11-hot-expert-final-blocked.json +334 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c3-selected-moe-profile.json +421 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c4-q4-hot-fulltile-v1-rejected.json +53 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c5-q4-sidemeta-v1-rejected.json +32 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c7-q5-opt-v1-rejected.json +40 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c8-q6-retain-legacy.json +56 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-p9_c9-tail-no-padding-not-retained.json +67 -0
- hipengine-0.2.0/benchmarks/results/2026-05-18-hipengine-qwen36-35b-a3b-q4km-prefill-q8-wmma-p8_1.json +371 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_b7-decode-gemv-rejected.json +880 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_c12-q4t16-repack-design.json +102 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_c13-q4t16-materializer.json +48 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_e2-e2e-correctness-rejected.json +303 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_h1-fastpath-safety-correctness-accepted.json +333 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_h2-decode-repack-design.json +50 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_h3-t16-512x128-bench.json +1279 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_h3-t16-e2e-correctness-accepted.json +334 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-35b-a3b-q4km-p9_h3-t16-rocprof-512x16-summary.json +1699 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-hipengine-qwen36-packed-int8-kv-readme-memory-diagnostic.json +225 -0
- hipengine-0.2.0/benchmarks/results/2026-05-19-llamacpp-mtp-qwen36-27b-diagnostic.json +537 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p10-b5-p9-e2e-gate-blocked.json +351 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p10-b6-acceptance-blocked.json +98 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p10-wave1-blocked.json +121 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_c14-q4t16-selected-wmma-prototype.json +60 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_c15-q4t16-replay-rejected.json +100 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_c16-selected-moe-alternatives.json +135 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_c17-no-q4-redesign-blocked.json +74 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d1-router-split-coop.json +82 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d10-q8t16-dual-split-64.json +110 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d11-rejected-q8t16-shared-silu.json +109 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d15-dense-dual-alpha-beta.json +115 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d16-q8t16-f32-ssm-out.json +121 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d17-key-bf16-rope.json +127 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d18-splitk-gqa-gate.json +130 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d2-bf16-key-rope-rejected.json +90 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d4-q4t16-silu-decode.json +94 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d6-q8t16-pair-dispatch.json +82 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d7-q8t16-qkv-gate-pair.json +74 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d8-q8t16-dcache-rejected.json +62 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_d9-q8t16-triple-qkv.json +80 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_g1-final-acceptance-blocked.json +183 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_h3-rejected-attn-gate-fusion.json +51 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_h3-rejected-attn128.json +57 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_h3-rejected-q4t16-silu256.json +50 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_h3-rejected-q5q6-direct-probes.json +78 -0
- hipengine-0.2.0/benchmarks/results/2026-05-20-hipengine-qwen36-35b-a3b-q4km-p9_h3-rejected-q6dense-dpreload.json +50 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-aotriton-v2-v3-diagnostic.json +155 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-d10-splitk-rocprof.json +472 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-d11-comparison-review.json +191 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-d8-splitk-decode.json +62 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-d9-splitk-sweep.json +102 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-decode-repack-residency-audit.json +96 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-long-context-chunked-smoke.json +111 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-long-context-preflight-blocked.json +65 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-memory-decode-pass-review.json +94 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-no-prefill-scratch-kv.json +44 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-q8-t16-scale-broadcast-rejected.json +55 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-r0-rocprof-baseline.json +1013 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-r1-post-x1-rocprof.json +509 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-retained-safe-mode.json +62 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-selected-moe-t16-launchbounds-rejected.json +79 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-task39-32k-smoke.json +109 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-task40-prefill-vs-paro-diagnosis.json +233 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-hipengine-qwen36-35b-a3b-q4km-p10-x1-correctness-plus-x2-wmma-blocker.json +176 -0
- hipengine-0.2.0/benchmarks/results/2026-05-21-local-rx7900xtx-gguf-vs-paro-memory-comparison.json +403 -0
- hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-after-memory-decode-pass-review.json +453 -0
- hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-direct-selected-moe-c1-4k128-accepted.json +69 -0
- hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-final-gate-4k128-accepted.json +134 -0
- hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-q8-t16-decode-probes-rejected.json +65 -0
- hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-q8-t16-second-pass-rejected.json +68 -0
- hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-router256-4k128-accepted.json +69 -0
- hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-selected-moe-down64-rejected.json +70 -0
- hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-selected-moe-t16-qk256-rejected.json +37 -0
- hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-small-kernel-second-pass-rejected.json +110 -0
- hipengine-0.2.0/benchmarks/results/2026-05-22-hipengine-qwen36-35b-a3b-q4km-q4ks-small-kernel-third-pass-rejected.json +102 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-hipengine-qwen36-35b-a3b-q4ks-persistent-session-w7900-gap-review.json +525 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-hipengine-qwen36-35b-a3b-q4ks-w7900-cold-start-diagnostic.json +144 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-hipengine-qwen36-35b-a3b-q4ks-w7900-readme-sweep-accepted.json +200 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-hipengine-qwen36-35b-a3b-q4ks-w7900-therock713-diagnostic.json +136 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-hipengine-w7900-therock713-gguf-paro-512-4k-spot-diagnostic.json +186 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-paro-512-prefill-workspace-overlap-rootcause-diagnostic.json +244 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-paro-prefill-workspace-overlap-threshold-diagnostic.json +271 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-1024-128.json +495 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-131072-128-blocked.json +22 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-32768-128.json +495 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-4096-128.json +495 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-512-128-rerun.json +479 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4km-65536-128-blocked.json +22 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-gguf-q4ks-4096-128.json +495 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-1024-128.json +6976 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-131072-128.json +6984 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-32768-128.json +6984 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-4096-128.json +6984 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-512-128.json +6976 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-bf16kv-65536-128.json +6984 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-int8kv-131072-128.json +11784 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-hipengine-paro-int8kv-65536-128.json +11784 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-llamacpp-hip-q4km-f16kv-sweep.json +1395 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-llamacpp-hip-q4km-q8kv-maxctx.json +499 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-llamacpp-vulkan-q4km-f16kv-sweep.json +1395 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-llamacpp-vulkan-q4km-q8kv-maxctx.json +499 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-paro-v011-current-regression-check.json +62 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-rocm7130423-current-head-paro-512-4k-check.json +176 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-rocm7130423-pure-current-head-paro-512-4k-check.json +196 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-rocm7130423-pure-packed-qwen36-paro-512-4k-check.json +395 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-rx7900xtx-rocm714-current-head-paro-512-4k-check.json +169 -0
- hipengine-0.2.0/benchmarks/results/2026-05-23-w7900-hipengine-therock713-paro-gguf-sweep-diagnostic.json +352 -0
- hipengine-0.2.0/benchmarks/results/2026-05-25-w7900-hipengine-gguf-q4ks-readme-persistent-5run.json +2768 -0
- hipengine-0.2.0/benchmarks/results/2026-05-25-w7900-hipengine-paro-readme-persistent-5run.json +2947 -0
- hipengine-0.2.0/benchmarks/results/2026-05-25-w7900-hipengine-readme-persistent-5run-diagnostic.json +313 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/BENCHMARK.md +61 -0
- hipengine-0.2.0/docs/ENVS.md +156 -0
- hipengine-0.2.0/docs/GGUF.md +3023 -0
- hipengine-0.2.0/docs/GGUF_DECODE_REPACK.md +333 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/KERNELS.md +121 -10
- hipengine-0.2.0/docs/KVCACHE.md +448 -0
- hipengine-0.2.0/docs/LESSONS-LEARNED.md +1068 -0
- hipengine-0.2.0/docs/OPTIMIZE-DENSE.md +423 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/PUBLISH.md +6 -1
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/README.md +5 -2
- hipengine-0.2.0/docs/RELAXED.md +651 -0
- hipengine-0.2.0/docs/ROOFLINE-gfx1151.md +805 -0
- hipengine-0.2.0/docs/SPECULATIVE-DECODE.md +1465 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/TESTING.md +55 -1
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/source_lineage.json +7 -7
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/dtype.py +2 -0
- hipengine-0.2.0/hipengine/dispatch/__init__.py +38 -0
- hipengine-0.2.0/hipengine/dispatch/kv.py +259 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/generation/__init__.py +1 -0
- hipengine-0.2.0/hipengine/generation/qwen35_gguf.py +101 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/generation/qwen35_paro.py +24 -3
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/generation/registry.py +3 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/cpu_reference/__init__.py +24 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/cpu_reference/ops.py +383 -1
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/__init__.py +16 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_wrap.py +5 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/paged_attn_decode.hip +696 -20
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/paged_attn_decode.py +579 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/paged_kv_write.hip +374 -1
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/attention/paged_kv_write.py +888 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/fused/__init__.py +24 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/fused/gguf_ops.hip +621 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/fused/gguf_ops.py +683 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear/dense_gemv.hip +84 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear/dense_gemv.py +48 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear_attn/conv.hip +12 -3
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear_attn/gdn.hip +259 -6
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear_attn/gdn.py +93 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/__init__.py +4 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/router.hip +155 -22
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/router.py +141 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/__init__.py +68 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_expert_pack8_gemv.hip +517 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_expert_pack8_gemv.py +274 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_gemv.hip +767 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_gemv.py +458 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_selected_pack8_gemv.hip +524 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_selected_pack8_gemv.py +344 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_selected_prefill.hip +800 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_selected_prefill.py +419 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_t16_selected_prefill.hip +498 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_k_t16_selected_prefill.py +289 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_gemv.hip +1078 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_gemv.py +806 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_pack8_gemv.hip +336 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_pack8_gemv.py +182 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_prefill.hip +649 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_prefill.py +319 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_selected_pack8_gemv.hip +423 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_selected_pack8_gemv.py +289 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_selected_prefill.hip +1400 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_selected_prefill.py +785 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_t16_selected_prefill.hip +419 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q4_k_t16_selected_prefill.py +314 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_embedding.hip +190 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_embedding.py +198 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_pack8_gemv.hip +349 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_pack8_gemv.py +170 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_t16_gemv.hip +185 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q6_k_t16_gemv.py +190 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_pack8_gemv.hip +514 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_pack8_gemv.py +352 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_prefill.hip +642 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_prefill.py +363 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_t16_gemv.hip +559 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_t16_gemv.py +652 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_t16_prefill.hip +333 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_q8_0_t16_prefill.py +274 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_t16_selected_gemv.hip +919 -0
- hipengine-0.2.0/hipengine/kernels/hip_gfx1100/quant/gguf_t16_selected_gemv.py +1068 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/runtime/state.py +5 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/registry.py +13 -0
- hipengine-0.2.0/hipengine/kvcache/__init__.py +30 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kvcache/policy.py +128 -3
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kvcache/spans.py +43 -1
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/llm.py +27 -6
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/loading/__init__.py +90 -0
- hipengine-0.2.0/hipengine/loading/gguf.py +327 -0
- hipengine-0.2.0/hipengine/loading/qwen35_gguf.py +443 -0
- hipengine-0.2.0/hipengine/loading/qwen35_gguf_expert_sidecar.py +577 -0
- hipengine-0.2.0/hipengine/loading/qwen35_gguf_materialize.py +514 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/models/__init__.py +12 -1
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/models/qwen35.py +58 -0
- hipengine-0.2.0/hipengine/quant/__init__.py +164 -0
- hipengine-0.2.0/hipengine/quant/gguf.py +678 -0
- hipengine-0.2.0/hipengine/quant/gguf_k.py +93 -0
- hipengine-0.2.0/hipengine/quant/gguf_q4_k.py +343 -0
- hipengine-0.2.0/hipengine/quant/gguf_t16.py +507 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/runtime/__init__.py +22 -0
- hipengine-0.2.0/hipengine/runtime/gguf_embedding.py +145 -0
- hipengine-0.2.0/hipengine/runtime/gguf_linear.py +1040 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/runtime/prefill.py +12 -0
- hipengine-0.2.0/hipengine/runtime/qwen35_gguf_runner.py +5718 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/runtime/qwen35_paro.py +316 -44
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/runtime/qwen35_paro_runner.py +590 -50
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/server/api.py +19 -0
- hipengine-0.2.0/hipengine/tokenization/__init__.py +5 -0
- hipengine-0.2.0/hipengine/tokenization/gguf.py +166 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/pyproject.toml +1 -1
- hipengine-0.2.0/scripts/gguf_k_gemv_smoke.py +304 -0
- hipengine-0.2.0/scripts/gguf_prefill_projection_smoke.py +308 -0
- hipengine-0.2.0/scripts/gguf_q6_k_embedding_smoke.py +151 -0
- hipengine-0.2.0/scripts/inspect_gguf.py +188 -0
- hipengine-0.2.0/scripts/llamacpp_mtp_bench.py +518 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_batch_packed_prefill_correctness.py +17 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_batch_serial_bench.py +14 -2
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_batch_serial_correctness.py +22 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_decode_graph_fixture_gate.py +13 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_e2e_correctness.py +16 -1
- hipengine-0.2.0/scripts/qwen35_gguf_aotriton_prefill_sweep.py +112 -0
- hipengine-0.2.0/scripts/qwen35_gguf_bench.py +920 -0
- hipengine-0.2.0/scripts/qwen35_gguf_build_expert_sidecar.py +111 -0
- hipengine-0.2.0/scripts/qwen35_gguf_bulk_parity.py +357 -0
- hipengine-0.2.0/scripts/qwen35_gguf_decode_graph_smoke.py +377 -0
- hipengine-0.2.0/scripts/qwen35_gguf_e2e_correctness.py +357 -0
- hipengine-0.2.0/scripts/qwen35_gguf_expert_pack8_smoke.py +315 -0
- hipengine-0.2.0/scripts/qwen35_gguf_moe_replay.py +801 -0
- hipengine-0.2.0/scripts/qwen35_gguf_p9_e2e_correctness.py +442 -0
- hipengine-0.2.0/scripts/qwen35_gguf_rocprof_summary.py +648 -0
- hipengine-0.2.0/scripts/qwen35_kv_e2e_fixture_gate.py +459 -0
- hipengine-0.2.0/scripts/qwen35_kv_int8_accuracy.py +870 -0
- hipengine-0.2.0/scripts/qwen35_kv_policy_args.py +80 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_correctness.py +39 -4
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_fixture_gate.py +55 -4
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_fullattn_stage_probe.py +3 -5
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_stage_probe.py +2 -7
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_paro_bench.py +38 -1
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_paro_packed_bench.py +6 -0
- hipengine-0.2.0/scripts/qwen35_readme_sweep.py +574 -0
- hipengine-0.2.0/scripts/resolve_worklog_conflict.py +186 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/smoke.py +971 -2
- hipengine-0.2.0/tests/_gguf_synthetic_weights.py +187 -0
- hipengine-0.2.0/tests/conftest.py +39 -0
- hipengine-0.2.0/tests/fixtures/cpu_reference/kv_int8_dequant_per_token_head.json +33 -0
- hipengine-0.2.0/tests/fixtures/cpu_reference/paged_attn_decode_int8_per_token_head.json +48 -0
- hipengine-0.2.0/tests/fixtures/gguf/qwen35_0_8b_q4_1_e2e.json +45 -0
- hipengine-0.2.0/tests/fixtures/gguf/qwen35_0_8b_q4_k_m_e2e.json +46 -0
- hipengine-0.2.0/tests/fixtures/gguf/qwen35_0_8b_q8_0_e2e.json +41 -0
- hipengine-0.2.0/tests/fixtures/gguf/qwen35_0_8b_ud_q4_k_xl_e2e.json +47 -0
- hipengine-0.2.0/tests/fixtures/gguf/qwen36_35b_a3b_q4km_p9_e2e.json +54 -0
- hipengine-0.2.0/tests/fixtures/gguf/qwen36_35b_a3b_q4km_smoke.json +39 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_aotriton_discovery.py +9 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_cpu_reference.py +93 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_generation_qwen35_paro.py +51 -2
- hipengine-0.2.0/tests/test_gguf_e2e_acceptance.py +131 -0
- hipengine-0.2.0/tests/test_gguf_embedding_dispatch.py +66 -0
- hipengine-0.2.0/tests/test_gguf_expert_pack8_gemv.py +261 -0
- hipengine-0.2.0/tests/test_gguf_gemv_decode_dispatch.py +344 -0
- hipengine-0.2.0/tests/test_gguf_k_gemv.py +238 -0
- hipengine-0.2.0/tests/test_gguf_k_selected_pack8_gemv_decode.py +272 -0
- hipengine-0.2.0/tests/test_gguf_k_selected_wmma_prefill.py +469 -0
- hipengine-0.2.0/tests/test_gguf_k_t16_selected_wmma_prefill.py +448 -0
- hipengine-0.2.0/tests/test_gguf_linear_dispatch.py +781 -0
- hipengine-0.2.0/tests/test_gguf_ops.py +260 -0
- hipengine-0.2.0/tests/test_gguf_q4_k_gemv.py +193 -0
- hipengine-0.2.0/tests/test_gguf_q4_k_pack8_gemv_decode.py +316 -0
- hipengine-0.2.0/tests/test_gguf_q4_k_selected_dual_pack8_gemv_decode.py +287 -0
- hipengine-0.2.0/tests/test_gguf_q4_k_selected_wmma_prefill.py +603 -0
- hipengine-0.2.0/tests/test_gguf_q4_k_t16_selected_wmma_prefill.py +235 -0
- hipengine-0.2.0/tests/test_gguf_q4_k_tile16_repack.py +74 -0
- hipengine-0.2.0/tests/test_gguf_q4_k_wmma_prefill.py +528 -0
- hipengine-0.2.0/tests/test_gguf_q6_k_embedding.py +176 -0
- hipengine-0.2.0/tests/test_gguf_q6_k_pack8_gemv_decode.py +204 -0
- hipengine-0.2.0/tests/test_gguf_q6_k_t16_gemv_decode.py +125 -0
- hipengine-0.2.0/tests/test_gguf_q8_0_pack8_gemv_decode.py +271 -0
- hipengine-0.2.0/tests/test_gguf_q8_0_t16_gemv_decode.py +590 -0
- hipengine-0.2.0/tests/test_gguf_q8_0_t16_wmma_prefill.py +315 -0
- hipengine-0.2.0/tests/test_gguf_q8_0_wmma_prefill.py +533 -0
- hipengine-0.2.0/tests/test_gguf_q8_0_wmma_prefill_dual.py +232 -0
- hipengine-0.2.0/tests/test_gguf_quant_layout.py +96 -0
- hipengine-0.2.0/tests/test_gguf_reader.py +161 -0
- hipengine-0.2.0/tests/test_gguf_t16_repack.py +179 -0
- hipengine-0.2.0/tests/test_gguf_t16_selected_gemv_decode.py +715 -0
- hipengine-0.2.0/tests/test_kv_dispatch.py +163 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_kvcache_policy.py +66 -1
- hipengine-0.2.0/tests/test_kvcache_spans.py +130 -0
- hipengine-0.2.0/tests/test_llm_gguf_generate_path.py +118 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_model_quant_and_imports.py +30 -0
- hipengine-0.2.0/tests/test_qwen35_bench_memory_audit.py +37 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_decode_state.py +222 -6
- hipengine-0.2.0/tests/test_qwen35_gguf_chunked_prefill.py +72 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_compact_moe_gemv_routing.py +468 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_compact_moe_wmma_resolver.py +140 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_compact_moe_wmma_routing.py +263 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_decode_graph_policy.py +173 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_decode_repack_dispatch.py +221 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_decode_repack_semantics.py +36 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_expert_sidecar.py +121 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_fastpath_safety.py +150 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_full_attention_gpu.py +407 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_gdn_prefill_correctness.py +695 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_gdn_prefill_routing.py +290 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_mapping.py +138 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_materialize.py +273 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_moe_replay.py +57 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_p10_x2_layer_correctness.py +163 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_p9_e2e_correctness.py +145 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_rocprof_summary.py +388 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_runner.py +154 -0
- hipengine-0.2.0/tests/test_qwen35_gguf_tokenizer.py +68 -0
- hipengine-0.2.0/tests/test_qwen35_kv_e2e_fixture_gate.py +207 -0
- hipengine-0.2.0/tests/test_qwen35_kv_int8_accuracy.py +82 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_linear_attn_gdn_plan.py +53 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_paged_attn_decode_plan.py +112 -1
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_paged_kv_write_plan.py +127 -1
- hipengine-0.2.0/tests/test_qwen35_prefill_workspace_policy.py +33 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_resident_batch_layout.py +371 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_router_plan.py +124 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_server_api.py +10 -1
- hipengine-0.1.0/.gitattributes +0 -3
- hipengine-0.1.0/CHANGELOG.md +0 -38
- hipengine-0.1.0/benchmarks/CHANGELOG.md +0 -112
- hipengine-0.1.0/benchmarks/README.md +0 -175
- hipengine-0.1.0/docs/GGUF.md +0 -576
- hipengine-0.1.0/docs/KVCACHE.md +0 -307
- hipengine-0.1.0/docs/LESSONS-LEARNED.md +0 -99
- hipengine-0.1.0/hipengine/dispatch/__init__.py +0 -18
- hipengine-0.1.0/hipengine/kernels/hip_gfx1100/attention/paged_kv_write.py +0 -453
- hipengine-0.1.0/hipengine/kvcache/__init__.py +0 -6
- hipengine-0.1.0/hipengine/quant/__init__.py +0 -28
- hipengine-0.1.0/tests/test_kvcache_spans.py +0 -62
- hipengine-0.1.0/uv.lock +0 -1602
- {hipengine-0.1.0 → hipengine-0.2.0}/.gitignore +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/CLAUDE.md +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/LICENSE +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/.gitkeep +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-13-hipengine-qwen35-paro-optimal-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-13-source-lineage-qwen35-paro-optimal-4k-128.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-13-source-lineage-qwen35-paro-optimal-512-128.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-c1-parent-fixture-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-cn-correctness-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-ab-fused-lmhead128-graph-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-c1-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-graph-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-linear-qkv-z-full-qk-fused-graph-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-linear-qkv-z-fused-graph-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-lmhead128-qk-qkvz-fused-graph-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-14-hipengine-qwen35-paro-512-128-tokenizer-cache-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-0p8b-paro-512-128-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c1-parent-fixture-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c1-parent-mixed-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c1-router-qnorm-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c1-scheduler-serial-bench-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c2-native-compact-prefill-correctness-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c2-scheduler-serial-bench-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c2-scheduler-serial-runner-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c2-serial-slot-runner-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c4-native-compact-prefill-correctness-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c4-scheduler-serial-bench-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c4-scheduler-serial-runner-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c8-native-compact-prefill-correctness-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c8-scheduler-serial-bench-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-c8-scheduler-serial-runner-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-cn-generated-equality-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-dflash-ddtree-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-linear-attn-segment-prefill-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefill-compact-c8-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefill-full-attn-boundary-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefill-full-single-request-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefill-multiloop-512-4k-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefill-plan-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer0-attention-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer0-attn-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer0-decode-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer0-gated-recurrent-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer0-stage-bisect-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-layer3-fullattn-stage-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-prefill-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-scratch-restore-sweep.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-serial-fullattn-layer4-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-serial-suffix-full40-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-native-prefix-sweep-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-prefix-bisect-blocked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-15-hipengine-qwen35-varlen-full-attn-prefill-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-aotriton-cast-glue-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-aotriton-gate-rotate-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-aotriton-threshold-sweep-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-aotriton-v3-memory-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-aotriton-v3-prefill-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-comparison-tables-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-decode-graph-replay-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-long-checkpoint-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-16-hipengine-qwen35-prefill-chunking-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-gfx1151-shisa-qwen36-packed-canonical-sweep-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-gfx1151-shisa-qwen36-packed-chunk256-sweep-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-08b-gfx1151-dense-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-35b-qwen36-27b-gfx1151-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d11-rotate-dual-pack8-fusion-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d12-rmsnorm-producer-fusion-deferred.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d13-same-input-projection-fusions-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d14-selected-moe-postop-fold-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d15-router-coop-fold-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d16-kv-pack8-fusion-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d21-marlin-k-qweight-neutral-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d31-d33-grouped-gqa-long-context-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d42-dispatch-cap-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d44-launch-bounds-deferred.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d51-gdn-decode-audit.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-d52-w8a16-decode-audit.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-p33-moe-metadata-fanout-deferred.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-p52-prefill-chunk-autotune-accepted.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p11-rocblas-ab-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p12-shared-gate-up-token-tile-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p13-shared-down-combine-token-tile-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p14-moe-wmma-threshold-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p16-prefill-mcumode-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p31-gdn-rotate-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-p32-router-sigmoid-rejected.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-paro-dual-format-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-qwen36-w1-unroll600-ablation-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen35-rocprof-amdahl-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen36-packed-shared-decode-fusion-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen36-shisa-force-legacy-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-hipengine-qwen36-shisa-packed-vs-legacy-refresh-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-llamacpp-hip-qwen36-peak.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-llamacpp-upstream-gfx1151-qwen36-gguf-rerun-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-17-llamacpp-vulkan-qwen36-peak.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-18-hipengine-gfx1100-shisa-qwen36-packed-gt1k-default-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/benchmarks/results/2026-05-18-hipengine-qwen35-gt1k-prefill-chunk-policy-diagnostic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/API.md +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/DFLASH.md +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/IMPLEMENTATION.md +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/MARLIN.md +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/MTP.md +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/OPTIMIZE.md +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/PLAN.md +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/PREFILL.md +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/docs/ROOFLINE.md +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/fixtures/qwen35_paro/parent_512_32_seed1234.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hatch_build.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/benchmark/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/benchmark/correctness.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/build.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/device.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/hip.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/memory.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/rocblas.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/core/tensor.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/dispatch/batch.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/dispatch/fusion.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/distributed/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/generation/batch_scheduler.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/backends.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/cpu_reference/fixtures.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/cuda_sm86/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_release.toml +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/MANIFEST.vendor.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/aiter_hip_common.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/flash/aiter.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/kernel_cluster.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/lazy_tensor_internal.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/packed_kernel.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/triton_kernel.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/_internal/util.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/config.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/cpp_tune.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/dtypes.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/flash.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/runtime.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/util.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/include/aotriton/v2/flash.h +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_F_0_0___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_F_0_1___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_F_3_0___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_T_0_0___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_T_0_1___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_F_T_3_0___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_F_0_0___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_F_0_1___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_F_3_0___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_T_0_0___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_T_0_1___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/aotriton.images/amd-gfx11xx/flash/attn_fwd/FONLY__/357/274/212bf16@16_256_T_T_3_0___gfx11xx.aks2" +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/libaotriton_v2.so +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/0.11.2b/lib/libaotriton_v2.so.0.11.2 +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/attention/aotriton_wrap.cc +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/common/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/convert/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/convert/cast.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/convert/cast.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/fused/paro_combine.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/fused/paro_combine.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/fused/paro_silu.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/fused/paro_silu.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear/lm_head.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear/lm_head.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear_attn/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/linear_attn/conv.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/group_scatter.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/group_scatter.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/moe/prefill.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/norm/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/norm/rmsnorm.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/norm/rmsnorm.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/paro_awq_gemv.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/paro_awq_gemv.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/paro_marlin_k.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/paro_marlin_k.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/w8a16_linear.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/quant/w8a16_linear.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/rotary/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/rotary/paro_rotate.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/rotary/paro_rotate.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/rotary/qwen35_rotary.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/rotary/qwen35_rotary.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/runtime/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/runtime/state.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/smoke/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/smoke/smoke_add.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/smoke/smoke_add.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/wmma/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/wmma/paro_awq_wmma.hip +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1100/wmma/paro_awq_wmma.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/kernels/hip_gfx1151/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/layers/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/layers/base.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/loading/materialize.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/loading/qwen35_paro.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/loading/safetensors.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/models/base.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/models/registry.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/models/toy.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/quant/base.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/quant/bf16.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/quant/fp16.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/quant/registry.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/quant/w4_paro.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/runtime/workspace.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/server/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/server/__main__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/speculative/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/speculative/interfaces.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/util/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/hipengine/util/amdgpu_vram.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/__init__.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/check_fixtures.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/check_lineage.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/fetch_aotriton.sh +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/gdn_decode_probe.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/llamacpp_bench_with_peak.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_batch_correctness.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_compare_tables.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_dflash_ddtree_blocker.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_compact_prefill_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_boundary.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_native_prefill_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_paro_next_token.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/qwen35_rocprof_audit.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/strip_paro_safetensors.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/vendor_aotriton.sh +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/scripts/w8a16_decode_probe.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/fixtures/cpu_reference/attention_decode_masked.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/fixtures/cpu_reference/full_attn_prefill_causal_gqa_gate.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/fixtures/cpu_reference/linear_basic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/fixtures/cpu_reference/rmsnorm_basic.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/fixtures/cpu_reference/rotate_split_half.json +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_build.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_cast_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_check_lineage.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_dense_gemv_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_dispatch_batch.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_fusion_spike.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_generation_batch_scheduler.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_gfx1151_backend.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_hip_runtime.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_kernel_registry.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_llm_generate.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_lm_head_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_loading_materialize.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_loading_safetensors.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_memory_stats.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_paro_awq_gemv_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_paro_awq_wmma_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_paro_combine_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_paro_rotate_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_paro_silu_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_linear_attn_conv_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_moe_group_scatter_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_native_prefill_boundary.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_native_prefill_fullattn_stage_probe.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_paro_layout.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_paro_marlin_k.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_rmsnorm_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_qwen35_rotary_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_runtime_state_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_runtime_workspace.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_smoke_add_plan.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_speculative_interfaces.py +0 -0
- {hipengine-0.1.0 → hipengine-0.2.0}/tests/test_w8a16_linear_plan.py +0 -0
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
# WORKLOG.md is append-only: when two branches both add new entries, git's
|
|
2
|
+
# built-in union driver concatenates both sides instead of emitting conflict
|
|
3
|
+
# markers. For an append-only file this is the correct semantics (common
|
|
4
|
+
# prefix + ours-tail + theirs-tail). See scripts/resolve_worklog_conflict.py
|
|
5
|
+
# for the recovery path when conflict markers already exist in the file.
|
|
6
|
+
WORKLOG.md merge=union
|
|
7
|
+
|
|
8
|
+
hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/** -whitespace
|
|
9
|
+
hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/**/*.aks2 filter=lfs diff=lfs merge=lfs -text
|
|
10
|
+
hipengine/kernels/hip_gfx1100/attention/aotriton_runtime/**/*.so* filter=lfs diff=lfs merge=lfs -text
|
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
name: Publish to PyPI
|
|
2
|
+
|
|
3
|
+
on:
|
|
4
|
+
push:
|
|
5
|
+
tags:
|
|
6
|
+
- "v*"
|
|
7
|
+
|
|
8
|
+
permissions: {}
|
|
9
|
+
|
|
10
|
+
jobs:
|
|
11
|
+
build:
|
|
12
|
+
runs-on: ubuntu-24.04
|
|
13
|
+
permissions:
|
|
14
|
+
contents: read
|
|
15
|
+
outputs:
|
|
16
|
+
version: ${{ steps.meta.outputs.version }}
|
|
17
|
+
|
|
18
|
+
steps:
|
|
19
|
+
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
20
|
+
with:
|
|
21
|
+
persist-credentials: false
|
|
22
|
+
lfs: true
|
|
23
|
+
|
|
24
|
+
- name: Install uv
|
|
25
|
+
uses: astral-sh/setup-uv@cec208311dfd045dd5311c1add060b2062131d57 # v8.0.0
|
|
26
|
+
|
|
27
|
+
- name: Set up Python 3.12
|
|
28
|
+
run: |
|
|
29
|
+
uv python install 3.12
|
|
30
|
+
uv venv --python 3.12
|
|
31
|
+
|
|
32
|
+
- name: Extract version metadata
|
|
33
|
+
id: meta
|
|
34
|
+
run: |
|
|
35
|
+
version="$(uv run --no-sync python - <<'PY'
|
|
36
|
+
from pathlib import Path
|
|
37
|
+
import tomllib
|
|
38
|
+
|
|
39
|
+
data = tomllib.loads(Path("pyproject.toml").read_text())
|
|
40
|
+
print(data["project"]["version"])
|
|
41
|
+
PY
|
|
42
|
+
)"
|
|
43
|
+
echo "version=${version}" >> "$GITHUB_OUTPUT"
|
|
44
|
+
|
|
45
|
+
- name: Verify tag matches package version
|
|
46
|
+
env:
|
|
47
|
+
RELEASE_TAG: ${{ github.ref_name }}
|
|
48
|
+
RELEASE_VERSION: ${{ steps.meta.outputs.version }}
|
|
49
|
+
run: |
|
|
50
|
+
expected="v${RELEASE_VERSION}"
|
|
51
|
+
if [ "${RELEASE_TAG}" != "${expected}" ]; then
|
|
52
|
+
echo "::error::Tag '${RELEASE_TAG}' does not match package version '${expected}'"
|
|
53
|
+
exit 1
|
|
54
|
+
fi
|
|
55
|
+
|
|
56
|
+
- name: Install dependencies
|
|
57
|
+
run: uv pip install -e '.[dev]' build twine
|
|
58
|
+
|
|
59
|
+
- name: Run release validation
|
|
60
|
+
run: uv run --no-sync pytest -q
|
|
61
|
+
|
|
62
|
+
- name: Remove stale build artifacts
|
|
63
|
+
run: rm -rf dist/
|
|
64
|
+
|
|
65
|
+
- name: Build
|
|
66
|
+
run: uv run --no-sync python -m build
|
|
67
|
+
|
|
68
|
+
- name: Verify package metadata
|
|
69
|
+
run: uv run --no-sync python -m twine check dist/*
|
|
70
|
+
|
|
71
|
+
- name: Upload build artifacts
|
|
72
|
+
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
|
|
73
|
+
with:
|
|
74
|
+
name: dist
|
|
75
|
+
path: |
|
|
76
|
+
dist/hipengine-*.whl
|
|
77
|
+
dist/hipengine-*.tar.gz
|
|
78
|
+
|
|
79
|
+
publish:
|
|
80
|
+
needs: build
|
|
81
|
+
runs-on: ubuntu-24.04
|
|
82
|
+
environment: pypi-publish
|
|
83
|
+
permissions:
|
|
84
|
+
id-token: write
|
|
85
|
+
contents: read
|
|
86
|
+
attestations: write
|
|
87
|
+
|
|
88
|
+
steps:
|
|
89
|
+
- name: Download build artifacts
|
|
90
|
+
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
|
|
91
|
+
with:
|
|
92
|
+
name: dist
|
|
93
|
+
path: dist/
|
|
94
|
+
|
|
95
|
+
- name: Generate artifact attestations
|
|
96
|
+
uses: actions/attest-build-provenance@a2bbfa25375fe432b6a289bc6b6cd05ecd0c4c32 # v4.1.0
|
|
97
|
+
with:
|
|
98
|
+
subject-path: |
|
|
99
|
+
dist/hipengine-*.whl
|
|
100
|
+
dist/hipengine-*.tar.gz
|
|
101
|
+
|
|
102
|
+
- name: Publish to PyPI (trusted publishing)
|
|
103
|
+
uses: pypa/gh-action-pypi-publish@cef221092ed1bacb1cc03d23a2d87d1d172e277b # v1.14.0
|
|
@@ -64,6 +64,7 @@ Do not drift these casually. They define what hipEngine is.
|
|
|
64
64
|
|
|
65
65
|
- Keep changes scoped to one logical unit (one kernel family, one plugin, one doc, one phase milestone).
|
|
66
66
|
- Write or update the targeted test/fixture before implementation when behavior or math changes. If RED-first is impractical, record why in `WORKLOG.md`.
|
|
67
|
+
- When adding tests that call HIP/ROCm runtime, `hipcc`, or GPU kernels, add an explicit HIP-availability guard (for example `ctypes.CDLL("libamdhip64.so")` + `pytest.skip`) so no-ROCm CI/publish runners skip them instead of failing release validation.
|
|
67
68
|
- Log non-trivial decisions, measurements, and dependency additions in `WORKLOG.md` as they happen.
|
|
68
69
|
- When profiling Python/ctypes JIT-built kernels with `rocprofv3`, prebuild the `.so` outside the profiler and run the profiled command with a precomputed compiler-version file plus `require_cached`; do not let the profiled process spawn `hipcc`/clang.
|
|
69
70
|
- Do not silently add `import torch`, `flash_attn`, or other CUDA-only deps to hot-path modules.
|
|
@@ -144,6 +145,7 @@ Working tree is shared state. Other agents or the human may be editing concurren
|
|
|
144
145
|
- **High-conflict files:** `AGENTS.md`, `CLAUDE.md`, `docs/PLAN.md`, `docs/BENCHMARK.md`, `docs/TESTING.md`, `docs/KERNELS.md`, `docs/IMPLEMENTATION.md`, `WORKLOG.md`, `pyproject.toml`, `hipengine/kernels/registry.py`, `hipengine/quant/registry.py`, `hipengine/models/registry.py`, `hipengine/dispatch/fusion.py`, `hipengine/core/*`.
|
|
145
146
|
- Same-file contention: stop and coordinate. The designated agent stages and commits their scoped hunks first to unblock others.
|
|
146
147
|
- `WORKLOG.md` appends are expected and not a conflict unless there are actual conflict markers or interleaved garbled lines. Re-read the live tail, append after it, commit with your logical unit.
|
|
148
|
+
- `WORKLOG.md` is configured with git's built-in `merge=union` driver (see `.gitattributes`), so concurrent appends auto-resolve as `common prefix + ours-tail + theirs-tail` with no conflict markers. If markers do appear (e.g. a stash or a rebase started before this was configured), run `python3 scripts/resolve_worklog_conflict.py WORKLOG.md` (add `--sort-by-date` to re-order `## YYYY-MM-DD` sections in each resolved block; `--check` for a pre-commit gate). The script only touches conflict blocks; content outside markers is left exactly as-is.
|
|
147
149
|
- Do not clean up another agent's benchmark outputs, staged files, or local artifacts unless the task explicitly asks for that cleanup.
|
|
148
150
|
|
|
149
151
|
## External Reference Repos
|
|
@@ -0,0 +1,112 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable user-facing changes for hipEngine releases are documented here.
|
|
4
|
+
|
|
5
|
+
This changelog is for package/API releases. Performance rollup history remains in
|
|
6
|
+
[`benchmarks/CHANGELOG.md`](benchmarks/CHANGELOG.md), with detailed benchmark
|
|
7
|
+
evidence under [`benchmarks/results/`](benchmarks/results/).
|
|
8
|
+
|
|
9
|
+
## v0.2.0 - 2026-05-25
|
|
10
|
+
|
|
11
|
+
Minor release for the GGUF runtime path and W7900 benchmark refresh. GGUF is a
|
|
12
|
+
meaningful new model-loading surface rather than a patch-level fix, so this
|
|
13
|
+
supersedes the previously planned v0.1.2 patch.
|
|
14
|
+
|
|
15
|
+
### Added
|
|
16
|
+
|
|
17
|
+
- Added Qwen3.6 35B MoE GGUF support for `Q4_K_M` and `Q4_K_S` model files,
|
|
18
|
+
including resident GGUF loading, bulk prefill, graph-replay decode,
|
|
19
|
+
decode-repacked T16 layouts, and WMMA/GEMV fast-path controls used by the
|
|
20
|
+
W7900 benchmark profile.
|
|
21
|
+
- Added `docs/ENVS.md` as the canonical environment-variable reference, including
|
|
22
|
+
TheRock ROCm process setup, cached-build profiling guidance, and safe GGUF
|
|
23
|
+
benchmark profiles.
|
|
24
|
+
- Added a persistent README sweep harness that loads each hipEngine model once
|
|
25
|
+
and runs repeated in-session workload measurements, matching llama-bench-style
|
|
26
|
+
repetition without multiplying model load/decode-repack time by every shape.
|
|
27
|
+
|
|
28
|
+
### Changed
|
|
29
|
+
|
|
30
|
+
- Refreshed W7900 README performance tables with 5-run persistent-session medians
|
|
31
|
+
for packed PARO and GGUF Q4_K_S while keeping the existing llama.cpp HIP/Vulkan
|
|
32
|
+
comparison rows unchanged.
|
|
33
|
+
- Documented the current GGUF tradeoffs: higher one-time load cost and resident
|
|
34
|
+
memory from decode-repack, Q4_K_S preferred for tighter VRAM budgets, and
|
|
35
|
+
performance still behind PARO on some shapes while already competitive in the
|
|
36
|
+
broader W7900 comparison.
|
|
37
|
+
|
|
38
|
+
### Fixed
|
|
39
|
+
|
|
40
|
+
- Fixed the PARO resident prefill workspace-overlap regression that shipped in
|
|
41
|
+
v0.1.1: short and mid prompts now keep prefill workspaces resident through
|
|
42
|
+
32K tokens, restoring 512/128-class prefill throughput while retaining the
|
|
43
|
+
long-context memory-saving path for prompts above 32K when active chunking
|
|
44
|
+
splits the prompt.
|
|
45
|
+
- Fixed GGUF non-split full-attention decode in max-context persistent sessions
|
|
46
|
+
by launching the context kernel with the active decode context instead of the
|
|
47
|
+
session's maximum allocation length.
|
|
48
|
+
|
|
49
|
+
### Known limitations
|
|
50
|
+
|
|
51
|
+
- GGUF support remains alpha: production correctness and performance coverage is
|
|
52
|
+
strongest for the documented Qwen3.6 35B MoE Q4_K_M/Q4_K_S paths on gfx1100,
|
|
53
|
+
and other GGUF quants/models require local validation.
|
|
54
|
+
- GGUF model load is slower than packed PARO on the same host because current
|
|
55
|
+
decode-repack happens on load and is not yet cached on disk.
|
|
56
|
+
|
|
57
|
+
## v0.1.1 - 2026-05-19
|
|
58
|
+
|
|
59
|
+
Patch release focused on long-context memory documentation and the INT8 KV cache
|
|
60
|
+
bring-up that landed after v0.1.0.
|
|
61
|
+
|
|
62
|
+
### Added
|
|
63
|
+
|
|
64
|
+
- INT8 KV cache policy controls and dispatch coverage for Qwen/PARO resident
|
|
65
|
+
inference paths, including CPU/layer/E2E correctness gates and memory audits.
|
|
66
|
+
- Documented Qwen3.6 packed PARO memory rows for 128K BF16 KV, 128K INT8 KV, and
|
|
67
|
+
256K INT8 KV on W7900/gfx1100, with retained-KV and loaded-weight VRAM notes.
|
|
68
|
+
|
|
69
|
+
### Changed
|
|
70
|
+
|
|
71
|
+
- Reduced the 256K INT8 KV tracked allocator high-water mark below the 24 GiB
|
|
72
|
+
class target by releasing/reusing prefill scratch and AOTriton query buffers.
|
|
73
|
+
- Clarified that packed vs unstripped PARO checkpoint size does not translate to
|
|
74
|
+
meaningfully different resident model-weight VRAM for the current text runtime.
|
|
75
|
+
|
|
76
|
+
### Known limitations
|
|
77
|
+
|
|
78
|
+
- INT8 KV correctness is gated by deterministic fixtures and layer probes; it is
|
|
79
|
+
not yet a long-rollout perplexity or compounding-error study.
|
|
80
|
+
- Qwen3.6 packed throughput rows remain diagnostic pending a promoted public
|
|
81
|
+
`LLM.generate()` correctness/repetition gate.
|
|
82
|
+
|
|
83
|
+
## v0.1.0 - 2026-05-18
|
|
84
|
+
|
|
85
|
+
Initial public alpha release.
|
|
86
|
+
|
|
87
|
+
### Added
|
|
88
|
+
|
|
89
|
+
- Torch-free Python runtime hot path for local ROCm inference bring-up.
|
|
90
|
+
- Plugin registries keyed by model/backend/quant/layer variants.
|
|
91
|
+
- HIP backends for `gfx1100` and `gfx1151`, plus `backend="auto"` detection with
|
|
92
|
+
`HIPENGINE_BACKEND` force override guidance for nearby targets.
|
|
93
|
+
- Qwen3.5/Qwen3.6 PARO W4 runtime path, JIT HIP build/cache plumbing, AOTriton
|
|
94
|
+
prefill runtime packaging, and OpenAI-compatible server entry point.
|
|
95
|
+
- CPU reference kernels and focused correctness/performance documentation.
|
|
96
|
+
|
|
97
|
+
### Packaging
|
|
98
|
+
|
|
99
|
+
- PyPI project name: `hipengine`.
|
|
100
|
+
- Python import package: `hipengine`.
|
|
101
|
+
- Canonical repository/wordmark: `hipEngine`.
|
|
102
|
+
- Release wheels are Linux x86-64 `manylinux_2_39` platform wheels because the
|
|
103
|
+
package bundles a ROCm/AOTriton shared-library runtime; ROCm runtime libraries
|
|
104
|
+
remain external system dependencies.
|
|
105
|
+
|
|
106
|
+
### Known limitations
|
|
107
|
+
|
|
108
|
+
- Alpha-quality API and model coverage; expect sharp edges outside the documented
|
|
109
|
+
Qwen/PARO paths.
|
|
110
|
+
- Default supported GPU targets are `gfx1100` and `gfx1151`; other AMD targets
|
|
111
|
+
require explicit backend forcing and local validation.
|
|
112
|
+
- Model weights are not distributed with the package.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: hipengine
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.2.0
|
|
4
4
|
Summary: ROCm-native local LLM inference engine with a torch-free runtime hot path
|
|
5
5
|
Project-URL: Homepage, https://github.com/shisa-ai/hipEngine
|
|
6
6
|
Project-URL: Repository, https://github.com/shisa-ai/hipEngine
|
|
@@ -75,52 +75,143 @@ supported GPUs and models.
|
|
|
75
75
|
|
|
76
76
|
## Status
|
|
77
77
|
|
|
78
|
-
**v0.
|
|
79
|
-
|
|
80
|
-
[`hipengine/kernels/hip_gfx1100/`](hipengine/kernels/hip_gfx1100/). Current
|
|
81
|
-
single-model tuning targets
|
|
78
|
+
**v0.2.0 alpha.** The runtime hot path is torch-free by construction, and the
|
|
79
|
+
first two 35B-class model-loading surfaces are now available on gfx1100:
|
|
82
80
|
[shisa-ai/Qwen3.6-35B-A3B-PARO-full4096-e5-packed](https://huggingface.co/shisa-ai/Qwen3.6-35B-A3B-PARO-full4096-e5-packed)
|
|
83
81
|
(19.07 GiB, 4.68 bpw) in packed
|
|
84
|
-
[ParoQuant](https://github.com/shisa-ai/paroquant) format.
|
|
82
|
+
[ParoQuant](https://github.com/shisa-ai/paroquant) format, plus Qwen3.6 GGUF
|
|
83
|
+
`Q4_K_M` / `Q4_K_S` files through the new resident GGUF path.
|
|
85
84
|
|
|
86
|
-
|
|
85
|
+
- INT8 KV cache support has been added for PARO. Qwen 3 MoE's full 256K context window can fit in <24GB tracked memory; see [Memory Usage](#memory-usage).
|
|
86
|
+
- Qwen 3.6 [Q4_K_M](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF?show_file_info=Qwen3.6-35B-A3B-UD-Q4_K_M.gguf) and [Q4_K_S](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF?show_file_info=Qwen3.6-35B-A3B-UD-IQ4_XS.gguf) GGUF support has landed (W7900 Q4_K_S sweep is in [Performance](#performance) alongside packed PARO and llama.cpp Q4_K_M HIP/Vulkan baselines). GGUF uses a substantial GGUF-specific runtime path with bulk prefill, graph decode, and on-load decode-repack into T16 tile layouts. Q4_K_S is recommended on 24 GiB cards because Q4_K_M is bigger; on the 48 GiB W7900 Q4_K_S fits all the way to 128K context, while on 24 GiB cards expect roughly 64K. GGUF also has a higher per-session load cost (~60 s vs ~24 s for PARO packed on the same hardware) for the same decode-repack reason.
|
|
87
|
+
- Current gfx1100 performance snapshots are summarized in [Performance](#performance) and compared against recent llama.cpp Q4_K_M baselines.
|
|
87
88
|
|
|
88
|
-
While we are far from [gfx1100 roofline](https://github.com/shisa-ai/hipEngine/blob/main/docs/ROOFLINE.md), the current gfx1100 implementation does well compared to Q4_K_M quants of recent llama.cpp builds (`b9042`) on the same model family. The latest W7900 packed rows use the default prefill policy: 512-token prompts stay unchunked and prompts above 1K use `1024/1024/4096/1024/1024` chunks.
|
|
89
89
|
|
|
90
|
-
|
|
90
|
+
## Hardware targets
|
|
91
|
+
|
|
92
|
+
| Backend | Hardware | Status |
|
|
93
|
+
| --- | --- | --- |
|
|
94
|
+
| `cpu_reference` | Any CPU, numpy | Correctness oracle; CI without GPU |
|
|
95
|
+
| `hip_gfx1100` | AMD Radeon Pro W7900 / RX 7900 XTX (RDNA3) | Active backend |
|
|
96
|
+
| `hip_gfx1151` | AMD Ryzen AI MAX+ 395 / Radeon 8060S (Strix Halo, RDNA3.5) | Active backend |
|
|
97
|
+
| `cuda_sm86` | NVIDIA Ampere consumer (3090-class) | Planned peer backend |
|
|
91
98
|
|
|
92
|
-
|
|
99
|
+
`backend="auto"` is the public API/server default. It maps exact `gfx1100` and
|
|
100
|
+
`gfx1151` detections to the matching HIP backend; unknown ROCm targets warn and
|
|
101
|
+
select `cpu_reference` where a CPU implementation exists. Users on nearby targets
|
|
102
|
+
such as `gfx1101`/`gfx1102` can force a backend with `backend="hip_gfx1100"`,
|
|
103
|
+
`--backend hip_gfx1100`, or `HIPENGINE_BACKEND=hip_gfx1100` after validating
|
|
104
|
+
correctness/performance.
|
|
105
|
+
|
|
106
|
+
Wave32 is the default for `hip_gfx1100` device code; wave64 is treated as an
|
|
107
|
+
isolated experiment with its own gates (see
|
|
108
|
+
[`docs/PLAN.md`](docs/PLAN.md#rdna3-wavefront-and-scheduling-caveat)).
|
|
109
|
+
|
|
110
|
+
## Memory Usage
|
|
111
|
+
|
|
112
|
+
With BF16 KV cache, hipEngine running the packed Qwen 3.6 PARO model fits a
|
|
113
|
+
128K context window in a 24GB-class memory budget. The INT8 KV cache option
|
|
114
|
+
(with FP16 per-token/per-head scales) uses the
|
|
115
|
+
`--kv-storage int8_per_token_head` flag and lets the **full 256K context** fit
|
|
116
|
+
under 24 GiB tracked allocator peak.
|
|
117
|
+
|
|
118
|
+
The numbers below are for
|
|
119
|
+
`shisa-ai/Qwen3.6-35B-A3B-PARO-full4096-e5-packed` on W7900/gfx1100 with q3072
|
|
120
|
+
full-attention prefill chunks:
|
|
121
|
+
|
|
122
|
+
| Model | Context | KV cache | Sampled peak | Allocator peak | Retained KV | Prefill | Decode |
|
|
123
|
+
| -------------------- | ------: | -------- | -----------: | -------------: | ----------: | -----------: | ---------: |
|
|
124
|
+
| Qwen3.6 35B-A3B PARO | 128K | BF16 | 21.04 GiB | 21.88 GiB | 2.69 GiB | 1091.9 tok/s | 62.2 tok/s |
|
|
125
|
+
| Qwen3.6 35B-A3B PARO | 128K | INT8 | 19.80 GiB | 20.89 GiB | 1.36 GiB | 1076.5 tok/s | 60.0 tok/s |
|
|
126
|
+
| Qwen3.6 35B-A3B PARO | 256K | INT8 | 21.96 GiB | 23.71 GiB | 2.71 GiB | 670.2 tok/s | 40.3 tok/s |
|
|
127
|
+
|
|
128
|
+
Regardless of the difference in PARO weight storage (legacy or packed),
|
|
129
|
+
loaded-weight memory is about the same — approximately 16.4 GiB in VRAM.
|
|
130
|
+
|
|
131
|
+
The INT8 KV correctness gate is currently the deterministic Qwen3.5 PARO
|
|
132
|
+
fixture `fixtures/qwen35_paro/parent_512_32_seed1234.json` (512-token prompt,
|
|
133
|
+
32 greedy decode tokens): `max_kl=0.015328`, `mean_kl=0.001639`, top-1 agreement
|
|
134
|
+
100%, and generated IDs match BF16 KV exactly. Layer attention probes at context
|
|
135
|
+
64 and 520 also had top-1 agreement 100% with max quantized-vs-BF16 KL
|
|
136
|
+
`2.34e-7`. This is a fixture/regression gate, not a long-rollout perplexity
|
|
137
|
+
study, so long context generations may have unmeasured compounding errors.
|
|
138
|
+
|
|
139
|
+
The same 128K/128 Qwen3.5 BF16-vs-INT8 run measured -0.99% prefill tok/s and
|
|
140
|
+
-3.20% decode tok/s for INT8 KV, so speed loss is also very small.
|
|
141
|
+
|
|
142
|
+
See
|
|
143
|
+
[`benchmarks/results/2026-05-19-hipengine-qwen36-packed-int8-kv-readme-memory-diagnostic.json`](benchmarks/results/2026-05-19-hipengine-qwen36-packed-int8-kv-readme-memory-diagnostic.json),
|
|
144
|
+
[`benchmarks/README.md`](benchmarks/README.md#blocked--diagnostic-benchmark-attempts),
|
|
145
|
+
and [`docs/KVCACHE.md`](docs/KVCACHE.md) for commands, artifacts, and the full
|
|
146
|
+
no-shadow memory audit.
|
|
147
|
+
|
|
148
|
+
### llama.cpp
|
|
149
|
+
|
|
150
|
+
When run with `q8_0` kvcache, llama.cpp can also fit in 24GB:
|
|
151
|
+
|
|
152
|
+
```bash
|
|
153
|
+
--flash-attn on -ctk q8_0 -ctv q8_0 -c 262144 -b 128 -ub 128
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
Results:
|
|
157
|
+
|
|
158
|
+
| Model | llama.cpp model buffer | KV cache | Compute buffer | rocm-smi VRAM used | Free VRAM |
|
|
159
|
+
| --- | ---: | ---: | ---: | ---: | ---: |
|
|
160
|
+
| Q4_K_M | 20583 MiB | 2720 MiB | 203 MiB | 24017 MiB / 23.45 GiB | ~543 MiB |
|
|
161
|
+
| Q4_K_S | 19399 MiB | 2720 MiB | 203 MiB | 22832 MiB / 22.30 GiB | ~1728 MiB |
|
|
162
|
+
|
|
163
|
+
With `-ub 512`:
|
|
164
|
+
|
|
165
|
+
| Model | Compute buffer | rocm-smi VRAM used | Free VRAM |
|
|
93
166
|
| --- | ---: | ---: | ---: |
|
|
94
|
-
|
|
|
95
|
-
|
|
|
96
|
-
|
|
97
|
-
|
|
167
|
+
| Q4_K_M | 812 MiB | 24540 MiB | ~20 MiB |
|
|
168
|
+
| Q4_K_S | 812 MiB | 23443 MiB | ~1117 MiB |
|
|
169
|
+
|
|
170
|
+
- Note Q4_K_M is incredibly tight with only 20 MiB of headroom and you may either need to resize down or set `-b 512 -ub 128`.
|
|
171
|
+
- Q4_K_S does not need small `-b`/`-ub`; `-ub 512` fits fine, and can even increase to `-b 2048` (but `-ub` is the more important VRAM knob that controls the physical microbatch / compute buffer size for llama.cpp).
|
|
172
|
+
|
|
173
|
+
## Performance
|
|
174
|
+
|
|
175
|
+
### gfx1100 (Radeon RX 7900 XTX / Radeon Pro W7900)
|
|
176
|
+
|
|
177
|
+
While we are far from [gfx1100 roofline](https://github.com/shisa-ai/hipEngine/blob/main/docs/ROOFLINE.md), the current gfx1100 implementation does well compared to Q4_K_M quants of recent llama.cpp builds (`b9042`) on the same model family. The latest W7900 hipEngine rows use TheRock ROCm 7.13 and load each resident model once for 1 warmup + 5 measured in-session repetitions per shape. PARO uses the default prefill policy: 512-token prompts stay unchunked and prompts above 1K use `1024/1024/4096/1024/1024` chunks. The `hipEngine GGUF Q4_K_S` column uses the same chunked-prefill policy plus the WMMA prefill + GEMV decode fast paths and the persistent on-load decode-repack into T16 tile layouts.
|
|
178
|
+
|
|
179
|
+
### Prefill tok/s
|
|
180
|
+
|
|
181
|
+
| Workload | hipEngine PARO | hipEngine GGUF Q4_K_S | llama.cpp HIP | llama.cpp Vulkan |
|
|
182
|
+
| --- | ---: | ---: | ---: | ---: |
|
|
183
|
+
| 512/128 | **2718.497** | 2258.847 | 2436.049 | 1816.927 |
|
|
184
|
+
| 4K/128 | **2838.773** | 2576.673 | 2176.905 | 1705.093 |
|
|
185
|
+
| 32K/128 | **2074.699** | 1893.967 | 1496.409 | 1128.554 |
|
|
186
|
+
| 128K/128 | **1055.454** | 998.143 | 710.213 | 480.539 |
|
|
98
187
|
|
|
99
188
|
### Decode tok/s
|
|
100
189
|
|
|
101
|
-
| Workload | hipEngine
|
|
102
|
-
| --- | ---: | ---: | ---: |
|
|
103
|
-
| 512/128 |
|
|
104
|
-
| 4K/128 |
|
|
105
|
-
| 32K/128 |
|
|
106
|
-
| 128K/128 |
|
|
190
|
+
| Workload | hipEngine PARO | hipEngine GGUF Q4_K_S | llama.cpp HIP | llama.cpp Vulkan |
|
|
191
|
+
| --- | ---: | ---: | ---: | ---: |
|
|
192
|
+
| 512/128 | 103.460 | 109.152 | 85.487 | **127.515** |
|
|
193
|
+
| 4K/128 | 101.964 | 100.048 | 87.375 | **120.163** |
|
|
194
|
+
| 32K/128 | 90.438 | 86.774 | 76.994 | **98.073** |
|
|
195
|
+
| 128K/128 | 59.598 | 57.954 | 57.341 | **64.478** |
|
|
107
196
|
|
|
108
197
|
### Peak GiB
|
|
109
198
|
|
|
110
|
-
| Workload | hipEngine
|
|
111
|
-
| --- | ---: | ---: | ---: |
|
|
112
|
-
| 512/128 |
|
|
113
|
-
| 4K/128 |
|
|
114
|
-
| 32K/128 |
|
|
115
|
-
| 128K/128 | **
|
|
199
|
+
| Workload | hipEngine PARO | hipEngine GGUF Q4_K_S | llama.cpp HIP | llama.cpp Vulkan |
|
|
200
|
+
| --- | ---: | ---: | ---: | ---: |
|
|
201
|
+
| 512/128 | 20.962 | 25.108 | 21.125 | **20.844** |
|
|
202
|
+
| 4K/128 | 21.906 | 25.108 | 21.197 | **20.969** |
|
|
203
|
+
| 32K/128 | 22.016 | 25.108 | 21.738 | **21.533** |
|
|
204
|
+
| 128K/128 | **22.122** | 25.108 | 23.605 | 23.596 |
|
|
205
|
+
|
|
206
|
+
hipEngine W7900 row source: [`benchmarks/results/2026-05-25-w7900-hipengine-readme-persistent-5run-diagnostic.json`](benchmarks/results/2026-05-25-w7900-hipengine-readme-persistent-5run-diagnostic.json). Both hipEngine columns are 5-run medians from one resident session allocated for the maximum requested context (`128K/128`), so the peak-memory column is a max-context persistent-session high-water mark rather than each shape's minimum allocation. Existing W7900 llama.cpp HIP/Vulkan Q4_K_M rows are reused unchanged. The hipEngine GGUF Q4_K_S column is compared against the existing llama.cpp Q4_K_M baselines because that is the lineage of measured baselines we have on this host; cross-quant comparisons should be read as approximate.
|
|
116
207
|
|
|
117
|
-
|
|
208
|
+
### gfx1151 (AMD Ryzen AI MAX+ 395 / Radeon 8060S)
|
|
118
209
|
|
|
119
210
|
The gfx1151 backend is a native `--offload-arch=gfx1151` peer backend using the same registry-keyed kernel surface. The Strix Halo snapshot below uses 256-row prefill chunks, which removed the 4K prefill gap without hurting long-context decode.
|
|
120
211
|
|
|
121
212
|
### Prefill tok/s
|
|
122
213
|
|
|
123
|
-
| Workload | hipEngine
|
|
214
|
+
| Workload | hipEngine PARO | llama.cpp HIP | llama.cpp Vulkan |
|
|
124
215
|
| --- | ---: | ---: | ---: |
|
|
125
216
|
| 512/128 | 983.206 | **1058.738** | 638.008 |
|
|
126
217
|
| 4K/128 | **1029.402** | 1004.220 | 595.400 |
|
|
@@ -129,7 +220,7 @@ The gfx1151 backend is a native `--offload-arch=gfx1151` peer backend using the
|
|
|
129
220
|
|
|
130
221
|
### Decode tok/s
|
|
131
222
|
|
|
132
|
-
| Workload | hipEngine
|
|
223
|
+
| Workload | hipEngine PARO | llama.cpp HIP | llama.cpp Vulkan |
|
|
133
224
|
| --- | ---: | ---: | ---: |
|
|
134
225
|
| 512/128 | **62.060** | 50.537 | 57.615 |
|
|
135
226
|
| 4K/128 | **63.605** | 49.379 | 55.027 |
|
|
@@ -141,25 +232,26 @@ On Strix Halo, `rocm-smi` / sysfs expose only a 512 MiB VRAM aperture, so cross-
|
|
|
141
232
|
See [`benchmarks/README.md`](benchmarks/README.md) for full protocol details,
|
|
142
233
|
correctness status, source-lineage targets, and external comparison baselines.
|
|
143
234
|
|
|
144
|
-
##
|
|
235
|
+
## GGUF Support
|
|
145
236
|
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
| `cuda_sm86` | NVIDIA Ampere consumer (3090-class) | Planned peer backend |
|
|
151
|
-
| `cpu_reference` | Any CPU, numpy | Correctness oracle; CI without GPU |
|
|
237
|
+
As of v0.2.0, hipEngine includes resident Qwen3.6 GGUF support for `Q4_K_M` and
|
|
238
|
+
`Q4_K_S` model files (with more formats planned). This is a major runtime path,
|
|
239
|
+
not just a loader shim: GGUF has its own quant readers, bulk-prefill path,
|
|
240
|
+
decode-repacked T16 layouts, and fast-path controls.
|
|
152
241
|
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
242
|
+
Current caveats:
|
|
243
|
+
|
|
244
|
+
- PARO models take ~24s to load on the W7900 test host; GGUF currently takes
|
|
245
|
+
about 60s because decode-repack happens on load. On-disk caching could reduce
|
|
246
|
+
startup time later, but would require additional storage for repacked layouts.
|
|
247
|
+
- GGUF has higher resident memory than packed PARO. In the current W7900 README
|
|
248
|
+
sweep, the max-context Q4_K_S session peaks at ~25.1 GiB tracked, so 128K is
|
|
249
|
+
W7900/48 GiB territory; on 24 GiB cards, expect roughly 64K context with
|
|
250
|
+
Q4_K_S.
|
|
251
|
+
- GGUF is close enough to PARO to share some high-level scheduling ideas, but in
|
|
252
|
+
practice it needs substantial GGUF-only kernels and dispatch. The goal for
|
|
253
|
+
future releases is to keep closing the remaining PARO/GGUF speed gap.
|
|
159
254
|
|
|
160
|
-
Wave32 is the default for `hip_gfx1100` device code; wave64 is treated as an
|
|
161
|
-
isolated experiment with its own gates (see
|
|
162
|
-
[`docs/PLAN.md`](docs/PLAN.md#rdna3-wavefront-and-scheduling-caveat)).
|
|
163
255
|
|
|
164
256
|
## Architecture at a glance
|
|
165
257
|
|
|
@@ -185,7 +277,7 @@ isolated experiment with its own gates (see
|
|
|
185
277
|
│ KERNELS (backend-keyed, 120 __global__ in the Qwen/PARO port) │
|
|
186
278
|
│ kernels/hip_gfx1100/ attention / linear_attn / moe / quant │
|
|
187
279
|
│ wmma / norm / rotary / fused │
|
|
188
|
-
│ kernels/hip_gfx1151/ native target-arch peer backend
|
|
280
|
+
│ kernels/hip_gfx1151/ native target-arch peer backend │
|
|
189
281
|
│ kernels/cuda_sm86/ (future) │
|
|
190
282
|
│ kernels/cpu_reference/ correctness oracle, no GPU required │
|
|
191
283
|
└─────────────────────────────────────────────────────────────────┘
|
|
@@ -262,6 +354,7 @@ current limitations.
|
|
|
262
354
|
| [`docs/BENCHMARK.md`](docs/BENCHMARK.md) | Benchmark protocols, baselines, correctness gate, artifact format |
|
|
263
355
|
| [`docs/TESTING.md`](docs/TESTING.md) | RED/GREEN workflow, correctness oracles, fixture policy |
|
|
264
356
|
| [`docs/KERNELS.md`](docs/KERNELS.md) | Kernel catalog, source-lineage drift workflow, JIT cache gotchas, build profiles |
|
|
357
|
+
| [`docs/ENVS.md`](docs/ENVS.md) | Environment variables, TheRock setup, benchmark/profiling profiles |
|
|
265
358
|
| [`docs/ROOFLINE.md`](docs/ROOFLINE.md) | RDNA3 / W7900 performance model and decision tree |
|
|
266
359
|
| [`docs/IMPLEMENTATION.md`](docs/IMPLEMENTATION.md) | Implementation status and concrete milestones |
|
|
267
360
|
| [`docs/API.md`](docs/API.md) | OpenAI-compatible server usage and endpoint support |
|