22:21
64d ago
→hipEngine: Fast Native Qwen 3.6 Inference for RDNA3
hipEngine released an AGPLv3 ROCm-native inference engine for Qwen3.6 on RDNA3 GPUs; on Qwen3.6 35B-A3B at 128K context with INT8 KV cache, it reports 20.89 GiB allocator peak, 1076.5 tok/s prefill, and 60.0 tok/s decode.
70
SCORE
H1·K1·R1