Latest benchmark: Radeon AI PRO R9700
One Radeon. Regular vLLM. More possibilities.
Qwen3.8 delivers 400.7 aggregate output tok/s on one 32 GB R9700 with ROCm 10 and vLLM 0.29. The 65K-preset run covered 290 requests including warmups, with thinking and prefix caching off. This is a separate workload, not a matched speedup over the earlier release.
Benchmark + demo


