A clear step-by-step guide to building llama.cpp from source with ROCm on Ubuntu — installing the dev packages, the exact CMake flags for gfx1201, and a runtime environment that just works.
How I squeezed the new Qwen 3.8 27B completely onto a single 16 GB GPU — and what it takes to get almost 42 tok/s with a 64K-token KV cache. Two quantizations, four llama.cpp configurations, full commands and results.
A follow-up guide to running Qwen 3.8 27B — the state-of-the-art open model that fits a consumer GPU — on Windows with ROCm, built from source for maximum tokens per second.