In short
Vulkan is a useful llama.cpp route for Radeon integrated graphics. ROCm is another option on supported combinations. Choose a working backend, then compare speed with the same model and settings.
The command expects the downloaded model file named model.gguf. This documentation-based procedure is not presented as a hardware validation run.
Vulkan or ROCm?
Support differs by operating system and build. Vulkan provides a broad graphics-compute route; ROCm requires a supported driver and hardware combination. A working Linux setup does not guarantee the same installation route on Windows.
Build or download llama.cpp with Vulkan
Install Git, Visual Studio C++ build tools, CMake and the Vulkan SDK. In a developer PowerShell, build the official repository:
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
cmake -B build -DGGML_VULKAN=ON
cmake --build build --config ReleaseAn official Windows Vulkan release package is an alternative; keep its DLLs beside the executables.
Run your first model
Use a small publisher-supplied GGUF and save it as model.gguf in the project folder:
.\build\bin\Release\llama-cli.exe -m .\model.gguf -ngl 99 -c 2048 -n 128 -p "Explain unified memory in three sentences."Check startup output for Vulkan detection and layer offload. Reduce offload or model size if allocation fails. Record the model file and build when comparing speed.
Frequently asked questions
Is Vulkan or ROCm faster on Strix Halo?
There is no universal winner. Match model, quantisation, context, driver and power settings on supported builds.