# llama.cpp with Vulkan on a Radeon iGPU

How-to · Local AI · 2 min read. By [David Wilson](https://testbenchlab.com/authors/david-wilson/). Updated 2026-10-02.

Vulkan is a useful llama.cpp route for Radeon integrated graphics. ROCm is another option on supported combinations. Choose a working backend, then compare speed with the same model and settings.

The command expects the downloaded model file named `model.gguf`. This documentation-based procedure is not presented as a hardware validation run.

## Vulkan or ROCm?

Support differs by operating system and build. Vulkan provides a broad graphics-compute route; ROCm requires a supported driver and hardware combination. A working Linux setup does not guarantee the same installation route on Windows.

## Build or download llama.cpp with Vulkan

Install Git, Visual Studio C++ build tools, CMake and the Vulkan SDK. In a developer PowerShell, build the official repository:

```powershell
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
cmake -B build -DGGML_VULKAN=ON
cmake --build build --config Release
```

An official Windows Vulkan release package is an alternative; keep its DLLs beside the executables.

## Run your first model

Use a small publisher-supplied GGUF and save it as `model.gguf` in the project folder:

```powershell
.\build\bin\Release\llama-cli.exe -m .\model.gguf -ngl 99 -c 2048 -n 128 -p "Explain unified memory in three sentences."
```

Check startup output for Vulkan detection and layer offload. Reduce offload or model size if allocation fails. Record the model file and build when comparing speed.

## Frequently asked questions

### Is Vulkan or ROCm faster on Strix Halo?

There is no universal winner. Match model, quantisation, context, driver and power settings on supported builds.

## Related reading

- [How to install Ollama on Windows 11 (Ryzen AI mini PC)](https://testbenchlab.com/guides/install-ollama-windows-11-ryzen-ai-mini-pc/)
- [How much RAM do you need to run a local LLM?](https://testbenchlab.com/guides/how-much-ram-for-local-llm/)

## Sources

1. [llama.cpp build documentation](https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md)
2. [Ollama hardware support](https://docs.ollama.com/gpu)
