# How much RAM do you need to run a local LLM?

Explainer · Local AI · 2 min read. By [David Wilson](https://testbenchlab.com/authors/david-wilson/). Updated 2026-10-02.

Budget local-model memory for weights, context cache and runtime, not just the download. 32GB is a useful smaller-model starting point; larger dense models can require 64GB or 128GB depending on quantisation and context.

| Dense model | Approximate Q4 weights | System-memory planning range |
|---|---:|---:|
| 7B | 4–5GB | 16–32GB |
| 14B | 8–10GB | 24–32GB |
| 32B | 18–22GB | 32–64GB |
| 70B | 40–48GB | 64–128GB |

Editorial estimates for typical four-bit formats and moderate context. Architecture, cache precision and runtime change the requirement.

## How much RAM for 7B, 14B, 32B and 70B models?

Use the table as a planning range, not a guaranteed minimum. Four-bit storage is roughly half a byte per parameter before overhead. Longer context, multimodal inputs and concurrent sessions can exceed these suggested ranges.

## What quantisation does to memory

Quantisation reduces weight precision and file size, with possible quality and speed trade-offs. A theoretical four-bit 70B payload is about 35GB before overhead. Actual formats also store scales, metadata and higher-precision tensors; runtime and context require more capacity.

## Which mini PC tier fits which model?

We prefer Geekom A9 Max for smaller-model desktop work and A9 Mega for larger shared-memory workloads. Verify accelerator-visible capacity. Fitting weights in system RAM is different from keeping the intended workload on the GPU at useful speed.

## Frequently asked questions

### Can 32GB of RAM run a 70B model?

A conventional four-bit dense 70B workload exceeds that capacity after weights and working memory are included. More aggressive formats change the trade-offs.

### Is 64GB enough for local AI?

It is enough for many quantised workloads, but not every model or context. Leave room for system work and check GPU-visible memory.

## Related reading

- [What is unified memory, and why does it matter for local AI?](https://testbenchlab.com/guides/what-is-unified-memory/)
- [The best mini PCs for AI in 2026](https://testbenchlab.com/best/best-mini-pcs-for-ai/)
- [The best Strix Halo mini PCs (Ryzen AI Max+ 395)](https://testbenchlab.com/best/best-strix-halo-mini-pcs/)
- [AI mini PCs](https://testbenchlab.com/categories/ai-mini-pcs/)

## Sources

1. [llama.cpp build documentation](https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md)
2. [Ollama hardware support](https://docs.ollama.com/gpu)
3. [Geekom A9 Max specifications](https://www.geekompc.com/geekom-a9-max-mini-pc/)
4. [Geekom A9 Mega specifications](https://www.geekompc.com/geekom-a9-mega-ai-mini-pc/)
