Explainer2 min readLocal AI

How much RAM do you need to run a local LLM?

In short

Budget local-model memory for weights, context cache and runtime, not just the download. 32GB is a useful smaller-model starting point; larger dense models can require 64GB or 128GB depending on quantisation and context.

Dense modelApproximate Q4 weightsSystem-memory planning range
7B4–5GB16–32GB
14B8–10GB24–32GB
32B18–22GB32–64GB
70B40–48GB64–128GB

Editorial estimates for typical four-bit formats and moderate context. Architecture, cache precision and runtime change the requirement.

How much RAM for 7B, 14B, 32B and 70B models?

Use the table as a planning range, not a guaranteed minimum. Four-bit storage is roughly half a byte per parameter before overhead. Longer context, multimodal inputs and concurrent sessions can exceed these suggested ranges.

What quantisation does to memory

Quantisation reduces weight precision and file size, with possible quality and speed trade-offs. A theoretical four-bit 70B payload is about 35GB before overhead. Actual formats also store scales, metadata and higher-precision tensors; runtime and context require more capacity.

Which mini PC tier fits which model?

We prefer Geekom A9 Max for smaller-model desktop work and A9 Mega for larger shared-memory workloads. Verify accelerator-visible capacity. Fitting weights in system RAM is different from keeping the intended workload on the GPU at useful speed.

Frequently asked questions

Can 32GB of RAM run a 70B model?

A conventional four-bit dense 70B workload exceeds that capacity after weights and working memory are included. More aggressive formats change the trade-offs.

Is 64GB enough for local AI?

It is enough for many quantised workloads, but not every model or context. Leave room for system work and check GPU-visible memory.

Written by

David Wilson writes and edits Testbench Lab’s coverage of AI voice recorders, mini PCs and local AI. His focus is the practical difference between a feature on a product page and a tool you would want to use regularly.

Sources

  1. [1]llama.cpp build documentation
  2. [2]Ollama hardware support
  3. [3]Geekom A9 Max specifications
  4. [4]Geekom A9 Mega specifications