Whether a local model runs on your hardware mostly comes down to one number: how much memory it needs once it is loaded. Pick a model size, a quantization level and a context length below, and the calculator estimates the total VRAM (or unified memory) required, then checks it against the GPUs and AI machines people actually buy in 2026.
LLM VRAM calculator
| GPU / device | Verdict |
|---|
Fits = estimate uses 90% or less of usable memory. Tight = 90-100%: it may load, but a longer chat, a bigger batch or another app on the GPU can push it over. Doesn't fit = you would need a smaller quant, shorter context, or CPU offload (much slower). Estimates, not benchmarks.
How the estimate works
The total is the sum of three parts:
- Weights. Parameter count multiplied by bytes per parameter. An 8B model at FP16 needs about 16 GB just for weights; at 4-bit it drops to under 5 GB. This is almost always the biggest slice.
- KV cache. While the model reads your prompt and writes its answer, it stores attention keys and values for every token in the context. The calculator uses a per-1,000-token figure for each model size, based on the layer counts and grouped-query attention used by current Llama, Qwen and Mistral-style models. It grows in a straight line with context, so 128K tokens costs 32 times more than 4K.
- Overhead. About 1.5 GB for the runtime, CUDA or Metal buffers and scratch space.
Each device is then compared against its usable memory, not the number on the box. A discrete GPU loses a few hundred megabytes to the driver and your desktop. Apple Silicon Macs and other unified-memory machines share RAM with the operating system, so by default only about two thirds to three quarters of it can go to the GPU. If the estimate uses 90% or less of usable memory it is marked Fits; 90-100% is Tight; anything above is Doesn’t fit.
Quantization explained
Quantization stores each weight with fewer bits. FP16 uses 16 bits per weight and is the reference quality. Q8 halves that with almost no measurable loss. Q6 and Q5 are near-lossless for chat and coding. Q4 (usually Q4_K_M in GGUF files) cuts memory to roughly a quarter of FP16 and is the most popular choice for local use, with a small quality drop that is most noticeable on small models and on tasks like maths. As a rule of thumb, a bigger model at Q4 usually beats a smaller model at Q8 that uses the same memory.
Caveats
- Mixture-of-experts models still need memory for all experts, even though only a few are active per token. Use the total parameter count.
- Some runtimes can quantize the KV cache to 8-bit or 4-bit, which roughly halves or quarters the context cost. The calculator assumes a standard 16-bit cache.
- Older models without grouped-query attention can need several times more KV cache than shown.
- Offloading layers to system RAM lets a model run when it doesn’t fit, but generation becomes much slower.
Estimates, not benchmarks. These numbers are a sizing guide built from public model architectures and typical runtime behaviour. We have not loaded every model on every device listed, and real usage varies by runtime (llama.cpp, Ollama, LM Studio, vLLM, MLX), driver version and settings. Leave some headroom.
FAQ
Can I run a 70B model on a single 24 GB GPU?
Not fully on the GPU. A 70B model at 4-bit needs around 40 GB for weights alone, so a 24 GB card has to offload layers to system RAM, which slows output sharply. A 64 GB Mac or a 128 GB unified-memory machine can hold it.
Is 8 GB of VRAM enough for local LLMs?
For 7-8B models at 4-bit with modest context, yes, though it is tight. For 13-14B models or long context, 12-16 GB is a much more comfortable floor.
Why does context length matter so much?
Every token in the context window adds to the KV cache. At 128K tokens, the cache for a 70B-class model can approach 40 GB on its own, more than many GPUs have in total.
Do AMD Radeon cards work for local AI?
Yes. Memory needs are the same, and llama.cpp, Ollama and LM Studio support Radeon cards through ROCm or Vulkan. NVIDIA’s CUDA still has the widest tool support.