LLM VRAM Calculator: Size the GPU before you download the model.
Runs in your browser — nothing you paste leaves this page. How we prove that
LLM VRAM Calculator playground
KV cache is counted at FP16. Overhead is a fixed 1.5 GiB estimate for the runtime and compute buffers. A custom size borrows the architecture of the nearest preset.
Results update as you type — press Enter to run now.
GPU tiers
- 8 GiBfitsRTX 4060, RTX 3070, RX 7600
- 12 GiBfitsRTX 4070, RTX 3060 12GB, RTX 5070
- 16 GiBfitsRTX 4060 Ti 16GB, RTX 4080, RTX 5080
- 24 GiBfitsRTX 4090, RTX 3090, RX 7900 XTX
- 32 GiBfitsRTX 5090, V100 32GB
- 48 GiBfitsRTX A6000, RTX 6000 Ada, L40S
- 80 GiBfitsA100 80GB, H100 80GB
Pick a model preset like Llama 3.1 8B or type a parameter count, choose a GGUF quantization such as Q4_K_M and a context length, and get the VRAM a local LLM needs — weights, KV cache and runtime overhead in GiB, with a verdict against common GPU sizes — all computed in your browser, no signup required.
The Gap
The weights fit. Then the context window does not.
GPU memory for a local LLM is three bills, not one. The weights are the part everyone sizes: parameters times bits per weight, which is why a 4-bit quantization of an 8B model is a 4.5 GiB file. The KV cache is the part everyone forgets — two tensors per layer that grow linearly with the context length, so the same 8B model wants 1 GiB of cache at 8k tokens and 16 GiB at 128k. The runtime overhead is the rest: compute buffers, CUDA or Metal context and the framework itself, a fixed cost of around 1.5 GiB.
Two unit systems make the edge cases worse. A GPU sold as “24 GB” exposes 24 GiB (2³⁰ bytes each), while model cards and file sizes often quote GB (10⁹ bytes) — a 7.4 % gap. This calculator works in GiB throughout, so a fit is judged against what the driver actually reports, and the reference below shows both figures for the worked example. On Apple silicon the unified memory is shared with the system; plan on about 75 % of the Mac’s RAM being available to the model.
Asking an AI how much VRAM a model needs is the risky shortcut: a language model recalls plausible numbers from forum posts about different quantizations and context lengths. This calculator computes the real per-layer KV cache from each model’s config.json shape and the measured bits per weight from llama.cpp — numbers you can verify an AI’s answer against, not more plausible text.
Converting between GiB and GB? The Data Size Converter shows the exact byte counts.
The Pipeline
How it works.
Four deterministic steps run on every change — all inside your browser tab, with the model shapes checked against each config.json.
-
Pick the shape.
A preset supplies the model’s layer count, KV heads and head dimension from its config.json; a custom parameter count borrows the nearest preset’s shape and is labelled as estimated.
-
Weigh the weights.
Parameters × effective bits per weight for the chosen quantization, ÷ 8, ÷ 2³⁰. The bits per weight come from llama.cpp’s measured table, not the nominal 4 or 8.
-
Size the KV cache.
Two tensors per layer (K and V) × KV heads × head dim × context tokens × 2 bytes at FP16 — the part of the bill that grows with your context window.
-
Add overhead, then verdict.
A fixed 1.5 GiB for the runtime and compute buffers, then the total is placed against the common card sizes and the smallest one that fits is named.
Formula Reference
The two formulas, worked.
Everything the panel shows comes from these two lines plus a fixed overhead. Here they are with real model shapes, so you can check the result by hand.
KV cache grows with context
Grouped-query attention keeps the cache small: Llama 3.1 has 32 attention heads but only 8 KV heads per layer, which is the number that counts here.
KV cache (FP16) = 2 × layers × kv_heads × head_dim × context × 2 bytes
Llama 3.1 8B at 128k context
layers 32 · kv_heads 8 · head_dim 128 · context 131,072
= 2 × 32 × 8 × 128 × 131,072 × 2 B
= 17,179,869,184 B
= 16.0 GiB (17.2 GB — the KV cache alone outgrows a 16 GiB card)
Same model at 8k context
= 1.0 GiB Weights by quantization
A “4-bit” GGUF is not 4 bits per weight: Q4_K_M averages 4.89 once its scales and the higher-precision tensors are counted, and Q8_0 averages 8.5.
Weights = params × bits_per_weight ÷ 8 ÷ 2^30
Llama 3.1 70B, Q4_K_M (4.89 bpw measured)
weights 70e9 × 4.89 ÷ 8 ÷ 2^30 = 39.85 GiB
KV @ 8k 2 × 80 × 8 × 128 × 8,192 × 2 B = 2.50 GiB
overhead = 1.50 GiB
total = 43.85 GiB → fits a 48 GiB card, not 24 or 32
Llama 3.1 8B, Q4_K_M
weights 8e9 × 4.89 ÷ 8 ÷ 2^30 = 4.55 GiB
KV @ 8k = 1.00 GiB
total (with overhead) = 7.05 GiB → fits an 8 GiB card Next Step
Serving the model on a cluster? Size the pod too.
Once the card is chosen, the Kubernetes Resource Calculator turns your replica count and per-pod CPU, memory and GPU requests into node totals and headroom — handy before you write the Deployment that mounts the model.
Llama 3.1 8B Q4_K_M 8k → 7.05 GiB fits 8 GiB
Llama 3.1 8B FP16 128k → 32.4 GiB fits 48 GiB
Llama 3.1 70B Q4_K_M 8k → 43.8 GiB fits 48 GiB Common models
What the open-weight models people run most often need in VRAM, one page each.
- Llama 3.2 3B VRAM
- Llama 3.1 8B VRAM
- Llama 3.1 70B VRAM
- Llama 3.1 405B VRAM
- Qwen2.5 7B VRAM
- Qwen2.5 14B VRAM
- Qwen2.5 32B VRAM
- Qwen2.5 72B VRAM
- Mistral 7B v0.3 VRAM
- Mistral Nemo 12B VRAM
- Gemma 2 9B VRAM
- Gemma 2 27B VRAM
- Mixtral 8x7B VRAM
- DeepSeek-R1-Distill-Qwen 7B VRAM
- DeepSeek-R1-Distill-Qwen 14B VRAM
- DeepSeek-R1-Distill-Qwen 32B VRAM
- DeepSeek-R1-Distill-Llama 8B VRAM
- DeepSeek-R1-Distill-Llama 70B VRAM
FAQ
Questions, answered.
Tap a question to expand the answer.
How much VRAM does a model need?
Add three parts. Weights are the parameter count multiplied by the bits per weight of your quantization, divided by 8 to get bytes and by 2^30 to get GiB — an 8B model is about 15 GiB at FP16, 8 GiB at Q8_0 and 4.6 GiB at Q4_K_M. The KV cache grows with context length (next question), and the runtime adds a fixed buffer, which this calculator sets at 1.5 GiB. The total has to fit the card's memory with a little to spare; if it does not, pick a smaller quantization or a shorter context before reaching for a bigger GPU.
What does quantization do, and how do Q4_K_M, Q8_0 and FP16 differ?
Quantization stores each weight in fewer bits than the 16 it was trained in. FP16 keeps 16 bits per weight (bpw); Q8_0 uses about 8.5 bpw and is practically lossless; Q4_K_M uses about 4.89 bpw — roughly a third of the FP16 size — with a small quality cost most people cannot notice in chat. The fractional values come from llama.cpp's mixed layouts, which keep a few sensitive tensors at higher precision and store a scale factor per block, which is why a "4-bit" file is not exactly half of an 8-bit one.
How does context length affect VRAM?
Every token in the context keeps a key and a value vector in every layer, so the KV cache is 2 × layers × KV heads × head dimension × tokens × 2 bytes at FP16. For Llama 3.1 8B (32 layers, 8 KV heads, head dimension 128) that is 128 KiB per token: 1 GiB at 8,192 tokens and 16 GiB at the full 131,072 — more than the Q4_K_M weights themselves. Grouped-query attention (8 KV heads instead of 32) cuts the cache fourfold, which is why Llama 3 fits long contexts where Llama 2 did not. The context slider draws the cache as its own bar segment so you can see where the memory goes.
Why does the calculator report GiB instead of GB?
GPU memory is sold in binary gigabytes: a "24 GB" RTX 4090 holds 24 GiB, which is 25.77 GB. A GB is 10^9 bytes and a GiB is 2^30 = 1,073,741,824 bytes, 7.37% larger, so dividing by 10^9 would make every model look 7.4% bigger than it is and mis-call fits right at the edge. The calculator divides by 2^30 everywhere, so a verdict compares like with like.
What is the 1.5 GiB overhead?
Weights and KV cache are not the whole picture: the runtime (llama.cpp, Ollama, vLLM) allocates compute buffers for activations, the CUDA or Metal context takes a few hundred MiB, and the driver reserves some memory of its own. 1.5 GiB is a round estimate that covers these for most models on a single card; the real value varies by runtime, batch size and GPU. Treat it as a margin, not a measurement — if a verdict lands within half a GiB of a tier, expect trouble.
Can a 70B model run on a 24 GB GPU?
Not fully on the GPU. Llama 3.1 70B at Q4_K_M needs about 40 GiB for the weights alone, before the KV cache and runtime overhead, so a single 24 GiB card cannot hold it. It fits on a 48 GiB card (RTX 6000 Ada, A6000) or across two 24 GiB cards with layer splitting; Q2_K at about 3.16 bpw brings the weights down to roughly 26 GiB, still too much for one card and with a visible quality cost. Runtimes can offload the remaining layers to system RAM, but token speed then drops to what the CPU's memory bandwidth allows, typically a few tokens per second.
How does Apple unified memory compare?
Apple Silicon shares one pool of memory between CPU and GPU, so a Mac with 64 GB can load models no 24 GiB card can. By default macOS lets the GPU use roughly 75% of unified memory (about 48 GiB on a 64 GB machine) and keeps the rest for the system; the iogpu.wired_limit_mb sysctl can raise the limit at your own risk. Enter about three quarters of the Mac's memory as the budget, and remember that bandwidth, not capacity, sets token speed: an M-series Max moves roughly 400 GB/s, an RTX 4090 about 1 TB/s.
Does my selection ever leave my browser?
No. The calculator runs 100% client-side: the architecture presets and the arithmetic ship with the page, and every number is computed in your browser tab. Nothing is uploaded to a server, there is no account or signup, and the share link encodes your selection in the URL fragment, which browsers never send to a server.
How much VRAM do I need for a 7B, 14B or 70B model?
For the weights alone: an 8B model needs about 4.6 GiB in Q4_K_M, 7.9 GiB in Q8_0 and 14.9 GiB in FP16 (a 7B about 4.0, 6.9 and 13.0 GiB); a 14B needs about 8.0, 13.9 and 26.1 GiB; a 70B about 39.9, 69.3 and 130.4 GiB. Add the KV cache for your context — at 8,192 tokens that is 1 GiB for Llama 3.1 8B and 2.5 GiB for Llama 3.1 70B — plus the 1.5 GiB runtime overhead, so an 8B in Q4_K_M at 8k context lands near 7.1 GiB in total. Each popular model has its own page with its real layer count and KV heads filled in, so you can read the verdict for that model directly.
Why is the KV cache sometimes bigger than the weights?
The weights are a fixed size, but the KV cache grows with every token: it costs 2 × layers × KV heads × head dimension × 2 bytes per token in FP16. For Llama 3.1 8B that is 2 × 32 × 8 × 128 × 2 = 128 KiB per token, so the full 131,072-token context needs 16 GiB of cache against about 4.6 GiB of Q4_K_M weights. Grouped-query attention is what keeps it that small — 8 KV heads instead of 32 cuts the cache to a quarter — while older multi-head-attention models with one KV head per query head run out of memory at far shorter contexts.
More free, private DevOps tools.
The LLM VRAM Calculator is one tool in OpsCanopy — a growing canopy of browser-based validators, converters and testers that never touch a server.
Related tools
New to AI & local LLMs? Read the AI & local LLMs guide →
42 free tools, every one offline-capable — opscanopy.com works with no signup and nothing uploaded.
More sizing: the Kubernetes Resource Calculator and the Data Size Converter, or browse the full tools directory.
Estimates only; the KV cache is sized at FP16 and the overhead is a fixed allowance, so always confirm against your own runtime’s report before buying hardware. OpsCanopy is free and open.