Skip to content

LLM VRAM Calculator: Size the GPU before you download the model.

Runs in your browser — nothing you paste leaves this page. How we prove that

LLM VRAM Calculator playground

Model
Quantization

KV cache is counted at FP16. Overhead is a fixed 1.5 GiB estimate for the runtime and compute buffers. A custom size borrows the architecture of the nearest preset.

Results update as you type — press Enter to run now.

fig. 40 — llm-vram-calculator · ai 7.05 GiB total at Q4_K_M, 8k context — fits an 8 GiB GPU.
Result
Total VRAM7.05 GiBfits an 8 GiB card
Weights4.55 GiB8 B × 4.89 bpw (Q4_K_M)
KV cache1.00 GiB8k ctx · 32 layers · 8 KV heads · 128 dim · FP16
Overhead1.50 GiBruntime + compute buffers, estimate
ArchitectureLlama 3.1 8B

GPU tiers

  • 8 GiBfitsRTX 4060, RTX 3070, RX 7600
  • 12 GiBfitsRTX 4070, RTX 3060 12GB, RTX 5070
  • 16 GiBfitsRTX 4060 Ti 16GB, RTX 4080, RTX 5080
  • 24 GiBfitsRTX 4090, RTX 3090, RX 7900 XTX
  • 32 GiBfitsRTX 5090, V100 32GB
  • 48 GiBfitsRTX A6000, RTX 6000 Ada, L40S
  • 80 GiBfitsA100 80GB, H100 80GB

Pick a model preset like Llama 3.1 8B or type a parameter count, choose a GGUF quantization such as Q4_K_M and a context length, and get the VRAM a local LLM needs — weights, KV cache and runtime overhead in GiB, with a verdict against common GPU sizes — all computed in your browser, no signup required.

The Gap

The weights fit. Then the context window does not.

GPU memory for a local LLM is three bills, not one. The weights are the part everyone sizes: parameters times bits per weight, which is why a 4-bit quantization of an 8B model is a 4.5 GiB file. The KV cache is the part everyone forgets — two tensors per layer that grow linearly with the context length, so the same 8B model wants 1 GiB of cache at 8k tokens and 16 GiB at 128k. The runtime overhead is the rest: compute buffers, CUDA or Metal context and the framework itself, a fixed cost of around 1.5 GiB.

Two unit systems make the edge cases worse. A GPU sold as “24 GB” exposes 24 GiB (2³⁰ bytes each), while model cards and file sizes often quote GB (10⁹ bytes) — a 7.4 % gap. This calculator works in GiB throughout, so a fit is judged against what the driver actually reports, and the reference below shows both figures for the worked example. On Apple silicon the unified memory is shared with the system; plan on about 75 % of the Mac’s RAM being available to the model.

Asking an AI how much VRAM a model needs is the risky shortcut: a language model recalls plausible numbers from forum posts about different quantizations and context lengths. This calculator computes the real per-layer KV cache from each model’s config.json shape and the measured bits per weight from llama.cpp — numbers you can verify an AI’s answer against, not more plausible text.

Converting between GiB and GB? The Data Size Converter shows the exact byte counts.

The Pipeline

How it works.

Four deterministic steps run on every change — all inside your browser tab, with the model shapes checked against each config.json.

  1. Pick the shape.

    A preset supplies the model’s layer count, KV heads and head dimension from its config.json; a custom parameter count borrows the nearest preset’s shape and is labelled as estimated.

  2. Weigh the weights.

    Parameters × effective bits per weight for the chosen quantization, ÷ 8, ÷ 2³⁰. The bits per weight come from llama.cpp’s measured table, not the nominal 4 or 8.

  3. Size the KV cache.

    Two tensors per layer (K and V) × KV heads × head dim × context tokens × 2 bytes at FP16 — the part of the bill that grows with your context window.

  4. Add overhead, then verdict.

    A fixed 1.5 GiB for the runtime and compute buffers, then the total is placed against the common card sizes and the smallest one that fits is named.

Formula Reference

The two formulas, worked.

Everything the panel shows comes from these two lines plus a fixed overhead. Here they are with real model shapes, so you can check the result by hand.

KV cache grows with context

Grouped-query attention keeps the cache small: Llama 3.1 has 32 attention heads but only 8 KV heads per layer, which is the number that counts here.

fig. 40.1 — kv-cache-llama-3-1-8b
KV cache (FP16) = 2 × layers × kv_heads × head_dim × context × 2 bytes

Llama 3.1 8B at 128k context
  layers 32 · kv_heads 8 · head_dim 128 · context 131,072
  = 2 × 32 × 8 × 128 × 131,072 × 2 B
  = 17,179,869,184 B
  = 16.0 GiB            (17.2 GB — the KV cache alone outgrows a 16 GiB card)

Same model at 8k context
  = 1.0 GiB

Weights by quantization

A “4-bit” GGUF is not 4 bits per weight: Q4_K_M averages 4.89 once its scales and the higher-precision tensors are counted, and Q8_0 averages 8.5.

fig. 40.2 — weights-llama-3-1-70b
Weights = params × bits_per_weight ÷ 8 ÷ 2^30

Llama 3.1 70B, Q4_K_M (4.89 bpw measured)
  weights   70e9 × 4.89 ÷ 8 ÷ 2^30  = 39.85 GiB
  KV @ 8k   2 × 80 × 8 × 128 × 8,192 × 2 B  =  2.50 GiB
  overhead                                 =  1.50 GiB
  total                                    = 43.85 GiB   → fits a 48 GiB card, not 24 or 32

Llama 3.1 8B, Q4_K_M
  weights   8e9 × 4.89 ÷ 8 ÷ 2^30   = 4.55 GiB
  KV @ 8k                            = 1.00 GiB
  total (with overhead)              = 7.05 GiB   → fits an 8 GiB card

Next Step

Serving the model on a cluster? Size the pod too.

Once the card is chosen, the Kubernetes Resource Calculator turns your replica count and per-pod CPU, memory and GPU requests into node totals and headroom — handy before you write the Deployment that mounts the model.

fit.txt
Llama 3.1 8B   Q4_K_M    8k   →  7.05 GiB   fits  8 GiB
Llama 3.1 8B   FP16    128k   →  32.4 GiB   fits 48 GiB
Llama 3.1 70B  Q4_K_M    8k   →  43.8 GiB   fits 48 GiB

FAQ

Questions, answered.

Tap a question to expand the answer.

Add three parts. Weights are the parameter count multiplied by the bits per weight of your quantization, divided by 8 to get bytes and by 2^30 to get GiB — an 8B model is about 15 GiB at FP16, 8 GiB at Q8_0 and 4.6 GiB at Q4_K_M. The KV cache grows with context length (next question), and the runtime adds a fixed buffer, which this calculator sets at 1.5 GiB. The total has to fit the card's memory with a little to spare; if it does not, pick a smaller quantization or a shorter context before reaching for a bigger GPU.

Quantization stores each weight in fewer bits than the 16 it was trained in. FP16 keeps 16 bits per weight (bpw); Q8_0 uses about 8.5 bpw and is practically lossless; Q4_K_M uses about 4.89 bpw — roughly a third of the FP16 size — with a small quality cost most people cannot notice in chat. The fractional values come from llama.cpp's mixed layouts, which keep a few sensitive tensors at higher precision and store a scale factor per block, which is why a "4-bit" file is not exactly half of an 8-bit one.

Every token in the context keeps a key and a value vector in every layer, so the KV cache is 2 × layers × KV heads × head dimension × tokens × 2 bytes at FP16. For Llama 3.1 8B (32 layers, 8 KV heads, head dimension 128) that is 128 KiB per token: 1 GiB at 8,192 tokens and 16 GiB at the full 131,072 — more than the Q4_K_M weights themselves. Grouped-query attention (8 KV heads instead of 32) cuts the cache fourfold, which is why Llama 3 fits long contexts where Llama 2 did not. The context slider draws the cache as its own bar segment so you can see where the memory goes.

GPU memory is sold in binary gigabytes: a "24 GB" RTX 4090 holds 24 GiB, which is 25.77 GB. A GB is 10^9 bytes and a GiB is 2^30 = 1,073,741,824 bytes, 7.37% larger, so dividing by 10^9 would make every model look 7.4% bigger than it is and mis-call fits right at the edge. The calculator divides by 2^30 everywhere, so a verdict compares like with like.

Weights and KV cache are not the whole picture: the runtime (llama.cpp, Ollama, vLLM) allocates compute buffers for activations, the CUDA or Metal context takes a few hundred MiB, and the driver reserves some memory of its own. 1.5 GiB is a round estimate that covers these for most models on a single card; the real value varies by runtime, batch size and GPU. Treat it as a margin, not a measurement — if a verdict lands within half a GiB of a tier, expect trouble.

Not fully on the GPU. Llama 3.1 70B at Q4_K_M needs about 40 GiB for the weights alone, before the KV cache and runtime overhead, so a single 24 GiB card cannot hold it. It fits on a 48 GiB card (RTX 6000 Ada, A6000) or across two 24 GiB cards with layer splitting; Q2_K at about 3.16 bpw brings the weights down to roughly 26 GiB, still too much for one card and with a visible quality cost. Runtimes can offload the remaining layers to system RAM, but token speed then drops to what the CPU's memory bandwidth allows, typically a few tokens per second.

Apple Silicon shares one pool of memory between CPU and GPU, so a Mac with 64 GB can load models no 24 GiB card can. By default macOS lets the GPU use roughly 75% of unified memory (about 48 GiB on a 64 GB machine) and keeps the rest for the system; the iogpu.wired_limit_mb sysctl can raise the limit at your own risk. Enter about three quarters of the Mac's memory as the budget, and remember that bandwidth, not capacity, sets token speed: an M-series Max moves roughly 400 GB/s, an RTX 4090 about 1 TB/s.

No. The calculator runs 100% client-side: the architecture presets and the arithmetic ship with the page, and every number is computed in your browser tab. Nothing is uploaded to a server, there is no account or signup, and the share link encodes your selection in the URL fragment, which browsers never send to a server.

For the weights alone: an 8B model needs about 4.6 GiB in Q4_K_M, 7.9 GiB in Q8_0 and 14.9 GiB in FP16 (a 7B about 4.0, 6.9 and 13.0 GiB); a 14B needs about 8.0, 13.9 and 26.1 GiB; a 70B about 39.9, 69.3 and 130.4 GiB. Add the KV cache for your context — at 8,192 tokens that is 1 GiB for Llama 3.1 8B and 2.5 GiB for Llama 3.1 70B — plus the 1.5 GiB runtime overhead, so an 8B in Q4_K_M at 8k context lands near 7.1 GiB in total. Each popular model has its own page with its real layer count and KV heads filled in, so you can read the verdict for that model directly.

The weights are a fixed size, but the KV cache grows with every token: it costs 2 × layers × KV heads × head dimension × 2 bytes per token in FP16. For Llama 3.1 8B that is 2 × 32 × 8 × 128 × 2 = 128 KiB per token, so the full 131,072-token context needs 16 GiB of cache against about 4.6 GiB of Q4_K_M weights. Grouped-query attention is what keeps it that small — 8 KV heads instead of 32 cuts the cache to a quarter — while older multi-head-attention models with one KV head per query head run out of memory at far shorter contexts.

More free, private DevOps tools.

The LLM VRAM Calculator is one tool in OpsCanopy — a growing canopy of browser-based validators, converters and testers that never touch a server.

New to AI & local LLMs?  Read the AI & local LLMs guide →

42 free tools, every one offline-capable — opscanopy.com works with no signup and nothing uploaded.

More sizing: the Kubernetes Resource Calculator and the Data Size Converter, or browse the full tools directory.

Estimates only; the KV cache is sized at FP16 and the overhead is a fixed allowance, so always confirm against your own runtime’s report before buying hardware. OpsCanopy is free and open.