Llama 3.1 405B VRAM: what it needs at Q4_K_M, Q8_0 and FP16
Runs in your browser — nothing you paste leaves this page. How we prove that
LLM VRAM calculator playground
KV cache is counted at FP16. Overhead is a fixed 1.5 GiB estimate for the runtime and compute buffers. A custom size borrows the architecture of the nearest preset.
Results update as you type — press Enter to run now.
GPU tiers
- 8 GiBtoo smallRTX 4060, RTX 3070, RX 7600
- 12 GiBtoo smallRTX 4070, RTX 3060 12GB, RTX 5070
- 16 GiBtoo smallRTX 4060 Ti 16GB, RTX 4080, RTX 5080
- 24 GiBtoo smallRTX 4090, RTX 3090, RX 7900 XTX
- 32 GiBtoo smallRTX 5090, V100 32GB
- 48 GiBtoo smallRTX A6000, RTX 6000 Ada, L40S
- 80 GiBtoo smallA100 80GB, H100 80GB
At Q4_K_M with an 8,192-token context, Llama 3.1 405B needs about 236 GiB of VRAM: 231 GiB of weights, 3.94 GiB of KV cache and 1.50 GiB of runtime overhead. The smallest common GPU tier that holds it is beyond any single 80 GiB card, so it needs several GPUs or a large unified-memory machine.
How much VRAM Llama 3.1 405B needs
Llama 3.1 405B is the largest dense open-weight model Meta released in July 2024; it is a datacentre model, and running it locally is a multi-GPU or very-large-Mac project.
With 8,192 tokens of context the estimate is 760 GiB at FP16, 406 GiB at Q8_0 and 236 GiB at Q4_K_M. FP16 needs more than a single 80 GiB card, Q8_0 needs more than a single 80 GiB card, and Q4_K_M needs more than a single 80 GiB card. The weights are 405 billion parameters times the bits per weight of each format (16, 8.5 and 4.89), divided by eight; everything is in GiB, the unit GPU memory is sold in.
Context length and the KV cache
The KV cache stores a key and a value vector for every token in every layer. Llama 3.1 405B has 126 layers, 8 KV heads and a head dimension of 128, read from its Hugging Face config.json, so at FP16 each token costs 504 KiB. That is 3.94 GiB at 8,192 tokens, and 63.0 GiB at the maximum of 131,072, where the Q4_K_M total becomes 295 GiB and needs more than a single 80 GiB card. The cache does not shrink with weight quantization; only a runtime option that quantizes the cache itself does.
Which quantization to pick
Even Q4_K_M needs 236 GiB, so every quantization here exceeds a single 80 GiB card. Q8_0 at 406 GiB is a cluster job, and FP16 at 760 GiB is what Meta's own FP8 release was made to avoid. Q4_K_M across four 80 GiB GPUs, or Q3_K_M at 194 GiB, are the practical routes.
Pitfalls
The KV cache is the small part here: 3.94 GiB at 8,192 tokens and 63.0 GiB at 131,072, thanks to only 8 KV heads. Splitting across GPUs adds per-device buffers this estimate does not include, and across machines the network becomes the bottleneck, not memory.
For most local work Llama 3.1 70B or Qwen2.5 72B get much of the quality at a fraction of the memory.
FAQ
Questions, answered.
Tap a question to expand the answer.
How many GPUs does Llama 3.1 405B need?
At Q4_K_M the estimate is 236 GiB, so three 80 GiB cards on paper and four in practice once per-GPU buffers are counted; at Q8_0 it is 406 GiB, six or more. FP16 at 760 GiB needs ten 80 GiB cards before any per-GPU overhead.
Can Llama 3.1 405B run on a Mac?
Only on the largest unified-memory configurations. Q4_K_M needs 236 GiB, and macOS lets the GPU use only part of unified memory by default (the limit can be raised with the `iogpu.wired_limit_mb` sysctl), so a 256 GB machine is too small and a 512 GB one is the realistic minimum.
More free, private DevOps tools.
The LLM VRAM Calculator is one tool in OpsCanopy — a growing canopy of browser-based validators, converters and testers that never touch a server.
Related tools
New to AI & local LLMs? Read the AI & local LLMs guide →
42 free tools, every one offline-capable — opscanopy.com works with no signup and nothing uploaded.
Related: the full LLM VRAM Calculator, the LLM Token Counter, or every AI tool.
Estimates only; the KV cache is sized at FP16 and the overhead is a fixed allowance, so always confirm against your own runtime’s report before buying hardware. OpsCanopy is free and open.