Mixtral 8x7B VRAM: what it needs at Q4_K_M, Q8_0 and FP16
Runs in your browser — nothing you paste leaves this page. How we prove that
LLM VRAM calculator playground
KV cache is counted at FP16. Overhead is a fixed 1.5 GiB estimate for the runtime and compute buffers. A custom size borrows the architecture of the nearest preset.
Results update as you type — press Enter to run now.
GPU tiers
- 8 GiBtoo smallRTX 4060, RTX 3070, RX 7600
- 12 GiBtoo smallRTX 4070, RTX 3060 12GB, RTX 5070
- 16 GiBtoo smallRTX 4060 Ti 16GB, RTX 4080, RTX 5080
- 24 GiBtoo smallRTX 4090, RTX 3090, RX 7900 XTX
- 32 GiBfitsRTX 5090, V100 32GB
- 48 GiBfitsRTX A6000, RTX 6000 Ada, L40S
- 80 GiBfitsA100 80GB, H100 80GB
At Q4_K_M with an 8,192-token context, Mixtral 8x7B needs about 29.1 GiB of VRAM: 26.6 GiB of weights, 1.00 GiB of KV cache and 1.50 GiB of runtime overhead. The smallest common GPU tier that holds it is 32 GiB (RTX 5090, V100 32GB).
How much VRAM Mixtral 8x7B needs
Mixtral 8x7B, released by Mistral AI in December 2023, is a sparse mixture-of-experts model: eight experts per layer, two active for each token, about 12.9B parameters doing work out of 46.7B in total.
With 8,192 tokens of context the estimate is 89.5 GiB at FP16, 48.7 GiB at Q8_0 and 29.1 GiB at Q4_K_M. FP16 needs more than a single 80 GiB card, Q8_0 needs an 80 GiB card, and Q4_K_M needs a 32 GiB card. The weights are 46.7 billion parameters times the bits per weight of each format (16, 8.5 and 4.89), divided by eight; everything is in GiB, the unit GPU memory is sold in.
Context length and the KV cache
The KV cache stores a key and a value vector for every token in every layer. Mixtral 8x7B has 32 layers, 8 KV heads and a head dimension of 128, read from its Hugging Face config.json, so at FP16 each token costs 128 KiB. That is 1.00 GiB at 8,192 tokens, and 4.00 GiB at the maximum of 32,768, where the Q4_K_M total becomes 32.1 GiB and needs a 48 GiB card. The cache does not shrink with weight quantization; only a runtime option that quantizes the cache itself does.
Which quantization to pick
Q4_K_M at 29.1 GiB needs a 32 GiB card; Q3_K_M at 24.2 GiB narrowly misses a 24 GiB card at 8,192 tokens. Q8_0 at 48.7 GiB needs an 80 GiB card.
Pitfalls
The trap is sizing by active parameters. Every expert has to be resident, so memory follows the 46.7B total, while speed follows the 12.9B active. That is why it runs like a 13B but needs memory like a 47B. Partial offload of experts to system RAM works better than with dense models. Context tops out at 32,768 tokens, 4.00 GiB of cache.
Mistral 7B v0.3 shares its attention shape; Qwen2.5 32B and Gemma 2 27B are dense models in the same memory class.
FAQ
Questions, answered.
Tap a question to expand the answer.
How much VRAM does Mixtral 8x7B need?
About 29.1 GiB at Q4_K_M with 8,192 tokens, 48.7 GiB at Q8_0 and 89.5 GiB at FP16. All 46.7B parameters count, not the 12.9B active per token.
Can Mixtral 8x7B run on a 24 GB GPU?
Not entirely in VRAM at Q4_K_M (29.1 GiB). Q3_K_M with 4,096 tokens comes to 23.7 GiB, which fits; otherwise offload some experts to system RAM and accept slower generation.
More free, private DevOps tools.
The LLM VRAM Calculator is one tool in OpsCanopy — a growing canopy of browser-based validators, converters and testers that never touch a server.
Related tools
New to AI & local LLMs? Read the AI & local LLMs guide →
42 free tools, every one offline-capable — opscanopy.com works with no signup and nothing uploaded.
Related: the full LLM VRAM Calculator, the LLM Token Counter, or every AI tool.
Estimates only; the KV cache is sized at FP16 and the overhead is a fixed allowance, so always confirm against your own runtime’s report before buying hardware. OpsCanopy is free and open.