Category
AI tools.
Size local LLM deployments — weights, KV cache and GPU memory — from real model architectures, in your browser.
AI tools
- AI fig. 40
LLM VRAM Calculator
Estimate the GPU memory a local LLM needs — weights, KV cache and runtime overhead, per quantization and context length.
Pick a model preset or type a parameter count, choose a GGUF quantization and a context length, and get the VRAM in GiB split into weights, KV cache and runtime overhead — with a verdict against 8, 12, 16, 24, 32, 48 and 80 GiB cards. Real per-layer KV-cache math from each model's config.json, not a rule of thumb.
Open tool - AI fig. 41
LLM Token Counter
Count GPT tokens offline — o200k_base and cl100k_base, with the exact token boundaries.
Paste a prompt or a document and count its tokens with the real OpenAI BPE encodings — o200k_base (GPT-5, GPT-4.1, GPT-4o and the o-series) or cl100k_base (GPT-4, GPT-3.5 and text-embedding-3) — with every token boundary highlighted, character and word counts, and an optional cost estimate at your own price per million tokens.
Open tool
When you reach for AI tools.
Running a model locally starts with one question — will it fit — and the usual answer is a rule of thumb that is wrong by a factor of three. Parameters times bytes per weight gets you the weights; it says nothing about the KV cache, which at a 128k context on an 8B model is twice the size of the Q4 weights it sits beside.
The cache is the part people underestimate because it depends on things the model card does not advertise: layer count, key-value heads, head dimension. A model with grouped-query attention needs a quarter of the cache of one without, at the same parameter count. Quantization labels add their own confusion — Q4_K_M is not four bits per weight, it is closer to five.
This computes the estimate from the model's real architecture, read from its published configuration, and shows the three parts separately — weights, KV cache, runtime overhead — against the memory tiers GPUs are actually sold in. It runs in your browser, so the numbers come from arithmetic you can check, not from a server you have to trust.