Skip to main content

VRAM Calculator

The VRAM Calculator answers one question before you deploy anything: will this model, with these settings, fit on this GPU? It breaks the requirement into model weights, KV cache, activation memory and overhead, and tells you what is left over.

LLM Launcher → VRAM Calculator, or the Vram Calc tab inside a device, where it is pre-filled with that device's real GPUs.

Qwen3-8B on RTX 4090: raising the context to 16K breaks the fit, a second GPU restores it

 


Configuration​

The left panel holds every input. Results update as you change them.

Model​

Search Hugging Face text-generation models. Type at least a few characters — the field says "Type to search models" until you do, and "Searching…" while it works. Selecting a model automatically fetches its parameters (hidden size, layers, attention heads, intermediate size), so the calculation uses the model's real architecture rather than an estimate.

You can paste either the model id (meta-llama/Llama-2-7b-hf) or its full URL.

GGUF Quantization​

When the selected repository ships GGUF files, an extra field appears with a GGUF badge: Select the GGUF quantization file you want to run. Each entry shows its approximate bits-per-weight (~4.5 bits/weight), and lower-bit quantizations use less VRAM but may reduce quality. If a file's bit width cannot be determined, it is listed as bits/weight unknown.

If the repository has no GGUF files, the field says "No GGUF files found" and the ordinary quantization selector is used instead.

GPU​

Pick a GPU from the catalogue, or use the real hardware:

  • Select Device GPUs reads the GPUs off one of your connected devices. Those entries are marked with a Device badge in the list.
  • GPU Count — how many of that GPU to use. More GPUs allow larger models through model parallelism.

Quantization​

The precision format for the model weights. Lower precision (for example INT4) uses less VRAM but may reduce accuracy; FP16/BF16 is full precision.

Sequence Length​

The maximum context length in tokens. Longer sequences require more VRAM for KV cache storage — this is usually the setting that turns a comfortable fit into a tight one.

Batch Size​

How many requests are processed simultaneously. Higher batch sizes increase throughput but require more VRAM for the KV cache.

GPU Memory Utilization​

The percentage of GPU VRAM made available to the model. Reserve 10–20% for system overhead and CUDA operations.

Calculation Type​

OptionWhat it does
20% OverheadSimple estimate: weights plus a 20% buffer. Fast, deliberately conservative.
With ActivationMore accurate: includes the activation memory actually used during inference, computed from sequence length, batch size, hidden size and intermediate size.

Use With Activation whenever the margin matters.


Results​

Summary tiles​

Across the top: Model, Quantization, Batch Size, Required VRAM and Available VRAM.

VRAM Breakdown​

A doughnut chart of where the memory goes.

Memory Details​

ComponentMeaning
Model WeightsThe memory occupied by the model's parameters at the chosen quantization.
KV CacheThe key/value cache, driven by sequence length and batch size.
OverheadThe flat buffer, in 20% Overhead mode.
ActivationsActivation memory during inference, in With Activation mode.
Free VRAMWhat is left.

Usage​

A bar showing what fraction of the available VRAM the configuration needs, and the verdict:

  • VRAM is Sufficient — it fits.
  • VRAM is Insufficient — it does not. Lower the quantization, shorten the sequence length, reduce the batch size, or add GPUs.
tip

Leave 20–30% free VRAM in production. A configuration that fits at exactly 100% will fail as soon as a request arrives with a longer prompt than you tested with.


How this relates to the launch wizard​

The wizard has its own, lighter check: the GPU fit advisor in the Model pane, which estimates from weight size only and says how many GPUs a model needs. Use it for a quick verdict while choosing a model.

Come here when you need the full picture — the effect of context length and batch size on the KV cache, or a comparison between two quantizations on hardware you do not own yet.

See also: Launch Guide · AI Launch


Typical uses​

  • Before buying hardware — try the model against GPUs you do not have.
  • Before a production deploy — confirm the context length and batch size you intend to serve actually fit.
  • Choosing a quantization — compare BF16, FP8 and INT4 for the same model on the same card.
  • Planning multi-GPU — raise GPU Count until the model fits, and see how much headroom is left.