VRAM Calculator
The VRAM Calculator answers one question before you deploy anything: will this model, with these settings, fit on this GPU? It breaks the requirement into model weights, KV cache, activation memory and overhead, and tells you what is left over.
LLM Launcher → VRAM Calculator, or the Vram Calc tab inside a device, where it is pre-filled with that device's real GPUs.
Configuration
The left panel holds every input. Results update as you change them.
Model
Search Hugging Face text-generation models. Type at least a few characters — the field says "Type to search models" until you do, and "Searching…" while it works. Selecting a model automatically fetches its parameters (hidden size, layers, attention heads, intermediate size), so the calculation uses the model's real architecture rather than an estimate.
You can paste either the model id (meta-llama/Llama-2-7b-hf) or its full URL.
GGUF Quantization
When the selected repository ships GGUF files, an extra field appears with a GGUF badge:
Select the GGUF quantization file you want to run. Each entry shows its approximate
bits-per-weight (~4.5 bits/weight), and lower-bit quantizations use less VRAM but may reduce
quality. If a file's bit width cannot be determined, it is listed as bits/weight unknown.
If the repository has no GGUF files, the field says "No GGUF files found" and the ordinary quantization selector is used instead.
GPU
Pick a GPU from the catalogue, or use the real hardware:
- Select Device GPUs reads the GPUs off one of your connected devices. Those entries are marked with a Device badge in the list.
- GPU Count — how many of that GPU to use. More GPUs allow larger models through model parallelism.
Quantization
The precision format for the model weights. Lower precision (for example INT4) uses less VRAM but may reduce accuracy; FP16/BF16 is full precision.
Sequence Length
The maximum context length in tokens. Longer sequences require more VRAM for KV cache storage — this is usually the setting that turns a comfortable fit into a tight one.
Batch Size
How many requests are processed simultaneously. Higher batch sizes increase throughput but require more VRAM for the KV cache.
GPU Memory Utilization
The percentage of GPU VRAM made available to the model. Reserve 10–20% for system overhead and CUDA operations.
Calculation Type
| Option | What it does |
|---|---|
| 20% Overhead | Simple estimate: weights plus a 20% buffer. Fast, deliberately conservative. |
| With Activation | More accurate: includes the activation memory actually used during inference, computed from sequence length, batch size, hidden size and intermediate size. |
Use With Activation whenever the margin matters.
Results
Summary tiles
Across the top: Model, Quantization, Batch Size, Required VRAM and Available VRAM.
VRAM Breakdown
A doughnut chart of where the memory goes.
Memory Details
| Component | Meaning |
|---|---|
| Model Weights | The memory occupied by the model's parameters at the chosen quantization. |
| KV Cache | The key/value cache, driven by sequence length and batch size. |
| Overhead | The flat buffer, in 20% Overhead mode. |
| Activations | Activation memory during inference, in With Activation mode. |
| Free VRAM | What is left. |
Usage
A bar showing what fraction of the available VRAM the configuration needs, and the verdict:
- VRAM is Sufficient — it fits.
- VRAM is Insufficient — it does not. Lower the quantization, shorten the sequence length, reduce the batch size, or add GPUs.
Leave 20–30% free VRAM in production. A configuration that fits at exactly 100% will fail as soon as a request arrives with a longer prompt than you tested with.
How this relates to the launch wizard
The wizard has its own, lighter check: the GPU fit advisor in the Model pane, which estimates from weight size only and says how many GPUs a model needs. Use it for a quick verdict while choosing a model.
Come here when you need the full picture — the effect of context length and batch size on the KV cache, or a comparison between two quantizations on hardware you do not own yet.
See also: Launch Guide · AI Launch