Calculate exact GPU VRAM / Unified Memory required to run 1.5B to 671B open-weight LLMs across FP16, Q8_0, Q6_K, Q4_K_M, and IQ2_XXS quantizations with KV-cache overhead and estimated tokens/sec.
To estimate GPU VRAM for running a local LLM in Ollama, llama.cpp, or vLLM, multiply the parameter count (in billions) by `bytes per parameter` (`2.0` for FP16, `1.05` for Q8_0, `0.62` for Q4_K_M), then add `1.5 GB – 4.0 GB` for the KV context cache (`8k–32k` tokens) and CUDA/Metal runtime buffers.
Test fit and memory bandwidth speed (tok/s) on RTX 3060 12GB, RTX 4070 Ti 16GB, RTX 4090 24GB, RTX 5090 32GB, and Mac M4 Max 64GB/128GB.
Ready-to-Run Ollama / llama.cpp CLI Generator
Outputs the exact ollama run and llama-server command flags (--ctx-size, -ngl, --cache-type-k).
Practical Use Cases
Choosing the Right Quantization Before Downloading
Know whether a 32B Q4_K_M model with a 32K context window fits inside 24GB VRAM before downloading a 20GB file.
Planning Local AI Workstation Hardware
Compare token generation speed (memory-bandwidth bound) between NVIDIA RTX GPUs and Apple Silicon Unified Memory.
Frequently Asked Questions (FAQs)
Why is Q4_K_M the most popular quantization for local LLMs?+
Q4_K_M uses ~4.85 bits per weight on average by keeping critical attention/feed-forward tensors at 6-bit while quantizing remaining weights to 4-bit. It cuts VRAM usage by ~70% compared to FP16 with less than 1% perplexity degradation.
How is local LLM token generation speed (tokens/sec) determined?+
Single-batch autoregressive token generation is memory-bandwidth bound: every generated token requires reading all active model weights from VRAM once. Estimated tok/s ≈ (GPU Memory Bandwidth in GB/s × 0.72) / Model Weight Size in GB.
Why does increasing context length from 4K to 64K use so much extra VRAM?+
Long context windows require storing Key-Value (KV) attention states for every layer and token. Enabling Q8_0 KV-cache quantization (OLLAMA_KV_CACHE_TYPE=q8_0) cuts context memory usage in half with virtually zero quality loss.
What happens if a model slightly exceeds my GPU's VRAM?+
If layers spill over into system DDR5 RAM (partial CPU offload), token generation speed typically drops by 5x to 10x because PCIe/system RAM bandwidth is much slower than GDDR6X/GDDR7 VRAM.
How much Unified Memory does Apple Silicon reserve for macOS?+
By default, macOS allows Metal GPU workloads to allocate roughly 75% of total Unified Memory (e.g., ~48GB on a 64GB Mac), though this limit can be raised via sysctl iogpu.wired_limit_mb.