Free LLM calculators that run in your browser
How much GPU memory a model needs, how fast it writes on your hardware and what fits on your GPU, read from the model files on Hugging Face. Nothing to install, no sign-up.
- LLM Speed Calculator Estimate how many tokens per second an LLM writes on your GPU or Mac, from memory bandwidth, the active parameters and the context length. Works with any model on Hugging Face.
- LLM VRAM Calculator Estimate how much GPU memory an LLM needs to run: weights, KV cache and overhead for any quantization and context length. Load any model from Hugging Face.
VRAM requirements by model
- DeepSeek V4 Flash 0731
- Qwen3.6 35B-A3B
- Qwen3.6 27B
- Qwen3.8 27B
- Qwen3.8 Flash Next
- Qwen3.8 2.4T-A95B
- Qwen3.5 9B
- Qwen3.5 122B-A10B
- Qwen3-Coder-Next
- DeepSeek V4.1 Flash
- DeepSeek V4 Flash
- DeepSeek V4 Pro
- DeepSeek V3.2
- GLM-5.3
- GLM-5.3 Flash
- GLM-5.2
- GLM-4.7 Flash
- Gemma 4 31B
- Gemma 4 26B-A4B
- Gemma 4 12B
- Gemma 4 E4B
- Kimi K3
- MiniMax M3
- MiniMax M2.7
- MiMo V2.6 Flash
- MiMo V2.6 Pro
- Mistral Medium 3.5 128B
- Nemotron 3 Nano 4B
- Nemotron 3 Nano 30B-A3B
- Nemotron 3 Super 120B-A12B
- Xing 4.0 29B-A4B
- Muse Glimmer 30B
- MiniCPM5 2B
- Llama 3.1 8B
- Llama 3.1 70B
- Qwen3 8B
- Qwen3 30B-A3B
- gpt-oss-20b
- gpt-oss-120b
- DeepSeek V3 / R1
Guides
What can your GPU run?
- RTX 4060 8GB
- RTX 3060 12GB
- Arc B580 12GB
- RTX 4070 12GB
- RTX 5070 12GB
- RTX 4060 Ti 16GB
- RTX 5060 Ti 16GB
- RX 9070 XT 16GB
- RTX 4070 Ti Super 16GB
- RTX 4080 Super 16GB
- RTX 5070 Ti 16GB
- RTX 5080 16GB
- RTX 3090
- RTX 4090
- RX 7900 XTX
- RTX 5090
- M4 Pro Mac (64 GB)
- M4 Max Mac (128 GB)
- M3 Ultra Mac Studio (512 GB)
- Ryzen AI Max+ 395 (128 GB)
- DGX Spark (128 GB)
- L40S
- RTX PRO 6000 Blackwell
- A100 80GB
- H100 SXM
- H200
- B200
- 2× RTX 3060 12GB
- 2× RTX 3090
- 4× RTX 3090
- 2× RTX 4090
- 2× RTX 5090
- 8× H100 SXM
- 8× H200