Llama 2 Models — Hardware Requirements
10 Llama 2 models from Meta and the community, from the smallest that runs in 0.3 GB of VRAM up to 69.0B parameters. Every row links to full quantization tables and GPU compatibility.
All Llama 2 Models by Size
| Model | Params | Runs from | Context | Publisher | Quant downloads |
|---|---|---|---|---|---|
| Llama2 0B Unit Test | 770940 | 0.3 GB | 1K | ||
| Llama 2 7B Chat HF | 6.7B | 3.1 GB | — | ||
| Llama 2 7B HF | 6.7B | 3.1 GB | — | ||
| HarmBench Llama 2 13B Cls | 13.0B | 5.6 GB | 2K | ||
| Llama 2 13B Chat HF | 13.0B | 6.1 GB | — | ||
| Llama 2 13B HF | 13.0B | 6.1 GB | — | ||
| Llama 2 70B HF | 69.0B | 151.8 GB | — | ||
| Llama 2 70B Chat HF | 69.0B | 151.8 GB | — |
Frequently Asked Questions
- How much VRAM do I need to run a Llama 2 model?
- The smallest Llama 2 model, Llama2 0B Unit Test, runs from 0.3 GB of VRAM at an aggressive quantization. Larger family members need proportionally more — see the table above for every model.
- Which Llama 2 models can I run on a 16 GB GPU?
- 6 of 8 Llama 2 models fit in 16 GB of VRAM at some quantization, including Llama 2 7B Chat HF, HarmBench Llama 2 13B Cls, Llama 2 7B HF.
- What is the most popular Llama 2 model to run locally?
- Llama 2 7B Chat HF is the most downloaded Llama 2 model in local-friendly quantized formats. It runs from 3.1 GB of VRAM.