Llama 2 Models — Hardware Requirements

10 Llama 2 models from Meta and the community, from the smallest that runs in 0.3 GB of VRAM up to 69.0B parameters. Every row links to full quantization tables and GPU compatibility.

All Llama 2 Models by Size

ModelParamsContext
Llama2 0B Unit Test7709401K
Llama 2 7B Chat HF6.7B—
Llama 2 7B HF6.7B—
HarmBench Llama 2 13B Cls13.0B2K
Llama 2 13B Chat HF13.0B—
Llama 2 13B HF13.0B—
Llama 2 70B HF69.0B—
Llama 2 70B Chat HF69.0B—

Frequently Asked Questions

How much VRAM do I need to run a Llama 2 model?
The smallest Llama 2 model, Llama2 0B Unit Test, runs from 0.3 GB of VRAM at an aggressive quantization. Larger family members need proportionally more — see the table above for every model.
Which Llama 2 models can I run on a 16 GB GPU?
6 of 8 Llama 2 models fit in 16 GB of VRAM at some quantization, including Llama 2 7B Chat HF, HarmBench Llama 2 13B Cls, Llama 2 7B HF.
What is the most popular Llama 2 model to run locally?
Llama 2 7B Chat HF is the most downloaded Llama 2 model in local-friendly quantized formats. It runs from 3.1 GB of VRAM.