Code Llama Models — Hardware Requirements

4 Code Llama models from Meta and the community, from the smallest that runs in 4.2 GB of VRAM up to 69.0B parameters. Every row links to full quantization tables and GPU compatibility.

All Code Llama Models by Size

ModelParamsContext
CodeLlama 7B Instruct HF6.7B16K
CodeLlama 7B HF6.7B16K
CodeLlama 34B Instruct HF33.7B16K
CodeLlama 34B HF33.7B16K
CodeLlama 70B Instruct HF69.0B4K

Frequently Asked Questions

How much VRAM do I need to run a Code Llama model?
The smallest Code Llama model, CodeLlama 7B Instruct HF, runs from 4.2 GB of VRAM at an aggressive quantization. Larger family members need proportionally more — see the table above for every model.
Which Code Llama models can I run on a 16 GB GPU?
4 of 5 Code Llama models fit in 16 GB of VRAM at some quantization, including CodeLlama 34B Instruct HF, CodeLlama 7B Instruct HF, CodeLlama 7B HF.
What is the most popular Code Llama model to run locally?
CodeLlama 34B Instruct HF is the most downloaded Code Llama model in local-friendly quantized formats. It runs from 10.0 GB of VRAM.