Granite Models — Hardware Requirements

20 Granite models from IBM and the community, from the smallest that runs in 1.0 GB of VRAM up to 32.2B parameters. Every row links to full quantization tables and GPU compatibility.

All Granite Models by Size

ModelParamsContext
Granite 3.0 1B A400m Instruct1.3B4K
Granite 4.0 1B Base1.6B131K
Granite 3.1 2B Instruct2.5B131K
Granite 3.3 2B Instruct2.5B131K
Granite 3.0 2B Instruct2.6B4K
Granite 4.0 Micro3.4B131K
Granite 4.1 3B3.4B131K
Granite 3B Code Base 2k3.5B2K
Granite 4.2 3B3.7B131K
Granite Vision 4.1 4B4.0B131K
Granite Switch 4.1 3B Preview4.1B131K
Granite 4.0 Tiny Preview6.7B131K
Granite 4.0 H Tiny6.9B131K
Granite 3.1 8B Instruct8.2B131K
Granite 3.0 8B Instruct8.2B4K
Granite 3.2 8B Instruct8.2B131K
Granite Guardian 3.2 8B Factuality Detection8.2B131K
Granite 3.3 8B Instruct8.2B131K
Granite Guardian 3.3 8B8.2B131K
Parable Granite 4.1 8B Claude Fable 58.4B131K
Granite 4.2 8B8.8B131K
Granite 4.1 8B8.8B131K
Granite Switch 4.1 8B Preview9.6B131K
Granite 4.1 30B28.9B131K
Granite 4.2 30B29.3B131K
Granite 4.0 H Small32.2B131K
Granite Switch 4.1 30B Preview32.2B131K

Frequently Asked Questions

How much VRAM do I need to run a Granite model?
The smallest Granite model, Granite 3.0 1B A400m Instruct, runs from 1.0 GB of VRAM at an aggressive quantization. Larger family members need proportionally more — see the table above for every model.
Which Granite models can I run on a 16 GB GPU?
25 of 27 Granite models fit in 16 GB of VRAM at some quantization, including Granite 4.2 8B, Granite 4.0 H Tiny, Granite 4.1 30B.
What is the most popular Granite model to run locally?
Granite 4.2 8B is the most downloaded Granite model in local-friendly quantized formats. It runs from 3.0 GB of VRAM.
Granite Models — VRAM & Hardware Requirements | llmrun