OLMo Models — Hardware Requirements

20 OLMo models from Allen AI and the community, from the smallest that runs in 0.3 GB of VRAM up to 33.3B parameters. Every row links to full quantization tables and GPU compatibility.

All OLMo Models by Size

ModelParamsContext
OLMo 1B HF1.2B2K
OLMo 2 0425 1B Instruct1.5B4K
OLMo 2 0425 1B1.5B4K
Molmo2 4B4.9B37K
OLMoE 1B 7B 0924 Instruct6.9B4K
OLMoE 1B 7B 09246.9B4K
OLMoE 1B 7B 0125 Instruct6.9B4K
Olmo Hybrid 7B7B66K
Olmo 3 7B Instruct7.3B66K
Olmo 3 7B Think7.3B66K
Olmo 3 1025 7B7.3B66K
Olmo 3 7B Instruct SFT7.3B66K
OLMo 2 1124 7B Instruct7.3B4K
OlmOCR 2 7B 10258.3B128K
Molmo2 8B8.7B37K
Olmo 3.1 32B Instruct32.2B66K
Olmo 3 32B Think32.2B66K
Olmo 3 32B Think SFT32.2B66K
Olmo 3 1125 32B32.2B66K
Olmo 3.1 32B Think32.2B66K
OLMo 2 0325 32B Instruct32.2B4K
FlexOlmo 7x7B 1T33.3B4K

Frequently Asked Questions

How much VRAM do I need to run a OLMo model?
The smallest OLMo model, OLMo 2 0425 1B Instruct, runs from 1.0 GB of VRAM at an aggressive quantization. Larger family members need proportionally more — see the table above for every model.
Which OLMo models can I run on a 16 GB GPU?
19 of 22 OLMo models fit in 16 GB of VRAM at some quantization, including OlmOCR 2 7B 1025, OLMoE 1B 7B 0924 Instruct, Olmo 3.1 32B Instruct.
What is the most popular OLMo model to run locally?
OlmOCR 2 7B 1025 is the most downloaded OLMo model in local-friendly quantized formats. It runs from 2.7 GB of VRAM.