OLMo Models — Hardware Requirements
20 OLMo models from Allen AI and the community, from the smallest that runs in 0.3 GB of VRAM up to 33.3B parameters. Every row links to full quantization tables and GPU compatibility.
All OLMo Models by Size
| Model | Params | Runs from | Context | Publisher | Quant downloads |
|---|---|---|---|---|---|
| OLMo 1B HF | 1.2B | 1.1 GB | 2K | ||
| OLMo 2 0425 1B Instruct | 1.5B | 1.0 GB | 4K | ||
| OLMo 2 0425 1B | 1.5B | 1.2 GB | 4K | ||
| Molmo2 4B | 4.9B | 2.5 GB | 37K | ||
| OLMoE 1B 7B 0924 Instruct | 6.9B | 3.5 GB | 4K | ||
| OLMoE 1B 7B 0924 | 6.9B | 3.5 GB | 4K | ||
| OLMoE 1B 7B 0125 Instruct | 6.9B | 2.5 GB | 4K | ||
| Olmo Hybrid 7B | 7B | 15.3 GB | 66K | ||
| Olmo 3 7B Instruct | 7.3B | 3.4 GB | 66K | ||
| Olmo 3 7B Think | 7.3B | 3.4 GB | 66K | ||
| Olmo 3 1025 7B | 7.3B | 3.4 GB | 66K | ||
| Olmo 3 7B Instruct SFT | 7.3B | 4.5 GB | 66K | ||
| OLMo 2 1124 7B Instruct | 7.3B | 4.5 GB | 4K | ||
| OlmOCR 2 7B 1025 | 8.3B | 2.7 GB | 128K | ||
| Molmo2 8B | 8.7B | 4.3 GB | 37K | ||
| Olmo 3.1 32B Instruct | 32.2B | 9.7 GB | 66K | ||
| Olmo 3 32B Think | 32.2B | 9.7 GB | 66K | ||
| Olmo 3 32B Think SFT | 32.2B | 14.5 GB | 66K | ||
| Olmo 3 1125 32B | 32.2B | 65.3 GB | 66K | ||
| Olmo 3.1 32B Think | 32.2B | 65.3 GB | 66K | ||
| OLMo 2 0325 32B Instruct | 32.2B | 14.5 GB | 4K | ||
| FlexOlmo 7x7B 1T | 33.3B | 67.9 GB | 4K |
Frequently Asked Questions
- How much VRAM do I need to run a OLMo model?
- The smallest OLMo model, OLMo 2 0425 1B Instruct, runs from 1.0 GB of VRAM at an aggressive quantization. Larger family members need proportionally more — see the table above for every model.
- Which OLMo models can I run on a 16 GB GPU?
- 19 of 22 OLMo models fit in 16 GB of VRAM at some quantization, including OlmOCR 2 7B 1025, OLMoE 1B 7B 0924 Instruct, Olmo 3.1 32B Instruct.
- What is the most popular OLMo model to run locally?
- OlmOCR 2 7B 1025 is the most downloaded OLMo model in local-friendly quantized formats. It runs from 2.7 GB of VRAM.