GLM Models — Hardware Requirements

3 GLM models from Z.ai and the community, from the smallest that runs in 0.3 GB of VRAM up to 6.2B parameters. Every row links to full quantization tables and GPU compatibility.

All GLM Models by Size

ModelParamsContext
Xglm 564M564M2K
Chatglm2 6B6B33K
Chatglm3 6B6.2B8K
GLM Z1 32B 041432.6B33K

Frequently Asked Questions

How much VRAM do I need to run a GLM model?
The smallest GLM model, Xglm 564M, runs from 0.3 GB of VRAM at an aggressive quantization. Larger family members need proportionally more — see the table above for every model.
Which GLM models can I run on a 16 GB GPU?
4 of 4 GLM models fit in 16 GB of VRAM at some quantization, including Chatglm3 6B, Chatglm2 6B, Xglm 564M.
What is the most popular GLM model to run locally?
Chatglm3 6B is the most downloaded GLM model in local-friendly quantized formats. It runs from 2.9 GB of VRAM.
GLM Models — VRAM & Hardware Requirements | llmrun