All LLM Models
Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Fable Traces
AliesTaha · 4.0B · runs from 2.2 GB
Fable Traces is a 4.0B-parameter open language model from AliesTaha. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Meditron3 70B
EPFLiGHT · 70.6B · runs from 155.2 GB
Meditron3 70B is a 70.6B-parameter open language model from EPFLiGHT. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 2 70B Chat HF
Meta · 69.0B · runs from 151.8 GB
Llama 2 70B Chat HF is a 69.0B-parameter open language model from Meta in the Llama 2 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
OFFELLIA Gemma 4 E4B 8B Claude 4.6 Opus Reasoning MTP
Brunobkr · 4B · runs from 1.9 GB
OFFELLIA Gemma 4 E4B 8B Claude 4.6 Opus Reasoning MTP is a 4B-parameter open language model from Brunobkr in the Gemma 4 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Edge0 8B A1B Preview
Edge0 · 7.9B · runs from 16.4 GB
Edge0 8B A1B Preview is a 7.9B-parameter open language model from Edge0. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Sarashina2.2 3B Instruct v0.1
sbintuitions · 3.4B · runs from 2.1 GB
Sarashina2.2 3B Instruct v0.1 is a 3.4B-parameter open language model from sbintuitions. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 3 1 Nemotron Ultra 253B V1
NVIDIA · 253.4B · runs from 557.5 GB
Llama 3 1 Nemotron Ultra 253B V1 is a 253.4B-parameter open language model from NVIDIA in the Llama 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Muse Glimmer 30B Heretic
darkc0de · 29.8B · runs from 13.1 GB
Muse Glimmer 30B Heretic is a 29.8B-parameter open language model from darkc0de in the Muse Glimmer family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
TinyLlama 1.1B Chat V0.6
TinyLlama · 1.1B · runs from 0.8 GB
TinyLlama 1.1B Chat V0.6 is a 1.1B-parameter open language model from TinyLlama in the TinyLlama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Apertus 70B Instruct 2509
swiss-ai · 70B · runs from 30.7 GB
Apertus 70B Instruct 2509 is a 70B-parameter open language model from swiss-ai in the Apertus family. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gollem V4 250M Pl
SlayerLab · 250M · runs from 0.6 GB
Gollem V4 250M Pl is a 250M-parameter open language model from SlayerLab. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Neural Chat 7B v3 3
Intel · 7.2B · runs from 3.6 GB
Neural Chat 7B v3 3 is a 7.2B-parameter open language model from Intel. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Baguettotron
PleIAs · 321M · runs from 0.6 GB
Baguettotron is a 321M-parameter open language model from PleIAs. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
DeepSeek Coder v2 Lite Base
DeepSeek · 15.7B · runs from 7.4 GB
DeepSeek Coder v2 Lite Base is a 15.7B-parameter open language model from DeepSeek in the DeepSeek Coder family. It supports a context window of up to 163,840 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Pollux 4B Judge
ai-forever · 4.0B · runs from 2.2 GB
Pollux 4B Judge is a 4.0B-parameter open language model from ai-forever. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 3n E2B IT Litert Lm
Google · 2B · runs from 0.9 GB
Gemma 3n E2B IT Litert Lm is a 2B-parameter open language model from Google in the Gemma 3 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Shieldgemma 2B
Google · 2.6B · runs from 1.2 GB
Shieldgemma 2B is a 2.6B-parameter open language model from Google in the Gemma 2 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Tiny Aya Global
Cohere · 3.3B · runs from 7.4 GB
Tiny Aya Global is a 3.3B-parameter open language model from Cohere in the Aya family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
DeepSeek V4 Pro DSpark
DeepSeek · 1650.5B · runs from 701.8 GB
DeepSeek-V4-Pro-DSpark is not a distinct model but the DeepSeek-V4-Pro checkpoint — a 1.6-trillion-parameter mixture-of-experts model — packaged with an additional DSpark speculative-decoding module bolted on for faster inference; DeepSeek's own card lists 49 billion parameters active per token for the underlying DeepSeek-V4-Pro checkpoint. The V4 series introduces a hybrid attention design combining Compressed Sparse Attention and Heavily Compressed Attention for long-context efficiency, Manifold-Constrained Hyper-Connections for training stability, and the Muon optimizer, and was pretrained on more than 32 trillion tokens. It sits alongside a smaller 284B-total/13B-active V4-Flash sibling that shares the same architecture. Given its enormous total size, it requires a multi-GPU server even heavily quantized. Context length is 1,048,576 tokens (1M). It is released under the MIT license, permitting unrestricted commercial and research use, and was published in June 2026.
MiniCPM5 2B SFT
OpenBMB · 2.5B · runs from 1.5 GB
MiniCPM5 2B SFT is a 2.5B-parameter open language model from OpenBMB in the MiniCPM family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GigaChat 20B A3B Base
ai-sage · 20B · runs from 9.0 GB
GigaChat 20B A3B Base is a 20B-parameter open language model from ai-sage. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama Krikri 8B Instruct
ilsp · 8.2B · runs from 4.0 GB
Llama Krikri 8B Instruct is a 8.2B-parameter open language model from ilsp in the Llama family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
DeepSeek V3.2 Speciale
DeepSeek · 685.4B · runs from 295.2 GB
DeepSeek-V3.2-Speciale is DeepSeek's high-compute reasoning variant of DeepSeek-V3.2, a mixture-of-experts model with roughly 41 billion active parameters out of about 685 billion total, fine-tuned from DeepSeek-V3.2-Exp-Base. It uses DeepSeek Sparse Attention (DSA), an efficient attention mechanism aimed at long-context scenarios, together with heavy reinforcement-learning post-training; the card reports gold-medal-level performance at the 2025 International Mathematical Olympiad and International Olympiad in Informatics, and claims it surpasses GPT-5 with reasoning on par with Gemini 3.0 Pro. Unlike the standard DeepSeek-V3.2 checkpoint, Speciale is dedicated purely to deep reasoning and does not support tool-calling. At this scale it requires a multi-GPU server cluster to run even quantized. Context length is 163,840 tokens. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in November 2025.
Deeplm 108M
samcheng0 · 108M · runs from 0.2 GB
Deeplm 108M is a 108M-parameter open language model from samcheng0. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GLM 5.3 DFlash2
incoai · 2.5B · runs from 1.4 GB
GLM 5.3 DFlash2 is a 2.5B-parameter open language model from incoai in the GLM 5 family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 4B Heretic
DreamFast · 4.0B · runs from 2.2 GB
Qwen3 4B Heretic is a 4.0B-parameter open language model from DreamFast in the Qwen 3 family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Vaultgemma 1B
Google · 1.0B · runs from 2.3 GB
Vaultgemma 1B is a 1.0B-parameter open language model from Google in the Gemma family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 12B IT Abliterated Uncensored
OpenYourMind · 12.0B · runs from 6.1 GB
Gemma 4 12B IT Abliterated Uncensored is a 12.0B-parameter open language model from OpenYourMind in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
MiMo V2.6 Pro RL
Xiaomi · 1020B · runs from 434.0 GB
MiMo V2.6 Pro RL is Xiaomi's largest open-weight language model, a sparse mixture-of-experts design with roughly 1.02 trillion total parameters and about 42 billion active per token. It is tuned for chat and function-calling, with heavy reinforcement-learning post-training behind its name. Only the active parameters compute per token, keeping decoding efficient, but the full weight set must fit in memory, putting it in server-class territory best reached through a hosted endpoint. The model supports a 1,048,576 token context window, suited to very long documents or agent sessions. It is released under the MIT license, one of the most permissive open licenses, with no restrictions on commercial use. Published in September 2026, its post-training relied on large-scale, fully asynchronous reinforcement learning (GRPO).
Mistral 7B v0.2
mistral-community · 7.2B · runs from 3.6 GB
Mistral 7B v0.2 is a 7.2B-parameter open language model from mistral-community in the Mistral family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.