All LLM Models

Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Fable Traces

AliesTaha · 4.0B · runs from 2.2 GB

5.3K 208

Fable Traces is a 4.0B-parameter open language model from AliesTaha. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Meditron3 70B

EPFLiGHT · 70.6B · runs from 155.2 GB

5.3K 26

Meditron3 70B is a 70.6B-parameter open language model from EPFLiGHT. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 2 70B Chat HF

Meta · 69.0B · runs from 151.8 GB

5.2K 2.2K

Llama 2 70B Chat HF is a 69.0B-parameter open language model from Meta in the Llama 2 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

OFFELLIA Gemma 4 E4B 8B Claude 4.6 Opus Reasoning MTP

Brunobkr · 4B · runs from 1.9 GB

5.1K 2

OFFELLIA Gemma 4 E4B 8B Claude 4.6 Opus Reasoning MTP is a 4B-parameter open language model from Brunobkr in the Gemma 4 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Edge0 8B A1B Preview

Edge0 · 7.9B · runs from 16.4 GB

5.1K 79

Edge0 8B A1B Preview is a 7.9B-parameter open language model from Edge0. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Sarashina2.2 3B Instruct v0.1

sbintuitions · 3.4B · runs from 2.1 GB

5.0K 38

Sarashina2.2 3B Instruct v0.1 is a 3.4B-parameter open language model from sbintuitions. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3 1 Nemotron Ultra 253B V1

NVIDIA · 253.4B · runs from 557.5 GB

5.0K 352

Llama 3 1 Nemotron Ultra 253B V1 is a 253.4B-parameter open language model from NVIDIA in the Llama 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Muse Glimmer 30B Heretic

darkc0de · 29.8B · runs from 13.1 GB

5.0K 15

Muse Glimmer 30B Heretic is a 29.8B-parameter open language model from darkc0de in the Muse Glimmer family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

TinyLlama 1.1B Chat V0.6

TinyLlama · 1.1B · runs from 0.8 GB

4.9K 113

TinyLlama 1.1B Chat V0.6 is a 1.1B-parameter open language model from TinyLlama in the TinyLlama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Apertus 70B Instruct 2509

swiss-ai · 70B · runs from 30.7 GB

4.9K 182

Apertus 70B Instruct 2509 is a 70B-parameter open language model from swiss-ai in the Apertus family. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gollem V4 250M Pl

SlayerLab · 250M · runs from 0.6 GB

4.9K 2

Gollem V4 250M Pl is a 250M-parameter open language model from SlayerLab. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Neural Chat 7B v3 3

Intel · 7.2B · runs from 3.6 GB

4.8K 83

Neural Chat 7B v3 3 is a 7.2B-parameter open language model from Intel. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatMath

Baguettotron

PleIAs · 321M · runs from 0.6 GB

4.8K 240

Baguettotron is a 321M-parameter open language model from PleIAs. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek Coder v2 Lite Base

DeepSeek · 15.7B · runs from 7.4 GB

4.8K 105

DeepSeek Coder v2 Lite Base is a 15.7B-parameter open language model from DeepSeek in the DeepSeek Coder family. It supports a context window of up to 163,840 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Pollux 4B Judge

ai-forever · 4.0B · runs from 2.2 GB

4.7K 4

Pollux 4B Judge is a 4.0B-parameter open language model from ai-forever. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 3n E2B IT Litert Lm

Google · 2B · runs from 0.9 GB

4.7K 556

Gemma 3n E2B IT Litert Lm is a 2B-parameter open language model from Google in the Gemma 3 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Shieldgemma 2B

Google · 2.6B · runs from 1.2 GB

4.6K 122

Shieldgemma 2B is a 2.6B-parameter open language model from Google in the Gemma 2 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Tiny Aya Global

Cohere · 3.3B · runs from 7.4 GB

4.6K 169

Tiny Aya Global is a 3.3B-parameter open language model from Cohere in the Aya family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek V4 Pro DSpark

DeepSeek · 1650.5B · runs from 701.8 GB

4.5K 541

DeepSeek-V4-Pro-DSpark is not a distinct model but the DeepSeek-V4-Pro checkpoint — a 1.6-trillion-parameter mixture-of-experts model — packaged with an additional DSpark speculative-decoding module bolted on for faster inference; DeepSeek's own card lists 49 billion parameters active per token for the underlying DeepSeek-V4-Pro checkpoint. The V4 series introduces a hybrid attention design combining Compressed Sparse Attention and Heavily Compressed Attention for long-context efficiency, Manifold-Constrained Hyper-Connections for training stability, and the Muon optimizer, and was pretrained on more than 32 trillion tokens. It sits alongside a smaller 284B-total/13B-active V4-Flash sibling that shares the same architecture. Given its enormous total size, it requires a multi-GPU server even heavily quantized. Context length is 1,048,576 tokens (1M). It is released under the MIT license, permitting unrestricted commercial and research use, and was published in June 2026.

Chat

MiniCPM5 2B SFT

OpenBMB · 2.5B · runs from 1.5 GB

4.5K 27

MiniCPM5 2B SFT is a 2.5B-parameter open language model from OpenBMB in the MiniCPM family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GigaChat 20B A3B Base

ai-sage · 20B · runs from 9.0 GB

4.4K 16

GigaChat 20B A3B Base is a 20B-parameter open language model from ai-sage. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama Krikri 8B Instruct

ilsp · 8.2B · runs from 4.0 GB

4.3K 32

Llama Krikri 8B Instruct is a 8.2B-parameter open language model from ilsp in the Llama family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek V3.2 Speciale

DeepSeek · 685.4B · runs from 295.2 GB

4.3K 726

DeepSeek-V3.2-Speciale is DeepSeek's high-compute reasoning variant of DeepSeek-V3.2, a mixture-of-experts model with roughly 41 billion active parameters out of about 685 billion total, fine-tuned from DeepSeek-V3.2-Exp-Base. It uses DeepSeek Sparse Attention (DSA), an efficient attention mechanism aimed at long-context scenarios, together with heavy reinforcement-learning post-training; the card reports gold-medal-level performance at the 2025 International Mathematical Olympiad and International Olympiad in Informatics, and claims it surpasses GPT-5 with reasoning on par with Gemini 3.0 Pro. Unlike the standard DeepSeek-V3.2 checkpoint, Speciale is dedicated purely to deep reasoning and does not support tool-calling. At this scale it requires a multi-GPU server cluster to run even quantized. Context length is 163,840 tokens. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in November 2025.

Chat

Deeplm 108M

samcheng0 · 108M · runs from 0.2 GB

4.3K 5

Deeplm 108M is a 108M-parameter open language model from samcheng0. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GLM 5.3 DFlash2

incoai · 2.5B · runs from 1.4 GB

4.3K 15

GLM 5.3 DFlash2 is a 2.5B-parameter open language model from incoai in the GLM 5 family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen3 4B Heretic

DreamFast · 4.0B · runs from 2.2 GB

4.2K 38

Qwen3 4B Heretic is a 4.0B-parameter open language model from DreamFast in the Qwen 3 family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Vaultgemma 1B

Google · 1.0B · runs from 2.3 GB

4.2K 240

Vaultgemma 1B is a 1.0B-parameter open language model from Google in the Gemma family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 4 12B IT Abliterated Uncensored

OpenYourMind · 12.0B · runs from 6.1 GB

4.1K 48

Gemma 4 12B IT Abliterated Uncensored is a 12.0B-parameter open language model from OpenYourMind in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

MiMo V2.6 Pro RL

Xiaomi · 1020B · runs from 434.0 GB

4.1K 450

MiMo V2.6 Pro RL is Xiaomi's largest open-weight language model, a sparse mixture-of-experts design with roughly 1.02 trillion total parameters and about 42 billion active per token. It is tuned for chat and function-calling, with heavy reinforcement-learning post-training behind its name. Only the active parameters compute per token, keeping decoding efficient, but the full weight set must fit in memory, putting it in server-class territory best reached through a hosted endpoint. The model supports a 1,048,576 token context window, suited to very long documents or agent sessions. It is released under the MIT license, one of the most permissive open licenses, with no restrictions on commercial use. Published in September 2026, its post-training relied on large-scale, fully asynchronous reinforcement learning (GRPO).

ChatFunctions

Mistral 7B v0.2

mistral-community · 7.2B · runs from 3.6 GB

4.1K 229

Mistral 7B v0.2 is a 7.2B-parameter open language model from mistral-community in the Mistral family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat