All LLM Models

Browse 856 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Featured only

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

NVIDIA Nemotron Nano 9B v2 Japanese

NVIDIA · 8.9B · runs from 4.4 GB

281.4K 124

NVIDIA Nemotron Nano 9B v2 Japanese is a specialized variant of the Nemotron Nano 9B v2, fine-tuned for Japanese language understanding and generation. At 8.9 billion parameters, it maintains the same hardware-friendly footprint as the English version while delivering natural Japanese conversational ability. For users looking to run a Japanese-language assistant locally, this model offers a rare combination of compact size and dedicated language optimization from a major hardware vendor. It handles Japanese text with the fluency you'd expect from a purpose-built model rather than a multilingual afterthought.

Chat

Cydonia 24B V4.3

TheDrummer · 23.6B · runs from 7.8 GB

6.0K 118

Cydonia 24B V4.3 is a 23.6B-parameter open language model from TheDrummer. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 4 E2B IT Qat Mobile Transformers

Google · 2.3B · runs from 1.4 GB

1.7K 28

Gemma 4 E2B IT Qat Mobile Transformers is a 2.3B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen3.6 27B DFlash

z-lab · 27B · runs from 12.6 GB

68.7K 345

Qwen3.6 27B DFlash is a 27B-parameter open language model from z-lab in the Qwen 3.6 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3.3 70B Instruct Abliterated

huihui-ai · 70.6B · runs from 20.4 GB

4.3K 74

Llama 3.3 70B Instruct Abliterated is a 70.6B-parameter open language model from huihui-ai in the Llama 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Mixtral 8x7B Instruct v0.1

Mistral AI · 46.7B · runs from 20.4 GB

806.7K 4.7K

Mixtral 8x7B Instruct v0.1 is Mistral AI's flagship Mixture-of-Experts model, combining eight expert networks of 7 billion parameters each for a 46.7B total weight count while activating only about 12.9 billion parameters per token. This sparse architecture delivers performance that rivals much larger dense models at a fraction of the inference cost, excelling across reasoning, code generation, and multilingual tasks. Because the full weights must still be loaded into memory, you will need around 24–48 GB of VRAM depending on quantization level, making it best suited for multi-GPU desktop setups or high-VRAM workstation cards. If your hardware can accommodate it, Mixtral offers one of the best performance-per-active-parameter ratios available for local deployment.

Chat

Gemma 4 12B OBLITERATED

OBLITERATUS · 12.0B · runs from 4.3 GB

43.6K 250

Gemma 4 12B OBLITERATED is a 12.0B-parameter open language model from OBLITERATUS in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 4 26B A4B IT Uncensored Heretic

llmfan46 · 25.8B · runs from 11.6 GB

2.1K 12

Gemma 4 26B A4B IT Uncensored Heretic is a 25.8B-parameter open language model from llmfan46 in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Gemma 4 26B A4B IT Assistant

Google · 26B · runs from 11.4 GB

126.5K 162

Gemma 4 26B A4B IT Assistant is a 26B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

MiniMax M2.1

MiniMaxAI · 228.7B · runs from 63.5 GB

10.0K 1.4K

MiniMax M2.1 is an earlier generation of MiniMax's large mixture-of-experts model series, featuring the same 228 billion total parameter architecture as its successor. It offers strong multilingual performance across Chinese and English tasks, including conversation, reasoning, and content generation. While M2.5 refines the formula, M2.1 remains a capable option for users with the multi-GPU hardware needed to host a model of this scale locally.

Chat

LFM2 8B A1B

LiquidAI · 8.3B · runs from 2.7 GB

46.0K 367

LFM2 8B A1B is Liquid AI's larger mixture-of-experts model, combining the company's novel hybrid architecture with approximately 8 billion total parameters. It uses a MoE design to keep active compute per token low while maintaining strong general performance across chat and reasoning tasks. For local users, it offers an intriguing alternative to conventional 8B transformers, with Liquid AI's architecture promising improved efficiency and throughput on consumer-grade hardware.

Chat

Deepseek Coder 33B Instruct

DeepSeek · 33.3B · runs from 14.6 GB

6.2K 573

Deepseek Coder 33B Instruct is a 33.3B-parameter open language model from DeepSeek in the DeepSeek Coder family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

NVIDIA Nemotron 3 Ultra 550B A55B BF16

NVIDIA · 560.5B · runs from 169.6 GB

67.2K 203

NVIDIA Nemotron 3 Ultra 550B A55B BF16 is a 560.5B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

SmolLM3 3B

Hugging Face · 3.1B · runs from 1.3 GB

516.4K 970

SmolLM3 3B is Hugging Face's latest-generation compact language model, representing a significant step up from the SmolLM2 series. At 3 billion parameters, it delivers considerably stronger reasoning, instruction following, and general language understanding while maintaining modest hardware requirements that keep it accessible on most consumer GPUs. This model benefits from improved training data, architectural refinements, and lessons learned from previous SmolLM generations. It is well positioned for local chatbot applications, coding assistance, and content generation tasks where you want strong performance without dedicating the resources required by 7B-class models.

Chat

GLM 4.5 Air

zai-org · 110.5B · runs from 30.8 GB

310.3K 608

GLM 4.5 Air is a 110.5B-parameter open language model from zai-org in the GLM 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek V3.1

DeepSeek · 684.5B · runs from 192.1 GB

217.9K 824

DeepSeek V3.1 is a 684.5B-parameter open language model from DeepSeek in the DeepSeek V3 family. It supports a context window of up to 163,840 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 3n E2B IT

Google · 5.4B · runs from 1.6 GB

355.6K 305

Gemma 3n E2B IT is a 5.4B-parameter open language model from Google in the Gemma 3 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Gemma 4 E4B IT Assistant

Google · 4B · runs from 2 GB

52.2K 108

Gemma 4 E4B IT Assistant is a 4B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3.2 11B Vision Instruct

Meta · 10.7B · runs from 5.0 GB

138.3K 1.6K

Llama 3.2 11B Vision Instruct is a 10.7B-parameter open language model from Meta in the Llama 3 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Jan Code 4B

janhq · 4.4B · runs from 2.4 GB

2.1K 68

Jan Code 4B is a 4.4B-parameter open language model from janhq. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatFunctionsCode

Qwen2.5 Coder 32B

Alibaba · 32.8B · runs from 9.8 GB

3.6K 156

Qwen2.5 Coder 32B is a 32.8B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

MiniMax M2.7 BF16 Ultra Uncensored Heretic

llmfan46 · 228.7B · runs from 97.8 GB

619 7

MiniMax M2.7 BF16 Ultra Uncensored Heretic is a 228.7B-parameter open language model from llmfan46 in the MiniMax family. It supports a context window of up to 204,800 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 3n E4B IT

Google · 7.8B · runs from 2.4 GB

16.9K 917

Gemma 3n E4B IT is a 7.8B-parameter open language model from Google in the Gemma 3 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Qwen3 235B A22B

Alibaba · 235.1B · runs from 100.4 GB

539.2K 1.1K

Qwen3 235B A22B is the largest model in Alibaba Cloud's Qwen 3 series, a Mixture of Experts (MoE) model with 235 billion total parameters and approximately 22 billion active parameters per forward pass. The MoE architecture enables it to deliver performance competitive with the best available open-weight models while requiring significantly less compute per token than a comparably sized dense model. It supports hybrid thinking mode for flexible chain-of-thought reasoning. Due to its massive total parameter count, running Qwen3 235B A22B locally requires substantial VRAM to load all expert weights, typically needing multiple high-end professional GPUs even at reduced precision. In heavily quantized formats it becomes accessible on workstation-class multi-GPU setups. Released under the Apache 2.0 license.

Chat

Hermes 4 14B

Nous Research · 14.8B · runs from 5.1 GB

37.6K 156

Hermes 4 14B is a 14.8B-parameter open language model from Nous Research in the Hermes family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoningRoleplay

GLM 4.6

zai-org · 356.8B · runs from 98.7 GB

13.8K 1.2K

GLM 4.6 is a 356.8B-parameter open language model from zai-org in the GLM 4 family. It supports a context window of up to 202,752 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

LFM2.5 350M

LiquidAI · 354M · runs from 0.5 GB

75.9K 332

LFM2.5 350M is a 354M-parameter open language model from LiquidAI. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

LFM2.5 1.2B Thinking

LiquidAI · 1.2B · runs from 0.9 GB

30.9K 361

LFM2.5 1.2B Thinking is a 1.2B-parameter open language model from LiquidAI. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen2.5 Coder 7B

Alibaba · 7.6B · runs from 3.6 GB

205.3K 139

Qwen2.5 Coder 7B is a 7.6-billion parameter code-specialized base (pretrained) model from Alibaba Cloud's Qwen 2.5 Coder series. It is trained on a large dataset of source code and natural language but is not instruction-tuned, making it suitable for fine-tuning, code-related research, and custom downstream applications. The model supports a 128K token context window and runs efficiently on consumer GPUs. It serves as the foundation for the Qwen2.5 Coder 7B Instruct variant and community fine-tunes targeting specific programming languages or workflows. Released under the Apache 2.0 license.

ChatCode

Hermes 3 Llama 3.1 8B

Nous Research · 8.0B · runs from 3.3 GB

240.7K 453

Hermes 3 Llama 3.1 8B is an 8-billion parameter instruction-tuned model by Nous Research, built on Meta's Llama 3.1 8B base. It is fine-tuned for advanced instruction following, multi-turn conversation, structured output, and creative roleplay scenarios. The Hermes series is known for producing highly steerable models that respond well to system prompts. This model supports a 128K token context window inherited from the Llama 3.1 architecture and runs efficiently on consumer GPUs with 8GB or more of VRAM. It is a popular choice among local inference enthusiasts who value strong instruction adherence and versatile conversational ability.

ChatRoleplay

All LLM Models

Understanding LLM VRAM Requirements

Model List

NVIDIA Nemotron Nano 9B v2 Japanese

Cydonia 24B V4.3

Gemma 4 E2B IT Qat Mobile Transformers

Qwen3.6 27B DFlash

Llama 3.3 70B Instruct Abliterated

Mixtral 8x7B Instruct v0.1

Gemma 4 12B OBLITERATED

Gemma 4 26B A4B IT Uncensored Heretic

Gemma 4 26B A4B IT Assistant

MiniMax M2.1

LFM2 8B A1B

Deepseek Coder 33B Instruct

NVIDIA Nemotron 3 Ultra 550B A55B BF16

SmolLM3 3B

GLM 4.5 Air

DeepSeek V3.1

Gemma 3n E2B IT

Gemma 4 E4B IT Assistant

Llama 3.2 11B Vision Instruct

Jan Code 4B

Qwen2.5 Coder 32B

MiniMax M2.7 BF16 Ultra Uncensored Heretic

Gemma 3n E4B IT

Qwen3 235B A22B

Hermes 4 14B

GLM 4.6

LFM2.5 350M

LFM2.5 1.2B Thinking

Qwen2.5 Coder 7B

Hermes 3 Llama 3.1 8B