All LLM Models

Browse 1214 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

DeepSeek v2 Lite Chat

DeepSeek · 15.7B · runs from 5.1 GB

143.5K 148

DeepSeek v2 Lite Chat is a 15.7B-parameter open language model from DeepSeek in the DeepSeek V2 family. It supports a context window of up to 163,840 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Ternary Bonsai 4B Unpacked

prism-ml · 4.0B · runs from 2.2 GB

1.5K 7

Ternary Bonsai 4B Unpacked is a 4.0B-parameter open language model from prism-ml. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 4 31B IT Qat Q4 0 Unquantized Assistant

Google · 31B · runs from 13.5 GB

16.0K 14

Gemma 4 31B IT Qat Q4 0 Unquantized Assistant is a 31B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Zephyr 7B Beta

Hugging Face · 7.2B · runs from 3.6 GB

61.1K 1.9K

Zephyr 7B Beta is a 7.2B-parameter open language model from Hugging Face. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Yi 1.5 34B Chat

01.AI · 34.4B · runs from 12.4 GB

17.1K 278

Yi 1.5 34B Chat is a 34.4-billion parameter instruction-tuned model by 01.AI, the Chinese AI lab founded by Kai-Fu Lee. It is a bilingual model with strong performance in both English and Chinese, making it particularly well suited for users who need high-quality generation in either language. Yi 1.5 represents an improved iteration of the Yi model family with enhanced reasoning and coding ability. The 34B size requires a GPU with at least 24GB of VRAM for quantized inference, placing it within reach of high-end consumer cards like the RTX 4090. Released under the Yi License.

Chat

CodeLlama 34B Instruct HF

Meta · 33.7B · runs from 10.0 GB

21.2K 305

CodeLlama-34b-Instruct-hf is Meta's 34-billion-parameter instruction-tuned Code Llama model, fine-tuned from the base Code Llama checkpoint for safer, more reliable code-assistant use, as opposed to the plain base and Python-specialized variants in the same family (which also ships in 7B, 13B, and 70B sizes). It is an early, first-generation code model built on the original Llama 2 architecture and predates the later Code Llama 70B release. At 34B parameters, it needs a multi-GPU setup or aggressive quantization for local use. Context length is 16,384 tokens. It is released under the Llama 2 Community License, a custom license that restricts commercial use above 700 million monthly active users, requiring a separate license from Meta, and was published in August 2023.

ChatCode

Pantheon Reasoning 27B

Gryphe · 27.8B · runs from 8.4 GB

175 23

Pantheon Reasoning 27B is a 27.8B-parameter open language model from Gryphe. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatRoleplayReasoning

Starcoder2 15B

BigCode · 16.0B · runs from 7.3 GB

4.9K 675

Starcoder2 15B is a 16.0B-parameter open language model from BigCode in the StarCoder family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Huihui Qwen3.6 35B A3B Claude 4.7 Opus Abliterated

huihui-ai · 36.0B · runs from 15.7 GB

5.7K 217

Huihui Qwen3.6 35B A3B Claude 4.7 Opus Abliterated is a 36.0B-parameter open language model from huihui-ai in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Qwopus3.5 9B v3

Jackrong · 9.7B · runs from 19.9 GB

4.0K 92

Qwopus3.5 9B v3 is a 9.7B-parameter open language model from Jackrong. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

VisionReasoning

Olmo 3 32B Think

Allen AI · 32.2B · runs from 9.7 GB

10.7K 177

Olmo 3 32B Think is Allen Institute for AI's reasoning-focused model in the fully open Olmo 3 family, trained to produce long chains of thought before answering so it can work through math and coding problems step by step. It is pretrained on Ai2's Dolma 3 corpus and then carried through a full post-training pipeline of supervised fine-tuning, direct preference optimization, and reinforcement learning with verifiable rewards on the Dolci-Think-RL dataset; Ai2 releases the code, intermediate checkpoints, and training logs for each stage, not just the final weights. It sits alongside a 7B Think and a 7B Instruct sibling in the same release. At 32 billion parameters it needs a high-end consumer GPU or a multi-GPU setup once quantized. Context length is 65,536 tokens. It is released under the Apache 2.0 license, though Ai2 asks that it be used for research and educational purposes in line with its Responsible Use Guidelines. It was published in November 2025 and has since been superseded by Olmo 3.1 32B Think.

Chat

Granite 3.3 2B Instruct

IBM · 2.5B · runs from 1.2 GB

56.3K 87

Granite 3.3 2B Instruct is a 2.5B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nemotron Labs Audex 30B A3B

NVIDIA · 30B · runs from 14.0 GB

2.2K 163

Nemotron Labs Audex 30B A3B is a 30B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Ministral 8B Instruct 2410

Mistral AI · 8.0B · runs from 3.3 GB

152.8K 586

Ministral 8B Instruct 2410 is Mistral AI's 8-billion-parameter instruction-tuned model, part of the "Ministraux" family aimed at edge and on-device deployment. It uses a 36-layer dense transformer with interleaved sliding-window attention, trained heavily on multilingual and code data, suiting chat, function calling, and general assistant tasks across ten languages. Its compact size means it runs comfortably on a single mainstream consumer GPU once quantized. The model supports a 32K token context window. It is released under Mistral's own license, listed as "Other," which restricts use to research purposes and requires a separate commercial license from Mistral AI, so review the terms before commercial use. It was published in October 2024 alongside the smaller Ministral 3B.

Chat

EXAONE 3.5 2.4B Instruct

LGAI-EXAONE · 2.4B · runs from 0.9 GB

61.3K 196

EXAONE 3.5 2.4B Instruct is a 2.4B-parameter open language model from LGAI-EXAONE in the EXAONE family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Lucid V1 Nemo

dreamgen · 12.2B · runs from 4.1 GB

217.4K 57

Lucid V1 Nemo is a 12.2B-parameter open language model from dreamgen. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

LocateAnything 3B

NVIDIA · 3.8B · runs from 2 GB

87.4K 3.0K

LocateAnything-3B is NVIDIA's 3.8-billion-parameter vision-language model for visual grounding rather than open-ended chat: referring-expression grounding, dense multi-object detection, GUI element grounding, point-based localization, and document or OCR layout grounding. Its core contribution is Parallel Box Decoding, which predicts a complete bounding box in one parallel step instead of token-by-token autoregressive decoding, giving up to 2.5x higher throughput while preserving geometric consistency. It combines a Qwen2.5-3B-Instruct language backbone with a MoonViT-SO-400M vision encoder and was trained on 12 million images with over 138 million grounding queries; NVIDIA has folded it into the Nemotron 3 Nano Omni model for agentic and computer-use grounding. At under 4 billion parameters, it runs on a single consumer GPU. Context length is 32,768 tokens. It is released under NVIDIA's non-commercial research license, permitting academic and non-profit use only; commercial use requires a separate license from NVIDIA. It was published in May 2026.

Vision

QwQ 32B Preview

Alibaba · 32.8B · runs from 10.7 GB

15.4K 1.7K

QwQ 32B Preview is a 32.8B-parameter open language model from Alibaba in the QwQ family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Ternary Bonsai 1.7B Unpacked

prism-ml · 1.7B · runs from 1.3 GB

1.5K 8

Ternary Bonsai 1.7B Unpacked is a 1.7B-parameter open language model from prism-ml. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Uyu 2 28B

mente-ai · 28.2B · runs from 13.6 GB

2.7K 21

Uyu 2 28B is a 28.2B-parameter open language model from mente-ai. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatRoleplay

Dream V0 Instruct 7B

Dream-org · 7.6B · runs from 2.5 GB

98.5K 160

Dream V0 Instruct 7B is a 7.6B-parameter open language model from Dream-org. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

CodeLlama 7B Instruct HF

Meta · 6.7B · runs from 4.2 GB

41.0K 258

CodeLlama 7B Instruct HF is a 6.7B-parameter open language model from Meta in the Code Llama family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Hy MT2 30B A3B

Tencent · 30.1B · runs from 13.2 GB

23.0K 501

Hy-MT2-30B-A3B is Tencent's largest "fast-thinking" multilingual translation model in the Hy-MT2 family, alongside smaller 1.8B and 7B siblings, built specifically for translation rather than general chat and tuned to follow translation instructions across 33 languages. It is a mixture-of-experts model with roughly 30 billion total and 3.5 billion active parameters per token, and Tencent reports it beating open models such as DeepSeek-V4-Pro and Kimi K2.6 on translation quality in fast-thinking mode. The release also ships an FP8-quantized checkpoint, and the smaller 1.8B sibling gets extreme sub-2-bit GGUF quantizations for on-device use. With roughly 3.5 billion active parameters, the 30B-A3B checkpoint is light enough to run on a single consumer GPU once quantized. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2026, alongside the smaller Hy-MT2-1.8B and Hy-MT2-7B models and the IFMTBench translation-instruction benchmark.

Translation

Yi 34B

01.AI · 34.4B · runs from 15.0 GB

11.1K 1.3K

Yi-34B is 01.AI's base pretrained language model, not instruction-tuned, with about 34.4 billion dense parameters trained on a 3-trillion-token bilingual Chinese-English corpus. At release it ranked first among open-source base models on benchmarks including the Hugging Face Open LLM Leaderboard and C-Eval, ahead of larger models such as Falcon-180B and Llama-2-70B. A long-context Yi-34B-200K variant and an instruction-tuned Yi-34B-Chat were released alongside it. Running the dense 34B model requires a high-end consumer GPU or multi-GPU setup, especially unquantized. Context length is 4,096 tokens, the default window for the 34B series before the extended 200K variant. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in November 2023; by current standards it is an older-generation base model, since superseded by 01.AI's Yi-1.5 and later families.

Chat

Supra Router 51M

SupraLabs · 52M · runs from 0.3 GB

2.8K 132

Supra Router 51M is a 52M-parameter open language model from SupraLabs. It supports a context window of up to 5,120 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 4 31B

Google · 32.7B · runs from 15.5 GB

683.3K 539

Gemma 4 31B is Google DeepMind's largest dense model in the Gemma 4 family, roughly 32.7 billion parameters, handling text and image input. This is the pretrained base checkpoint rather than an instruction-tuned model, meant as a foundation for fine-tuning rather than direct chat use. Gemma 4 introduces a hybrid attention design interleaving local sliding-window attention with occasional full global attention, plus configurable reasoning modes. At this size, local inference needs a high-end consumer or prosumer GPU, especially once quantized. It supports a 262,144 token context window, among the largest in the family. It carries Google's Gemma 4 license terms, published as Apache 2.0 on Hugging Face with additional Gemma-specific usage terms linked from the card. Published in March 2026, it is the largest of five Gemma 4 sizes (E2B, E4B, 12B, 26B-A4B MoE, 31B dense).

Vision

OLMoE 1B 7B 0924

Allen AI · 6.9B · runs from 3.5 GB

126.3K 153

OLMoE 1B 7B 0924 is a 6.9B-parameter open language model from Allen AI in the OLMo family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Hunyuan A13B Instruct

Tencent · 80.4B · runs from 22.7 GB

51.2K 795

Hunyuan A13B Instruct is a 80.4B-parameter open language model from Tencent in the Hunyuan family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Ornith 1.0 35B AEON Ultimate Uncensored BF16

AEON-7 · 35.1B · runs from 15.3 GB

416.7K 29

Ornith 1.0 35B AEON Ultimate Uncensored BF16 is a 35.1B-parameter open language model from AEON-7 in the Ornith family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoningVision

Granite 3.0 2B Instruct

IBM · 2.6B · runs from 1.3 GB

5.2K 48

Granite-3.0-2B-Instruct is IBM's 2-billion-parameter chat model, fine-tuned from Granite-3.0-2B-Base on a mix of permissively licensed open instruction datasets and IBM's own synthetic data, using supervised fine-tuning, reinforcement-learning alignment, and model merging. It is one of several Granite 3.0 sizes IBM released together, aimed at enterprise use cases such as summarization, question answering, and retrieval-augmented generation. It supports dialogue in twelve languages, primarily English, though multilingual performance trails English performance. At 2B parameters, it is small enough to run on a laptop CPU or any consumer GPU, including edge deployments. Context length is 4,096 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2024. It was later superseded by the Granite 3.1 model family.

Chat