All LLM Models
Browse 1138 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
QwQ 32B Preview
Alibaba · 32.8B · runs from 10.7 GB
QwQ 32B Preview is a 32.8B-parameter open language model from Alibaba in the QwQ family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ternary Bonsai 1.7B Unpacked
prism-ml · 1.7B · runs from 1.3 GB
Ternary Bonsai 1.7B Unpacked is a 1.7B-parameter open language model from prism-ml. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Uyu 2 28B
mente-ai · 28.2B · runs from 13.6 GB
Uyu 2 28B is a 28.2B-parameter open language model from mente-ai. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Dream V0 Instruct 7B
Dream-org · 7.6B · runs from 2.5 GB
Dream V0 Instruct 7B is a 7.6B-parameter open language model from Dream-org. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
CodeLlama 7B Instruct HF
Meta · 6.7B · runs from 4.2 GB
CodeLlama 7B Instruct HF is a 6.7B-parameter open language model from Meta in the Code Llama family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Hy MT2 30B A3B
Tencent · 30.1B · runs from 13.2 GB
Hy-MT2-30B-A3B is Tencent's largest "fast-thinking" multilingual translation model in the Hy-MT2 family, alongside smaller 1.8B and 7B siblings, built specifically for translation rather than general chat and tuned to follow translation instructions across 33 languages. It is a mixture-of-experts model with roughly 30 billion total and 3.5 billion active parameters per token, and Tencent reports it beating open models such as DeepSeek-V4-Pro and Kimi K2.6 on translation quality in fast-thinking mode. The release also ships an FP8-quantized checkpoint, and the smaller 1.8B sibling gets extreme sub-2-bit GGUF quantizations for on-device use. With roughly 3.5 billion active parameters, the 30B-A3B checkpoint is light enough to run on a single consumer GPU once quantized. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2026, alongside the smaller Hy-MT2-1.8B and Hy-MT2-7B models and the IFMTBench translation-instruction benchmark.
Yi 34B
01.AI · 34.4B · runs from 15.0 GB
Yi-34B is 01.AI's base pretrained language model, not instruction-tuned, with about 34.4 billion dense parameters trained on a 3-trillion-token bilingual Chinese-English corpus. At release it ranked first among open-source base models on benchmarks including the Hugging Face Open LLM Leaderboard and C-Eval, ahead of larger models such as Falcon-180B and Llama-2-70B. A long-context Yi-34B-200K variant and an instruction-tuned Yi-34B-Chat were released alongside it. Running the dense 34B model requires a high-end consumer GPU or multi-GPU setup, especially unquantized. Context length is 4,096 tokens, the default window for the 34B series before the extended 200K variant. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in November 2023; by current standards it is an older-generation base model, since superseded by 01.AI's Yi-1.5 and later families.
Supra Router 51M
SupraLabs · 52M · runs from 0.3 GB
Supra Router 51M is a 52M-parameter open language model from SupraLabs. It supports a context window of up to 5,120 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 31B
Google · 32.7B · runs from 15.5 GB
Gemma 4 31B is Google DeepMind's largest dense model in the Gemma 4 family, roughly 32.7 billion parameters, handling text and image input. This is the pretrained base checkpoint rather than an instruction-tuned model, meant as a foundation for fine-tuning rather than direct chat use. Gemma 4 introduces a hybrid attention design interleaving local sliding-window attention with occasional full global attention, plus configurable reasoning modes. At this size, local inference needs a high-end consumer or prosumer GPU, especially once quantized. It supports a 262,144 token context window, among the largest in the family. It carries Google's Gemma 4 license terms, published as Apache 2.0 on Hugging Face with additional Gemma-specific usage terms linked from the card. Published in March 2026, it is the largest of five Gemma 4 sizes (E2B, E4B, 12B, 26B-A4B MoE, 31B dense).
OLMoE 1B 7B 0924
Allen AI · 6.9B · runs from 3.5 GB
OLMoE 1B 7B 0924 is a 6.9B-parameter open language model from Allen AI in the OLMo family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ornith 1.0 35B AEON Ultimate Uncensored BF16
AEON-7 · 35.1B · runs from 15.3 GB
Ornith 1.0 35B AEON Ultimate Uncensored BF16 is a 35.1B-parameter open language model from AEON-7 in the Ornith family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Granite 3.0 2B Instruct
IBM · 2.6B · runs from 1.3 GB
Granite-3.0-2B-Instruct is IBM's 2-billion-parameter chat model, fine-tuned from Granite-3.0-2B-Base on a mix of permissively licensed open instruction datasets and IBM's own synthetic data, using supervised fine-tuning, reinforcement-learning alignment, and model merging. It is one of several Granite 3.0 sizes IBM released together, aimed at enterprise use cases such as summarization, question answering, and retrieval-augmented generation. It supports dialogue in twelve languages, primarily English, though multilingual performance trails English performance. At 2B parameters, it is small enough to run on a laptop CPU or any consumer GPU, including edge deployments. Context length is 4,096 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2024. It was later superseded by the Granite 3.1 model family.
TinyLlama 1.1B Intermediate Step 1431k 3T
TinyLlama · 1.1B · runs from 0.8 GB
TinyLlama 1.1B Intermediate Step 1431k 3T is a 1.1B-parameter open language model from TinyLlama in the TinyLlama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Starcoder2 3B
BigCode · 3.0B · runs from 1.6 GB
StarCoder2-3B is BigCode's 3-billion-parameter base code model, trained from scratch on 17 programming languages drawn from The Stack v2 using a fill-in-the-middle objective over more than 3 trillion tokens. It is a pretrained checkpoint rather than an instruction-tuned assistant, so it completes and continues code rather than following natural-language commands, and it uses grouped-query attention with a sliding-window mechanism for efficiency. It has a 16,384 token context window and is released under the BigCode OpenRAIL-M license, which carries use-based restrictions rather than being fully permissive. At 3 billion parameters, it runs easily on almost any modern laptop or consumer GPU, even without a high-end card, whether quantized or run at half precision.
EuroLLM 22B Instruct 2512
utter-project · 22.6B · runs from 7.0 GB
EuroLLM 22B Instruct 2512 is a 22.6B-parameter open language model from utter-project. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 31B IT Speculator.eagle3
RedHatAI · 31B · runs from 14.5 GB
Gemma 4 31B IT Speculator.eagle3 is a 31B-parameter open language model from RedHatAI in the Gemma 4 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
H2o Danube3 500M Chat
h2oai · 514M · runs from 0.6 GB
H2o Danube3 500M Chat is a 514M-parameter open language model from h2oai. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3 Mini 4k Instruct
Microsoft · 3.8B · runs from 2.7 GB
Microsoft Phi 3 Mini 4K Instruct is a 3.8-billion parameter instruction-tuned model from Microsoft Research's Phi 3 generation, with a 4K token context window. The Phi 3 family demonstrated that small models trained on carefully curated, high-quality data can achieve performance competitive with models several times their size. The model runs on consumer GPUs with as little as 4-6GB of VRAM when quantized, making it one of the most accessible capable chat models for local deployment. Released under the MIT license.
Vikhr Nemo 12B Instruct R 21 09 24
Vikhrmodels · 12.2B · runs from 4.8 GB
Vikhr Nemo 12B Instruct R 21 09 24 is a 12.2B-parameter open language model from Vikhrmodels. It supports a context window of up to 1,024,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Nanbeige4.1 3B Heretic
heretic-org · 3.9B · runs from 2.1 GB
Nanbeige4.1 3B Heretic is a 3.9B-parameter open language model from heretic-org. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 7B
Alibaba · 7.6B · runs from 3.6 GB
Qwen2.5-7B is Alibaba's 7.6-billion-parameter base pretrained model in the Qwen2.5 series, one step up from the 1.5B checkpoint. Like its smaller sibling it is a raw causal language model rather than an instruction-tuned chat model: the card recommends applying supervised fine-tuning, RLHF, or further pretraining before using it conversationally. It uses a dense transformer with grouped-query attention, and at 7.6B parameters it fits comfortably on a single consumer GPU once quantized, or unquantized on a higher-memory card. Context length is 131,072 tokens, one of the longer windows in the 0.5B-to-72B Qwen2.5 lineup. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in September 2024. Qwen2.5 as a whole brought major gains in coding, math, and structured-data understanding over Qwen2.
INTELLECT 1 Instruct
PrimeIntellect · 10.2B · runs from 3.7 GB
INTELLECT-1-Instruct is the instruction-tuned version of INTELLECT-1, a 10-billion-parameter language model that Prime Intellect describes as the first collaboratively, globally distributed pretraining run of its scale: 1 trillion tokens of English text and code trained across up to 14 concurrent nodes on three continents contributed by 30 independent community participants, using the DiLoCo distributed-training algorithm. Post-training was handled by Arcee AI through supervised fine-tuning, eight rounds of direct preference optimization, and model merging, using the Llama-3 tokenizer and logits distilled from Llama-3.1-405B to improve alignment. At 10B parameters it fits a single consumer GPU once quantized, more comfortably at full precision on a higher-end card. Context length is 8,192 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in November 2024, alongside the base INTELLECT-1 model.
Pythia 70M Deduped
EleutherAI · 96M · runs from 0.0 GB
Pythia 70M Deduped is a 96M-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Openchat 3.5 0106
OpenChat · 7.2B · runs from 3.6 GB
OpenChat-3.5-0106 is a 7.2-billion-parameter chat model fine-tuned from Mistral-7B-v0.1 using C-RLFT (Conditioned Reinforcement Learning Fine-Tuning), a method that lets the model learn from mixed-quality data by conditioning on data source rather than requiring uniformly high-quality demonstrations. At release, OpenChat billed it as the best-performing open 7B chat model, adding a dedicated coding mode alongside its generalist and math-reasoning mode plus experimental evaluator and feedback capabilities. At 7B parameters it runs on a single consumer GPU, and on modest cards once quantized. Context length is 8,192 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in January 2024, an update to the earlier OpenChat-3.5 and OpenChat-3.5-1210 checkpoints.
Tmax 9B
Allen AI · 9.0B · runs from 4.4 GB
Tmax 9B is Allen Institute for AI's terminal-agent model, fine-tuned from Alibaba's Qwen3.5-9B using DPPO, a reinforcement-learning method, on a small curated dataset of terminal-agent trajectories, with the vision head removed so it operates as a text-only, language-model-only checkpoint. It is built specifically to drive command-line and terminal tasks rather than general chat, and clearly outperforms its Qwen3.5-9B base on Terminal-Bench evaluations. Being a single dense 9-billion-parameter model, it runs on a single consumer GPU. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, though Ai2 asks that it be used in line with its Responsible Use Guidelines. It was published in June 2026, as the 9B member of a family of terminal agents that also includes 2B, 4B, and 27B sizes.
Cogito V1 Preview Qwen 32B
deepcogito · 32B · runs from 10.4 GB
Cogito V1 Preview Qwen 32B is a 32B-parameter open language model from deepcogito in the Qwen family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2 0.5B
Alibaba · 494M · runs from 0.5 GB
Qwen2-0.5B is Alibaba's roughly 494-million-parameter dense base language model, the smallest member of the Qwen2 family that scales up to 72B and includes one Mixture-of-Experts variant. This is a pretrained base model, not an instruction-tuned chat model; Alibaba explicitly recommends applying supervised fine-tuning, RLHF, or further pretraining before using it for open-ended generation. Given its size, it runs easily on nearly any GPU and even on CPUs. Its configuration specifies a context window of up to 131,072 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2024. It was superseded within the same year by Qwen2.5-0.5B, part of Alibaba's fast release cadence for its small-model line.
Hermes 4.3 36B
Nous Research · 36.2B · runs from 10.5 GB
Hermes 4.3 36B is a 36.2B-parameter open language model from Nous Research in the Hermes family. It supports a context window of up to 524,288 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.6 27B DFlash
z-lab · 27B · runs from 11.8 GB
Qwen3.6 27B DFlash is a 27B-parameter open language model from z-lab in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 4 Reasoning
Microsoft · 14.7B · runs from 4.8 GB
Phi-4-reasoning is Microsoft's reasoning-focused fine-tune of Phi-4, a 14-billion-parameter dense decoder-only transformer, produced through supervised fine-tuning on chain-of-thought traces followed by reinforcement learning, concentrated on math, science, and coding tasks plus safety alignment data. The fine-tuning set combined synthetic prompts with high-quality filtered public web data, and responses are structured in two parts: a reasoning trace followed by a final summarized answer. Microsoft designed it for memory- and compute-constrained, latency-bound scenarios rather than frontier-scale serving, and at 14B parameters it can run on a single high-end consumer GPU once quantized. Context length is 32,768 tokens. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in April 2025.