All LLM Models
Browse 1242 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Supra Router 51M
SupraLabs · 52M · runs from 0.3 GB
Supra Router 51M is a 52M-parameter open language model from SupraLabs. It supports a context window of up to 5,120 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 31B
Google · 32.7B · runs from 15.5 GB
Gemma 4 31B is Google DeepMind's largest dense model in the Gemma 4 family, roughly 32.7 billion parameters, handling text and image input. This is the pretrained base checkpoint rather than an instruction-tuned model, meant as a foundation for fine-tuning rather than direct chat use. Gemma 4 introduces a hybrid attention design interleaving local sliding-window attention with occasional full global attention, plus configurable reasoning modes. At this size, local inference needs a high-end consumer or prosumer GPU, especially once quantized. It supports a 262,144 token context window, among the largest in the family. It carries Google's Gemma 4 license terms, published as Apache 2.0 on Hugging Face with additional Gemma-specific usage terms linked from the card. Published in March 2026, it is the largest of five Gemma 4 sizes (E2B, E4B, 12B, 26B-A4B MoE, 31B dense).
OLMoE 1B 7B 0924
Allen AI · 6.9B · runs from 3.5 GB
OLMoE 1B 7B 0924 is a 6.9B-parameter open language model from Allen AI in the OLMo family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Hunyuan A13B Instruct
Tencent · 80.4B · runs from 22.7 GB
Hunyuan A13B Instruct is a 80.4B-parameter open language model from Tencent in the Hunyuan family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ornith 1.0 35B AEON Ultimate Uncensored BF16
AEON-7 · 35.1B · runs from 15.3 GB
Ornith 1.0 35B AEON Ultimate Uncensored BF16 is a 35.1B-parameter open language model from AEON-7 in the Ornith family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Granite 3.0 2B Instruct
IBM · 2.6B · runs from 1.3 GB
Granite-3.0-2B-Instruct is IBM's 2-billion-parameter chat model, fine-tuned from Granite-3.0-2B-Base on a mix of permissively licensed open instruction datasets and IBM's own synthetic data, using supervised fine-tuning, reinforcement-learning alignment, and model merging. It is one of several Granite 3.0 sizes IBM released together, aimed at enterprise use cases such as summarization, question answering, and retrieval-augmented generation. It supports dialogue in twelve languages, primarily English, though multilingual performance trails English performance. At 2B parameters, it is small enough to run on a laptop CPU or any consumer GPU, including edge deployments. Context length is 4,096 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2024. It was later superseded by the Granite 3.1 model family.
TinyLlama 1.1B Intermediate Step 1431k 3T
TinyLlama · 1.1B · runs from 0.8 GB
TinyLlama 1.1B Intermediate Step 1431k 3T is a 1.1B-parameter open language model from TinyLlama in the TinyLlama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Starcoder2 3B
BigCode · 3.0B · runs from 1.6 GB
StarCoder2-3B is BigCode's 3-billion-parameter base code model, trained from scratch on 17 programming languages drawn from The Stack v2 using a fill-in-the-middle objective over more than 3 trillion tokens. It is a pretrained checkpoint rather than an instruction-tuned assistant, so it completes and continues code rather than following natural-language commands, and it uses grouped-query attention with a sliding-window mechanism for efficiency. It has a 16,384 token context window and is released under the BigCode OpenRAIL-M license, which carries use-based restrictions rather than being fully permissive. At 3 billion parameters, it runs easily on almost any modern laptop or consumer GPU, even without a high-end card, whether quantized or run at half precision.
EuroLLM 22B Instruct 2512
utter-project · 22.6B · runs from 7.0 GB
EuroLLM 22B Instruct 2512 is a 22.6B-parameter open language model from utter-project. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 31B IT Speculator.eagle3
RedHatAI · 31B · runs from 14.5 GB
Gemma 4 31B IT Speculator.eagle3 is a 31B-parameter open language model from RedHatAI in the Gemma 4 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Hermes 2 Theta Llama 3 70B
Nous Research · 70.6B · runs from 20.4 GB
Hermes 2 Theta Llama-3 70B is Nous Research's merged model combining its own Hermes 2 Pro with Meta's Llama-3-70B-Instruct, built with Charles Goddard and Arcee AI's MergeKit tool and then further RLHF-tuned on top of the merge. It uses the ChatML prompt format and is specifically trained for function calling, structured JSON outputs, and feature extraction from retrieval-augmented (RAG) documents, aiming at agentic and tool-using workflows rather than plain chat. At 70 billion parameters it needs a multi-GPU workstation or heavy quantization to run locally. Context length is 8,192 tokens, inherited from its Llama-3 base. It is released under the Llama 3 Community License, Meta's custom license permitting commercial use below 700 million monthly active users. It was published in June 2024.
H2o Danube3 500M Chat
h2oai · 514M · runs from 0.6 GB
H2o Danube3 500M Chat is a 514M-parameter open language model from h2oai. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3 Mini 4k Instruct
Microsoft · 3.8B · runs from 2.7 GB
Microsoft Phi 3 Mini 4K Instruct is a 3.8-billion parameter instruction-tuned model from Microsoft Research's Phi 3 generation, with a 4K token context window. The Phi 3 family demonstrated that small models trained on carefully curated, high-quality data can achieve performance competitive with models several times their size. The model runs on consumer GPUs with as little as 4-6GB of VRAM when quantized, making it one of the most accessible capable chat models for local deployment. Released under the MIT license.
Vikhr Nemo 12B Instruct R 21 09 24
Vikhrmodels · 12.2B · runs from 4.8 GB
Vikhr Nemo 12B Instruct R 21 09 24 is a 12.2B-parameter open language model from Vikhrmodels. It supports a context window of up to 1,024,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ling Flash 2.0
Inclusion AI · 102.9B · runs from 28.7 GB
Ling-flash-2.0 is Inclusion AI's non-reasoning chat and code model, a mixture-of-experts system with about 102.9 billion total parameters but only roughly 6.15 billion activated per token (4.8 billion non-embedding), built on the Ling 2.0 architecture with a roughly 1/32 expert activation ratio, aux-loss-free routing, multi-token-prediction layers, and partial RoPE. Trained on more than 20 trillion tokens with supervised fine-tuning and multi-stage reinforcement learning, Inclusion AI reports it matches dense models of around 40 billion parameters on complex reasoning, code generation, and frontend-development benchmarks despite its small active-parameter count, while running several times faster thanks to its sparsity. Its small active-parameter footprint keeps generation fast, but all of its parameters must stay in memory, so it needs a multi-GPU setup or a high-memory workstation even once quantized. Context length is natively 32,768 tokens, extendable to 128,000 tokens with YaRN. It is released under the MIT license, permitting unrestricted commercial and research use. It was published in September 2025, alongside the reasoning-focused Ring-flash-2.0 built on the same base.
Nanbeige4.1 3B Heretic
heretic-org · 3.9B · runs from 2.1 GB
Nanbeige4.1 3B Heretic is a 3.9B-parameter open language model from heretic-org. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 72B
Alibaba · 72.7B · runs from 31.0 GB
Qwen2.5-72B is the base, pretrained 72.7-billion-parameter language model from Alibaba's Qwen2.5 series — it is not instruction-tuned and is not intended for direct conversational use; Alibaba recommends applying SFT, RLHF, or further pretraining on top of it. It uses a Transformer architecture with RoPE, SwiGLU, RMSNorm, QKV attention bias, 80 layers, and grouped-query attention (64 query heads, 8 key/value heads), with stronger coding, math, and structured-output ability than the earlier Qwen2 line, plus multilingual coverage across 29-plus languages. At 72.7 billion parameters, it needs a multi-GPU workstation to run at full precision, though quantized versions fit fewer cards. Context length is 131,072 tokens. It is released under Alibaba's custom Qwen license, which permits research and commercial use but requires a separate license from Alibaba once a product or service passes 100 million monthly active users. It was published in September 2024.
Qwen2.5 7B
Alibaba · 7.6B · runs from 3.6 GB
Qwen2.5-7B is Alibaba's 7.6-billion-parameter base pretrained model in the Qwen2.5 series, one step up from the 1.5B checkpoint. Like its smaller sibling it is a raw causal language model rather than an instruction-tuned chat model: the card recommends applying supervised fine-tuning, RLHF, or further pretraining before using it conversationally. It uses a dense transformer with grouped-query attention, and at 7.6B parameters it fits comfortably on a single consumer GPU once quantized, or unquantized on a higher-memory card. Context length is 131,072 tokens, one of the longer windows in the 0.5B-to-72B Qwen2.5 lineup. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in September 2024. Qwen2.5 as a whole brought major gains in coding, math, and structured-data understanding over Qwen2.
INTELLECT 1 Instruct
PrimeIntellect · 10.2B · runs from 3.7 GB
INTELLECT-1-Instruct is the instruction-tuned version of INTELLECT-1, a 10-billion-parameter language model that Prime Intellect describes as the first collaboratively, globally distributed pretraining run of its scale: 1 trillion tokens of English text and code trained across up to 14 concurrent nodes on three continents contributed by 30 independent community participants, using the DiLoCo distributed-training algorithm. Post-training was handled by Arcee AI through supervised fine-tuning, eight rounds of direct preference optimization, and model merging, using the Llama-3 tokenizer and logits distilled from Llama-3.1-405B to improve alignment. At 10B parameters it fits a single consumer GPU once quantized, more comfortably at full precision on a higher-end card. Context length is 8,192 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in November 2024, alongside the base INTELLECT-1 model.
Pythia 70M Deduped
EleutherAI · 96M · runs from 0.0 GB
Pythia 70M Deduped is a 96M-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 3.1 Tulu 3 70B DPO
Allen AI · 70.6B · runs from 20.4 GB
Llama-3.1-Tulu-3-70B-DPO is Allen Institute for AI's (Ai2's) 70.6-billion-parameter instruction-following model, fine-tuned from Meta's Llama 3.1 70B base through Ai2's fully open Tulu 3 post-training recipe. This checkpoint is the direct preference optimization (DPO) stage, trained on Ai2's published preference mixture on top of the Tulu-3-70B-SFT model; a further RLVR stage produces the final Tulu-3-70B release. Tulu 3 targets strong performance not just on chat but on tasks like MATH, GSM8K, and IFEval, with all training data, code, and recipes released openly as a reference post-training pipeline. At 70.6 billion parameters, it needs a multi-GPU workstation or server to run, even once quantized. Context length is 131,072 tokens, inherited from the Llama 3.1 base. It is released under Meta's Llama 3.1 Community License, a custom license that is free for most commercial and research use but requires organizations with more than 700 million monthly active users to request separate permission from Meta. It was published in November 2024.
Openchat 3.5 0106
OpenChat · 7.2B · runs from 3.6 GB
OpenChat-3.5-0106 is a 7.2-billion-parameter chat model fine-tuned from Mistral-7B-v0.1 using C-RLFT (Conditioned Reinforcement Learning Fine-Tuning), a method that lets the model learn from mixed-quality data by conditioning on data source rather than requiring uniformly high-quality demonstrations. At release, OpenChat billed it as the best-performing open 7B chat model, adding a dedicated coding mode alongside its generalist and math-reasoning mode plus experimental evaluator and feedback capabilities. At 7B parameters it runs on a single consumer GPU, and on modest cards once quantized. Context length is 8,192 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in January 2024, an update to the earlier OpenChat-3.5 and OpenChat-3.5-1210 checkpoints.
Tmax 9B
Allen AI · 9.0B · runs from 4.4 GB
Tmax 9B is Allen Institute for AI's terminal-agent model, fine-tuned from Alibaba's Qwen3.5-9B using DPPO, a reinforcement-learning method, on a small curated dataset of terminal-agent trajectories, with the vision head removed so it operates as a text-only, language-model-only checkpoint. It is built specifically to drive command-line and terminal tasks rather than general chat, and clearly outperforms its Qwen3.5-9B base on Terminal-Bench evaluations. Being a single dense 9-billion-parameter model, it runs on a single consumer GPU. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, though Ai2 asks that it be used in line with its Responsible Use Guidelines. It was published in June 2026, as the 9B member of a family of terminal agents that also includes 2B, 4B, and 27B sizes.
Cogito V1 Preview Qwen 32B
deepcogito · 32B · runs from 10.4 GB
Cogito V1 Preview Qwen 32B is a 32B-parameter open language model from deepcogito in the Qwen family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2 0.5B
Alibaba · 494M · runs from 0.5 GB
Qwen2-0.5B is Alibaba's roughly 494-million-parameter dense base language model, the smallest member of the Qwen2 family that scales up to 72B and includes one Mixture-of-Experts variant. This is a pretrained base model, not an instruction-tuned chat model; Alibaba explicitly recommends applying supervised fine-tuning, RLHF, or further pretraining before using it for open-ended generation. Given its size, it runs easily on nearly any GPU and even on CPUs. Its configuration specifies a context window of up to 131,072 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2024. It was superseded within the same year by Qwen2.5-0.5B, part of Alibaba's fast release cadence for its small-model line.
Hermes 4.3 36B
Nous Research · 36.2B · runs from 10.5 GB
Hermes 4.3 36B is a 36.2B-parameter open language model from Nous Research in the Hermes family. It supports a context window of up to 524,288 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.6 27B DFlash
z-lab · 27B · runs from 11.8 GB
Qwen3.6 27B DFlash is a 27B-parameter open language model from z-lab in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 4 Reasoning
Microsoft · 14.7B · runs from 4.8 GB
Phi-4-reasoning is Microsoft's reasoning-focused fine-tune of Phi-4, a 14-billion-parameter dense decoder-only transformer, produced through supervised fine-tuning on chain-of-thought traces followed by reinforcement learning, concentrated on math, science, and coding tasks plus safety alignment data. The fine-tuning set combined synthetic prompts with high-quality filtered public web data, and responses are structured in two parts: a reasoning trace followed by a final summarized answer. Microsoft designed it for memory- and compute-constrained, latency-bound scenarios rather than frontier-scale serving, and at 14B parameters it can run on a single high-end consumer GPU once quantized. Context length is 32,768 tokens. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in April 2025.
Qwen3 0.6B Base
Alibaba · 596M · runs from 0.7 GB
Qwen3 0.6B Base is the smallest pretrained foundation model in Alibaba Cloud's Qwen 3 family, with approximately 600 million parameters. As a base model, it is not tuned for chat or instructions and is intended for fine-tuning, research, and experimentation. Its minimal size makes it suitable for rapid prototyping and resource-constrained training experiments. The model runs on virtually any hardware, including CPU-only setups. It is useful for educational purposes, architecture exploration, and as a compact foundation for task-specific fine-tuning where model size is a primary constraint. Released under the Apache 2.0 license.
GPT OSS 20B RichardErkhov Heresy
MuXodious · 21.5B · runs from 6.3 GB
GPT OSS 20B RichardErkhov Heresy is a 21.5B-parameter open language model from MuXodious in the GPT-OSS family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.