All LLM Models
Browse 982 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Phi 4 Mini Instruct
Microsoft · 3.8B · runs from 1.9 GB
Microsoft Phi 4 Mini Instruct is a 3.8-billion parameter instruction-tuned model from Microsoft Research's Phi 4 family. It applies the Phi series' data-centric training philosophy to a compact model, delivering strong performance in coding, reasoning, and chat tasks relative to its small footprint. The model runs on consumer GPUs with as little as 4-6GB of VRAM when quantized, making it accessible on mainstream and even entry-level hardware. Released under the MIT license.
Mistral 7B Instruct v0.3
Mistral AI · 7.2B · runs from 2.7 GB
Mistral 7B Instruct v0.3 is the latest instruction-tuned release of Mistral AI's original 7-billion-parameter model, delivering meaningful improvements in instruction following, function calling, and multilingual support over its predecessors. With an extended 32K-token vocabulary and refined chat capabilities, v0.3 remains one of the most capable sub-10B models available. At 7.2 billion parameters it sits comfortably in the sweet spot for local inference, running well on GPUs with 6–8 GB of VRAM at full precision and even on 4 GB cards with 4-bit quantization. It is an excellent default choice for anyone getting started with local LLMs who wants strong conversational performance without heavy hardware.
Qwen2.5 14B
Alibaba · 14.8B · runs from 6.8 GB
Qwen2.5 14B is a 14.8B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Jan v3 4B Base Instruct
janhq · 4.4B · runs from 2.4 GB
Jan v3 4B Base Instruct is a 4.4B-parameter open language model from janhq. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 0.5B Instruct
Alibaba · 494M · runs from 0.5 GB
Qwen2.5 0.5B Instruct is the smallest instruction-tuned model in Alibaba Cloud's Qwen 2.5 family, with just 494 million parameters. It is designed for ultra-lightweight deployment scenarios where minimal hardware resources are available, running comfortably on virtually any modern GPU or even CPU-only configurations. Despite its tiny footprint, the model supports a 128K token context window and can handle basic chat, simple summarization, and lightweight instruction following. It is primarily useful for edge deployment, experimentation, and prototyping where model size is a critical constraint. Released under the Apache 2.0 license.
Gemma 2 2B IT
Google · 2.6B · runs from 0.9 GB
Google Gemma 2 2B IT is a 2-billion parameter instruction-tuned model from Google's Gemma 2 family, the smallest variant in the Gemma 2 series. It is designed for efficient local inference on resource-constrained hardware, handling basic conversational tasks and simple instruction following at minimal compute cost. The model can run on GPUs with as little as 4GB of VRAM when quantized, and even on CPU-only setups. Released under the Gemma license.
Agents A1 4B
InternScience · 4.5B · runs from 2 GB
Agents A1 4B is a 4.5B-parameter open language model from InternScience in the Agents-A1 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mellum2 12B A2.5B Instruct
JetBrains · 12.1B · runs from 4.0 GB
Mellum2 12B A2.5B Instruct is JetBrains' 12-billion-parameter mixture-of-experts coding model, with about 2.4 billion parameters active per token across 64 experts. Because only a fraction of the weights are computed for each token, it runs noticeably faster than a dense model of similar size, while all the experts still need to fit in memory. Unlike JetBrains' earlier Mellum, a narrow fill-in-the-middle completion model, Mellum2 is a full coding assistant that can generate and edit code, call tools, hold multi-turn conversations, and reason through problems. Its modest active-parameter count makes it practical to run on a single mainstream consumer GPU once quantized. The model supports a 131K token context window and is released under the Apache 2.0 license, allowing unrestricted commercial and research use. It was published in May 2026.
Ling 3.0 Tiny
Inclusion AI · 7.9B · runs from 2.8 GB
Ling 3.0 Tiny is InclusionAI's smallest mixture-of-experts model in Ant Group's Ling 3.0 family, totaling roughly 7.9 billion parameters with about 1.4 billion active per token. It shares the family's hybrid-linear-attention MoE architecture, also used at far larger scale in Ling-3.0-Flash, scaled down here for lightweight chat and agent use. Only the active parameters compute per token, so it responds quickly while still needing its full parameter set in memory; under 8 billion total, it runs on a mainstream consumer GPU once quantized. The model provides a 131K token context window for extended conversations or agent traces. It is released under the MIT license, allowing unrestricted commercial and research use. Published in August 2026, it extends Ant Group's Ling 3.0 lineup, built for efficient agent workloads.
Yi Coder 9B Chat
01.AI · 8.8B · runs from 3.1 GB
Yi Coder 9B Chat is a 8.8B-parameter open language model from 01.AI in the Yi family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3.5 Mini Instruct
Microsoft · 3.8B · runs from 2.3 GB
Phi 3.5 Mini Instruct is a 3.8B-parameter open language model from Microsoft in the Phi 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Meta Llama 3 8B Instruct
Meta · 8.0B · runs from 2.6 GB
Meta Llama 3 8B Instruct is the instruction-tuned version of Meta's Llama 3 8B base model, with 8 billion parameters. It is fine-tuned for dialogue and chat use cases using supervised fine-tuning and RLHF, making it ready for conversational applications out of the box. The model supports an 8K token context window and performs well across coding, reasoning, and general knowledge tasks. Its efficient size makes it one of the most popular models for local inference on consumer hardware. Released under the Meta Llama 3 Community License.
S1 Mini
superwhisper · 752M · runs from 0.7 GB
S1 Mini is a 752M-parameter open language model from superwhisper. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NVIDIA Nemotron Nano 9B v2
NVIDIA · 8.9B · runs from 3.5 GB
NVIDIA Nemotron Nano 9B v2 is a compact yet capable chat model from NVIDIA, packing 8.9 billion parameters into a size that runs comfortably on a wide range of consumer GPUs. Built on NVIDIA's Nemotron architecture, it delivers strong instruction-following and conversational performance while keeping VRAM requirements modest. This second-generation Nano model reflects NVIDIA's push to make high-quality language models accessible on local hardware. It's an excellent starting point for users who want a responsive, general-purpose assistant without needing top-tier GPU memory.
Qwen3 4B Thinking 2507
Alibaba · 4.0B · runs from 2.2 GB
Qwen3 4B Thinking 2507 is the reasoning-optimized variant of Alibaba's compact 4-billion-parameter Qwen3 model, released in the July 2025 update cycle. Despite its small size, this thinking variant is tuned to produce chain-of-thought reasoning and step-by-step problem solving, making it a surprisingly capable lightweight reasoner. This model is ideal for users who want basic reasoning and analytical capabilities on very modest hardware. It can run on most consumer GPUs and even some CPU-only setups when quantized, providing an accessible entry point for experimenting with reasoning-style models without any significant hardware investment.
LFM2.5 2.6B DSpark
Liquid AI · 328M · runs from 0.5 GB
LFM2.5 2.6B DSpark is a 328M-parameter open language model from Liquid AI in the LFM2.5 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 3 4B IT
Google · 4.3B · runs from 2.0 GB
Gemma 3 4B IT is a small, roughly 4.3-billion-parameter model from Google's earlier Gemma 3 generation, tuned for instruction-following and chat. It accepts image input alongside text, so it can handle visual question answering and image-grounded prompts in addition to text-only conversation. At this size, the model is comfortable on modest consumer hardware, running even on integrated graphics at low quantization, making it one of the most accessible vision-capable models for local use. Released on February 20, 2025, Gemma 3 4B IT predates the Gemma 4 family and is distributed under Google's Gemma license terms rather than a standard open-source license. Its compact size and image-input support make it a practical choice for lightweight, on-device multimodal applications.
Phi 4
Microsoft · 14.7B · runs from 7.0 GB
Microsoft Phi 4 is a 14-billion parameter language model from Microsoft Research's Phi series, designed to deliver strong reasoning, mathematical, and coding performance at an efficient size. Phi 4 continues the Phi family's focus on maximizing capability per parameter through high-quality training data curation, achieving benchmark scores that rival much larger models on reasoning and STEM tasks. The model runs well on consumer GPUs with 12-16GB of VRAM in quantized formats. It excels at mathematical problem solving, code generation, and structured reasoning. Released under the MIT license.
Mistral Nemo Instruct 2407
Mistral AI · 12.2B · runs from 5.9 GB
Mistral Nemo Instruct 2407 is a 12-billion-parameter instruction-tuned chat model from Mistral AI, a dated release from July 2024. It targets general dialogue and instruction-following use cases, sitting in a practical middle ground between lightweight and large-scale models in terms of both capability and resource demands. The model offers a 128K token context window, generous for its parameter class, and is released under the Apache 2.0 license, allowing unrestricted local and commercial use. At around 12 billion parameters, Mistral Nemo Instruct 2407 fits on a single consumer GPU when quantized, putting it within easy reach of local enthusiasts running their own hardware.
Qwen2 7B Instruct
Alibaba · 7.6B · runs from 2.7 GB
Qwen2 7B Instruct is a 7.6B-parameter open language model from Alibaba in the Qwen 2 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mistral Small Instruct 2409
Mistral AI · 22.2B · runs from 7.4 GB
Mistral Small Instruct 2409 is a 22.2B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Yi 1.5 6B Chat
01.AI · 6.1B · runs from 3.0 GB
Yi 1.5 6B Chat is a 6.1B-parameter open language model from 01.AI in the Yi 1.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ministral 3 3B Reasoning 2512
Mistral AI · 4.3B · runs from 2.3 GB
Ministral 3 3B Reasoning 2512 is a 4.3B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 3 12B IT
Google · 12.2B · runs from 5.7 GB
Google Gemma 3 12B IT is a 12-billion parameter multimodal instruction-tuned model from Google's Gemma 3 series. It supports both text and image inputs, offering vision-language capabilities at a more accessible size point than the 27B variant. Gemma 3 12B IT runs on consumer GPUs with 12-16GB of VRAM in quantized formats, making it a practical choice for local multimodal inference without requiring top-tier hardware. Released under the Gemma license.
Qwen2.5 Coder 1.5B Instruct
Alibaba · 1.5B · runs from 0.9 GB
Qwen2.5 Coder 1.5B Instruct is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Bonsai 8B Unpacked
prism-ml · 8.2B · runs from 4.1 GB
Bonsai 8B Unpacked is a 8.2B-parameter open language model from prism-ml. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Spark X2.5 1.7B
XHToken · 1.7B · runs from 1.1 GB
Spark X2.5 1.7B is a 1.7B-parameter open language model from XHToken. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 3.2 1B
Meta · 1.2B · runs from 0.6 GB
Meta Llama 3.2 1B is a 1.2-billion parameter base (pretrained) model from Meta's Llama 3.2 release. It is the smallest model in the Llama 3.2 family and is designed for research, fine-tuning, and embedding into resource-constrained environments. It supports a 128K token context window. As a base model, it is not optimized for conversational use without further fine-tuning. Its minimal resource requirements make it suitable for experimentation, edge deployment, and as a starting point for domain-specific fine-tuning. Released under the Llama 3.2 Community License.
Gemma 3 270M IT
Google · 268M · runs from 0.1 GB
Google Gemma 3 270M IT is a 270-million parameter instruction-tuned model from Google's Gemma 3 family, an experimental release pushing the boundaries of how small an effective chat model can be. The model runs on virtually any hardware, including entry-level GPUs and CPU-only setups, making it useful for experimentation, education, and exploring the limits of small-scale language modeling. Released under the Gemma license.
GLM 4.6V Flash
Z.ai · 10.3B · runs from 3.2 GB
GLM 4.6V Flash is a 10.3B-parameter open language model from Z.ai in the GLM 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.