All LLM Models
Browse 880 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Qwen2.5 7B Instruct
Alibaba · 7.6B · runs from 2.7 GB
Qwen2.5 7B Instruct is a 7.6-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It supports a 128K token context window and is fine-tuned for conversational AI, instruction following, and general assistant tasks. Its efficient size makes it well-suited for local deployment on consumer GPUs with 8GB or more of VRAM. The model delivers strong performance for its parameter class across reasoning, multilingual understanding, and coding tasks. It benefits from the improved pretraining data and techniques of the Qwen 2.5 generation. Released under the Apache 2.0 license and widely supported by inference frameworks such as llama.cpp, vLLM, and Ollama.
Qwen3 14B
Alibaba · 14.8B · runs from 4.7 GB
Qwen3 14B is a 14-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 series. It occupies a practical middle ground in the Qwen 3 lineup, offering stronger reasoning and generation quality than the 8B variant while remaining manageable on GPUs with 16GB or more of VRAM in quantized formats. The model supports hybrid thinking mode for flexible reasoning depth. Qwen3 14B is well suited for chat, instruction following, coding assistance, and multilingual tasks. It benefits from the generational improvements of Qwen 3 in pretraining data and alignment techniques, delivering performance that competes with larger models from previous generations. Released under the Apache 2.0 license.
Qwen3 8B
Alibaba · 8.2B · runs from 2.9 GB
Qwen3 8B is an 8.2-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 series. It is a general-purpose chat model that delivers strong performance across reasoning, multilingual understanding, and coding tasks while remaining efficient enough to run on consumer GPUs with 8GB or more of VRAM. Like other Qwen 3 models, it supports hybrid thinking mode for flexible reasoning depth. The model benefits from the improved pretraining data and training methodology of the Qwen 3 generation, offering notable quality gains over Qwen 2.5 at the same parameter count. It is widely supported by inference frameworks including llama.cpp, vLLM, and Ollama. Released under the Apache 2.0 license.
Qwen2.5 Coder 7B Instruct
Alibaba · 7.6B · runs from 3.0 GB
Qwen2.5 Coder 7B Instruct is a 7.6-billion parameter code-specialized instruction-tuned model from Alibaba Cloud. It is trained on a large corpus of source code and natural language, fine-tuned for programming assistance tasks such as code generation, completion, debugging, and code explanation. The model supports a 128K token context window and runs efficiently on consumer GPUs with 8GB or more of VRAM. It provides a good balance between coding capability and hardware requirements for developers looking to run a local coding assistant. Released under the Apache 2.0 license.
Gemma 4 E4B IT
Google · 8.0B · runs from 3.2 GB
Gemma 4 E4B IT packs Google's Gemma 4 architecture into a compact, roughly 8-billion-parameter footprint, positioned as the mid-sized option in Gemma 4's efficiency-focused E-series alongside the smaller E2B variant. It is tuned for chat and instruction-following rather than multimodal input, focusing on dialogue quality within a lightweight package. Parameter counts in this range are well suited to local inference on a single consumer GPU, even at higher precision, and comfortably so once quantized. The model provides a 128K token context window, sufficient for most chat and document-assistance use cases without needing the largest Gemma 4 variants. It is released under the Apache 2.0 license, and was published on March 2, 2026, shortly before the larger dense and mixture-of-experts Gemma 4 models that followed.
Qwen2.5 14B Instruct
Alibaba · 14.8B · runs from 5.1 GB
Qwen2.5 14B Instruct is a 14-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It supports a 128K token context window and provides a balanced tradeoff between quality and hardware requirements, running well on GPUs with 16GB of VRAM in quantized formats. The model is fine-tuned for chat, instruction following, and general-purpose assistant tasks. It performs well across reasoning, coding, and multilingual benchmarks for its size class, making it a practical option for local deployment when larger models are not feasible. Released under the Apache 2.0 license.
Qwen3.5 9B
Alibaba · 9.7B · runs from 3.2 GB
Qwen3.5-9B is Alibaba's dense 9-billion-parameter language model with a built-in vision encoder, part of the Qwen3.5 generation. This is the post-trained, instruction-tuned release, combining Gated DeltaNet and gated-attention layers for efficient long-context inference, and it can take images alongside text for tasks like visual question answering as well as general chat, coding, and tool use. It offers a 262,144-token native context window, extensible up to roughly 1,010,000 tokens, and is released under the Apache 2.0 license. At around 9.7 billion parameters, 4-bit quantization needs only about 5.6GB of memory, so it fits comfortably on a single mid-range consumer GPU.
Qwen3.5 9B The Defiant Fable Uncensored Heretic NEO IMATRIX MAX MTP
DavidAU · 9.7B · runs from 3.8 GB
Qwen3.5 9B The Defiant Fable Uncensored Heretic NEO IMATRIX MAX MTP is a 9.7B-parameter open language model from DavidAU in the Qwen 3.5 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 12B IT Qat Q4 0 Unquantized
Google · 12.0B · runs from 6.1 GB
Gemma 4 12B IT Qat Q4 0 Unquantized is a 12.0B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 3.1 8B Instruct
Meta · 8.0B · runs from 3.6 GB
Meta Llama 3.1 8B Instruct is an 8-billion parameter instruction-tuned language model from Meta. Part of the Llama 3.1 release, it supports a 128K token context window and is fine-tuned for conversational use, tool calling, and general assistant tasks. Its compact size makes it well-suited for local deployment on modern consumer GPUs with 8GB or more of VRAM. Llama 3.1 8B Instruct delivers strong performance for its parameter class across benchmarks in reasoning, coding, and multilingual understanding. It is released under the Llama 3.1 Community License and is widely supported by inference frameworks such as llama.cpp, vLLM, and Ollama.
Qwen3 4B
Alibaba · 4.0B · runs from 1.6 GB
Qwen3 4B is a compact 4-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 family. It is designed for efficient local inference on consumer hardware, supporting chat and general assistant tasks while fitting comfortably on GPUs with 6GB or more of VRAM in quantized formats. The model supports hybrid thinking mode, allowing it to balance reasoning depth and response speed. Despite its small footprint, Qwen3 4B delivers quality competitive with larger models from previous generations, making it a practical choice for lightweight local deployments and resource-constrained environments. Released under the Apache 2.0 license.
Qwen2.5 Coder 14B Instruct
Alibaba · 14.8B · runs from 5.1 GB
Qwen2.5 Coder 14B Instruct is a 14.8B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 4B Instruct 2507
Alibaba · 4.0B · runs from 1.9 GB
Qwen3 4B Instruct 2507 is a July 2025 refresh of Alibaba's compact 4-billion-parameter chat model from the Qwen3 family. This updated release brings improved instruction following and conversational quality while remaining lightweight enough to run on most modern GPUs and even some higher-end integrated graphics setups. With its modest size, the 4B Instruct 2507 strikes a practical balance between capability and resource efficiency. It is well suited for everyday chat, summarization, and light assistant tasks on consumer hardware, making it one of the more accessible entry points into the Qwen3 lineup.
LFM2.5 2.6B
Liquid AI · 2.7B · runs from 1.6 GB
LFM2.5-2.6B is Liquid AI's on-device model in the LFM2.5 family, a hybrid architecture mixing gated short-convolution blocks with grouped-query attention layers, at about 2.7 billion parameters. It is post-trained for agentic workloads, tool use, data extraction, retrieval-augmented generation, and long-context tasks, though the publisher does not recommend it for heavy agentic coding or knowledge-intensive question answering. It supports a 131K token context window and is released under Liquid AI's custom LFM1.0 license. At around 2.7 billion parameters, it runs easily on almost any modern laptop or consumer GPU at 4-bit quantization, and the publisher reports it using well under 2.5 GB of memory during on-device inference.
Qwen3 0.6B
Alibaba · 752M · runs from 0.6 GB
Qwen3 0.6B is the smallest instruction-tuned model in Alibaba Cloud's Qwen 3 family, with approximately 752 million parameters. It is designed for ultra-lightweight deployment where minimal hardware resources are available, running comfortably on virtually any modern GPU or CPU-only setups. The model supports hybrid thinking mode despite its tiny footprint. While limited in reasoning depth compared to larger variants, Qwen3 0.6B handles basic chat, simple summarization, and lightweight instruction following. It is primarily useful for edge deployment, rapid prototyping, and experimentation where model size is a critical constraint. Released under the Apache 2.0 license.
Gemma 4 E2B IT Qat Q4 0 Unquantized
Google · 5.1B · runs from 2.5 GB
Gemma 4 E2B IT Qat Q4 0 Unquantized is a 5.1B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 12B IT
Google · 12.0B · runs from 5.1 GB
Sitting between Gemma 4's compact E-series and its larger 31B sibling, Gemma 4 12B IT is a 12-billion-parameter instruction-tuned model built for general chat and dialogue use. It does not take image input, focusing instead on text-based reasoning, coding help, and conversational tasks. At this parameter count, the model fits comfortably on a single consumer GPU without necessarily requiring the heaviest quantization, making it a reasonable middle-ground choice for local setups. The model supports a generous 256K token context window, useful for long documents or extended sessions, and is released under the Apache 2.0 license. It arrived in May 2026, after the initial March 2026 releases of the Gemma 4 E-series and 31B dense model, suggesting continued iteration within the family after the launch wave.
Llama 3.2 3B Instruct
Meta · 3.2B · runs from 1.0 GB
Meta Llama 3.2 3B Instruct is a 3-billion parameter instruction-tuned model from Meta's Llama 3.2 release, designed for efficient local inference on resource-constrained hardware. It supports a 128K token context window and is optimized for conversational AI, summarization, and general assistant tasks. Despite its small footprint, Llama 3.2 3B Instruct delivers competitive performance for its size class and can run on GPUs with as little as 4GB of VRAM when quantized. It is released under the Llama 3.2 Community License and is a practical choice for edge deployment and lightweight local inference.
DeepSeek R1 0528 Qwen3 8B
DeepSeek · 8.2B · runs from 2.9 GB
With about 8 billion parameters, DeepSeek R1 0528 Qwen3 8B is a dated May 2025 release in DeepSeek's R1 line, built on a Qwen3-based architecture rather than DeepSeek's own base model. It supports explicit step-by-step reasoning, generating intermediate reasoning traces before producing a final answer, which sets it apart from standard instruction-tuned chat models that answer directly. Its modest size means it runs comfortably on modest consumer GPUs, making it accessible for local experimentation with reasoning-style outputs without specialized hardware. It offers a 128K token context window, ample for most reasoning and coding tasks, and is released under the MIT license, allowing broad local and commercial use.
Qwen2.5 3B Instruct
Alibaba · 3.1B · runs from 1.4 GB
Qwen2.5 3B Instruct is a 3.1-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 family. It is designed for efficient local inference on consumer hardware, supporting a 128K token context window despite its compact footprint. The model can run on GPUs with as little as 4GB of VRAM when quantized. Despite its small size, Qwen2.5 3B Instruct delivers competitive performance for basic conversational tasks, summarization, and simple instruction following. It is a good option for edge deployment and resource-constrained environments. Released under the Apache 2.0 license.
Llama 3.2 1B Instruct
Meta · 1.2B · runs from 0.4 GB
Meta Llama 3.2 1B Instruct is a 1-billion parameter instruction-tuned model from Meta, the smallest in the Llama 3.2 family. It is designed for ultra-lightweight deployment scenarios where minimal hardware resources are available, supporting a 128K token context window despite its compact size. This model is suitable for basic conversational tasks, text summarization, and simple instruction following. It can run on virtually any modern GPU and even on CPU-only setups with acceptable performance. Released under the Llama 3.2 Community License.
Qwen2.5 1.5B Instruct
Alibaba · 1.5B · runs from 0.8 GB
Qwen2.5 1.5B Instruct is a 1.5-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It is a lightweight model suitable for deployment on minimal hardware, including low-VRAM GPUs and even CPU-only setups with acceptable latency. It supports a 128K token context window. The model handles basic conversational tasks, simple question answering, and text generation. While limited in reasoning depth compared to larger variants, it is useful for applications where fast response times and minimal resource consumption are priorities. Released under the Apache 2.0 license.
Gemma 4 E2B IT
Google · 5.1B · runs from 2.1 GB
The smallest model discussed here from Google's Gemma 4 E-series, Gemma 4 E2B IT, comes in at roughly 5 billion parameters. It is tuned for chat and instruction-following rather than multimodal tasks, prioritizing a minimal footprint over broad input support. Models at this size are comfortable on typical consumer GPUs, and can even run on integrated graphics when quantized aggressively, making it one of the more accessible options for local deployment. It ships with a 128K token context window, matching its larger E4B sibling, and is released under the Apache 2.0 license for open commercial and research use. Google published it on March 2, 2026, alongside the E4B variant, as the entry point into the Gemma 4 generation.
Qwen3.5 2B
Alibaba · 2.3B · runs from 1.0 GB
Qwen3.5 2B is a 2.3-billion-parameter model from Alibaba's Qwen team, one of the smaller entries in the Qwen3.5 lineup (0.8B–9B) built to handle text and image input together. It uses a hybrid architecture mixing linear-attention layers with periodic full-attention layers to keep long-context inference efficient. As a vision-capable model it can read and reason about images alongside written prompts, and its small size lets it run on a laptop or even a phone once quantized, without a dedicated GPU. It supports an unusually long 262K token context window for its size. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in late February 2026 as part of Alibaba's push toward compact, natively multimodal edge models.
Gemma 4 E4B IT Qat Q4 0 Unquantized
Google · 7.9B · runs from 3.9 GB
Gemma 4 E4B IT Qat Q4 0 Unquantized is a 7.9B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
LFM2.5 8B A1B
Liquid AI · 8.5B · runs from 4 GB
LFM2.5-8B-A1B is Liquid AI's Mixture-of-Experts model for on-device deployment, with 8.3 billion total parameters but only about 1.5 billion active per token, built on the hybrid LFM2 architecture that mixes gated convolution blocks with grouped-query attention. It is reasoning-tuned for on-device personal-assistant use, chaining tool calls and following complex instructions, though it is not the best fit for heavy programming or knowledge-intensive question answering without retrieval. It has a 128K token context window and is released under Liquid AI's custom LFM1.0 license. Because only 1.5 billion parameters activate per token it runs quickly, though the full 8.3 billion still need to be held in memory, roughly 5 GB at 4-bit, well within reach of almost any modern laptop or GPU.
Qwen3.5 0.8B
Alibaba · 873M · runs from 0.6 GB
Qwen3.5-0.8B is the smallest model in Alibaba's Qwen3.5 generation, a dense sub-billion-parameter vision-language model using the same hybrid Gated DeltaNet and gated-attention design as the rest of the family. It is a post-trained, instruction-tuned release, and its publisher notes that at this scale it is best suited to prototyping, task-specific fine-tuning, and other research or development purposes rather than production chat. It supports a 262,144-token native context window and is released under the Apache 2.0 license. At under one billion parameters, it needs well under a gigabyte of memory at 4-bit, so it runs on essentially any modern laptop or GPU.
LFM2.5 230M
Liquid AI · 230M · runs from 0.5 GB
LFM2.5-230M is Liquid AI's smallest model in the LFM2.5 family, a 230-million-parameter checkpoint distilled from the larger LFM2.5-350M and refined with multi-stage reinforcement learning for lightweight agentic tasks. It is instruction-tuned for data extraction and simple tool-use pipelines under the tightest memory and compute budgets, and the publisher does not recommend it for reasoning-heavy work like advanced math, code generation, or creative writing. It is released under Liquid AI's custom LFM1.0 license. At 230 million parameters, it runs on essentially any modern device, including phones and low-power boards like a Raspberry Pi, without needing a dedicated GPU at all.
LFM2.5 1.2B Instruct
Liquid AI · 1.2B · runs from 0.9 GB
LFM2.5 1.2B Instruct is an instruction-tuned model from Liquid AI that uses a novel hybrid architecture combining state-space models with attention mechanisms. At just 1.2 billion parameters, it is exceptionally lightweight and can run on virtually any hardware, including laptops and edge devices. Liquid AI's unconventional architecture aims to deliver better efficiency and longer context handling than traditional transformer models at this scale, making it an interesting option for users exploring alternatives to standard transformer-based LLMs.
Gemma 3 1B IT
Google · 1000M · runs from 0.3 GB
Google Gemma 3 1B IT is a 1-billion parameter instruction-tuned model from Google's Gemma 3 family. It is an ultra-compact text-only chat model designed for deployment on minimal hardware, including low-VRAM GPUs and edge devices. The model handles basic conversational tasks, simple instruction following, and lightweight text generation. It can run on virtually any modern GPU and even on CPU-only setups with acceptable latency. Released under the Gemma license.