All LLM Models
Browse 982 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Llama 3.2 1B Instruct
Meta · 1.2B · runs from 0.4 GB
Meta Llama 3.2 1B Instruct is a 1-billion parameter instruction-tuned model from Meta, the smallest in the Llama 3.2 family. It is designed for ultra-lightweight deployment scenarios where minimal hardware resources are available, supporting a 128K token context window despite its compact size. This model is suitable for basic conversational tasks, text summarization, and simple instruction following. It can run on virtually any modern GPU and even on CPU-only setups with acceptable performance. Released under the Llama 3.2 Community License.
Qwen3 30B A3B
Alibaba · 30.5B · runs from 8.8 GB
Qwen3 30B A3B is a Mixture of Experts (MoE) model from Alibaba Cloud's Qwen 3 series, with 30 billion total parameters and approximately 3 billion active parameters per forward pass. The MoE architecture delivers quality significantly above what a standard 3B dense model could achieve, while keeping per-token compute costs low. It supports hybrid thinking mode for flexible reasoning. The model requires VRAM proportional to its full 30B parameter count for weight loading, but its low active parameter count results in fast inference throughput. It is an efficient option for users who want quality beyond dense small models without the full cost of larger architectures. Released under the Apache 2.0 license.
Qwen2.5 1.5B Instruct
Alibaba · 1.5B · runs from 0.8 GB
Qwen2.5 1.5B Instruct is a 1.5-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It is a lightweight model suitable for deployment on minimal hardware, including low-VRAM GPUs and even CPU-only setups with acceptable latency. It supports a 128K token context window. The model handles basic conversational tasks, simple question answering, and text generation. While limited in reasoning depth compared to larger variants, it is useful for applications where fast response times and minimal resource consumption are priorities. Released under the Apache 2.0 license.
Gemma 4 E2B IT
Google · 5.1B · runs from 2.1 GB
The smallest model discussed here from Google's Gemma 4 E-series, Gemma 4 E2B IT, comes in at roughly 5 billion parameters. It is tuned for chat and instruction-following rather than multimodal tasks, prioritizing a minimal footprint over broad input support. Models at this size are comfortable on typical consumer GPUs, and can even run on integrated graphics when quantized aggressively, making it one of the more accessible options for local deployment. It ships with a 128K token context window, matching its larger E4B sibling, and is released under the Apache 2.0 license for open commercial and research use. Google published it on March 2, 2026, alongside the E4B variant, as the entry point into the Gemma 4 generation.
GLM 4.7 Flash
Z.ai · 31.2B · runs from 9.7 GB
GLM-4.7-Flash is a Mixture-of-Experts language model from Z.ai, part of the GLM 4 family, with about 31 billion total parameters and roughly 3 billion active per token. It is a chat and agentic model, positioned by Z.ai as a lightweight option for coding, tool use, and multi-turn agent tasks, with configurable reasoning-effort settings and support for extended thinking on harder problems. It supports a 202,752-token context window and is released under the MIT license. At this size, 4-bit quantization needs roughly 18GB of memory, putting it within reach of a single high-end consumer GPU in the 16-24GB range.
Qwen3.5 2B
Alibaba · 2.3B · runs from 1.0 GB
Qwen3.5 2B is a 2.3-billion-parameter model from Alibaba's Qwen team, one of the smaller entries in the Qwen3.5 lineup (0.8B–9B) built to handle text and image input together. It uses a hybrid architecture mixing linear-attention layers with periodic full-attention layers to keep long-context inference efficient. As a vision-capable model it can read and reason about images alongside written prompts, and its small size lets it run on a laptop or even a phone once quantized, without a dedicated GPU. It supports an unusually long 262K token context window for its size. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in late February 2026 as part of Alibaba's push toward compact, natively multimodal edge models.
Gemma 4 E4B IT Qat Q4 0 Unquantized
Google · 7.9B · runs from 3.9 GB
Gemma 4 E4B IT Qat Q4 0 Unquantized is a 7.9B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
LFM2.5 8B A1B
Liquid AI · 8.5B · runs from 4 GB
LFM2.5-8B-A1B is Liquid AI's Mixture-of-Experts model for on-device deployment, with 8.3 billion total parameters but only about 1.5 billion active per token, built on the hybrid LFM2 architecture that mixes gated convolution blocks with grouped-query attention. It is reasoning-tuned for on-device personal-assistant use, chaining tool calls and following complex instructions, though it is not the best fit for heavy programming or knowledge-intensive question answering without retrieval. It has a 128K token context window and is released under Liquid AI's custom LFM1.0 license. Because only 1.5 billion parameters activate per token it runs quickly, though the full 8.3 billion still need to be held in memory, roughly 5 GB at 4-bit, well within reach of almost any modern laptop or GPU.
Qwen3.5 0.8B
Alibaba · 873M · runs from 0.6 GB
Qwen3.5-0.8B is the smallest model in Alibaba's Qwen3.5 generation, a dense sub-billion-parameter vision-language model using the same hybrid Gated DeltaNet and gated-attention design as the rest of the family. It is a post-trained, instruction-tuned release, and its publisher notes that at this scale it is best suited to prototyping, task-specific fine-tuning, and other research or development purposes rather than production chat. It supports a 262,144-token native context window and is released under the Apache 2.0 license. At under one billion parameters, it needs well under a gigabyte of memory at 4-bit, so it runs on essentially any modern laptop or GPU.
LFM2.5 230M
Liquid AI · 230M · runs from 0.5 GB
LFM2.5-230M is Liquid AI's smallest model in the LFM2.5 family, a 230-million-parameter checkpoint distilled from the larger LFM2.5-350M and refined with multi-stage reinforcement learning for lightweight agentic tasks. It is instruction-tuned for data extraction and simple tool-use pipelines under the tightest memory and compute budgets, and the publisher does not recommend it for reasoning-heavy work like advanced math, code generation, or creative writing. It is released under Liquid AI's custom LFM1.0 license. At 230 million parameters, it runs on essentially any modern device, including phones and low-power boards like a Raspberry Pi, without needing a dedicated GPU at all.
LFM2.5 1.2B Instruct
Liquid AI · 1.2B · runs from 0.9 GB
LFM2.5 1.2B Instruct is an instruction-tuned model from Liquid AI that uses a novel hybrid architecture combining state-space models with attention mechanisms. At just 1.2 billion parameters, it is exceptionally lightweight and can run on virtually any hardware, including laptops and edge devices. Liquid AI's unconventional architecture aims to deliver better efficiency and longer context handling than traditional transformer models at this scale, making it an interesting option for users exploring alternatives to standard transformer-based LLMs.
Gemma 3 1B IT
Google · 1000M · runs from 0.3 GB
Google Gemma 3 1B IT is a 1-billion parameter instruction-tuned model from Google's Gemma 3 family. It is an ultra-compact text-only chat model designed for deployment on minimal hardware, including low-VRAM GPUs and edge devices. The model handles basic conversational tasks, simple instruction following, and lightweight text generation. It can run on virtually any modern GPU and even on CPU-only setups with acceptable latency. Released under the Gemma license.
Ornith 1.0 9B
Deep Reinforce · 9.4B · runs from 3.7 GB
Ornith 1.0 9B is Deep Reinforce's smaller, dense sibling to Ornith 1.0 35B, with roughly 9.4 billion parameters and no mixture-of-experts routing. It is tuned for general chat and instruction-following use. Models in this parameter range run comfortably on modest consumer GPUs, and remain usable on lower-VRAM setups at higher quantization, making it a practical option for local deployment without high-end hardware. The model supports a 256K token context window, matching its larger sibling despite the smaller parameter count. It is released under the MIT license and was published on June 21, 2026, giving Deep Reinforce two size options in the Ornith 1.0 line for different hardware budgets.
Qwen3 VL 4B Instruct
Alibaba · 4.4B · runs from 1.7 GB
Qwen3 VL 4B Instruct is Alibaba's 4.4-billion-parameter vision-language model from the Qwen3-VL family, built to read and reason over images and video alongside text in one conversation. It suits multimodal chat, document and chart understanding, and simple visual-agent tasks like describing screenshots or pulling structured data out of pictures. At this size the model runs comfortably on a mainstream consumer GPU once quantized, and is light enough for many laptops. The model supports a 262K token context window, giving room for long documents or extended visual conversations. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2025 as part of the Qwen3-VL generation, built for deeper visual perception and stronger agentic behavior than the earlier Qwen2-VL series.
Qwen3.5 27B
Alibaba · 27.8B · runs from 8.4 GB
Qwen3.5 27B is a 27.8-billion-parameter dense model from Alibaba's Qwen team, built for both text and image input using the same hybrid linear/full-attention architecture as the rest of the Qwen3.5 family. As a vision-capable model it can describe and reason about images alongside prompts, suiting multimodal chat and document tasks. At this size, local inference calls for quantization and a single high-end 24GB-plus consumer or workstation GPU rather than budget hardware. It supports a 262K token context window for long documents and multi-turn conversations, and is released under the Apache 2.0 license for unrestricted commercial and research use. Published in late February 2026, it sits alongside a separate small-model tier (0.8B–9B) in the same family, giving developers a mid-sized, single-GPU-friendly dense option.
Spark X2.5 4B
XHToken · 4.1B · runs from 2.2 GB
Spark X2.5 4B is a 4.1B-parameter open language model from XHToken. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 1.7B
Alibaba · 2.0B · runs from 1.1 GB
Qwen3 1.7B is a 1.7-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 series. It is a lightweight model designed for deployment on minimal hardware, including low-VRAM GPUs and even CPU-only configurations with acceptable latency. Despite its compact size, it supports hybrid thinking mode and handles basic conversational tasks, simple question answering, and text generation. The model is useful for edge deployment, embedded applications, and scenarios where fast inference with minimal resource consumption is the priority. It represents a significant quality improvement over Qwen 2.5 at the sub-2B scale. Released under the Apache 2.0 license.
Mistral Small 24B Instruct 2501
Mistral AI · 23.6B · runs from 10.7 GB
Mistral Small 24B Instruct is Mistral AI's January 2025 release targeting the mid-range parameter sweet spot. At 24 billion parameters it sits between lightweight 7B models and heavier 70B-class offerings, delivering strong instruction-following, reasoning, and coding performance without demanding top-tier hardware. This model fits comfortably on a single GPU with 16–24 GB of VRAM at common quantization levels, making it an attractive option for users with cards like the RTX 4090 or RTX 3090 who want a noticeable step up from 7B models. It strikes an appealing balance between quality and resource requirements for serious local use.
Qwen3.6 27B
Alibaba · 27.8B · runs from 11.5 GB
Released a few months before its Qwen3.8 sibling, Qwen3.6 27B is Alibaba's dense, roughly 28-billion-parameter model from the Qwen 3.6 generation. Vision input is supported alongside text, so the model can handle image-based prompts as well as standard chat, coding, and reasoning tasks. Like other models in this size class, it is best run locally on a single high-end consumer GPU with quantization rather than lower-end hardware. Qwen3.6 27B carries a 256K token context window for handling lengthy inputs, and is released under the Apache 2.0 license, permitting free commercial and research use. It sits alongside the larger Qwen3.6 35B A3B mixture-of-experts model as one of two Qwen 3.6 options with local-deployment potential.
Qwen AgentWorld 35B A3B
Alibaba · 34.7B · runs from 9.9 GB
Qwen AgentWorld 35B A3B is Alibaba's 35-billion-parameter mixture-of-experts model, with about 3 billion parameters active per token (the A3B in its name). Unlike a typical chat model, it is a language world model: given an agent's action and history, it predicts what the environment does next, covering domains such as terminal, web, Android, and software engineering. Only the active experts run per token, so inference stays fast even though all the weights must fit in memory. At this size, local inference calls for quantization and a single high-end consumer GPU. The model supports a 262K token context window and is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in June 2026, it is built on a Qwen3.5-35B-A3B base and is meant for simulating agents, not direct conversation.
Ternary Bonsai 8B Unpacked
prism-ml · 8.2B · runs from 4.1 GB
Ternary Bonsai 8B Unpacked is a 8.2B-parameter open language model from prism-ml. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Hy MT2 1.8B
Tencent · 2.0B · runs from 1.1 GB
Hy MT2 1.8B is a 2-billion-parameter translation model from Tencent's Hunyuan team, part of a three-size Hy-MT2 family (1.8B, 7B, and a 30B mixture-of-experts variant) built specifically for machine translation rather than general chat. It supports translation across roughly three dozen languages and follows translation-specific instructions, such as adjusting tone or terminology. It runs on modest consumer hardware, including laptops, and Tencent has also released extreme-quantized versions shrunk to a few hundred megabytes for on-device use. It supports a 262K token context window, letting it translate long documents in a single pass. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in May 2026, it is the smallest member of a translation-focused family Tencent positions as a self-hostable alternative to commercial translation APIs.
Granite 4.2 8B
IBM · 8.8B · runs from 3.0 GB
Granite 4.2 8B is a 8.8B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
KAT Coder V2.5 Dev
Kwaipilot · 34.7B · runs from 9.9 GB
KAT Coder V2.5 Dev is a 34.7B-parameter open language model from Kwaipilot. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mistral 7B v0.3
Mistral AI · 7.2B · runs from 3.6 GB
Mistral 7B v0.3 is a 7.2B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 VL 8B Instruct
Alibaba · 8.8B · runs from 3.0 GB
Qwen3 VL 8B Instruct is Alibaba's 8.8-billion-parameter vision-language model in the Qwen3-VL lineup, built to process images and text together in one conversation. Beyond image description and visual question answering, it is tuned as a visual agent that can read GUI screenshots and reason about on-screen elements, useful for document analysis and early automation tasks. At this size, local inference is practical on a single mainstream-to-high-end consumer GPU once quantized. The model supports a 262,144 token context window, enough for long documents or extended chat histories. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. Published in October 2025 alongside 2B and 32B siblings, Qwen3-VL adds video understanding with fine-grained event indexing beyond earlier Qwen vision-language releases.
Ornith 1.0 35B
Deep Reinforce · 35.1B · runs from 10.0 GB
Ornith 1.0 35B, from Deep Reinforce, is a 35-billion-parameter model built on a mixture-of-experts architecture, meaning only a subset of its parameters activate for any given token even though the full model must be held in memory. Because the full expert set still needs to fit in memory much like a dense model of the same size, local use is best suited to a single high-end consumer GPU once quantized rather than lower-end hardware. It is tuned for general chat and instruction-following. The model provides a 256K token context window for handling long inputs, and is released under the MIT license, one of the more permissive options in this catalogue. It was published on June 21, 2026, the same day as its smaller sibling, Ornith 1.0 9B.
MiniCPM5 1B Claude Opus Fable5 v2 Thinking
GnLOLot · 1.1B · runs from 0.8 GB
MiniCPM5 1B Claude Opus Fable5 v2 Thinking is a 1.1B-parameter open language model from GnLOLot in the MiniCPM family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
North Mini Code 1.0
Cohere · 30.5B · runs from 8.8 GB
North Mini Code 1.0 is Cohere's first open-weight coding model, a mixture-of-experts design with about 30.5 billion total parameters and roughly 3.3 billion active per token. It is tuned for agentic coding: tool calling, terminal workflows, and software-engineering tasks rather than general chat. Only the active parameters compute per token, so it decodes quickly for its size, though all weights must fit in memory; it runs on a single high-end consumer GPU once quantized. The model supports a 500,000 token context window, useful for large codebases or long agent transcripts. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in June 2026 as the debut release in Cohere's North family, it is designed to run on a single H100-class GPU in FP8.
Gemma 2 9B IT
Google · 9.2B · runs from 3.0 GB
Google Gemma 2 9B IT is a 9.2-billion parameter instruction-tuned model from Google's Gemma 2 series. It is a text-only chat model optimized for conversational tasks, instruction following, and general-purpose assistance. At release, it was recognized for delivering unusually strong performance relative to its parameter count. The model runs efficiently on consumer GPUs with 8-12GB of VRAM in quantized formats, making it accessible on mainstream hardware. It is a popular choice for local inference among users who want strong quality without the VRAM demands of larger models. Released under the Gemma license.