All LLM Models
Browse 880 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Ornith 1.0 9B
Deep Reinforce · 9.4B · runs from 3.7 GB
Ornith 1.0 9B is Deep Reinforce's smaller, dense sibling to Ornith 1.0 35B, with roughly 9.4 billion parameters and no mixture-of-experts routing. It is tuned for general chat and instruction-following use. Models in this parameter range run comfortably on modest consumer GPUs, and remain usable on lower-VRAM setups at higher quantization, making it a practical option for local deployment without high-end hardware. The model supports a 256K token context window, matching its larger sibling despite the smaller parameter count. It is released under the MIT license and was published on June 21, 2026, giving Deep Reinforce two size options in the Ornith 1.0 line for different hardware budgets.
Qwen3 VL 4B Instruct
Alibaba · 4.4B · runs from 1.7 GB
Qwen3 VL 4B Instruct is Alibaba's 4.4-billion-parameter vision-language model from the Qwen3-VL family, built to read and reason over images and video alongside text in one conversation. It suits multimodal chat, document and chart understanding, and simple visual-agent tasks like describing screenshots or pulling structured data out of pictures. At this size the model runs comfortably on a mainstream consumer GPU once quantized, and is light enough for many laptops. The model supports a 262K token context window, giving room for long documents or extended visual conversations. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2025 as part of the Qwen3-VL generation, built for deeper visual perception and stronger agentic behavior than the earlier Qwen2-VL series.
Spark X2.5 4B
XHToken · 4.1B · runs from 2.2 GB
Spark X2.5 4B is a 4.1B-parameter open language model from XHToken. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 1.7B
Alibaba · 2.0B · runs from 1.1 GB
Qwen3 1.7B is a 1.7-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 series. It is a lightweight model designed for deployment on minimal hardware, including low-VRAM GPUs and even CPU-only configurations with acceptable latency. Despite its compact size, it supports hybrid thinking mode and handles basic conversational tasks, simple question answering, and text generation. The model is useful for edge deployment, embedded applications, and scenarios where fast inference with minimal resource consumption is the priority. It represents a significant quality improvement over Qwen 2.5 at the sub-2B scale. Released under the Apache 2.0 license.
Ternary Bonsai 8B Unpacked
prism-ml · 8.2B · runs from 4.1 GB
Ternary Bonsai 8B Unpacked is a 8.2B-parameter open language model from prism-ml. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Hy MT2 1.8B
Tencent · 2.0B · runs from 1.1 GB
Hy MT2 1.8B is a 2-billion-parameter translation model from Tencent's Hunyuan team, part of a three-size Hy-MT2 family (1.8B, 7B, and a 30B mixture-of-experts variant) built specifically for machine translation rather than general chat. It supports translation across roughly three dozen languages and follows translation-specific instructions, such as adjusting tone or terminology. It runs on modest consumer hardware, including laptops, and Tencent has also released extreme-quantized versions shrunk to a few hundred megabytes for on-device use. It supports a 262K token context window, letting it translate long documents in a single pass. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in May 2026, it is the smallest member of a translation-focused family Tencent positions as a self-hostable alternative to commercial translation APIs.
Granite 4.2 8B
IBM · 8.8B · runs from 3.0 GB
Granite 4.2 8B is a 8.8B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mistral 7B v0.3
Mistral AI · 7.2B · runs from 3.6 GB
Mistral 7B v0.3 is a 7.2B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 VL 8B Instruct
Alibaba · 8.8B · runs from 3.0 GB
Qwen3 VL 8B Instruct is Alibaba's 8.8-billion-parameter vision-language model in the Qwen3-VL lineup, built to process images and text together in one conversation. Beyond image description and visual question answering, it is tuned as a visual agent that can read GUI screenshots and reason about on-screen elements, useful for document analysis and early automation tasks. At this size, local inference is practical on a single mainstream-to-high-end consumer GPU once quantized. The model supports a 262,144 token context window, enough for long documents or extended chat histories. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. Published in October 2025 alongside 2B and 32B siblings, Qwen3-VL adds video understanding with fine-grained event indexing beyond earlier Qwen vision-language releases.
MiniCPM5 1B Claude Opus Fable5 v2 Thinking
GnLOLot · 1.1B · runs from 0.8 GB
MiniCPM5 1B Claude Opus Fable5 v2 Thinking is a 1.1B-parameter open language model from GnLOLot in the MiniCPM family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 2 9B IT
Google · 9.2B · runs from 3.0 GB
Google Gemma 2 9B IT is a 9.2-billion parameter instruction-tuned model from Google's Gemma 2 series. It is a text-only chat model optimized for conversational tasks, instruction following, and general-purpose assistance. At release, it was recognized for delivering unusually strong performance relative to its parameter count. The model runs efficiently on consumer GPUs with 8-12GB of VRAM in quantized formats, making it accessible on mainstream hardware. It is a popular choice for local inference among users who want strong quality without the VRAM demands of larger models. Released under the Gemma license.
Phi 4 Mini Instruct
Microsoft · 3.8B · runs from 1.9 GB
Microsoft Phi 4 Mini Instruct is a 3.8-billion parameter instruction-tuned model from Microsoft Research's Phi 4 family. It applies the Phi series' data-centric training philosophy to a compact model, delivering strong performance in coding, reasoning, and chat tasks relative to its small footprint. The model runs on consumer GPUs with as little as 4-6GB of VRAM when quantized, making it accessible on mainstream and even entry-level hardware. Released under the MIT license.
Mistral 7B Instruct v0.3
Mistral AI · 7.2B · runs from 2.7 GB
Mistral 7B Instruct v0.3 is the latest instruction-tuned release of Mistral AI's original 7-billion-parameter model, delivering meaningful improvements in instruction following, function calling, and multilingual support over its predecessors. With an extended 32K-token vocabulary and refined chat capabilities, v0.3 remains one of the most capable sub-10B models available. At 7.2 billion parameters it sits comfortably in the sweet spot for local inference, running well on GPUs with 6–8 GB of VRAM at full precision and even on 4 GB cards with 4-bit quantization. It is an excellent default choice for anyone getting started with local LLMs who wants strong conversational performance without heavy hardware.
Qwen2.5 14B
Alibaba · 14.8B · runs from 6.8 GB
Qwen2.5 14B is a 14.8B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Jan v3 4B Base Instruct
janhq · 4.4B · runs from 2.4 GB
Jan v3 4B Base Instruct is a 4.4B-parameter open language model from janhq. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 0.5B Instruct
Alibaba · 494M · runs from 0.5 GB
Qwen2.5 0.5B Instruct is the smallest instruction-tuned model in Alibaba Cloud's Qwen 2.5 family, with just 494 million parameters. It is designed for ultra-lightweight deployment scenarios where minimal hardware resources are available, running comfortably on virtually any modern GPU or even CPU-only configurations. Despite its tiny footprint, the model supports a 128K token context window and can handle basic chat, simple summarization, and lightweight instruction following. It is primarily useful for edge deployment, experimentation, and prototyping where model size is a critical constraint. Released under the Apache 2.0 license.
Gemma 2 2B IT
Google · 2.6B · runs from 0.9 GB
Google Gemma 2 2B IT is a 2-billion parameter instruction-tuned model from Google's Gemma 2 family, the smallest variant in the Gemma 2 series. It is designed for efficient local inference on resource-constrained hardware, handling basic conversational tasks and simple instruction following at minimal compute cost. The model can run on GPUs with as little as 4GB of VRAM when quantized, and even on CPU-only setups. Released under the Gemma license.
Agents A1 4B
InternScience · 4.5B · runs from 2 GB
Agents A1 4B is a 4.5B-parameter open language model from InternScience in the Agents-A1 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mellum2 12B A2.5B Instruct
JetBrains · 12.1B · runs from 4.0 GB
Mellum2 12B A2.5B Instruct is JetBrains' 12-billion-parameter mixture-of-experts coding model, with about 2.4 billion parameters active per token across 64 experts. Because only a fraction of the weights are computed for each token, it runs noticeably faster than a dense model of similar size, while all the experts still need to fit in memory. Unlike JetBrains' earlier Mellum, a narrow fill-in-the-middle completion model, Mellum2 is a full coding assistant that can generate and edit code, call tools, hold multi-turn conversations, and reason through problems. Its modest active-parameter count makes it practical to run on a single mainstream consumer GPU once quantized. The model supports a 131K token context window and is released under the Apache 2.0 license, allowing unrestricted commercial and research use. It was published in May 2026.
Ling 3.0 Tiny
Inclusion AI · 7.9B · runs from 2.8 GB
Ling 3.0 Tiny is InclusionAI's smallest mixture-of-experts model in Ant Group's Ling 3.0 family, totaling roughly 7.9 billion parameters with about 1.4 billion active per token. It shares the family's hybrid-linear-attention MoE architecture, also used at far larger scale in Ling-3.0-Flash, scaled down here for lightweight chat and agent use. Only the active parameters compute per token, so it responds quickly while still needing its full parameter set in memory; under 8 billion total, it runs on a mainstream consumer GPU once quantized. The model provides a 131K token context window for extended conversations or agent traces. It is released under the MIT license, allowing unrestricted commercial and research use. Published in August 2026, it extends Ant Group's Ling 3.0 lineup, built for efficient agent workloads.
Yi Coder 9B Chat
01.AI · 8.8B · runs from 3.1 GB
Yi Coder 9B Chat is a 8.8B-parameter open language model from 01.AI in the Yi family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3.5 Mini Instruct
Microsoft · 3.8B · runs from 2.3 GB
Phi 3.5 Mini Instruct is a 3.8B-parameter open language model from Microsoft in the Phi 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Meta Llama 3 8B Instruct
Meta · 8.0B · runs from 2.6 GB
Meta Llama 3 8B Instruct is the instruction-tuned version of Meta's Llama 3 8B base model, with 8 billion parameters. It is fine-tuned for dialogue and chat use cases using supervised fine-tuning and RLHF, making it ready for conversational applications out of the box. The model supports an 8K token context window and performs well across coding, reasoning, and general knowledge tasks. Its efficient size makes it one of the most popular models for local inference on consumer hardware. Released under the Meta Llama 3 Community License.
S1 Mini
superwhisper · 752M · runs from 0.7 GB
S1 Mini is a 752M-parameter open language model from superwhisper. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NVIDIA Nemotron Nano 9B v2
NVIDIA · 8.9B · runs from 3.5 GB
NVIDIA Nemotron Nano 9B v2 is a compact yet capable chat model from NVIDIA, packing 8.9 billion parameters into a size that runs comfortably on a wide range of consumer GPUs. Built on NVIDIA's Nemotron architecture, it delivers strong instruction-following and conversational performance while keeping VRAM requirements modest. This second-generation Nano model reflects NVIDIA's push to make high-quality language models accessible on local hardware. It's an excellent starting point for users who want a responsive, general-purpose assistant without needing top-tier GPU memory.
Qwen3 4B Thinking 2507
Alibaba · 4.0B · runs from 2.2 GB
Qwen3 4B Thinking 2507 is the reasoning-optimized variant of Alibaba's compact 4-billion-parameter Qwen3 model, released in the July 2025 update cycle. Despite its small size, this thinking variant is tuned to produce chain-of-thought reasoning and step-by-step problem solving, making it a surprisingly capable lightweight reasoner. This model is ideal for users who want basic reasoning and analytical capabilities on very modest hardware. It can run on most consumer GPUs and even some CPU-only setups when quantized, providing an accessible entry point for experimenting with reasoning-style models without any significant hardware investment.
LFM2.5 2.6B DSpark
Liquid AI · 328M · runs from 0.5 GB
LFM2.5 2.6B DSpark is a 328M-parameter open language model from Liquid AI in the LFM2.5 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 3 4B IT
Google · 4.3B · runs from 2.0 GB
Gemma 3 4B IT is a small, roughly 4.3-billion-parameter model from Google's earlier Gemma 3 generation, tuned for instruction-following and chat. It accepts image input alongside text, so it can handle visual question answering and image-grounded prompts in addition to text-only conversation. At this size, the model is comfortable on modest consumer hardware, running even on integrated graphics at low quantization, making it one of the most accessible vision-capable models for local use. Released on February 20, 2025, Gemma 3 4B IT predates the Gemma 4 family and is distributed under Google's Gemma license terms rather than a standard open-source license. Its compact size and image-input support make it a practical choice for lightweight, on-device multimodal applications.
Phi 4
Microsoft · 14.7B · runs from 7.0 GB
Microsoft Phi 4 is a 14-billion parameter language model from Microsoft Research's Phi series, designed to deliver strong reasoning, mathematical, and coding performance at an efficient size. Phi 4 continues the Phi family's focus on maximizing capability per parameter through high-quality training data curation, achieving benchmark scores that rival much larger models on reasoning and STEM tasks. The model runs well on consumer GPUs with 12-16GB of VRAM in quantized formats. It excels at mathematical problem solving, code generation, and structured reasoning. Released under the MIT license.
Mistral Nemo Instruct 2407
Mistral AI · 12.2B · runs from 5.9 GB
Mistral Nemo Instruct 2407 is a 12-billion-parameter instruction-tuned chat model from Mistral AI, a dated release from July 2024. It targets general dialogue and instruction-following use cases, sitting in a practical middle ground between lightweight and large-scale models in terms of both capability and resource demands. The model offers a 128K token context window, generous for its parameter class, and is released under the Apache 2.0 license, allowing unrestricted local and commercial use. At around 12 billion parameters, Mistral Nemo Instruct 2407 fits on a single consumer GPU when quantized, putting it within easy reach of local enthusiasts running their own hardware.