All LLM Models
Browse 880 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Ministral 3 14B Reasoning 2512
Mistral AI · 13.9B · runs from 6.7 GB
Ministral 3 14B Reasoning 2512 is a 13.9B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NVIDIA Nemotron Nano 12B v2
NVIDIA · 12.3B · runs from 5.8 GB
NVIDIA Nemotron Nano 12B v2 is a 12.3B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 VL 2B Instruct
Alibaba · 2.1B · runs from 1.1 GB
Qwen3 VL 2B Instruct is Alibaba's smallest model in the Qwen3-VL lineup, a 2.1-billion-parameter vision-language model built to handle images and text together. It performs image captioning, visual question answering, and document reading, and its visual-agent tuning lets it interpret GUI screenshots for simple automation. Its small size suits on-device and edge deployment, running comfortably on modest consumer GPUs or even some laptops once quantized. The model supports a 262,144 token context window, enough for lengthy documents or multi-turn conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in October 2025 alongside the 8B and 32B models, it targets edge and mobile use cases, trading some visual reasoning depth for a much smaller footprint.
MN 12B Mag Mell R1
inflatebot · 12.2B · runs from 4.1 GB
MN 12B Mag Mell R1 is a 12.2B-parameter open language model from inflatebot. It supports a context window of up to 1,024,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Deepseek Coder 6.7B Instruct
DeepSeek · 6.7B · runs from 4.2 GB
DeepSeek Coder 6.7B Instruct is a first-generation code-specialized model trained on a large corpus of source code and programming-related data. At 6.7 billion parameters, it provides solid code completion, generation, and explanation capabilities across popular programming languages while remaining small enough to run on most consumer GPUs. While newer models in the DeepSeek lineup have surpassed it in raw capability, this model remains a practical choice for users who need a lightweight local coding assistant with minimal hardware requirements. It runs well on GPUs with as little as 6 GB of VRAM when quantized.
Llama 2 7B Chat HF
Meta · 6.7B · runs from 3.1 GB
Meta Llama 2 7B Chat is a 7-billion parameter instruction-tuned model from Meta's Llama 2 family, optimized for dialogue use cases. It was fine-tuned using supervised fine-tuning and RLHF on top of the Llama 2 7B base model, with a 4K token context window. This model is suitable for basic conversational AI tasks and runs efficiently on consumer GPUs. While newer Llama generations offer improved performance, Llama 2 7B Chat remains a well-understood and widely-supported option for local inference. Released under the Llama 2 Community License.
Hy MT2 7B
Tencent · 8.0B · runs from 2.8 GB
Hy-MT2-7B is Tencent's mid-size entry in the Hy-MT2 family of dedicated, "fast-thinking" multilingual translation models, alongside 1.8B and 30B-A3B (mixture-of-experts) siblings. Rather than being a general-purpose chat assistant, it is trained specifically for translation across 33 languages and to follow multilingual translation instructions, and the card reports it outperforming open models such as DeepSeek-V4-Pro and Kimi K2.6 in fast-thinking mode. Tencent also ships FP8 and GGUF quantized builds for on-device use with llama.cpp. At roughly 8 billion dense parameters, it fits on a single consumer GPU once quantized. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2026, alongside the open-sourced IFMTBench instruction-following translation benchmark.
SmolLM3 3B
Hugging Face · 3.1B · runs from 1.3 GB
SmolLM3 3B is Hugging Face's latest-generation compact language model, representing a significant step up from the SmolLM2 series. At 3 billion parameters, it delivers considerably stronger reasoning, instruction following, and general language understanding while maintaining modest hardware requirements that keep it accessible on most consumer GPUs. This model benefits from improved training data, architectural refinements, and lessons learned from previous SmolLM generations. It is well positioned for local chatbot applications, coding assistance, and content generation tasks where you want strong performance without dedicating the resources required by 7B-class models.
DeepSeek R1 Distill Qwen 14B
DeepSeek · 14.8B · runs from 5.1 GB
DeepSeek R1 Distill Qwen 14B sits in a sweet spot between the smaller 7B distill and the more demanding 32B version, offering strong reasoning performance at 14.8 billion parameters on the Qwen 2.5 architecture. It captures a meaningful share of the full R1's chain-of-thought capabilities while keeping resource requirements within the range of mainstream consumer GPUs. Quantized to 4-bit, it fits comfortably on GPUs with 12 GB of VRAM, delivering reliable step-by-step reasoning for math, logic, and analytical problems.
MiniCPM O 4 5
OpenBMB · 9.4B · runs from 4.6 GB
MiniCPM-o 4.5 is OpenBMB's roughly 9.4-billion-parameter omni-modal model, successor to MiniCPM-o 2.6, built from a SigLip2 vision encoder, a Whisper-medium audio encoder, a CosyVoice2 speech decoder, and a Qwen3-8B language backbone. It processes images, video, and audio and can generate spoken output, adding full-duplex live streaming so it can see, listen, and speak at once, plus a proactive mode that lets it comment unprompted, and bundled instruct and thinking modes. At this size it needs a high-end consumer GPU once quantized. Context length is 40,960 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in February 2026. Its defining upgrade over 2.6 is full-duplex interaction: earlier omni models process input and generate output in alternating turns, while 4.5 streams both directions at once without blocking.
Qwen2.5 Omni 7B
Alibaba · 10.7B · runs from 3.3 GB
Qwen2.5-Omni-7B is Alibaba's 10.7-billion-parameter flagship in the Qwen2.5-Omni family, an end-to-end model that perceives text, images, audio, and video and generates both text and natural streaming speech in response. It shares the family's Thinker-Talker design and TMRoPE positional scheme for synchronizing audio and video timestamps, tuned for low-latency, real-time conversation rather than turn-based chat alone. At just under 11 billion parameters, it needs a single mainstream-to-high-end consumer GPU once quantized. Its language backbone carries a 32,768 token context window. It is released under Alibaba's Qwen Research License, a non-commercial license limited to research and evaluation rather than the Apache 2.0 used for Qwen's text-only models. Published in March 2025, it was the first Omni release, with the smaller 3B variant following about a month later.
MiMo V2.6 Distill Qwen 9B
Xiaomi · 9.4B · runs from 3.4 GB
MiMo V2.6 Distill Qwen 9B is Xiaomi's 9.4-billion-parameter agentic model, built by fine-tuning Alibaba's Qwen3.5-9B on data generated by Xiaomi's larger MiMo-V2.6-Pro and MiMo-V2.6-Flash models. It handles vision, code, and function-calling tasks, and is positioned as the practical local-hardware option next to Xiaomi's much larger frontier MiMo-V2.6 models, which are designed for hosted use. At under 10 billion parameters, it runs well on a single mainstream consumer GPU once quantized. The model offers a 262K token context window, suited to long agent traces or multi-file coding sessions. Published in September 2026, it is one of the newest releases in Xiaomi's MiMo line, distilling the larger models' agentic and coding behaviour into a size that fits ordinary hardware.
Nanbeige4.1 3B
Nanbeige · 3.9B · runs from 2.1 GB
Nanbeige4.1 3B is a compact chat model from Nanbeige, a Chinese AI startup focused on building efficient small-scale language models. At just under 4 billion parameters, it is designed to run on virtually any modern GPU or even on CPU, making it one of the more accessible options for users with limited hardware. Despite its small size, it handles basic conversation, simple reasoning, and Chinese-English bilingual tasks, serving as a practical entry point for local LLM experimentation.
DeepSeek R1 Distill Llama 8B
DeepSeek · 8.0B · runs from 2.8 GB
DeepSeek R1 Distill Llama 8B brings R1's reinforcement-learned reasoning capabilities to the widely supported Llama 3.1 8B architecture. By distilling the full 684.5B R1 model's reasoning patterns into this 8 billion parameter dense model, DeepSeek created a version that benefits from the extensive Llama ecosystem of tools, quantizations, and inference engines. For users who prefer the Llama architecture or already have tooling built around it, this model offers a plug-and-play path to chain-of-thought reasoning. Its hardware requirements are very approachable, running well on consumer GPUs with 8 GB or more of VRAM at common quantization levels.
Mistral 7B Instruct v0.2
Mistral AI · 7.2B · runs from 3.6 GB
Mistral 7B Instruct v0.2 is a 7.2-billion-parameter instruction-tuned language model from Mistral AI, built for general chat, question answering, and instruction following. It refines the original Mistral 7B with better adherence to complex prompts and improved handling of longer inputs. Its compact size suits local deployment on modern consumer GPUs, running comfortably on mainstream hardware once quantized. The model supports a 32,768 token context window, a result of raising the RoPE base frequency used for positional encoding, which improved handling of longer inputs versus the original v0.1. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in December 2023, it became one of the most widely adopted 7B open-weight chat models, still supported by tools such as llama.cpp, vLLM, and Ollama.
MiniCPM V 4 5
OpenBMB · 8.7B · runs from 4.3 GB
MiniCPM-V 4.5 is an 8.7-billion-parameter vision-language model from OpenBMB, combining a Qwen3-8B language backbone with a SigLIP2 vision encoder. It is designed to run efficiently on modest hardware, including phones and laptops, while handling image, multi-image, and video understanding alongside text chat; a unified resampler compresses video frames heavily so longer clips don't demand proportionally more compute. At under 9 billion parameters, it fits comfortably on a single mainstream consumer GPU once quantized. It supports a 40,960 token context window, sufficient for long documents or multi-turn multimodal conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in August 2025, it distinguishes itself with a switchable "fast" and "deep" thinking mode, trading response speed for more deliberate reasoning on harder problems.
Gemma 3n E2B IT
Google · 5.4B · runs from 1.6 GB
Gemma 3n E2B is the smaller of Google's two Gemma 3n instruction-tuned models, designed for phones, laptops and other on-device use. "E2B" stands for an effective size of about 2 billion parameters: the raw checkpoint holds around 5.4 billion, but techniques such as per-layer embeddings let it run with a memory footprint closer to a 2B model. It accepts text, image and audio input and generates text. The model has a 32K-token context window and is released under Google's Gemma terms of use. At 4-bit quantization it needs only a few gigabytes of memory, so it runs on almost any modern GPU, laptop or recent phone.
Qwen2.5 Math 7B Instruct
Alibaba · 7.6B · runs from 3.0 GB
Qwen2.5 Math 7B Instruct is a 7.6B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 Coder 7B Instruct Abliterated
huihui-ai · 7.6B · runs from 3.0 GB
Qwen2.5 Coder 7B Instruct Abliterated is a 7.6B-parameter open language model from huihui-ai in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mistral 7B Instruct v0.1
Mistral AI · 7.2B · runs from 3.6 GB
Mistral 7B Instruct v0.1 was the first instruction-tuned variant of the original Mistral 7B, fine-tuned for conversational and instruction-following tasks. While it has since been superseded by v0.2 and v0.3, it remains a solid lightweight chat model and an important milestone in the open-weight model ecosystem. Its hardware requirements are identical to the base Mistral 7B, running smoothly on GPUs with as little as 6 GB of VRAM when quantized. Users seeking the best Mistral 7B experience should generally prefer the newer v0.3 release, but v0.1 is still useful for reproducibility and benchmarking purposes.
Gemma 4 12B IT Heretic
igorls · 12.0B · runs from 6.1 GB
Gemma 4 12B IT Heretic is a 12.0B-parameter open language model from igorls in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 E2B IT Qat Mobile Transformers
Google · 2.3B · runs from 1.4 GB
Gemma 4 E2B IT Qat Mobile Transformers is a 2.3B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 2 2B IT Abliterated
IlyaGusev · 2.6B · runs from 1.6 GB
Gemma 2 2B IT Abliterated is a 2.6B-parameter open language model from IlyaGusev in the Gemma 2 family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 3n E4B IT
Google · 7.8B · runs from 2.4 GB
Gemma 3n E4B IT is the instruction-tuned variant of Google's Gemma 3n E4B, a multimodal model built for on-device use on phones, laptops, and tablets that accepts text, image, audio, and video input and can perform automatic speech recognition and speech translation alongside text chat. It uses Google's MatFormer (Matryoshka Transformer) architecture, which nests a smaller sub-model inside the full network so the same checkpoint can run at reduced effective capacity, and Per-Layer Embedding, which caches embedding parameters to fast local storage instead of holding them all in memory. The "E4B" designation refers to an effective parameter count of around 4 billion at inference, even though the checkpoint's total parameters are larger; either way it is light enough for a single consumer GPU or even a high-end phone. Context length is 32,768 tokens. It is released under Google's Gemma Terms of Use, a custom license permitting broad commercial and research use alongside a prohibited-use policy, and was published in June 2025.
NeoHorse 1 4B
TokenRhythm · 4.2B · runs from 2.3 GB
NeoHorse 1 4B is a 4.2B-parameter open language model from TokenRhythm. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 E2B IT Uncensored
TrevorJS · 5.1B · runs from 2.5 GB
Gemma 4 E2B IT Uncensored is a 5.1B-parameter open language model from TrevorJS in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2 VL 2B Instruct
Alibaba · 2.2B · runs from 1.1 GB
Qwen2 VL 2B Instruct is a 2.2-billion-parameter vision-language model from Alibaba's Qwen2-VL series, able to process images, multi-image comparisons, and video alongside text prompts. It targets visual question answering, document and chart reading, and basic agentic tasks such as interpreting a screenshot to plan a next action. Its small size makes it well suited to laptops and even some phones, running comfortably on modest consumer hardware once quantized. The model supports a 32K token context window, enough for moderate documents or extended chat. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in August 2024. Qwen2-VL introduced Naive Dynamic Resolution and Multimodal Rotary Position Embedding, letting it handle arbitrary image resolutions and understand videos well over twenty minutes long.
Qwen2.5 Coder 0.5B Instruct
Alibaba · 494M · runs from 0.5 GB
Qwen2.5 Coder 0.5B Instruct is a 494M-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 Math 1.5B Instruct
Alibaba · 1.5B · runs from 1 GB
Qwen2.5 Math 1.5B Instruct is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GLM 4.7 Flash REAP 23B A3B
Cerebras · 23.0B · runs from 7.4 GB
GLM 4.7 Flash REAP 23B A3B is a 23.0B-parameter open language model from Cerebras in the GLM 4 family. It supports a context window of up to 202,752 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.