All LLM Models
Browse 1138 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Qwen3 4B DFlash B16
z-lab · 537M · runs from 0.6 GB
Qwen3 4B DFlash B16 is a 537M-parameter open language model from z-lab in the Qwen 3 family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 4 Mini Reasoning
Microsoft · 3.8B · runs from 1.6 GB
Phi-4-mini-reasoning is Microsoft's compact 3.8-billion-parameter model in the Phi-4 family, fine-tuned specifically for multi-step, logic-intensive mathematical reasoning using synthetic math data distilled from a larger teacher model. It targets tasks like formal proof generation, symbolic computation, and advanced word problems, and is designed to run in memory- and latency-constrained environments such as educational tools or embedded tutoring systems, rather than for general knowledge or open-domain chat. It supports a 128K token context window and is released under the MIT license. At 3.8 billion parameters, it runs comfortably on almost any modern laptop or consumer GPU, even at higher precision, making it one of the more accessible reasoning-focused models to run locally.
TinyLlama 1.1B Chat v1.0
TinyLlama · 1.1B · runs from 0.8 GB
TinyLlama 1.1B Chat is a 1.1-billion parameter chat model built on the Llama 2 architecture and trained on approximately 3 trillion tokens, an unusually large dataset for a model of its size. The TinyLlama project demonstrated that small models can achieve strong performance when given sufficient training compute, making it a standout in the sub-2B parameter class. The Chat variant is fine-tuned for conversational use and runs on virtually any modern GPU, including entry-level cards with 4GB of VRAM or less. It is a practical choice for lightweight local inference, edge deployment, and experimentation where hardware resources are limited.
Qwen3.5 9B DFlash
z-lab · 1.3B · runs from 0.9 GB
Qwen3.5 9B DFlash is a 1.3B-parameter open language model from z-lab in the Qwen 3.5 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.5 4B DFlash
z-lab · 634M · runs from 0.6 GB
Qwen3.5 4B DFlash is a 634M-parameter open language model from z-lab in the Qwen 3.5 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 31B IT Uncensored Heretic
llmfan46 · 31.3B · runs from 14.9 GB
Gemma 4 31B IT Uncensored Heretic is a 31.3B-parameter open language model from llmfan46 in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ornith 1.5 9B
ornith-ai · 9.7B · runs from 3.2 GB
Ornith-1.5-9B is the lightest model in Ornith AI's Ornith-1.5 family, a dense 9.7-billion-parameter model built for agentic coding tasks and designed for efficient single-GPU deployment. It shares the family's self-improving training approach, in which reinforcement learning jointly optimizes the tasks, agent scaffolding, and solution rollouts used to train it, and a quantized variant is available for edge and mobile deployment. It offers a 262,144-token context window and is released under the MIT license. At roughly 9.7 billion parameters, 4-bit quantization needs only about 5.6GB of memory, so it runs comfortably on a single mid-range consumer GPU.
Diffusiongemma 26B A4B IT
Google · 25.8B · runs from 11.6 GB
Diffusiongemma 26B A4B IT takes a different approach from the rest of the Gemma 4 family: rather than generating tokens autoregressively, it uses a block-diffusion architecture. It still follows a mixture-of-experts design, with roughly 26 billion total parameters and about 4 billion active per token. The active-parameter figure is the more relevant one for inference cost, while the total parameter count determines memory needs, putting local use in single high-end consumer GPU territory once quantized. It also accepts image input alongside text for multimodal use. The model offers a 256K token context window and is released under the Apache 2.0 license. Published on June 9, 2026, it represents Google's exploration of diffusion-based language modeling within the Gemma 4 generation, alongside the more conventional autoregressive Gemma 4 variants also released that year.
Bonsai 1.7B Unpacked
prism-ml · 1.7B · runs from 1.3 GB
Bonsai 1.7B Unpacked is a 1.7B-parameter open language model from prism-ml. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
DeepSeek R1 Distill Qwen 32B
DeepSeek · 32.8B · runs from 9.8 GB
DeepSeek R1 Distill Qwen 32B takes the reasoning capabilities developed in the full 684.5B R1 model and distills them into the 32.8 billion parameter Qwen 2.5 architecture. The result is a dense model that punches well above its weight class on math, science, and coding reasoning tasks, often matching models two to three times its size. At around 32.8 billion parameters, this model fits comfortably on a single high-end consumer GPU when quantized to 4-bit precision, making it one of the most capable reasoning models you can run on a desktop workstation.
Ministral 3 14B Reasoning 2512
Mistral AI · 13.9B · runs from 6.7 GB
Ministral 3 14B Reasoning 2512 is a 13.9B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NVIDIA Nemotron Nano 12B v2
NVIDIA · 12.3B · runs from 5.8 GB
NVIDIA Nemotron Nano 12B v2 is a 12.3B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 VL 2B Instruct
Alibaba · 2.1B · runs from 1.1 GB
Qwen3 VL 2B Instruct is Alibaba's smallest model in the Qwen3-VL lineup, a 2.1-billion-parameter vision-language model built to handle images and text together. It performs image captioning, visual question answering, and document reading, and its visual-agent tuning lets it interpret GUI screenshots for simple automation. Its small size suits on-device and edge deployment, running comfortably on modest consumer GPUs or even some laptops once quantized. The model supports a 262,144 token context window, enough for lengthy documents or multi-turn conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in October 2025 alongside the 8B and 32B models, it targets edge and mobile use cases, trading some visual reasoning depth for a much smaller footprint.
Qwen3 30B A3B Thinking 2507
Alibaba · 30.5B · runs from 8.8 GB
Qwen3 30B A3B Thinking 2507 is the reasoning-focused variant of Alibaba's 30-billion-parameter mixture-of-experts model, updated in July 2025. Like its instruct sibling, it activates only about 3 billion parameters per token, keeping resource demands low while enabling multi-step reasoning and chain-of-thought problem solving. This thinking variant is designed for tasks that benefit from deliberate, step-by-step logic such as math, coding puzzles, and analytical questions. Its efficient MoE design means users with modest GPUs can still access strong reasoning capabilities without needing datacenter-class hardware.
Qwen3.6 35B A3B Uncensored Heretic Native MTP Preserved
llmfan46 · 35.1B · runs from 15.3 GB
Qwen3.6 35B A3B Uncensored Heretic Native MTP Preserved is a 35.1B-parameter open language model from llmfan46 in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
MN 12B Mag Mell R1
inflatebot · 12.2B · runs from 4.1 GB
MN 12B Mag Mell R1 is a 12.2B-parameter open language model from inflatebot. It supports a context window of up to 1,024,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Deepseek Coder 6.7B Instruct
DeepSeek · 6.7B · runs from 4.2 GB
DeepSeek Coder 6.7B Instruct is a first-generation code-specialized model trained on a large corpus of source code and programming-related data. At 6.7 billion parameters, it provides solid code completion, generation, and explanation capabilities across popular programming languages while remaining small enough to run on most consumer GPUs. While newer models in the DeepSeek lineup have surpassed it in raw capability, this model remains a practical choice for users who need a lightweight local coding assistant with minimal hardware requirements. It runs well on GPUs with as little as 6 GB of VRAM when quantized.
Llama 2 7B Chat HF
Meta · 6.7B · runs from 3.1 GB
Meta Llama 2 7B Chat is a 7-billion parameter instruction-tuned model from Meta's Llama 2 family, optimized for dialogue use cases. It was fine-tuned using supervised fine-tuning and RLHF on top of the Llama 2 7B base model, with a 4K token context window. This model is suitable for basic conversational AI tasks and runs efficiently on consumer GPUs. While newer Llama generations offer improved performance, Llama 2 7B Chat remains a well-understood and widely-supported option for local inference. Released under the Llama 2 Community License.
Hy MT2 7B
Tencent · 8.0B · runs from 2.8 GB
Hy-MT2-7B is Tencent's mid-size entry in the Hy-MT2 family of dedicated, "fast-thinking" multilingual translation models, alongside 1.8B and 30B-A3B (mixture-of-experts) siblings. Rather than being a general-purpose chat assistant, it is trained specifically for translation across 33 languages and to follow multilingual translation instructions, and the card reports it outperforming open models such as DeepSeek-V4-Pro and Kimi K2.6 in fast-thinking mode. Tencent also ships FP8 and GGUF quantized builds for on-device use with llama.cpp. At roughly 8 billion dense parameters, it fits on a single consumer GPU once quantized. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2026, alongside the open-sourced IFMTBench instruction-following translation benchmark.
SmolLM3 3B
Hugging Face · 3.1B · runs from 1.3 GB
SmolLM3 3B is Hugging Face's latest-generation compact language model, representing a significant step up from the SmolLM2 series. At 3 billion parameters, it delivers considerably stronger reasoning, instruction following, and general language understanding while maintaining modest hardware requirements that keep it accessible on most consumer GPUs. This model benefits from improved training data, architectural refinements, and lessons learned from previous SmolLM generations. It is well positioned for local chatbot applications, coding assistance, and content generation tasks where you want strong performance without dedicating the resources required by 7B-class models.
Nex N2.5 Mini
nex-agi · 35.1B · runs from 10.0 GB
Nex N2.5 Mini is a 35.1B-parameter open language model from nex-agi. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
DeepSeek R1 Distill Qwen 14B
DeepSeek · 14.8B · runs from 5.1 GB
DeepSeek R1 Distill Qwen 14B sits in a sweet spot between the smaller 7B distill and the more demanding 32B version, offering strong reasoning performance at 14.8 billion parameters on the Qwen 2.5 architecture. It captures a meaningful share of the full R1's chain-of-thought capabilities while keeping resource requirements within the range of mainstream consumer GPUs. Quantized to 4-bit, it fits comfortably on GPUs with 12 GB of VRAM, delivering reliable step-by-step reasoning for math, logic, and analytical problems.
MiniCPM O 4 5
OpenBMB · 9.4B · runs from 4.6 GB
MiniCPM-o 4.5 is OpenBMB's roughly 9.4-billion-parameter omni-modal model, successor to MiniCPM-o 2.6, built from a SigLip2 vision encoder, a Whisper-medium audio encoder, a CosyVoice2 speech decoder, and a Qwen3-8B language backbone. It processes images, video, and audio and can generate spoken output, adding full-duplex live streaming so it can see, listen, and speak at once, plus a proactive mode that lets it comment unprompted, and bundled instruct and thinking modes. At this size it needs a high-end consumer GPU once quantized. Context length is 40,960 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in February 2026. Its defining upgrade over 2.6 is full-duplex interaction: earlier omni models process input and generate output in alternating turns, while 4.5 streams both directions at once without blocking.
Qwen2.5 Omni 7B
Alibaba · 10.7B · runs from 3.3 GB
Qwen2.5-Omni-7B is Alibaba's 10.7-billion-parameter flagship in the Qwen2.5-Omni family, an end-to-end model that perceives text, images, audio, and video and generates both text and natural streaming speech in response. It shares the family's Thinker-Talker design and TMRoPE positional scheme for synchronizing audio and video timestamps, tuned for low-latency, real-time conversation rather than turn-based chat alone. At just under 11 billion parameters, it needs a single mainstream-to-high-end consumer GPU once quantized. Its language backbone carries a 32,768 token context window. It is released under Alibaba's Qwen Research License, a non-commercial license limited to research and evaluation rather than the Apache 2.0 used for Qwen's text-only models. Published in March 2025, it was the first Omni release, with the smaller 3B variant following about a month later.
MiMo V2.6 Distill Qwen 9B
Xiaomi · 9.4B · runs from 3.4 GB
MiMo V2.6 Distill Qwen 9B is Xiaomi's 9.4-billion-parameter agentic model, built by fine-tuning Alibaba's Qwen3.5-9B on data generated by Xiaomi's larger MiMo-V2.6-Pro and MiMo-V2.6-Flash models. It handles vision, code, and function-calling tasks, and is positioned as the practical local-hardware option next to Xiaomi's much larger frontier MiMo-V2.6 models, which are designed for hosted use. At under 10 billion parameters, it runs well on a single mainstream consumer GPU once quantized. The model offers a 262K token context window, suited to long agent traces or multi-file coding sessions. Published in September 2026, it is one of the newest releases in Xiaomi's MiMo line, distilling the larger models' agentic and coding behaviour into a size that fits ordinary hardware.
Nanbeige4.1 3B
Nanbeige · 3.9B · runs from 2.1 GB
Nanbeige4.1 3B is a compact chat model from Nanbeige, a Chinese AI startup focused on building efficient small-scale language models. At just under 4 billion parameters, it is designed to run on virtually any modern GPU or even on CPU, making it one of the more accessible options for users with limited hardware. Despite its small size, it handles basic conversation, simple reasoning, and Chinese-English bilingual tasks, serving as a practical entry point for local LLM experimentation.
K2 Horizon MoVA 36B A4B
IFM · 37.4B · runs from 10.8 GB
K2 Horizon MoVA 36B A4B is a 37.4B-parameter open language model from IFM. It supports a context window of up to 524,288 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.8 Whittle MoE 27B A17.8B
logic65 · 26.9B · runs from 12.2 GB
Qwen3.8 Whittle MoE 27B A17.8B is a 26.9B-parameter open language model from logic65 in the Qwen 3.8 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.5 27B Claude 4.6 Opus Reasoning Distilled
Jackrong · 27.8B · runs from 8.4 GB
The full-precision version of Jackrong's Qwen3.5 27B reasoning distillation from Claude 4.6 Opus. With 27.8 billion parameters in unquantized form, this model preserves the maximum quality from the distillation process but requires significantly more VRAM, typically 56 GB or more in BF16. It is primarily intended for users with professional-grade GPUs or multi-GPU setups. This variant is ideal for further fine-tuning, experimentation, or running at full fidelity when hardware allows. Most users looking to run the model locally for inference should consider the GGUF-quantized version instead, which offers a much better tradeoff between quality and resource usage.
DeepSeek R1 Distill Llama 8B
DeepSeek · 8.0B · runs from 2.8 GB
DeepSeek R1 Distill Llama 8B brings R1's reinforcement-learned reasoning capabilities to the widely supported Llama 3.1 8B architecture. By distilling the full 684.5B R1 model's reasoning patterns into this 8 billion parameter dense model, DeepSeek created a version that benefits from the extensive Llama ecosystem of tools, quantizations, and inference engines. For users who prefer the Llama architecture or already have tooling built around it, this model offers a plug-and-play path to chain-of-thought reasoning. Its hardware requirements are very approachable, running well on consumer GPUs with 8 GB or more of VRAM at common quantization levels.