All LLM Models

Browse 982 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

LFM2.5 8B A1B DSpark

Liquid AI · 8B · runs from 3.7 GB

4.2K 38

LFM2.5 8B A1B DSpark is a 8B-parameter open language model from Liquid AI in the LFM2.5 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen2.5 VL 3B Instruct

Alibaba · 3.8B · runs from 1.4 GB

2.4M 704

Qwen2.5 VL 3B Instruct is Alibaba's 3.8-billion-parameter vision-language model in the Qwen 2.5 lineup, built to process images and text together in a single conversation. It can describe images, answer questions about visual content, read charts and documents, and locate objects within a scene, making it a compact option for on-device or edge multimodal applications. Its small size means it runs comfortably on modest consumer GPUs, and even on laptops or lower-end hardware once quantized, without requiring a workstation-class card. The model supports a 128K token context window, enough for lengthy documents or extended visual conversations. It was published in January 2025 alongside the larger Qwen2.5-VL models, sharing the same architecture and vision encoder scaled down for lighter-weight, latency-sensitive deployments.

Vision

SmolLM2 135M Instruct

Hugging Face · 135M · runs from 0.4 GB

1.6M 423

SmolLM2 135M Instruct is the instruction-tuned variant of Hugging Face's 135-million-parameter SmolLM2 model. Fine-tuned to follow user prompts and engage in basic conversational exchanges, it delivers surprisingly coherent responses given its minimal size, making it ideal for testing chat interfaces or running on extremely constrained devices. This model is a practical choice when you need an instruction-following model that fits comfortably in under 1 GB of memory. It works well for simple question answering, text reformatting, and lightweight assistant tasks where response quality can be traded for instant inference speed.

Chat

Qwen2.5 Coder 3B Instruct

Alibaba · 3.1B · runs from 1.7 GB

620.3K 129

Qwen2.5 Coder 3B Instruct is a 3.1B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Devstral Small 2 24B Instruct 2512

Mistral AI · 24.0B · runs from 7.3 GB

301.1K 669

Devstral Small 2 24B Instruct is Mistral AI's dense 24-billion-parameter model for agentic software-engineering work, fine-tuned to follow instructions for chat, coding agents, and tool-heavy workflows. Built on the same architecture as Ministral 3, it adds vision capabilities for analyzing images alongside code and text, and its publisher designed it specifically to be lightweight enough for local, on-device use rather than requiring a large server. It supports a context window of roughly 384,000 tokens and is released under the Apache 2.0 license. Mistral notes it is light enough to run on a single RTX 4090 or a Mac with 32GB of RAM, consistent with its 4-bit memory needs of around 14GB.

Chat

Nanbeige4.2 3B

Nanbeige · 4.2B · runs from 2.2 GB

40.3K 801

Nanbeige4.2 3B is a 4.2B-parameter open language model from Nanbeige. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek R1 Distill Qwen 1.5B

DeepSeek · 1.8B · runs from 0.8 GB

445.0K 1.6K

DeepSeek R1 Distill Qwen 1.5B is the smallest model in the R1 distillation family, packing chain-of-thought reasoning capabilities into just 1.5 billion parameters using the Qwen 2.5 architecture. It represents an ambitious attempt to bring structured reasoning to the smallest practical model size. At this scale, the model can run on virtually any modern GPU and even on CPU-only setups with acceptable speed. While its reasoning depth is naturally limited compared to its larger siblings, it still demonstrates structured thinking patterns that set it apart from generic models of similar size.

ChatReasoning

Nemotron Orchestrator 8B

NVIDIA · 8.2B · runs from 4.1 GB

2.7K 598

Nemotron Orchestrator 8B is a 8.2B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Trinity Mini

Arcee AI · 26.1B · runs from 11.5 GB

25.3K 205

Trinity Mini is a 26.1B-parameter open language model from Arcee AI. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek R1 Distill Qwen 7B

DeepSeek · 7.6B · runs from 3.0 GB

275.7K 897

DeepSeek R1 Distill Qwen 7B compresses the reasoning techniques from DeepSeek's full R1 model into a compact 7.6 billion parameter dense model built on the Qwen 2.5 architecture. Despite its small footprint, it demonstrates surprisingly capable step-by-step reasoning on math and logic problems that would stump many models several times its size. This is one of the most accessible reasoning models available for local use, fitting comfortably on GPUs with 6 GB or more of VRAM when quantized. It strikes a practical balance between genuine chain-of-thought reasoning ability and the hardware constraints of a typical consumer setup.

ChatReasoning

Llama 3.2 3B

Meta · 3.2B · runs from 1.5 GB

366.5K 965

Meta Llama 3.2 3B is a 3.2-billion parameter base (pretrained) model from Meta's Llama 3.2 family. It supports a 128K token context window and is intended for fine-tuning, research, and custom applications rather than direct conversational use. The model provides a good balance between capability and efficiency at the small model scale. It is popular as a foundation for community fine-tunes and domain-specific adaptations. Released under the Llama 3.2 Community License.

Chat

Qwen3 VL 32B Instruct

Alibaba · 33.4B · runs from 9.8 GB

360.2K 240

Qwen3 VL 32B Instruct is Alibaba's largest dense model in the initial Qwen3-VL lineup, a 33.4-billion-parameter vision-language model built to process images and text in one pass. It handles image description, visual question answering, and document understanding, and its visual-agent tuning lets it read GUI screenshots and reason about on-screen elements. Local inference calls for quantization and a capable GPU, fitting on a single high-end consumer or workstation card. The model supports a 262,144 token context window, enough for long documents or extended multi-turn conversations. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. Published in October 2025 alongside the 2B and 8B models, it shares the family's long native context and video-understanding capabilities, giving stronger multimodal reasoning than the smaller variants.

Vision

Gemma 4 26B A4B IT Uncensored

TrevorJS · 25.8B · runs from 11.6 GB

2.1K 69

Gemma 4 26B A4B IT Uncensored is a 25.8B-parameter open language model from TrevorJS in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Ornith 1.5 35B A3B

ornith-ai · 36.0B · runs from 10.3 GB

363.7K 658

Ornith-1.5-35B-A3B is the mid-size model in Ornith AI's Ornith-1.5 family, a 36-billion-parameter Mixture-of-Experts model that activates only about 3 billion parameters per token. Like its siblings, it is built for agentic coding and inherits the family's self-improving training loop, which uses reinforcement learning to jointly refine training tasks, agent scaffolding, and solution rollouts rather than relying on a fixed, hand-built curriculum. It supports a 262,144-token context window and is released under the MIT license. Because MoE models keep all experts in memory, 4-bit quantization needs roughly 21GB, fitting a single GPU in the 16-24GB range.

Chat

Mistral Small 3.2 24B Instruct 2506

Mistral AI · 24.0B · runs from 7.3 GB

197.8K 619

Mistral-Small-3.2-24B-Instruct-2506 is a dense 24-billion-parameter vision-language model from Mistral AI, a minor refinement of Mistral-Small-3.1-24B-Instruct-2503. The update focuses on following precise instructions more reliably, cutting down on repetitive or runaway generations, and making function calling more robust, while still handling both text and image inputs for tasks like chart and document understanding. It supports a 131,072-token context window and is released under the Apache 2.0 license. At this size, 4-bit quantization needs roughly 14GB of memory, so it runs on a single consumer GPU in the 16-24GB range, such as an RTX 4090.

Chat

LFM2.5 350M

Liquid AI · 354M · runs from 0.5 GB

82.0K 424

LFM2.5 350M is a 354M-parameter open language model from Liquid AI in the LFM2.5 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen3 4B DFlash B16

z-lab · 537M · runs from 0.6 GB

25.9K 31

Qwen3 4B DFlash B16 is a 537M-parameter open language model from z-lab in the Qwen 3 family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Phi 4 Mini Reasoning

Microsoft · 3.8B · runs from 1.6 GB

61.4K 240

Phi-4-mini-reasoning is Microsoft's compact 3.8-billion-parameter model in the Phi-4 family, fine-tuned specifically for multi-step, logic-intensive mathematical reasoning using synthetic math data distilled from a larger teacher model. It targets tasks like formal proof generation, symbolic computation, and advanced word problems, and is designed to run in memory- and latency-constrained environments such as educational tools or embedded tutoring systems, rather than for general knowledge or open-domain chat. It supports a 128K token context window and is released under the MIT license. At 3.8 billion parameters, it runs comfortably on almost any modern laptop or consumer GPU, even at higher precision, making it one of the more accessible reasoning-focused models to run locally.

ChatMathCodeReasoning

TinyLlama 1.1B Chat v1.0

TinyLlama · 1.1B · runs from 0.8 GB

1.5M 1.8K

TinyLlama 1.1B Chat is a 1.1-billion parameter chat model built on the Llama 2 architecture and trained on approximately 3 trillion tokens, an unusually large dataset for a model of its size. The TinyLlama project demonstrated that small models can achieve strong performance when given sufficient training compute, making it a standout in the sub-2B parameter class. The Chat variant is fine-tuned for conversational use and runs on virtually any modern GPU, including entry-level cards with 4GB of VRAM or less. It is a practical choice for lightweight local inference, edge deployment, and experimentation where hardware resources are limited.

Chat

Qwen3.5 9B DFlash

z-lab · 1.3B · runs from 0.9 GB

10.6K 44

Qwen3.5 9B DFlash is a 1.3B-parameter open language model from z-lab in the Qwen 3.5 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen3.5 4B DFlash

z-lab · 634M · runs from 0.6 GB

5.6K 40

Qwen3.5 4B DFlash is a 634M-parameter open language model from z-lab in the Qwen 3.5 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Ornith 1.5 9B

ornith-ai · 9.7B · runs from 3.2 GB

483.3K 330

Ornith-1.5-9B is the lightest model in Ornith AI's Ornith-1.5 family, a dense 9.7-billion-parameter model built for agentic coding tasks and designed for efficient single-GPU deployment. It shares the family's self-improving training approach, in which reinforcement learning jointly optimizes the tasks, agent scaffolding, and solution rollouts used to train it, and a quantized variant is available for edge and mobile deployment. It offers a 262,144-token context window and is released under the MIT license. At roughly 9.7 billion parameters, 4-bit quantization needs only about 5.6GB of memory, so it runs comfortably on a single mid-range consumer GPU.

Chat

Diffusiongemma 26B A4B IT

Google · 25.8B · runs from 11.6 GB

628.5K 1.2K

Diffusiongemma 26B A4B IT takes a different approach from the rest of the Gemma 4 family: rather than generating tokens autoregressively, it uses a block-diffusion architecture. It still follows a mixture-of-experts design, with roughly 26 billion total parameters and about 4 billion active per token. The active-parameter figure is the more relevant one for inference cost, while the total parameter count determines memory needs, putting local use in single high-end consumer GPU territory once quantized. It also accepts image input alongside text for multimodal use. The model offers a 256K token context window and is released under the Apache 2.0 license. Published on June 9, 2026, it represents Google's exploration of diffusion-based language modeling within the Gemma 4 generation, alongside the more conventional autoregressive Gemma 4 variants also released that year.

Vision

Bonsai 1.7B Unpacked

prism-ml · 1.7B · runs from 1.3 GB

3.5K 13

Bonsai 1.7B Unpacked is a 1.7B-parameter open language model from prism-ml. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek R1 Distill Qwen 32B

DeepSeek · 32.8B · runs from 9.8 GB

478.1K 1.6K

DeepSeek R1 Distill Qwen 32B takes the reasoning capabilities developed in the full 684.5B R1 model and distills them into the 32.8 billion parameter Qwen 2.5 architecture. The result is a dense model that punches well above its weight class on math, science, and coding reasoning tasks, often matching models two to three times its size. At around 32.8 billion parameters, this model fits comfortably on a single high-end consumer GPU when quantized to 4-bit precision, making it one of the most capable reasoning models you can run on a desktop workstation.

ChatReasoning

Ministral 3 14B Reasoning 2512

Mistral AI · 13.9B · runs from 6.7 GB

109.5K 152

Ministral 3 14B Reasoning 2512 is a 13.9B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Reasoning

NVIDIA Nemotron Nano 12B v2

NVIDIA · 12.3B · runs from 5.8 GB

5.7K 164

NVIDIA Nemotron Nano 12B v2 is a 12.3B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen3 VL 2B Instruct

Alibaba · 2.1B · runs from 1.1 GB

3.1M 472

Qwen3 VL 2B Instruct is Alibaba's smallest model in the Qwen3-VL lineup, a 2.1-billion-parameter vision-language model built to handle images and text together. It performs image captioning, visual question answering, and document reading, and its visual-agent tuning lets it interpret GUI screenshots for simple automation. Its small size suits on-device and edge deployment, running comfortably on modest consumer GPUs or even some laptops once quantized. The model supports a 262,144 token context window, enough for lengthy documents or multi-turn conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in October 2025 alongside the 8B and 32B models, it targets edge and mobile use cases, trading some visual reasoning depth for a much smaller footprint.

Vision

Qwen3 30B A3B Thinking 2507

Alibaba · 30.5B · runs from 8.8 GB

89.1K 380

Qwen3 30B A3B Thinking 2507 is the reasoning-focused variant of Alibaba's 30-billion-parameter mixture-of-experts model, updated in July 2025. Like its instruct sibling, it activates only about 3 billion parameters per token, keeping resource demands low while enabling multi-step reasoning and chain-of-thought problem solving. This thinking variant is designed for tasks that benefit from deliberate, step-by-step logic such as math, coding puzzles, and analytical questions. Its efficient MoE design means users with modest GPUs can still access strong reasoning capabilities without needing datacenter-class hardware.

Chat

MN 12B Mag Mell R1

inflatebot · 12.2B · runs from 4.1 GB

40.0K 239

MN 12B Mag Mell R1 is a 12.2B-parameter open language model from inflatebot. It supports a context window of up to 1,024,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning