All LLM Models

Browse 880 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Qwen2 7B Instruct

Alibaba · 7.6B · runs from 2.7 GB

331.9K 688

Qwen2 7B Instruct is a 7.6B-parameter open language model from Alibaba in the Qwen 2 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Mistral Small Instruct 2409

Mistral AI · 22.2B · runs from 7.4 GB

4.0K 394

Mistral Small Instruct 2409 is a 22.2B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Yi 1.5 6B Chat

01.AI · 6.1B · runs from 3.0 GB

8.8K 42

Yi 1.5 6B Chat is a 6.1B-parameter open language model from 01.AI in the Yi 1.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Ministral 3 3B Reasoning 2512

Mistral AI · 4.3B · runs from 2.3 GB

73.4K 120

Ministral 3 3B Reasoning 2512 is a 4.3B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Reasoning

Gemma 3 12B IT

Google · 12.2B · runs from 5.7 GB

593.7K 840

Google Gemma 3 12B IT is a 12-billion parameter multimodal instruction-tuned model from Google's Gemma 3 series. It supports both text and image inputs, offering vision-language capabilities at a more accessible size point than the 27B variant. Gemma 3 12B IT runs on consumer GPUs with 12-16GB of VRAM in quantized formats, making it a practical choice for local multimodal inference without requiring top-tier hardware. Released under the Gemma license.

Vision

Qwen2.5 Coder 1.5B Instruct

Alibaba · 1.5B · runs from 0.9 GB

555.0K 146

Qwen2.5 Coder 1.5B Instruct is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Bonsai 8B Unpacked

prism-ml · 8.2B · runs from 4.1 GB

9.0K 17

Bonsai 8B Unpacked is a 8.2B-parameter open language model from prism-ml. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Spark X2.5 1.7B

XHToken · 1.7B · runs from 1.1 GB

8.6K 142

Spark X2.5 1.7B is a 1.7B-parameter open language model from XHToken. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3.2 1B

Meta · 1.2B · runs from 0.6 GB

840.0K 2.6K

Meta Llama 3.2 1B is a 1.2-billion parameter base (pretrained) model from Meta's Llama 3.2 release. It is the smallest model in the Llama 3.2 family and is designed for research, fine-tuning, and embedding into resource-constrained environments. It supports a 128K token context window. As a base model, it is not optimized for conversational use without further fine-tuning. Its minimal resource requirements make it suitable for experimentation, edge deployment, and as a starting point for domain-specific fine-tuning. Released under the Llama 3.2 Community License.

Chat

Gemma 3 270M IT

Google · 268M · runs from 0.1 GB

73.8K 640

Google Gemma 3 270M IT is a 270-million parameter instruction-tuned model from Google's Gemma 3 family, an experimental release pushing the boundaries of how small an effective chat model can be. The model runs on virtually any hardware, including entry-level GPUs and CPU-only setups, making it useful for experimentation, education, and exploring the limits of small-scale language modeling. Released under the Gemma license.

Chat

GLM 4.6V Flash

Z.ai · 10.3B · runs from 3.2 GB

68.5K 627

GLM 4.6V Flash is a 10.3B-parameter open language model from Z.ai in the GLM 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

LFM2.5 8B A1B DSpark

Liquid AI · 8B · runs from 3.7 GB

4.2K 38

LFM2.5 8B A1B DSpark is a 8B-parameter open language model from Liquid AI in the LFM2.5 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen2.5 VL 3B Instruct

Alibaba · 3.8B · runs from 1.4 GB

2.4M 704

Qwen2.5 VL 3B Instruct is Alibaba's 3.8-billion-parameter vision-language model in the Qwen 2.5 lineup, built to process images and text together in a single conversation. It can describe images, answer questions about visual content, read charts and documents, and locate objects within a scene, making it a compact option for on-device or edge multimodal applications. Its small size means it runs comfortably on modest consumer GPUs, and even on laptops or lower-end hardware once quantized, without requiring a workstation-class card. The model supports a 128K token context window, enough for lengthy documents or extended visual conversations. It was published in January 2025 alongside the larger Qwen2.5-VL models, sharing the same architecture and vision encoder scaled down for lighter-weight, latency-sensitive deployments.

Vision

SmolLM2 135M Instruct

Hugging Face · 135M · runs from 0.4 GB

1.6M 423

SmolLM2 135M Instruct is the instruction-tuned variant of Hugging Face's 135-million-parameter SmolLM2 model. Fine-tuned to follow user prompts and engage in basic conversational exchanges, it delivers surprisingly coherent responses given its minimal size, making it ideal for testing chat interfaces or running on extremely constrained devices. This model is a practical choice when you need an instruction-following model that fits comfortably in under 1 GB of memory. It works well for simple question answering, text reformatting, and lightweight assistant tasks where response quality can be traded for instant inference speed.

Chat

Qwen2.5 Coder 3B Instruct

Alibaba · 3.1B · runs from 1.7 GB

620.3K 129

Qwen2.5 Coder 3B Instruct is a 3.1B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Devstral Small 2 24B Instruct 2512

Mistral AI · 24.0B · runs from 7.3 GB

301.1K 669

Devstral Small 2 24B Instruct is Mistral AI's dense 24-billion-parameter model for agentic software-engineering work, fine-tuned to follow instructions for chat, coding agents, and tool-heavy workflows. Built on the same architecture as Ministral 3, it adds vision capabilities for analyzing images alongside code and text, and its publisher designed it specifically to be lightweight enough for local, on-device use rather than requiring a large server. It supports a context window of roughly 384,000 tokens and is released under the Apache 2.0 license. Mistral notes it is light enough to run on a single RTX 4090 or a Mac with 32GB of RAM, consistent with its 4-bit memory needs of around 14GB.

Chat

Nanbeige4.2 3B

Nanbeige · 4.2B · runs from 2.2 GB

40.3K 801

Nanbeige4.2 3B is a 4.2B-parameter open language model from Nanbeige. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek R1 Distill Qwen 1.5B

DeepSeek · 1.8B · runs from 0.8 GB

445.0K 1.6K

DeepSeek R1 Distill Qwen 1.5B is the smallest model in the R1 distillation family, packing chain-of-thought reasoning capabilities into just 1.5 billion parameters using the Qwen 2.5 architecture. It represents an ambitious attempt to bring structured reasoning to the smallest practical model size. At this scale, the model can run on virtually any modern GPU and even on CPU-only setups with acceptable speed. While its reasoning depth is naturally limited compared to its larger siblings, it still demonstrates structured thinking patterns that set it apart from generic models of similar size.

ChatReasoning

Nemotron Orchestrator 8B

NVIDIA · 8.2B · runs from 4.1 GB

2.7K 598

Nemotron Orchestrator 8B is a 8.2B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek R1 Distill Qwen 7B

DeepSeek · 7.6B · runs from 3.0 GB

275.7K 897

DeepSeek R1 Distill Qwen 7B compresses the reasoning techniques from DeepSeek's full R1 model into a compact 7.6 billion parameter dense model built on the Qwen 2.5 architecture. Despite its small footprint, it demonstrates surprisingly capable step-by-step reasoning on math and logic problems that would stump many models several times its size. This is one of the most accessible reasoning models available for local use, fitting comfortably on GPUs with 6 GB or more of VRAM when quantized. It strikes a practical balance between genuine chain-of-thought reasoning ability and the hardware constraints of a typical consumer setup.

ChatReasoning

Llama 3.2 3B

Meta · 3.2B · runs from 1.5 GB

366.5K 965

Meta Llama 3.2 3B is a 3.2-billion parameter base (pretrained) model from Meta's Llama 3.2 family. It supports a 128K token context window and is intended for fine-tuning, research, and custom applications rather than direct conversational use. The model provides a good balance between capability and efficiency at the small model scale. It is popular as a foundation for community fine-tunes and domain-specific adaptations. Released under the Llama 3.2 Community License.

Chat

Mistral Small 3.2 24B Instruct 2506

Mistral AI · 24.0B · runs from 7.3 GB

197.8K 619

Mistral-Small-3.2-24B-Instruct-2506 is a dense 24-billion-parameter vision-language model from Mistral AI, a minor refinement of Mistral-Small-3.1-24B-Instruct-2503. The update focuses on following precise instructions more reliably, cutting down on repetitive or runaway generations, and making function calling more robust, while still handling both text and image inputs for tasks like chart and document understanding. It supports a 131,072-token context window and is released under the Apache 2.0 license. At this size, 4-bit quantization needs roughly 14GB of memory, so it runs on a single consumer GPU in the 16-24GB range, such as an RTX 4090.

Chat

LFM2.5 350M

Liquid AI · 354M · runs from 0.5 GB

82.0K 424

LFM2.5 350M is a 354M-parameter open language model from Liquid AI in the LFM2.5 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen3 4B DFlash B16

z-lab · 537M · runs from 0.6 GB

25.9K 31

Qwen3 4B DFlash B16 is a 537M-parameter open language model from z-lab in the Qwen 3 family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Phi 4 Mini Reasoning

Microsoft · 3.8B · runs from 1.6 GB

61.4K 240

Phi-4-mini-reasoning is Microsoft's compact 3.8-billion-parameter model in the Phi-4 family, fine-tuned specifically for multi-step, logic-intensive mathematical reasoning using synthetic math data distilled from a larger teacher model. It targets tasks like formal proof generation, symbolic computation, and advanced word problems, and is designed to run in memory- and latency-constrained environments such as educational tools or embedded tutoring systems, rather than for general knowledge or open-domain chat. It supports a 128K token context window and is released under the MIT license. At 3.8 billion parameters, it runs comfortably on almost any modern laptop or consumer GPU, even at higher precision, making it one of the more accessible reasoning-focused models to run locally.

ChatMathCodeReasoning

TinyLlama 1.1B Chat v1.0

TinyLlama · 1.1B · runs from 0.8 GB

1.5M 1.8K

TinyLlama 1.1B Chat is a 1.1-billion parameter chat model built on the Llama 2 architecture and trained on approximately 3 trillion tokens, an unusually large dataset for a model of its size. The TinyLlama project demonstrated that small models can achieve strong performance when given sufficient training compute, making it a standout in the sub-2B parameter class. The Chat variant is fine-tuned for conversational use and runs on virtually any modern GPU, including entry-level cards with 4GB of VRAM or less. It is a practical choice for lightweight local inference, edge deployment, and experimentation where hardware resources are limited.

Chat

Qwen3.5 9B DFlash

z-lab · 1.3B · runs from 0.9 GB

10.6K 44

Qwen3.5 9B DFlash is a 1.3B-parameter open language model from z-lab in the Qwen 3.5 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen3.5 4B DFlash

z-lab · 634M · runs from 0.6 GB

5.6K 40

Qwen3.5 4B DFlash is a 634M-parameter open language model from z-lab in the Qwen 3.5 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Ornith 1.5 9B

ornith-ai · 9.7B · runs from 3.2 GB

483.3K 330

Ornith-1.5-9B is the lightest model in Ornith AI's Ornith-1.5 family, a dense 9.7-billion-parameter model built for agentic coding tasks and designed for efficient single-GPU deployment. It shares the family's self-improving training approach, in which reinforcement learning jointly optimizes the tasks, agent scaffolding, and solution rollouts used to train it, and a quantized variant is available for edge and mobile deployment. It offers a 262,144-token context window and is released under the MIT license. At roughly 9.7 billion parameters, 4-bit quantization needs only about 5.6GB of memory, so it runs comfortably on a single mid-range consumer GPU.

Chat

Bonsai 1.7B Unpacked

prism-ml · 1.7B · runs from 1.3 GB

3.5K 13

Bonsai 1.7B Unpacked is a 1.7B-parameter open language model from prism-ml. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat