All LLM Models

Browse 1214 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

NVIDIA Nemotron Nano 9B v2

NVIDIA · 8.9B · runs from 3.5 GB

384.9K 520

NVIDIA Nemotron Nano 9B v2 is a compact yet capable chat model from NVIDIA, packing 8.9 billion parameters into a size that runs comfortably on a wide range of consumer GPUs. Built on NVIDIA's Nemotron architecture, it delivers strong instruction-following and conversational performance while keeping VRAM requirements modest. This second-generation Nano model reflects NVIDIA's push to make high-quality language models accessible on local hardware. It's an excellent starting point for users who want a responsive, general-purpose assistant without needing top-tier GPU memory.

Chat

Qwen3 4B Thinking 2507

Alibaba · 4.0B · runs from 2.2 GB

406.8K 617

Qwen3 4B Thinking 2507 is the reasoning-optimized variant of Alibaba's compact 4-billion-parameter Qwen3 model, released in the July 2025 update cycle. Despite its small size, this thinking variant is tuned to produce chain-of-thought reasoning and step-by-step problem solving, making it a surprisingly capable lightweight reasoner. This model is ideal for users who want basic reasoning and analytical capabilities on very modest hardware. It can run on most consumer GPUs and even some CPU-only setups when quantized, providing an accessible entry point for experimenting with reasoning-style models without any significant hardware investment.

Chat

LFM2.5 2.6B DSpark

Liquid AI · 328M · runs from 0.5 GB

4.9K 36

LFM2.5 2.6B DSpark is a 328M-parameter open language model from Liquid AI in the LFM2.5 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 3 4B IT

Google · 4.3B · runs from 2.0 GB

1.7M 1.5K

Gemma 3 4B IT is a small, roughly 4.3-billion-parameter model from Google's earlier Gemma 3 generation, tuned for instruction-following and chat. It accepts image input alongside text, so it can handle visual question answering and image-grounded prompts in addition to text-only conversation. At this size, the model is comfortable on modest consumer hardware, running even on integrated graphics at low quantization, making it one of the most accessible vision-capable models for local use. Released on February 20, 2025, Gemma 3 4B IT predates the Gemma 4 family and is distributed under Google's Gemma license terms rather than a standard open-source license. Its compact size and image-input support make it a practical choice for lightweight, on-device multimodal applications.

Vision

Hermes 3 Llama 3.1 70B

Nous Research · 70.6B · runs from 20.4 GB

3.5K 138

Hermes 3 Llama 3.1 70B is a 70.6B-parameter open language model from Nous Research in the Llama 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatRoleplay

QwQ 32B

Alibaba · 32.8B · runs from 14.8 GB

73.2K 3.0K

QwQ 32B is a 32-billion parameter reasoning-focused model from Alibaba Cloud's Qwen family. Unlike standard chat models, QwQ is specifically optimized for step-by-step logical reasoning, complex problem solving, and mathematical tasks. It employs extended chain-of-thought processing, generating detailed internal reasoning before producing final answers, which significantly improves accuracy on challenging analytical problems. The model requires a GPU with at least 24GB of VRAM for quantized inference and delivers reasoning performance competitive with much larger models. It is particularly well suited for users who need strong analytical capabilities for math, science, coding logic, and multi-step problem solving. Released under the Apache 2.0 license.

ChatReasoning

DeepSeek R1 Distill Llama 70B

DeepSeek · 70.6B · runs from 20.4 GB

78.5K 807

DeepSeek R1 Distill Llama 70B is the largest model in the R1 distillation lineup, combining the reasoning capabilities developed in the full 684.5B R1 with the robust Llama 3.1 70B architecture. At 70 billion parameters, it delivers the strongest reasoning performance of any dense R1 distill, approaching the full R1's quality on many math and coding benchmarks. Running this model locally requires a multi-GPU setup or a single GPU with very high VRAM capacity, though quantized versions can fit on hardware with 48 GB or more. For users who need top-tier open-weight reasoning and have the hardware to support a 70B dense model, this is one of the strongest options available.

ChatReasoning

Phi 4

Microsoft · 14.7B · runs from 7.0 GB

624.1K 2.3K

Microsoft Phi 4 is a 14-billion parameter language model from Microsoft Research's Phi series, designed to deliver strong reasoning, mathematical, and coding performance at an efficient size. Phi 4 continues the Phi family's focus on maximizing capability per parameter through high-quality training data curation, achieving benchmark scores that rival much larger models on reasoning and STEM tasks. The model runs well on consumer GPUs with 12-16GB of VRAM in quantized formats. It excels at mathematical problem solving, code generation, and structured reasoning. Released under the MIT license.

ChatMathCode

Mistral Nemo Instruct 2407

Mistral AI · 12.2B · runs from 5.9 GB

359.1K 1.7K

Mistral Nemo Instruct 2407 is a 12-billion-parameter instruction-tuned chat model from Mistral AI, a dated release from July 2024. It targets general dialogue and instruction-following use cases, sitting in a practical middle ground between lightweight and large-scale models in terms of both capability and resource demands. The model offers a 128K token context window, generous for its parameter class, and is released under the Apache 2.0 license, allowing unrestricted local and commercial use. At around 12 billion parameters, Mistral Nemo Instruct 2407 fits on a single consumer GPU when quantized, putting it within easy reach of local enthusiasts running their own hardware.

Chat

Qwen2 7B Instruct

Alibaba · 7.6B · runs from 2.7 GB

331.9K 688

Qwen2 7B Instruct is a 7.6B-parameter open language model from Alibaba in the Qwen 2 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Mistral Small Instruct 2409

Mistral AI · 22.2B · runs from 7.4 GB

4.0K 394

Mistral Small Instruct 2409 is a 22.2B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Yi 1.5 6B Chat

01.AI · 6.1B · runs from 3.0 GB

8.8K 42

Yi 1.5 6B Chat is a 6.1B-parameter open language model from 01.AI in the Yi 1.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Ministral 3 3B Reasoning 2512

Mistral AI · 4.3B · runs from 2.3 GB

73.4K 120

Ministral 3 3B Reasoning 2512 is a 4.3B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Reasoning

Gemma 3 12B IT

Google · 12.2B · runs from 5.7 GB

593.7K 840

Google Gemma 3 12B IT is a 12-billion parameter multimodal instruction-tuned model from Google's Gemma 3 series. It supports both text and image inputs, offering vision-language capabilities at a more accessible size point than the 27B variant. Gemma 3 12B IT runs on consumer GPUs with 12-16GB of VRAM in quantized formats, making it a practical choice for local multimodal inference without requiring top-tier hardware. Released under the Gemma license.

Vision

Gemma 3 27B IT

Google · 27.4B · runs from 12.8 GB

387.3K 2.0K

Google Gemma 3 27B IT is a 27.4-billion parameter multimodal instruction-tuned model from Google's Gemma 3 family. It supports both text and image inputs, making it one of the most capable openly available vision-language models for local inference. The model handles conversational AI, visual question answering, image description, and complex reasoning tasks across modalities. Gemma 3 27B IT requires a GPU with at least 24GB of VRAM for quantized inference, placing it within reach of high-end consumer cards like the RTX 4090. It uses a dense Transformer architecture with a large context window and benefits from Google's extensive pretraining pipeline. Released under the Gemma license.

Vision

Qwen2.5 Coder 1.5B Instruct

Alibaba · 1.5B · runs from 0.9 GB

555.0K 146

Qwen2.5 Coder 1.5B Instruct is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Bonsai 8B Unpacked

prism-ml · 8.2B · runs from 4.1 GB

9.0K 17

Bonsai 8B Unpacked is a 8.2B-parameter open language model from prism-ml. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Spark X2.5 1.7B

XHToken · 1.7B · runs from 1.1 GB

8.6K 142

Spark X2.5 1.7B is a 1.7B-parameter open language model from XHToken. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3.2 1B

Meta · 1.2B · runs from 0.6 GB

840.0K 2.6K

Meta Llama 3.2 1B is a 1.2-billion parameter base (pretrained) model from Meta's Llama 3.2 release. It is the smallest model in the Llama 3.2 family and is designed for research, fine-tuning, and embedding into resource-constrained environments. It supports a 128K token context window. As a base model, it is not optimized for conversational use without further fine-tuning. Its minimal resource requirements make it suitable for experimentation, edge deployment, and as a starting point for domain-specific fine-tuning. Released under the Llama 3.2 Community License.

Chat

Gemma 3 270M IT

Google · 268M · runs from 0.1 GB

73.8K 640

Google Gemma 3 270M IT is a 270-million parameter instruction-tuned model from Google's Gemma 3 family, an experimental release pushing the boundaries of how small an effective chat model can be. The model runs on virtually any hardware, including entry-level GPUs and CPU-only setups, making it useful for experimentation, education, and exploring the limits of small-scale language modeling. Released under the Gemma license.

Chat

GLM 4.6V Flash

Z.ai · 10.3B · runs from 3.2 GB

68.5K 627

GLM 4.6V Flash is a 10.3B-parameter open language model from Z.ai in the GLM 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

LFM2.5 8B A1B DSpark

Liquid AI · 8B · runs from 3.7 GB

4.2K 38

LFM2.5 8B A1B DSpark is a 8B-parameter open language model from Liquid AI in the LFM2.5 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen2.5 VL 3B Instruct

Alibaba · 3.8B · runs from 1.4 GB

2.4M 704

Qwen2.5 VL 3B Instruct is Alibaba's 3.8-billion-parameter vision-language model in the Qwen 2.5 lineup, built to process images and text together in a single conversation. It can describe images, answer questions about visual content, read charts and documents, and locate objects within a scene, making it a compact option for on-device or edge multimodal applications. Its small size means it runs comfortably on modest consumer GPUs, and even on laptops or lower-end hardware once quantized, without requiring a workstation-class card. The model supports a 128K token context window, enough for lengthy documents or extended visual conversations. It was published in January 2025 alongside the larger Qwen2.5-VL models, sharing the same architecture and vision encoder scaled down for lighter-weight, latency-sensitive deployments.

Vision

SmolLM2 135M Instruct

Hugging Face · 135M · runs from 0.4 GB

1.6M 423

SmolLM2 135M Instruct is the instruction-tuned variant of Hugging Face's 135-million-parameter SmolLM2 model. Fine-tuned to follow user prompts and engage in basic conversational exchanges, it delivers surprisingly coherent responses given its minimal size, making it ideal for testing chat interfaces or running on extremely constrained devices. This model is a practical choice when you need an instruction-following model that fits comfortably in under 1 GB of memory. It works well for simple question answering, text reformatting, and lightweight assistant tasks where response quality can be traded for instant inference speed.

Chat

Qwen2.5 Coder 3B Instruct

Alibaba · 3.1B · runs from 1.7 GB

620.3K 129

Qwen2.5 Coder 3B Instruct is a 3.1B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Qwen3.8 35B A3B Distill

empero-ai · 35.1B · runs from 12.2 GB

4.6K 58

Qwen3.8 35B A3B Distill is a 35.1B-parameter open language model from empero-ai in the Qwen 3.8 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatVisionReasoningFunctions

Devstral Small 2 24B Instruct 2512

Mistral AI · 24.0B · runs from 7.3 GB

301.1K 669

Devstral Small 2 24B Instruct is Mistral AI's dense 24-billion-parameter model for agentic software-engineering work, fine-tuned to follow instructions for chat, coding agents, and tool-heavy workflows. Built on the same architecture as Ministral 3, it adds vision capabilities for analyzing images alongside code and text, and its publisher designed it specifically to be lightweight enough for local, on-device use rather than requiring a large server. It supports a context window of roughly 384,000 tokens and is released under the Apache 2.0 license. Mistral notes it is light enough to run on a single RTX 4090 or a Mac with 32GB of RAM, consistent with its 4-bit memory needs of around 14GB.

Chat

Nanbeige4.2 3B

Nanbeige · 4.2B · runs from 2.2 GB

40.3K 801

Nanbeige4.2 3B is a 4.2B-parameter open language model from Nanbeige. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek R1 Distill Qwen 1.5B

DeepSeek · 1.8B · runs from 0.8 GB

445.0K 1.6K

DeepSeek R1 Distill Qwen 1.5B is the smallest model in the R1 distillation family, packing chain-of-thought reasoning capabilities into just 1.5 billion parameters using the Qwen 2.5 architecture. It represents an ambitious attempt to bring structured reasoning to the smallest practical model size. At this scale, the model can run on virtually any modern GPU and even on CPU-only setups with acceptable speed. While its reasoning depth is naturally limited compared to its larger siblings, it still demonstrates structured thinking patterns that set it apart from generic models of similar size.

ChatReasoning

Nemotron Orchestrator 8B

NVIDIA · 8.2B · runs from 4.1 GB

2.7K 598

Nemotron Orchestrator 8B is a 8.2B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat