All LLM Models

Browse 82 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Qwen2.5 7B Instruct

Alibaba · 7.6B · runs from 2.7 GB

9.7M 2.2K

Qwen2.5 7B Instruct is a 7.6-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It supports a 128K token context window and is fine-tuned for conversational AI, instruction following, and general assistant tasks. Its efficient size makes it well-suited for local deployment on consumer GPUs with 8GB or more of VRAM. The model delivers strong performance for its parameter class across reasoning, multilingual understanding, and coding tasks. It benefits from the improved pretraining data and techniques of the Qwen 2.5 generation. Released under the Apache 2.0 license and widely supported by inference frameworks such as llama.cpp, vLLM, and Ollama.

Chat

Qwen3 14B

Alibaba · 14.8B · runs from 4.7 GB

2.2M 478

Qwen3 14B is a 14-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 series. It occupies a practical middle ground in the Qwen 3 lineup, offering stronger reasoning and generation quality than the 8B variant while remaining manageable on GPUs with 16GB or more of VRAM in quantized formats. The model supports hybrid thinking mode for flexible reasoning depth. Qwen3 14B is well suited for chat, instruction following, coding assistance, and multilingual tasks. It benefits from the generational improvements of Qwen 3 in pretraining data and alignment techniques, delivering performance that competes with larger models from previous generations. Released under the Apache 2.0 license.

Chat

Qwen3 8B

Alibaba · 8.2B · runs from 2.9 GB

12.7M 2.0K

Qwen3 8B is an 8.2-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 series. It is a general-purpose chat model that delivers strong performance across reasoning, multilingual understanding, and coding tasks while remaining efficient enough to run on consumer GPUs with 8GB or more of VRAM. Like other Qwen 3 models, it supports hybrid thinking mode for flexible reasoning depth. The model benefits from the improved pretraining data and training methodology of the Qwen 3 generation, offering notable quality gains over Qwen 2.5 at the same parameter count. It is widely supported by inference frameworks including llama.cpp, vLLM, and Ollama. Released under the Apache 2.0 license.

Chat

Qwen2.5 Coder 7B Instruct

Alibaba · 7.6B · runs from 3.0 GB

2.5M 803

Qwen2.5 Coder 7B Instruct is a 7.6-billion parameter code-specialized instruction-tuned model from Alibaba Cloud. It is trained on a large corpus of source code and natural language, fine-tuned for programming assistance tasks such as code generation, completion, debugging, and code explanation. The model supports a 128K token context window and runs efficiently on consumer GPUs with 8GB or more of VRAM. It provides a good balance between coding capability and hardware requirements for developers looking to run a local coding assistant. Released under the Apache 2.0 license.

ChatCode

Qwen2.5 14B Instruct

Alibaba · 14.8B · runs from 5.1 GB

2.1M 370

Qwen2.5 14B Instruct is a 14-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It supports a 128K token context window and provides a balanced tradeoff between quality and hardware requirements, running well on GPUs with 16GB of VRAM in quantized formats. The model is fine-tuned for chat, instruction following, and general-purpose assistant tasks. It performs well across reasoning, coding, and multilingual benchmarks for its size class, making it a practical option for local deployment when larger models are not feasible. Released under the Apache 2.0 license.

Chat

Qwen3.5 9B

Alibaba · 9.7B · runs from 3.2 GB

9.6M 2.0K

Qwen3.5-9B is Alibaba's dense 9-billion-parameter language model with a built-in vision encoder, part of the Qwen3.5 generation. This is the post-trained, instruction-tuned release, combining Gated DeltaNet and gated-attention layers for efficient long-context inference, and it can take images alongside text for tasks like visual question answering as well as general chat, coding, and tool use. It offers a 262,144-token native context window, extensible up to roughly 1,010,000 tokens, and is released under the Apache 2.0 license. At around 9.7 billion parameters, 4-bit quantization needs only about 5.6GB of memory, so it fits comfortably on a single mid-range consumer GPU.

Vision

Qwen3 4B

Alibaba · 4.0B · runs from 1.6 GB

7.0M 712

Qwen3 4B is a compact 4-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 family. It is designed for efficient local inference on consumer hardware, supporting chat and general assistant tasks while fitting comfortably on GPUs with 6GB or more of VRAM in quantized formats. The model supports hybrid thinking mode, allowing it to balance reasoning depth and response speed. Despite its small footprint, Qwen3 4B delivers quality competitive with larger models from previous generations, making it a practical choice for lightweight local deployments and resource-constrained environments. Released under the Apache 2.0 license.

Chat

Qwen2.5 Coder 14B Instruct

Alibaba · 14.8B · runs from 5.1 GB

1.9M 187

Qwen2.5 Coder 14B Instruct is a 14.8B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Qwen3 4B Instruct 2507

Alibaba · 4.0B · runs from 1.9 GB

4.0M 975

Qwen3 4B Instruct 2507 is a July 2025 refresh of Alibaba's compact 4-billion-parameter chat model from the Qwen3 family. This updated release brings improved instruction following and conversational quality while remaining lightweight enough to run on most modern GPUs and even some higher-end integrated graphics setups. With its modest size, the 4B Instruct 2507 strikes a practical balance between capability and resource efficiency. It is well suited for everyday chat, summarization, and light assistant tasks on consumer hardware, making it one of the more accessible entry points into the Qwen3 lineup.

Chat

Qwen3 0.6B

Alibaba · 752M · runs from 0.6 GB

26.7M 1.7K

Qwen3 0.6B is the smallest instruction-tuned model in Alibaba Cloud's Qwen 3 family, with approximately 752 million parameters. It is designed for ultra-lightweight deployment where minimal hardware resources are available, running comfortably on virtually any modern GPU or CPU-only setups. The model supports hybrid thinking mode despite its tiny footprint. While limited in reasoning depth compared to larger variants, Qwen3 0.6B handles basic chat, simple summarization, and lightweight instruction following. It is primarily useful for edge deployment, rapid prototyping, and experimentation where model size is a critical constraint. Released under the Apache 2.0 license.

Chat

Qwen2.5 3B Instruct

Alibaba · 3.1B · runs from 1.4 GB

4.4M 575

Qwen2.5 3B Instruct is a 3.1-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 family. It is designed for efficient local inference on consumer hardware, supporting a 128K token context window despite its compact footprint. The model can run on GPUs with as little as 4GB of VRAM when quantized. Despite its small size, Qwen2.5 3B Instruct delivers competitive performance for basic conversational tasks, summarization, and simple instruction following. It is a good option for edge deployment and resource-constrained environments. Released under the Apache 2.0 license.

Chat

Qwen2.5 1.5B Instruct

Alibaba · 1.5B · runs from 0.8 GB

7.9M 851

Qwen2.5 1.5B Instruct is a 1.5-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It is a lightweight model suitable for deployment on minimal hardware, including low-VRAM GPUs and even CPU-only setups with acceptable latency. It supports a 128K token context window. The model handles basic conversational tasks, simple question answering, and text generation. While limited in reasoning depth compared to larger variants, it is useful for applications where fast response times and minimal resource consumption are priorities. Released under the Apache 2.0 license.

Chat

Qwen3.5 2B

Alibaba · 2.3B · runs from 1.0 GB

4.8M 415

Qwen3.5 2B is a 2.3-billion-parameter model from Alibaba's Qwen team, one of the smaller entries in the Qwen3.5 lineup (0.8B–9B) built to handle text and image input together. It uses a hybrid architecture mixing linear-attention layers with periodic full-attention layers to keep long-context inference efficient. As a vision-capable model it can read and reason about images alongside written prompts, and its small size lets it run on a laptop or even a phone once quantized, without a dedicated GPU. It supports an unusually long 262K token context window for its size. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in late February 2026 as part of Alibaba's push toward compact, natively multimodal edge models.

Vision

Qwen3.5 0.8B

Alibaba · 873M · runs from 0.6 GB

2.4M 728

Qwen3.5-0.8B is the smallest model in Alibaba's Qwen3.5 generation, a dense sub-billion-parameter vision-language model using the same hybrid Gated DeltaNet and gated-attention design as the rest of the family. It is a post-trained, instruction-tuned release, and its publisher notes that at this scale it is best suited to prototyping, task-specific fine-tuning, and other research or development purposes rather than production chat. It supports a 262,144-token native context window and is released under the Apache 2.0 license. At under one billion parameters, it needs well under a gigabyte of memory at 4-bit, so it runs on essentially any modern laptop or GPU.

Vision

Qwen3 VL 4B Instruct

Alibaba · 4.4B · runs from 1.7 GB

3.5M 473

Qwen3 VL 4B Instruct is Alibaba's 4.4-billion-parameter vision-language model from the Qwen3-VL family, built to read and reason over images and video alongside text in one conversation. It suits multimodal chat, document and chart understanding, and simple visual-agent tasks like describing screenshots or pulling structured data out of pictures. At this size the model runs comfortably on a mainstream consumer GPU once quantized, and is light enough for many laptops. The model supports a 262K token context window, giving room for long documents or extended visual conversations. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2025 as part of the Qwen3-VL generation, built for deeper visual perception and stronger agentic behavior than the earlier Qwen2-VL series.

Vision

Qwen3 1.7B

Alibaba · 2.0B · runs from 1.1 GB

3.6M 555

Qwen3 1.7B is a 1.7-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 series. It is a lightweight model designed for deployment on minimal hardware, including low-VRAM GPUs and even CPU-only configurations with acceptable latency. Despite its compact size, it supports hybrid thinking mode and handles basic conversational tasks, simple question answering, and text generation. The model is useful for edge deployment, embedded applications, and scenarios where fast inference with minimal resource consumption is the priority. It represents a significant quality improvement over Qwen 2.5 at the sub-2B scale. Released under the Apache 2.0 license.

Chat

Qwen3 VL 8B Instruct

Alibaba · 8.8B · runs from 3.0 GB

20.1M 1.1K

Qwen3 VL 8B Instruct is Alibaba's 8.8-billion-parameter vision-language model in the Qwen3-VL lineup, built to process images and text together in one conversation. Beyond image description and visual question answering, it is tuned as a visual agent that can read GUI screenshots and reason about on-screen elements, useful for document analysis and early automation tasks. At this size, local inference is practical on a single mainstream-to-high-end consumer GPU once quantized. The model supports a 262,144 token context window, enough for long documents or extended chat histories. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. Published in October 2025 alongside 2B and 32B siblings, Qwen3-VL adds video understanding with fine-grained event indexing beyond earlier Qwen vision-language releases.

Vision

Qwen2.5 14B

Alibaba · 14.8B · runs from 6.8 GB

78.2K 157

Qwen2.5 14B is a 14.8B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen2.5 0.5B Instruct

Alibaba · 494M · runs from 0.5 GB

8.7M 635

Qwen2.5 0.5B Instruct is the smallest instruction-tuned model in Alibaba Cloud's Qwen 2.5 family, with just 494 million parameters. It is designed for ultra-lightweight deployment scenarios where minimal hardware resources are available, running comfortably on virtually any modern GPU or even CPU-only configurations. Despite its tiny footprint, the model supports a 128K token context window and can handle basic chat, simple summarization, and lightweight instruction following. It is primarily useful for edge deployment, experimentation, and prototyping where model size is a critical constraint. Released under the Apache 2.0 license.

Chat

Qwen3 4B Thinking 2507

Alibaba · 4.0B · runs from 2.2 GB

406.8K 617

Qwen3 4B Thinking 2507 is the reasoning-optimized variant of Alibaba's compact 4-billion-parameter Qwen3 model, released in the July 2025 update cycle. Despite its small size, this thinking variant is tuned to produce chain-of-thought reasoning and step-by-step problem solving, making it a surprisingly capable lightweight reasoner. This model is ideal for users who want basic reasoning and analytical capabilities on very modest hardware. It can run on most consumer GPUs and even some CPU-only setups when quantized, providing an accessible entry point for experimenting with reasoning-style models without any significant hardware investment.

Chat

Qwen2 7B Instruct

Alibaba · 7.6B · runs from 2.7 GB

331.9K 688

Qwen2 7B Instruct is a 7.6B-parameter open language model from Alibaba in the Qwen 2 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen2.5 Coder 1.5B Instruct

Alibaba · 1.5B · runs from 0.9 GB

555.0K 146

Qwen2.5 Coder 1.5B Instruct is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Qwen2.5 VL 3B Instruct

Alibaba · 3.8B · runs from 1.4 GB

2.4M 704

Qwen2.5 VL 3B Instruct is Alibaba's 3.8-billion-parameter vision-language model in the Qwen 2.5 lineup, built to process images and text together in a single conversation. It can describe images, answer questions about visual content, read charts and documents, and locate objects within a scene, making it a compact option for on-device or edge multimodal applications. Its small size means it runs comfortably on modest consumer GPUs, and even on laptops or lower-end hardware once quantized, without requiring a workstation-class card. The model supports a 128K token context window, enough for lengthy documents or extended visual conversations. It was published in January 2025 alongside the larger Qwen2.5-VL models, sharing the same architecture and vision encoder scaled down for lighter-weight, latency-sensitive deployments.

Vision

Qwen2.5 Coder 3B Instruct

Alibaba · 3.1B · runs from 1.7 GB

620.3K 129

Qwen2.5 Coder 3B Instruct is a 3.1B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Qwen3 VL 2B Instruct

Alibaba · 2.1B · runs from 1.1 GB

3.1M 472

Qwen3 VL 2B Instruct is Alibaba's smallest model in the Qwen3-VL lineup, a 2.1-billion-parameter vision-language model built to handle images and text together. It performs image captioning, visual question answering, and document reading, and its visual-agent tuning lets it interpret GUI screenshots for simple automation. Its small size suits on-device and edge deployment, running comfortably on modest consumer GPUs or even some laptops once quantized. The model supports a 262,144 token context window, enough for lengthy documents or multi-turn conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in October 2025 alongside the 8B and 32B models, it targets edge and mobile use cases, trading some visual reasoning depth for a much smaller footprint.

Vision

Qwen2.5 Omni 7B

Alibaba · 10.7B · runs from 3.3 GB

338.0K 1.9K

Qwen2.5-Omni-7B is Alibaba's 10.7-billion-parameter flagship in the Qwen2.5-Omni family, an end-to-end model that perceives text, images, audio, and video and generates both text and natural streaming speech in response. It shares the family's Thinker-Talker design and TMRoPE positional scheme for synchronizing audio and video timestamps, tuned for low-latency, real-time conversation rather than turn-based chat alone. At just under 11 billion parameters, it needs a single mainstream-to-high-end consumer GPU once quantized. Its language backbone carries a 32,768 token context window. It is released under Alibaba's Qwen Research License, a non-commercial license limited to research and evaluation rather than the Apache 2.0 used for Qwen's text-only models. Published in March 2025, it was the first Omni release, with the smaller 3B variant following about a month later.

Chat

Qwen2.5 Math 7B Instruct

Alibaba · 7.6B · runs from 3.0 GB

84.0K 92

Qwen2.5 Math 7B Instruct is a 7.6B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatMath

Qwen2 VL 2B Instruct

Alibaba · 2.2B · runs from 1.1 GB

2.0M 520

Qwen2 VL 2B Instruct is a 2.2-billion-parameter vision-language model from Alibaba's Qwen2-VL series, able to process images, multi-image comparisons, and video alongside text prompts. It targets visual question answering, document and chart reading, and basic agentic tasks such as interpreting a screenshot to plan a next action. Its small size makes it well suited to laptops and even some phones, running comfortably on modest consumer hardware once quantized. The model supports a 32K token context window, enough for moderate documents or extended chat. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in August 2024. Qwen2-VL introduced Naive Dynamic Resolution and Multimodal Rotary Position Embedding, letting it handle arbitrary image resolutions and understand videos well over twenty minutes long.

Vision

Qwen2.5 Coder 0.5B Instruct

Alibaba · 494M · runs from 0.5 GB

140.1K 78

Qwen2.5 Coder 0.5B Instruct is a 494M-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Qwen2.5 Math 1.5B Instruct

Alibaba · 1.5B · runs from 1 GB

321.9K 58

Qwen2.5 Math 1.5B Instruct is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatMath