All LLM Models

Browse 130 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Qwen3 Coder 30B A3B Instruct

Alibaba · 30.5B · runs from 8.8 GB

564.9K 1.3K

Qwen3 Coder 30B A3B Instruct is a code-specialized Mixture of Experts (MoE) model from Alibaba Cloud's Qwen 3 Coder series, with 30 billion total parameters and approximately 3 billion active parameters per forward pass. The MoE architecture allows it to deliver strong coding performance while keeping per-token compute costs low, making it faster at inference than comparably capable dense models. The model is instruction-tuned for programming assistance, code generation, debugging, and software engineering conversation. It requires VRAM proportional to its total 30B parameter count for loading weights, but benefits from efficient inference throughput due to its low active parameter count. Released under the Apache 2.0 license.

ChatCode

Qwen3.6 35B A3B

Alibaba · 36.0B · runs from 15.7 GB

3.2M 2.9K

Qwen3.6 35B A3B pairs a mixture-of-experts design with roughly 35 billion total parameters, only about 3 billion of which are active per token. That small active footprint keeps generation speed close to a 3B-class dense model, even though the full parameter set must still be held in memory, similar to a dense model of that size, meaning local use still calls for a single high-end consumer GPU once quantized. The model also accepts image input alongside text, extending it to visual question answering and image-grounded chat. It supports a 256K token context window for long documents and extended conversations, and is released under the Apache 2.0 license. Published in April 2026, it gives Qwen 3.6 users a faster-inference alternative to the dense Qwen3.6 27B model released the same generation.

Vision

Qwen2.5 7B Instruct

Alibaba · 7.6B · runs from 2.7 GB

9.7M 2.2K

Qwen2.5 7B Instruct is a 7.6-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It supports a 128K token context window and is fine-tuned for conversational AI, instruction following, and general assistant tasks. Its efficient size makes it well-suited for local deployment on consumer GPUs with 8GB or more of VRAM. The model delivers strong performance for its parameter class across reasoning, multilingual understanding, and coding tasks. It benefits from the improved pretraining data and techniques of the Qwen 2.5 generation. Released under the Apache 2.0 license and widely supported by inference frameworks such as llama.cpp, vLLM, and Ollama.

Chat

Qwen3 14B

Alibaba · 14.8B · runs from 4.7 GB

2.2M 478

Qwen3 14B is a 14-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 series. It occupies a practical middle ground in the Qwen 3 lineup, offering stronger reasoning and generation quality than the 8B variant while remaining manageable on GPUs with 16GB or more of VRAM in quantized formats. The model supports hybrid thinking mode for flexible reasoning depth. Qwen3 14B is well suited for chat, instruction following, coding assistance, and multilingual tasks. It benefits from the generational improvements of Qwen 3 in pretraining data and alignment techniques, delivering performance that competes with larger models from previous generations. Released under the Apache 2.0 license.

Chat

Qwen3 8B

Alibaba · 8.2B · runs from 2.9 GB

12.7M 2.0K

Qwen3 8B is an 8.2-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 series. It is a general-purpose chat model that delivers strong performance across reasoning, multilingual understanding, and coding tasks while remaining efficient enough to run on consumer GPUs with 8GB or more of VRAM. Like other Qwen 3 models, it supports hybrid thinking mode for flexible reasoning depth. The model benefits from the improved pretraining data and training methodology of the Qwen 3 generation, offering notable quality gains over Qwen 2.5 at the same parameter count. It is widely supported by inference frameworks including llama.cpp, vLLM, and Ollama. Released under the Apache 2.0 license.

Chat

Qwen3.8 27B

Alibaba · 27.8B · runs from 12.6 GB

6.9M 16.1K

Alibaba's Qwen3.8 27B is a dense, vision-capable model with around 28 billion parameters, continuing the Qwen line into its 3.8 release. It can process images alongside text prompts, supporting tasks like visual question answering and image-grounded chat in addition to general-purpose reasoning and coding assistance. At this size, the model fits on a single high-end consumer GPU once quantized, putting it within reach of enthusiast local setups rather than requiring server-class hardware. It ships with a 256K token context window, enough for long-form documents or extended multi-turn sessions, and is distributed under the Apache 2.0 license. Released in August 2026, it represents one of the more recent entries in Alibaba's Qwen 3.8 family of open-weight models.

Vision

Qwen2.5 Coder 7B Instruct

Alibaba · 7.6B · runs from 3.0 GB

2.5M 803

Qwen2.5 Coder 7B Instruct is a 7.6-billion parameter code-specialized instruction-tuned model from Alibaba Cloud. It is trained on a large corpus of source code and natural language, fine-tuned for programming assistance tasks such as code generation, completion, debugging, and code explanation. The model supports a 128K token context window and runs efficiently on consumer GPUs with 8GB or more of VRAM. It provides a good balance between coding capability and hardware requirements for developers looking to run a local coding assistant. Released under the Apache 2.0 license.

ChatCode

Qwen2.5 14B Instruct

Alibaba · 14.8B · runs from 5.1 GB

2.1M 370

Qwen2.5 14B Instruct is a 14-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It supports a 128K token context window and provides a balanced tradeoff between quality and hardware requirements, running well on GPUs with 16GB of VRAM in quantized formats. The model is fine-tuned for chat, instruction following, and general-purpose assistant tasks. It performs well across reasoning, coding, and multilingual benchmarks for its size class, making it a practical option for local deployment when larger models are not feasible. Released under the Apache 2.0 license.

Chat

Qwen3.5 9B

Alibaba · 9.7B · runs from 3.2 GB

9.6M 2.0K

Qwen3.5-9B is Alibaba's dense 9-billion-parameter language model with a built-in vision encoder, part of the Qwen3.5 generation. This is the post-trained, instruction-tuned release, combining Gated DeltaNet and gated-attention layers for efficient long-context inference, and it can take images alongside text for tasks like visual question answering as well as general chat, coding, and tool use. It offers a 262,144-token native context window, extensible up to roughly 1,010,000 tokens, and is released under the Apache 2.0 license. At around 9.7 billion parameters, 4-bit quantization needs only about 5.6GB of memory, so it fits comfortably on a single mid-range consumer GPU.

Vision

Qwen2.5 32B Instruct

Alibaba · 32.8B · runs from 9.8 GB

2.1M 362

Qwen2.5 32B Instruct is a 32-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 family. It occupies a practical sweet spot between the 14B and 72B variants, offering strong reasoning and multilingual capabilities while remaining feasible to run on a single high-end consumer GPU with 24GB or more of VRAM at reduced precision. The model supports a 128K token context window and is optimized for conversational use, instruction following, and structured output generation. It is a popular choice for local inference when the 72B model is too demanding but users need more capability than the 14B variant. Released under the Apache 2.0 license.

Chat

Qwen3 32B

Alibaba · 32.8B · runs from 9.7 GB

5.0M 749

Qwen3 32B is the flagship dense model in Alibaba Cloud's Qwen 3 series, with 32 billion parameters. It is instruction-tuned for chat and delivers strong performance across reasoning, coding, mathematics, and multilingual tasks. Qwen3 32B supports a hybrid thinking mode that allows the model to engage in extended chain-of-thought reasoning or respond quickly depending on the task, giving users flexibility between depth and speed. The model requires a GPU with at least 24GB of VRAM for quantized inference, placing it within reach of high-end consumer cards like the RTX 4090. It represents a significant generational improvement over Qwen 2.5 in both instruction following and knowledge breadth. Released under the Apache 2.0 license.

Chat

Qwen2.5 Coder 32B Instruct

Alibaba · 32.8B · runs from 9.8 GB

1.2M 2.2K

Qwen2.5 Coder 32B Instruct is a 32.8-billion parameter code-specialized model from Alibaba Cloud, instruction-tuned for programming assistance and code generation. It is trained on a large corpus of source code alongside natural language data, making it highly capable for tasks such as code completion, debugging, code explanation, and software engineering dialogue. The model supports a 128K token context window and delivers code generation quality competitive with the best open-weight coding models at any scale. It requires a GPU with at least 24GB of VRAM for quantized inference. Released under the Apache 2.0 license.

ChatCode

Qwen3 4B

Alibaba · 4.0B · runs from 1.6 GB

7.0M 712

Qwen3 4B is a compact 4-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 family. It is designed for efficient local inference on consumer hardware, supporting chat and general assistant tasks while fitting comfortably on GPUs with 6GB or more of VRAM in quantized formats. The model supports hybrid thinking mode, allowing it to balance reasoning depth and response speed. Despite its small footprint, Qwen3 4B delivers quality competitive with larger models from previous generations, making it a practical choice for lightweight local deployments and resource-constrained environments. Released under the Apache 2.0 license.

Chat

Qwen2.5 Coder 14B Instruct

Alibaba · 14.8B · runs from 5.1 GB

1.9M 187

Qwen2.5 Coder 14B Instruct is a 14.8B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Qwen3.5 122B A10B

Alibaba · 125.1B · runs from 250.6 GB

331.9K 622

Qwen3.5 122B A10B is Alibaba's 125-billion-parameter mixture-of-experts model from the Qwen 3.5 medium series, with about 10 billion parameters active per token (the A10B in its name). Because only the active experts run per token, inference is markedly faster than a dense model of comparable size, though all the weights still need to be held in memory. It is a unified vision-language model, handling images alongside text for reasoning, coding, and agentic workflows. Given its size, local inference needs a multi-GPU workstation or a large unified-memory machine; most people access a model this size through a hosted endpoint instead. It supports a 262K token context window and is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in February 2026, it sits alongside Qwen3.5-Flash, Qwen3.5-27B, and Qwen3.5-35B-A3B in Alibaba's medium-model series.

Vision

Qwen3 4B Instruct 2507

Alibaba · 4.0B · runs from 1.9 GB

4.0M 975

Qwen3 4B Instruct 2507 is a July 2025 refresh of Alibaba's compact 4-billion-parameter chat model from the Qwen3 family. This updated release brings improved instruction following and conversational quality while remaining lightweight enough to run on most modern GPUs and even some higher-end integrated graphics setups. With its modest size, the 4B Instruct 2507 strikes a practical balance between capability and resource efficiency. It is well suited for everyday chat, summarization, and light assistant tasks on consumer hardware, making it one of the more accessible entry points into the Qwen3 lineup.

Chat

Qwen3 0.6B

Alibaba · 752M · runs from 0.6 GB

26.7M 1.7K

Qwen3 0.6B is the smallest instruction-tuned model in Alibaba Cloud's Qwen 3 family, with approximately 752 million parameters. It is designed for ultra-lightweight deployment where minimal hardware resources are available, running comfortably on virtually any modern GPU or CPU-only setups. The model supports hybrid thinking mode despite its tiny footprint. While limited in reasoning depth compared to larger variants, Qwen3 0.6B handles basic chat, simple summarization, and lightweight instruction following. It is primarily useful for edge deployment, rapid prototyping, and experimentation where model size is a critical constraint. Released under the Apache 2.0 license.

Chat

Qwen3 VL 30B A3B Instruct

Alibaba · 31.1B · runs from 13.6 GB

457.6K 604

Qwen3 VL 30B A3B Instruct is Alibaba's 31-billion-parameter mixture-of-experts vision-language model in the Qwen 3 lineup, with about 3 billion parameters active per token (the A3B in its name). Because only the active experts run for each token, inference is faster than a similarly sized dense model, while all the weights still need to fit in memory. It handles images and video alongside text, supporting visual question answering, document and chart reading, and multi-image reasoning. At this size, local inference calls for quantization and a single high-end consumer GPU. The model supports a 262K token context window, suited to long documents, video transcripts, or extended visual conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in September 2025 as part of Qwen's third-generation vision-language family.

Vision

Qwen3 30B A3B Instruct 2507

Alibaba · 30.5B · runs from 13.4 GB

784.5K 840

Qwen3 30B A3B Instruct 2507 is a July 2025 updated mixture-of-experts model from Alibaba with 30 billion total parameters but only around 3 billion active during inference. This MoE architecture gives it a remarkably small memory and compute footprint relative to its total parameter count, letting users run a model with broad knowledge on mid-range hardware. The 2507 instruct refresh improves alignment and instruction-following quality over the original release. Because only a fraction of the weights are active at any given time, this model can often run on a single consumer GPU with 8 GB or more of VRAM when quantized, making it an excellent choice for users who want strong chat performance without heavyweight hardware.

Chat

Qwen2.5 3B Instruct

Alibaba · 3.1B · runs from 1.4 GB

4.4M 575

Qwen2.5 3B Instruct is a 3.1-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 family. It is designed for efficient local inference on consumer hardware, supporting a 128K token context window despite its compact footprint. The model can run on GPUs with as little as 4GB of VRAM when quantized. Despite its small size, Qwen2.5 3B Instruct delivers competitive performance for basic conversational tasks, summarization, and simple instruction following. It is a good option for edge deployment and resource-constrained environments. Released under the Apache 2.0 license.

Chat

Qwen3 30B A3B

Alibaba · 30.5B · runs from 8.8 GB

1.7M 948

Qwen3 30B A3B is a Mixture of Experts (MoE) model from Alibaba Cloud's Qwen 3 series, with 30 billion total parameters and approximately 3 billion active parameters per forward pass. The MoE architecture delivers quality significantly above what a standard 3B dense model could achieve, while keeping per-token compute costs low. It supports hybrid thinking mode for flexible reasoning. The model requires VRAM proportional to its full 30B parameter count for weight loading, but its low active parameter count results in fast inference throughput. It is an efficient option for users who want quality beyond dense small models without the full cost of larger architectures. Released under the Apache 2.0 license.

Chat

Qwen2.5 1.5B Instruct

Alibaba · 1.5B · runs from 0.8 GB

7.9M 851

Qwen2.5 1.5B Instruct is a 1.5-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It is a lightweight model suitable for deployment on minimal hardware, including low-VRAM GPUs and even CPU-only setups with acceptable latency. It supports a 128K token context window. The model handles basic conversational tasks, simple question answering, and text generation. While limited in reasoning depth compared to larger variants, it is useful for applications where fast response times and minimal resource consumption are priorities. Released under the Apache 2.0 license.

Chat

Qwen3.5 2B

Alibaba · 2.3B · runs from 1.0 GB

4.8M 415

Qwen3.5 2B is a 2.3-billion-parameter model from Alibaba's Qwen team, one of the smaller entries in the Qwen3.5 lineup (0.8B–9B) built to handle text and image input together. It uses a hybrid architecture mixing linear-attention layers with periodic full-attention layers to keep long-context inference efficient. As a vision-capable model it can read and reason about images alongside written prompts, and its small size lets it run on a laptop or even a phone once quantized, without a dedicated GPU. It supports an unusually long 262K token context window for its size. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in late February 2026 as part of Alibaba's push toward compact, natively multimodal edge models.

Vision

Qwen2.5 72B Instruct

Alibaba · 72.7B · runs from 21.0 GB

326.0K 994

Qwen2.5 72B Instruct is the flagship model of the Qwen 2.5 series from Alibaba Cloud, with 72.7 billion parameters. It is instruction-tuned for conversational use and excels across reasoning, coding, mathematics, and multilingual tasks. Qwen2.5 72B delivers performance competitive with leading open-weight 70B-class models while supporting a 128K token context window and structured output generation. The model uses a Transformer architecture with grouped-query attention and was pretrained on a diverse multilingual corpus of over 18 trillion tokens. Running it locally requires high-VRAM GPUs or multi-GPU setups, though quantized formats make it accessible on workstation-class hardware. Released under the Apache 2.0 license.

Chat

Qwen3.5 0.8B

Alibaba · 873M · runs from 0.6 GB

2.4M 728

Qwen3.5-0.8B is the smallest model in Alibaba's Qwen3.5 generation, a dense sub-billion-parameter vision-language model using the same hybrid Gated DeltaNet and gated-attention design as the rest of the family. It is a post-trained, instruction-tuned release, and its publisher notes that at this scale it is best suited to prototyping, task-specific fine-tuning, and other research or development purposes rather than production chat. It supports a 262,144-token native context window and is released under the Apache 2.0 license. At under one billion parameters, it needs well under a gigabyte of memory at 4-bit, so it runs on essentially any modern laptop or GPU.

Vision

Qwen2.5 VL 7B Instruct

Alibaba · 8.3B · runs from 17 GB

6.8M 1.7K

Qwen2.5 VL 7B Instruct is Alibaba's 8.3-billion-parameter vision-language model in the Qwen 2.5 lineup, built to handle images and video alongside text in a single pass. It can read documents, charts, and screenshots, describe and reason about visual content, and point out object locations within an image, useful for document understanding, visual question answering, and lightweight visual-agent tasks. At this parameter count, local inference is practical on a single mainstream or high-end consumer GPU once the weights are quantized. The model supports a 128K token context window, enough for long documents or extended multi-turn visual conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in January 2025 as part of Qwen's second-generation vision-language line, later succeeded by Qwen3-VL.

Vision

Qwen3 Coder Next

Alibaba · 79.7B · runs from 22.3 GB

520.7K 1.7K

Qwen3 Coder Next is a 79.7-billion parameter code-specialized instruction-tuned model from Alibaba Cloud, the next generation of the Qwen Coder series. It is trained extensively on source code and programming-related data, delivering strong performance across code generation, completion, debugging, refactoring, and software engineering dialogue. The model represents a significant step up in coding capability within the Qwen family. Due to its large parameter count, running Qwen3 Coder Next locally requires substantial VRAM, typically 48GB or more at reduced precision, placing it in the territory of professional GPUs or multi-GPU consumer setups. It is a top-tier choice for developers who need the most capable local coding assistant available. Released under the Apache 2.0 license.

ChatCode

Qwen3 VL 4B Instruct

Alibaba · 4.4B · runs from 1.7 GB

3.5M 473

Qwen3 VL 4B Instruct is Alibaba's 4.4-billion-parameter vision-language model from the Qwen3-VL family, built to read and reason over images and video alongside text in one conversation. It suits multimodal chat, document and chart understanding, and simple visual-agent tasks like describing screenshots or pulling structured data out of pictures. At this size the model runs comfortably on a mainstream consumer GPU once quantized, and is light enough for many laptops. The model supports a 262K token context window, giving room for long documents or extended visual conversations. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2025 as part of the Qwen3-VL generation, built for deeper visual perception and stronger agentic behavior than the earlier Qwen2-VL series.

Vision

Qwen3.8 Flash Next

Alibaba · 180.0B · runs from 76.9 GB

807.5K 5.6K

Qwen3.8 Flash Next is a vision-capable model in Alibaba's Qwen 3.8 line, with roughly 180 billion parameters. The Flash Next naming points to a fast-inference variant within the family, and the model accepts image input alongside text for combined visual and textual tasks, extending the Qwen line into multimodal territory. It provides a 256K token context window for long documents or extended sessions, and is released under a custom license outside standard open-source terms. At around 180 billion parameters, running Qwen3.8 Flash Next locally needs multi-GPU or server-class hardware rather than a single consumer GPU, even for quantized deployments.

Vision

Qwen3.5 27B

Alibaba · 27.8B · runs from 8.4 GB

1.9M 1.1K

Qwen3.5 27B is a 27.8-billion-parameter dense model from Alibaba's Qwen team, built for both text and image input using the same hybrid linear/full-attention architecture as the rest of the Qwen3.5 family. As a vision-capable model it can describe and reason about images alongside prompts, suiting multimodal chat and document tasks. At this size, local inference calls for quantization and a single high-end 24GB-plus consumer or workstation GPU rather than budget hardware. It supports a 262K token context window for long documents and multi-turn conversations, and is released under the Apache 2.0 license for unrestricted commercial and research use. Published in late February 2026, it sits alongside a separate small-model tier (0.8B–9B) in the same family, giving developers a mid-sized, single-GPU-friendly dense option.

Vision