All LLM Models

Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Qwen3 VL 4B Instruct

Alibaba · 4.4B · runs from 1.7 GB

3.5M 473

Qwen3 VL 4B Instruct is Alibaba's 4.4-billion-parameter vision-language model from the Qwen3-VL family, built to read and reason over images and video alongside text in one conversation. It suits multimodal chat, document and chart understanding, and simple visual-agent tasks like describing screenshots or pulling structured data out of pictures. At this size the model runs comfortably on a mainstream consumer GPU once quantized, and is light enough for many laptops. The model supports a 262K token context window, giving room for long documents or extended visual conversations. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2025 as part of the Qwen3-VL generation, built for deeper visual perception and stronger agentic behavior than the earlier Qwen2-VL series.

Vision

Qwen3.8 Flash Next

Alibaba · 180.0B · runs from 76.9 GB

807.5K 5.6K

Qwen3.8 Flash Next is a vision-capable model in Alibaba's Qwen 3.8 line, with roughly 180 billion parameters. The Flash Next naming points to a fast-inference variant within the family, and the model accepts image input alongside text for combined visual and textual tasks, extending the Qwen line into multimodal territory. It provides a 256K token context window for long documents or extended sessions, and is released under a custom license outside standard open-source terms. At around 180 billion parameters, running Qwen3.8 Flash Next locally needs multi-GPU or server-class hardware rather than a single consumer GPU, even for quantized deployments.

Vision

Qwen3.5 27B

Alibaba · 27.8B · runs from 8.4 GB

1.9M 1.1K

Qwen3.5 27B is a 27.8-billion-parameter dense model from Alibaba's Qwen team, built for both text and image input using the same hybrid linear/full-attention architecture as the rest of the Qwen3.5 family. As a vision-capable model it can describe and reason about images alongside prompts, suiting multimodal chat and document tasks. At this size, local inference calls for quantization and a single high-end 24GB-plus consumer or workstation GPU rather than budget hardware. It supports a 262K token context window for long documents and multi-turn conversations, and is released under the Apache 2.0 license for unrestricted commercial and research use. Published in late February 2026, it sits alongside a separate small-model tier (0.8B–9B) in the same family, giving developers a mid-sized, single-GPU-friendly dense option.

Vision

Spark X2.5 4B

XHToken · 4.1B · runs from 2.2 GB

33.3K 1.3K

Spark X2.5 4B is a 4.1B-parameter open language model from XHToken. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatFunctions

Qwen3 1.7B

Alibaba · 2.0B · runs from 1.1 GB

3.6M 555

Qwen3 1.7B is a 1.7-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 3 series. It is a lightweight model designed for deployment on minimal hardware, including low-VRAM GPUs and even CPU-only configurations with acceptable latency. Despite its compact size, it supports hybrid thinking mode and handles basic conversational tasks, simple question answering, and text generation. The model is useful for edge deployment, embedded applications, and scenarios where fast inference with minimal resource consumption is the priority. It represents a significant quality improvement over Qwen 2.5 at the sub-2B scale. Released under the Apache 2.0 license.

Chat

Kimi K3

Moonshot AI · 2779.9B · runs from 770.2 GB

1.9M 11.5K

Kimi K3 from Moonshot AI is an enormous model, weighing in at roughly 2.8 trillion parameters, among the largest open-weight releases in this catalogue. It accepts image input alongside text, supporting multimodal chat and visual reasoning in addition to standard dialogue. At this scale, local deployment is not realistic for most users; it needs multi-GPU or server-class hardware, and most people will access it through a hosted endpoint rather than running it on their own machines. Kimi K3 supports a 1M token context window, allowing it to process very long documents, codebases, or conversation histories in one pass. It was released on June 13, 2026, under its own custom license terms rather than a standard Apache or MIT license, so users should review the license text before commercial use.

VisionChat

Laguna S 2.1

poolside · 117.6B · runs from 50.5 GB

26.0K 1.0K

Laguna S 2.1 is a 117.6B-parameter open language model from poolside. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Laguna XS 2.1

poolside · 33.4B · runs from 14.6 GB

18.0K 248

Laguna XS 2.1 is a 33.4B-parameter open language model from poolside. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Mistral Small 24B Instruct 2501

Mistral AI · 23.6B · runs from 10.7 GB

50.4K 970

Mistral Small 24B Instruct is Mistral AI's January 2025 release targeting the mid-range parameter sweet spot. At 24 billion parameters it sits between lightweight 7B models and heavier 70B-class offerings, delivering strong instruction-following, reasoning, and coding performance without demanding top-tier hardware. This model fits comfortably on a single GPU with 16–24 GB of VRAM at common quantization levels, making it an attractive option for users with cards like the RTX 4090 or RTX 3090 who want a noticeable step up from 7B models. It strikes an appealing balance between quality and resource requirements for serious local use.

Chat

Qwen3 235B A22B Instruct 2507

Alibaba · 235.1B · runs from 71.0 GB

206.2K 801

Qwen3 235B A22B Instruct 2507 is Alibaba's flagship instruction-tuned model from the July 2025 update, featuring 235 billion total parameters with approximately 22 billion active during inference. As the largest instruct model in the Qwen3 lineup, it delivers top-tier conversational quality, knowledge depth, and instruction following. Despite its massive total parameter count, the MoE architecture keeps active compute manageable. Running this model locally still requires substantial hardware, typically multi-GPU setups with 48 GB or more of total VRAM, but the 2507 refresh makes it one of the most capable open-weight models available for users with high-end local infrastructure.

Chat

Qwen3.6 27B

Alibaba · 27.8B · runs from 11.5 GB

3.1M 2.3K

Released a few months before its Qwen3.8 sibling, Qwen3.6 27B is Alibaba's dense, roughly 28-billion-parameter model from the Qwen 3.6 generation. Vision input is supported alongside text, so the model can handle image-based prompts as well as standard chat, coding, and reasoning tasks. Like other models in this size class, it is best run locally on a single high-end consumer GPU with quantization rather than lower-end hardware. Qwen3.6 27B carries a 256K token context window for handling lengthy inputs, and is released under the Apache 2.0 license, permitting free commercial and research use. It sits alongside the larger Qwen3.6 35B A3B mixture-of-experts model as one of two Qwen 3.6 options with local-deployment potential.

Vision

Qwen AgentWorld 35B A3B

Alibaba · 34.7B · runs from 9.9 GB

18.7K 722

Qwen AgentWorld 35B A3B is Alibaba's 35-billion-parameter mixture-of-experts model, with about 3 billion parameters active per token (the A3B in its name). Unlike a typical chat model, it is a language world model: given an agent's action and history, it predicts what the environment does next, covering domains such as terminal, web, Android, and software engineering. Only the active experts run per token, so inference stays fast even though all the weights must fit in memory. At this size, local inference calls for quantization and a single high-end consumer GPU. The model supports a 262K token context window and is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in June 2026, it is built on a Qwen3.5-35B-A3B base and is meant for simulating agents, not direct conversation.

ChatFunctions

Kimi K2.5

Moonshot AI · 1026.9B · runs from 286.3 GB

312.8K 2.9K

Kimi K2.5 is an earlier release in Moonshot AI's Kimi K2 series, sharing the roughly 1-trillion-parameter scale of later checkpoints in the line. Released at the start of 2026, it accepts image input alongside text, giving it multimodal capability on top of standard chat use. Its context window reaches 256K tokens, enough for lengthy documents or long conversation histories, and it is released under a custom license outside the standard open-source terms. Because of its size, Kimi K2.5 needs multi-GPU or server-class hardware; most people will use it through a hosted endpoint rather than running it themselves.

Vision

Ternary Bonsai 8B Unpacked

prism-ml · 8.2B · runs from 4.1 GB

1.1K 16

Ternary Bonsai 8B Unpacked is a 8.2B-parameter open language model from prism-ml. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Hy MT2 1.8B

Tencent · 2.0B · runs from 1.1 GB

30.4K 1.2K

Hy MT2 1.8B is a 2-billion-parameter translation model from Tencent's Hunyuan team, part of a three-size Hy-MT2 family (1.8B, 7B, and a 30B mixture-of-experts variant) built specifically for machine translation rather than general chat. It supports translation across roughly three dozen languages and follows translation-specific instructions, such as adjusting tone or terminology. It runs on modest consumer hardware, including laptops, and Tencent has also released extreme-quantized versions shrunk to a few hundred megabytes for on-device use. It supports a 262K token context window, letting it translate long documents in a single pass. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in May 2026, it is the smallest member of a translation-focused family Tencent positions as a self-hostable alternative to commercial translation APIs.

Translation

Granite 4.2 8B

IBM · 8.8B · runs from 3.0 GB

105.5K 82

Granite 4.2 8B is a 8.8B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

KAT Coder V2.5 Dev

Kwaipilot · 34.7B · runs from 9.9 GB

9.6K 649

KAT Coder V2.5 Dev is a 34.7B-parameter open language model from Kwaipilot. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCodeFunctions

Mistral 7B v0.3

Mistral AI · 7.2B · runs from 3.6 GB

176.7K 594

Mistral 7B v0.3 is a 7.2B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen3 VL 8B Instruct

Alibaba · 8.8B · runs from 3.0 GB

20.1M 1.1K

Qwen3 VL 8B Instruct is Alibaba's 8.8-billion-parameter vision-language model in the Qwen3-VL lineup, built to process images and text together in one conversation. Beyond image description and visual question answering, it is tuned as a visual agent that can read GUI screenshots and reason about on-screen elements, useful for document analysis and early automation tasks. At this size, local inference is practical on a single mainstream-to-high-end consumer GPU once quantized. The model supports a 262,144 token context window, enough for long documents or extended chat histories. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. Published in October 2025 alongside 2B and 32B siblings, Qwen3-VL adds video understanding with fine-grained event indexing beyond earlier Qwen vision-language releases.

Vision

Kimi K2.6

Moonshot AI · 1026.9B · runs from 286.3 GB

441.0K 1.6K

Kimi K2.6, released by Moonshot AI in April 2026, is a large vision-capable model in the Kimi K2 line with roughly 1 trillion parameters. It accepts image input alongside text, extending the K2 series into multimodal use, and stands as one of the largest openly released checkpoints in this catalogue. The model provides a 256K token context window, suited to long documents and extended multi-turn sessions, and is distributed under a custom license rather than standard open-source terms. Given its scale, running Kimi K2.6 locally needs multi-GPU or server-class hardware; most users will reach it through a hosted endpoint rather than locally.

Vision

Ornith 1.0 35B

Deep Reinforce · 35.1B · runs from 10.0 GB

1.6M 506

Ornith 1.0 35B, from Deep Reinforce, is a 35-billion-parameter model built on a mixture-of-experts architecture, meaning only a subset of its parameters activate for any given token even though the full model must be held in memory. Because the full expert set still needs to fit in memory much like a dense model of the same size, local use is best suited to a single high-end consumer GPU once quantized rather than lower-end hardware. It is tuned for general chat and instruction-following. The model provides a 256K token context window for handling long inputs, and is released under the MIT license, one of the more permissive options in this catalogue. It was published on June 21, 2026, the same day as its smaller sibling, Ornith 1.0 9B.

Chat

Kimi K2.7 Code

Moonshot AI · 1026.9B · runs from 286.3 GB

104.8K 1.4K

Kimi-K2.7-Code is Moonshot AI's coding-focused agentic model, built on top of Kimi K2.6 with improvements aimed at long-horizon software engineering workflows and reduced reasoning-token usage. It is a Mixture-of-Experts model with about 1 trillion total parameters and roughly 32 billion active per token, and it also includes a vision encoder, so it can take image input alongside code and text while working through complex, multi-step coding tasks. The model supports a 256K token context window and is released under a modified MIT license. At roughly 1 trillion total parameters, even 4-bit quantization needs several hundred gigabytes of memory, so this is squarely server or multi-GPU territory; most people will use it through a hosted API rather than running it locally.

VisionCode

MiniCPM5 1B Claude Opus Fable5 v2 Thinking

GnLOLot · 1.1B · runs from 0.8 GB

1.2K 44

MiniCPM5 1B Claude Opus Fable5 v2 Thinking is a 1.1B-parameter open language model from GnLOLot in the MiniCPM family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatFunctions

North Mini Code 1.0

Cohere · 30.5B · runs from 8.8 GB

10.8K 568

North Mini Code 1.0 is Cohere's first open-weight coding model, a mixture-of-experts design with about 30.5 billion total parameters and roughly 3.3 billion active per token. It is tuned for agentic coding: tool calling, terminal workflows, and software-engineering tasks rather than general chat. Only the active parameters compute per token, so it decodes quickly for its size, though all weights must fit in memory; it runs on a single high-end consumer GPU once quantized. The model supports a 500,000 token context window, useful for large codebases or long agent transcripts. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in June 2026 as the debut release in Cohere's North family, it is designed to run on a single H100-class GPU in FP8.

ChatCodeFunctions

Llama 3.3 70B Instruct

Meta · 70.6B · runs from 21.3 GB

917.9K 3.1K

Meta Llama 3.3 70B Instruct is a 70-billion parameter large language model from Meta, released as part of the Llama 3.3 generation. It is an instruction-tuned model optimized for dialogue and chat use cases, offering strong performance across reasoning, coding, and multilingual tasks. Llama 3.3 70B delivers quality competitive with much larger models while remaining feasible to run on high-end consumer or workstation GPUs with sufficient VRAM. The model uses a grouped-query attention architecture with a 128K token context window and was trained on a massive multilingual corpus. It is released under the Llama 3.3 Community License, making it one of the most capable openly available models for local inference.

Chat

Step 3.7 Flash

StepFun · 201.4B · runs from 60.9 GB

23.1K 458

Step 3.7 Flash is StepFun's 201.4-billion-parameter mixture-of-experts vision-language model, built to handle images and text together in agentic workflows such as coding assistance and search. It supports tool use and selectable reasoning intensity, trading speed for deeper reasoning. At this scale, the full weight set must be held in memory, putting it in server-class territory: running it locally needs multiple high-memory GPUs, and most people will use a hosted API instead. The model supports a 262,144 token context window, enough for long documents or extended agentic sessions. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in May 2026, Step 3.7 Flash is positioned by StepFun as an efficiency-focused model for coding agents and search, adding native image understanding over the earlier Step 3.5 Flash.

Vision

Apodex 1.1 Mini

apodex · 36.0B · runs from 12.5 GB

11.0K 139

Apodex 1.1 Mini is a 36.0B-parameter open language model from apodex. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatFunctions

Gemma 2 9B IT

Google · 9.2B · runs from 3.0 GB

864.2K 954

Google Gemma 2 9B IT is a 9.2-billion parameter instruction-tuned model from Google's Gemma 2 series. It is a text-only chat model optimized for conversational tasks, instruction following, and general-purpose assistance. At release, it was recognized for delivering unusually strong performance relative to its parameter count. The model runs efficiently on consumer GPUs with 8-12GB of VRAM in quantized formats, making it accessible on mainstream hardware. It is a popular choice for local inference among users who want strong quality without the VRAM demands of larger models. Released under the Gemma license.

Chat

Ling 3.0 Flash

Inclusion AI · 127.5B · runs from 36.2 GB

16.7K 418

Ling 3.0 Flash is a 127.5-billion-parameter Mixture-of-Experts language model from InclusionAI, Ant Group's open-model initiative, with only about 5.6 billion parameters active per token, keeping per-token compute low even though all expert weights must still fit in memory. It is built for chat and agentic workloads such as tool use and multi-step tasks, blending Ant's fast "Flash" line with reasoning techniques from its "Ring" series. Local deployment needs multiple high-VRAM GPUs or a large unified-memory machine. It supports a 262K token context window for long documents and agent sessions. It is released under the MIT license, a permissive option for commercial and research use. Published in August 2026, it activates only a small fraction of its experts per token to approach larger-model quality at lower compute.

Chat

Phi 4 Mini Instruct

Microsoft · 3.8B · runs from 1.9 GB

379.2K 840

Microsoft Phi 4 Mini Instruct is a 3.8-billion parameter instruction-tuned model from Microsoft Research's Phi 4 family. It applies the Phi series' data-centric training philosophy to a compact model, delivering strong performance in coding, reasoning, and chat tasks relative to its small footprint. The model runs on consumer GPUs with as little as 4-6GB of VRAM when quantized, making it accessible on mainstream and even entry-level hardware. Released under the MIT license.

ChatCode