All LLM Models
Browse 1138 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Mistral 7B Instruct v0.2
Mistral AI · 7.2B · runs from 3.6 GB
Mistral 7B Instruct v0.2 is a 7.2-billion-parameter instruction-tuned language model from Mistral AI, built for general chat, question answering, and instruction following. It refines the original Mistral 7B with better adherence to complex prompts and improved handling of longer inputs. Its compact size suits local deployment on modern consumer GPUs, running comfortably on mainstream hardware once quantized. The model supports a 32,768 token context window, a result of raising the RoPE base frequency used for positional encoding, which improved handling of longer inputs versus the original v0.1. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in December 2023, it became one of the most widely adopted 7B open-weight chat models, still supported by tools such as llama.cpp, vLLM, and Ollama.
MiniCPM V 4 5
OpenBMB · 8.7B · runs from 4.3 GB
MiniCPM-V 4.5 is an 8.7-billion-parameter vision-language model from OpenBMB, combining a Qwen3-8B language backbone with a SigLIP2 vision encoder. It is designed to run efficiently on modest hardware, including phones and laptops, while handling image, multi-image, and video understanding alongside text chat; a unified resampler compresses video frames heavily so longer clips don't demand proportionally more compute. At under 9 billion parameters, it fits comfortably on a single mainstream consumer GPU once quantized. It supports a 40,960 token context window, sufficient for long documents or multi-turn multimodal conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in August 2025, it distinguishes itself with a switchable "fast" and "deep" thinking mode, trading response speed for more deliberate reasoning on harder problems.
Kimi Linear 48B A3B Instruct
Moonshot AI · 49.1B · runs from 14.3 GB
Kimi Linear 48B A3B Instruct is a 49.1B-parameter open language model from Moonshot AI in the Kimi family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 3n E2B IT
Google · 5.4B · runs from 1.6 GB
Gemma 3n E2B is the smaller of Google's two Gemma 3n instruction-tuned models, designed for phones, laptops and other on-device use. "E2B" stands for an effective size of about 2 billion parameters: the raw checkpoint holds around 5.4 billion, but techniques such as per-layer embeddings let it run with a memory footprint closer to a 2B model. It accepts text, image and audio input and generates text. The model has a 32K-token context window and is released under Google's Gemma terms of use. At 4-bit quantization it needs only a few gigabytes of memory, so it runs on almost any modern GPU, laptop or recent phone.
Qwen2.5 Math 7B Instruct
Alibaba · 7.6B · runs from 3.0 GB
Qwen2.5 Math 7B Instruct is a 7.6B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.6 27B AEON Ultimate Uncensored
AEON-7 · 27.4B · runs from 12.4 GB
Qwen3.6 27B AEON Ultimate Uncensored is a 27.4B-parameter open language model from AEON-7 in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 Coder 7B Instruct Abliterated
huihui-ai · 7.6B · runs from 3.0 GB
Qwen2.5 Coder 7B Instruct Abliterated is a 7.6B-parameter open language model from huihui-ai in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mistral 7B Instruct v0.1
Mistral AI · 7.2B · runs from 3.6 GB
Mistral 7B Instruct v0.1 was the first instruction-tuned variant of the original Mistral 7B, fine-tuned for conversational and instruction-following tasks. While it has since been superseded by v0.2 and v0.3, it remains a solid lightweight chat model and an important milestone in the open-weight model ecosystem. Its hardware requirements are identical to the base Mistral 7B, running smoothly on GPUs with as little as 6 GB of VRAM when quantized. Users seeking the best Mistral 7B experience should generally prefer the newer v0.3 release, but v0.1 is still useful for reproducibility and benchmarking purposes.
Gemma 4 12B IT Heretic
igorls · 12.0B · runs from 6.1 GB
Gemma 4 12B IT Heretic is a 12.0B-parameter open language model from igorls in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Huihui Qwen3 Coder 30B A3B Instruct Abliterated
huihui-ai · 30.5B · runs from 8.8 GB
Huihui Qwen3 Coder 30B A3B Instruct Abliterated is a 30.5B-parameter open language model from huihui-ai in the Qwen 3 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 E2B IT Qat Mobile Transformers
Google · 2.3B · runs from 1.4 GB
Gemma 4 E2B IT Qat Mobile Transformers is a 2.3B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 2 2B IT Abliterated
IlyaGusev · 2.6B · runs from 1.6 GB
Gemma 2 2B IT Abliterated is a 2.6B-parameter open language model from IlyaGusev in the Gemma 2 family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 3n E4B IT
Google · 7.8B · runs from 2.4 GB
Gemma 3n E4B IT is the instruction-tuned variant of Google's Gemma 3n E4B, a multimodal model built for on-device use on phones, laptops, and tablets that accepts text, image, audio, and video input and can perform automatic speech recognition and speech translation alongside text chat. It uses Google's MatFormer (Matryoshka Transformer) architecture, which nests a smaller sub-model inside the full network so the same checkpoint can run at reduced effective capacity, and Per-Layer Embedding, which caches embedding parameters to fast local storage instead of holding them all in memory. The "E4B" designation refers to an effective parameter count of around 4 billion at inference, even though the checkpoint's total parameters are larger; either way it is light enough for a single consumer GPU or even a high-end phone. Context length is 32,768 tokens. It is released under Google's Gemma Terms of Use, a custom license permitting broad commercial and research use alongside a prohibited-use policy, and was published in June 2025.
Qwen3.8 27B OBLITERATED
OBLITERATUS · 27.8B · runs from 12.6 GB
Qwen3.8 27B OBLITERATED is a 27.8B-parameter open language model from OBLITERATUS in the Qwen 3.8 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phonellm Alpha 1
pipecat-ai · 31.6B · runs from 9.1 GB
Phonellm Alpha 1 is a 31.6B-parameter open language model from pipecat-ai. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NeoHorse 1 4B
TokenRhythm · 4.2B · runs from 2.3 GB
NeoHorse 1 4B is a 4.2B-parameter open language model from TokenRhythm. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 E2B IT Uncensored
TrevorJS · 5.1B · runs from 2.5 GB
Gemma 4 E2B IT Uncensored is a 5.1B-parameter open language model from TrevorJS in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 VL 32B Instruct
Alibaba · 33.5B · runs from 10.0 GB
Qwen2.5 VL 32B Instruct is a 33.5-billion-parameter vision-language model from Alibaba's Qwen team, able to process images and text together for document parsing, chart reading, and visual question answering. It was tuned with reinforcement learning on top of the original Qwen2.5-VL release for more detailed, better-formatted answers and sharper accuracy on math and visual-logic problems. At this size, local inference needs quantization and a single high-end 24-32GB-class consumer or workstation GPU rather than lower-end hardware. The model supports a 128K token context window, enough for lengthy documents or multi-image inputs. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in March 2025 as a mid-sized addition to the Qwen2.5-VL lineup, between the smaller 7B model and the flagship 72B version.
Qwen2 VL 2B Instruct
Alibaba · 2.2B · runs from 1.1 GB
Qwen2 VL 2B Instruct is a 2.2-billion-parameter vision-language model from Alibaba's Qwen2-VL series, able to process images, multi-image comparisons, and video alongside text prompts. It targets visual question answering, document and chart reading, and basic agentic tasks such as interpreting a screenshot to plan a next action. Its small size makes it well suited to laptops and even some phones, running comfortably on modest consumer hardware once quantized. The model supports a 32K token context window, enough for moderate documents or extended chat. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in August 2024. Qwen2-VL introduced Naive Dynamic Resolution and Multimodal Rotary Position Embedding, letting it handle arbitrary image resolutions and understand videos well over twenty minutes long.
Qwen2.5 Coder 0.5B Instruct
Alibaba · 494M · runs from 0.5 GB
Qwen2.5 Coder 0.5B Instruct is a 494M-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 Math 1.5B Instruct
Alibaba · 1.5B · runs from 1 GB
Qwen2.5 Math 1.5B Instruct is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GLM 4.7 Flash REAP 23B A3B
Cerebras · 23.0B · runs from 7.4 GB
GLM 4.7 Flash REAP 23B A3B is a 23.0B-parameter open language model from Cerebras in the GLM 4 family. It supports a context window of up to 202,752 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NeoHorse 1 9B
TokenRhythm · 9.0B · runs from 4.4 GB
NeoHorse 1 9B is a 9.0B-parameter open language model from TokenRhythm. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mistral Small 3.1 24B Instruct 2503
Mistral AI · 24.0B · runs from 7.3 GB
Mistral Small 3.1 24B Instruct 2503 is a 24-billion-parameter model from Mistral AI, the French AI lab, built on the earlier text-only Mistral Small 3 with added support for image input alongside text. It can reason about images in the same conversation as written prompts, useful for document understanding and multimodal chat. At 24 billion parameters, it needs quantization and a single high-end 24GB-class consumer or workstation GPU for local inference rather than budget hardware. The model supports a 128K token context window for long documents or extended conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in March 2025, it added vision understanding and a longer context window to the earlier text-only Mistral Small while keeping the same 24B parameter budget.
Nemotron 3 Nano Omni 30B A3B Reasoning BF16
NVIDIA · 33.0B · runs from 10.0 GB
Nemotron 3 Nano Omni 30B A3B Reasoning BF16 is a 33.0B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
SmolVLM2 2.2B Instruct
Hugging Face · 2.2B · runs from 0.7 GB
SmolVLM2-2.2B Instruct is Hugging Face's 2.2-billion-parameter vision-language model, the largest member of the SmolVLM2 family, built to process interleaved text, images, and video within a single conversation. Beyond captioning and visual question answering, it is trained on additional video-instruction data for tasks like summarizing clips or answering questions about video content. At this size it still runs on a single consumer GPU, useful for on-device, resource-constrained deployment. The model supports a context window of 8,192 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in February 2025 alongside smaller 500M and 256M SmolVLM2 siblings. Unlike those variants, which favor raw efficiency, the 2.2B model is tuned to be the most capable of the three on both image and video benchmarks.
SmolLM2 360M Instruct
Hugging Face · 362M · runs from 0.5 GB
SmolLM2 360M Instruct is an instruction-tuned model from Hugging Face that occupies the sweet spot between the 135M and 1.7B entries in the SmolLM2 lineup. At 360 million parameters, it offers noticeably better coherence and instruction-following ability than the smallest variants while still running comfortably on virtually any modern GPU or even on CPU. This model is well suited for on-device assistants, embedded applications, and rapid prototyping where you need real conversational ability without dedicating significant hardware resources. It handles short-form generation, summarization, and basic reasoning tasks with reasonable quality.
MiniCPM V 4.6
OpenBMB · 1.3B · runs from 0.9 GB
MiniCPM-V 4.6 is OpenBMB's 1.3-billion-parameter vision-language model, pairing a SigLIP2 image encoder with a small Qwen3.5-based backbone to read images, multi-image sequences, and video alongside text. It targets on-device and edge deployment above all else, with adapted builds for iOS, Android, and HarmonyOS, small enough to run on phones and any modern laptop without heavy hardware, quantized or not. The model offers a 262K token context window, unusually long for its size. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in April 2026, introducing visual token compression that cuts visual-encoding computation by more than half versus earlier MiniCPM-V releases.
Cydonia 24B V4.3
TheDrummer · 23.6B · runs from 7.8 GB
Cydonia 24B V4.3 is a 23.6B-parameter open language model from TheDrummer. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 Omni 3B
Alibaba · 5.5B · runs from 1.7 GB
Qwen2.5-Omni-3B is Alibaba's 5.5-billion-parameter end-to-end omni-modal model, built to perceive text, images, audio, and video and respond with both text and natural streaming speech. It uses a Thinker-Talker architecture with a time-aligned position embedding (TMRoPE) that synchronizes video frames with audio, enabling real-time voice and video chat with chunked input and immediate output. As the smaller of the two Omni checkpoints, it runs on a single consumer GPU once quantized. Its language backbone supports a 32,768 token context window. It is released under Alibaba's Qwen Research License, which permits only non-commercial research and evaluation, unlike the Apache 2.0 used for Qwen's text-only models. Published in April 2025, about a month after the larger Qwen2.5-Omni-7B, it brings real-time multimodal chat to smaller-scale deployments.