All LLM Models

Browse 1242 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Qwen2 VL 2B Instruct

Alibaba · 2.2B · runs from 1.1 GB

2.0M 520

Qwen2 VL 2B Instruct is a 2.2-billion-parameter vision-language model from Alibaba's Qwen2-VL series, able to process images, multi-image comparisons, and video alongside text prompts. It targets visual question answering, document and chart reading, and basic agentic tasks such as interpreting a screenshot to plan a next action. Its small size makes it well suited to laptops and even some phones, running comfortably on modest consumer hardware once quantized. The model supports a 32K token context window, enough for moderate documents or extended chat. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in August 2024. Qwen2-VL introduced Naive Dynamic Resolution and Multimodal Rotary Position Embedding, letting it handle arbitrary image resolutions and understand videos well over twenty minutes long.

Vision

Mixtral 8x7B Instruct v0.1

Mistral AI · 46.7B · runs from 20.4 GB

217.4K 4.8K

Mixtral 8x7B Instruct v0.1 is Mistral AI's flagship Mixture-of-Experts model, combining eight expert networks of 7 billion parameters each for a 46.7B total weight count while activating only about 12.9 billion parameters per token. This sparse architecture delivers performance that rivals much larger dense models at a fraction of the inference cost, excelling across reasoning, code generation, and multilingual tasks. Because the full weights must still be loaded into memory, you will need around 24–48 GB of VRAM depending on quantization level, making it best suited for multi-GPU desktop setups or high-VRAM workstation cards. If your hardware can accommodate it, Mixtral offers one of the best performance-per-active-parameter ratios available for local deployment.

Chat

Qwen2.5 Coder 0.5B Instruct

Alibaba · 494M · runs from 0.5 GB

140.1K 78

Qwen2.5 Coder 0.5B Instruct is a 494M-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Qwen2.5 Math 1.5B Instruct

Alibaba · 1.5B · runs from 1 GB

321.9K 58

Qwen2.5 Math 1.5B Instruct is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatMath

GLM 4.7 Flash REAP 23B A3B

Cerebras · 23.0B · runs from 7.4 GB

1.1K 78

GLM 4.7 Flash REAP 23B A3B is a 23.0B-parameter open language model from Cerebras in the GLM 4 family. It supports a context window of up to 202,752 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

NeoHorse 1 9B

TokenRhythm · 9.0B · runs from 4.4 GB

13.0K 1.0K

NeoHorse 1 9B is a 9.0B-parameter open language model from TokenRhythm. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatFunctionsReasoning

Mistral Small 3.1 24B Instruct 2503

Mistral AI · 24.0B · runs from 7.3 GB

483.4K 1.4K

Mistral Small 3.1 24B Instruct 2503 is a 24-billion-parameter model from Mistral AI, the French AI lab, built on the earlier text-only Mistral Small 3 with added support for image input alongside text. It can reason about images in the same conversation as written prompts, useful for document understanding and multimodal chat. At 24 billion parameters, it needs quantization and a single high-end 24GB-class consumer or workstation GPU for local inference rather than budget hardware. The model supports a 128K token context window for long documents or extended conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in March 2025, it added vision understanding and a longer context window to the earlier text-only Mistral Small while keeping the same 24B parameter budget.

Chat

Nemotron 3 Nano Omni 30B A3B Reasoning BF16

NVIDIA · 33.0B · runs from 10.0 GB

340.0K 343

Nemotron 3 Nano Omni 30B A3B Reasoning BF16 is a 33.0B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Reasoning

SmolVLM2 2.2B Instruct

Hugging Face · 2.2B · runs from 0.7 GB

139.3K 335

SmolVLM2-2.2B Instruct is Hugging Face's 2.2-billion-parameter vision-language model, the largest member of the SmolVLM2 family, built to process interleaved text, images, and video within a single conversation. Beyond captioning and visual question answering, it is trained on additional video-instruction data for tasks like summarizing clips or answering questions about video content. At this size it still runs on a single consumer GPU, useful for on-device, resource-constrained deployment. The model supports a context window of 8,192 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in February 2025 alongside smaller 500M and 256M SmolVLM2 siblings. Unlike those variants, which favor raw efficiency, the 2.2B model is tuned to be the most capable of the three on both image and video benchmarks.

Vision

SmolLM2 360M Instruct

Hugging Face · 362M · runs from 0.5 GB

306.9K 222

SmolLM2 360M Instruct is an instruction-tuned model from Hugging Face that occupies the sweet spot between the 135M and 1.7B entries in the SmolLM2 lineup. At 360 million parameters, it offers noticeably better coherence and instruction-following ability than the smallest variants while still running comfortably on virtually any modern GPU or even on CPU. This model is well suited for on-device assistants, embedded applications, and rapid prototyping where you need real conversational ability without dedicating significant hardware resources. It handles short-form generation, summarization, and basic reasoning tasks with reasonable quality.

Chat

Kimi Dev 72B

Moonshot AI · 72.7B · runs from 21.0 GB

7.1K 393

Kimi Dev 72B is Moonshot AI's developer-focused model built on the Qwen2.5-72B architecture, specifically optimized for coding tasks, tool use, and agentic workflows. It combines strong general-purpose chat abilities with specialized developer capabilities, making it a compelling choice for software engineering assistance. At 72 billion parameters it requires substantial hardware, typically needing 40+ GB of VRAM at 4-bit quantization, which puts it in reach of dual consumer GPU setups or single professional cards like the A100 or RTX 6000 Ada. If you are primarily looking for a local coding assistant with strong reasoning skills, Kimi Dev is a top-tier option in the 70B class.

ChatCode

MiniCPM V 4.6

OpenBMB · 1.3B · runs from 0.9 GB

401.1K 1.2K

MiniCPM-V 4.6 is OpenBMB's 1.3-billion-parameter vision-language model, pairing a SigLIP2 image encoder with a small Qwen3.5-based backbone to read images, multi-image sequences, and video alongside text. It targets on-device and edge deployment above all else, with adapted builds for iOS, Android, and HarmonyOS, small enough to run on phones and any modern laptop without heavy hardware, quantized or not. The model offers a 262K token context window, unusually long for its size. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in April 2026, introducing visual token compression that cuts visual-encoding computation by more than half versus earlier MiniCPM-V releases.

Vision

Cydonia 24B V4.3

TheDrummer · 23.6B · runs from 7.8 GB

4.2K 132

Cydonia 24B V4.3 is a 23.6B-parameter open language model from TheDrummer. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen2.5 Omni 3B

Alibaba · 5.5B · runs from 1.7 GB

327.9K 357

Qwen2.5-Omni-3B is Alibaba's 5.5-billion-parameter end-to-end omni-modal model, built to perceive text, images, audio, and video and respond with both text and natural streaming speech. It uses a Thinker-Talker architecture with a time-aligned position embedding (TMRoPE) that synchronizes video frames with audio, enabling real-time voice and video chat with chunked input and immediate output. As the smaller of the two Omni checkpoints, it runs on a single consumer GPU once quantized. Its language backbone supports a 32,768 token context window. It is released under Alibaba's Qwen Research License, which permits only non-commercial research and evaluation, unlike the Apache 2.0 used for Qwen's text-only models. Published in April 2025, about a month after the larger Qwen2.5-Omni-7B, it brings real-time multimodal chat to smaller-scale deployments.

Chat

Granite 4.0 H Tiny

IBM · 6.9B · runs from 2.4 GB

57.6K 209

Granite 4.0 H Tiny is a 6.9B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

OlmOCR 2 7B 1025

Allen AI · 8.3B · runs from 2.7 GB

204.9K 158

olmOCR-2-7B-1025 is Allen AI's 8.3-billion-parameter vision-language model, fine-tuned from Qwen2.5-VL-7B-Instruct specifically for document OCR rather than general chat. It converts scanned pages and PDFs into clean text, using GRPO reinforcement learning atop supervised fine-tuning to sharpen accuracy on math, tables, and tricky OCR cases. This is the full-precision BF16 release; Allen AI recommends its FP8 sibling for production use. At this size, it runs on a single mainstream-to-high-end consumer GPU once quantized. It supports a 128,000 token context window inherited from its Qwen2.5-VL base, useful for multi-page documents. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. Published in October 2025, it pairs with Allen AI's olmOCR toolkit, which handles page rendering and retries at scale via vLLM.

Vision

Deepseek Coder 6.7B Base

DeepSeek · 6.7B · runs from 3.2 GB

89.1K 128

Deepseek Coder 6.7B Base is a 6.7B-parameter open language model from DeepSeek in the DeepSeek Coder family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

MiniCPM5 2B

OpenBMB · 2.5B · runs from 1.2 GB

564.4K 1.7K

MiniCPM5-2B is the second model in OpenBMB's MiniCPM5 series, a dense 2.5-billion-parameter causal language model built on a Llama-style architecture. This release is the final version, post-trained with reinforcement learning and on-policy distillation, and is aimed at local assistants, coding agents, tool-use workflows, and reasoning tasks where a small footprint matters, including on-device and edge deployment. It supports a 131,072-token context window, unusually long for its size, and ships under the Apache 2.0 license, allowing free commercial and non-commercial use. At roughly 2.5 billion parameters, it needs well under 2GB of memory at 4-bit quantization, so it runs comfortably on almost any modern GPU or laptop.

Chat

GLM 4.5 Air

Z.ai · 110.5B · runs from 30.8 GB

80.9K 639

GLM-4.5-Air is Z.ai's more compact model in the GLM-4.5 series, a foundation model built for intelligent agents with 106 billion total parameters and about 12 billion active via its Mixture-of-Experts design. Like its larger sibling GLM-4.5, it offers both a thinking mode for complex reasoning and tool use and a non-thinking mode for quick replies, unifying reasoning, coding, and agentic capabilities in one checkpoint. It supports a 128K token context window and is released under the MIT license for commercial and research use. With around 106 billion total parameters, running it locally at 4-bit needs roughly 60 GB of memory, putting it in reach of a multi-GPU workstation or a high-memory Mac rather than a single card.

Chat

Medgemma 1.5 4B IT

Google · 4.3B · runs from 1.3 GB

326.4K 920

MedGemma 1.5 4B is Google's 4.3-billion-parameter instruction-tuned vision-language model in the MedGemma line of health-AI foundation models, succeeding the original MedGemma 4B at the same size. It is built for developers creating healthcare applications, covering tasks such as medical image interpretation and clinical text understanding, and is not a validated diagnostic tool: outputs need independent verification and further evaluation before any clinical use. At this size it runs on a single consumer GPU once quantized. It is released under Google's Health AI Developer Foundations terms, a use-restricted license that requires accepting specific health-AI conditions on Hugging Face rather than a fully open one. Published in January 2026, it is the second 4B release in the MedGemma series.

Vision

Qwen3.8 27B DSpark

RadixArk · 27B · runs from 11.8 GB

322.9K 85

Qwen3.8 27B DSpark is a 27B-parameter open language model from RadixArk in the Qwen 3.8 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Hemmingway 1

Altworld · 26.9B · runs from 8.1 GB

3.8K 577

Hemmingway 1 is a 26.9B-parameter open language model from Altworld. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

NVIDIA Nemotron Nano 9B v2 Japanese

NVIDIA · 8.9B · runs from 4.4 GB

281.4K 124

NVIDIA Nemotron Nano 9B v2 Japanese is a specialized variant of the Nemotron Nano 9B v2, fine-tuned for Japanese language understanding and generation. At 8.9 billion parameters, it maintains the same hardware-friendly footprint as the English version while delivering natural Japanese conversational ability. For users looking to run a Japanese-language assistant locally, this model offers a rare combination of compact size and dedicated language optimization from a major hardware vendor. It handles Japanese text with the fluency you'd expect from a purpose-built model rather than a multilingual afterthought.

Chat

Deepseek Coder 33B Instruct

DeepSeek · 33.3B · runs from 14.6 GB

5.4K 584

DeepSeek-Coder-33B-Instruct is DeepSeek's 33.3-billion-parameter instruction-tuned code model, initialized from DeepSeek-Coder-33B-Base and further fine-tuned on 2 billion tokens of instruction data for chat-style code generation, completion, and fixing. The underlying Deepseek Coder family was pretrained from scratch on 2 trillion tokens (87% code, 13% natural language in English and Chinese) with project-level context and a fill-in-the-blank training objective, giving it strong project-level completion and infilling ability alongside state-of-the-art open-model results on HumanEval, MBPP, and related coding benchmarks at the time of release. At 33.3 billion parameters, it needs a high-end consumer GPU or a multi-GPU setup once quantized. Context length is 16,384 tokens. Model weights are released under DeepSeek's custom model license, which permits commercial use. It was published in November 2023, alongside 1.3B, 5.7B, and 6.7B siblings and the corresponding base checkpoint.

ChatCode

Qwen3 VL 8B Thinking

Alibaba · 8.8B · runs from 3.0 GB

158.2K 224

Qwen3 VL 8B Thinking is an 8.8-billion-parameter vision-language model from Alibaba's Qwen team, the reasoning-focused counterpart to the Qwen3-VL-8B Instruct model. It processes images, video, and text together and is tuned to work through problems step by step before answering, which tends to help on multi-step visual reasoning, STEM problems, and chart or document analysis. At under 9 billion parameters, it runs well on a single mainstream consumer GPU once quantized, making local multimodal use practical without server-class hardware. It supports a 262K token context window for long documents, multi-image input, or extended video. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. Published in October 2025 alongside the non-thinking Instruct variant, it trades some response speed for deeper visual-language reasoning.

Vision

Medgemma 4B IT

Google · 4.3B · runs from 1.3 GB

1.1M 1.1K

MedGemma 4B is Google's 4.3-billion-parameter instruction-tuned vision-language model, built on Gemma 3 and further trained for medical text and image understanding. It adds a medical image encoder covering chest X-rays, dermatology photos, histopathology slides, and fundus images, and can generate findings and answer questions about medical images. It is a foundation for developers building healthcare applications, not a validated diagnostic tool, and needs further evaluation before clinical use. It runs on a single consumer GPU once quantized. It inherits Gemma 3's 128,000 token context window. It is released under Google's Health AI Developer Foundations terms, a use-restricted license rather than a fully open one. Published in May 2025, it was Google's first instruction-tuned multimodal MedGemma release, alongside a larger text-only 27B variant.

Vision

Devstral Small 2507

Mistral AI · 23.6B · runs from 7.2 GB

72.3K 368

Devstral Small 2507 is Mistral AI's agentic coding model, developed with All Hands AI and fine-tuned from the 24-billion-parameter Mistral Small 3.1 with its vision encoder removed to keep it text-only. It is built to explore codebases, edit multiple files, and drive software-engineering agents, using Mistral's function-calling format and a Tekken tokenizer with a 131K-token vocabulary. It supports a 128K token context window and is released under the Apache 2.0 license. At 24 billion parameters, Devstral is light enough to run on a single RTX 4090 or a Mac with around 32 GB of unified memory once quantized to 4-bit, making it practical for local coding-agent setups.

Chat

Phi 2

Microsoft · 2.8B · runs from 2.1 GB

586.8K 3.5K

Microsoft Phi 2 is a 2.8-billion parameter language model from Microsoft Research that pioneered the concept of small but highly capable language models. Released in late 2023, Phi 2 demonstrated that strategic data curation and training methodology could allow a sub-3B model to outperform many 7B and 13B models on reasoning and coding benchmarks. The model runs on virtually any modern GPU and even on CPU-only setups. While succeeded by Phi 3 and Phi 4, Phi 2 remains historically significant as the model that proved small-scale language models could be genuinely useful for practical tasks. Released under the MIT license.

ChatCode

DeepSeek R1 Distill Qwen 32B Abliterated

huihui-ai · 32.8B · runs from 9.8 GB

1.1K 253

DeepSeek R1 Distill Qwen 32B Abliterated is a 32.8B-parameter open language model from huihui-ai in the DeepSeek R1 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Gemma 4 26B A4B IT Uncensored Heretic

llmfan46 · 25.8B · runs from 11.6 GB

2.1K 12

Gemma 4 26B A4B IT Uncensored Heretic is a 25.8B-parameter open language model from llmfan46 in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision