All LLM Models
Browse 1214 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Carnice V1 9B Hermes Agent Stage2 Merged
kai-os · 9.0B · runs from 4.4 GB
Carnice V1 9B Hermes Agent Stage2 Merged is a 9.0B-parameter open language model from kai-os in the Hermes family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
EuroLLM 1.7B Instruct
utter-project · 1.7B · runs from 1.2 GB
EuroLLM 1.7B Instruct is a 1.7B-parameter open language model from utter-project. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Dobby Mini Unhinged Llama 3.1 8B
SentientAGI · 8.0B · runs from 2.8 GB
Dobby Mini Unhinged Llama 3.1 8B is a 8.0B-parameter open language model from SentientAGI in the Llama 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Athene v2 Chat
Nexusflow · 72.7B · runs from 21.0 GB
Athene-V2-Chat-72B is Nexusflow's 72-billion-parameter chat model, fine-tuned from Qwen2.5-72B-Instruct through reinforcement learning from human feedback (RLHF) to specialize in chat, math, and coding. On Chatbot Arena the card reports it beating GPT-4o-0513 in the hard-prompts and math categories and matching it on coding, instruction-following, and multi-turn chat. A sister model, Athene-V2-Agent-72B, is fine-tuned separately for function calling and agentic tasks. At 72B parameters, it needs a multi-GPU workstation or server to run, even quantized. Context length is 32,768 tokens. It is released under the Nexusflow Research License, a custom license restricted to personal, non-profit, non-commercial use, with commercial use requiring separate permission from Nexusflow, and was published in November 2024.
Qwen2 7B
Alibaba · 7.6B · runs from 3.6 GB
Qwen2 7B is a 7.6B-parameter open language model from Alibaba in the Qwen 2 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Meta Llama 3 70B Instruct
Meta · 70.6B · runs from 23.3 GB
Meta Llama 3 70B Instruct is a 70.6-billion parameter instruction-tuned model from Meta's Llama 3 release. It is fine-tuned for dialogue, coding assistance, and complex reasoning tasks using supervised fine-tuning and RLHF. At the time of release, it was among the most capable openly available models. The model supports an 8K token context window and requires substantial VRAM for local inference, typically needing multi-GPU setups or high-VRAM professional GPUs. It has been widely adopted for local deployment in quantized formats. Released under the Meta Llama 3 Community License.
Qwen2 72B Instruct
Alibaba · 72.7B · runs from 21.0 GB
Qwen2-72B-Instruct is the instruction-tuned, chat-ready 72.7-billion-parameter model from Alibaba's Qwen2 series, built on a Transformer architecture with SwiGLU activation, QKV attention bias, and grouped-query attention, and trained with supervised fine-tuning plus direct preference optimization. It generally surpassed the prior Qwen1.5 line and competed with proprietary models on language understanding, generation, multilingual tasks, coding, and math benchmarks at release, though it has since been superseded by Qwen2.5-72B-Instruct. At 72.7 billion parameters, it needs a multi-GPU workstation to run in full precision. Context length is natively 32,768 tokens, extendable to 131,072 tokens using YaRN scaling, as documented on the model card. It is released under Alibaba's Tongyi Qianwen custom license, which permits commercial use but requires a separate license once a deployment passes 100 million monthly active users. It was published in May 2024.
Eurus 2 7B PRIME
PRIME-RL · 7.6B · runs from 3.0 GB
Eurus-2-7B-PRIME is a 7.6-billion-parameter reasoning model from the PRIME-RL project, built on Qwen2.5-Math-7B-Base and trained with PRIME (Process Reinforcement through Implicit Rewards), an open-source online reinforcement-learning method that scores intermediate reasoning steps rather than only final answers. It starts from the Eurus-2-7B-SFT checkpoint and is trained further on the Eurus-2-RL-Data set, focused on math and coding problem-solving. The card reports a 16.7% average improvement over the SFT starting point, with over 20% gains on AMC and AIME competition math, enough to surpass the larger Qwen2.5-Math-7B-Instruct on several reasoning benchmarks. At this size it fits on a single consumer GPU. Context length is 4,096 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in December 2024.
InternVL3 8B
OpenGVLab · 7.9B · runs from 2.4 GB
InternVL3-8B is OpenGVLab's roughly 7.9-billion-parameter vision-language model, pairing an InternViT-300M vision encoder with a Qwen2.5-7B language backbone in a ViT-MLP-LLM architecture. It handles general image and video understanding and document analysis, extending into tool use, GUI agent tasks, and 3D scene perception beyond typical captioning. It is comfortably runnable on a single mainstream-to-high-end consumer GPU once quantized. Its language backbone supports a 32,768 token context window. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in April 2025 as part of a 1B-to-78B InternVL3 family sharing the same vision encoder. Its key change versus InternVL2.5 is Native Multimodal Pre-Training, which trains vision and language jointly from the start instead of adapting a language-only model afterward.
LFM2.5 1.2B JP 202606
Liquid AI · 1.2B · runs from 0.9 GB
LFM2.5 1.2B JP 202606 is a 1.2B-parameter open language model from Liquid AI in the LFM2.5 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Saiga Llama3 8B
IlyaGusev · 8.0B · runs from 4.0 GB
Saiga Llama3 8B is a 8.0B-parameter open language model from IlyaGusev in the Llama 3 family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Yi 9B
01.AI · 8.8B · runs from 4.1 GB
Yi-9B is 01.AI's 8.8-billion-parameter base language model, continuously pretrained from Yi-6B on an additional 0.8 trillion tokens as part of the bilingual English/Chinese Yi series (trained on 3 trillion tokens overall). It is a pretrained model, not instruction-tuned; a separate long-context Yi-9B-200K variant exists for extended-context use. The card reports it as the strongest model in its size class among Mistral-7B, SOLAR-10.7B, Gemma-7B, and DeepSeek-Coder-7B-Base, particularly in code, math, common-sense reasoning, and reading comprehension. It fits on a single consumer GPU. Context length is 4,096 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in March 2024.
Pythia 1B
EleutherAI · 1.1B · runs from 0.5 GB
Pythia 1B is a 1.1B-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gpt2 Medium
OpenAI · 380M · runs from 0.2 GB
GPT-2 Medium scales the original GPT-2 architecture to 380 million parameters, offering noticeably improved text generation quality over the base 137M variant while remaining extremely lightweight by current standards. It supports the same autoregressive language modeling tasks as its smaller and larger siblings. Like all GPT-2 variants, it runs comfortably on virtually any modern hardware including CPU-only setups, making it an accessible option for learning, prototyping, and lightweight text generation experiments without needing a dedicated GPU.
Granite 4.0 Tiny Preview
IBM · 6.7B · runs from 2.7 GB
Granite 4.0 Tiny Preview is a 6.7B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Granite Guardian 3.3 8B
IBM · 8.2B · runs from 2.9 GB
Granite Guardian 3.3 8B is a 8.2B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
PaddleOCR VL 1.6
PaddlePaddle · 959M · runs from 0.6 GB
PaddleOCR-VL-1.6 is PaddlePaddle's compact, roughly 0.9-billion-parameter vision-language model for document parsing, built on the ERNIE 4.5 line and specialized for OCR, table, formula, chart, and seal/stamp recognition plus text spotting rather than open-domain chat. It upgrades PaddleOCR-VL-1.5 with a region-aware data optimization framework that targets the earlier model's weak spots and a progressive post-training recipe combining curated data selection with reinforcement learning, while staying architecture-compatible with 1.5 for drop-in migration. The card reports a new state-of-the-art 96.33% on OmniDocBench v1.6. At under a billion parameters it runs on a single modest consumer GPU. Context length is 131,072 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2026.
C4ai Command R V01
Cohere · 35.0B · runs from 15.9 GB
C4AI Command R is a 35-billion-parameter instruction-tuned model released by Cohere For AI in March 2024. It was built for conversational use with a particular focus on retrieval-augmented generation and tool use, including grounded answers that cite the source documents they draw on, and it was trained to work across ten major languages. It supports a 128K-token context window and is released under a Creative Commons Attribution-NonCommercial 4.0 license, so commercial use is not permitted. At roughly 35 billion parameters, 4-bit quantization needs around 20GB of memory, within reach of a single 24GB consumer GPU.
WhiteRabbitNeo 13B V1
WhiteRabbitNeo · 13B · runs from 7.5 GB
WhiteRabbitNeo 13B V1 is a 13B-parameter open language model from WhiteRabbitNeo. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Laguna XS.2
poolside · 33.4B · runs from 14.6 GB
Laguna XS.2 is a 33.4B-parameter open language model from poolside. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
CodeLlama 7B HF
Meta · 6.7B · runs from 4.2 GB
CodeLlama 7B HF is a 6.7B-parameter open language model from Meta in the Code Llama family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
SmolLM 135M
Hugging Face · 135M · runs from 0.4 GB
SmolLM 135M is the original first-generation small language model from Hugging Face, designed to push the boundaries of what is achievable at extremely low parameter counts. With just 135 million parameters, it was a pioneering effort in making capable language models accessible on the most resource-constrained hardware. While the SmolLM2 and SmolLM3 families have since surpassed it in quality, the original SmolLM 135M remains a useful reference point for research and a practical option for ultra-lightweight deployment scenarios where every megabyte of memory counts.
Pythia 160M
EleutherAI · 213M · runs from 0.1 GB
Pythia 160M is part of EleutherAI's Pythia training suite, a collection of models trained on the same data in the same order at multiple scales to enable rigorous scientific research into how language models learn. At 160 million parameters, it is the smallest model in the suite and runs on virtually any hardware. This model is primarily valuable for researchers studying scaling laws, training dynamics, and emergent capabilities across model sizes. EleutherAI released full training checkpoints, data, and code, making Pythia 160M one of the most transparent and reproducible models available for academic study.
Granite 3.0 1B A400m Instruct
IBM · 1.3B · runs from 1.0 GB
Granite 3.0 1B A400m Instruct is a 1.3B-parameter open language model from IBM in the Granite family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mamba 130M HF
State Spaces · 129M · runs from 0.1 GB
Mamba 130M is a state-space model developed by State Spaces that offers a fundamentally different architecture from the Transformer-based models that dominate the LLM landscape. Using selective state-space layers instead of attention, Mamba achieves linear-time inference scaling with sequence length, making it particularly efficient for processing long inputs. At 130 million parameters this is primarily a research and demonstration model, but it showcases the potential of state-space architectures for local deployment. Users interested in exploring alternatives to Transformer-based language models will find Mamba 130M a lightweight and accessible entry point for experimentation.
Cosmos Reason2 8B
NVIDIA · 8.8B · runs from 4.1 GB
Cosmos Reason2-8B is NVIDIA's 8.8-billion-parameter open reasoning vision-language model for physical AI, built on a Qwen3-VL-8B-Instruct backbone and tuned to reason step by step about video and images the way a human would when planning actions in the real world. Rather than just labeling objects, it applies physics, spatio-temporal understanding, and common sense to tasks like robot planning, autonomous-vehicle video captioning, and video-analytics annotation, producing structured outputs such as 2D/3D point localization, bounding boxes, trajectory coordinates, and on-screen OCR text. It ships alongside a smaller 2B variant for edge deployment, while the 8B model needs a capable single GPU or more, less once quantized. Context length is roughly 256,000 tokens, up sharply from 16,000 tokens in the original Cosmos Reason 1. It is released under the NVIDIA Open Model License, a custom license that permits commercial use and derivative models but requires attribution ("Built on NVIDIA Cosmos") and prohibits removing its safety guardrails. It was published in December 2025.
OLMoE 1B 7B 0125 Instruct
Allen AI · 6.9B · runs from 2.5 GB
OLMoE 1B 7B 0125 Instruct is a 6.9B-parameter open language model from Allen AI in the OLMo family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
JiRackUltra 7B
CMSManhattan · 7.6B · runs from 2.5 GB
JiRackUltra 7B is a 7.6B-parameter open language model from CMSManhattan. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Olmo 3 1025 7B
Allen AI · 7.3B · runs from 3.4 GB
Olmo 3 1025 7B is a 7.3B-parameter open language model from Allen AI in the OLMo family. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Yi 6B Chat
01.AI · 6.1B · runs from 2.9 GB
Yi-6B-Chat is 01.AI's 6-billion-parameter bilingual (English/Chinese) chat model, instruction-tuned from the Yi-6B base model, part of the first-generation Yi series trained from scratch on a 3-trillion-token multilingual corpus. It uses the same Transformer structure popularized by Llama, though 01.AI states it is an independently trained model rather than a Llama derivative, and it was competitive with much larger contemporaries on benchmarks like the Hugging Face Open LLM Leaderboard and C-Eval at release. At 6B parameters it fits a single consumer GPU in half precision, and a much smaller card once 4-bit or 8-bit quantized. Context length is 4,096 tokens; a separate 200K-context variant of the base model is also available for longer documents. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in November 2023.