All LLM Models
Browse 1214 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Granite 4.1 3B
IBM · 3.4B · runs from 1.9 GB
Granite 4.1 3B is a 3.4B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma4 12B QAT Uncensored HauhauCS Balanced
HauhauCS · 12B · runs from 5.6 GB
Gemma4 12B QAT Uncensored HauhauCS Balanced is a 12B-parameter open language model from HauhauCS in the Gemma 4 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GPT OSS 20B NPU2
FastFlowLM · 20B · runs from 8.9 GB
GPT OSS 20B NPU2 is a 20B-parameter open language model from FastFlowLM in the GPT-OSS family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Vlt5 Base Keywords
Voicelab · 275M · runs from 0.6 GB
Vlt5 Base Keywords is a 275M-parameter open language model from Voicelab. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.8 27B DFlash2
incoai · 27B · runs from 11.8 GB
Qwen3.8 27B DFlash2 is a 27B-parameter open language model from incoai in the Qwen 3.8 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GPT J 6B
EleutherAI · 6B · runs from 2.8 GB
GPT J 6B is a 6B-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GigaChat3.1 Audio 10B A1.8B
ai-sage · 10B · runs from 20.6 GB
GigaChat3.1 Audio 10B A1.8B is a 10B-parameter open language model from ai-sage. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Xflux Text Encoders
XLabs-AI · 4.8B · runs from 10.5 GB
Xflux Text Encoders is a 4.8B-parameter open language model from XLabs-AI. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 4B Base Dapo V4
ReliquaryForge · 4.0B · runs from 2.2 GB
Qwen3 4B Base Dapo V4 is a 4.0B-parameter open language model from ReliquaryForge in the Qwen 3 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Apertus V1.5 8B
swiss-ai · 8.9B · runs from 4.2 GB
Apertus V1.5 8B is an 8.9-billion-parameter multimodal model from the Swiss AI Initiative, a collaboration between ETH Zurich, EPFL, and the Swiss National Supercomputing Centre. Extended from the text-only Apertus 1 through continued pretraining, it accepts images alongside text and can process spoken audio experimentally, letting it reason about visual content in a chat setting. Its compact size suits local deployment on a single consumer GPU once quantized. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, alongside a public acceptable-use policy governing deployment. Published in July 2026, Apertus V1.5 stands out for being fully open in a way few large models are: not just the weights but the training data, code, and development process are published, with an optional reasoning mode for harder prompts.
Opt 350M
Meta · 350M · runs from 0.8 GB
Meta OPT 350M is a 350-million parameter language model from Meta's Open Pre-trained Transformer (OPT) project, released in 2022 as part of a suite of models ranging from 125M to 175B parameters. It was designed to provide researchers with open access to models comparable to GPT-3 at various scales. The 350M variant runs on minimal hardware and is suitable for research, prototyping, and educational use. While it has been surpassed by modern architectures in terms of capability, it remains a lightweight option for basic text generation experiments and as a benchmark baseline.
DeepSeek v2 Lite
DeepSeek · 15.7B · runs from 7.4 GB
DeepSeek V2 Lite is a compact mixture-of-experts model with 15.7 billion total parameters, designed to deliver a strong quality-to-compute ratio for general chat and instruction following. It uses the same innovative MLA (Multi-Head Latent Attention) architecture as the larger V2, which reduces memory requirements during inference. With its modest parameter count, V2 Lite runs comfortably on a single consumer GPU, making it accessible to users who want to try DeepSeek's MoE approach without needing specialized hardware. It handles everyday conversational tasks, summarization, and light analysis well, offering a practical entry point into the DeepSeek model family.
Openai GPT
OpenAI · 120M · runs from 0.1 GB
OpenAI GPT is the original 2018 transformer-based language model that started the GPT lineage, based on the paper "Improving Language Understanding by Generative Pre-Training." At just 120 million parameters, it is a historically significant model that demonstrated the power of unsupervised pretraining followed by supervised fine-tuning. This model is primarily of academic and historical interest today. It runs on essentially any hardware and can be useful for educational exploration of transformer architectures, but it should not be compared to modern instruction-tuned models in terms of practical capability.
The GuageLLM 23M
Hai929 · 23M · runs from 0.0 GB
The GuageLLM 23M is a 23M-parameter open language model from Hai929. It supports a context window of up to 64 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 7B
huggyllama · 6.7B · runs from 3.1 GB
This is a community reupload of Meta's original Llama 1 7B model, published by the huggyllama account on Hugging Face. The original Llama 1 was a 6.7-billion parameter base model released by Meta in early 2023, trained on 1 trillion tokens of publicly available data. It pioneered the wave of open-weight large language models. As a first-generation Llama model, it has been superseded by Llama 2 and Llama 3 in terms of quality and capability. It remains of historical and research interest as the model that catalyzed the open-source LLM ecosystem. This upload provides convenient access in Hugging Face Transformers format.
LLaDA2.0 Mini
Inclusion AI · 16.3B · runs from 7.3 GB
LLaDA2.0 Mini is a 16.3B-parameter open language model from Inclusion AI. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Hyenadna Medium 450k Seqlen HF
LongSafari · 28M · runs from 0.1 GB
Hyenadna Medium 450k Seqlen HF is a 28M-parameter open language model from LongSafari. It supports a context window of up to 450,002 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
TinyStories 1M
roneneldan · 1M · runs from 0.0 GB
TinyStories 1M is a 1M-parameter open language model from roneneldan. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Zamba2 1.2B Instruct
Zyphra · 1.2B · runs from 1.4 GB
Zamba2 1.2B Instruct is a 1.2B-parameter open language model from Zyphra. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Granite 4.1 8B
IBM · 8.8B · runs from 4.4 GB
Granite-4.1-8B is IBM's 8-billion-parameter long-context instruct model, fine-tuned from Granite-4.1-8B-Base on a mix of permissively licensed open datasets and internally collected synthetic data. An improved post-training pipeline of supervised fine-tuning and reinforcement learning gives it stronger tool calling, instruction following, and chat behavior, and IBM positions it for business assistants, retrieval-augmented generation, summarization, classification, and code tasks including fill-in-the-middle completions, across a dozen supported languages. It offers a 131,072-token context window and is released under the Apache 2.0 license. At around 8.8 billion parameters, 4-bit quantization needs roughly 5GB of memory, so it fits comfortably on an 8-12GB consumer GPU.
CAJAL 4B
Agnuxo · 4B · runs from 1.9 GB
CAJAL 4B is a 4B-parameter open language model from Agnuxo. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
LLaMmlein 1B Prerelease
LSX-UniWue · 1.1B · runs from 0.8 GB
LLaMmlein 1B Prerelease is a 1.1B-parameter open language model from LSX-UniWue. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Wildguard
Allen AI · 7.2B · runs from 3.4 GB
Wildguard is a 7.2B-parameter open language model from Allen AI. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
VLM2Vec Full
TIGER-Lab · 4.1B · runs from 9.4 GB
VLM2Vec Full is a 4.1B-parameter open language model from TIGER-Lab. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 2 2B
Google · 2.6B · runs from 1.2 GB
Google Gemma 2 2B is a 2-billion parameter base (pretrained) model from Google's Gemma 2 family. As a base model, it is not instruction-tuned and is intended for fine-tuning, research, and custom downstream applications. Its compact size makes it suitable for experimentation, rapid prototyping, and domain-specific fine-tuning on consumer hardware with minimal VRAM. Released under the Gemma license.
North Micro Vision Instruct
Cohere · 2.5B · runs from 5.5 GB
North Micro Vision Instruct is Cohere's 2.5-billion-parameter vision-language model, built to handle native-resolution images alongside text in multi-turn conversations. It pairs a custom vision encoder with a small in-house language model, tuned for document understanding, chart and table reading, spatial reasoning, and visual grounding across a dozen-plus languages. Its small size suits local deployment on modest consumer GPUs once quantized. The model supports a very large 500,000 token context window, well beyond most models its size, useful for many-page documents or long image-heavy conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in August 2026 as part of Cohere's North family, it accepts inputs at resolutions up to roughly an A4 page at 200 dpi, aimed at document-heavy enterprise use.
Gpt2 Small Dutch
GroNLP · 129M · runs from 0.1 GB
Gpt2 Small Dutch is a 129M-parameter open language model from GroNLP. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
MiMo 7B Base
Xiaomi · 7.8B · runs from 3.9 GB
MiMo 7B Base is a 7.8B-parameter open language model from Xiaomi. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 1.5B Quantized.w8a8
RedHatAI · 1.8B · runs from 1.1 GB
Qwen2.5 1.5B Quantized.w8a8 is a 1.8B-parameter open language model from RedHatAI in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 2 13B Chat HF
Meta · 13.0B · runs from 6.1 GB
Meta Llama 2 13B Chat is a 13-billion parameter instruction-tuned model from Meta's Llama 2 family, fine-tuned for dialogue and chat applications. It offers improved reasoning and generation quality over the 7B variant while maintaining manageable hardware requirements with a 4K token context window. The model was fine-tuned using supervised fine-tuning and RLHF. It can run on consumer GPUs with 16GB or more of VRAM at reduced precision. Released under the Llama 2 Community License.