All LLM Models
Browse 982 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Meta Llama 3 8B
Meta · 8.0B · runs from 3.8 GB
Meta Llama 3 8B is an 8-billion parameter base (pretrained) language model from Meta's Llama 3 release. As a base model, it is not fine-tuned for chat or instructions and is intended for further fine-tuning, research, or as a foundation for custom applications. It uses grouped-query attention and was trained on over 15 trillion tokens. Llama 3 8B supports an 8K token context window and delivers strong benchmark performance across language understanding, reasoning, and coding tasks for its size. It is released under the Meta Llama 3 Community License and runs efficiently on consumer GPUs with 8GB or more of VRAM.
Kappa 20B 131k
eousphoros · 20.9B · runs from 9.3 GB
Kappa 20B 131k is a 20.9B-parameter open language model from eousphoros. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama Guard 3 1B
alpindale · 1.5B · runs from 0.8 GB
Llama Guard 3 1B is a 1.5B-parameter open language model from alpindale in the Llama family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Falcon Mamba 7B
TII UAE · 7.3B · runs from 2.7 GB
Falcon Mamba 7B is a 7.3B-parameter open language model from TII UAE in the Falcon family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.5 2B Base
Alibaba · 2.3B · runs from 1.4 GB
Qwen3.5-2B-Base is a 2.3-billion-parameter dense checkpoint in Alibaba's Qwen3.5 base model family, a native vision-language foundation model rather than a text-only model with vision bolted on. It is pretrain-only: fine-tuning, in-context-learning, or further research, not direct conversation, though its control tokens are compatible with the official chat template for efficient LoRA-style adaptation. Its hybrid architecture pairs Gated DeltaNet linear attention with periodic gated full-attention layers. At 2.3B parameters, it runs comfortably on a single consumer GPU, even unquantized. The model supports a native 262,144 token context window, extensible up to 1,010,000 tokens. It is released under the Apache 2.0 license, and was published in February 2026. Qwen3.5 introduced early-fusion multimodal pretraining that Alibaba says outperforms the separate Qwen3-VL models on reasoning, coding, and visual understanding.
SmolLM3 3B Base
Hugging Face · 3.1B · runs from 1.3 GB
SmolLM3 3B Base is the pretrained foundation model from Hugging Face's third-generation SmolLM family. Without instruction tuning or chat alignment, it serves as a versatile starting point for researchers and developers who want to fine-tune the model for specific domains, tasks, or behavioral profiles. With 3 billion parameters and the architectural improvements introduced in SmolLM3, this base model offers strong general language capabilities in a package that remains practical to train and adapt on consumer-grade hardware. It is an excellent choice for custom fine-tuning projects where off-the-shelf chat behavior is not needed.
Llama 3.1 8B
Meta · 8.0B · runs from 3.6 GB
Meta Llama 3.1 8B is an 8-billion parameter base (pretrained) model from the Llama 3.1 family. It is not instruction-tuned and is intended for fine-tuning, research, and custom downstream applications. Compared to Llama 3 8B, it extends the context window to 128K tokens and benefits from improved training data and methodology. The model uses grouped-query attention and was trained on a multilingual corpus. It is released under the Llama 3.1 Community License and is widely used as a foundation for community fine-tunes and specialized models.
Qwen3 0.6B Heretic Abliterated Uncensored
DavidAU · 596M · runs from 0.7 GB
Qwen3 0.6B Heretic Abliterated Uncensored is a 596M-parameter open language model from DavidAU in the Qwen 3 family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qianfan OCR
Baidu · 4.7B · runs from 2.5 GB
Qianfan-OCR is Baidu's 4.7-billion-parameter vision-language model for document intelligence rather than general chat. It pairs a Qianfan-ViT vision encoder with a Qwen3-4B language backbone, doing direct image-to-Markdown conversion alongside table extraction, chart understanding, and document Q&A in one end-to-end model instead of a multi-stage pipeline. It also has an optional "Layout-as-Thought" mode that reasons about page layout before producing output. At under 5 billion parameters, it runs on a single consumer or prosumer GPU once quantized. Context length is 32,768 tokens, extendable further per Baidu's documentation. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. Published in March 2026, Baidu reports it as its top-scoring end-to-end model on public document-parsing benchmarks, supporting 192 languages.
SmolLM 135M Instruct
Hugging Face · 135M · runs from 0.4 GB
SmolLM 135M Instruct is a 135M-parameter open language model from Hugging Face in the SmolLM family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
OneRec 1.7B
OpenOneRec · 2.1B · runs from 1.1 GB
OneRec 1.7B is a 2.1B-parameter open language model from OpenOneRec. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
LFM2 1.2B
Liquid AI · 1.2B · runs from 0.9 GB
LFM2 1.2B is a 1.2B-parameter open language model from Liquid AI in the LFM2 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3Guard Gen 0.6B
Alibaba · 752M · runs from 0.7 GB
Qwen3Guard Gen 0.6B is a 752M-parameter open language model from Alibaba in the Qwen 3 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
JiRackUltra 32B
CMSManhattan · 32.8B · runs from 9.8 GB
JiRackUltra 32B is a 32.8B-parameter open language model from CMSManhattan. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen1.5 0.5B Chat
Alibaba · 620M · runs from 0.8 GB
Qwen1.5 0.5B Chat is an early-generation small language model from Alibaba's Qwen series with just 620 million parameters. As one of the smallest models in the Qwen family, it was designed to demonstrate that useful conversational ability is possible even at sub-billion parameter scales. This model runs easily on virtually any hardware including CPUs, older GPUs, and even mobile devices. While its capabilities are limited compared to larger Qwen models, it remains a useful option for embedded applications, rapid prototyping, or situations where minimal resource consumption is the top priority.
Qwen3Guard Gen 4B
Alibaba · 4.4B · runs from 2.4 GB
Qwen3Guard Gen 4B is a 4.4B-parameter open language model from Alibaba in the Qwen 3 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 1.1 2B IT
Google · 2.5B · runs from 1.1 GB
Gemma 1.1 2B IT is a 2.5B-parameter open language model from Google in the Gemma family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Starling LM 7B Beta
Nexusflow · 7.2B · runs from 3.6 GB
Starling-LM-7B-beta is Nexusflow's chat model, fine-tuned from Openchat-3.5-0106 (itself based on Mistral-7B-v0.1) using reinforcement learning from AI feedback (RLAIF). It was trained with Nexusflow's own 34B reward model and a PPO-based policy optimization pipeline on the Nectar preference dataset, and scored 8.12 on MT-Bench with GPT-4 as judge, an improvement over the team's earlier Starling model. It uses OpenChat's exact chat template rather than a generic one. At 7 billion parameters it runs comfortably on a single consumer GPU. Context length is 8,192 tokens. It is released under the Apache 2.0 license with an added condition that it not be used to compete with OpenAI, reflecting that its Nectar training data was generated with GPT-4. It was published in March 2024.
Chatglm3 6B
Z.ai · 6.2B · runs from 2.9 GB
ChatGLM3-6B is the third generation of Zhipu AI's (now Z.ai) open bilingual Chinese-English chat model, built on the GLM architecture at around 6.2 billion parameters. Compared with its predecessors it adds native support for function calling, a code interpreter, and agent-style tasks through a newly designed prompt format, alongside a stronger base model, ChatGLM3-6B-Base, trained on more diverse data. A companion long-context variant, ChatGLM3-6B-32K, was released alongside it. Its modest size lets it run on a single consumer GPU. Context length is 8,192 tokens. The code is released under Apache 2.0, but the model weights use a separate Model License that is free for academic research and allows commercial use only after completing Zhipu's registration questionnaire. It was published in October 2023, and has since been superseded by the GLM-4 series.
JiRackUltra 1B
CMSManhattan · 1.8B · runs from 0.8 GB
JiRackUltra 1B is a 1.8B-parameter open language model from CMSManhattan. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 Math 1.5B
Alibaba · 1.5B · runs from 1 GB
Qwen2.5 Math 1.5B is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Fanar 2 27B Instruct
QCRI · 27.0B · runs from 9.1 GB
Fanar 2 27B Instruct is a 27.0B-parameter open language model from QCRI. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
EXAONE 4.0 1.2B
LGAI-EXAONE · 1.3B · runs from 1.0 GB
EXAONE 4.0 1.2B is a 1.3B-parameter open language model from LGAI-EXAONE in the EXAONE family. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen1.5 1.8B Chat
Alibaba · 1.8B · runs from 1.5 GB
Qwen1.5 1.8B Chat is a 1.8B-parameter open language model from Alibaba in the Qwen family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Bloom 560M
BigScience · 559M · runs from 0.3 GB
Bloom 560M is a 559M-parameter open language model from BigScience. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
SmolLM 360M Instruct
Hugging Face · 362M · runs from 0.5 GB
SmolLM 360M Instruct is a 362M-parameter open language model from Hugging Face in the SmolLM family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
LightOnOCR 2 1B
lightonai · 1.0B · runs from 0.8 GB
LightOnOCR-2-1B is LightOn's flagship end-to-end OCR vision-language model, a 1-billion-parameter system that converts documents such as PDFs, scans, and photos directly into clean, naturally ordered markdown without a separate OCR pipeline. It handles tables, receipts, forms, multi-column layouts, and math notation, is trained on a large multilingual corpus with particular strength in French, arXiv papers, and scanned documents, and this second-generation release adds RLVR (reinforcement learning from verifiable rewards) training on top of the base model for extra accuracy. LightOn reports state-of-the-art results on the OlmOCR-Bench benchmark while being roughly nine times smaller and several times faster than competing OCR systems. At just 1 billion parameters, it runs comfortably even on a modest single consumer GPU. Context length is 16,384 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in January 2026, alongside base, bounding-box, and merged "soup" variants of the same model.
Deepseek Coder 7B Instruct V1.5
DeepSeek · 6.9B · runs from 4.2 GB
Deepseek Coder 7B Instruct V1.5 is a 6.9-billion-parameter code-focused language model from DeepSeek, fine-tuned from the DeepSeek-LLM 7B base for programming assistance, code generation, and general chat. It continues DeepSeek's original Coder line, trained on roughly 2 trillion tokens of code-heavy text before instruction tuning, and its default system prompt frames it specifically as a programming assistant. Its size makes it well suited to local deployment on a single mainstream consumer GPU with around 8GB or more of VRAM once quantized. Context length is limited to 4,096 tokens, short by current standards. It is released under DeepSeek's own custom license (listed as "Other"), so commercial users should review the license terms before deployment. Published in January 2024, it is one of DeepSeek's earlier widely adopted open-weight coding assistants.
TIPO 500M Ft
KBlueLeaf · 508M · runs from 0.7 GB
TIPO 500M Ft is a 508M-parameter open language model from KBlueLeaf. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Racka 4B
elte-nlp · 4.0B · runs from 1.9 GB
Racka 4B is a 4.0B-parameter open language model from elte-nlp. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.