All LLM Models
Browse 1214 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Qwen2 1.5B Instruct
Alibaba · 1.5B · runs from 0.8 GB
Qwen2 1.5B Instruct is a 1.5B-parameter open language model from Alibaba in the Qwen 2 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NVIDIA Nemotron 3 Nano 4B BF16
NVIDIA · 4.0B · runs from 2.2 GB
NVIDIA Nemotron 3 Nano 4B BF16 is a 4.0B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 E2B
Google · 5.1B · runs from 2.5 GB
Gemma 4 E2B is Google DeepMind's smallest model in the Gemma 4 family, a dense architecture with roughly 5.1 billion total parameters, of which Google describes about 2.3 billion as its effective footprint at inference. This is the pretrained base checkpoint, not an instruction-tuned chat model, meant as a starting point for fine-tuning. The Gemma 4 family is multimodal — text, image, and at this size natively audio — and E2B targets efficient on-device execution on phones and laptops. It is easy to run locally, even on modest hardware once quantized. It supports a 131,072 token context window. It carries Google's Gemma 4 license terms, published as Apache 2.0 on Hugging Face with additional Gemma-specific usage terms linked from the card. Published in March 2026, it is the smallest of five sizes in the Gemma 4 lineup, aimed at mobile and edge deployment.
NeuralDaredevil 8B Abliterated
mlabonne · 8.0B · runs from 4.0 GB
NeuralDaredevil 8B Abliterated is a 8.0B-parameter open language model from mlabonne. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 14B Base
Alibaba · 14.8B · runs from 4.7 GB
Qwen3 14B Base is a 14.8B-parameter open language model from Alibaba in the Qwen 3 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Huihui MiniCPM5 1B Abliterated
huihui-ai · 1.1B · runs from 0.6 GB
Huihui MiniCPM5 1B Abliterated is a 1.1B-parameter open language model from huihui-ai in the MiniCPM family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2 1.5B
Alibaba · 1.5B · runs from 1 GB
Qwen2 1.5B is a 1.5-billion parameter base (pretrained) model from Alibaba Cloud's older Qwen 2 generation. It was trained on a multilingual corpus and supports a context window of up to 32K tokens. As a base model, it is designed for fine-tuning and research rather than direct conversational use. While superseded by the Qwen 2.5 series in terms of training data quality and benchmark performance, Qwen2 1.5B remains a lightweight option for experimentation and as a baseline for comparison. Released under the Apache 2.0 license.
MiniCPM5 1B
OpenBMB · 1.1B · runs from 0.6 GB
MiniCPM5 1B is a compact 1-billion-parameter chat model from OpenBMB, part of the MiniCPM family and built on a Llama-style architecture. Its small size makes it a practical choice for scenarios where footprint, latency, and power draw matter more than maximum capability, such as edge or on-device deployment. The model offers a 128K token context window, unusually long for a model this size, and is released under the Apache 2.0 license, permitting free local and commercial use without restriction. At around 1 billion parameters, it runs easily on modest consumer hardware, including machines without a dedicated high-end GPU.
LLaDA 8B Instruct
GSAI-ML · 8.0B · runs from 2.4 GB
LLaDA 8B Instruct is a 8.0B-parameter open language model from GSAI-ML. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Nanonets OCR2 3B
nanonets · 3.8B · runs from 1.4 GB
Nanonets-OCR2-3B is Nanonets' image-to-markdown OCR model, a 3.8-billion-parameter vision-language model fine-tuned from Qwen2.5-VL-3B-Instruct to turn documents into structured markdown with semantic tagging rather than plain extracted text. It converts equations into LaTeX, isolates signatures and watermarks into dedicated tags, converts form checkboxes into standard Unicode symbols, extracts complex tables as markdown or HTML, renders flow and organizational charts as Mermaid diagrams, and handles handwritten and multilingual documents across a dozen or more languages; it can also answer direct questions about a document's contents. It is an OCR and document-understanding tool rather than a general chat assistant. Its small size lets it run on a single consumer GPU. Context length is 128,000 tokens, inherited from its Qwen2.5-VL base. Nanonets has not published a specific open-source license for the model on Hugging Face. It was published in October 2025, alongside a smaller Nanonets-OCR2-1.5B-exp variant and a hosted Nanonets-OCR2-Plus service.
LFM2 1.2B RAG
Liquid AI · 1.2B · runs from 0.9 GB
LFM2 1.2B RAG is a 1.2B-parameter open language model from Liquid AI in the LFM2 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Yi 6B
01.AI · 6.1B · runs from 2.9 GB
Yi-6B is 01.AI's 6.1-billion-parameter base language model — pretrained, not instruction-tuned — from the first Yi series, one of the company's earliest large bilingual (English/Chinese) open models trained from scratch. It was trained on roughly 3 trillion multilingual tokens and is positioned as a general-purpose foundation for fine-tuning rather than direct chat use; a separate Yi-6B-Chat release and a long-context Yi-6B-200K variant followed later. As a 6-billion-parameter dense model, it fits easily on a single consumer GPU. Context length is 4,096 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in November 2023, making it one of the older entries in the now much-expanded Yi model family.
Huihui Gemma 4 E2B IT Abliterated
huihui-ai · 5.1B · runs from 2.5 GB
Huihui Gemma 4 E2B IT Abliterated is a 5.1B-parameter open language model from huihui-ai in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 3.1 Tulu 3 70B
Allen AI · 70.6B · runs from 20.4 GB
Llama-3.1-Tulu-3-70B is Allen Institute for AI's 70.6-billion-parameter instruction-following chat model, built by post-training Meta's Llama-3.1-70B base through a fully open pipeline of supervised fine-tuning, direct preference optimization, and a final reinforcement-learning-with-verifiable-rewards stage, with every dataset, script, and recipe published alongside the weights. It targets strong performance on chat as well as math (MATH, GSM8K) and instruction-following (IFEval) benchmarks, and is one entry in a full Tulu 3 family spanning 8B, 70B, and 405B sizes with released SFT, DPO, RLVR, and reward-model checkpoints at each size. At 70.6 billion parameters, it needs a multi-GPU workstation to run in full precision. Context length is 131,072 tokens, inherited from its Llama 3.1 base. It is released under the Llama 3.1 Community License, which requires organizations above 700 million monthly active users to obtain a separate license from Meta. It was published in November 2024.
Cosmos Reason2 2B
NVIDIA · 2.4B · runs from 1.1 GB
Cosmos Reason2 2B is NVIDIA's 2.4-billion-parameter open reasoning vision-language model for physical AI, built on the Qwen3-VL-2B-Instruct backbone and designed to let robots and embodied agents perceive a scene and reason step by step about space, time, and physical common sense before acting. It generates long chains of thought and can output structured spatial data such as 2D/3D point localization, bounding boxes, and trajectories, plus OCR reading of scene text, going beyond a general chat assistant into physical-world planning. At just 2.4 billion parameters it is small enough to run on a single consumer GPU. Context length is 256,000 tokens, up sharply from the 16,000 tokens of the original Cosmos Reason1. It is released under the NVIDIA Open Model License, a custom license permitting commercial use and redistribution provided derivatives carry a "Built on NVIDIA Cosmos" attribution notice. It was published in December 2025, alongside a larger 8B sibling.
Qwen3.5 0.8B Base
Alibaba · 873M · runs from 0.6 GB
Qwen3.5-0.8B-Base is Alibaba's smallest base checkpoint in the Qwen3.5 family, at 0.87 billion parameters, a native vision-language foundation model that fuses image and text tokens during pretraining rather than bolting a vision encoder onto a text-only model. Like the rest of the -Base line, it ships as pretrained-only weights for fine-tuning and research, not direct interaction, though its control tokens support efficient LoRA-style adaptation with the official chat template. Its hybrid architecture pairs Gated DeltaNet linear attention with periodic full attention layers. It runs easily on a single modest consumer GPU, even unquantized. Context length is 262,144 tokens natively, extensible up to 1,010,000 tokens. It is released under the Apache 2.0 license, and was published in February 2026 as the smallest of five Qwen3.5-Base sizes, from 0.8B dense up to a 35B mixture-of-experts model.
OLMo 2 0425 1B Instruct
Allen AI · 1.5B · runs from 1.0 GB
OLMo 2 0425 1B Instruct is a 1.5B-parameter open language model from Allen AI in the OLMo family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gpt2 Large
OpenAI · 812M · runs from 0.4 GB
Gpt2 Large is a 812M-parameter open language model from OpenAI. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
MiniCPM5 2B DSpark
OpenBMB · 324M · runs from 0.5 GB
MiniCPM5 2B DSpark is a 324M-parameter open language model from OpenBMB in the MiniCPM family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
ERNIE 4.5 21B A3B PT
Baidu · 21B · runs from 6.2 GB
ERNIE 4.5 21B A3B PT is a 21B-parameter open language model from Baidu in the ERNIE family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Granite Vision 4.1 4B
IBM · 4.0B · runs from 2.2 GB
Granite Vision 4.1 4B is IBM's 4-billion-parameter vision-language model in the Granite 4.1 family, purpose-built for enterprise document understanding rather than general chat. It is tailored for reading tables, charts, and key-value data out of scanned documents and images, and is delivered as an adapter built on top of the smaller Granite-4.1-3B text model. Its small size means it runs comfortably on modest consumer hardware once quantized, without needing a high-end GPU, making it practical for local or on-premises document-processing pipelines. The model supports a 131K token context window, enough for multi-page documents. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in April 2026 as part of IBM's broader Granite 4.1 release, which also added speech, embedding, and guardrail models.
L3 8B Stheno V3.2
Sao10K · 8.0B · runs from 4.0 GB
L3 8B Stheno V3.2 is a 8.0B-parameter open language model from Sao10K. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 1.5B
Alibaba · 1.5B · runs from 1 GB
Qwen2.5-1.5B is Alibaba's 1.5-billion-parameter entry in the Qwen2.5 family, released as a base pretrained language model rather than an instruction-tuned chat model. Qwen2.5 improved knowledge, coding, and math capabilities over its predecessor through an expanded pretraining corpus, but this checkpoint is the raw base: the model card explicitly discourages using it directly for conversation and recommends fine-tuning (SFT, RLHF, or continued pretraining) first. At this size, it runs comfortably on almost any modern GPU, or even a CPU, without quantization. The model supports a 131,072 token context window and is released under the Apache 2.0 license, allowing unrestricted commercial and research use. It was published in September 2024 alongside the rest of the Qwen2.5 line, which spans from 0.5B to 72B parameters and adds multilingual support for over 29 languages.
HarmBench Llama 2 13B Cls
cais · 13.0B · runs from 5.6 GB
HarmBench Llama 2 13B Cls is a 13.0B-parameter open language model from cais in the Llama 2 family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Falcon 11B
TII UAE · 11.1B · runs from 5.0 GB
Falcon2-11B is TII's base pretrained causal decoder model, not instruction-tuned, with 11 billion parameters trained on over 5,000 billion tokens of the RefinedWeb dataset plus curated corpora spanning English and nine other European languages. Training proceeded in stages that grew the context window from 2,048 to 8,192 tokens before a final high-quality data stage. TII recommends further fine-tuning before production use. At 11 billion parameters it runs on a single consumer GPU, especially once quantized. Context length is 8,192 tokens. It is released under the TII Falcon License 2.0, a custom Apache 2.0-based license that adds an acceptable-use policy for responsible use of the model. It was published in May 2024, and has since been superseded by the Falcon3 series.
Starcoder2 7B
BigCode · 7.2B · runs from 3.5 GB
Starcoder2 7B is a 7.2B-parameter open language model from BigCode in the StarCoder family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.5 9B Claude 4.6 Opus Reasoning Distilled
Jackrong · 9.7B · runs from 4.7 GB
Qwen3.5 9B Claude 4.6 Opus Reasoning Distilled is a 9.7B-parameter open language model from Jackrong in the Qwen 3.5 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.5 35B A3B Base
Alibaba · 36.0B · runs from 10.3 GB
Qwen3.5-35B-A3B-Base is the mixture-of-experts member of Alibaba's Qwen3.5 base family, pairing a 256-expert MoE layer (8 routed plus 1 shared expert per token) with the same Gated DeltaNet and gated-attention hybrid backbone used across the line. It totals roughly 36 billion parameters but activates only about 3 billion per token (the "A3B" in its name), so decoding stays fast even though the full expert set must stay resident in memory. Like its dense siblings it is pretrain-only, for fine-tuning and research, not direct chat; unlike them, its card lists only a pretraining stage, with no post-training pass. It needs a high-end consumer GPU once quantized. Context length is 262,144 tokens natively, extensible to 1,010,000 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in February 2026, alongside four smaller dense Qwen3.5-Base models from 0.8B to 9B parameters.
SmolLM 360M
Hugging Face · 362M · runs from 0.5 GB
SmolLM-360M is Hugging Face's 360-million-parameter base language model, one of three sizes (135M, 360M, and 1.7B) in the original SmolLM family. It is a pretrained model, not instruction-tuned or chat-formatted, trained on 600 billion tokens from the curated Cosmo-Corpus dataset: Cosmopedia v2 synthetic textbooks and stories, Python-Edu educational code samples, and FineWeb-Edu deduplicated educational web text. At this size it is intended as a lightweight research and fine-tuning base rather than a ready chat assistant, and it runs comfortably on a CPU or any consumer GPU. Context length is 2,048 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in July 2024.
InternVL3 1B
OpenGVLab · 938M · runs from 0.3 GB
InternVL3-1B is OpenGVLab's smallest vision-language model in the InternVL3 series, combining an InternViT-300M-448px-V2.5 vision encoder with a Qwen2.5-0.5B language backbone through an MLP projector. Unlike earlier InternVL generations, it uses Native Multimodal Pre-Training, interleaving image-text and video-text data with text-only corpora in a single pretraining stage rather than adapting a finished text model, and adds Variable Visual Position Encoding for better long-context handling. The series extends its capabilities to tool use, GUI agents, industrial image analysis, and 3D vision perception. At under 1 billion parameters, it is small enough to run on a modest consumer GPU or even a CPU. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in April 2025, as the smallest member of a family that scales up to InternVL3-78B.