All LLM Models
Browse 982 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
UI TARS 1.5 7B
ByteDance-Seed · 8.3B · runs from 2.7 GB
UI-TARS-1.5-7B is ByteDance's 8.3-billion-parameter vision-language model, built on a Qwen2.5-VL-7B foundation and tuned as a GUI and computer-use agent rather than a chatbot. It reasons through its thoughts before acting, using reinforcement learning to control desktop and browser interfaces, play games, and complete multi-step tasks from screenshots. ByteDance evaluates it against agents like OpenAI's CUA and Claude on benchmarks such as OSWorld and Windows Agent Arena. At this size, local inference is practical on a single mainstream-to-high-end consumer GPU once quantized. It supports a 128,000 token context window, useful for long action histories in agent tasks. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. Published in April 2025, it builds on the original UI-TARS architecture with added inference-time reasoning scaling.
Gemma 4 E4B
Google · 8.0B · runs from 3.9 GB
Gemma 4 E4B is Google DeepMind's second-smallest model in the Gemma 4 family, a dense architecture with roughly 8 billion total parameters, of which Google describes about 4.5 billion as its effective footprint at inference. This is the pretrained base checkpoint rather than an instruction-tuned model, meant for fine-tuning rather than direct chat use. Like the rest of the family it is multimodal — text, image, and natively audio at this size — targeting efficient on-device deployment on laptops and higher-end phones. It runs comfortably on a single consumer GPU, or on-device once quantized. It supports a 131,072 token context window. It carries Google's Gemma 4 license terms, published as Apache 2.0 on Hugging Face with additional Gemma-specific usage terms linked from the card. Published in March 2026, it sits between the E2B on-device model and the larger 12B, 26B-A4B, and 31B tiers.
Yi 1.5 9B Chat
01.AI · 8.8B · runs from 4.1 GB
Yi-1.5-9B-Chat is 01.AI's second-generation 9-billion-parameter bilingual (English/Chinese) chat model, continually pretrained from the original Yi series on a further 500 billion high-quality tokens and then fine-tuned on 3 million diverse instruction samples. Compared with the original Yi, Yi-1.5 delivers stronger coding, math, and reasoning while keeping the same language understanding and commonsense reasoning, and 01.AI reports it as the top performer among similarly sized open models on its benchmark suite. At 9B parameters it fits comfortably on a single consumer GPU, or a much smaller card once quantized. Context length is 4,096 tokens; separate 16K- and 32K-context variants of the same model are also available for longer documents. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2024.
Meta Llama 3.1 8B Instruct Abliterated
mlabonne · 8.0B · runs from 3.3 GB
Meta Llama 3.1 8B Instruct Abliterated is a 8.0B-parameter open language model from mlabonne in the Llama 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
EXAONE 3.5 7.8B Instruct
LGAI-EXAONE · 7.8B · runs from 2.7 GB
EXAONE 3.5 7.8B Instruct is a 7.8B-parameter open language model from LGAI-EXAONE in the EXAONE family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 3.1 Tulu 3 8B
Allen AI · 8.0B · runs from 3.3 GB
Llama-3.1-Tulu-3-8B is Allen Institute for AI's fully open instruction-following model, built on Meta's Llama 3.1 8B base through a public pipeline of supervised fine-tuning, direct preference optimization, and a final reinforcement learning stage with verifiable rewards (RLVR). It is tuned for chat as well as harder tasks like math, GSM8K, and instruction-following (IFEval), with all training data, code, and recipes released openly as part of the Tulu 3 project. Its 8B size lets it run on a single consumer GPU. Context length is 131,072 tokens. It is released under the Llama 3.1 Community License Agreement, which restricts use above 700 million monthly active users and imposes Meta's acceptable-use policy. It was published in November 2024, alongside a 70B and later a 405B sibling trained the same way.
Moonlight 16B A3B Instruct
Moonshot AI · 16.0B · runs from 5.1 GB
Moonlight 16B A3B Instruct is a 16.0B-parameter open language model from Moonshot AI in the Moonlight family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
MiniCPM V 4.6 Thinking
OpenBMB · 1.3B · runs from 0.7 GB
MiniCPM-V 4.6 Thinking is OpenBMB's 1.3-billion-parameter vision-language model, a long chain-of-thought reasoning variant of MiniCPM-V 4.6 built on a SigLIP2-400M vision encoder paired with a small Qwen3.5-0.8B language backbone. It generates an explicit reasoning trace before answering, aimed at multimodal reasoning, math, and OCR-heavy document tasks rather than quick captioning, keeping the same edge-friendly, phone-oriented architecture. Its small size lets it run on a single modest consumer GPU. The model supports a 262,144 token context window. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2026. Its distinguishing trait versus base 4.6 is the thinking mode, which trades some latency for better performance on reasoning-heavy visual tasks while reusing the same mixed 4x/16x visual token compression.
Gpt2
OpenAI · 137M · runs from 0.1 GB
GPT-2 is the landmark 2019 language model from OpenAI that helped ignite widespread interest in large-scale text generation. At only 137 million parameters it is tiny by modern standards, but it holds an important place in AI history as the model that was initially deemed too dangerous to release in full. Today GPT-2 runs effortlessly on virtually any hardware, including CPUs, making it ideal for educational purposes, experimentation, and understanding transformer fundamentals. It should not be expected to match the quality of modern instruction-tuned models, but it remains a useful teaching tool and conversation starter.
Apertus 8B Instruct 2509
swiss-ai · 8.1B · runs from 2.8 GB
Apertus 8B Instruct is an open-source instruction-tuned model from Swiss AI, a collaborative research initiative. Built on an 8 billion parameter base, it emphasizes transparency, open data, and European AI sovereignty. For local users, it delivers solid general-purpose chat and instruction-following in a standard 8B footprint that runs well on consumer GPUs with 8 to 10 GB of VRAM, making it a practical choice for those who value open, community-driven model development.
Nemotron Mini 4B Instruct
NVIDIA · 4B · runs from 1.8 GB
Nemotron Mini 4B Instruct is a 4B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
DeepSeek v2 Lite Chat
DeepSeek · 15.7B · runs from 5.1 GB
DeepSeek v2 Lite Chat is a 15.7B-parameter open language model from DeepSeek in the DeepSeek V2 family. It supports a context window of up to 163,840 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ternary Bonsai 4B Unpacked
prism-ml · 4.0B · runs from 2.2 GB
Ternary Bonsai 4B Unpacked is a 4.0B-parameter open language model from prism-ml. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Zephyr 7B Beta
Hugging Face · 7.2B · runs from 3.6 GB
Zephyr 7B Beta is a 7.2B-parameter open language model from Hugging Face. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
CodeLlama 34B Instruct HF
Meta · 33.7B · runs from 10.0 GB
CodeLlama-34b-Instruct-hf is Meta's 34-billion-parameter instruction-tuned Code Llama model, fine-tuned from the base Code Llama checkpoint for safer, more reliable code-assistant use, as opposed to the plain base and Python-specialized variants in the same family (which also ships in 7B, 13B, and 70B sizes). It is an early, first-generation code model built on the original Llama 2 architecture and predates the later Code Llama 70B release. At 34B parameters, it needs a multi-GPU setup or aggressive quantization for local use. Context length is 16,384 tokens. It is released under the Llama 2 Community License, a custom license that restricts commercial use above 700 million monthly active users, requiring a separate license from Meta, and was published in August 2023.
Pantheon Reasoning 27B
Gryphe · 27.8B · runs from 8.4 GB
Pantheon Reasoning 27B is a 27.8B-parameter open language model from Gryphe. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Starcoder2 15B
BigCode · 16.0B · runs from 7.3 GB
Starcoder2 15B is a 16.0B-parameter open language model from BigCode in the StarCoder family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Olmo 3 32B Think
Allen AI · 32.2B · runs from 9.7 GB
Olmo 3 32B Think is Allen Institute for AI's reasoning-focused model in the fully open Olmo 3 family, trained to produce long chains of thought before answering so it can work through math and coding problems step by step. It is pretrained on Ai2's Dolma 3 corpus and then carried through a full post-training pipeline of supervised fine-tuning, direct preference optimization, and reinforcement learning with verifiable rewards on the Dolci-Think-RL dataset; Ai2 releases the code, intermediate checkpoints, and training logs for each stage, not just the final weights. It sits alongside a 7B Think and a 7B Instruct sibling in the same release. At 32 billion parameters it needs a high-end consumer GPU or a multi-GPU setup once quantized. Context length is 65,536 tokens. It is released under the Apache 2.0 license, though Ai2 asks that it be used for research and educational purposes in line with its Responsible Use Guidelines. It was published in November 2025 and has since been superseded by Olmo 3.1 32B Think.
Granite 3.3 2B Instruct
IBM · 2.5B · runs from 1.2 GB
Granite 3.3 2B Instruct is a 2.5B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ministral 8B Instruct 2410
Mistral AI · 8.0B · runs from 3.3 GB
Ministral 8B Instruct 2410 is Mistral AI's 8-billion-parameter instruction-tuned model, part of the "Ministraux" family aimed at edge and on-device deployment. It uses a 36-layer dense transformer with interleaved sliding-window attention, trained heavily on multilingual and code data, suiting chat, function calling, and general assistant tasks across ten languages. Its compact size means it runs comfortably on a single mainstream consumer GPU once quantized. The model supports a 32K token context window. It is released under Mistral's own license, listed as "Other," which restricts use to research purposes and requires a separate commercial license from Mistral AI, so review the terms before commercial use. It was published in October 2024 alongside the smaller Ministral 3B.
EXAONE 3.5 2.4B Instruct
LGAI-EXAONE · 2.4B · runs from 0.9 GB
EXAONE 3.5 2.4B Instruct is a 2.4B-parameter open language model from LGAI-EXAONE in the EXAONE family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Lucid V1 Nemo
dreamgen · 12.2B · runs from 4.1 GB
Lucid V1 Nemo is a 12.2B-parameter open language model from dreamgen. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
LocateAnything 3B
NVIDIA · 3.8B · runs from 2 GB
LocateAnything-3B is NVIDIA's 3.8-billion-parameter vision-language model for visual grounding rather than open-ended chat: referring-expression grounding, dense multi-object detection, GUI element grounding, point-based localization, and document or OCR layout grounding. Its core contribution is Parallel Box Decoding, which predicts a complete bounding box in one parallel step instead of token-by-token autoregressive decoding, giving up to 2.5x higher throughput while preserving geometric consistency. It combines a Qwen2.5-3B-Instruct language backbone with a MoonViT-SO-400M vision encoder and was trained on 12 million images with over 138 million grounding queries; NVIDIA has folded it into the Nemotron 3 Nano Omni model for agentic and computer-use grounding. At under 4 billion parameters, it runs on a single consumer GPU. Context length is 32,768 tokens. It is released under NVIDIA's non-commercial research license, permitting academic and non-profit use only; commercial use requires a separate license from NVIDIA. It was published in May 2026.
QwQ 32B Preview
Alibaba · 32.8B · runs from 10.7 GB
QwQ 32B Preview is a 32.8B-parameter open language model from Alibaba in the QwQ family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ternary Bonsai 1.7B Unpacked
prism-ml · 1.7B · runs from 1.3 GB
Ternary Bonsai 1.7B Unpacked is a 1.7B-parameter open language model from prism-ml. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Dream V0 Instruct 7B
Dream-org · 7.6B · runs from 2.5 GB
Dream V0 Instruct 7B is a 7.6B-parameter open language model from Dream-org. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
CodeLlama 7B Instruct HF
Meta · 6.7B · runs from 4.2 GB
CodeLlama 7B Instruct HF is a 6.7B-parameter open language model from Meta in the Code Llama family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Supra Router 51M
SupraLabs · 52M · runs from 0.3 GB
Supra Router 51M is a 52M-parameter open language model from SupraLabs. It supports a context window of up to 5,120 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
OLMoE 1B 7B 0924
Allen AI · 6.9B · runs from 3.5 GB
OLMoE 1B 7B 0924 is a 6.9B-parameter open language model from Allen AI in the OLMo family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Granite 3.0 2B Instruct
IBM · 2.6B · runs from 1.3 GB
Granite-3.0-2B-Instruct is IBM's 2-billion-parameter chat model, fine-tuned from Granite-3.0-2B-Base on a mix of permissively licensed open instruction datasets and IBM's own synthetic data, using supervised fine-tuning, reinforcement-learning alignment, and model merging. It is one of several Granite 3.0 sizes IBM released together, aimed at enterprise use cases such as summarization, question answering, and retrieval-augmented generation. It supports dialogue in twelve languages, primarily English, though multilingual performance trails English performance. At 2B parameters, it is small enough to run on a laptop CPU or any consumer GPU, including edge deployments. Context length is 4,096 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2024. It was later superseded by the Granite 3.1 model family.