All LLM Models
Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
HF Moshiko
kmhf · 7.8B · runs from 16.9 GB
HF Moshiko is a 7.8B-parameter open language model from kmhf. It supports a context window of up to 3,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Deepseek V4 Flash 0731 Spark
0xSero · 60.3B · runs from 26.0 GB
Deepseek V4 Flash 0731 Spark is a 60.3B-parameter open language model from 0xSero in the DeepSeek V4 family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Xglm 564M
Meta · 564M · runs from 0.3 GB
Xglm 564M is a 564M-parameter open language model from Meta in the GLM family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 3.1 405B
Meta · 405.9B · runs from 189.7 GB
Meta Llama 3.1 405B is the largest model in the Llama family with 405 billion parameters. It represents Meta's most capable open-weight model, delivering performance competitive with leading proprietary models across reasoning, coding, math, and multilingual tasks. It features a 128K token context window. Due to its massive size, running Llama 3.1 405B locally requires significant hardware, typically multiple high-end professional GPUs with a combined VRAM of 200GB or more at reduced precision. It is primarily used in quantized formats for local inference or via multi-node setups. Released under the Llama 3.1 Community License.
MiMo 7B RL
Xiaomi · 7.8B · runs from 3.9 GB
MiMo 7B RL is a 7.8B-parameter open language model from Xiaomi. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NVIDIA Nemotron Parse 2.0
NVIDIA · 903M · runs from 2.0 GB
NVIDIA Nemotron Parse 2.0 is a sub-1-billion-parameter vision-encoder-decoder model purpose-built for document parsing rather than open-ended chat: given a page image, it outputs structured text with layout classes, bounding boxes, and reading order for elements like titles, paragraphs, tables, charts, and footnotes. It pairs a ViT-H vision encoder based on NVIDIA's C-RADIO with a 10-block mBART decoder, and over its predecessor v1.2 it adds roughly 20,000 new vocabulary tokens for more efficient multilingual (especially CJK and Indic-script) OCR, a dedicated chart class for chart-to-table parsing, and stronger table detection and text extraction. At under a billion parameters it runs on a single modest GPU. License is the OpenMDW License Agreement version 1.1, permitting both commercial and non-commercial use; the bundled tokenizer is separately licensed under CC-BY-4.0. It was published in August 2026.
Gemma 3 270M
Google · 268M · runs from 0.1 GB
Google Gemma 3 270M is a 270-million parameter base (pretrained) model from Google's Gemma 3 family. It is an experimental release intended for research, fine-tuning, and exploring the capabilities of ultra-small language models. The model runs on virtually any hardware with negligible resource requirements. Released under the Gemma license.
GPT Neox Japanese 2.7B
abeja · 2.7B · runs from 5.9 GB
GPT Neox Japanese 2.7B is a 2.7B-parameter open language model from abeja. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Open Llama 7B
openlm-research · 7B · runs from 3.3 GB
Open Llama 7B is a 7B-parameter open language model from openlm-research in the Llama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Deepseek Vl2 Tiny
DeepSeek · 3.4B · runs from 7.4 GB
DeepSeek-VL2-Tiny is the smallest in DeepSeek's VL2 series of Mixture-of-Experts vision-language models, with about 3.4 billion total parameters but only roughly 1.2 billion activated per token. It handles vision-language tasks — visual question answering, OCR, document/table/chart understanding, and visual grounding — below the larger VL2-Small and full VL2 variants. Only a fraction of parameters activate per token, so decoding stays fast even though every expert must still be loaded into memory. It is compact enough to run on a single consumer GPU once quantized. Context length is limited to 4,096 tokens. It is released under DeepSeek's own model license, a custom permissive license with an acceptable-use policy rather than a fully open license like MIT or Apache 2.0. Published in December 2024, it pairs a small DeepSeekMoE-3B backbone with a dynamic image-tiling vision encoder.
Glm 4 9B Chat
Z.ai · 9.4B · runs from 4.4 GB
Glm 4 9B Chat is a 9.4B-parameter open language model from Z.ai in the GLM 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Opt 1.3B
Meta · 1.3B · runs from 2.9 GB
Opt 1.3B is a 1.3B-parameter open language model from Meta. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Stories15M MOE
ggml-org · 36M · runs from 0.3 GB
Stories15M MOE is a 36M-parameter open language model from ggml-org. It supports a context window of up to 256 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Inkling Small DSpark Preview
RadixArk · 1.4B · runs from 0.9 GB
Inkling Small DSpark Preview is a 1.4B-parameter open language model from RadixArk in the Inkling family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen Base Invoicev1.01 1.5B
LaaP-ai · 1.5B · runs from 1.0 GB
Qwen Base Invoicev1.01 1.5B is a 1.5B-parameter open language model from LaaP-ai in the Qwen family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Chandra
datalab-to · 8.8B · runs from 4.3 GB
Chandra is Datalab's 8.8-billion-parameter vision-language OCR model, built on a Qwen3VL backbone, that converts document images and PDFs into markdown, HTML, or JSON while preserving layout, tables, math, and forms with checkboxes. It handles handwriting, multi-column and complex layouts, and extracts images and diagrams with captions and structured data across more than 40 languages, aiming at document digitization rather than general chat. On the olmOCR benchmark Chandra scored highest overall among compared models, ahead of Datalab's own Marker pipeline, Mistral's OCR API, DeepSeek-OCR, and anchored GPT-4o and Gemini Flash 2 baselines. At under 9 billion parameters it runs on a single consumer GPU. Context length is 262,144 tokens. It is released under an OpenRAIL license, a responsible-AI license that permits broad use but restricts certain harmful applications. It was published in October 2025 and has since been superseded by a newer Chandra OCR 2 model.
Nemotron Labs Diffusion 8B Base
NVIDIA · 8.5B · runs from 17.6 GB
Nemotron Labs Diffusion 8B Base is a 8.5B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
MobileLLaMA 1.4B Chat
mtgv · 1.4B · runs from 1.3 GB
MobileLLaMA 1.4B Chat is a 1.4B-parameter open language model from mtgv in the Llama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
TinyLLama V0
Maykeye · 5M · runs from 0.0 GB
TinyLLama V0 is a 5M-parameter open language model from Maykeye in the TinyLlama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Hy3 Preview
Tencent · 298.8B · runs from 127.6 GB
Hy3 Preview is a 298.8B-parameter open language model from Tencent in the Hunyuan 3 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama2 0B Unit Test
MaxJeblick · 770940 · runs from 0.3 GB
Llama2 0B Unit Test is a 770940-parameter open language model from MaxJeblick in the Llama 2 family. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Edge0 35B A3B Preview
Edge0 · 34.7B · runs from 69.7 GB
Edge0 35B A3B Preview is a 34.7B-parameter open language model from Edge0. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Meta Llama 3 70B
Meta · 70.6B · runs from 33.0 GB
Meta Llama 3 70B is a 70.6B-parameter open language model from Meta in the Llama 3 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Nl2sh 1.5B Q4 K M
ThorOdinson246 · 1.5B · runs from 0.7 GB
Nl2sh 1.5B Q4 K M is a 1.5B-parameter open language model from ThorOdinson246. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Apertus 70B 2509
swiss-ai · 70.6B · runs from 31.0 GB
Apertus 70B 2509 is a 70.6B-parameter open language model from swiss-ai in the Apertus family. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Moonlight 16B A3B
Moonshot AI · 16.0B · runs from 7.5 GB
Moonlight 16B A3B is a compact Mixture-of-Experts model from Moonshot AI that packs 16 billion total parameters while activating only around 3 billion per token. This efficient sparse design lets it punch well above its active parameter count, delivering surprisingly strong chat performance for its effective inference cost. The small active parameter count means Moonlight runs briskly on modest hardware, fitting comfortably on GPUs with 8–12 GB of VRAM at common quantization levels. It is an appealing choice for users who want MoE-level performance diversity without the heavy memory footprint typically associated with mixture models.
Qwen3 30B A3B.w8a8
nytopop · 30.6B · runs from 13.4 GB
Qwen3 30B A3B.w8a8 is a 30.6B-parameter open language model from nytopop in the Qwen 3 family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Transformer 1.3B 100B
fla-hub · 1.4B · runs from 3 GB
Transformer 1.3B 100B is a 1.4B-parameter open language model from fla-hub. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Nemotron Labs Diffusion 8B
NVIDIA · 8.5B · runs from 17.6 GB
Nemotron Labs Diffusion 8B is a 8.5B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3 Vision 128k Instruct
Microsoft · 4.1B · runs from 9.4 GB
Phi 3 Vision 128k Instruct is a 4.1B-parameter open language model from Microsoft in the Phi 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.