All LLM Models
Browse 982 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
MiMo 7B RL
Xiaomi · 7.8B · runs from 3.9 GB
MiMo 7B RL is a 7.8B-parameter open language model from Xiaomi. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NVIDIA Nemotron Parse 2.0
NVIDIA · 903M · runs from 2.0 GB
NVIDIA Nemotron Parse 2.0 is a sub-1-billion-parameter vision-encoder-decoder model purpose-built for document parsing rather than open-ended chat: given a page image, it outputs structured text with layout classes, bounding boxes, and reading order for elements like titles, paragraphs, tables, charts, and footnotes. It pairs a ViT-H vision encoder based on NVIDIA's C-RADIO with a 10-block mBART decoder, and over its predecessor v1.2 it adds roughly 20,000 new vocabulary tokens for more efficient multilingual (especially CJK and Indic-script) OCR, a dedicated chart class for chart-to-table parsing, and stronger table detection and text extraction. At under a billion parameters it runs on a single modest GPU. License is the OpenMDW License Agreement version 1.1, permitting both commercial and non-commercial use; the bundled tokenizer is separately licensed under CC-BY-4.0. It was published in August 2026.
Gemma 3 270M
Google · 268M · runs from 0.1 GB
Google Gemma 3 270M is a 270-million parameter base (pretrained) model from Google's Gemma 3 family. It is an experimental release intended for research, fine-tuning, and exploring the capabilities of ultra-small language models. The model runs on virtually any hardware with negligible resource requirements. Released under the Gemma license.
GPT Neox Japanese 2.7B
abeja · 2.7B · runs from 5.9 GB
GPT Neox Japanese 2.7B is a 2.7B-parameter open language model from abeja. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Open Llama 7B
openlm-research · 7B · runs from 3.3 GB
Open Llama 7B is a 7B-parameter open language model from openlm-research in the Llama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Deepseek Vl2 Tiny
DeepSeek · 3.4B · runs from 7.4 GB
DeepSeek-VL2-Tiny is the smallest in DeepSeek's VL2 series of Mixture-of-Experts vision-language models, with about 3.4 billion total parameters but only roughly 1.2 billion activated per token. It handles vision-language tasks — visual question answering, OCR, document/table/chart understanding, and visual grounding — below the larger VL2-Small and full VL2 variants. Only a fraction of parameters activate per token, so decoding stays fast even though every expert must still be loaded into memory. It is compact enough to run on a single consumer GPU once quantized. Context length is limited to 4,096 tokens. It is released under DeepSeek's own model license, a custom permissive license with an acceptable-use policy rather than a fully open license like MIT or Apache 2.0. Published in December 2024, it pairs a small DeepSeekMoE-3B backbone with a dynamic image-tiling vision encoder.
Glm 4 9B Chat
Z.ai · 9.4B · runs from 4.4 GB
Glm 4 9B Chat is a 9.4B-parameter open language model from Z.ai in the GLM 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Opt 1.3B
Meta · 1.3B · runs from 2.9 GB
Opt 1.3B is a 1.3B-parameter open language model from Meta. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Stories15M MOE
ggml-org · 36M · runs from 0.3 GB
Stories15M MOE is a 36M-parameter open language model from ggml-org. It supports a context window of up to 256 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Inkling Small DSpark Preview
RadixArk · 1.4B · runs from 0.9 GB
Inkling Small DSpark Preview is a 1.4B-parameter open language model from RadixArk in the Inkling family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen Base Invoicev1.01 1.5B
LaaP-ai · 1.5B · runs from 1.0 GB
Qwen Base Invoicev1.01 1.5B is a 1.5B-parameter open language model from LaaP-ai in the Qwen family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Chandra
datalab-to · 8.8B · runs from 4.3 GB
Chandra is Datalab's 8.8-billion-parameter vision-language OCR model, built on a Qwen3VL backbone, that converts document images and PDFs into markdown, HTML, or JSON while preserving layout, tables, math, and forms with checkboxes. It handles handwriting, multi-column and complex layouts, and extracts images and diagrams with captions and structured data across more than 40 languages, aiming at document digitization rather than general chat. On the olmOCR benchmark Chandra scored highest overall among compared models, ahead of Datalab's own Marker pipeline, Mistral's OCR API, DeepSeek-OCR, and anchored GPT-4o and Gemini Flash 2 baselines. At under 9 billion parameters it runs on a single consumer GPU. Context length is 262,144 tokens. It is released under an OpenRAIL license, a responsible-AI license that permits broad use but restricts certain harmful applications. It was published in October 2025 and has since been superseded by a newer Chandra OCR 2 model.
MobileLLaMA 1.4B Chat
mtgv · 1.4B · runs from 1.3 GB
MobileLLaMA 1.4B Chat is a 1.4B-parameter open language model from mtgv in the Llama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
TinyLLama V0
Maykeye · 5M · runs from 0.0 GB
TinyLLama V0 is a 5M-parameter open language model from Maykeye in the TinyLlama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama2 0B Unit Test
MaxJeblick · 770940 · runs from 0.3 GB
Llama2 0B Unit Test is a 770940-parameter open language model from MaxJeblick in the Llama 2 family. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Nl2sh 1.5B Q4 K M
ThorOdinson246 · 1.5B · runs from 0.7 GB
Nl2sh 1.5B Q4 K M is a 1.5B-parameter open language model from ThorOdinson246. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Moonlight 16B A3B
Moonshot AI · 16.0B · runs from 7.5 GB
Moonlight 16B A3B is a compact Mixture-of-Experts model from Moonshot AI that packs 16 billion total parameters while activating only around 3 billion per token. This efficient sparse design lets it punch well above its active parameter count, delivering surprisingly strong chat performance for its effective inference cost. The small active parameter count means Moonlight runs briskly on modest hardware, fitting comfortably on GPUs with 8–12 GB of VRAM at common quantization levels. It is an appealing choice for users who want MoE-level performance diversity without the heavy memory footprint typically associated with mixture models.
Transformer 1.3B 100B
fla-hub · 1.4B · runs from 3 GB
Transformer 1.3B 100B is a 1.4B-parameter open language model from fla-hub. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3 Vision 128k Instruct
Microsoft · 4.1B · runs from 9.4 GB
Phi 3 Vision 128k Instruct is a 4.1B-parameter open language model from Microsoft in the Phi 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Falcon Mamba Tiny Dev
TII UAE · 9M · runs from 0.0 GB
Falcon Mamba Tiny Dev is a 9M-parameter open language model from TII UAE in the Falcon family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
K2 Horizon 7B Uno
IFM · 7B · runs from 3.3 GB
K2 Horizon 7B Uno is a 7B-parameter open language model from IFM. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Distil Lfm25 Shellper
distil-labs · 354M · runs from 0.5 GB
Distil Lfm25 Shellper is a 354M-parameter open language model from distil-labs. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Granite 4.2 3B
IBM · 3.7B · runs from 2.0 GB
Granite 4.2 3B is a 3.7B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
DialoGPT Small
Microsoft · 176M · runs from 0.1 GB
DialoGPT Small is a 176M-parameter open language model from Microsoft. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GLM 5.3 Flash DFlash2
incoai · 1.2B · runs from 0.8 GB
GLM 5.3 Flash DFlash2 is a 1.2B-parameter open language model from incoai in the GLM 5 family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GPT Neo 1.3B
EleutherAI · 1.4B · runs from 3 GB
GPT Neo 1.3B is a 1.4B-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
ERNIE 4.5 21B A3B Thinking
Baidu · 21.8B · runs from 9.7 GB
ERNIE 4.5 21B A3B Thinking is a 21.8B-parameter open language model from Baidu in the ERNIE family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mistral NeMo Minitron 8B Instruct
NVIDIA · 8.4B · runs from 4.2 GB
Mistral NeMo Minitron 8B Instruct is a 8.4B-parameter open language model from NVIDIA in the Mistral family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GPT Bigcode Santacoder
BigCode · 1.1B · runs from 0.5 GB
GPT Bigcode Santacoder is a 1.1B-parameter open language model from BigCode. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Beetle Monolingual Fineweb3 Eng
Beetle-FineWeb3-24B · 194M · runs from 0.4 GB
Beetle Monolingual Fineweb3 Eng is a 194M-parameter open language model from Beetle-FineWeb3-24B. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.