All LLM Models

Browse 1214 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Gemma 4 12B

Google · 12.0B · runs from 6.1 GB

146.8K 749

Gemma 4 12B Unified is a base checkpoint in Google DeepMind's Gemma 4 family, an open-weight multimodal model with roughly 12 billion parameters that takes text, image, and audio input and produces text output. Its "Unified" design skips separate encoders per modality, projecting raw image patches and audio directly into the language model, cutting multimodal latency. As a pretrained release, it's a base for fine-tuning or research rather than direct chat use. It supports a 256K token context window and multilingual pretraining across 140+ languages, and is released under the Apache 2.0 license. At roughly 12 billion parameters, it runs on a single consumer GPU with around 8-12 GB of VRAM once quantized to 4-bit.

Chat

Llm Jp 3 150M

llm-jp · 152M · runs from 0.4 GB

141.5K 8

Llm Jp 3 150M is a 152M-parameter open language model from llm-jp. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Idefics3 8B Llama3

HuggingFaceM4 · 8.5B · runs from 4.2 GB

130.1K 306

Idefics3-8B-Llama3 is Hugging Face's 8.5-billion-parameter open vision-language chat model, accepting arbitrary interleaved sequences of images and text and answering in text: image captioning, visual question answering, multi-image storytelling, or plain text-only chat. It combines a SigLIP-SO400M vision encoder with a Llama-3.1-8B-Instruct language backbone, tiling each image into 364x364 sub-images encoded as 169 visual tokens apiece, which sharply improves OCR and document-understanding scores over the earlier Idefics2. Its post-training is supervised fine-tuning only, without an RLHF stage, so it can give terse answers unless prompted further. At 8.5 billion parameters, it fits on a single consumer GPU once quantized. Context length is 131,072 tokens, inherited from its Llama 3.1 backbone. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in August 2024, succeeding Idefics1 and Idefics2 in the same open multimodal model family.

Vision

Moondream3 Preview

moondream · 9.3B · runs from 20.4 GB

124.4K 685

Moondream 3 (Preview) is a 9.3-billion-parameter vision-language model from Moondream, built as a sparse Mixture-of-Experts with 64 experts per layer and only 2 billion parameters active per token. It targets visual reasoning such as object detection, counting, pointing, and visual question answering, aiming for accuracy closer to larger closed models while keeping inference light. All expert weights must fit in memory, but the active footprint keeps it usable on a single high-end consumer GPU once quantized. Moondream reports a usable context window of around 32,000 tokens. It ships under a Business Source License 1.1 with an additional use grant: free for personal, research, and most commercial use, but resale as a hosted service is blocked. Published in September 2025 as a preview, it introduces a mixture-of-experts design not present in earlier, dense Moondream models.

Vision

Codegen 350M Mono

Salesforce · 350M · runs from 0.8 GB

123.3K 101

Codegen 350M Mono is a 350M-parameter open language model from Salesforce. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Ilama 3.2 1B

hmellor · 1.2B · runs from 2.8 GB

117.7K0

Ilama 3.2 1B is a 1.2B-parameter open language model from hmellor. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

HF Moshiko

kmhf · 7.8B · runs from 16.9 GB

117.2K0

HF Moshiko is a 7.8B-parameter open language model from kmhf. It supports a context window of up to 3,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Xglm 564M

Meta · 564M · runs from 0.3 GB

111.7K 54

Xglm 564M is a 564M-parameter open language model from Meta in the GLM family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

MiMo 7B RL

Xiaomi · 7.8B · runs from 3.9 GB

102.5K 276

MiMo 7B RL is a 7.8B-parameter open language model from Xiaomi. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

NVIDIA Nemotron Parse 2.0

NVIDIA · 903M · runs from 2.0 GB

102.4K 116

NVIDIA Nemotron Parse 2.0 is a sub-1-billion-parameter vision-encoder-decoder model purpose-built for document parsing rather than open-ended chat: given a page image, it outputs structured text with layout classes, bounding boxes, and reading order for elements like titles, paragraphs, tables, charts, and footnotes. It pairs a ViT-H vision encoder based on NVIDIA's C-RADIO with a 10-block mBART decoder, and over its predecessor v1.2 it adds roughly 20,000 new vocabulary tokens for more efficient multilingual (especially CJK and Indic-script) OCR, a dedicated chart class for chart-to-table parsing, and stronger table detection and text extraction. At under a billion parameters it runs on a single modest GPU. License is the OpenMDW License Agreement version 1.1, permitting both commercial and non-commercial use; the bundled tokenizer is separately licensed under CC-BY-4.0. It was published in August 2026.

Vision

Gemma 3 270M

Google · 268M · runs from 0.1 GB

99.2K 1.1K

Google Gemma 3 270M is a 270-million parameter base (pretrained) model from Google's Gemma 3 family. It is an experimental release intended for research, fine-tuning, and exploring the capabilities of ultra-small language models. The model runs on virtually any hardware with negligible resource requirements. Released under the Gemma license.

Chat

GPT Neox Japanese 2.7B

abeja · 2.7B · runs from 5.9 GB

98.3K 59

GPT Neox Japanese 2.7B is a 2.7B-parameter open language model from abeja. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Open Llama 7B

openlm-research · 7B · runs from 3.3 GB

97.3K 139

Open Llama 7B is a 7B-parameter open language model from openlm-research in the Llama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Deepseek Vl2 Tiny

DeepSeek · 3.4B · runs from 7.4 GB

93.9K 253

DeepSeek-VL2-Tiny is the smallest in DeepSeek's VL2 series of Mixture-of-Experts vision-language models, with about 3.4 billion total parameters but only roughly 1.2 billion activated per token. It handles vision-language tasks — visual question answering, OCR, document/table/chart understanding, and visual grounding — below the larger VL2-Small and full VL2 variants. Only a fraction of parameters activate per token, so decoding stays fast even though every expert must still be loaded into memory. It is compact enough to run on a single consumer GPU once quantized. Context length is limited to 4,096 tokens. It is released under DeepSeek's own model license, a custom permissive license with an acceptable-use policy rather than a fully open license like MIT or Apache 2.0. Published in December 2024, it pairs a small DeepSeekMoE-3B backbone with a dynamic image-tiling vision encoder.

Vision

Glm 4 9B Chat

Z.ai · 9.4B · runs from 4.4 GB

92.7K 1.3K

Glm 4 9B Chat is a 9.4B-parameter open language model from Z.ai in the GLM 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Opt 1.3B

Meta · 1.3B · runs from 2.9 GB

92.5K 186

Opt 1.3B is a 1.3B-parameter open language model from Meta. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Stories15M MOE

ggml-org · 36M · runs from 0.3 GB

88.1K 8

Stories15M MOE is a 36M-parameter open language model from ggml-org. It supports a context window of up to 256 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Inkling Small DSpark Preview

RadixArk · 1.4B · runs from 0.9 GB

87.6K 2

Inkling Small DSpark Preview is a 1.4B-parameter open language model from RadixArk in the Inkling family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen Base Invoicev1.01 1.5B

LaaP-ai · 1.5B · runs from 1.0 GB

87.5K0

Qwen Base Invoicev1.01 1.5B is a 1.5B-parameter open language model from LaaP-ai in the Qwen family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Chandra

datalab-to · 8.8B · runs from 4.3 GB

87.1K 532

Chandra is Datalab's 8.8-billion-parameter vision-language OCR model, built on a Qwen3VL backbone, that converts document images and PDFs into markdown, HTML, or JSON while preserving layout, tables, math, and forms with checkboxes. It handles handwriting, multi-column and complex layouts, and extracts images and diagrams with captions and structured data across more than 40 languages, aiming at document digitization rather than general chat. On the olmOCR benchmark Chandra scored highest overall among compared models, ahead of Datalab's own Marker pipeline, Mistral's OCR API, DeepSeek-OCR, and anchored GPT-4o and Gemini Flash 2 baselines. At under 9 billion parameters it runs on a single consumer GPU. Context length is 262,144 tokens. It is released under an OpenRAIL license, a responsible-AI license that permits broad use but restricts certain harmful applications. It was published in October 2025 and has since been superseded by a newer Chandra OCR 2 model.

Vision

Nemotron Labs Diffusion 8B Base

NVIDIA · 8.5B · runs from 17.6 GB

86.5K 8

Nemotron Labs Diffusion 8B Base is a 8.5B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

MobileLLaMA 1.4B Chat

mtgv · 1.4B · runs from 1.3 GB

82.4K 21

MobileLLaMA 1.4B Chat is a 1.4B-parameter open language model from mtgv in the Llama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

TinyLLama V0

Maykeye · 5M · runs from 0.0 GB

81.5K 45

TinyLLama V0 is a 5M-parameter open language model from Maykeye in the TinyLlama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama2 0B Unit Test

MaxJeblick · 770940 · runs from 0.3 GB

79.4K 2

Llama2 0B Unit Test is a 770940-parameter open language model from MaxJeblick in the Llama 2 family. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nl2sh 1.5B Q4 K M

ThorOdinson246 · 1.5B · runs from 0.7 GB

75.4K 63

Nl2sh 1.5B Q4 K M is a 1.5B-parameter open language model from ThorOdinson246. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Moonlight 16B A3B

Moonshot AI · 16.0B · runs from 7.5 GB

72.7K 109

Moonlight 16B A3B is a compact Mixture-of-Experts model from Moonshot AI that packs 16 billion total parameters while activating only around 3 billion per token. This efficient sparse design lets it punch well above its active parameter count, delivering surprisingly strong chat performance for its effective inference cost. The small active parameter count means Moonlight runs briskly on modest hardware, fitting comfortably on GPUs with 8–12 GB of VRAM at common quantization levels. It is an appealing choice for users who want MoE-level performance diversity without the heavy memory footprint typically associated with mixture models.

Chat

Qwen3 30B A3B.w8a8

nytopop · 30.6B · runs from 13.4 GB

72.0K 2

Qwen3 30B A3B.w8a8 is a 30.6B-parameter open language model from nytopop in the Qwen 3 family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Transformer 1.3B 100B

fla-hub · 1.4B · runs from 3 GB

71.4K0

Transformer 1.3B 100B is a 1.4B-parameter open language model from fla-hub. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nemotron Labs Diffusion 8B

NVIDIA · 8.5B · runs from 17.6 GB

70.3K 49

Nemotron Labs Diffusion 8B is a 8.5B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Phi 3 Vision 128k Instruct

Microsoft · 4.1B · runs from 9.4 GB

67.2K 973

Phi 3 Vision 128k Instruct is a 4.1B-parameter open language model from Microsoft in the Phi 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCodeVision