All LLM Models

Browse 880 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Gemma 4 E4B IT Assistant

Google · 4B · runs from 2 GB

52.2K 108

Gemma 4 E4B IT Assistant is a 4B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

MiniCPM5 1B Agentic Tooluse Merged FP16

ewinregirgojr · 1.1B · runs from 0.8 GB

1.7K 3

MiniCPM5 1B Agentic Tooluse Merged FP16 is a 1.1B-parameter open language model from ewinregirgojr in the MiniCPM family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatFunctions

Qwen3 VL 2B Thinking

Alibaba · 2.1B · runs from 1.1 GB

98.9K 117

Qwen3-VL-2B-Thinking is Alibaba's 2.1-billion-parameter vision-language model, the reasoning-enhanced "Thinking" edition of the smallest Qwen3-VL checkpoint, alongside a matching Instruct edition. Beyond image description and visual question answering, it acts as a visual agent capable of operating PC and mobile GUIs, generating code from screenshots or diagrams, and reasoning about 2D and 3D spatial relationships. Its OCR pipeline covers 32 languages and is tuned for low light, blur, and tilted text. At just over 2 billion parameters, it runs easily on a single modest consumer GPU, even unquantized. The model natively supports a 262,144 token context window, expandable to 1,010,000 tokens for long documents or hours of video. It is released under the Apache 2.0 license, and was published in October 2025. Compared with Qwen2.5-VL, Qwen3-VL adds DeepStack vision-feature fusion and timestamp-grounded video event localization.

Vision

Llama 3.2 11B Vision Instruct

Meta · 10.7B · runs from 5.0 GB

65.1K 1.7K

Llama 3.2 11B Vision Instruct is a smaller vision-language model in Meta's Llama 3 family, with roughly 11 billion parameters, built to process images together with text and follow chat-style instructions. It brings multimodal capability, such as visual question answering and image description, to a size that is far more approachable than the 90-billion-parameter variant from the same release. Released in September 2024 under Meta's Llama 3.2 community license, it does not carry a fully permissive open license but does allow broad use under Meta's terms. At around 11 billion parameters, it fits comfortably on a single mid-range to high-end consumer GPU once quantized to 4-bit, putting it within reach of enthusiast local setups.

Vision

OLMoE 1B 7B 0924 Instruct

Allen AI · 6.9B · runs from 3.5 GB

30.9K 96

OLMoE 1B 7B 0924 Instruct is a 6.9B-parameter open language model from Allen AI in the OLMo family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nanonets OCR S

nanonets · 3.8B · runs from 1.4 GB

209.2K 1.6K

Nanonets-OCR-s is Nanonets' image-to-markdown OCR model, built on top of Qwen2.5-VL-3B-Instruct, that goes beyond plain text extraction to produce structured markdown for downstream processing by other language models. It converts mathematical equations and formulas into LaTeX, describes embedded images and charts inside structured tags, isolates signatures and watermarks into their own tags, converts checkboxes into standard Unicode symbols, and extracts complex tables into both markdown and HTML formats. At under 4 billion parameters it is light enough to run on a single consumer GPU. Context length is 128,000 tokens. Nanonets has not published a license for the model on its Hugging Face model card. It was published in June 2025.

Vision

Gemma 4 E4B IT Uncensored

TrevorJS · 8.0B · runs from 3.9 GB

4.4K 33

Gemma 4 E4B IT Uncensored is a 8.0B-parameter open language model from TrevorJS in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

HunyuanOCR

Tencent · 1.1B · runs from 0.7 GB

713.3K 823

HunyuanOCR is Tencent's 1.1-billion-parameter vision-language model built for OCR and document understanding, unifying document parsing, text spotting, information extraction, and text-image translation in one model rather than a general chat assistant. The current checkpoint, HunyuanOCR-1.5, adds DFlash speculative decoding, where a lightweight draft model proposes tokens that the main model verifies in one pass, cutting latency on long structured outputs like tables and formulas. It also ships GGUF weights for llama.cpp, running on CPUs and laptop or consumer GPUs, not just server-grade vLLM. At just over 1 billion parameters, it is light enough for a single modest consumer GPU. Context length is 131,072 tokens; pretraining supports image resolutions up to 4K. It is released under the Tencent Hunyuan Community License, a custom license permitting commercial use outside the EU, UK, and South Korea and below 100 million monthly active users. It was published in November 2025.

Vision

Qwen2 VL 7B Instruct

Alibaba · 8.3B · runs from 3.2 GB

747.3K 1.3K

Qwen2 VL 7B Instruct is Alibaba's 8.3-billion-parameter vision-language model from the original Qwen 2 generation, built to process images and video alongside text in a single conversation. It can describe images, answer visual questions, and reason over multi-image or video input, suiting it to multimodal chat and visual document tasks. At this parameter count, local inference is practical on a single mainstream or high-end consumer GPU once the weights are quantized, rather than requiring multi-GPU hardware. The model supports a 32K token context window, enough for moderate-length documents or multi-turn visual conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in August 2024, it was later superseded by Qwen2.5-VL-7B-Instruct, which its own model card lists as its successor.

Vision

SmolVLM 256M Instruct

Hugging Face · 256M · runs from 0.4 GB

441.5K 411

SmolVLM-256M Instruct is Hugging Face's 256-million-parameter vision-language model, built on a compact SmolLM2-135M language backbone paired with a small 93-million-parameter SigLIP vision encoder. Hugging Face calls it the smallest publicly available multimodal model, designed for image captioning, visual question answering, and basic text transcription from images rather than open-ended chat. Its size makes it practical to run even on CPUs or entry-level GPUs, and on-device deployment is a stated use case. Context length is limited to 8,192 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in January 2025. Compared to the larger 2.2B SmolVLM2 sibling, it uses a smaller vision encoder and more aggressive image-token compression to further cut memory and latency, at some cost to accuracy.

Vision

Granite 4.0 Micro

IBM · 3.4B · runs from 1.4 GB

48.0K 275

Granite-4.0-Micro is IBM's 3-billion-parameter long-context instruct model, fine-tuned from Granite-4.0-Micro-Base on a mix of permissively licensed open datasets and IBM's own synthetic data, using supervised fine-tuning, reinforcement-learning alignment, and model merging. Unlike some of its Granite 4.0 siblings (H Micro, H Tiny MoE, H Small MoE), which mix in Mamba2 layers or mixture-of-experts routing, Granite-4.0-Micro is a plain decoder-only dense transformer with grouped-query attention, RoPE, SwiGLU, and RMSNorm. It targets enterprise use cases such as summarization, RAG, tool-calling, and code completion, and at 3B parameters it runs comfortably on a single consumer GPU, even quantized on a laptop. Context length is 131,072 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2025.

Chat

LFM2.5 1.2B Thinking

Liquid AI · 1.2B · runs from 0.9 GB

10.1K 400

LFM2.5 1.2B Thinking is a 1.2B-parameter open language model from Liquid AI in the LFM2.5 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

K2 Horizon 7B

IFM · 9.0B · runs from 4.4 GB

18.9K 230

K2 Horizon 7B is a 9.0B-parameter open language model from IFM. It supports a context window of up to 524,288 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Mistral 7B v0.1

Mistral AI · 7.2B · runs from 3.6 GB

436.0K 4.2K

Mistral 7B v0.1 is the original base model from Mistral AI that helped reshape expectations for small open-weight language models when it launched in late 2023. As a pretrained foundation model without instruction tuning, it is designed for fine-tuning, research, and custom downstream tasks rather than direct conversational use. With 7 billion parameters and support for grouped-query attention and sliding-window attention, it remains a popular starting point for practitioners building specialized models. Its modest VRAM requirements of roughly 6 GB at 4-bit quantization keep it accessible on a wide range of consumer GPUs.

Chat

DeepSeek R1 Distill Qwen 14B Abliterated v2

huihui-ai · 14.8B · runs from 7.0 GB

692 162

DeepSeek R1 Distill Qwen 14B Abliterated v2 is a 14.8B-parameter open language model from huihui-ai in the DeepSeek R1 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Granite 3.3 8B Instruct

IBM · 8.2B · runs from 2.9 GB

60.2K 161

Granite 3.3 8B Instruct is a 8.2B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Chandra Ocr 2

datalab-to · 5.3B · runs from 2.7 GB

2.5M 518

Chandra OCR 2 is Datalab's 5.3-billion-parameter vision-language model for OCR and document conversion, turning scanned pages and PDFs into markdown, HTML, or JSON while preserving layout. It runs on a Qwen3.5-based backbone mixing linear attention with periodic full-attention layers, and handles handwriting, forms, tables, and math across 90-plus languages. It is the second generation of Datalab's Chandra model, improved over the first release. At just over 5 billion parameters, it runs on a single consumer GPU once quantized. Context extends to 262,144 tokens, generous for multi-page documents. It carries a modified OpenRAIL-M license: free for research, personal use, and startups under $2 million in funding or revenue, but it bars competing against Datalab's own API, more restrictive than a standard permissive license. Published in March 2026, following the original Chandra.

Vision

Gemma 4 12B IT AEON Abliterated K4 BF16

AEON-7 · 12.0B · runs from 6.1 GB

2.3K 25

Gemma 4 12B IT AEON Abliterated K4 BF16 is a 12.0B-parameter open language model from AEON-7 in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoningFunctions

Qwen2.5 Coder 14B

Alibaba · 14.8B · runs from 5.1 GB

50.2K 94

Qwen2.5 Coder 14B is a 14.8B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

OvisOCR2

ATH-MaaS · 853M · runs from 0.6 GB

217.4K 470

OvisOCR2 is a compact 853-million-parameter end-to-end vision-language model for page-level document OCR, built by post-training Qwen3.5-0.8B with a data engine that mixes real and synthetic documents and a multi-stage supervised fine-tuning, reinforcement learning, and OPD training recipe. Given a document page image it outputs a single Markdown document that reproduces the natural reading order, including body text, formulas rendered as LaTeX, tables as HTML, and captioned image regions, replacing traditional multi-stage OCR pipelines with one model. It set a new state of the art on the OmniDocBench v1.6 leaderboard, the first end-to-end model to top a benchmark long dominated by pipeline-based approaches, and also leads PureDocBench. At under a billion parameters it runs comfortably on a single consumer GPU or even a CPU. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in July 2026.

Vision

Hermes 3 Llama 3.2 3B

Nous Research · 3B · runs from 1.6 GB

77.3K 175

Hermes 3 Llama 3.2 3B is a 3-billion parameter instruction-tuned model by Nous Research, fine-tuned from Meta's Llama 3.2 3B base. It applies the Hermes training methodology to a compact model, targeting strong instruction following and conversational quality at minimal hardware cost. Despite its small size, this model benefits from the Hermes fine-tuning approach that emphasizes system prompt adherence and structured output. It can run on GPUs with as little as 4GB of VRAM when quantized, making it suitable for lightweight local deployments and resource-constrained environments.

ChatRoleplay

Ornith 1.5 9B OBLITERATED

OBLITERATUS · 9.7B · runs from 4.7 GB

411.8K 74

Ornith 1.5 9B OBLITERATED is a 9.7B-parameter open language model from OBLITERATUS in the Ornith family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Rocinante X 12B V1

TheDrummer · 12.2B · runs from 4.5 GB

1.3K 92

Rocinante X 12B V1 is a 12.2B-parameter open language model from TheDrummer. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Phi 3 Medium 128k Instruct

Microsoft · 14.0B · runs from 4.6 GB

12.6K 389

Phi-3-Medium-128K-Instruct is Microsoft's 14-billion-parameter instruction-tuned model in the Phi-3 family, trained on a mix of synthetic data and filtered high-quality web content chosen for reasoning density, then post-trained with supervised fine-tuning and direct preference optimization for instruction-following and safety. It is the long-context variant of Phi-3-Medium, alongside a 4K-context sibling, and is aimed at memory- and latency-constrained deployments needing strong code, math, and logical reasoning rather than frontier-scale serving. At 14B parameters, it runs on a single high-end consumer GPU once quantized. Context length is 131,072 tokens. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in May 2024.

ChatCode

OpenHermes 2.5 Mistral 7B

Teknium · 7B · runs from 3.5 GB

151.8K 888

OpenHermes 2.5 is a community-driven fine-tune of Mistral 7B created by Teknium, trained on over 900,000 entries of high-quality synthetic data generated primarily by GPT-4. It quickly became one of the most popular open chat models of its era, consistently topping community benchmarks for 7B-class models. For local users, it offers strong instruction-following, creative writing, and coding assistance in a package that runs comfortably on a single consumer GPU with 8 GB of VRAM.

Chat

Kimi VL A3B Instruct

Moonshot AI · 16.4B · runs from 5.3 GB

200.0K 284

Kimi VL A3B Instruct is Moonshot AI's 16.4-billion-parameter mixture-of-experts vision-language model, with about 3 billion parameters active per token (the A3B in its name). Because only the active experts run for each token, it runs faster than a dense model of similar total size, though the full set of weights must still fit in memory. It pairs a native-resolution vision encoder with a language core built on Moonshot's Moonlight-16B-A3B model, supporting image and video understanding, OCR, multi-image reasoning, and agent-style tool use. At this size, local inference is practical on a single high-end consumer GPU once quantized. The model supports a 131K token context window, useful for long documents or video transcripts. It is released under the MIT license, one of the most permissive open licenses, and was published in April 2025.

VisionFunctions

Internlm3 8B Instruct

InternLM · 8.8B · runs from 3.4 GB

72.5K 235

Internlm3 8B Instruct is a 8.8B-parameter open language model from InternLM in the InternLM family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek OCR 2

DeepSeek · 3.4B · runs from 1.9 GB

863.7K 1.1K

DeepSeek OCR 2 is DeepSeek's compact vision-language model built for optical character recognition and document understanding, totaling 3.4 billion parameters with about 1.2 billion active per token through its mixture-of-experts design. Only the active parameters compute per token, keeping inference fast, though the full weight set still needs to fit in memory; at this size that's within reach of a consumer GPU or laptop once quantized. It uses a DeepEncoder V2 vision architecture that reasons over document layout semantically instead of scanning images in a fixed pattern. The model has an 8K token context window, suited to single-document OCR passes. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in January 2026 as DeepSeek's OCR successor.

Vision

ZDTaichu5.0 9B

TaichuAI · 9.8B · runs from 3.2 GB

6.9K 761

ZDTaichu5.0-9B is TaichuAI's 9.8-billion-parameter multimodal foundation model, combining a Qwen3.5-9B language backbone with an NVIDIA C-RADIOv4-H vision encoder to accept text, images, and video at any resolution. Beyond general image, document, and OCR understanding, it is built for spatial reasoning (2D/3D relations, viewpoint and depth, affordances), embodied-AI planning, and multi-step agentic tool use, using an "Entropy-Gated Adaptive Recurrent Reasoning" mechanism that allocates extra recurrent computation to harder tokens. The card reports it leading spatial and agentic benchmarks among comparable 10B-scale open vision-language models. At under 10 billion parameters, it fits on a single consumer GPU once quantized. Context length is 131,072 tokens. It is released under the NVIDIA Open Model License Agreement, with the underlying Qwen3.5 component retaining its Apache 2.0 license, and was published in September 2026.

VisionFunctions

Devstral Small 2505

Mistral AI · 23.6B · runs from 7.2 GB

2.2K 868

Devstral Small 2505, internally called Devstral Small 1.0, is an agentic coding model built by Mistral AI in collaboration with All Hands AI, finetuned from Mistral-Small-3.1-24B-Base-2503 with its vision encoder removed, so it is text-only. It is designed to explore codebases and edit multiple files as a software engineering agent, and reached 46.8% on SWE-Bench Verified under the OpenHands scaffold, the top score among open models at release, ahead of GPT-4.1-mini and Claude 3.5 Haiku on the same benchmark. At 24 billion parameters it is explicitly built to be lightweight enough for a single consumer GPU or an Apple Silicon Mac, making it a real local-deployment option rather than a server-only model. Context length is 131,072 tokens (a 128k window). It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2025; a larger 123-billion-parameter Devstral 2 later succeeded it.

Chat