All LLM Models

Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Granite 3.3 8B Instruct

IBM · 8.2B · runs from 2.9 GB

60.2K 161

Granite 3.3 8B Instruct is a 8.2B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Chandra Ocr 2

datalab-to · 5.3B · runs from 2.7 GB

2.5M 518

Chandra OCR 2 is Datalab's 5.3-billion-parameter vision-language model for OCR and document conversion, turning scanned pages and PDFs into markdown, HTML, or JSON while preserving layout. It runs on a Qwen3.5-based backbone mixing linear attention with periodic full-attention layers, and handles handwriting, forms, tables, and math across 90-plus languages. It is the second generation of Datalab's Chandra model, improved over the first release. At just over 5 billion parameters, it runs on a single consumer GPU once quantized. Context extends to 262,144 tokens, generous for multi-page documents. It carries a modified OpenRAIL-M license: free for research, personal use, and startups under $2 million in funding or revenue, but it bars competing against Datalab's own API, more restrictive than a standard permissive license. Published in March 2026, following the original Chandra.

Vision

Gemma 4 12B IT AEON Abliterated K4 BF16

AEON-7 · 12.0B · runs from 6.1 GB

2.3K 25

Gemma 4 12B IT AEON Abliterated K4 BF16 is a 12.0B-parameter open language model from AEON-7 in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoningFunctions

Falcon 40B Instruct

TII UAE · 40B · runs from 12.1 GB

5.1K 1.2K

Falcon-40B-Instruct is TII's 40-billion-parameter causal decoder-only chat model, fine-tuned from the Falcon-40B base model on the Baize chat dataset so it is ready to use as an assistant out of the box. Falcon-40B was, at release, the top-ranked fully open model on the Open LLM Leaderboard, and the architecture is optimized for inference with FlashAttention and multi-query attention. It is trained and used primarily in English, with some French capability, and TII recommends further safety evaluation before production use. At 40B parameters it needs a multi-GPU workstation in half precision, while a 4-bit quantization brings it within reach of a single high-memory GPU. License is Apache 2.0, permitting unrestricted commercial and research use. It was published in May 2023, one day after the Falcon-40B base model it is fine-tuned from.

Chat

GLM 4.6V

Z.ai · 107.7B · runs from 30.1 GB

3.6K 396

GLM-4.6V is Z.ai's 107.7-billion-parameter mixture-of-experts vision-language model, the larger of two GLM-4.6V variants (alongside a 9B "Flash" model for local, low-latency deployment) built for visual reasoning, document understanding, and multimodal agents rather than text-only chat. It routes to roughly 14.3 billion active parameters per token and is the first in the GLM-V line to add native multimodal function calling, letting it take images, screenshots, or document pages directly as tool inputs and interpret visual tool outputs like charts or rendered pages as part of its reasoning. It also supports interleaved image-text generation and pixel-accurate UI-to-HTML/CSS reconstruction, and needs a multi-GPU server to run even quantized given its size. Context length is 131,072 tokens. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in December 2025.

Vision

Qwen2.5 Coder 14B

Alibaba · 14.8B · runs from 5.1 GB

50.2K 94

Qwen2.5 Coder 14B is a 14.8B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Mistral Large Instruct 2407

Mistral AI · 122.6B · runs from 37.1 GB

2.5K 865

Mistral-Large-Instruct-2407, also known as Mistral Large 2, is Mistral AI's flagship dense instruction-tuned model of about 123 billion parameters, designed to run efficient single-node inference despite its size. It targets state-of-the-art reasoning, coding, and knowledge tasks with native function calling and JSON output for agentic use, and covers a dozen or more natural languages including French, German, Spanish, Chinese, Japanese, Russian, and Korean alongside more than 80 programming languages. Mistral reports reduced hallucination rates and stronger reasoning compared with the original Mistral Large. Even tuned for single-node deployment, a dense model of this size needs a multi-GPU workstation to run. Context length is 131,072 tokens (a 128k window). It is released under the Mistral AI Research License, a custom license restricting use to non-commercial research purposes; commercial deployment requires a separate license from Mistral AI. It was published in July 2024.

Chat

OvisOCR2

ATH-MaaS · 853M · runs from 0.6 GB

217.4K 470

OvisOCR2 is a compact 853-million-parameter end-to-end vision-language model for page-level document OCR, built by post-training Qwen3.5-0.8B with a data engine that mixes real and synthetic documents and a multi-stage supervised fine-tuning, reinforcement learning, and OPD training recipe. Given a document page image it outputs a single Markdown document that reproduces the natural reading order, including body text, formulas rendered as LaTeX, tables as HTML, and captioned image regions, replacing traditional multi-stage OCR pipelines with one model. It set a new state of the art on the OmniDocBench v1.6 leaderboard, the first end-to-end model to top a benchmark long dominated by pipeline-based approaches, and also leads PureDocBench. At under a billion parameters it runs comfortably on a single consumer GPU or even a CPU. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in July 2026.

Vision

Gemma 4 31B IT Heretic

coder3101 · 31.3B · runs from 10.2 GB

4.9K 66

Gemma 4 31B IT Heretic is a 31.3B-parameter open language model from coder3101 in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

VisionCode

Hermes 3 Llama 3.2 3B

Nous Research · 3B · runs from 1.6 GB

77.3K 175

Hermes 3 Llama 3.2 3B is a 3-billion parameter instruction-tuned model by Nous Research, fine-tuned from Meta's Llama 3.2 3B base. It applies the Hermes training methodology to a compact model, targeting strong instruction following and conversational quality at minimal hardware cost. Despite its small size, this model benefits from the Hermes fine-tuning approach that emphasizes system prompt adherence and structured output. It can run on GPUs with as little as 4GB of VRAM when quantized, making it suitable for lightweight local deployments and resource-constrained environments.

ChatRoleplay

Step 3.5 Flash

StepFun · 199.4B · runs from 60.3 GB

124.0K 837

Step 3.5 Flash is an efficient mixture-of-experts model from StepFun AI, a Chinese AI startup, featuring roughly 199 billion total parameters. The Flash designation signals its focus on speed and low-latency inference, making it well-suited for interactive applications despite its large total parameter count. Running it locally requires a multi-GPU setup, but its MoE architecture means only a portion of the model activates per token, delivering strong multilingual performance with better throughput than a comparably sized dense model.

Chat

GLM 4.6

Z.ai · 356.8B · runs from 98.7 GB

18.1K 1.2K

GLM-4.6 is Z.ai's (Zhipu AI's) flagship chat and agentic model, a mixture-of-experts system with about 356.8 billion total parameters and roughly 34 billion active per token. Compared with the prior GLM-4.5, it extends the context window, improves coding performance in real-world tools such as Claude Code and Cline, strengthens tool-use and search-agent capability during inference, and refines writing style and role-play naturalism; Z.ai reports it is competitive with models like DeepSeek-V3.1-Terminus and Claude Sonnet 4 on agent, reasoning, and coding benchmarks. It shares its inference method and underlying architecture with GLM-4.5. Given its total parameter count, it needs a multi-GPU server-class setup even once quantized. Context length is 200,000 tokens, up from 128K in GLM-4.5. It is released under the MIT license, permitting unrestricted commercial and research use. It was published in September 2025.

Chat

Ornith 1.5 9B OBLITERATED

OBLITERATUS · 9.7B · runs from 4.7 GB

411.8K 74

Ornith 1.5 9B OBLITERATED is a 9.7B-parameter open language model from OBLITERATUS in the Ornith family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3.1 Nemotron 70B Instruct HF

NVIDIA · 70.6B · runs from 20.4 GB

15.6K 2.1K

Llama 3.1 Nemotron 70B Instruct is a 70-billion parameter chat model by NVIDIA, created by applying reinforcement learning from human feedback (RLHF) to Meta's Llama 3.1 70B base model. NVIDIA's Nemotron training pipeline focuses on improving helpfulness, accuracy, and response quality beyond the standard Llama instruction tuning. The model requires substantial VRAM for local inference, typically needing multi-GPU setups or high-end professional GPUs. In quantized formats it becomes accessible on workstation-class hardware. It is available in Hugging Face Transformers format and is supported by popular inference engines.

Chat

DeepSeek V3.1

DeepSeek · 684.5B · runs from 192.1 GB

300.9K 833

DeepSeek-V3.1 is a hybrid instruct model from DeepSeek that supports both thinking and non-thinking modes within a single set of weights, switched via its chat template. It is a roughly 684-billion-parameter Mixture-of-Experts model with 37 billion parameters activated per token, post-trained on top of DeepSeek-V3.1-Base with an extended long-context training phase, adding improved tool calling and dedicated formats for code and search agents. It has a 163,840-token context window and is released under the MIT license. Even at 4-bit quantization it needs roughly 394GB of memory, so local deployment requires server-class or multi-GPU hardware rather than a single machine.

Chat

Rocinante X 12B V1

TheDrummer · 12.2B · runs from 4.5 GB

1.3K 92

Rocinante X 12B V1 is a 12.2B-parameter open language model from TheDrummer. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nemotron Cascade 2 30B A3B

NVIDIA · 31.6B · runs from 9.1 GB

92.0K 515

Nemotron Cascade 2 30B A3B is a 31.6B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Phi 3 Medium 128k Instruct

Microsoft · 14.0B · runs from 4.6 GB

12.6K 389

Phi-3-Medium-128K-Instruct is Microsoft's 14-billion-parameter instruction-tuned model in the Phi-3 family, trained on a mix of synthetic data and filtered high-quality web content chosen for reasoning density, then post-trained with supervised fine-tuning and direct preference optimization for instruction-following and safety. It is the long-context variant of Phi-3-Medium, alongside a 4K-context sibling, and is aimed at memory- and latency-constrained deployments needing strong code, math, and logical reasoning rather than frontier-scale serving. At 14B parameters, it runs on a single high-end consumer GPU once quantized. Context length is 131,072 tokens. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in May 2024.

ChatCode

OpenHermes 2.5 Mistral 7B

Teknium · 7B · runs from 3.5 GB

151.8K 888

OpenHermes 2.5 is a community-driven fine-tune of Mistral 7B created by Teknium, trained on over 900,000 entries of high-quality synthetic data generated primarily by GPT-4. It quickly became one of the most popular open chat models of its era, consistently topping community benchmarks for 7B-class models. For local users, it offers strong instruction-following, creative writing, and coding assistance in a package that runs comfortably on a single consumer GPU with 8 GB of VRAM.

Chat

Kimi VL A3B Instruct

Moonshot AI · 16.4B · runs from 5.3 GB

200.0K 284

Kimi VL A3B Instruct is Moonshot AI's 16.4-billion-parameter mixture-of-experts vision-language model, with about 3 billion parameters active per token (the A3B in its name). Because only the active experts run for each token, it runs faster than a dense model of similar total size, though the full set of weights must still fit in memory. It pairs a native-resolution vision encoder with a language core built on Moonshot's Moonlight-16B-A3B model, supporting image and video understanding, OCR, multi-image reasoning, and agent-style tool use. At this size, local inference is practical on a single high-end consumer GPU once quantized. The model supports a 131K token context window, useful for long documents or video transcripts. It is released under the MIT license, one of the most permissive open licenses, and was published in April 2025.

VisionFunctions

Qwen3.6 35B A3B Abliterated V4

Bahushruth · 34.7B · runs from 12.1 GB

863.6K 9

Qwen3.6 35B A3B Abliterated V4 is a 34.7B-parameter open language model from Bahushruth in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Internlm3 8B Instruct

InternLM · 8.8B · runs from 3.4 GB

72.5K 235

Internlm3 8B Instruct is a 8.8B-parameter open language model from InternLM in the InternLM family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Huihui GPT OSS 20B BF16 Abliterated

huihui-ai · 20.9B · runs from 9.3 GB

13.7K 222

Huihui GPT OSS 20B BF16 Abliterated is a 20.9B-parameter open language model from huihui-ai in the GPT-OSS family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Dolphin 2.9.1 Yi 1.5 34B

dphn · 34.4B · runs from 10.3 GB

4.8M 69

Dolphin 2.9.1 Yi 1.5 34B is a 34.4-billion parameter chat model created by Eric Hartford's Dolphin project, fine-tuned from 01.AI's Yi 1.5 34B base. The Dolphin series is known for producing uncensored fine-tunes that remove alignment-based refusals, giving users more direct and unrestricted model responses. This model combines the strong bilingual capabilities of Yi 1.5 with Dolphin's open fine-tuning approach. It requires a GPU with at least 24GB of VRAM for quantized local inference and is popular among users who prefer models without built-in content restrictions.

Chat

Qwen3 Coder 480B A35B Instruct

Alibaba · 480.2B · runs from 144.6 GB

34.6K 1.4K

Qwen3 Coder 480B A35B Instruct is Alibaba's largest code-specialized model, a massive 480.2-billion-parameter mixture-of-experts system with roughly 35 billion parameters active per token. This is the most powerful open-weight coding model in the Qwen3 family, designed for professional-grade code generation, analysis, and software engineering tasks. Running this model locally is a serious undertaking that requires multi-GPU server-class hardware with several hundred gigabytes of combined VRAM. For users with access to such infrastructure, it offers exceptional code quality and understanding that rivals leading proprietary coding assistants, all while keeping data and computation entirely under local control.

ChatCode

Qwen3.8 27B TURBO Fable Cold Fusion 735 882 Heretic Uncensored NM DAU

DavidAU · 27.8B · runs from 56.3 GB

2.6K 159

Qwen3.8 27B TURBO Fable Cold Fusion 735 882 Heretic Uncensored NM DAU is a 27.8B-parameter open language model from DavidAU in the Qwen 3.8 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

DeepSeek OCR 2

DeepSeek · 3.4B · runs from 1.9 GB

863.7K 1.1K

DeepSeek OCR 2 is DeepSeek's compact vision-language model built for optical character recognition and document understanding, totaling 3.4 billion parameters with about 1.2 billion active per token through its mixture-of-experts design. Only the active parameters compute per token, keeping inference fast, though the full weight set still needs to fit in memory; at this size that's within reach of a consumer GPU or laptop once quantized. It uses a DeepEncoder V2 vision architecture that reasons over document layout semantically instead of scanning images in a fixed pattern. The model has an 8K token context window, suited to single-document OCR passes. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in January 2026 as DeepSeek's OCR successor.

Vision

ZDTaichu5.0 9B

TaichuAI · 9.8B · runs from 3.2 GB

6.9K 761

ZDTaichu5.0-9B is TaichuAI's 9.8-billion-parameter multimodal foundation model, combining a Qwen3.5-9B language backbone with an NVIDIA C-RADIOv4-H vision encoder to accept text, images, and video at any resolution. Beyond general image, document, and OCR understanding, it is built for spatial reasoning (2D/3D relations, viewpoint and depth, affordances), embodied-AI planning, and multi-step agentic tool use, using an "Entropy-Gated Adaptive Recurrent Reasoning" mechanism that allocates extra recurrent computation to harder tokens. The card reports it leading spatial and agentic benchmarks among comparable 10B-scale open vision-language models. At under 10 billion parameters, it fits on a single consumer GPU once quantized. Context length is 131,072 tokens. It is released under the NVIDIA Open Model License Agreement, with the underlying Qwen3.5 component retaining its Apache 2.0 license, and was published in September 2026.

VisionFunctions

Devstral Small 2505

Mistral AI · 23.6B · runs from 7.2 GB

2.2K 868

Devstral Small 2505, internally called Devstral Small 1.0, is an agentic coding model built by Mistral AI in collaboration with All Hands AI, finetuned from Mistral-Small-3.1-24B-Base-2503 with its vision encoder removed, so it is text-only. It is designed to explore codebases and edit multiple files as a software engineering agent, and reached 46.8% on SWE-Bench Verified under the OpenHands scaffold, the top score among open models at release, ahead of GPT-4.1-mini and Claude 3.5 Haiku on the same benchmark. At 24 billion parameters it is explicitly built to be lightweight enough for a single consumer GPU or an Apple Silicon Mac, making it a real local-deployment option rather than a server-only model. Context length is 131,072 tokens (a 128k window). It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2025; a larger 123-billion-parameter Devstral 2 later succeeded it.

Chat

DeepSeek V2.5

DeepSeek · 235.7B · runs from 67.7 GB

6.8K 735

DeepSeek-V2.5 is DeepSeek's merged general-purpose and coding chat model, combining DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct into a single mixture-of-experts checkpoint with about 235.7 billion total parameters and roughly 21.4 billion active per token. DeepSeek reports it better aligns with human preferences and improves writing and instruction-following over its two predecessors, with gains on AlpacaEval, ArenaHard, and coding benchmarks like HumanEval and LiveCodeBench. It is a chat, not base, model. Given its total parameter count, DeepSeek states that running it in BF16 requires eight 80GB-class GPUs, so it needs a substantial multi-GPU server, not a single consumer machine, even once quantized. Context length is 163,840 tokens. Model weights are released under DeepSeek's custom model license, which permits commercial use, while the surrounding code is MIT-licensed. It was published in September 2024, unifying the separate V2-Chat and Coder-V2-Instruct lines into one model.

Chat