All LLM Models
Browse 982 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Gemma 4 12B IT AEON Abliterated K4 BF16
AEON-7 · 12.0B · runs from 6.1 GB
Gemma 4 12B IT AEON Abliterated K4 BF16 is a 12.0B-parameter open language model from AEON-7 in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 Coder 14B
Alibaba · 14.8B · runs from 5.1 GB
Qwen2.5 Coder 14B is a 14.8B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
OvisOCR2
ATH-MaaS · 853M · runs from 0.6 GB
OvisOCR2 is a compact 853-million-parameter end-to-end vision-language model for page-level document OCR, built by post-training Qwen3.5-0.8B with a data engine that mixes real and synthetic documents and a multi-stage supervised fine-tuning, reinforcement learning, and OPD training recipe. Given a document page image it outputs a single Markdown document that reproduces the natural reading order, including body text, formulas rendered as LaTeX, tables as HTML, and captioned image regions, replacing traditional multi-stage OCR pipelines with one model. It set a new state of the art on the OmniDocBench v1.6 leaderboard, the first end-to-end model to top a benchmark long dominated by pipeline-based approaches, and also leads PureDocBench. At under a billion parameters it runs comfortably on a single consumer GPU or even a CPU. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in July 2026.
Gemma 4 31B IT Heretic
coder3101 · 31.3B · runs from 10.2 GB
Gemma 4 31B IT Heretic is a 31.3B-parameter open language model from coder3101 in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Hermes 3 Llama 3.2 3B
Nous Research · 3B · runs from 1.6 GB
Hermes 3 Llama 3.2 3B is a 3-billion parameter instruction-tuned model by Nous Research, fine-tuned from Meta's Llama 3.2 3B base. It applies the Hermes training methodology to a compact model, targeting strong instruction following and conversational quality at minimal hardware cost. Despite its small size, this model benefits from the Hermes fine-tuning approach that emphasizes system prompt adherence and structured output. It can run on GPUs with as little as 4GB of VRAM when quantized, making it suitable for lightweight local deployments and resource-constrained environments.
Ornith 1.5 9B OBLITERATED
OBLITERATUS · 9.7B · runs from 4.7 GB
Ornith 1.5 9B OBLITERATED is a 9.7B-parameter open language model from OBLITERATUS in the Ornith family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Rocinante X 12B V1
TheDrummer · 12.2B · runs from 4.5 GB
Rocinante X 12B V1 is a 12.2B-parameter open language model from TheDrummer. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Nemotron Cascade 2 30B A3B
NVIDIA · 31.6B · runs from 9.1 GB
Nemotron Cascade 2 30B A3B is a 31.6B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3 Medium 128k Instruct
Microsoft · 14.0B · runs from 4.6 GB
Phi-3-Medium-128K-Instruct is Microsoft's 14-billion-parameter instruction-tuned model in the Phi-3 family, trained on a mix of synthetic data and filtered high-quality web content chosen for reasoning density, then post-trained with supervised fine-tuning and direct preference optimization for instruction-following and safety. It is the long-context variant of Phi-3-Medium, alongside a 4K-context sibling, and is aimed at memory- and latency-constrained deployments needing strong code, math, and logical reasoning rather than frontier-scale serving. At 14B parameters, it runs on a single high-end consumer GPU once quantized. Context length is 131,072 tokens. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in May 2024.
OpenHermes 2.5 Mistral 7B
Teknium · 7B · runs from 3.5 GB
OpenHermes 2.5 is a community-driven fine-tune of Mistral 7B created by Teknium, trained on over 900,000 entries of high-quality synthetic data generated primarily by GPT-4. It quickly became one of the most popular open chat models of its era, consistently topping community benchmarks for 7B-class models. For local users, it offers strong instruction-following, creative writing, and coding assistance in a package that runs comfortably on a single consumer GPU with 8 GB of VRAM.
Kimi VL A3B Instruct
Moonshot AI · 16.4B · runs from 5.3 GB
Kimi VL A3B Instruct is Moonshot AI's 16.4-billion-parameter mixture-of-experts vision-language model, with about 3 billion parameters active per token (the A3B in its name). Because only the active experts run for each token, it runs faster than a dense model of similar total size, though the full set of weights must still fit in memory. It pairs a native-resolution vision encoder with a language core built on Moonshot's Moonlight-16B-A3B model, supporting image and video understanding, OCR, multi-image reasoning, and agent-style tool use. At this size, local inference is practical on a single high-end consumer GPU once quantized. The model supports a 131K token context window, useful for long documents or video transcripts. It is released under the MIT license, one of the most permissive open licenses, and was published in April 2025.
Internlm3 8B Instruct
InternLM · 8.8B · runs from 3.4 GB
Internlm3 8B Instruct is a 8.8B-parameter open language model from InternLM in the InternLM family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Huihui GPT OSS 20B BF16 Abliterated
huihui-ai · 20.9B · runs from 9.3 GB
Huihui GPT OSS 20B BF16 Abliterated is a 20.9B-parameter open language model from huihui-ai in the GPT-OSS family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Dolphin 2.9.1 Yi 1.5 34B
dphn · 34.4B · runs from 10.3 GB
Dolphin 2.9.1 Yi 1.5 34B is a 34.4-billion parameter chat model created by Eric Hartford's Dolphin project, fine-tuned from 01.AI's Yi 1.5 34B base. The Dolphin series is known for producing uncensored fine-tunes that remove alignment-based refusals, giving users more direct and unrestricted model responses. This model combines the strong bilingual capabilities of Yi 1.5 with Dolphin's open fine-tuning approach. It requires a GPU with at least 24GB of VRAM for quantized local inference and is popular among users who prefer models without built-in content restrictions.
DeepSeek OCR 2
DeepSeek · 3.4B · runs from 1.9 GB
DeepSeek OCR 2 is DeepSeek's compact vision-language model built for optical character recognition and document understanding, totaling 3.4 billion parameters with about 1.2 billion active per token through its mixture-of-experts design. Only the active parameters compute per token, keeping inference fast, though the full weight set still needs to fit in memory; at this size that's within reach of a consumer GPU or laptop once quantized. It uses a DeepEncoder V2 vision architecture that reasons over document layout semantically instead of scanning images in a fixed pattern. The model has an 8K token context window, suited to single-document OCR passes. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and was published in January 2026 as DeepSeek's OCR successor.
ZDTaichu5.0 9B
TaichuAI · 9.8B · runs from 3.2 GB
ZDTaichu5.0-9B is TaichuAI's 9.8-billion-parameter multimodal foundation model, combining a Qwen3.5-9B language backbone with an NVIDIA C-RADIOv4-H vision encoder to accept text, images, and video at any resolution. Beyond general image, document, and OCR understanding, it is built for spatial reasoning (2D/3D relations, viewpoint and depth, affordances), embodied-AI planning, and multi-step agentic tool use, using an "Entropy-Gated Adaptive Recurrent Reasoning" mechanism that allocates extra recurrent computation to harder tokens. The card reports it leading spatial and agentic benchmarks among comparable 10B-scale open vision-language models. At under 10 billion parameters, it fits on a single consumer GPU once quantized. Context length is 131,072 tokens. It is released under the NVIDIA Open Model License Agreement, with the underlying Qwen3.5 component retaining its Apache 2.0 license, and was published in September 2026.
Devstral Small 2505
Mistral AI · 23.6B · runs from 7.2 GB
Devstral Small 2505, internally called Devstral Small 1.0, is an agentic coding model built by Mistral AI in collaboration with All Hands AI, finetuned from Mistral-Small-3.1-24B-Base-2503 with its vision encoder removed, so it is text-only. It is designed to explore codebases and edit multiple files as a software engineering agent, and reached 46.8% on SWE-Bench Verified under the OpenHands scaffold, the top score among open models at release, ahead of GPT-4.1-mini and Claude 3.5 Haiku on the same benchmark. At 24 billion parameters it is explicitly built to be lightweight enough for a single consumer GPU or an Apple Silicon Mac, making it a real local-deployment option rather than a server-only model. Context length is 131,072 tokens (a 128k window). It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2025; a larger 123-billion-parameter Devstral 2 later succeeded it.
Seed OSS 36B Instruct
ByteDance-Seed · 36.2B · runs from 10.5 GB
Seed OSS 36B Instruct is a 36.2B-parameter open language model from ByteDance-Seed in the Seed-OSS family. It supports a context window of up to 524,288 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Hunyuan 7B Instruct
Tencent · 7.5B · runs from 2.6 GB
Hunyuan 7B Instruct is a 7.5B-parameter open language model from Tencent in the Hunyuan family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Moondream2
vikhyatk · 1.9B · runs from 0.9 GB
Moondream2 is a compact 1.9-billion-parameter vision-language model from the Moondream project, published under developer Vikhyat Korrapati's vikhyatk account, designed for lightweight image understanding rather than open-ended chat. It handles image captioning, visual question answering, object detection, and point-based visual grounding, useful for document pipelines and robotics perception where a full-size multimodal model is unnecessary. Its small size means it runs comfortably on modest consumer GPUs and even CPU-only or edge setups. Because Moondream2 focuses on single-image and short-prompt tasks, it does not advertise an extended context window like larger chat models. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. First published in March 2024, it remains a popular choice for on-device and edge vision applications, with built-in object detection and pointing alongside captioning.
Falcon H1 7B Instruct
TII UAE · 7.6B · runs from 2.6 GB
Falcon H1 7B Instruct is TII UAE's instruction-tuned chat model built on Falcon-H1-7B-Base, using a hybrid architecture that combines standard Transformer attention with Mamba state-space layers in a causal decoder-only design. That hybrid-head approach is meant to keep long-context inference efficient while retaining Transformer-level quality, and the model supports 18 languages including Arabic, Chinese, Japanese, and several European languages. On released benchmarks it beats similarly sized Qwen2.5-7B, Llama-3.1-8B, and Falcon3-7B/10B on most general knowledge, science, and instruction-following tasks. At 7.6 billion parameters it fits a single consumer GPU. Context length is 262,144 tokens. It is released under the Falcon LLM License, a custom TII license that permits commercial use subject to an acceptable-use policy and an attribution requirement in derivative works, and was published in May 2025.
Pythia 410M Deduped
EleutherAI · 506M · runs from 0.1 GB
Pythia 410M Deduped is a 506M-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Magistral Small 2506
Mistral AI · 23.6B · runs from 7.2 GB
Magistral Small 2506 is Mistral AI's small, efficient reasoning model, built on the 24-billion-parameter Mistral Small 3.1 and further trained with supervised fine-tuning on Magistral Medium's reasoning traces plus reinforcement learning. It produces long chains of reasoning before answering and supports dozens of languages, making it suited to multi-step problem solving rather than simple chat. The architecture supports up to 128K tokens, but Mistral recommends keeping usage around 40K since output quality can degrade past that point. Released under the Apache 2.0 license, it is compact enough to fit on a single RTX 4090 or a 32 GB RAM Mac once quantized to 4-bit, in line with what a 24-billion-parameter model needs at that precision.
Qwen2.5 Coder 3B
Alibaba · 3.1B · runs from 1.4 GB
Qwen2.5-Coder-3B is Alibaba's 3.1-billion-parameter base (pretrained, not instruction-tuned) code language model, one of six sizes in the Qwen2.5-Coder family built on Qwen2.5-3B. It was pretrained on 5.5 trillion tokens of source code, text-code grounding data, and synthetic data, and is intended as a foundation for further fine-tuning, reinforcement learning, or fill-in-the-middle code-completion tasks rather than direct chat use; Alibaba explicitly advises against using it for conversations as-is. The family's larger 32B model is described as matching GPT-4o-level coding ability, though this 3B checkpoint targets lightweight, resource-constrained deployment. At just over 3 billion parameters, it runs easily on a single modest consumer GPU or even a CPU. Context length is 32,768 tokens. It is released under the Qwen Research License, a non-commercial license restricted to research and evaluation use; commercial use requires a separate license from Alibaba Cloud. It was published in November 2024.
Pythia 6.9B
EleutherAI · 7.0B · runs from 2.1 GB
Pythia 6.9B is a 7.0B-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 4 Reasoning Plus
Microsoft · 14.7B · runs from 4.8 GB
Phi 4 Reasoning Plus is a 14.7B-parameter open language model from Microsoft in the Phi 4 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Pythia 160M Deduped
EleutherAI · 213M · runs from 0.1 GB
Pythia 160M Deduped is a 213M-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Olmo 3 7B Instruct
Allen AI · 7.3B · runs from 3.4 GB
OLMo 3 7B Instruct is an instruction-tuned language model from the Allen Institute for AI, built as part of their Open Language Model initiative. Like all OLMo releases, it comes with fully open training data, code, and intermediate checkpoints, setting a high standard for reproducibility and scientific transparency in the LLM space. At roughly 7 billion parameters, this model delivers competitive performance on instruction following, reasoning, and general knowledge tasks while remaining runnable on consumer GPUs with 8 GB or more of VRAM. It is an excellent choice for users who value open science and want a capable, well-documented model for local chat and assistant applications.
MiniCPM O 2 6
OpenBMB · 8.7B · runs from 3.4 GB
MiniCPM-o 2.6 is OpenBMB's 8.7-billion-parameter omni-modal model, combining a SigLip-400M vision encoder, a Whisper-medium audio encoder, a ChatTTS speech decoder, and a Qwen2.5-7B language backbone trained end to end. Unlike a pure vision-language model, it takes images, video, and audio as input and can generate spoken output too, supporting real-time bilingual speech conversation, voice cloning, and live-streaming understanding. At this size it is practical on a single high-end consumer GPU once quantized. Context length is 32,768 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in January 2025. Its distinguishing feature is the end-to-end omni-modal architecture, using time-division multiplexing that lets it process continuous video and audio streams and reply with synthesized speech, not just text.
GPT Neox 20B
EleutherAI · 20.7B · runs from 6.3 GB
GPT Neox 20B is a 20.7B-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.