All LLM Models
Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Ornith 1.0 397B
Deep Reinforce · 396.8B · runs from 109.5 GB
Ornith-1.0-397B is a Mixture-of-Experts model from Deep Reinforce, the largest member of the original Ornith family, built for agentic coding. It was post-trained on top of Google's Gemma 4 and Alibaba's Qwen3.5 using a self-improving reinforcement-learning framework that learns to generate not just solution rollouts but the scaffolding that drives them, aiming for better search trajectories on tasks like terminal use and repository-scale software engineering. It supports a 262,144-token context window and is released under the MIT license, described by its publisher as globally accessible with no regional limits. At roughly 397 billion parameters, 4-bit quantization needs around 228GB of memory, putting local use in server or multi-GPU territory.
QwQ 32B Preview
Alibaba · 32.8B · runs from 10.7 GB
QwQ 32B Preview is a 32.8B-parameter open language model from Alibaba in the QwQ family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ternary Bonsai 1.7B Unpacked
prism-ml · 1.7B · runs from 1.3 GB
Ternary Bonsai 1.7B Unpacked is a 1.7B-parameter open language model from prism-ml. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Uyu 2 28B
mente-ai · 28.2B · runs from 13.6 GB
Uyu 2 28B is a 28.2B-parameter open language model from mente-ai. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mixtral 8x22B Instruct v0.1
Mistral AI · 140.6B · runs from 58.8 GB
Mixtral-8x22B-Instruct-v0.1 is Mistral AI's instruction-tuned chat model, fine-tuned from the Mixtral-8x22B-v0.1 base model. It is a sparse Mixture-of-Experts model with 8 experts per layer and 2 active per token, giving roughly 39.2 billion active parameters out of about 140.6 billion total, and supports function calling for agentic and tool-use workflows. Because all experts must stay in memory even though only two run per token, it needs a multi-GPU workstation or a high-memory machine even once quantized. Context length is 65,536 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in April 2024, as the instruct variant of the larger successor to Mixtral 8x7B.
DeepSeek V3.2
DeepSeek · 685.4B · runs from 192.3 GB
DeepSeek V3.2 is the latest iteration of DeepSeek's general-purpose flagship, building on the V3 architecture with 685.4 billion total parameters in a mixture-of-experts configuration. This update refines the model's conversational abilities, instruction following, and multilingual performance compared to earlier V3 releases. Running V3.2 locally requires significant GPU resources due to the large total parameter count, though the MoE design means only a subset of parameters are active for any given token. Users with multi-GPU workstations or servers can run quantized versions effectively, making this one of the most powerful open-weight chat models available for self-hosted deployment.
Dream V0 Instruct 7B
Dream-org · 7.6B · runs from 2.5 GB
Dream V0 Instruct 7B is a 7.6B-parameter open language model from Dream-org. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
CodeLlama 7B Instruct HF
Meta · 6.7B · runs from 4.2 GB
CodeLlama 7B Instruct HF is a 6.7B-parameter open language model from Meta in the Code Llama family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NVIDIA Nemotron 3 Super 120B A12B BF16 Heretic
trohrbaugh · 120.7B · runs from 51.8 GB
NVIDIA Nemotron 3 Super 120B A12B BF16 Heretic is a 120.7B-parameter open language model from trohrbaugh in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
MiMo V2.6 Flash RL
Xiaomi · 309B · runs from 93.1 GB
MiMo V2.6 Flash RL is a 309-billion-parameter Mixture-of-Experts language model from Xiaomi's MiMo team, with about 15 billion parameters active per token, keeping generation efficient relative to its total size even though the full weight set must still fit in memory. It is built for chat and agentic use with native function-calling support, and this "Flash" checkpoint is the efficiency-oriented sibling to Xiaomi's larger MiMo-V2.6-Pro. At this scale, local inference needs server-class hardware, so most users rely on a hosted endpoint. It supports a 1,048,576 token (1M) context window. It is released under the MIT license, allowing unrestricted commercial and research use. Published on September 21, 2026, its post-training relied on large-scale, fully asynchronous reinforcement learning (GRPO).
Hy MT2 30B A3B
Tencent · 30.1B · runs from 13.2 GB
Hy-MT2-30B-A3B is Tencent's largest "fast-thinking" multilingual translation model in the Hy-MT2 family, alongside smaller 1.8B and 7B siblings, built specifically for translation rather than general chat and tuned to follow translation instructions across 33 languages. It is a mixture-of-experts model with roughly 30 billion total and 3.5 billion active parameters per token, and Tencent reports it beating open models such as DeepSeek-V4-Pro and Kimi K2.6 on translation quality in fast-thinking mode. The release also ships an FP8-quantized checkpoint, and the smaller 1.8B sibling gets extreme sub-2-bit GGUF quantizations for on-device use. With roughly 3.5 billion active parameters, the 30B-A3B checkpoint is light enough to run on a single consumer GPU once quantized. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2026, alongside the smaller Hy-MT2-1.8B and Hy-MT2-7B models and the IFMTBench translation-instruction benchmark.
Yi 34B
01.AI · 34.4B · runs from 15.0 GB
Yi-34B is 01.AI's base pretrained language model, not instruction-tuned, with about 34.4 billion dense parameters trained on a 3-trillion-token bilingual Chinese-English corpus. At release it ranked first among open-source base models on benchmarks including the Hugging Face Open LLM Leaderboard and C-Eval, ahead of larger models such as Falcon-180B and Llama-2-70B. A long-context Yi-34B-200K variant and an instruction-tuned Yi-34B-Chat were released alongside it. Running the dense 34B model requires a high-end consumer GPU or multi-GPU setup, especially unquantized. Context length is 4,096 tokens, the default window for the 34B series before the extended 200K variant. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in November 2023; by current standards it is an older-generation base model, since superseded by 01.AI's Yi-1.5 and later families.
Supra Router 51M
SupraLabs · 52M · runs from 0.3 GB
Supra Router 51M is a 52M-parameter open language model from SupraLabs. It supports a context window of up to 5,120 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 31B
Google · 32.7B · runs from 15.5 GB
Gemma 4 31B is Google DeepMind's largest dense model in the Gemma 4 family, roughly 32.7 billion parameters, handling text and image input. This is the pretrained base checkpoint rather than an instruction-tuned model, meant as a foundation for fine-tuning rather than direct chat use. Gemma 4 introduces a hybrid attention design interleaving local sliding-window attention with occasional full global attention, plus configurable reasoning modes. At this size, local inference needs a high-end consumer or prosumer GPU, especially once quantized. It supports a 262,144 token context window, among the largest in the family. It carries Google's Gemma 4 license terms, published as Apache 2.0 on Hugging Face with additional Gemma-specific usage terms linked from the card. Published in March 2026, it is the largest of five Gemma 4 sizes (E2B, E4B, 12B, 26B-A4B MoE, 31B dense).
Motif 3
Motif-Technologies · 314.8B · runs from 630.3 GB
Motif 3 is a 314.8B-parameter open language model from Motif-Technologies. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
OLMoE 1B 7B 0924
Allen AI · 6.9B · runs from 3.5 GB
OLMoE 1B 7B 0924 is a 6.9B-parameter open language model from Allen AI in the OLMo family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Kimi K2 Instruct 0905
Moonshot AI · 1026.5B · runs from 286.2 GB
Kimi-K2-Instruct-0905 is Moonshot AI's updated flagship chat and agentic-coding model, a mixture-of-experts system with roughly 1.03 trillion total parameters and about 32.9 billion active per token, routed across 384 experts with 8 selected plus one shared expert per token. It is tuned for agentic coding, tool calling, and long-horizon agent workflows, using multi-head latent attention and SwiGLU activations, and Moonshot reports clear gains on coding-agent benchmarks like SWE-bench and Terminal-Bench, plus improved frontend code aesthetics, over the prior K2-Instruct-0711 release. With a trillion-parameter total footprint it needs a multi-GPU server-class setup even once quantized. Context length is 262,144 tokens, extended from 128K in the previous release. It is released under Moonshot's Modified MIT License, which is otherwise permissive but requires products with more than 100 million monthly active users or $20 million in monthly revenue to display "Kimi K2" in their interface. It was published in September 2025.
Hunyuan A13B Instruct
Tencent · 80.4B · runs from 22.7 GB
Hunyuan A13B Instruct is a 80.4B-parameter open language model from Tencent in the Hunyuan family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ornith 1.0 35B AEON Ultimate Uncensored BF16
AEON-7 · 35.1B · runs from 15.3 GB
Ornith 1.0 35B AEON Ultimate Uncensored BF16 is a 35.1B-parameter open language model from AEON-7 in the Ornith family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Granite 3.0 2B Instruct
IBM · 2.6B · runs from 1.3 GB
Granite-3.0-2B-Instruct is IBM's 2-billion-parameter chat model, fine-tuned from Granite-3.0-2B-Base on a mix of permissively licensed open instruction datasets and IBM's own synthetic data, using supervised fine-tuning, reinforcement-learning alignment, and model merging. It is one of several Granite 3.0 sizes IBM released together, aimed at enterprise use cases such as summarization, question answering, and retrieval-augmented generation. It supports dialogue in twelve languages, primarily English, though multilingual performance trails English performance. At 2B parameters, it is small enough to run on a laptop CPU or any consumer GPU, including edge deployments. Context length is 4,096 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in October 2024. It was later superseded by the Granite 3.1 model family.
TinyLlama 1.1B Intermediate Step 1431k 3T
TinyLlama · 1.1B · runs from 0.8 GB
TinyLlama 1.1B Intermediate Step 1431k 3T is a 1.1B-parameter open language model from TinyLlama in the TinyLlama family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Starcoder2 3B
BigCode · 3.0B · runs from 1.6 GB
StarCoder2-3B is BigCode's 3-billion-parameter base code model, trained from scratch on 17 programming languages drawn from The Stack v2 using a fill-in-the-middle objective over more than 3 trillion tokens. It is a pretrained checkpoint rather than an instruction-tuned assistant, so it completes and continues code rather than following natural-language commands, and it uses grouped-query attention with a sliding-window mechanism for efficiency. It has a 16,384 token context window and is released under the BigCode OpenRAIL-M license, which carries use-based restrictions rather than being fully permissive. At 3 billion parameters, it runs easily on almost any modern laptop or consumer GPU, even without a high-end card, whether quantized or run at half precision.
EuroLLM 22B Instruct 2512
utter-project · 22.6B · runs from 7.0 GB
EuroLLM 22B Instruct 2512 is a 22.6B-parameter open language model from utter-project. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 31B IT Speculator.eagle3
RedHatAI · 31B · runs from 14.5 GB
Gemma 4 31B IT Speculator.eagle3 is a 31B-parameter open language model from RedHatAI in the Gemma 4 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Hermes 2 Theta Llama 3 70B
Nous Research · 70.6B · runs from 20.4 GB
Hermes 2 Theta Llama-3 70B is Nous Research's merged model combining its own Hermes 2 Pro with Meta's Llama-3-70B-Instruct, built with Charles Goddard and Arcee AI's MergeKit tool and then further RLHF-tuned on top of the merge. It uses the ChatML prompt format and is specifically trained for function calling, structured JSON outputs, and feature extraction from retrieval-augmented (RAG) documents, aiming at agentic and tool-using workflows rather than plain chat. At 70 billion parameters it needs a multi-GPU workstation or heavy quantization to run locally. Context length is 8,192 tokens, inherited from its Llama-3 base. It is released under the Llama 3 Community License, Meta's custom license permitting commercial use below 700 million monthly active users. It was published in June 2024.
H2o Danube3 500M Chat
h2oai · 514M · runs from 0.6 GB
H2o Danube3 500M Chat is a 514M-parameter open language model from h2oai. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3 Mini 4k Instruct
Microsoft · 3.8B · runs from 2.7 GB
Microsoft Phi 3 Mini 4K Instruct is a 3.8-billion parameter instruction-tuned model from Microsoft Research's Phi 3 generation, with a 4K token context window. The Phi 3 family demonstrated that small models trained on carefully curated, high-quality data can achieve performance competitive with models several times their size. The model runs on consumer GPUs with as little as 4-6GB of VRAM when quantized, making it one of the most accessible capable chat models for local deployment. Released under the MIT license.
Vikhr Nemo 12B Instruct R 21 09 24
Vikhrmodels · 12.2B · runs from 4.8 GB
Vikhr Nemo 12B Instruct R 21 09 24 is a 12.2B-parameter open language model from Vikhrmodels. It supports a context window of up to 1,024,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ling Flash 2.0
Inclusion AI · 102.9B · runs from 28.7 GB
Ling-flash-2.0 is Inclusion AI's non-reasoning chat and code model, a mixture-of-experts system with about 102.9 billion total parameters but only roughly 6.15 billion activated per token (4.8 billion non-embedding), built on the Ling 2.0 architecture with a roughly 1/32 expert activation ratio, aux-loss-free routing, multi-token-prediction layers, and partial RoPE. Trained on more than 20 trillion tokens with supervised fine-tuning and multi-stage reinforcement learning, Inclusion AI reports it matches dense models of around 40 billion parameters on complex reasoning, code generation, and frontend-development benchmarks despite its small active-parameter count, while running several times faster thanks to its sparsity. Its small active-parameter footprint keeps generation fast, but all of its parameters must stay in memory, so it needs a multi-GPU setup or a high-memory workstation even once quantized. Context length is natively 32,768 tokens, extendable to 128,000 tokens with YaRN. It is released under the MIT license, permitting unrestricted commercial and research use. It was published in September 2025, alongside the reasoning-focused Ring-flash-2.0 built on the same base.
Nanbeige4.1 3B Heretic
heretic-org · 3.9B · runs from 2.1 GB
Nanbeige4.1 3B Heretic is a 3.9B-parameter open language model from heretic-org. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.