All LLM Models
Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Hunyuan 0.5B Instruct
Tencent · 539M · runs from 0.6 GB
Hunyuan 0.5B Instruct is a 539M-parameter open language model from Tencent in the Hunyuan family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Supra 50M Reasoning
SupraLabs · 52M · runs from 0.3 GB
Supra 50M Reasoning is a 52M-parameter open language model from SupraLabs. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Internlm 20B
InternLM · 20B · runs from 42.8 GB
InternLM-20B is a 20-billion-parameter base language model released by the Shanghai Artificial Intelligence Laboratory with SenseTime, the Chinese University of Hong Kong, and Fudan University, pretrained on over 2.3 trillion tokens of English, Chinese, and code data. It is a pretrained model, not instruction-tuned; a separate InternLM-20B-Chat checkpoint underwent additional SFT and RLHF. Unlike typical 13B models that use 32-40 layers, it uses an unusually deep 60-layer architecture, which the card credits for its strong language, reasoning, and comprehension scores against Llama-13B, Llama2-13B, and Baichuan2-13B, approaching some 33B-70B models on several benchmarks. As an early open Chinese LLM from 2023, it predates today's more capable open models of similar size, and it needs a high-end consumer GPU or multi-GPU setup to run comfortably. Context length is 4,096 tokens, though the card states it can support 16,384 tokens through inference-time extrapolation. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in September 2023.
Kappa 20B 131k Mxfp4
eousphoros · 20.9B · runs from 9.3 GB
Kappa 20B 131k Mxfp4 is a 20.9B-parameter open language model from eousphoros. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Darwin 4B Genesis
FINAL-Bench · 7.5B · runs from 15.6 GB
Darwin 4B Genesis is a 7.5B-parameter open language model from FINAL-Bench. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 4 Quantized.w8a8
RedHatAI · 14.7B · runs from 7.0 GB
Phi 4 Quantized.w8a8 is a 14.7B-parameter open language model from RedHatAI in the Phi 4 family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
UniScientist 30B A3B
UnipatAI · 30.5B · runs from 13.4 GB
UniScientist 30B A3B is a 30.5B-parameter open language model from UnipatAI. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 19B
0xSero · 19.0B · runs from 38.7 GB
Gemma 4 19B is a 19.0B-parameter open language model from 0xSero in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Falcon3 Mamba 7B Base
TII UAE · 7.3B · runs from 16 GB
Falcon3 Mamba 7B Base is a 7.3B-parameter open language model from TII UAE in the Falcon family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Dolphin X1 Trinity Nano
dphn · 6.1B · runs from 3.0 GB
Dolphin X1 Trinity Nano is a 6.1B-parameter open language model from dphn in the Phi family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Apertus 8B MeditronFO
EPFLiGHT · 8.1B · runs from 4.0 GB
Apertus 8B MeditronFO is a 8.1B-parameter open language model from EPFLiGHT in the Apertus family. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Penguin VL 2B
Tencent · 2.2B · runs from 4.9 GB
Penguin VL 2B is a 2.2B-parameter open language model from Tencent. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
QUEST 35B RL
osunlp · 35.1B · runs from 70.6 GB
QUEST 35B RL is a 35.1B-parameter open language model from osunlp. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Intern S2 Preview 397B
InternLM · 404.1B · runs from 808.6 GB
Intern-S2-Preview-397B is InternLM's largest scientific multimodal foundation model, a 404.1-billion-parameter mixture-of-experts model (512 experts, 10 active per token, about 25.1 billion active parameters) positioned as the flagship above the smaller 36.1-billion-parameter Intern-S2-Preview. It introduces a vision-language pretraining paradigm that learns directly from raw pages of scientific literature rather than pre-parsed text, jointly modeling symbolic and visual relationships, and scales scientific reinforcement learning across more than 20 domains alongside long-horizon agent reinforcement learning in large sandboxed environments. InternLM reports leading general-reasoning results among open models plus strong biomolecular-interaction and material-structure generation performance. At its size it requires a multi-GPU server even quantized. Context length is 262,144 tokens, with text-reasoning benchmarks in the card evaluated at up to 256,000 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in July 2026.
Hebatron
HebArabNlpProject · 31.6B · runs from 13.8 GB
Hebatron is a 31.6B-parameter open language model from HebArabNlpProject. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Jan V3.5 4B
janhq · 4.4B · runs from 2.4 GB
Jan V3.5 4B is a 4.4B-parameter open language model from janhq. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
CyberPal2.0 20B
cyber-pal-security · 20.9B · runs from 9.3 GB
CyberPal2.0 20B is a 20.9B-parameter open language model from cyber-pal-security. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Darwin 36B Opus
FINAL-Bench · 34.7B · runs from 69.7 GB
Darwin 36B Opus is a 34.7B-parameter open language model from FINAL-Bench. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Kimi K2.6 Eagle3
NVIDIA · 1.8B · runs from 1.1 GB
Kimi K2.6 Eagle3 is a 1.8B-parameter open language model from NVIDIA in the Kimi K2 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwopus3.5 9B V3.5
Jackrong · 9.7B · runs from 19.9 GB
Qwopus3.5 9B V3.5 is a 9.7B-parameter open language model from Jackrong. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Nemotron H 47B Reasoning 128K
NVIDIA · 46.8B · runs from 102.9 GB
Nemotron H 47B Reasoning 128K is a 46.8B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
II Medical 8B 1706
Intelligent-Internet · 8.2B · runs from 4.1 GB
II Medical 8B 1706 is a 8.2B-parameter open language model from Intelligent-Internet. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Quark 135M
ThingAI · 135M · runs from 0.4 GB
Quark 135M is a 135M-parameter open language model from ThingAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 Omni 3B MNN
taobao-mnn · 3B · runs from 6.6 GB
Qwen2.5 Omni 3B MNN is a 3B-parameter open language model from taobao-mnn in the Qwen 2.5 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Dhara 70M
codelion · 71M · runs from 0.5 GB
Dhara 70M is a 71M-parameter open language model from codelion. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma3 1B CulturaViva ITA
nickprock · 1000M · runs from 0.8 GB
Gemma3 1B CulturaViva ITA is a 1000M-parameter open language model from nickprock in the Gemma 3 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Supra Mini V5 8M
SupraLabs · 8M · runs from 0.3 GB
Supra Mini V5 8M is a 8M-parameter open language model from SupraLabs. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3 14B PARO
z-lab · 1.6B · runs from 1.3 GB
Qwen3 14B PARO is a 1.6B-parameter open language model from z-lab in the Qwen 3 family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Granite Switch 4.1 30B Preview
IBM · 32.2B · runs from 65.3 GB
Granite Switch 4.1 30B Preview is a 32.2B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Claim Extractor 2B Q 2605
principled-intelligence · 2.3B · runs from 5.0 GB
Claim Extractor 2B Q 2605 is a 2.3B-parameter open language model from principled-intelligence. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.