All LLM Models
Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
NCP ArchPreview Dolma3 8.9B Stage2 v3
ArchSpace-Collection · 8.9B · runs from 19.7 GB
NCP ArchPreview Dolma3 8.9B Stage2 v3 is a 8.9B-parameter open language model from ArchSpace-Collection. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Zagreus 0.4B Ita
mii-llm · 438M · runs from 0.6 GB
Zagreus 0.4B Ita is a 438M-parameter open language model from mii-llm. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Xgen 7B 8k Base
Salesforce · 7B · runs from 3.3 GB
XGen-7B-8K-Base is Salesforce AI Research's 7-billion-parameter pretrained base language model, introduced in the 2023 paper "Long Sequence Modeling with XGen: A 7B LLM Trained on 8K Input Sequence Length" as one of the earlier open 7B models built specifically for longer input sequences. It is not instruction-tuned; a separate XGen-7B-8K-Inst checkpoint, released for research purposes only, adds supervised instruction fine-tuning on top of the same base, and a sibling XGen-7B-4K-Base uses a shorter 4K training sequence length. It uses OpenAI's Tiktoken tokenizer rather than a custom vocabulary. At 7 billion parameters it runs easily on a single consumer GPU. Context length is 8,192 tokens, the model's namesake feature. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in June 2023, predating the wave of 7B open models that followed later that year such as Mistral 7B.
Supergemma4 E4b Abliterated
Jiunsong · 7.5B · runs from 3.7 GB
Supergemma4 E4b Abliterated is a 7.5B-parameter open language model from Jiunsong in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Sweep Next Edit v2 7B
sweepai · 7.6B · runs from 3.6 GB
Sweep Next Edit v2 7B is a 7.6B-parameter open language model from sweepai. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Functiongemma 270M Ft Mobile Actions
litert-community · 270M · runs from 0.6 GB
Functiongemma 270M Ft Mobile Actions is a 270M-parameter open language model from litert-community in the Gemma family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
RedPajama INCITE 7B Base
togethercomputer · 7B · runs from 3.3 GB
RedPajama-INCITE-7B-Base is Together Computer's open base pretrained language model, not instruction-tuned, with roughly 6.9 billion parameters trained on the RedPajama-Data-1T dataset, an open reproduction of the corpus used to train Meta's original LLaMA. It was developed with a consortium including Ontocord.ai, ETH DS3Lab, Stanford CRFM and Hazy Research, and LAION, using compute awarded through the 2023 INCITE program. Instruction-tuned and chat variants, RedPajama-INCITE-7B-Instruct and RedPajama-INCITE-7B-Chat, were released alongside it. At under 7 billion parameters it runs on a single consumer GPU. Context length is 2,048 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in May 2023, as one of the first fully open, commercially usable base models trained on openly licensed data.
Nidum Gemma 2B Uncensored
VibeStudio · 2.5B · runs from 1.4 GB
Nidum Gemma 2B Uncensored is a 2.5B-parameter open language model from VibeStudio in the Gemma 2 family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GPT X2 125M
AxiomicLabs · 144M · runs from 0.6 GB
GPT X2 125M is a 144M-parameter open language model from AxiomicLabs. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GPT OSS 20B Heretic
p-e-w · 20.9B · runs from 9.3 GB
GPT OSS 20B Heretic is a 20.9B-parameter open language model from p-e-w in the GPT-OSS family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Nex N2 Pro
nex-agi · 396.8B · runs from 794.0 GB
Nex N2 Pro is a 396.8B-parameter open language model from nex-agi. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
VibeThinker 1.5B
WeiboAI · 1.8B · runs from 1.1 GB
VibeThinker 1.5B is a 1.8B-parameter open language model from WeiboAI. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NVIDIA Nemotron Labs 3 Puzzle 75B A9B BF16
NVIDIA · 75.4B · runs from 165.8 GB
NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 is a deployment-optimized compression of NVIDIA's Nemotron-3-Super-120B-A12B, produced with Iterative Puzzle, a neural-architecture-search-based post-training compression framework, cutting the model from 120.7B total / 12.8B active parameters down to 75.3B total / roughly 9.3B active. It keeps the parent's hybrid architecture of interleaved Mamba, mixture-of-experts, and attention layers with multi-token prediction for faster generation, and the card reports about 2x higher server throughput on a single 8-GPU B200 node while only slightly trailing Nemotron-3-Super on reasoning, coding, agentic, and long-context benchmarks. It is a general-purpose reasoning and chat model intended for agents, chatbots, and RAG systems across seven languages, and needs a multi-GPU server to run. Context length defaults to 262,144 tokens in the published Hugging Face configuration, though the card states the architecture supports up to 1,048,576 tokens at higher memory cost. It is released under the OpenMDW License, version 1.1, a custom license for open model weights, and the card states it was published on July 6, 2026.
SKT OMNI SUPREME
Shrijanagain · 170.7B · runs from 375.6 GB
SKT OMNI SUPREME is a 170.7B-parameter open language model from Shrijanagain. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.6 27B Uncensored HauhauCS Aggressive Safetensor Benchmark
DreamFast · 27.8B · runs from 12.6 GB
Qwen3.6 27B Uncensored HauhauCS Aggressive Safetensor Benchmark is a 27.8B-parameter open language model from DreamFast in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
DeepSeek V4 Flash 180B
0xSero · 101.6B · runs from 43.5 GB
DeepSeek V4 Flash 180B is a 101.6B-parameter open language model from 0xSero in the DeepSeek V4 family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
MiMo V2.5 Pro Base
Xiaomi · 1023.2B · runs from 435.4 GB
MiMo V2.5 Pro Base is a 1023.2B-parameter open language model from Xiaomi. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Chadrock 35B Ace Saber Rocmfp4 Mtp
jcbtc · 35B · runs from 16.4 GB
Chadrock 35B Ace Saber Rocmfp4 Mtp is a 35B-parameter open language model from jcbtc. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
CyberSecQwen 4B
lablab-ai-amd-developer-hackathon · 4.0B · runs from 2.2 GB
CyberSecQwen 4B is a 4.0B-parameter open language model from lablab-ai-amd-developer-hackathon in the Qwen family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama3 OpenBioLLM 70B
aaditya · 70B · runs from 30.7 GB
Llama3 OpenBioLLM 70B is a 70B-parameter open language model from aaditya in the Llama 3 family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Rnj 1.5 Instruct
EssentialAI · 8.3B · runs from 17.2 GB
Rnj 1.5 Instruct is a 8.3B-parameter open language model from EssentialAI. It supports a context window of up to 163,840 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Intern S2 Preview
InternLM · 36.1B · runs from 72.6 GB
Intern-S2-Preview is InternLM's 36.1-billion-parameter scientific multimodal foundation model, a mixture-of-experts model continued-pretrained from Qwen3.5 (256 experts, 8 active per token, about 4.9 billion active parameters) and further trained through a full pipeline from pretraining to reinforcement learning on hundreds of professional scientific tasks. InternLM reports it matches its own trillion-scale Intern-S1-Pro on several core scientific benchmarks despite its much smaller size, adds material crystal-structure generation and small-molecule spatial modeling, and improves agentic tool use for scientific workflows; it also models long, heterogeneous time-series signals. At its size it needs a high-end consumer GPU or multi-GPU setup once quantized. Context length is 262,144 tokens, though the card evaluates text-reasoning benchmarks at up to 128,000 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2026.
MiniMax M1 80k
MiniMax · 456.1B · runs from 913.0 GB
MiniMax M1 80k is a 456.1B-parameter open language model from MiniMax in the MiniMax family. It supports a context window of up to 10,240,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
T Pro IT 2.0
t-tech · 32.8B · runs from 14.6 GB
T Pro IT 2.0 is a 32.8B-parameter open language model from t-tech. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Recursive Language Model 198M
Girinath11 · 198M · runs from 0.4 GB
Recursive Language Model 198M is a 198M-parameter open language model from Girinath11. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
ErniePEUnleashed
Kezmark · 3.4B · runs from 1.9 GB
ErniePEUnleashed is a 3.4B-parameter open language model from Kezmark in the ERNIE family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
GPT S 5M
AxiomicLabs · 5M · runs from 0.3 GB
GPT S 5M is a 5M-parameter open language model from AxiomicLabs. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Falcon H1 7B Base
TII UAE · 7.6B · runs from 3.7 GB
Falcon H1 7B Base is a 7.6B-parameter open language model from TII UAE in the Falcon family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
OmniCoder 9B
Tesslate · 9.4B · runs from 19.4 GB
OmniCoder 9B is a 9.4B-parameter open language model from Tesslate. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
XortronCriminalComputingConfig
darkc0de · 23.6B · runs from 10.7 GB
XortronCriminalComputingConfig is a 23.6B-parameter open language model from darkc0de. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.