All LLM Models

Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

NCP ArchPreview Dolma3 8.9B Stage2 v3

ArchSpace-Collection · 8.9B · runs from 19.7 GB

1.0K 4

NCP ArchPreview Dolma3 8.9B Stage2 v3 is a 8.9B-parameter open language model from ArchSpace-Collection. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Zagreus 0.4B Ita

mii-llm · 438M · runs from 0.6 GB

1.0K 10

Zagreus 0.4B Ita is a 438M-parameter open language model from mii-llm. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Xgen 7B 8k Base

Salesforce · 7B · runs from 3.3 GB

997 317

XGen-7B-8K-Base is Salesforce AI Research's 7-billion-parameter pretrained base language model, introduced in the 2023 paper "Long Sequence Modeling with XGen: A 7B LLM Trained on 8K Input Sequence Length" as one of the earlier open 7B models built specifically for longer input sequences. It is not instruction-tuned; a separate XGen-7B-8K-Inst checkpoint, released for research purposes only, adds supervised instruction fine-tuning on top of the same base, and a sibling XGen-7B-4K-Base uses a shorter 4K training sequence length. It uses OpenAI's Tiktoken tokenizer rather than a custom vocabulary. At 7 billion parameters it runs easily on a single consumer GPU. Context length is 8,192 tokens, the model's namesake feature. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in June 2023, predating the wave of 7B open models that followed later that year such as Mistral 7B.

Chat

Supergemma4 E4b Abliterated

Jiunsong · 7.5B · runs from 3.7 GB

996 72

Supergemma4 E4b Abliterated is a 7.5B-parameter open language model from Jiunsong in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Sweep Next Edit v2 7B

sweepai · 7.6B · runs from 3.6 GB

995 32

Sweep Next Edit v2 7B is a 7.6B-parameter open language model from sweepai. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Functiongemma 270M Ft Mobile Actions

litert-community · 270M · runs from 0.6 GB

994 232

Functiongemma 270M Ft Mobile Actions is a 270M-parameter open language model from litert-community in the Gemma family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

RedPajama INCITE 7B Base

togethercomputer · 7B · runs from 3.3 GB

988 92

RedPajama-INCITE-7B-Base is Together Computer's open base pretrained language model, not instruction-tuned, with roughly 6.9 billion parameters trained on the RedPajama-Data-1T dataset, an open reproduction of the corpus used to train Meta's original LLaMA. It was developed with a consortium including Ontocord.ai, ETH DS3Lab, Stanford CRFM and Hazy Research, and LAION, using compute awarded through the 2023 INCITE program. Instruction-tuned and chat variants, RedPajama-INCITE-7B-Instruct and RedPajama-INCITE-7B-Chat, were released alongside it. At under 7 billion parameters it runs on a single consumer GPU. Context length is 2,048 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in May 2023, as one of the first fully open, commercially usable base models trained on openly licensed data.

Chat

Nidum Gemma 2B Uncensored

VibeStudio · 2.5B · runs from 1.4 GB

980 5

Nidum Gemma 2B Uncensored is a 2.5B-parameter open language model from VibeStudio in the Gemma 2 family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GPT X2 125M

AxiomicLabs · 144M · runs from 0.6 GB

979 23

GPT X2 125M is a 144M-parameter open language model from AxiomicLabs. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GPT OSS 20B Heretic

p-e-w · 20.9B · runs from 9.3 GB

978 132

GPT OSS 20B Heretic is a 20.9B-parameter open language model from p-e-w in the GPT-OSS family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nex N2 Pro

nex-agi · 396.8B · runs from 794.0 GB

973 376

Nex N2 Pro is a 396.8B-parameter open language model from nex-agi. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

VibeThinker 1.5B

WeiboAI · 1.8B · runs from 1.1 GB

968 524

VibeThinker 1.5B is a 1.8B-parameter open language model from WeiboAI. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatMathCode

NVIDIA Nemotron Labs 3 Puzzle 75B A9B BF16

NVIDIA · 75.4B · runs from 165.8 GB

962 68

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 is a deployment-optimized compression of NVIDIA's Nemotron-3-Super-120B-A12B, produced with Iterative Puzzle, a neural-architecture-search-based post-training compression framework, cutting the model from 120.7B total / 12.8B active parameters down to 75.3B total / roughly 9.3B active. It keeps the parent's hybrid architecture of interleaved Mamba, mixture-of-experts, and attention layers with multi-token prediction for faster generation, and the card reports about 2x higher server throughput on a single 8-GPU B200 node while only slightly trailing Nemotron-3-Super on reasoning, coding, agentic, and long-context benchmarks. It is a general-purpose reasoning and chat model intended for agents, chatbots, and RAG systems across seven languages, and needs a multi-GPU server to run. Context length defaults to 262,144 tokens in the published Hugging Face configuration, though the card states the architecture supports up to 1,048,576 tokens at higher memory cost. It is released under the OpenMDW License, version 1.1, a custom license for open model weights, and the card states it was published on July 6, 2026.

Chat

SKT OMNI SUPREME

Shrijanagain · 170.7B · runs from 375.6 GB

962 5

SKT OMNI SUPREME is a 170.7B-parameter open language model from Shrijanagain. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatFunctions

Qwen3.6 27B Uncensored HauhauCS Aggressive Safetensor Benchmark

DreamFast · 27.8B · runs from 12.6 GB

937 5

Qwen3.6 27B Uncensored HauhauCS Aggressive Safetensor Benchmark is a 27.8B-parameter open language model from DreamFast in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek V4 Flash 180B

0xSero · 101.6B · runs from 43.5 GB

924 10

DeepSeek V4 Flash 180B is a 101.6B-parameter open language model from 0xSero in the DeepSeek V4 family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

MiMo V2.5 Pro Base

Xiaomi · 1023.2B · runs from 435.4 GB

899 40

MiMo V2.5 Pro Base is a 1023.2B-parameter open language model from Xiaomi. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatFunctionsCode

Chadrock 35B Ace Saber Rocmfp4 Mtp

jcbtc · 35B · runs from 16.4 GB

896 7

Chadrock 35B Ace Saber Rocmfp4 Mtp is a 35B-parameter open language model from jcbtc. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

CyberSecQwen 4B

lablab-ai-amd-developer-hackathon · 4.0B · runs from 2.2 GB

884 14

CyberSecQwen 4B is a 4.0B-parameter open language model from lablab-ai-amd-developer-hackathon in the Qwen family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama3 OpenBioLLM 70B

aaditya · 70B · runs from 30.7 GB

862 517

Llama3 OpenBioLLM 70B is a 70B-parameter open language model from aaditya in the Llama 3 family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Rnj 1.5 Instruct

EssentialAI · 8.3B · runs from 17.2 GB

844 20

Rnj 1.5 Instruct is a 8.3B-parameter open language model from EssentialAI. It supports a context window of up to 163,840 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Intern S2 Preview

InternLM · 36.1B · runs from 72.6 GB

842 119

Intern-S2-Preview is InternLM's 36.1-billion-parameter scientific multimodal foundation model, a mixture-of-experts model continued-pretrained from Qwen3.5 (256 experts, 8 active per token, about 4.9 billion active parameters) and further trained through a full pipeline from pretraining to reinforcement learning on hundreds of professional scientific tasks. InternLM reports it matches its own trillion-scale Intern-S1-Pro on several core scientific benchmarks despite its much smaller size, adds material crystal-structure generation and small-molecule spatial modeling, and improves agentic tool use for scientific workflows; it also models long, heterogeneous time-series signals. At its size it needs a high-end consumer GPU or multi-GPU setup once quantized. Context length is 262,144 tokens, though the card evaluates text-reasoning benchmarks at up to 128,000 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2026.

Vision

MiniMax M1 80k

MiniMax · 456.1B · runs from 913.0 GB

834 692

MiniMax M1 80k is a 456.1B-parameter open language model from MiniMax in the MiniMax family. It supports a context window of up to 10,240,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

T Pro IT 2.0

t-tech · 32.8B · runs from 14.6 GB

834 126

T Pro IT 2.0 is a 32.8B-parameter open language model from t-tech. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Recursive Language Model 198M

Girinath11 · 198M · runs from 0.4 GB

787 10

Recursive Language Model 198M is a 198M-parameter open language model from Girinath11. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

ErniePEUnleashed

Kezmark · 3.4B · runs from 1.9 GB

781 4

ErniePEUnleashed is a 3.4B-parameter open language model from Kezmark in the ERNIE family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GPT S 5M

AxiomicLabs · 5M · runs from 0.3 GB

769 12

GPT S 5M is a 5M-parameter open language model from AxiomicLabs. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Falcon H1 7B Base

TII UAE · 7.6B · runs from 3.7 GB

746 11

Falcon H1 7B Base is a 7.6B-parameter open language model from TII UAE in the Falcon family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

OmniCoder 9B

Tesslate · 9.4B · runs from 19.4 GB

744 683

OmniCoder 9B is a 9.4B-parameter open language model from Tesslate. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCodeFunctions

XortronCriminalComputingConfig

darkc0de · 23.6B · runs from 10.7 GB

721 137

XortronCriminalComputingConfig is a 23.6B-parameter open language model from darkc0de. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat