All LLM Models

Browse 1214 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Qwen3.8 27B Uncensored

orcarouter · 27.8B · runs from 13.0 GB

67.0K 232

Qwen3.8 27B Uncensored is a 27.8B-parameter open language model from orcarouter in the Qwen 3.8 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

VisionFunctionsReasoning

Falcon Mamba Tiny Dev

TII UAE · 9M · runs from 0.0 GB

62.4K 2

Falcon Mamba Tiny Dev is a 9M-parameter open language model from TII UAE in the Falcon family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nemotron H 8B Base 8K

NVIDIA · 8.1B · runs from 17.8 GB

60.9K 60

Nemotron H 8B Base 8K is a 8.1B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Opt 6.7B

Meta · 6.7B · runs from 14.7 GB

58.7K 121

Opt 6.7B is a 6.7B-parameter open language model from Meta. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

K2 Horizon 7B Uno

IFM · 7B · runs from 3.3 GB

58.4K 96

K2 Horizon 7B Uno is a 7B-parameter open language model from IFM. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Distil Lfm25 Shellper

distil-labs · 354M · runs from 0.5 GB

57.0K 11

Distil Lfm25 Shellper is a 354M-parameter open language model from distil-labs. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatFunctions

XCurOS0.1 8B Instruct

XCurOS · 7.6B · runs from 15.7 GB

56.8K 4

XCurOS0.1 8B Instruct is a 7.6B-parameter open language model from XCurOS. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

NVIDIA Nemotron 3 Nano 30B A3B Base BF16

NVIDIA · 31.6B · runs from 14.8 GB

55.9K 132

NVIDIA Nemotron 3 Nano 30B A3B Base BF16 is the foundation model version of the Nemotron 3 Nano 30B, offered in full BF16 precision. Unlike the chat-tuned variants, this base model hasn't been instruction-tuned, making it suitable for fine-tuning, research, or custom alignment workflows. At 31.6 billion total parameters with a mixture-of-experts architecture, the base model gives developers and researchers a strong starting point for building specialized applications. It retains all the architectural benefits of the MoE design while leaving the behavioral layer open for customization.

Chat

Qwen3.8 27B EfficientThink Uncensored K3 Opus5 Grok4.6 GPT5.6Sol SFT SimPO DFlash2

nerkyor · 27B · runs from 12.6 GB

49.3K 44

Qwen3.8 27B EfficientThink Uncensored K3 Opus5 Grok4.6 GPT5.6Sol SFT SimPO DFlash2 is a 27B-parameter open language model from nerkyor in the Qwen 3.8 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoningFunctions

Granite 4.2 3B

IBM · 3.7B · runs from 2.0 GB

49.1K 97

Granite 4.2 3B is a 3.7B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

DialoGPT Small

Microsoft · 176M · runs from 0.1 GB

48.5K 147

DialoGPT Small is a 176M-parameter open language model from Microsoft. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GLM 5.3 Flash DFlash2

incoai · 1.2B · runs from 0.8 GB

47.7K 115

GLM 5.3 Flash DFlash2 is a 1.2B-parameter open language model from incoai in the GLM 5 family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GPT Neo 1.3B

EleutherAI · 1.4B · runs from 3 GB

47.2K 326

GPT Neo 1.3B is a 1.4B-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

ERNIE 4.5 21B A3B Thinking

Baidu · 21.8B · runs from 9.7 GB

46.9K 792

ERNIE 4.5 21B A3B Thinking is a 21.8B-parameter open language model from Baidu in the ERNIE family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama Guard 3 8B

Meta · 8.0B · runs from 17.7 GB

46.1K 327

Meta Llama Guard 3 8B is an 8-billion parameter safety classifier model built on the Llama 3.1 architecture. Unlike general-purpose chat models, Llama Guard is specifically designed to classify whether prompts or responses contain unsafe content across categories such as violence, sexual content, criminal planning, and other policy violations. The model is intended to be used as a moderation layer in LLM-based applications, providing input and output safety filtering. It follows a taxonomy-based classification approach and can be customized for different safety policies. Released under the Llama 3.1 Community License.

Chat

Mistral NeMo Minitron 8B Instruct

NVIDIA · 8.4B · runs from 4.2 GB

45.9K 85

Mistral NeMo Minitron 8B Instruct is a 8.4B-parameter open language model from NVIDIA in the Mistral family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GPT Bigcode Santacoder

BigCode · 1.1B · runs from 0.5 GB

45.8K 28

GPT Bigcode Santacoder is a 1.1B-parameter open language model from BigCode. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Beetle Monolingual Fineweb3 Eng

Beetle-FineWeb3-24B · 194M · runs from 0.4 GB

45.8K0

Beetle Monolingual Fineweb3 Eng is a 194M-parameter open language model from Beetle-FineWeb3-24B. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

PowerLM 3B

ibm-research · 3.5B · runs from 2.5 GB

45.7K 21

PowerLM 3B is a 3.5B-parameter open language model from ibm-research. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Opt 2.7B

Meta · 2.7B · runs from 5.9 GB

45.6K 89

Opt 2.7B is a 2.7B-parameter open language model from Meta. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Stablelm 3B 4e1t

Stability AI · 2.8B · runs from 2.2 GB

43.7K 316

Stablelm 3B 4e1t is a 2.8B-parameter open language model from Stability AI in the StableLM family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Tiny LLM

arnir0 · 13M · runs from 0.3 GB

43.6K 69

Tiny LLM is a 13M-parameter open language model from arnir0. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3 8B Instruct Gradient 1048k

gradientai · 8.0B · runs from 4.0 GB

43.5K 683

Llama 3 8B Instruct Gradient 1048k is a 8.0B-parameter open language model from gradientai in the Llama 3 family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Granite 3.1 8B Instruct

IBM · 8.2B · runs from 4.1 GB

43.4K 169

Granite 3.1 8B Instruct is a 8.2B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Beetle Monolingual Humanscale Eng

Beetle-HumanScale · 194M · runs from 0.4 GB

43.1K0

Beetle Monolingual Humanscale Eng is a 194M-parameter open language model from Beetle-HumanScale. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Typhoon2.5 Qwen3 4B

typhoon-ai · 4.0B · runs from 2.2 GB

42.9K 10

Typhoon2.5 Qwen3 4B is a 4.0B-parameter open language model from typhoon-ai in the Qwen 3 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GPT OSS Safeguard 20B

OpenAI · 21.5B · runs from 9.5 GB

42.0K 263

GPT OSS Safeguard 20B is a 21.5B-parameter open language model from OpenAI in the GPT-OSS family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Granite 3.0 8B Instruct

IBM · 8.2B · runs from 4.1 GB

41.9K 208

Granite 3.0 8B Instruct is a 8.2B-parameter open language model from IBM in the Granite family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

SedibaLM

Sedibaai · 1.5B · runs from 1.0 GB

41.7K0

SedibaLM is a 1.5B-parameter open language model from Sedibaai. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

EXAONE 3.0 7.8B Instruct

LGAI-EXAONE · 7.8B · runs from 17.2 GB

41.7K 420

EXAONE 3.0 7.8B Instruct is a 7.8B-parameter open language model from LGAI-EXAONE in the EXAONE family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat