All LLM Models

Browse 57 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Nemotron Cascade 8B

NVIDIA · 8B · runs from 4 GB

31.7K 65

Nemotron Cascade 8B is a 8B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Nemotron Labs Diffusion 14B

NVIDIA · 13.5B · runs from 6.5 GB

7.1K 143

Nemotron Labs Diffusion 14B is a 13.5B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3 1 Nemotron Ultra 253B V1

NVIDIA · 253.4B · runs from 557.5 GB

5.0K 352

Llama 3 1 Nemotron Ultra 253B V1 is a 253.4B-parameter open language model from NVIDIA in the Llama 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

NVIDIA Nemotron 3 Super 120B A12B Base BF16

NVIDIA · 123.6B · runs from 53.0 GB

3.8K 19

NVIDIA Nemotron 3 Super 120B A12B Base BF16 is a 123.6B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

OpenMath Nemotron 1.5B

NVIDIA · 1.5B · runs from 1.0 GB

3.5K 29

OpenMath Nemotron 1.5B is a 1.5B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatMath

Nemotron Research Reasoning Qwen 1.5B

NVIDIA · 1.8B · runs from 1.1 GB

2.6K 243

Nemotron Research Reasoning Qwen 1.5B is a 1.8B-parameter open language model from NVIDIA in the Qwen family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Nemotron Labs Audex 2B

NVIDIA · 2B · runs from 4.4 GB

2.3K 76

Nemotron Labs Audex 2B is a 2B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Nemotron Terminal 8B

NVIDIA · 8.2B · runs from 4.1 GB

2.3K 26

Nemotron Terminal 8B is a 8.2B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nemotron Content Safety Reasoning 4B

NVIDIA · 4.3B · runs from 2.5 GB

2.3K 19

Nemotron Content Safety Reasoning 4B is a 4.3B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Kimi K2.6 DFlash

NVIDIA · 3.5B · runs from 1.8 GB

2.1K 25

Kimi K2.6 DFlash is a 3.5B-parameter open language model from NVIDIA in the Kimi K2 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nemotron 3 Labs Ultra Math SFT

NVIDIA · 560.5B · runs from 238.8 GB

2.1K 7

Nemotron 3 Labs Ultra Math SFT is a 560.5B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatMath

Llama 3.1 Nemotron Safety Guard 8B v3

NVIDIA · 8.0B · runs from 4.0 GB

1.7K 13

Llama 3.1 Nemotron Safety Guard 8B v3 is a 8.0B-parameter open language model from NVIDIA in the Llama 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Kimi K2.7 Code DFlash

NVIDIA · 3.5B · runs from 1.8 GB

1.3K 10

Kimi K2.7 Code DFlash is a 3.5B-parameter open language model from NVIDIA in the Kimi K2 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Nemotron Terminal 32B

NVIDIA · 32.8B · runs from 14.6 GB

1.3K 30

Nemotron Terminal 32B is a 32.8B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

NVIDIA Nemotron Labs 3 Puzzle 75B A9B BF16

NVIDIA · 75.4B · runs from 165.8 GB

962 68

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 is a deployment-optimized compression of NVIDIA's Nemotron-3-Super-120B-A12B, produced with Iterative Puzzle, a neural-architecture-search-based post-training compression framework, cutting the model from 120.7B total / 12.8B active parameters down to 75.3B total / roughly 9.3B active. It keeps the parent's hybrid architecture of interleaved Mamba, mixture-of-experts, and attention layers with multi-token prediction for faster generation, and the card reports about 2x higher server throughput on a single 8-GPU B200 node while only slightly trailing Nemotron-3-Super on reasoning, coding, agentic, and long-context benchmarks. It is a general-purpose reasoning and chat model intended for agents, chatbots, and RAG systems across seven languages, and needs a multi-GPU server to run. Context length defaults to 262,144 tokens in the published Hugging Face configuration, though the card states the architecture supports up to 1,048,576 tokens at higher memory cost. It is released under the OpenMDW License, version 1.1, a custom license for open model weights, and the card states it was published on July 6, 2026.

Chat

OpenReasoning Nemotron 32B

NVIDIA · 32.8B · runs from 14.8 GB

702 126

OpenReasoning Nemotron 32B is a 32.8B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCodeReasoning

Nemotron H 8B Reasoning 128K

NVIDIA · 8.1B · runs from 17.8 GB

628 26

Nemotron H 8B Reasoning 128K is a 8.1B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Llama 3 1 Nemotron 51B Instruct

NVIDIA · 51B · runs from 112.2 GB

570 209

Llama 3 1 Nemotron 51B Instruct is a 51B-parameter open language model from NVIDIA in the Llama 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Kimi K2.6 Eagle3

NVIDIA · 1.8B · runs from 1.1 GB

381 7

Kimi K2.6 Eagle3 is a 1.8B-parameter open language model from NVIDIA in the Kimi K2 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nemotron H 47B Reasoning 128K

NVIDIA · 46.8B · runs from 102.9 GB

372 21

Nemotron H 47B Reasoning 128K is a 46.8B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Nemotron Terminal 14B

NVIDIA · 14.8B · runs from 6.9 GB

336 8

Nemotron Terminal 14B is a 14.8B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen3 Nemotron 235B A22B GenRM 2603

NVIDIA · 235.1B · runs from 100.4 GB

283 29

Qwen3 Nemotron 235B A22B GenRM 2603 is a 235.1B-parameter open language model from NVIDIA in the Qwen 3 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

NVIDIA Nemotron 3 Ultra 550B A55B GenRM

NVIDIA · 560.5B · runs from 262.1 GB

164 9

NVIDIA Nemotron 3 Ultra 550B A55B GenRM is a 560.5B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nemotron Flash 3B

NVIDIA · 2.7B · runs from 6.0 GB

157 17

Nemotron Flash 3B is a 2.7B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 29,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

OpenCodeReasoning Nemotron 1.1 32B

NVIDIA · 32.8B · runs from 14.8 GB

136 48

OpenCodeReasoning Nemotron 1.1 32B is a 32.8B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCodeReasoning

Riva Translate 4B Instruct

NVIDIA · 4.2B · runs from 2.3 GB

131 18

Riva Translate 4B Instruct is a 4.2B-parameter open language model from NVIDIA. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3.3 Nemotron 70B Reward

NVIDIA · 70.6B · runs from 31.0 GB

112 3

Llama 3.3 Nemotron 70B Reward is a 70.6B-parameter open language model from NVIDIA in the Llama 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat