All LLM Models

Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

GLM 5.1

Z.ai · 753.9B · runs from 211.5 GB

79.5K 1.8K

GLM 5.1 is a 753.9B-parameter open language model from Z.ai in the GLM 5 family. It supports a context window of up to 202,752 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3.1 Tulu 3 8B

Allen AI · 8.0B · runs from 3.3 GB

8.1K 178

Llama-3.1-Tulu-3-8B is Allen Institute for AI's fully open instruction-following model, built on Meta's Llama 3.1 8B base through a public pipeline of supervised fine-tuning, direct preference optimization, and a final reinforcement learning stage with verifiable rewards (RLVR). It is tuned for chat as well as harder tasks like math, GSM8K, and instruction-following (IFEval), with all training data, code, and recipes released openly as part of the Tulu 3 project. Its 8B size lets it run on a single consumer GPU. Context length is 131,072 tokens. It is released under the Llama 3.1 Community License Agreement, which restricts use above 700 million monthly active users and imposes Meta's acceptable-use policy. It was published in November 2024, alongside a 70B and later a 405B sibling trained the same way.

Chat

Moonlight 16B A3B Instruct

Moonshot AI · 16.0B · runs from 5.1 GB

49.7K 199

Moonlight 16B A3B Instruct is a 16.0B-parameter open language model from Moonshot AI in the Moonlight family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

MiniCPM V 4.6 Thinking

OpenBMB · 1.3B · runs from 0.7 GB

109.9K 31

MiniCPM-V 4.6 Thinking is OpenBMB's 1.3-billion-parameter vision-language model, a long chain-of-thought reasoning variant of MiniCPM-V 4.6 built on a SigLIP2-400M vision encoder paired with a small Qwen3.5-0.8B language backbone. It generates an explicit reasoning trace before answering, aimed at multimodal reasoning, math, and OCR-heavy document tasks rather than quick captioning, keeping the same edge-friendly, phone-oriented architecture. Its small size lets it run on a single modest consumer GPU. The model supports a 262,144 token context window. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2026. Its distinguishing trait versus base 4.6 is the thinking mode, which trades some latency for better performance on reasoning-heavy visual tasks while reusing the same mixed 4x/16x visual token compression.

Vision

Gpt2

OpenAI · 137M · runs from 0.1 GB

15.4M 4.2K

GPT-2 is the landmark 2019 language model from OpenAI that helped ignite widespread interest in large-scale text generation. At only 137 million parameters it is tiny by modern standards, but it holds an important place in AI history as the model that was initially deemed too dangerous to release in full. Today GPT-2 runs effortlessly on virtually any hardware, including CPUs, making it ideal for educational purposes, experimentation, and understanding transformer fundamentals. It should not be expected to match the quality of modern instruction-tuned models, but it remains a useful teaching tool and conversation starter.

Chat

Apertus 8B Instruct 2509

swiss-ai · 8.1B · runs from 2.8 GB

477.6K 492

Apertus 8B Instruct is an open-source instruction-tuned model from Swiss AI, a collaborative research initiative. Built on an 8 billion parameter base, it emphasizes transparency, open data, and European AI sovereignty. For local users, it delivers solid general-purpose chat and instruction-following in a standard 8B footprint that runs well on consumer GPUs with 8 to 10 GB of VRAM, making it a practical choice for those who value open, community-driven model development.

Chat

Nemotron Mini 4B Instruct

NVIDIA · 4B · runs from 1.8 GB

89.1K 187

Nemotron Mini 4B Instruct is a 4B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

DeepSeek v2 Lite Chat

DeepSeek · 15.7B · runs from 5.1 GB

143.5K 148

DeepSeek v2 Lite Chat is a 15.7B-parameter open language model from DeepSeek in the DeepSeek V2 family. It supports a context window of up to 163,840 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Ternary Bonsai 4B Unpacked

prism-ml · 4.0B · runs from 2.2 GB

1.5K 7

Ternary Bonsai 4B Unpacked is a 4.0B-parameter open language model from prism-ml. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 4 31B IT Qat Q4 0 Unquantized Assistant

Google · 31B · runs from 13.5 GB

16.0K 14

Gemma 4 31B IT Qat Q4 0 Unquantized Assistant is a 31B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Zephyr 7B Beta

Hugging Face · 7.2B · runs from 3.6 GB

61.1K 1.9K

Zephyr 7B Beta is a 7.2B-parameter open language model from Hugging Face. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

MiMo V2.5

Xiaomi · 310.8B · runs from 85.9 GB

282.1K 427

MiMo V2.5 is a 310.8B-parameter open language model from Xiaomi in the MiMo family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatFunctions

Yi 1.5 34B Chat

01.AI · 34.4B · runs from 12.4 GB

17.1K 278

Yi 1.5 34B Chat is a 34.4-billion parameter instruction-tuned model by 01.AI, the Chinese AI lab founded by Kai-Fu Lee. It is a bilingual model with strong performance in both English and Chinese, making it particularly well suited for users who need high-quality generation in either language. Yi 1.5 represents an improved iteration of the Yi model family with enhanced reasoning and coding ability. The 34B size requires a GPU with at least 24GB of VRAM for quantized inference, placing it within reach of high-end consumer cards like the RTX 4090. Released under the Yi License.

Chat

CodeLlama 34B Instruct HF

Meta · 33.7B · runs from 10.0 GB

21.2K 305

CodeLlama-34b-Instruct-hf is Meta's 34-billion-parameter instruction-tuned Code Llama model, fine-tuned from the base Code Llama checkpoint for safer, more reliable code-assistant use, as opposed to the plain base and Python-specialized variants in the same family (which also ships in 7B, 13B, and 70B sizes). It is an early, first-generation code model built on the original Llama 2 architecture and predates the later Code Llama 70B release. At 34B parameters, it needs a multi-GPU setup or aggressive quantization for local use. Context length is 16,384 tokens. It is released under the Llama 2 Community License, a custom license that restricts commercial use above 700 million monthly active users, requiring a separate license from Meta, and was published in August 2023.

ChatCode

Pantheon Reasoning 27B

Gryphe · 27.8B · runs from 8.4 GB

175 23

Pantheon Reasoning 27B is a 27.8B-parameter open language model from Gryphe. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatRoleplayReasoning

Starcoder2 15B

BigCode · 16.0B · runs from 7.3 GB

4.9K 675

Starcoder2 15B is a 16.0B-parameter open language model from BigCode in the StarCoder family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Huihui Qwen3.6 35B A3B Claude 4.7 Opus Abliterated

huihui-ai · 36.0B · runs from 15.7 GB

5.7K 217

Huihui Qwen3.6 35B A3B Claude 4.7 Opus Abliterated is a 36.0B-parameter open language model from huihui-ai in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Qwopus3.5 9B v3

Jackrong · 9.7B · runs from 19.9 GB

4.0K 92

Qwopus3.5 9B v3 is a 9.7B-parameter open language model from Jackrong. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

VisionReasoning

Olmo 3 32B Think

Allen AI · 32.2B · runs from 9.7 GB

10.7K 177

Olmo 3 32B Think is Allen Institute for AI's reasoning-focused model in the fully open Olmo 3 family, trained to produce long chains of thought before answering so it can work through math and coding problems step by step. It is pretrained on Ai2's Dolma 3 corpus and then carried through a full post-training pipeline of supervised fine-tuning, direct preference optimization, and reinforcement learning with verifiable rewards on the Dolci-Think-RL dataset; Ai2 releases the code, intermediate checkpoints, and training logs for each stage, not just the final weights. It sits alongside a 7B Think and a 7B Instruct sibling in the same release. At 32 billion parameters it needs a high-end consumer GPU or a multi-GPU setup once quantized. Context length is 65,536 tokens. It is released under the Apache 2.0 license, though Ai2 asks that it be used for research and educational purposes in line with its Responsible Use Guidelines. It was published in November 2025 and has since been superseded by Olmo 3.1 32B Think.

Chat

Granite 3.3 2B Instruct

IBM · 2.5B · runs from 1.2 GB

56.3K 87

Granite 3.3 2B Instruct is a 2.5B-parameter open language model from IBM in the Granite family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Ornith 1.5 397B

ornith-ai · 403.4B · runs from 111.4 GB

594.1K 89

Ornith-1.5-397B is the flagship model in Ornith AI's Ornith-1.5 family, a roughly 403-billion-parameter Mixture-of-Experts model built for agentic coding. It extends the earlier Ornith-1.0 line, developed on top of Qwen3.5 and Gemma 4, by widening its self-improvement loop to jointly optimize task generation, agent scaffolding, and solution rollouts through reinforcement learning rather than relying on fixed, human-curated training tasks. It supports a 262,144-token context window and is released under the MIT license. At this scale, 4-bit quantization needs roughly 232GB of memory, putting local use firmly in server or multi-GPU territory; most users will rely on a hosted endpoint.

Chat

Nemotron Labs Audex 30B A3B

NVIDIA · 30B · runs from 14.0 GB

2.2K 163

Nemotron Labs Audex 30B A3B is a 30B-parameter open language model from NVIDIA in the Nemotron family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatReasoning

Ministral 8B Instruct 2410

Mistral AI · 8.0B · runs from 3.3 GB

152.8K 586

Ministral 8B Instruct 2410 is Mistral AI's 8-billion-parameter instruction-tuned model, part of the "Ministraux" family aimed at edge and on-device deployment. It uses a 36-layer dense transformer with interleaved sliding-window attention, trained heavily on multilingual and code data, suiting chat, function calling, and general assistant tasks across ten languages. Its compact size means it runs comfortably on a single mainstream consumer GPU once quantized. The model supports a 32K token context window. It is released under Mistral's own license, listed as "Other," which restricts use to research purposes and requires a separate commercial license from Mistral AI, so review the terms before commercial use. It was published in October 2024 alongside the smaller Ministral 3B.

Chat

EXAONE 3.5 2.4B Instruct

LGAI-EXAONE · 2.4B · runs from 0.9 GB

61.3K 196

EXAONE 3.5 2.4B Instruct is a 2.4B-parameter open language model from LGAI-EXAONE in the EXAONE family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

CodeLlama 70B Instruct HF

Meta · 69.0B · runs from 29.4 GB

410 210

CodeLlama-70B-Instruct is Meta's largest Code Llama model, a 69-billion-parameter, instruction-tuned member of the Code Llama family (7B to 70B) built on Llama 2 for general code synthesis, completion, and chat-style coding assistance. Unlike the smaller Code Llama sizes, the 70B Instruct model uses a different chat prompt template, reflecting extra fine-tuning changes made specifically for this larger variant; separate base and Python-specialist versions are also available at the same size. It is not a Python-only or infilling-focused model, but a general instruction-following code assistant. At 69 billion parameters, it needs a multi-GPU workstation or server to run, even once quantized. Context length is 4,096 tokens. It is released under Meta's Llama 2 Community License, a custom license that is free for most commercial and research use but requires organizations with more than 700 million monthly active users to request separate permission from Meta. It was published in January 2024.

ChatCode

GLM 4.5

Z.ai · 358.3B · runs from 99.2 GB

101.3K 1.4K

GLM-4.5 is Z.ai's flagship model in the GLM-4.5 series, a foundation model built for intelligent-agent applications with 355 billion total parameters and about 32 billion active via its Mixture-of-Experts design. It is a hybrid reasoning model, offering a thinking mode for complex reasoning and tool use alongside a non-thinking mode for fast responses, and it unifies reasoning, coding, and agentic capabilities in one checkpoint alongside the smaller GLM-4.5-Air. It ships with a 128K token context window under the MIT license, permitting commercial and research use. At 355 billion total parameters, even 4-bit quantization needs roughly 200 GB of memory, which puts local use in multi-GPU server territory; most people will access it through a hosted endpoint instead.

Chat

DeepSeek V3.1 Terminus

DeepSeek · 684.5B · runs from 192.1 GB

24.7K 369

DeepSeek-V3.1-Terminus is an updated checkpoint of DeepSeek-V3.1, a Mixture-of-Experts chat and reasoning model with roughly 40.1 billion active parameters out of about 684.5 billion total, fine-tuned from DeepSeek-V3.1-Base. This revision keeps the same capabilities as V3.1 while fixing issues reported by users, chiefly reducing mixed Chinese-English text and stray characters in output and further improving the model's Code Agent and Search Agent performance, including gains on BrowseComp, SWE-bench Verified, and Terminal-bench. Its scale requires a multi-GPU server; it is not something that fits on consumer hardware. Context length is 163,840 tokens. It is released under the MIT license, permitting unrestricted commercial and research use. It was published in September 2025, as a refinement of DeepSeek-V3.1 rather than a new base model.

Chat

Laguna M.1

poolside · 225.8B · runs from 68.3 GB

1.9K 120

Laguna M.1 is a 225.8B-parameter open language model from poolside. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Lucid V1 Nemo

dreamgen · 12.2B · runs from 4.1 GB

217.4K 57

Lucid V1 Nemo is a 12.2B-parameter open language model from dreamgen. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

LocateAnything 3B

NVIDIA · 3.8B · runs from 2 GB

87.4K 3.0K

LocateAnything-3B is NVIDIA's 3.8-billion-parameter vision-language model for visual grounding rather than open-ended chat: referring-expression grounding, dense multi-object detection, GUI element grounding, point-based localization, and document or OCR layout grounding. Its core contribution is Parallel Box Decoding, which predicts a complete bounding box in one parallel step instead of token-by-token autoregressive decoding, giving up to 2.5x higher throughput while preserving geometric consistency. It combines a Qwen2.5-3B-Instruct language backbone with a MoonViT-SO-400M vision encoder and was trained on 12 million images with over 138 million grounding queries; NVIDIA has folded it into the Nemotron 3 Nano Omni model for agentic and computer-use grounding. At under 4 billion parameters, it runs on a single consumer GPU. Context length is 32,768 tokens. It is released under NVIDIA's non-commercial research license, permitting academic and non-profit use only; commercial use requires a separate license from NVIDIA. It was published in May 2026.

Vision