All LLM Models

Browse 1242 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Kumru 2B Base

vngrs-ai · 2.4B · runs from 1.4 GB

1.0K 18

Kumru 2B Base is a 2.4B-parameter open language model from vngrs-ai. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Turkish Gemma 9B v0.1

ytu-ce-cosmos · 9.2B · runs from 4.8 GB

1.0K 37

Turkish Gemma 9B v0.1 is a 9.2B-parameter open language model from ytu-ce-cosmos in the Gemma family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Surjo 50M

SurjoLabs · 54M · runs from 0.4 GB

1.0K 13

Surjo 50M is a 54M-parameter open language model from SurjoLabs. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

PicoLM 80M Instruct

aethertp · 90M · runs from 0.4 GB

1.0K 2

PicoLM 80M Instruct is a 90M-parameter open language model from aethertp. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Supra 50M Base

SupraLabs · 52M · runs from 0.3 GB

1.0K 44

Supra 50M Base is a 52M-parameter open language model from SupraLabs. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

T5gemma L L Ul2 IT

Google · 1.2B · runs from 2.7 GB

1.0K 6

T5gemma L L Ul2 IT is a 1.2B-parameter open language model from Google in the Gemma family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

CyberStrike OffSec 35B

oyildirim · 35.1B · runs from 15.3 GB

1.0K 76

CyberStrike OffSec 35B is a 35.1B-parameter open language model from oyildirim. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

NCP ArchPreview Dolma3 8.9B Stage2 v3

ArchSpace-Collection · 8.9B · runs from 19.7 GB

1.0K 4

NCP ArchPreview Dolma3 8.9B Stage2 v3 is a 8.9B-parameter open language model from ArchSpace-Collection. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Zagreus 0.4B Ita

mii-llm · 438M · runs from 0.6 GB

1.0K 10

Zagreus 0.4B Ita is a 438M-parameter open language model from mii-llm. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Xgen 7B 8k Base

Salesforce · 7B · runs from 3.3 GB

997 317

XGen-7B-8K-Base is Salesforce AI Research's 7-billion-parameter pretrained base language model, introduced in the 2023 paper "Long Sequence Modeling with XGen: A 7B LLM Trained on 8K Input Sequence Length" as one of the earlier open 7B models built specifically for longer input sequences. It is not instruction-tuned; a separate XGen-7B-8K-Inst checkpoint, released for research purposes only, adds supervised instruction fine-tuning on top of the same base, and a sibling XGen-7B-4K-Base uses a shorter 4K training sequence length. It uses OpenAI's Tiktoken tokenizer rather than a custom vocabulary. At 7 billion parameters it runs easily on a single consumer GPU. Context length is 8,192 tokens, the model's namesake feature. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in June 2023, predating the wave of 7B open models that followed later that year such as Mistral 7B.

Chat

Supergemma4 E4b Abliterated

Jiunsong · 7.5B · runs from 3.7 GB

996 72

Supergemma4 E4b Abliterated is a 7.5B-parameter open language model from Jiunsong in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Sweep Next Edit v2 7B

sweepai · 7.6B · runs from 3.6 GB

995 32

Sweep Next Edit v2 7B is a 7.6B-parameter open language model from sweepai. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Functiongemma 270M Ft Mobile Actions

litert-community · 270M · runs from 0.6 GB

994 232

Functiongemma 270M Ft Mobile Actions is a 270M-parameter open language model from litert-community in the Gemma family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

RedPajama INCITE 7B Base

togethercomputer · 7B · runs from 3.3 GB

988 92

RedPajama-INCITE-7B-Base is Together Computer's open base pretrained language model, not instruction-tuned, with roughly 6.9 billion parameters trained on the RedPajama-Data-1T dataset, an open reproduction of the corpus used to train Meta's original LLaMA. It was developed with a consortium including Ontocord.ai, ETH DS3Lab, Stanford CRFM and Hazy Research, and LAION, using compute awarded through the 2023 INCITE program. Instruction-tuned and chat variants, RedPajama-INCITE-7B-Instruct and RedPajama-INCITE-7B-Chat, were released alongside it. At under 7 billion parameters it runs on a single consumer GPU. Context length is 2,048 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in May 2023, as one of the first fully open, commercially usable base models trained on openly licensed data.

Chat

Nidum Gemma 2B Uncensored

VibeStudio · 2.5B · runs from 1.4 GB

980 5

Nidum Gemma 2B Uncensored is a 2.5B-parameter open language model from VibeStudio in the Gemma 2 family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GPT X2 125M

AxiomicLabs · 144M · runs from 0.6 GB

979 23

GPT X2 125M is a 144M-parameter open language model from AxiomicLabs. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GPT OSS 20B Heretic

p-e-w · 20.9B · runs from 9.3 GB

978 132

GPT OSS 20B Heretic is a 20.9B-parameter open language model from p-e-w in the GPT-OSS family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

VibeThinker 1.5B

WeiboAI · 1.8B · runs from 1.1 GB

968 524

VibeThinker 1.5B is a 1.8B-parameter open language model from WeiboAI. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatMathCode

Qwen3.6 27B Uncensored HauhauCS Aggressive Safetensor Benchmark

DreamFast · 27.8B · runs from 12.6 GB

937 5

Qwen3.6 27B Uncensored HauhauCS Aggressive Safetensor Benchmark is a 27.8B-parameter open language model from DreamFast in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Chadrock 35B Ace Saber Rocmfp4 Mtp

jcbtc · 35B · runs from 16.4 GB

896 7

Chadrock 35B Ace Saber Rocmfp4 Mtp is a 35B-parameter open language model from jcbtc. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

CyberSecQwen 4B

lablab-ai-amd-developer-hackathon · 4.0B · runs from 2.2 GB

884 14

CyberSecQwen 4B is a 4.0B-parameter open language model from lablab-ai-amd-developer-hackathon in the Qwen family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama3 OpenBioLLM 70B

aaditya · 70B · runs from 30.7 GB

862 517

Llama3 OpenBioLLM 70B is a 70B-parameter open language model from aaditya in the Llama 3 family. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Rnj 1.5 Instruct

EssentialAI · 8.3B · runs from 17.2 GB

844 20

Rnj 1.5 Instruct is a 8.3B-parameter open language model from EssentialAI. It supports a context window of up to 163,840 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

T Pro IT 2.0

t-tech · 32.8B · runs from 14.6 GB

834 126

T Pro IT 2.0 is a 32.8B-parameter open language model from t-tech. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Recursive Language Model 198M

Girinath11 · 198M · runs from 0.4 GB

787 10

Recursive Language Model 198M is a 198M-parameter open language model from Girinath11. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

ErniePEUnleashed

Kezmark · 3.4B · runs from 1.9 GB

781 4

ErniePEUnleashed is a 3.4B-parameter open language model from Kezmark in the ERNIE family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GPT S 5M

AxiomicLabs · 5M · runs from 0.3 GB

769 12

GPT S 5M is a 5M-parameter open language model from AxiomicLabs. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Falcon H1 7B Base

TII UAE · 7.6B · runs from 3.7 GB

746 11

Falcon H1 7B Base is a 7.6B-parameter open language model from TII UAE in the Falcon family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

OmniCoder 9B

Tesslate · 9.4B · runs from 19.4 GB

744 683

OmniCoder 9B is a 9.4B-parameter open language model from Tesslate. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCodeFunctions

XortronCriminalComputingConfig

darkc0de · 23.6B · runs from 10.7 GB

721 137

XortronCriminalComputingConfig is a 23.6B-parameter open language model from darkc0de. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat