All LLM Models

Browse 1475 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Qwen1.5 110B

Alibaba · 111.2B · runs from 46.9 GB

1.2K 104

Qwen1.5-110B is Alibaba's largest dense base model in the Qwen1.5 series, a pretrained language model that is not instruction-tuned, released as a beta preview ahead of Qwen2. Like the rest of the Qwen1.5 line it uses a SwiGLU Transformer with an improved multilingual tokenizer, and at 110B it is one of only two sizes in the series, alongside 32B, to include grouped-query attention. A matching aligned chat model was released alongside it. At 111 billion parameters it requires a multi-GPU workstation or server to run, even quantized. Context length is 32,768 tokens. It is released under Alibaba's custom Tongyi Qianwen license, which permits commercial use below 100 million monthly active users. It was published in April 2024; Qwen1.5 has since been superseded by the Qwen2, Qwen2.5, and Qwen3 series.

Chat

Supra2 100M Instruct

SupraLabs · 101M · runs from 0.4 GB

79.6K 53

Supra2 100M Instruct is a 101M-parameter open language model from SupraLabs. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gpt2 Mini

erwanf · 39M · runs from 0.0 GB

138.6K 7

Gpt2 Mini is a 39M-parameter open language model from erwanf. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Falcon 7B

TII UAE · 7.2B · runs from 3.4 GB

387.0K 1.1K

Falcon 7B was one of the first truly competitive open-source large language models, released in mid-2023 by the Technology Innovation Institute in Abu Dhabi. Trained on the massive RefinedWeb dataset, it demonstrated that carefully curated web data could rival models trained on more traditionally assembled corpora. At 7 billion parameters, Falcon 7B helped establish the 7B class as the sweet spot for local inference, offering genuine language understanding on consumer GPUs with as little as 6 GB of VRAM.

Chat

Bloomz 560M

BigScience · 559M · runs from 0.3 GB

845.9K 138

Bloomz 560M is a 559M-parameter open language model from BigScience. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Molmo2 4B

Allen AI · 4.9B · runs from 2.5 GB

105.7K 54

Molmo2-4B is Allen Institute for AI's (Ai2) vision-language model for image, video, and multi-image understanding and grounding, built on a Qwen3-4B-Instruct backbone with a SigLIP 2 vision encoder. It is trained on Ai2's own curated Molmo2 datasets rather than third-party captioning data of unclear provenance, and beyond ordinary visual question answering it supports pointing at and tracking objects across video frames. The card reports state-of-the-art results among open weight-and-data models on short-video understanding, counting, and captioning, with competitive results on long videos. At under 5 billion parameters, it fits on a single consumer GPU. Context length is 36,864 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in December 2025.

Vision

Pythia 14M

EleutherAI · 14M · runs from 0.0 GB

438.4K 6

Pythia 14M is a 14M-parameter open language model from EleutherAI. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

NousCoder 14B

Nous Research · 14.8B · runs from 6.9 GB

522 227

NousCoder 14B is a 14.8B-parameter open language model from Nous Research. It supports a context window of up to 81,920 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Qwen 1 8B

Alibaba · 1.8B · runs from 0.9 GB

2.3K 74

Qwen-1.8B is Alibaba's first-generation 1.8-billion-parameter base language model, pretrained from scratch on over 2.2 trillion tokens of Chinese, English, multilingual, code, and math data, with the same roughly 150,000-token vocabulary used across the Qwen family. It is a raw pretrained model rather than a chat assistant; Alibaba's aligned Qwen-1.8B-Chat is built on top of it. Its main selling point is low-cost deployment: the card reports int4/int8 quantized versions needing under 2GB of memory for inference and as little as 6GB for fine-tuning, so it runs comfortably on almost any consumer GPU or even a laptop. Context length is 8,192 tokens. It is released under a custom Tongyi Qianwen Research License, free for academic research, with commercial use requiring direct contact with Alibaba. It was published in November 2023.

Chat

Falcon 40B

TII UAE · 41.8B · runs from 19.6 GB

6.5K 2.4K

Falcon-40B is TII's 40-billion-parameter causal decoder-only base language model, pretrained from scratch on 1,000 billion tokens of the RefinedWeb dataset enhanced with curated corpora. At release TII described it as the best fully open model available, outperforming LLaMA, StableLM, and MPT on the Open LLM Leaderboard, with an inference-optimized architecture using FlashAttention and multi-query attention. It is a raw, pretrained model meant to be fine-tuned for chat or other downstream tasks rather than used directly; TII's own Falcon-40B-Instruct is the ready-to-use chat version built on top of it. Trained mainly on English, German, Spanish, and French, it needs a multi-GPU workstation in half precision, while a 4-bit quantization brings it within reach of a single high-memory GPU. License is Apache 2.0, permitting unrestricted commercial and research use. It was published in May 2023.

Chat

Llama 3.1 70B Instruct

Meta · 70.6B · runs from 33.0 GB

136.0K 987

Meta Llama 3.1 70B Instruct is a 70.6-billion parameter instruction-tuned model from Meta's Llama 3.1 family. It features a 128K token context window and is optimized for chat, tool use, and complex reasoning tasks. The 70B size offers a strong balance between capability and hardware requirements, running well on multi-GPU setups or high-VRAM workstation cards. This model was trained on over 15 trillion tokens and fine-tuned with reinforcement learning from human feedback (RLHF). It excels at coding assistance, mathematical reasoning, and multilingual dialogue. Released under the Llama 3.1 Community License.

Chat

OLMo 1B HF

Allen AI · 1.2B · runs from 1.1 GB

66.7K 29

OLMo 1B HF is a 1.2B-parameter open language model from Allen AI in the OLMo family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Qwen 14B

Alibaba · 14.2B · runs from 6.6 GB

5.1K 214

Qwen-14B is Alibaba's first-generation 14-billion-parameter base language model, pretrained from scratch on over 3 trillion tokens of Chinese, English, multilingual, code, and math data, with a roughly 150,000-token vocabulary for broader multilingual coverage. It is a raw pretrained Transformer, not tuned for conversation; Alibaba's aligned Qwen-14B-Chat is the assistant built on top of it. The card reports it beating other open models of similar size, and in some benchmarks larger models too, on Chinese and English evaluation suites. At 14B parameters it needs a capable consumer GPU at full precision, considerably less once quantized. Context length is 8,192 tokens, extendable further with the NTK-aware interpolation and window-attention techniques described in the card. It is released under a custom Tongyi Qianwen License Agreement that is free for research, with commercial use requiring a separate application to Alibaba. It was published in September 2023.

Chat

T5 Paraphrase Paws

Vamsi · 223M · runs from 0.1 GB

83.8K 44

T5 Paraphrase Paws is a 223M-parameter open language model from Vamsi. It supports a context window of up to 512 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Phi 3 Mini 128k Instruct

Microsoft · 3.8B · runs from 2.7 GB

178.4K 1.7K

Phi 3 Mini 128k Instruct is a 3.8B-parameter open language model from Microsoft in the Phi 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Sarvam 30B Uncensored

aoxo · 32.2B · runs from 14 GB

413 6

Sarvam 30B Uncensored is a 32.2B-parameter open language model from aoxo. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

GLM 4.6 Derestricted v3

ArliAI · 356.8B · runs from 152.3 GB

1.7K 134

GLM 4.6 Derestricted v3 is a 356.8B-parameter open language model from ArliAI in the GLM 4 family. It supports a context window of up to 202,752 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Gemma 4 12B IT Assistant

Google · 12B · runs from 5.4 GB

29.2K 82

Gemma 4 12B IT Assistant is a 12B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Phi Tiny MoE Instruct

Microsoft · 3.8B · runs from 2.2 GB

92.3K 43

Phi Tiny MoE Instruct is a 3.8B-parameter open language model from Microsoft in the Phi family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Nemotron Labs Diffusion 3B

NVIDIA · 3.8B · runs from 2.1 GB

67.8K 42

Nemotron Labs Diffusion 3B is a 3.8B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

A.X 4.0 Light

skt · 7.3B · runs from 2.4 GB

44.9K 113

A.X 4.0 Light is a 7.3B-parameter open language model from skt. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Tiny Mixtral

TitanML · 247M · runs from 0.4 GB

167.9K 2

Tiny Mixtral is a 247M-parameter open language model from TitanML in the Mixtral family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3.1 Nemotron Nano 8B V1

NVIDIA · 8B · runs from 2.8 GB

308.6K 219

Llama 3.1 Nemotron Nano 8B is an 8-billion parameter chat model by NVIDIA, a compact entry in the Nemotron family derived from Meta's Llama 3.1 architecture. It applies NVIDIA's alignment and fine-tuning techniques to deliver improved response quality over the base Llama 3.1 8B Instruct model at the same parameter count. The model runs on consumer GPUs with 8GB or more of VRAM and supports a 128K token context window. Its small footprint and NVIDIA-tuned quality make it a practical option for local inference on mainstream hardware.

Chat

C4ai Command A 03 2025

Cohere · 111.1B · runs from 51.9 GB

2.5K 395

Command A (command-a-03-2025) is Cohere and Cohere Labs' 111.1-billion-parameter dense (non-mixture-of-experts) language model tuned for enterprise agentic work: tool use, retrieval-augmented generation, and multilingual tasks across 23 languages. Its architecture repeats a pattern of three sliding-window attention layers (4,096-token window) followed by one global-attention layer without positional embeddings, a design later reused and scaled up in the larger Command A+. At release, Cohere reported roughly 150% higher throughput than its predecessor Command R+ 08-2024 and said the model could run on as few as two GPUs. At 111 billion parameters, it needs a multi-GPU server to run. Context length is 256,000 tokens. It is released under the CC BY-NC 4.0 license, restricting use to non-commercial purposes only. It was published in March 2025.

Chat

Baichuan 13B Base

baichuan-inc · 13B · runs from 6.1 GB

585 187

Baichuan-13B-Base is Baichuan Intelligence's 13-billion-parameter pretrained base model, the follow-up to Baichuan-7B, trained on 1.4 trillion tokens of bilingual Chinese/English data, about 40% more than LLaMA-13B at the time. It replaces the usual rotary position embeddings with ALiBi linear position biasing, which the authors report gives roughly 31.6% faster token generation than a comparable LLaMA-13B. Official int8 and int4 quantized versions are provided so the model deploys on consumer GPUs such as the RTX 3090 with little accuracy loss. A separately released Baichuan-13B-Chat provides the aligned, conversational counterpart to this base checkpoint. Context length is 4,096 tokens. It is released under a custom Community License for the Baichuan-13B Model: free for academic research, with commercial use requiring written authorization from Baichuan via email request. It was published in July 2023, an early bilingual open model that predates the Baichuan2 series.

Chat

NVIDIA Nemotron 3 Ultra 550B A55B Base BF16

NVIDIA · 560.5B · runs from 262.1 GB

2.0K 25

NVIDIA Nemotron 3 Ultra 550B A55B Base BF16 is a 560.5B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Baichuan2 13B Base

baichuan-inc · 13B · runs from 6.1 GB

716 82

Baichuan2-13B-Base is Baichuan Intelligence's 13-billion-parameter base language model — pretrained, not instruction-tuned — from the second-generation Baichuan series, trained on 2.6 trillion tokens of high-quality Chinese and English text. At release it topped same-size open models on Chinese and English benchmarks spanning general knowledge, law, medicine, math, code, and multilingual translation; a separate Baichuan2-13B-Chat model built on top of it adds instruction and safety tuning. As a 13-billion-parameter dense model, it fits on a single consumer GPU once quantized. It is released under a custom Baichuan2 community license: free for academic research and for commercial use once a developer applies for and receives a free commercial license from Baichuan by email. It was published in September 2023.

Chat

GLM 4.7 Flash Coder

whitecircle · 29.9B · runs from 13.8 GB

27.8K 14

GLM 4.7 Flash Coder is a 29.9B-parameter open language model from whitecircle in the GLM 4 family. It supports a context window of up to 202,752 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Qwen3.5 4B

Alibaba · 4.7B · runs from 2.5 GB

7.0M 956

Qwen3.5-4B is Alibaba's dense 4-billion-parameter vision-language model from the Qwen3.5 generation, sharing the same hybrid Gated DeltaNet and gated-attention architecture as its larger siblings. This is the post-trained, instruction-tuned release, able to process images together with text and handle general chat, coding assistance, and agentic tool use. It has a 262,144-token native context window, extensible up to roughly 1,010,000 tokens, and is released under the Apache 2.0 license. At around 4.7 billion parameters, 4-bit quantization needs under 3GB of memory, so it runs easily on almost any modern GPU or laptop.

Vision

Opt 125M

Meta · 125M · runs from 0.3 GB

6.8M 324

Meta OPT 125M is a 125-million parameter language model from Meta's Open Pre-trained Transformer (OPT) project. Released in 2022, it was part of Meta's effort to provide the research community with openly available large language models that replicate the performance of GPT-3 class models at various scales. As one of the smallest models in the OPT family, the 125M variant is primarily useful for research, experimentation, and educational purposes. It can run on virtually any hardware, including CPU-only setups. While significantly less capable than modern models, it remains a useful reference point in LLM research.

Chat