All LLM Models

Browse 41 LLM models with VRAM requirements, quantization options, and hardware compatibility.

Understanding LLM VRAM Requirements

How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.

Model List

Llama 3.1 8B Instruct

Meta · 8.0B · runs from 3.6 GB

6.2M 7.9K

Meta Llama 3.1 8B Instruct is an 8-billion parameter instruction-tuned language model from Meta. Part of the Llama 3.1 release, it supports a 128K token context window and is fine-tuned for conversational use, tool calling, and general assistant tasks. Its compact size makes it well-suited for local deployment on modern consumer GPUs with 8GB or more of VRAM. Llama 3.1 8B Instruct delivers strong performance for its parameter class across benchmarks in reasoning, coding, and multilingual understanding. It is released under the Llama 3.1 Community License and is widely supported by inference frameworks such as llama.cpp, vLLM, and Ollama.

Chat

Muse Glimmer 30B

Meta · 29.8B · runs from 8.7 GB

312.4K 1.9K

Muse Glimmer 30B is Meta's 30-billion-parameter dense vision-language model, purpose-built for local, always-on AI agents rather than general chat. It pairs a text decoder with a dedicated image encoder for reasoning over screenshots, charts, and documents, and includes native tool-calling with a separate reasoning channel so it can plan multi-step actions and recover from failures. At this size, local inference calls for quantization and a capable GPU; it fits a single high-end 24-32GB-class card, matching Meta's goal of running it entirely on consumer machines. The model supports a 131K token context window, enough for extended agent sessions. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in August 2026, it is Meta's first open-weight model since Llama 4, and ships without the Llama licenses' monthly-active-user cap.

Vision

Llama 3.2 3B Instruct

Meta · 3.2B · runs from 1.0 GB

1.9M 2.7K

Meta Llama 3.2 3B Instruct is a 3-billion parameter instruction-tuned model from Meta's Llama 3.2 release, designed for efficient local inference on resource-constrained hardware. It supports a 128K token context window and is optimized for conversational AI, summarization, and general assistant tasks. Despite its small footprint, Llama 3.2 3B Instruct delivers competitive performance for its size class and can run on GPUs with as little as 4GB of VRAM when quantized. It is released under the Llama 3.2 Community License and is a practical choice for edge deployment and lightweight local inference.

Chat

Llama 3.2 1B Instruct

Meta · 1.2B · runs from 0.4 GB

7.4M 1.7K

Meta Llama 3.2 1B Instruct is a 1-billion parameter instruction-tuned model from Meta, the smallest in the Llama 3.2 family. It is designed for ultra-lightweight deployment scenarios where minimal hardware resources are available, supporting a 128K token context window despite its compact size. This model is suitable for basic conversational tasks, text summarization, and simple instruction following. It can run on virtually any modern GPU and even on CPU-only setups with acceptable performance. Released under the Llama 3.2 Community License.

Chat

Llama 3.3 70B Instruct

Meta · 70.6B · runs from 21.3 GB

917.9K 3.1K

Meta Llama 3.3 70B Instruct is a 70-billion parameter large language model from Meta, released as part of the Llama 3.3 generation. It is an instruction-tuned model optimized for dialogue and chat use cases, offering strong performance across reasoning, coding, and multilingual tasks. Llama 3.3 70B delivers quality competitive with much larger models while remaining feasible to run on high-end consumer or workstation GPUs with sufficient VRAM. The model uses a grouped-query attention architecture with a 128K token context window and was trained on a massive multilingual corpus. It is released under the Llama 3.3 Community License, making it one of the most capable openly available models for local inference.

Chat

Meta Llama 3 8B Instruct

Meta · 8.0B · runs from 2.6 GB

1.2M 5.1K

Meta Llama 3 8B Instruct is the instruction-tuned version of Meta's Llama 3 8B base model, with 8 billion parameters. It is fine-tuned for dialogue and chat use cases using supervised fine-tuning and RLHF, making it ready for conversational applications out of the box. The model supports an 8K token context window and performs well across coding, reasoning, and general knowledge tasks. Its efficient size makes it one of the most popular models for local inference on consumer hardware. Released under the Meta Llama 3 Community License.

Chat

Llama 3.2 1B

Meta · 1.2B · runs from 0.6 GB

840.0K 2.6K

Meta Llama 3.2 1B is a 1.2-billion parameter base (pretrained) model from Meta's Llama 3.2 release. It is the smallest model in the Llama 3.2 family and is designed for research, fine-tuning, and embedding into resource-constrained environments. It supports a 128K token context window. As a base model, it is not optimized for conversational use without further fine-tuning. Its minimal resource requirements make it suitable for experimentation, edge deployment, and as a starting point for domain-specific fine-tuning. Released under the Llama 3.2 Community License.

Chat

Llama 3.2 3B

Meta · 3.2B · runs from 1.5 GB

366.5K 965

Meta Llama 3.2 3B is a 3.2-billion parameter base (pretrained) model from Meta's Llama 3.2 family. It supports a 128K token context window and is intended for fine-tuning, research, and custom applications rather than direct conversational use. The model provides a good balance between capability and efficiency at the small model scale. It is popular as a foundation for community fine-tunes and domain-specific adaptations. Released under the Llama 3.2 Community License.

Chat

Llama 2 7B Chat HF

Meta · 6.7B · runs from 3.1 GB

499.7K 4.9K

Meta Llama 2 7B Chat is a 7-billion parameter instruction-tuned model from Meta's Llama 2 family, optimized for dialogue use cases. It was fine-tuned using supervised fine-tuning and RLHF on top of the Llama 2 7B base model, with a 4K token context window. This model is suitable for basic conversational AI tasks and runs efficiently on consumer GPUs. While newer Llama generations offer improved performance, Llama 2 7B Chat remains a well-understood and widely-supported option for local inference. Released under the Llama 2 Community License.

Chat

Llama 4 Scout 17B 16E Instruct

Meta · 108.6B · runs from 32.9 GB

162.9K 1.4K

Llama 4 Scout is the smaller of Meta's first Llama 4 models, a natively multimodal Mixture-of-Experts model with 16 experts, about 109 billion total parameters and 17 billion active per token. This instruction-tuned checkpoint accepts text and images and generates text, and is aimed at assistant-style chat, visual understanding and long-document work. Meta advertises an unusually long context window of up to 10 million tokens, and the model is released under the Llama 4 Community License. Because every expert has to stay in memory, 4-bit quantization still needs roughly 60GB, which means a high-memory workstation, a large unified-memory Mac or a multi-GPU setup rather than a single consumer card.

Vision

Llama 3.2 11B Vision Instruct

Meta · 10.7B · runs from 5.0 GB

65.1K 1.7K

Llama 3.2 11B Vision Instruct is a smaller vision-language model in Meta's Llama 3 family, with roughly 11 billion parameters, built to process images together with text and follow chat-style instructions. It brings multimodal capability, such as visual question answering and image description, to a size that is far more approachable than the 90-billion-parameter variant from the same release. Released in September 2024 under Meta's Llama 3.2 community license, it does not carry a fully permissive open license but does allow broad use under Meta's terms. At around 11 billion parameters, it fits comfortably on a single mid-range to high-end consumer GPU once quantized to 4-bit, putting it within reach of enthusiast local setups.

Vision

Llama 4 Maverick 17B 128E Instruct

Meta · 401.6B · runs from 121.5 GB

10.7K 513

Llama 4 Maverick is Meta's instruction-tuned, natively multimodal mixture-of-experts model, with roughly 401.6 billion total parameters and about 17 billion active per token routed across 128 experts. It accepts multilingual text and image input and produces multilingual text and code output, using early fusion to integrate vision and language from pretraining rather than bolting on a separate vision encoder; Meta trained it on roughly 22 trillion tokens of multimodal data with an August 2024 knowledge cutoff. It is built for assistant-style chat and visual reasoning tasks. Despite its relatively small active-parameter count, its total size means it needs a multi-GPU server-class setup to run even once quantized. Context length is up to 1,000,000 tokens. It is released under the Llama 4 Community License, a custom license that is free for most commercial and research use but requires companies with more than 700 million monthly active users to request separate permission from Meta. It was published in April 2025.

Vision

CodeLlama 34B Instruct HF

Meta · 33.7B · runs from 10.0 GB

21.2K 305

CodeLlama-34b-Instruct-hf is Meta's 34-billion-parameter instruction-tuned Code Llama model, fine-tuned from the base Code Llama checkpoint for safer, more reliable code-assistant use, as opposed to the plain base and Python-specialized variants in the same family (which also ships in 7B, 13B, and 70B sizes). It is an early, first-generation code model built on the original Llama 2 architecture and predates the later Code Llama 70B release. At 34B parameters, it needs a multi-GPU setup or aggressive quantization for local use. Context length is 16,384 tokens. It is released under the Llama 2 Community License, a custom license that restricts commercial use above 700 million monthly active users, requiring a separate license from Meta, and was published in August 2023.

ChatCode

CodeLlama 70B Instruct HF

Meta · 69.0B · runs from 29.4 GB

410 210

CodeLlama-70B-Instruct is Meta's largest Code Llama model, a 69-billion-parameter, instruction-tuned member of the Code Llama family (7B to 70B) built on Llama 2 for general code synthesis, completion, and chat-style coding assistance. Unlike the smaller Code Llama sizes, the 70B Instruct model uses a different chat prompt template, reflecting extra fine-tuning changes made specifically for this larger variant; separate base and Python-specialist versions are also available at the same size. It is not a Python-only or infilling-focused model, but a general instruction-following code assistant. At 69 billion parameters, it needs a multi-GPU workstation or server to run, even once quantized. Context length is 4,096 tokens. It is released under Meta's Llama 2 Community License, a custom license that is free for most commercial and research use but requires organizations with more than 700 million monthly active users to request separate permission from Meta. It was published in January 2024.

ChatCode

CodeLlama 7B Instruct HF

Meta · 6.7B · runs from 4.2 GB

41.0K 258

CodeLlama 7B Instruct HF is a 6.7B-parameter open language model from Meta in the Code Llama family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Meta Llama 3 70B Instruct

Meta · 70.6B · runs from 23.3 GB

119.2K 1.5K

Meta Llama 3 70B Instruct is a 70.6-billion parameter instruction-tuned model from Meta's Llama 3 release. It is fine-tuned for dialogue, coding assistance, and complex reasoning tasks using supervised fine-tuning and RLHF. At the time of release, it was among the most capable openly available models. The model supports an 8K token context window and requires substantial VRAM for local inference, typically needing multi-GPU setups or high-VRAM professional GPUs. It has been widely adopted for local deployment in quantized formats. Released under the Meta Llama 3 Community License.

Chat

CodeLlama 7B HF

Meta · 6.7B · runs from 4.2 GB

325.5K 379

CodeLlama 7B HF is a 6.7B-parameter open language model from Meta in the Code Llama family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

ChatCode

Meta Llama 3 8B

Meta · 8.0B · runs from 3.8 GB

261.8K 6.7K

Meta Llama 3 8B is an 8-billion parameter base (pretrained) language model from Meta's Llama 3 release. As a base model, it is not fine-tuned for chat or instructions and is intended for further fine-tuning, research, or as a foundation for custom applications. It uses grouped-query attention and was trained on over 15 trillion tokens. Llama 3 8B supports an 8K token context window and delivers strong benchmark performance across language understanding, reasoning, and coding tasks for its size. It is released under the Meta Llama 3 Community License and runs efficiently on consumer GPUs with 8GB or more of VRAM.

Chat

Llama 3.1 8B

Meta · 8.0B · runs from 3.6 GB

555.3K 2.5K

Meta Llama 3.1 8B is an 8-billion parameter base (pretrained) model from the Llama 3.1 family. It is not instruction-tuned and is intended for fine-tuning, research, and custom downstream applications. Compared to Llama 3 8B, it extends the context window to 128K tokens and benefits from improved training data and methodology. The model uses grouped-query attention and was trained on a multilingual corpus. It is released under the Llama 3.1 Community License and is widely used as a foundation for community fine-tunes and specialized models.

Chat

Llama 3.1 70B Instruct

Meta · 70.6B · runs from 33.0 GB

136.0K 987

Meta Llama 3.1 70B Instruct is a 70.6-billion parameter instruction-tuned model from Meta's Llama 3.1 family. It features a 128K token context window and is optimized for chat, tool use, and complex reasoning tasks. The 70B size offers a strong balance between capability and hardware requirements, running well on multi-GPU setups or high-VRAM workstation cards. This model was trained on over 15 trillion tokens and fine-tuned with reinforcement learning from human feedback (RLHF). It excels at coding assistance, mathematical reasoning, and multilingual dialogue. Released under the Llama 3.1 Community License.

Chat

Opt 125M

Meta · 125M · runs from 0.3 GB

6.8M 324

Meta OPT 125M is a 125-million parameter language model from Meta's Open Pre-trained Transformer (OPT) project. Released in 2022, it was part of Meta's effort to provide the research community with openly available large language models that replicate the performance of GPT-3 class models at various scales. As one of the smallest models in the OPT family, the 125M variant is primarily useful for research, experimentation, and educational purposes. It can run on virtually any hardware, including CPU-only setups. While significantly less capable than modern models, it remains a useful reference point in LLM research.

Chat

Llama 2 7B HF

Meta · 6.7B · runs from 3.1 GB

854.1K 2.4K

Meta Llama 2 7B is a 6.7-billion parameter base (pretrained) language model from Meta's Llama 2 generation, provided in Hugging Face Transformers format. It was trained on 2 trillion tokens with a 4K token context window and represented a significant step in openly available large language models when released. As a base model, it is designed for further fine-tuning and research rather than direct chat use. While superseded by Llama 3 and later releases in terms of benchmark performance, Llama 2 7B remains widely used in the research community and as a baseline for comparison. Released under the Llama 2 Community License.

Chat

Opt 350M

Meta · 350M · runs from 0.8 GB

293.6K 150

Meta OPT 350M is a 350-million parameter language model from Meta's Open Pre-trained Transformer (OPT) project, released in 2022 as part of a suite of models ranging from 125M to 175B parameters. It was designed to provide researchers with open access to models comparable to GPT-3 at various scales. The 350M variant runs on minimal hardware and is suitable for research, prototyping, and educational use. While it has been surpassed by modern architectures in terms of capability, it remains a lightweight option for basic text generation experiments and as a benchmark baseline.

Chat

Llama 2 13B Chat HF

Meta · 13.0B · runs from 6.1 GB

157.3K 1.1K

Meta Llama 2 13B Chat is a 13-billion parameter instruction-tuned model from Meta's Llama 2 family, fine-tuned for dialogue and chat applications. It offers improved reasoning and generation quality over the 7B variant while maintaining manageable hardware requirements with a 4K token context window. The model was fine-tuned using supervised fine-tuning and RLHF. It can run on consumer GPUs with 16GB or more of VRAM at reduced precision. Released under the Llama 2 Community License.

Chat

Llama 3.2 90B Vision Instruct

Meta · 88.6B · runs from 194.9 GB

150.4K 362

Llama 3.2 90B Vision Instruct is a vision-language model in Meta's Llama 3 family, with around 89 billion parameters, that accepts image input alongside text and is tuned to follow chat-style instructions. It extends the Llama 3 lineup with multimodal understanding, aimed at tasks that combine visual and textual reasoning, such as describing images or answering questions about their content. It was released in September 2024 under Meta's Llama 3.2 community license, which permits use under Meta's specific terms rather than a fully open license. At around 89 billion parameters, local deployment at 4-bit precision needs a multi-GPU workstation or a machine with generous unified memory; most people will run it through a hosted API instead.

Vision

Xglm 564M

Meta · 564M · runs from 0.3 GB

111.7K 54

Xglm 564M is a 564M-parameter open language model from Meta in the GLM family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Llama 3.1 405B

Meta · 405.9B · runs from 189.7 GB

103.5K 989

Meta Llama 3.1 405B is the largest model in the Llama family with 405 billion parameters. It represents Meta's most capable open-weight model, delivering performance competitive with leading proprietary models across reasoning, coding, math, and multilingual tasks. It features a 128K token context window. Due to its massive size, running Llama 3.1 405B locally requires significant hardware, typically multiple high-end professional GPUs with a combined VRAM of 200GB or more at reduced precision. It is primarily used in quantized formats for local inference or via multi-node setups. Released under the Llama 3.1 Community License.

Chat

Opt 1.3B

Meta · 1.3B · runs from 2.9 GB

92.5K 186

Opt 1.3B is a 1.3B-parameter open language model from Meta. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Meta Llama 3 70B

Meta · 70.6B · runs from 33.0 GB

75.5K 881

Meta Llama 3 70B is a 70.6B-parameter open language model from Meta in the Llama 3 family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat

Opt 6.7B

Meta · 6.7B · runs from 14.7 GB

58.7K 121

Opt 6.7B is a 6.7B-parameter open language model from Meta. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Chat