All LLM Models
Browse 1242 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
SmolLM 135M Instruct
Hugging Face · 135M · runs from 0.4 GB
SmolLM 135M Instruct is a 135M-parameter open language model from Hugging Face in the SmolLM family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
OneRec 1.7B
OpenOneRec · 2.1B · runs from 1.1 GB
OneRec 1.7B is a 2.1B-parameter open language model from OpenOneRec. It supports a context window of up to 40,960 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
LFM2 1.2B
Liquid AI · 1.2B · runs from 0.9 GB
LFM2 1.2B is a 1.2B-parameter open language model from Liquid AI in the LFM2 family. It supports a context window of up to 128,000 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3Guard Gen 0.6B
Alibaba · 752M · runs from 0.7 GB
Qwen3Guard Gen 0.6B is a 752M-parameter open language model from Alibaba in the Qwen 3 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
JiRackUltra 32B
CMSManhattan · 32.8B · runs from 9.8 GB
JiRackUltra 32B is a 32.8B-parameter open language model from CMSManhattan. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Kimi Linear 48B A3B Base
Moonshot AI · 49.1B · runs from 14.3 GB
Kimi Linear 48B A3B Base is a 49.1B-parameter open language model from Moonshot AI in the Kimi family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen1.5 0.5B Chat
Alibaba · 620M · runs from 0.8 GB
Qwen1.5 0.5B Chat is an early-generation small language model from Alibaba's Qwen series with just 620 million parameters. As one of the smallest models in the Qwen family, it was designed to demonstrate that useful conversational ability is possible even at sub-billion parameter scales. This model runs easily on virtually any hardware including CPUs, older GPUs, and even mobile devices. While its capabilities are limited compared to larger Qwen models, it remains a useful option for embedded applications, rapid prototyping, or situations where minimal resource consumption is the top priority.
Qwen3Guard Gen 4B
Alibaba · 4.4B · runs from 2.4 GB
Qwen3Guard Gen 4B is a 4.4B-parameter open language model from Alibaba in the Qwen 3 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 1.1 2B IT
Google · 2.5B · runs from 1.1 GB
Gemma 1.1 2B IT is a 2.5B-parameter open language model from Google in the Gemma family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 3 3 Nemotron Super 49B V1 5
NVIDIA · 49.9B · runs from 15.1 GB
Llama 3.3 Nemotron Super 49B is a 49.9-billion parameter chat model by NVIDIA, built on a modified Llama 3.3 architecture. It occupies a unique size point between the common 70B and 8B tiers, offering strong reasoning and conversational ability while requiring less VRAM than full 70B models. NVIDIA's Nemotron Super training pipeline applies extensive alignment tuning to optimize helpfulness and factual accuracy. The model typically needs 32GB or more of VRAM for local inference at reduced precision, placing it within reach of high-end consumer GPUs like the RTX 4090 or professional workstation cards.
C4ai Command R 08 2024
Cohere · 32.3B · runs from 14.7 GB
C4AI Command R 08-2024 is Cohere Labs' (formerly Cohere For AI's) 32.3-billion-parameter chat model, an August 2024 refresh of the original Command R built for retrieval-augmented generation with citations, single-step tool use, and multi-step agentic tool use. It is an auto-regressive transformer using grouped-query attention for faster inference, trained and evaluated across a wide multilingual set including English, French, Spanish, German, Japanese, Korean, Arabic, and Simplified Chinese, among others. Cohere positions this release around stronger grounded RAG capability and broader multilingual coverage rather than a change in scale. At 32.3 billion parameters, it needs a high-end consumer GPU or a multi-GPU setup once quantized. Context length is 128,000 tokens. It is released under a Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license plus Cohere's Acceptable Use Policy, restricting the model to non-commercial use. It was published in August 2024.
Starling LM 7B Beta
Nexusflow · 7.2B · runs from 3.6 GB
Starling-LM-7B-beta is Nexusflow's chat model, fine-tuned from Openchat-3.5-0106 (itself based on Mistral-7B-v0.1) using reinforcement learning from AI feedback (RLAIF). It was trained with Nexusflow's own 34B reward model and a PPO-based policy optimization pipeline on the Nectar preference dataset, and scored 8.12 on MT-Bench with GPT-4 as judge, an improvement over the team's earlier Starling model. It uses OpenChat's exact chat template rather than a generic one. At 7 billion parameters it runs comfortably on a single consumer GPU. Context length is 8,192 tokens. It is released under the Apache 2.0 license with an added condition that it not be used to compete with OpenAI, reflecting that its Nectar training data was generated with GPT-4. It was published in March 2024.
Chatglm3 6B
Z.ai · 6.2B · runs from 2.9 GB
ChatGLM3-6B is the third generation of Zhipu AI's (now Z.ai) open bilingual Chinese-English chat model, built on the GLM architecture at around 6.2 billion parameters. Compared with its predecessors it adds native support for function calling, a code interpreter, and agent-style tasks through a newly designed prompt format, alongside a stronger base model, ChatGLM3-6B-Base, trained on more diverse data. A companion long-context variant, ChatGLM3-6B-32K, was released alongside it. Its modest size lets it run on a single consumer GPU. Context length is 8,192 tokens. The code is released under Apache 2.0, but the model weights use a separate Model License that is free for academic research and allows commercial use only after completing Zhipu's registration questionnaire. It was published in October 2023, and has since been superseded by the GLM-4 series.
JiRackUltra 1B
CMSManhattan · 1.8B · runs from 0.8 GB
JiRackUltra 1B is a 1.8B-parameter open language model from CMSManhattan. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 Math 1.5B
Alibaba · 1.5B · runs from 1 GB
Qwen2.5 Math 1.5B is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
StableBeluga2
Stability AI · 70B · runs from 20.2 GB
StableBeluga2 is Stability AI's chat-tuned, 70-billion-parameter model, an Orca-style supervised fine-tune of Meta's Llama 2 70B base rather than a model trained from scratch. It follows Stability's internal reimplementation of the Orca training recipe, using a structured System/User/Assistant prompt format, and was trained in mixed BF16 precision with AdamW; smaller Stable Beluga 7B and 13B siblings were released alongside it. It is a historically significant early Llama-2 fine-tune from the 2023 open-model wave rather than a current state-of-the-art model by 2026 standards. At 70 billion parameters, it needs a multi-GPU workstation or server to run, even once quantized. Context length is 4,096 tokens, inherited from the Llama 2 base. It is released under Stability AI's Stable Beluga Non-Commercial Community License, which restricts the fine-tuned weights to non-commercial use. It was published in July 2023.
Fanar 2 27B Instruct
QCRI · 27.0B · runs from 9.1 GB
Fanar 2 27B Instruct is a 27.0B-parameter open language model from QCRI. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
EXAONE 4.0 1.2B
LGAI-EXAONE · 1.3B · runs from 1.0 GB
EXAONE 4.0 1.2B is a 1.3B-parameter open language model from LGAI-EXAONE in the EXAONE family. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen1.5 1.8B Chat
Alibaba · 1.8B · runs from 1.5 GB
Qwen1.5 1.8B Chat is a 1.8B-parameter open language model from Alibaba in the Qwen family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Bloom 560M
BigScience · 559M · runs from 0.3 GB
Bloom 560M is a 559M-parameter open language model from BigScience. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
SmolLM 360M Instruct
Hugging Face · 362M · runs from 0.5 GB
SmolLM 360M Instruct is a 362M-parameter open language model from Hugging Face in the SmolLM family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
LightOnOCR 2 1B
lightonai · 1.0B · runs from 0.8 GB
LightOnOCR-2-1B is LightOn's flagship end-to-end OCR vision-language model, a 1-billion-parameter system that converts documents such as PDFs, scans, and photos directly into clean, naturally ordered markdown without a separate OCR pipeline. It handles tables, receipts, forms, multi-column layouts, and math notation, is trained on a large multilingual corpus with particular strength in French, arXiv papers, and scanned documents, and this second-generation release adds RLVR (reinforcement learning from verifiable rewards) training on top of the base model for extra accuracy. LightOn reports state-of-the-art results on the OlmOCR-Bench benchmark while being roughly nine times smaller and several times faster than competing OCR systems. At just 1 billion parameters, it runs comfortably even on a modest single consumer GPU. Context length is 16,384 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in January 2026, alongside base, bounding-box, and merged "soup" variants of the same model.
Deepseek Coder 7B Instruct V1.5
DeepSeek · 6.9B · runs from 4.2 GB
Deepseek Coder 7B Instruct V1.5 is a 6.9-billion-parameter code-focused language model from DeepSeek, fine-tuned from the DeepSeek-LLM 7B base for programming assistance, code generation, and general chat. It continues DeepSeek's original Coder line, trained on roughly 2 trillion tokens of code-heavy text before instruction tuning, and its default system prompt frames it specifically as a programming assistant. Its size makes it well suited to local deployment on a single mainstream consumer GPU with around 8GB or more of VRAM once quantized. Context length is limited to 4,096 tokens, short by current standards. It is released under DeepSeek's own custom license (listed as "Other"), so commercial users should review the license terms before deployment. Published in January 2024, it is one of DeepSeek's earlier widely adopted open-weight coding assistants.
NVIDIA Nemotron 3.5 Lightning 30B A3B Base BF16
NVIDIA · 31.6B · runs from 13.8 GB
NVIDIA Nemotron 3.5 Lightning 30B A3B Base BF16 is a 31.6B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
TIPO 500M Ft
KBlueLeaf · 508M · runs from 0.7 GB
TIPO 500M Ft is a 508M-parameter open language model from KBlueLeaf. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ring Flash 2.0
Inclusion AI · 102.9B · runs from 28.7 GB
Ring-flash-2.0 is Inclusion AI's reasoning ("thinking") model, a mixture-of-experts system built on the same Ling 2.0 architecture and Ling-flash-base-2.0 checkpoint as Ling-flash-2.0, sharing its roughly 102.9-billion total and about 6.15-billion active parameter footprint. It adds long chain-of-thought supervised fine-tuning followed by reinforcement learning with verifiable rewards and a further RLHF stage, using Inclusion AI's own IcePop algorithm to stabilize MoE reinforcement learning against training-inference probability mismatches during long rollouts. The company reports it leads comparable open thinking models on math, code, and logical-reasoning benchmarks while retaining the creative-writing strength of its non-thinking twin, Ling-flash-2.0. Its small active-parameter footprint keeps generation fast, but all of its parameters must stay in memory, so it needs a multi-GPU setup or a high-memory workstation even once quantized. Context length is natively 32,768 tokens, extendable to 128,000 tokens with YaRN. It is released under the MIT license, permitting unrestricted commercial and research use. It was published in September 2025, alongside its non-reasoning sibling Ling-flash-2.0.
Agents A1
InternScience · 35.1B · runs from 15.3 GB
Agents-A1 is InternScience's 35-billion-parameter mixture-of-experts agentic model, built to handle long-horizon tool use, engineering, scientific research, and instruction-following tasks rather than general chit-chat. It routes to roughly 3.9 billion active parameters per token and is trained with a three-stage recipe: broad supervised fine-tuning on agentic behaviors, domain-specialist teacher models, and multi-teacher multi-domain on-policy distillation that transfers that expertise back into one general model. The card reports it approaching or matching frontier-scale systems like GPT-5.5 and Kimi-K2.6 on agentic benchmarks such as BrowseComp, GAIA, and SciCode despite its far smaller size, and it needs a capable multi-GPU workstation to run at full precision, less once quantized. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in June 2026; the same family later added a smaller 4B model for local deployment.
Racka 4B
elte-nlp · 4.0B · runs from 1.9 GB
Racka 4B is a 4.0B-parameter open language model from elte-nlp. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Baichuan2 13B Chat
baichuan-inc · 13B · runs from 3.9 GB
Baichuan2-13B-Chat is Baichuan Intelligence's 13-billion-parameter bilingual (Chinese/English) chat model, instruction-aligned from the Baichuan2-13B-Base model that was pretrained from scratch on 2.6 trillion tokens. It was evaluated across general, legal, medical, math, code, and multilingual-translation benchmarks against contemporaries like LLaMA2-13B-Chat and Vicuna-13B, and a 4-bit quantized version is also distributed for lower-memory deployment. At 13B parameters it needs a capable consumer GPU at full precision, considerably less once quantized to 4 bits. License is a custom Baichuan2 Community License: free for academic research, and free for commercial use after obtaining a license via email request to the developers. It was published in August 2023, with a v2 revision issued in December 2023 that improved math, logical reasoning, and instruction-following.
Falcon H1 0.5B Base
TII UAE · 521M · runs from 0.5 GB
Falcon H1 0.5B Base is a 521M-parameter open language model from TII UAE in the Falcon family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.