All LLM Models
Browse 1214 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Gemma 1.1 2B IT
Google · 2.5B · runs from 1.1 GB
Gemma 1.1 2B IT is a 2.5B-parameter open language model from Google in the Gemma family. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Llama 3 3 Nemotron Super 49B V1 5
NVIDIA · 49.9B · runs from 15.1 GB
Llama 3.3 Nemotron Super 49B is a 49.9-billion parameter chat model by NVIDIA, built on a modified Llama 3.3 architecture. It occupies a unique size point between the common 70B and 8B tiers, offering strong reasoning and conversational ability while requiring less VRAM than full 70B models. NVIDIA's Nemotron Super training pipeline applies extensive alignment tuning to optimize helpfulness and factual accuracy. The model typically needs 32GB or more of VRAM for local inference at reduced precision, placing it within reach of high-end consumer GPUs like the RTX 4090 or professional workstation cards.
C4ai Command R 08 2024
Cohere · 32.3B · runs from 14.7 GB
C4AI Command R 08-2024 is Cohere Labs' (formerly Cohere For AI's) 32.3-billion-parameter chat model, an August 2024 refresh of the original Command R built for retrieval-augmented generation with citations, single-step tool use, and multi-step agentic tool use. It is an auto-regressive transformer using grouped-query attention for faster inference, trained and evaluated across a wide multilingual set including English, French, Spanish, German, Japanese, Korean, Arabic, and Simplified Chinese, among others. Cohere positions this release around stronger grounded RAG capability and broader multilingual coverage rather than a change in scale. At 32.3 billion parameters, it needs a high-end consumer GPU or a multi-GPU setup once quantized. Context length is 128,000 tokens. It is released under a Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license plus Cohere's Acceptable Use Policy, restricting the model to non-commercial use. It was published in August 2024.
Starling LM 7B Beta
Nexusflow · 7.2B · runs from 3.6 GB
Starling-LM-7B-beta is Nexusflow's chat model, fine-tuned from Openchat-3.5-0106 (itself based on Mistral-7B-v0.1) using reinforcement learning from AI feedback (RLAIF). It was trained with Nexusflow's own 34B reward model and a PPO-based policy optimization pipeline on the Nectar preference dataset, and scored 8.12 on MT-Bench with GPT-4 as judge, an improvement over the team's earlier Starling model. It uses OpenChat's exact chat template rather than a generic one. At 7 billion parameters it runs comfortably on a single consumer GPU. Context length is 8,192 tokens. It is released under the Apache 2.0 license with an added condition that it not be used to compete with OpenAI, reflecting that its Nectar training data was generated with GPT-4. It was published in March 2024.
Chatglm3 6B
Z.ai · 6.2B · runs from 2.9 GB
ChatGLM3-6B is the third generation of Zhipu AI's (now Z.ai) open bilingual Chinese-English chat model, built on the GLM architecture at around 6.2 billion parameters. Compared with its predecessors it adds native support for function calling, a code interpreter, and agent-style tasks through a newly designed prompt format, alongside a stronger base model, ChatGLM3-6B-Base, trained on more diverse data. A companion long-context variant, ChatGLM3-6B-32K, was released alongside it. Its modest size lets it run on a single consumer GPU. Context length is 8,192 tokens. The code is released under Apache 2.0, but the model weights use a separate Model License that is free for academic research and allows commercial use only after completing Zhipu's registration questionnaire. It was published in October 2023, and has since been superseded by the GLM-4 series.
JiRackUltra 1B
CMSManhattan · 1.8B · runs from 0.8 GB
JiRackUltra 1B is a 1.8B-parameter open language model from CMSManhattan. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen2.5 Math 1.5B
Alibaba · 1.5B · runs from 1 GB
Qwen2.5 Math 1.5B is a 1.5B-parameter open language model from Alibaba in the Qwen 2.5 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
StableBeluga2
Stability AI · 70B · runs from 20.2 GB
StableBeluga2 is Stability AI's chat-tuned, 70-billion-parameter model, an Orca-style supervised fine-tune of Meta's Llama 2 70B base rather than a model trained from scratch. It follows Stability's internal reimplementation of the Orca training recipe, using a structured System/User/Assistant prompt format, and was trained in mixed BF16 precision with AdamW; smaller Stable Beluga 7B and 13B siblings were released alongside it. It is a historically significant early Llama-2 fine-tune from the 2023 open-model wave rather than a current state-of-the-art model by 2026 standards. At 70 billion parameters, it needs a multi-GPU workstation or server to run, even once quantized. Context length is 4,096 tokens, inherited from the Llama 2 base. It is released under Stability AI's Stable Beluga Non-Commercial Community License, which restricts the fine-tuned weights to non-commercial use. It was published in July 2023.
Fanar 2 27B Instruct
QCRI · 27.0B · runs from 9.1 GB
Fanar 2 27B Instruct is a 27.0B-parameter open language model from QCRI. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
EXAONE 4.0 1.2B
LGAI-EXAONE · 1.3B · runs from 1.0 GB
EXAONE 4.0 1.2B is a 1.3B-parameter open language model from LGAI-EXAONE in the EXAONE family. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen1.5 1.8B Chat
Alibaba · 1.8B · runs from 1.5 GB
Qwen1.5 1.8B Chat is a 1.8B-parameter open language model from Alibaba in the Qwen family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Bloom 560M
BigScience · 559M · runs from 0.3 GB
Bloom 560M is a 559M-parameter open language model from BigScience. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
SmolLM 360M Instruct
Hugging Face · 362M · runs from 0.5 GB
SmolLM 360M Instruct is a 362M-parameter open language model from Hugging Face in the SmolLM family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
LightOnOCR 2 1B
lightonai · 1.0B · runs from 0.8 GB
LightOnOCR-2-1B is LightOn's flagship end-to-end OCR vision-language model, a 1-billion-parameter system that converts documents such as PDFs, scans, and photos directly into clean, naturally ordered markdown without a separate OCR pipeline. It handles tables, receipts, forms, multi-column layouts, and math notation, is trained on a large multilingual corpus with particular strength in French, arXiv papers, and scanned documents, and this second-generation release adds RLVR (reinforcement learning from verifiable rewards) training on top of the base model for extra accuracy. LightOn reports state-of-the-art results on the OlmOCR-Bench benchmark while being roughly nine times smaller and several times faster than competing OCR systems. At just 1 billion parameters, it runs comfortably even on a modest single consumer GPU. Context length is 16,384 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in January 2026, alongside base, bounding-box, and merged "soup" variants of the same model.
Deepseek Coder 7B Instruct V1.5
DeepSeek · 6.9B · runs from 4.2 GB
Deepseek Coder 7B Instruct V1.5 is a 6.9-billion-parameter code-focused language model from DeepSeek, fine-tuned from the DeepSeek-LLM 7B base for programming assistance, code generation, and general chat. It continues DeepSeek's original Coder line, trained on roughly 2 trillion tokens of code-heavy text before instruction tuning, and its default system prompt frames it specifically as a programming assistant. Its size makes it well suited to local deployment on a single mainstream consumer GPU with around 8GB or more of VRAM once quantized. Context length is limited to 4,096 tokens, short by current standards. It is released under DeepSeek's own custom license (listed as "Other"), so commercial users should review the license terms before deployment. Published in January 2024, it is one of DeepSeek's earlier widely adopted open-weight coding assistants.
NVIDIA Nemotron 3.5 Lightning 30B A3B Base BF16
NVIDIA · 31.6B · runs from 13.8 GB
NVIDIA Nemotron 3.5 Lightning 30B A3B Base BF16 is a 31.6B-parameter open language model from NVIDIA in the Nemotron family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
TIPO 500M Ft
KBlueLeaf · 508M · runs from 0.7 GB
TIPO 500M Ft is a 508M-parameter open language model from KBlueLeaf. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Agents A1
InternScience · 35.1B · runs from 15.3 GB
Agents-A1 is InternScience's 35-billion-parameter mixture-of-experts agentic model, built to handle long-horizon tool use, engineering, scientific research, and instruction-following tasks rather than general chit-chat. It routes to roughly 3.9 billion active parameters per token and is trained with a three-stage recipe: broad supervised fine-tuning on agentic behaviors, domain-specialist teacher models, and multi-teacher multi-domain on-policy distillation that transfers that expertise back into one general model. The card reports it approaching or matching frontier-scale systems like GPT-5.5 and Kimi-K2.6 on agentic benchmarks such as BrowseComp, GAIA, and SciCode despite its far smaller size, and it needs a capable multi-GPU workstation to run at full precision, less once quantized. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in June 2026; the same family later added a smaller 4B model for local deployment.
Racka 4B
elte-nlp · 4.0B · runs from 1.9 GB
Racka 4B is a 4.0B-parameter open language model from elte-nlp. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Baichuan2 13B Chat
baichuan-inc · 13B · runs from 3.9 GB
Baichuan2-13B-Chat is Baichuan Intelligence's 13-billion-parameter bilingual (Chinese/English) chat model, instruction-aligned from the Baichuan2-13B-Base model that was pretrained from scratch on 2.6 trillion tokens. It was evaluated across general, legal, medical, math, code, and multilingual-translation benchmarks against contemporaries like LLaMA2-13B-Chat and Vicuna-13B, and a 4-bit quantized version is also distributed for lower-memory deployment. At 13B parameters it needs a capable consumer GPU at full precision, considerably less once quantized to 4 bits. License is a custom Baichuan2 Community License: free for academic research, and free for commercial use after obtaining a license via email request to the developers. It was published in August 2023, with a v2 revision issued in December 2023 that improved math, logical reasoning, and instruction-following.
Falcon H1 0.5B Base
TII UAE · 521M · runs from 0.5 GB
Falcon H1 0.5B Base is a 521M-parameter open language model from TII UAE in the Falcon family. It supports a context window of up to 16,384 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
OLMo 2 0425 1B
Allen AI · 1.5B · runs from 1.2 GB
OLMo 2 1B is Allen AI's smallest base language model in the OLMo 2 family, a 1.48-billion-parameter dense transformer trained on 4 trillion tokens. This is the raw pretrained checkpoint, not an instruction-tuned assistant — Allen AI releases separate SFT, DPO, and RLVR1 versions for chat use. Unusually for an open-weight release, Allen AI also publishes the training data, code, and intermediate checkpoints, making it a reference point for reproducible LLM research. At 1.48B parameters it runs comfortably on modest hardware, even unquantized. It supports a 4,096 token context window, short by current standards. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. Published in April 2025, it is the smallest of four OLMo 2 sizes (1B, 7B, 13B, 32B) sharing the same training recipe and fully open data pipeline.
Qwen1.5 32B Chat
Alibaba · 32.5B · runs from 14.3 GB
Qwen1.5-32B-Chat is Alibaba's instruction-tuned, 32.5-billion-parameter chat model from the Qwen1.5 series, a beta release of the Qwen2 architecture that sits between the 14B and 72B dense models in the lineup. Qwen1.5 improved on the original Qwen with stable 32K context support across all model sizes, broader multilingual coverage, and no need for custom trust_remote_code, and this 32B checkpoint additionally uses grouped-query attention, unlike the smaller Qwen1.5 sizes, for faster inference. It was aligned on top of the pretrained base with supervised fine-tuning and direct preference optimization. At 32.5 billion parameters it needs a high-end consumer GPU, or a multi-GPU setup once quantized, to run comfortably. Context length is 32,768 tokens. It is released under Alibaba's Tongyi Qianwen license, a custom license that is free for most commercial and research use but requires a separate license from Alibaba once a deployment exceeds 100 million monthly active users. It was published in April 2024, part of Alibaba's second LLM generation, later followed by Qwen2 and Qwen2.5.
TwIL LM3
webAI-Official · 3.1B · runs from 1.5 GB
TwIL LM3 is a 3.1B-parameter open language model from webAI-Official. It supports a context window of up to 65,536 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Dots.ocr
dots-studio · 3.0B · runs from 1.6 GB
dots.ocr is a 3-billion-parameter vision-language model built around a compact 1.7-billion-parameter language backbone, purpose-built for multilingual document layout parsing and OCR. It unifies layout detection and text recognition in one model, switching tasks — table extraction, formula recognition, reading order — by changing the prompt, instead of chaining separate detection models. Despite its small footprint, its developers report state-of-the-art results on document-parsing benchmarks, including strong performance on low-resource languages. It runs comfortably on a single consumer GPU. Context length is 131,072 tokens, ample for long documents. It is released under the MIT license, a highly permissive option for commercial use. Published in July 2025, it shows a compact single VLM can rival dedicated layout-detection models like DocLayout-YOLO.
Qwen1.5 32B
Alibaba · 32.5B · runs from 14.3 GB
Qwen1.5-32B is Alibaba's 32.5-billion-parameter dense base language model, one of eight sizes (0.5B to 72B, plus a 14B mixture-of-experts variant) in the Qwen1.5 series, a beta preview of what became Qwen2. It is a raw pretrained Transformer with SwiGLU activation, QKV attention bias, and group-query attention (added specifically for the 32B and larger sizes), and Alibaba does not recommend using it directly for chat, only as a foundation for further fine-tuning or alignment. At 32.5B parameters it needs a high-end consumer GPU or multi-GPU setup once quantized, more at full precision. Context length is 32,768 tokens, stable across all Qwen1.5 model sizes. It is released under a custom Tongyi Qianwen Research License, free for research and academic use, with commercial deployment requiring a separate license from Alibaba. It was published in April 2024.
Kimi K3 DSpark
RadixArk · 2.2B · runs from 1.3 GB
Kimi K3 DSpark is a 2.2B-parameter open language model from RadixArk in the Kimi K3 family. It supports a context window of up to 1,048,576 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 E2B IT Qat Q4 0 Unquantized Heretic
coder3101 · 5.1B · runs from 2.5 GB
Gemma 4 E2B IT Qat Q4 0 Unquantized Heretic is a 5.1B-parameter open language model from coder3101 in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Qwen3.5 4B Base
Alibaba · 4.7B · runs from 2.5 GB
Qwen3.5-4B-Base is a 4.7-billion-parameter dense checkpoint in Alibaba's Qwen3.5 base family, a native vision-language foundation model trained with early fusion of image and text tokens rather than a text model with a bolted-on vision tower. It is pretrained-only weights meant for fine-tuning or research, not direct conversation, though its control tokens support efficient LoRA-style adaptation with the official chat template. It uses the same hybrid Gated DeltaNet plus gated-attention architecture as its siblings. At under 5B parameters, it runs comfortably on a single mainstream consumer GPU. Context length is a native 262,144 tokens, extensible up to 1,010,000 tokens. It is released under the Apache 2.0 license, and was published in February 2026. It sits mid-lineup among five Qwen3.5-Base sizes, between the 2B and 9B dense checkpoints, with a 35B mixture-of-experts variant at the top.
DeepSeek OCR
DeepSeek · 3.3B · runs from 1.8 GB
DeepSeek OCR is a 3.3-billion-parameter vision-language model from DeepSeek, purpose-built for optical character recognition and document parsing rather than general chat. It pairs a vision encoder with a Mixture-of-Experts decoder that activates roughly 1.1 billion parameters per token, keeping decoding fast while all expert weights still need to fit in memory. Its core idea is compressing a page of text into a much smaller set of image tokens before decoding, and it is small enough to run on a single consumer GPU once quantized. Context length is limited to 8,192 tokens, reflecting its page-oriented use case. It is released under the MIT license, a highly permissive option for commercial use, and was published in October 2025, introducing "optical context compression" to shrink the token count needed for OCR.