All LLM Models
Browse 54 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Gemma 4 E4B IT
Google · 8.0B · runs from 3.2 GB
Gemma 4 E4B IT packs Google's Gemma 4 architecture into a compact, roughly 8-billion-parameter footprint, positioned as the mid-sized option in Gemma 4's efficiency-focused E-series alongside the smaller E2B variant. It is tuned for chat and instruction-following rather than multimodal input, focusing on dialogue quality within a lightweight package. Parameter counts in this range are well suited to local inference on a single consumer GPU, even at higher precision, and comfortably so once quantized. The model provides a 128K token context window, sufficient for most chat and document-assistance use cases without needing the largest Gemma 4 variants. It is released under the Apache 2.0 license, and was published on March 2, 2026, shortly before the larger dense and mixture-of-experts Gemma 4 models that followed.
Gemma 4 31B IT
Google · 31.3B · runs from 10.2 GB
Gemma 4 31B IT is Google's 31-billion-parameter instruction-tuned model in the Gemma 4 lineup, built to handle both text and image input in a single pass. As a vision-capable model, it can describe, compare, or reason about images alongside written prompts, making it suitable for multimodal chat and document-understanding tasks. At this parameter count, local inference calls for quantization and a fairly capable GPU; it fits on a single high-end consumer or workstation card rather than lower-end hardware. The model supports a 256K token context window, enough for long documents, transcripts, or multi-turn conversations without aggressive truncation. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use, and reflects Google's continued push toward mid-sized, multimodal open-weight models following the earlier Gemma generations.
Gemma 4 26B A4B IT
Google · 25.8B · runs from 11.6 GB
Gemma 4 26B A4B IT applies a mixture-of-experts design to Google's Gemma 4 family, with roughly 26 billion total parameters but only about 4 billion active for any given token. That active-parameter count is what determines inference speed, so despite its total size the model can respond about as quickly as a much smaller dense model, though the full parameter set still needs to be held in memory, putting local use in single high-end consumer GPU territory once quantized. It also accepts image input alongside text, making it usable for multimodal chat and visual question answering. The model offers a 256K token context window, suited to long documents or extended conversations, and is released under the Apache 2.0 license for unrestricted use. Released alongside the dense 31B variant, it gives developers a faster-inference option within the same Gemma 4 generation.
Gemma 4 12B IT Qat Q4 0 Unquantized
Google · 12.0B · runs from 6.1 GB
Gemma 4 12B IT Qat Q4 0 Unquantized is a 12.0B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 E2B IT Qat Q4 0 Unquantized
Google · 5.1B · runs from 2.5 GB
Gemma 4 E2B IT Qat Q4 0 Unquantized is a 5.1B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 12B IT
Google · 12.0B · runs from 5.1 GB
Sitting between Gemma 4's compact E-series and its larger 31B sibling, Gemma 4 12B IT is a 12-billion-parameter instruction-tuned model built for general chat and dialogue use. It does not take image input, focusing instead on text-based reasoning, coding help, and conversational tasks. At this parameter count, the model fits comfortably on a single consumer GPU without necessarily requiring the heaviest quantization, making it a reasonable middle-ground choice for local setups. The model supports a generous 256K token context window, useful for long documents or extended sessions, and is released under the Apache 2.0 license. It arrived in May 2026, after the initial March 2026 releases of the Gemma 4 E-series and 31B dense model, suggesting continued iteration within the family after the launch wave.
Gemma 4 26B A4B IT Qat Q4 0 Unquantized
Google · 26.5B · runs from 11.9 GB
Gemma 4 26B A4B IT Qat Q4 0 Unquantized is a 26.5B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 E2B IT
Google · 5.1B · runs from 2.1 GB
The smallest model discussed here from Google's Gemma 4 E-series, Gemma 4 E2B IT, comes in at roughly 5 billion parameters. It is tuned for chat and instruction-following rather than multimodal tasks, prioritizing a minimal footprint over broad input support. Models at this size are comfortable on typical consumer GPUs, and can even run on integrated graphics when quantized aggressively, making it one of the more accessible options for local deployment. It ships with a 128K token context window, matching its larger E4B sibling, and is released under the Apache 2.0 license for open commercial and research use. Google published it on March 2, 2026, alongside the E4B variant, as the entry point into the Gemma 4 generation.
Gemma 4 E4B IT Qat Q4 0 Unquantized
Google · 7.9B · runs from 3.9 GB
Gemma 4 E4B IT Qat Q4 0 Unquantized is a 7.9B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 31B IT Qat Q4 0 Unquantized
Google · 32.7B · runs from 15.5 GB
Gemma 4 31B IT Qat Q4 0 Unquantized is a 32.7B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 3 1B IT
Google · 1000M · runs from 0.3 GB
Google Gemma 3 1B IT is a 1-billion parameter instruction-tuned model from Google's Gemma 3 family. It is an ultra-compact text-only chat model designed for deployment on minimal hardware, including low-VRAM GPUs and edge devices. The model handles basic conversational tasks, simple instruction following, and lightweight text generation. It can run on virtually any modern GPU and even on CPU-only setups with acceptable latency. Released under the Gemma license.
Gemma 2 9B IT
Google · 9.2B · runs from 3.0 GB
Google Gemma 2 9B IT is a 9.2-billion parameter instruction-tuned model from Google's Gemma 2 series. It is a text-only chat model optimized for conversational tasks, instruction following, and general-purpose assistance. At release, it was recognized for delivering unusually strong performance relative to its parameter count. The model runs efficiently on consumer GPUs with 8-12GB of VRAM in quantized formats, making it accessible on mainstream hardware. It is a popular choice for local inference among users who want strong quality without the VRAM demands of larger models. Released under the Gemma license.
Gemma 2 2B IT
Google · 2.6B · runs from 0.9 GB
Google Gemma 2 2B IT is a 2-billion parameter instruction-tuned model from Google's Gemma 2 family, the smallest variant in the Gemma 2 series. It is designed for efficient local inference on resource-constrained hardware, handling basic conversational tasks and simple instruction following at minimal compute cost. The model can run on GPUs with as little as 4GB of VRAM when quantized, and even on CPU-only setups. Released under the Gemma license.
Medgemma 27B Text IT
Google · 27.0B · runs from 59.4 GB
Google MedGemma 27B Text IT is a 27-billion parameter instruction-tuned model specialized for the medical domain, built on the Gemma architecture by Google. It is fine-tuned on medical and clinical text data to provide improved performance on healthcare-related tasks such as medical question answering, clinical reasoning, and health information summarization. The model requires a GPU with at least 24GB of VRAM for quantized inference. Its domain specialization makes it notably more capable than general models on clinical benchmarks, though it should not be used as a substitute for professional medical advice. Released under the Gemma license.
Gemma 3 4B IT
Google · 4.3B · runs from 2.0 GB
Gemma 3 4B IT is a small, roughly 4.3-billion-parameter model from Google's earlier Gemma 3 generation, tuned for instruction-following and chat. It accepts image input alongside text, so it can handle visual question answering and image-grounded prompts in addition to text-only conversation. At this size, the model is comfortable on modest consumer hardware, running even on integrated graphics at low quantization, making it one of the most accessible vision-capable models for local use. Released on February 20, 2025, Gemma 3 4B IT predates the Gemma 4 family and is distributed under Google's Gemma license terms rather than a standard open-source license. Its compact size and image-input support make it a practical choice for lightweight, on-device multimodal applications.
Gemma 3 12B IT
Google · 12.2B · runs from 5.7 GB
Google Gemma 3 12B IT is a 12-billion parameter multimodal instruction-tuned model from Google's Gemma 3 series. It supports both text and image inputs, offering vision-language capabilities at a more accessible size point than the 27B variant. Gemma 3 12B IT runs on consumer GPUs with 12-16GB of VRAM in quantized formats, making it a practical choice for local multimodal inference without requiring top-tier hardware. Released under the Gemma license.
Gemma 3 27B IT
Google · 27.4B · runs from 12.8 GB
Google Gemma 3 27B IT is a 27.4-billion parameter multimodal instruction-tuned model from Google's Gemma 3 family. It supports both text and image inputs, making it one of the most capable openly available vision-language models for local inference. The model handles conversational AI, visual question answering, image description, and complex reasoning tasks across modalities. Gemma 3 27B IT requires a GPU with at least 24GB of VRAM for quantized inference, placing it within reach of high-end consumer cards like the RTX 4090. It uses a dense Transformer architecture with a large context window and benefits from Google's extensive pretraining pipeline. Released under the Gemma license.
Gemma 3 270M IT
Google · 268M · runs from 0.1 GB
Google Gemma 3 270M IT is a 270-million parameter instruction-tuned model from Google's Gemma 3 family, an experimental release pushing the boundaries of how small an effective chat model can be. The model runs on virtually any hardware, including entry-level GPUs and CPU-only setups, making it useful for experimentation, education, and exploring the limits of small-scale language modeling. Released under the Gemma license.
Diffusiongemma 26B A4B IT
Google · 25.8B · runs from 11.6 GB
Diffusiongemma 26B A4B IT takes a different approach from the rest of the Gemma 4 family: rather than generating tokens autoregressively, it uses a block-diffusion architecture. It still follows a mixture-of-experts design, with roughly 26 billion total parameters and about 4 billion active per token. The active-parameter figure is the more relevant one for inference cost, while the total parameter count determines memory needs, putting local use in single high-end consumer GPU territory once quantized. It also accepts image input alongside text for multimodal use. The model offers a 256K token context window and is released under the Apache 2.0 license. Published on June 9, 2026, it represents Google's exploration of diffusion-based language modeling within the Gemma 4 generation, alongside the more conventional autoregressive Gemma 4 variants also released that year.
Gemma 3n E2B IT
Google · 5.4B · runs from 1.6 GB
Gemma 3n E2B is the smaller of Google's two Gemma 3n instruction-tuned models, designed for phones, laptops and other on-device use. "E2B" stands for an effective size of about 2 billion parameters: the raw checkpoint holds around 5.4 billion, but techniques such as per-layer embeddings let it run with a memory footprint closer to a 2B model. It accepts text, image and audio input and generates text. The model has a 32K-token context window and is released under Google's Gemma terms of use. At 4-bit quantization it needs only a few gigabytes of memory, so it runs on almost any modern GPU, laptop or recent phone.
Gemma 4 E2B IT Qat Mobile Transformers
Google · 2.3B · runs from 1.4 GB
Gemma 4 E2B IT Qat Mobile Transformers is a 2.3B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 3n E4B IT
Google · 7.8B · runs from 2.4 GB
Gemma 3n E4B IT is the instruction-tuned variant of Google's Gemma 3n E4B, a multimodal model built for on-device use on phones, laptops, and tablets that accepts text, image, audio, and video input and can perform automatic speech recognition and speech translation alongside text chat. It uses Google's MatFormer (Matryoshka Transformer) architecture, which nests a smaller sub-model inside the full network so the same checkpoint can run at reduced effective capacity, and Per-Layer Embedding, which caches embedding parameters to fast local storage instead of holding them all in memory. The "E4B" designation refers to an effective parameter count of around 4 billion at inference, even though the checkpoint's total parameters are larger; either way it is light enough for a single consumer GPU or even a high-end phone. Context length is 32,768 tokens. It is released under Google's Gemma Terms of Use, a custom license permitting broad commercial and research use alongside a prohibited-use policy, and was published in June 2025.
Medgemma 1.5 4B IT
Google · 4.3B · runs from 1.3 GB
MedGemma 1.5 4B is Google's 4.3-billion-parameter instruction-tuned vision-language model in the MedGemma line of health-AI foundation models, succeeding the original MedGemma 4B at the same size. It is built for developers creating healthcare applications, covering tasks such as medical image interpretation and clinical text understanding, and is not a validated diagnostic tool: outputs need independent verification and further evaluation before any clinical use. At this size it runs on a single consumer GPU once quantized. It is released under Google's Health AI Developer Foundations terms, a use-restricted license that requires accepting specific health-AI conditions on Hugging Face rather than a fully open one. Published in January 2026, it is the second 4B release in the MedGemma series.
Medgemma 4B IT
Google · 4.3B · runs from 1.3 GB
MedGemma 4B is Google's 4.3-billion-parameter instruction-tuned vision-language model, built on Gemma 3 and further trained for medical text and image understanding. It adds a medical image encoder covering chest X-rays, dermatology photos, histopathology slides, and fundus images, and can generate findings and answer questions about medical images. It is a foundation for developers building healthcare applications, not a validated diagnostic tool, and needs further evaluation before clinical use. It runs on a single consumer GPU once quantized. It inherits Gemma 3's 128,000 token context window. It is released under Google's Health AI Developer Foundations terms, a use-restricted license rather than a fully open one. Published in May 2025, it was Google's first instruction-tuned multimodal MedGemma release, alongside a larger text-only 27B variant.
Gemma 4 26B A4B IT Assistant
Google · 26B · runs from 11.4 GB
Gemma 4 26B A4B IT Assistant is a 26B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 2 27B IT
Google · 27.2B · runs from 9.0 GB
Google Gemma 2 27B IT is a 27.2-billion parameter instruction-tuned model from Google's Gemma 2 generation. It is a text-only chat model optimized for conversational use, reasoning, and instruction following. Gemma 2 27B IT was one of the strongest openly available models in its size class at release. The model requires a GPU with at least 24GB of VRAM for quantized local inference. It is widely supported by popular inference engines and remains a strong choice for users seeking high-quality local chat without needing 70B-class hardware. Released under the Gemma license.
Gemma 4 E4B IT Assistant
Google · 4B · runs from 2 GB
Gemma 4 E4B IT Assistant is a 4B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 E4B
Google · 8.0B · runs from 3.9 GB
Gemma 4 E4B is Google DeepMind's second-smallest model in the Gemma 4 family, a dense architecture with roughly 8 billion total parameters, of which Google describes about 4.5 billion as its effective footprint at inference. This is the pretrained base checkpoint rather than an instruction-tuned model, meant for fine-tuning rather than direct chat use. Like the rest of the family it is multimodal — text, image, and natively audio at this size — targeting efficient on-device deployment on laptops and higher-end phones. It runs comfortably on a single consumer GPU, or on-device once quantized. It supports a 131,072 token context window. It carries Google's Gemma 4 license terms, published as Apache 2.0 on Hugging Face with additional Gemma-specific usage terms linked from the card. Published in March 2026, it sits between the E2B on-device model and the larger 12B, 26B-A4B, and 31B tiers.
Gemma 4 31B IT Qat Q4 0 Unquantized Assistant
Google · 31B · runs from 13.5 GB
Gemma 4 31B IT Qat Q4 0 Unquantized Assistant is a 31B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Gemma 4 31B
Google · 32.7B · runs from 15.5 GB
Gemma 4 31B is Google DeepMind's largest dense model in the Gemma 4 family, roughly 32.7 billion parameters, handling text and image input. This is the pretrained base checkpoint rather than an instruction-tuned model, meant as a foundation for fine-tuning rather than direct chat use. Gemma 4 introduces a hybrid attention design interleaving local sliding-window attention with occasional full global attention, plus configurable reasoning modes. At this size, local inference needs a high-end consumer or prosumer GPU, especially once quantized. It supports a 262,144 token context window, among the largest in the family. It carries Google's Gemma 4 license terms, published as Apache 2.0 on Hugging Face with additional Gemma-specific usage terms linked from the card. Published in March 2026, it is the largest of five Gemma 4 sizes (E2B, E4B, 12B, 26B-A4B MoE, 31B dense).