All LLM Models
Browse 29 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Mistral Small 24B Instruct 2501
Mistral AI · 23.6B · runs from 10.7 GB
Mistral Small 24B Instruct is Mistral AI's January 2025 release targeting the mid-range parameter sweet spot. At 24 billion parameters it sits between lightweight 7B models and heavier 70B-class offerings, delivering strong instruction-following, reasoning, and coding performance without demanding top-tier hardware. This model fits comfortably on a single GPU with 16–24 GB of VRAM at common quantization levels, making it an attractive option for users with cards like the RTX 4090 or RTX 3090 who want a noticeable step up from 7B models. It strikes an appealing balance between quality and resource requirements for serious local use.
Mistral 7B v0.3
Mistral AI · 7.2B · runs from 3.6 GB
Mistral 7B v0.3 is a 7.2B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mistral 7B Instruct v0.3
Mistral AI · 7.2B · runs from 2.7 GB
Mistral 7B Instruct v0.3 is the latest instruction-tuned release of Mistral AI's original 7-billion-parameter model, delivering meaningful improvements in instruction following, function calling, and multilingual support over its predecessors. With an extended 32K-token vocabulary and refined chat capabilities, v0.3 remains one of the most capable sub-10B models available. At 7.2 billion parameters it sits comfortably in the sweet spot for local inference, running well on GPUs with 6–8 GB of VRAM at full precision and even on 4 GB cards with 4-bit quantization. It is an excellent default choice for anyone getting started with local LLMs who wants strong conversational performance without heavy hardware.
Mistral Nemo Instruct 2407
Mistral AI · 12.2B · runs from 5.9 GB
Mistral Nemo Instruct 2407 is a 12-billion-parameter instruction-tuned chat model from Mistral AI, a dated release from July 2024. It targets general dialogue and instruction-following use cases, sitting in a practical middle ground between lightweight and large-scale models in terms of both capability and resource demands. The model offers a 128K token context window, generous for its parameter class, and is released under the Apache 2.0 license, allowing unrestricted local and commercial use. At around 12 billion parameters, Mistral Nemo Instruct 2407 fits on a single consumer GPU when quantized, putting it within easy reach of local enthusiasts running their own hardware.
Mistral Small Instruct 2409
Mistral AI · 22.2B · runs from 7.4 GB
Mistral Small Instruct 2409 is a 22.2B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Ministral 3 3B Reasoning 2512
Mistral AI · 4.3B · runs from 2.3 GB
Ministral 3 3B Reasoning 2512 is a 4.3B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Devstral Small 2 24B Instruct 2512
Mistral AI · 24.0B · runs from 7.3 GB
Devstral Small 2 24B Instruct is Mistral AI's dense 24-billion-parameter model for agentic software-engineering work, fine-tuned to follow instructions for chat, coding agents, and tool-heavy workflows. Built on the same architecture as Ministral 3, it adds vision capabilities for analyzing images alongside code and text, and its publisher designed it specifically to be lightweight enough for local, on-device use rather than requiring a large server. It supports a context window of roughly 384,000 tokens and is released under the Apache 2.0 license. Mistral notes it is light enough to run on a single RTX 4090 or a Mac with 32GB of RAM, consistent with its 4-bit memory needs of around 14GB.
Mistral Small 3.2 24B Instruct 2506
Mistral AI · 24.0B · runs from 7.3 GB
Mistral-Small-3.2-24B-Instruct-2506 is a dense 24-billion-parameter vision-language model from Mistral AI, a minor refinement of Mistral-Small-3.1-24B-Instruct-2503. The update focuses on following precise instructions more reliably, cutting down on repetitive or runaway generations, and making function calling more robust, while still handling both text and image inputs for tasks like chart and document understanding. It supports a 131,072-token context window and is released under the Apache 2.0 license. At this size, 4-bit quantization needs roughly 14GB of memory, so it runs on a single consumer GPU in the 16-24GB range, such as an RTX 4090.
Ministral 3 14B Reasoning 2512
Mistral AI · 13.9B · runs from 6.7 GB
Ministral 3 14B Reasoning 2512 is a 13.9B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Mistral 7B Instruct v0.2
Mistral AI · 7.2B · runs from 3.6 GB
Mistral 7B Instruct v0.2 is a 7.2-billion-parameter instruction-tuned language model from Mistral AI, built for general chat, question answering, and instruction following. It refines the original Mistral 7B with better adherence to complex prompts and improved handling of longer inputs. Its compact size suits local deployment on modern consumer GPUs, running comfortably on mainstream hardware once quantized. The model supports a 32,768 token context window, a result of raising the RoPE base frequency used for positional encoding, which improved handling of longer inputs versus the original v0.1. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in December 2023, it became one of the most widely adopted 7B open-weight chat models, still supported by tools such as llama.cpp, vLLM, and Ollama.
Mistral 7B Instruct v0.1
Mistral AI · 7.2B · runs from 3.6 GB
Mistral 7B Instruct v0.1 was the first instruction-tuned variant of the original Mistral 7B, fine-tuned for conversational and instruction-following tasks. While it has since been superseded by v0.2 and v0.3, it remains a solid lightweight chat model and an important milestone in the open-weight model ecosystem. Its hardware requirements are identical to the base Mistral 7B, running smoothly on GPUs with as little as 6 GB of VRAM when quantized. Users seeking the best Mistral 7B experience should generally prefer the newer v0.3 release, but v0.1 is still useful for reproducibility and benchmarking purposes.
Mixtral 8x7B Instruct v0.1
Mistral AI · 46.7B · runs from 20.4 GB
Mixtral 8x7B Instruct v0.1 is Mistral AI's flagship Mixture-of-Experts model, combining eight expert networks of 7 billion parameters each for a 46.7B total weight count while activating only about 12.9 billion parameters per token. This sparse architecture delivers performance that rivals much larger dense models at a fraction of the inference cost, excelling across reasoning, code generation, and multilingual tasks. Because the full weights must still be loaded into memory, you will need around 24–48 GB of VRAM depending on quantization level, making it best suited for multi-GPU desktop setups or high-VRAM workstation cards. If your hardware can accommodate it, Mixtral offers one of the best performance-per-active-parameter ratios available for local deployment.
Mistral Small 3.1 24B Instruct 2503
Mistral AI · 24.0B · runs from 7.3 GB
Mistral Small 3.1 24B Instruct 2503 is a 24-billion-parameter model from Mistral AI, the French AI lab, built on the earlier text-only Mistral Small 3 with added support for image input alongside text. It can reason about images in the same conversation as written prompts, useful for document understanding and multimodal chat. At 24 billion parameters, it needs quantization and a single high-end 24GB-class consumer or workstation GPU for local inference rather than budget hardware. The model supports a 128K token context window for long documents or extended conversations. It is released under the Apache 2.0 license, allowing unrestricted commercial and research use. Published in March 2025, it added vision understanding and a longer context window to the earlier text-only Mistral Small while keeping the same 24B parameter budget.
Mistral Medium 3.5 128B
Mistral AI · 127.7B · runs from 55.3 GB
Mistral Medium 3.5 128B is a 127.7B-parameter open language model from Mistral AI in the Mistral family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Devstral Small 2507
Mistral AI · 23.6B · runs from 7.2 GB
Devstral Small 2507 is Mistral AI's agentic coding model, developed with All Hands AI and fine-tuned from the 24-billion-parameter Mistral Small 3.1 with its vision encoder removed to keep it text-only. It is built to explore codebases, edit multiple files, and drive software-engineering agents, using Mistral's function-calling format and a Tekken tokenizer with a 131K-token vocabulary. It supports a 128K token context window and is released under the Apache 2.0 license. At 24 billion parameters, Devstral is light enough to run on a single RTX 4090 or a Mac with around 32 GB of unified memory once quantized to 4-bit, making it practical for local coding-agent setups.
Magistral Small 2509
Mistral AI · 24.0B · runs from 7.3 GB
Magistral Small 2509 (also called Magistral Small 1.2) is Mistral AI's small reasoning model, built on Mistral Small 3.2 24B Instruct with added chain-of-thought reasoning trained through supervised fine-tuning on Magistral Medium traces followed by reinforcement learning. Unlike the text-only Magistral Small 1.1, this 1.2 release adds a vision encoder, so it can reason over images as well as text, wrapping its reasoning trace in dedicated [THINK]/[/THINK] tokens. It supports dozens of languages and is small enough to fit on a single consumer GPU once quantized. Context length is 131,072 tokens, though the card notes performance may degrade somewhat past 40,000 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in September 2025.
Devstral 2 123B Instruct 2512
Mistral AI · 125.0B · runs from 35.4 GB
Devstral 2 123B Instruct 2512 is Mistral AI's large agentic coding model for software engineering tasks such as exploring codebases, editing multiple files, and running autonomous coding agents. Released as an FP8-quantized checkpoint of a roughly 125-billion-parameter dense model, it is a direct step up from the smaller Devstral Small line and pairs with Mistral's own Vibe CLI as well as third-party scaffolds like OpenHands and Claude Code. It scores 72.2% on SWE-Bench Verified, 61.3% on SWE-Bench Multilingual, and 32.6% on Terminal-Bench 2, competitive with or ahead of several much larger open models. At this size it needs a multi-GPU workstation even when quantized. Context length is 262,144 tokens. It is released under a Modified MIT License that blocks companies with over $20 million in monthly consolidated revenue from using it without a separate commercial license from Mistral AI. It was published in November 2025.
Mixtral 8x7B v0.1
Mistral AI · 46.7B · runs from 19.8 GB
Mixtral-8x7B-v0.1 is Mistral AI's pretrained Sparse Mixture-of-Experts model, combining eight 7-billion-parameter experts for roughly 46.7 billion total parameters. As a base checkpoint, it has no built-in moderation or instruction-following behavior and is intended as a foundation for further fine-tuning rather than direct chat use. It supports a 32K token context window and is released under the Apache 2.0 license, allowing unrestricted commercial and research use. At around 47 billion total parameters, running it locally at 4-bit quantization needs roughly 24-32 GB of memory, putting it within reach of a single high-end consumer GPU or a dual-GPU setup.
Mistral 7B v0.1
Mistral AI · 7.2B · runs from 3.6 GB
Mistral 7B v0.1 is the original base model from Mistral AI that helped reshape expectations for small open-weight language models when it launched in late 2023. As a pretrained foundation model without instruction tuning, it is designed for fine-tuning, research, and custom downstream tasks rather than direct conversational use. With 7 billion parameters and support for grouped-query attention and sliding-window attention, it remains a popular starting point for practitioners building specialized models. Its modest VRAM requirements of roughly 6 GB at 4-bit quantization keep it accessible on a wide range of consumer GPUs.
Mistral Large Instruct 2407
Mistral AI · 122.6B · runs from 37.1 GB
Mistral-Large-Instruct-2407, also known as Mistral Large 2, is Mistral AI's flagship dense instruction-tuned model of about 123 billion parameters, designed to run efficient single-node inference despite its size. It targets state-of-the-art reasoning, coding, and knowledge tasks with native function calling and JSON output for agentic use, and covers a dozen or more natural languages including French, German, Spanish, Chinese, Japanese, Russian, and Korean alongside more than 80 programming languages. Mistral reports reduced hallucination rates and stronger reasoning compared with the original Mistral Large. Even tuned for single-node deployment, a dense model of this size needs a multi-GPU workstation to run. Context length is 131,072 tokens (a 128k window). It is released under the Mistral AI Research License, a custom license restricting use to non-commercial research purposes; commercial deployment requires a separate license from Mistral AI. It was published in July 2024.
Devstral Small 2505
Mistral AI · 23.6B · runs from 7.2 GB
Devstral Small 2505, internally called Devstral Small 1.0, is an agentic coding model built by Mistral AI in collaboration with All Hands AI, finetuned from Mistral-Small-3.1-24B-Base-2503 with its vision encoder removed, so it is text-only. It is designed to explore codebases and edit multiple files as a software engineering agent, and reached 46.8% on SWE-Bench Verified under the OpenHands scaffold, the top score among open models at release, ahead of GPT-4.1-mini and Claude 3.5 Haiku on the same benchmark. At 24 billion parameters it is explicitly built to be lightweight enough for a single consumer GPU or an Apple Silicon Mac, making it a real local-deployment option rather than a server-only model. Context length is 131,072 tokens (a 128k window). It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in May 2025; a larger 123-billion-parameter Devstral 2 later succeeded it.
Mistral Large 3 675B Instruct 2512
Mistral AI · 675B · runs from 204.2 GB
Mistral-Large-3-675B-Instruct-2512 is Mistral AI's flagship instruction-tuned model, a multimodal granular Mixture-of-Experts system pairing a roughly 673-billion-parameter, about 39-billion-active-parameter language backbone with a 2.5-billion-parameter vision encoder, for around 675 billion total and 41 billion active parameters overall. It handles vision alongside text, supports dozens of languages, and is built for agentic use with native function calling and JSON output, aimed at long-document understanding, coding, and enterprise knowledge work rather than dedicated step-by-step reasoning. This is a frontier-scale model that requires a full multi-GPU server node to run, even quantized. Context length is 262,144 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in November 2025, as the third generation of Mistral's Large model line, alongside FP8, NVFP4, and BF16 weight releases.
Magistral Small 2506
Mistral AI · 23.6B · runs from 7.2 GB
Magistral Small 2506 is Mistral AI's small, efficient reasoning model, built on the 24-billion-parameter Mistral Small 3.1 and further trained with supervised fine-tuning on Magistral Medium's reasoning traces plus reinforcement learning. It produces long chains of reasoning before answering and supports dozens of languages, making it suited to multi-step problem solving rather than simple chat. The architecture supports up to 128K tokens, but Mistral recommends keeping usage around 40K since output quality can degrade past that point. Released under the Apache 2.0 license, it is compact enough to fit on a single RTX 4090 or a 32 GB RAM Mac once quantized to 4-bit, in line with what a 24-billion-parameter model needs at that precision.
Ministral 8B Instruct 2410
Mistral AI · 8.0B · runs from 3.3 GB
Ministral 8B Instruct 2410 is Mistral AI's 8-billion-parameter instruction-tuned model, part of the "Ministraux" family aimed at edge and on-device deployment. It uses a 36-layer dense transformer with interleaved sliding-window attention, trained heavily on multilingual and code data, suiting chat, function calling, and general assistant tasks across ten languages. Its compact size means it runs comfortably on a single mainstream consumer GPU once quantized. The model supports a 32K token context window. It is released under Mistral's own license, listed as "Other," which restricts use to research purposes and requires a separate commercial license from Mistral AI, so review the terms before commercial use. It was published in October 2024 alongside the smaller Ministral 3B.
Mixtral 8x22B Instruct v0.1
Mistral AI · 140.6B · runs from 58.8 GB
Mixtral-8x22B-Instruct-v0.1 is Mistral AI's instruction-tuned chat model, fine-tuned from the Mixtral-8x22B-v0.1 base model. It is a sparse Mixture-of-Experts model with 8 experts per layer and 2 active per token, giving roughly 39.2 billion active parameters out of about 140.6 billion total, and supports function calling for agentic and tool-use workflows. Because all experts must stay in memory even though only two run per token, it needs a multi-GPU workstation or a high-memory machine even once quantized. Context length is 65,536 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in April 2024, as the instruct variant of the larger successor to Mixtral 8x7B.
Leanstral 1.5 119B A6B
Mistral AI · 119B · runs from 55.6 GB
Leanstral 1.5 119B A6B is Mistral AI's open-source code agent model specialized for Lean 4, the proof assistant used to formalize complex mathematics and software specifications. Built as part of the Mistral Small 4 family, it is a mixture-of-experts model with 128 experts and 4 active per token, totaling 119 billion parameters with about 6.5 billion active per token, and it accepts both text and image input while producing text output. It is designed to work through long formal-proof tasks over many hours via the Mistral Vibe CLI and the lean-lsp-mcp tool, with a "high" reasoning-effort mode for complex proofs. Because all 128 experts must be held in memory regardless of how few activate per token, it still needs a multi-GPU setup to run locally. Context length is 256,000 tokens, though Mistral recommends staying under 200,000 tokens in practice. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in July 2026 as an update to the original Leanstral model.
Mistral Nemo Base 2407
Mistral AI · 12.2B · runs from 5.8 GB
Mistral-Nemo-Base-2407 is a 12-billion-parameter pretrained base language model jointly developed by Mistral AI and NVIDIA, intended as a drop-in replacement for the earlier Mistral 7B rather than a chat assistant; it has not been instruction-tuned or aligned, and the card notes it carries no built-in moderation mechanisms. It uses a standard transformer architecture with grouped-query attention and SwiGLU activations, and was trained on a large share of multilingual and code data across nine languages. At 12 billion parameters it fits comfortably on a single consumer GPU once quantized. Context length is 131,072 tokens (a 128k window). It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in July 2024, alongside an instruction-tuned Nemo variant.
Mixtral 8x22B v0.1
Mistral AI · 140.6B · runs from 60.5 GB
Mixtral-8x22B-v0.1 is Mistral AI's large sparse mixture-of-experts base model, a pretrained checkpoint with no instruction tuning and no built-in moderation, intended as the foundation for fine-tuned or instruct derivatives rather than direct chat use. It routes each token through 2 of 8 experts, giving roughly 39 billion active parameters out of about 141 billion total, so its per-token compute is far lighter than its total size at the cost of having to hold every expert in memory. Its pretraining data covers English, French, German, Spanish, and Italian. Because all experts must be resident even though only a fraction activate per token, it still needs a multi-GPU workstation even when quantized. Context length is 65,536 tokens. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use, and was published in April 2024, ahead of an instruction-tuned Mixtral-8x22B-Instruct release.
Mathstral 7B v0.1
Mistral AI · 7.2B · runs from 3.6 GB
Mathstral 7B v0.1 is a 7.2B-parameter open language model from Mistral AI. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.