All LLM Models
Browse 25 LLM models with VRAM requirements, quantization options, and hardware compatibility.
Understanding LLM VRAM Requirements
How much VRAM you need depends on the model size and quantization level. Quantization reduces the precision of model weights, trading small quality losses for significantly lower VRAM usage. For example, a 7B parameter model needs ~14 GB at FP16 but only ~4 GB at Q4_K_M quantization.
Model List
Phi 4 Mini Instruct
Microsoft · 3.8B · runs from 1.9 GB
Microsoft Phi 4 Mini Instruct is a 3.8-billion parameter instruction-tuned model from Microsoft Research's Phi 4 family. It applies the Phi series' data-centric training philosophy to a compact model, delivering strong performance in coding, reasoning, and chat tasks relative to its small footprint. The model runs on consumer GPUs with as little as 4-6GB of VRAM when quantized, making it accessible on mainstream and even entry-level hardware. Released under the MIT license.
Phi 3.5 Mini Instruct
Microsoft · 3.8B · runs from 2.3 GB
Phi 3.5 Mini Instruct is a 3.8B-parameter open language model from Microsoft in the Phi 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 4
Microsoft · 14.7B · runs from 7.0 GB
Microsoft Phi 4 is a 14-billion parameter language model from Microsoft Research's Phi series, designed to deliver strong reasoning, mathematical, and coding performance at an efficient size. Phi 4 continues the Phi family's focus on maximizing capability per parameter through high-quality training data curation, achieving benchmark scores that rival much larger models on reasoning and STEM tasks. The model runs well on consumer GPUs with 12-16GB of VRAM in quantized formats. It excels at mathematical problem solving, code generation, and structured reasoning. Released under the MIT license.
Phi 4 Mini Reasoning
Microsoft · 3.8B · runs from 1.6 GB
Phi-4-mini-reasoning is Microsoft's compact 3.8-billion-parameter model in the Phi-4 family, fine-tuned specifically for multi-step, logic-intensive mathematical reasoning using synthetic math data distilled from a larger teacher model. It targets tasks like formal proof generation, symbolic computation, and advanced word problems, and is designed to run in memory- and latency-constrained environments such as educational tools or embedded tutoring systems, rather than for general knowledge or open-domain chat. It supports a 128K token context window and is released under the MIT license. At 3.8 billion parameters, it runs comfortably on almost any modern laptop or consumer GPU, even at higher precision, making it one of the more accessible reasoning-focused models to run locally.
Phi 2
Microsoft · 2.8B · runs from 2.1 GB
Microsoft Phi 2 is a 2.8-billion parameter language model from Microsoft Research that pioneered the concept of small but highly capable language models. Released in late 2023, Phi 2 demonstrated that strategic data curation and training methodology could allow a sub-3B model to outperform many 7B and 13B models on reasoning and coding benchmarks. The model runs on virtually any modern GPU and even on CPU-only setups. While succeeded by Phi 3 and Phi 4, Phi 2 remains historically significant as the model that proved small-scale language models could be genuinely useful for practical tasks. Released under the MIT license.
Phi 3.5 MoE Instruct
Microsoft · 41.9B · runs from 12.1 GB
Phi-3.5-MoE-instruct is Microsoft's Mixture-of-Experts model in the Phi-3.5 line, combining 16 experts of 3.8 billion parameters each, about 42 billion total, with only around 6.6 billion active per token. It is an instruction-tuned model refined with supervised fine-tuning, PPO, and DPO, built for multilingual reasoning tasks including code, math, and logic in memory- or latency-constrained deployments. It supports a 128K token context window and ships under the MIT license, allowing free commercial and research use. Because only about 6.6 billion parameters activate per token, it runs noticeably faster than its 42-billion-parameter size implies, though the full model still needs roughly 24 GB of memory at 4-bit quantization since all expert weights stay resident.
Phi 3 Medium 128k Instruct
Microsoft · 14.0B · runs from 4.6 GB
Phi-3-Medium-128K-Instruct is Microsoft's 14-billion-parameter instruction-tuned model in the Phi-3 family, trained on a mix of synthetic data and filtered high-quality web content chosen for reasoning density, then post-trained with supervised fine-tuning and direct preference optimization for instruction-following and safety. It is the long-context variant of Phi-3-Medium, alongside a 4K-context sibling, and is aimed at memory- and latency-constrained deployments needing strong code, math, and logical reasoning rather than frontier-scale serving. At 14B parameters, it runs on a single high-end consumer GPU once quantized. Context length is 131,072 tokens. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in May 2024.
Phi 4 Reasoning Plus
Microsoft · 14.7B · runs from 4.8 GB
Phi 4 Reasoning Plus is a 14.7B-parameter open language model from Microsoft in the Phi 4 family. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3 Mini 4k Instruct
Microsoft · 3.8B · runs from 2.7 GB
Microsoft Phi 3 Mini 4K Instruct is a 3.8-billion parameter instruction-tuned model from Microsoft Research's Phi 3 generation, with a 4K token context window. The Phi 3 family demonstrated that small models trained on carefully curated, high-quality data can achieve performance competitive with models several times their size. The model runs on consumer GPUs with as little as 4-6GB of VRAM when quantized, making it one of the most accessible capable chat models for local deployment. Released under the MIT license.
Phi 4 Reasoning
Microsoft · 14.7B · runs from 4.8 GB
Phi-4-reasoning is Microsoft's reasoning-focused fine-tune of Phi-4, a 14-billion-parameter dense decoder-only transformer, produced through supervised fine-tuning on chain-of-thought traces followed by reinforcement learning, concentrated on math, science, and coding tasks plus safety alignment data. The fine-tuning set combined synthetic prompts with high-quality filtered public web data, and responses are structured in two parts: a reasoning trace followed by a final summarized answer. Microsoft designed it for memory- and compute-constrained, latency-bound scenarios rather than frontier-scale serving, and at 14B parameters it can run on a single high-end consumer GPU once quantized. Context length is 32,768 tokens. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in April 2025.
Phi 1 5
Microsoft · 1.4B · runs from 0.7 GB
Phi 1 5 is a 1.4B-parameter open language model from Microsoft in the Phi family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3 Mini 128k Instruct
Microsoft · 3.8B · runs from 2.7 GB
Phi 3 Mini 128k Instruct is a 3.8B-parameter open language model from Microsoft in the Phi 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi Tiny MoE Instruct
Microsoft · 3.8B · runs from 2.2 GB
Phi Tiny MoE Instruct is a 3.8B-parameter open language model from Microsoft in the Phi family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3.5 Vision Instruct
Microsoft · 4.1B · runs from 2.9 GB
Phi 3.5 Vision Instruct is Microsoft's 4.1-billion-parameter multimodal model, pairing an image encoder and connector with the Phi-3 Mini language model to handle text and code alongside pictures. It suits visual question answering, chart and table reading, document OCR, and comparing details across multiple images in one prompt. Its small size suits laptops and modest consumer GPUs, running smoothly even on limited hardware once quantized. The model supports a 128K token context window, enough for lengthy documents. It is released under the MIT license, one of the most permissive options available, allowing unrestricted commercial and research use. Published in August 2024, it was trained on roughly 500 billion tokens of synthetic and filtered web data.
Phi 3 Vision 128k Instruct
Microsoft · 4.1B · runs from 9.4 GB
Phi 3 Vision 128k Instruct is a 4.1B-parameter open language model from Microsoft in the Phi 3 family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
DialoGPT Small
Microsoft · 176M · runs from 0.1 GB
DialoGPT Small is a 176M-parameter open language model from Microsoft. It supports a context window of up to 1,024 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3 Small 8k Instruct
Microsoft · 7.4B · runs from 15.3 GB
Phi-3-Small-8K-Instruct is Microsoft's 7-billion-parameter instruction-tuned chat model from the Phi-3 family, trained on a mix of synthetic data and heavily filtered public web text chosen for high reasoning density, then post-trained with supervised fine-tuning and direct preference optimization. It targets memory- and compute-constrained, latency-sensitive deployments while still emphasizing strong code, math, and logical reasoning, and Microsoft reports it holds up well against same-size and next-size-up models on common-sense, language, math, code, and long-context benchmarks. It sits alongside Mini and Medium-sized Phi-3 variants, plus a 128K-context Small sibling. At around 7 billion parameters it runs comfortably on a single consumer GPU. Context length is 8,192 tokens, as the name indicates. It is released under the MIT license, permitting unrestricted commercial and research use, and was published in May 2024.
Bitnet B1.58 2B 4T
Microsoft · 850M · runs from 2.2 GB
Bitnet B1.58 2B 4T is a 850M-parameter open language model from Microsoft. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 3 Medium 4k Instruct
Microsoft · 14.0B · runs from 6.7 GB
Phi 3 Medium 4k Instruct is a 14.0B-parameter open language model from Microsoft in the Phi 3 family. It supports a context window of up to 4,096 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 1
Microsoft · 1.4B · runs from 0.7 GB
Phi 1 is a 1.4B-parameter open language model from Microsoft in the Phi family. It supports a context window of up to 2,048 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
UserLM 8B
Microsoft · 8.0B · runs from 4.0 GB
UserLM 8B is a 8.0B-parameter open language model from Microsoft. It supports a context window of up to 8,192 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
MediPhi Instruct
Microsoft · 3.8B · runs from 2.7 GB
MediPhi Instruct is a 3.8B-parameter open language model from Microsoft in the Phi family. It supports a context window of up to 131,072 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
Phi 4 Mini Flash Reasoning
Microsoft · 3.9B · runs from 2.3 GB
Phi 4 Mini Flash Reasoning is a 3.9B-parameter open language model from Microsoft in the Phi 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
MAI DS R1
Microsoft · 671.0B · runs from 289.1 GB
MAI DS R1 is a 671.0B-parameter open language model from Microsoft. It supports a context window of up to 163,840 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.
NextCoder 7B
Microsoft · 7.6B · runs from 3.6 GB
NextCoder 7B is a 7.6B-parameter open language model from Microsoft. It supports a context window of up to 32,768 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.