Can I Run LLM Model Locally?

Find out which AI models your machine can actually run. Check GPU compatibility, VRAM requirements, and expected performance.

Model Rankings

Top models ranked by compatibility for 24 GB VRAM

178 models · 91 excellent · 15 good

LLM models ranked by compatibility and performance
ModelVRAMGrade
Qwen3.6 27B27.8B
Q4_K_M· tok/s·262K ctx·RUNS GREAT
17.4 GBS100
Gemma 3 27B IT27.4B
Q4_K_M· tok/s·131K ctx·RUNS GREAT
18.1 GBS100
Q4_K_M· tok/s·262K ctx·RUNS GREAT
16.6 GBS100
Q4_K_M· tok/s·262K ctx·RUNS GREAT
18.7 GBS100
Q4_K_M· tok/s·262K ctx·RUNS GREAT
18.7 GBS100
Q4_K_M· tok/s·8K ctx·RUNS GREAT
18.0 GBS100
Q4_K_M· tok/s·262K ctx·RUNS GREAT
16.1 GBS100
Hy MT2 30B A3B30.1B
Q4_K_M· tok/s·262K ctx·RUNS GREAT
18.4 GBS100
Q4_K_M· tok/s·262K ctx·RUNS GREAT
18.7 GBS100
North Mini Code 1.030.5B
Q4_K_M· tok/s·500K ctx·RUNS GREAT
18.7 GBS100

Browse by VRAM

Find the best models for your VRAM tier

Popular Devices

All Hardware →

GPUs, MacBooks, AI boxes, and more — find what runs AI best

Popular Models

View all →

Qwen3.6 35B A3B

Alibaba · 36.0B · runs from 10.3 GB

6.0M 2.4K

Qwen3.6 35B A3B is a 36.0B-parameter open language model from Alibaba in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Qwen3.6 27B

Alibaba · 27.8B · runs from 8.4 GB

5.2M 2.0K

Qwen3.6 27B is a 27.8B-parameter open language model from Alibaba in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Qwen2.5 7B Instruct

Alibaba · 7.6B · runs from 2.7 GB

11.5M 1.4K

Qwen2.5 7B Instruct is a 7.6-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It supports a 128K token context window and is fine-tuned for conversational AI, instruction following, and general assistant tasks. Its efficient size makes it well-suited for local deployment on consumer GPUs with 8GB or more of VRAM. The model delivers strong performance for its parameter class across reasoning, multilingual understanding, and coding tasks. It benefits from the improved pretraining data and techniques of the Qwen 2.5 generation. Released under the Apache 2.0 license and widely supported by inference frameworks such as llama.cpp, vLLM, and Ollama.

Chat

Gemma 4 26B A4B IT

Google · 26.5B · runs from 8.0 GB

13.2M 1.3K

Gemma 4 26B A4B IT is a 26.5B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Gemma 4 31B IT

Google · 32.7B · runs from 10.6 GB

12.0M 3.3K

Gemma 4 31B IT is a 32.7B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Llama 3.2 1B Instruct

Meta · 1.2B · runs from 0.4 GB

9.8M 1.5K

Meta Llama 3.2 1B Instruct is a 1-billion parameter instruction-tuned model from Meta, the smallest in the Llama 3.2 family. It is designed for ultra-lightweight deployment scenarios where minimal hardware resources are available, supporting a 128K token context window despite its compact size. This model is suitable for basic conversational tasks, text summarization, and simple instruction following. It can run on virtually any modern GPU and even on CPU-only setups with acceptable performance. Released under the Llama 3.2 Community License.

Chat

How It Works

Three steps to find your perfect local AI setup

1

Select Your Hardware

Pick your GPU or Apple Silicon device from the dropdown.

2

Check Compatibility

See which models fit in your VRAM with performance grades.

3

Run It

Install via Ollama, LM Studio, or download the GGUF file.