Run LLMs in the cloud

Live GPU prices from Vast.ai and RunPod, refreshed every hour. Rent a card by the hour, or start from a model and see the cheapest GPU it fits on.

By model size

Weights at Q4_K_M, on the cheapest NVIDIA GPU that fits with 20% VRAM headroom. Speed is for one user at a time; cost is per output token.

SizeVRAM at Q4Cheapest GPUSpeedPer 1M outputHourlyDeploy
7–8BLlama / Qwen-classVRAM at Q4~5.3 GBCheapest GPURTX 3060 12GBSpeed~44 tok/sPer 1M output tokens$0.38from $0.060/hr
12–14BQwen / Gemma-classVRAM at Q4~9.2 GBCheapest GPURTX 3060 12GBSpeed~25 tok/sPer 1M output tokens$0.66from $0.060/hr
27–32BQwen / Gemma-classVRAM at Q4~21 GBCheapest GPURTX A600048 GBSpeed~24 tok/sPer 1M output tokens$3.88from $0.33/hr
70BLlama / Qwen-classVRAM at Q4~46 GBCheapest GPUA100 80GBSpeed~29 tok/sPer 1M output tokens$10.22from $1.06/hr
100–120B MoE~12B activeVRAM at Q4~79 GBCheapest GPURTX PRO 600096 GBSpeed~106 tok/sPer 1M output tokens$3.54from $1.35/hr
200B+235B MoE, ~22B activeVRAM at Q4~155 GBCheapest GPUB200192 GBSpeed~184 tok/sPer 1M output tokens$9.02from $5.98/hr

Each row uses the top of its size range. A model gets a price here when it fits on one rentable GPU (the largest right now has 288 GB); anything bigger needs several GPUs.

The open models people run most, each on the cheapest NVIDIA GPU that fits it at Q4_K_M.

ModelParamsVRAMCheapest GPUSpeedPer 1M outputHourlyDeploy
Qwen3 Coder 30B A3B Instruct30.5B MoE, ~19 GBRTX 3090, ~166 tok/s, $0.28 per 1M output tokensfrom $0.17/hr
Qwen3.6 35B A3B36.0B MoE, ~22 GBRTX A6000, ~125 tok/s, $0.73 per 1M output tokensfrom $0.33/hr
Qwen3 8B8.2B, ~5.5 GBRTX 3060 12GB, ~42 tok/s, $0.40 per 1M output tokensfrom $0.060/hr
Qwen2.5 7B Instruct7.6B, ~5.0 GBRTX 3060 12GB, ~47 tok/s, $0.36 per 1M output tokensfrom $0.060/hr
Qwen3.8 27B27.8B, ~17 GBRTX 3090, ~35 tok/s, $1.33 per 1M output tokensfrom $0.17/hr
Gemma 4 E4B IT8.0B, ~5.3 GBRTX 3060 12GB, ~44 tok/s, $0.38 per 1M output tokensfrom $0.060/hr
Qwen3 32B32.8B, ~20 GBRTX A6000, ~25 tok/s, $3.73 per 1M output tokensfrom $0.33/hr
Gemma 4 31B IT31.3B, ~20 GBRTX A6000, ~25 tok/s, $3.74 per 1M output tokensfrom $0.33/hr
Qwen3.5 9B9.7B, ~6.4 GBRTX 3060 12GB, ~37 tok/s, $0.46 per 1M output tokensfrom $0.060/hr
Gemma 4 26B A4B IT25.8B MoE, ~16 GBRTX 3090, ~174 tok/s, $0.27 per 1M output tokensfrom $0.17/hr
Qwen3 4B4.0B, ~2.9 GBRTX 3060 12GB, ~81 tok/s, $0.21 per 1M output tokensfrom $0.060/hr
Qwen3 4B Instruct 25074.0B, ~2.9 GBRTX 3060 12GB, ~81 tok/s, $0.21 per 1M output tokensfrom $0.060/hr
DeepSeek V4 Flash290.9B MoE, ~175 GBB300, ~250 tok/s, $7.70 per 1M output tokensfrom $6.94/hr
Gemma 4 12B IT12.0B, ~8.2 GBRTX 3060 12GB, ~28 tok/s, $0.59 per 1M output tokensfrom $0.060/hr
Qwen2.5 Coder 32B Instruct32.8B, ~21 GBRTX A6000, ~24 tok/s, $3.76 per 1M output tokensfrom $0.33/hr
Qwen3 0.6B752M, ~0.9 GBRTX 3060 12GB, ~269 tok/s, $0.062 per 1M output tokensfrom $0.060/hr
Qwen3 30B A3B Instruct 250730.5B MoE, ~19 GBRTX 3090, ~166 tok/s, $0.28 per 1M output tokensfrom $0.17/hr
LFM2.5 2.6B2.7B, ~2.0 GBRTX 3060 12GB, ~115 tok/s, $0.15 per 1M output tokensfrom $0.060/hr
NVIDIA Nemotron 3.5 Lightning 30B A3B BF1631.6B MoE, ~19 GBRTX 3090, ~170 tok/s, $0.27 per 1M output tokensfrom $0.17/hr
Llama 3.1 8B Instruct8.0B, ~5.3 GBRTX 3060 12GB, ~44 tok/s, $0.38 per 1M output tokensfrom $0.060/hr
Inkling Small266.0B MoE, ~160 GBB300, ~399 tok/s, $4.84 per 1M output tokensfrom $6.94/hr
GLM 5.3 Flash321.3B MoE, ~195 GBB300, ~217 tok/s, $8.89 per 1M output tokensfrom $6.94/hr
Muse Glimmer 30B29.8B, ~18 GBRTX 3090, ~33 tok/s, $1.40 per 1M output tokensfrom $0.17/hr
Llama 3.2 3B Instruct3.2B, ~2.1 GBRTX 3060 12GB, ~110 tok/s, $0.15 per 1M output tokensfrom $0.060/hr
Gemma 4 E2B IT5.1B, ~3.4 GBRTX 3060 12GB, ~68 tok/s, $0.25 per 1M output tokensfrom $0.060/hr
Llama 3.2 1B Instruct1.2B, ~0.8 GBRTX 3060 12GB, ~285 tok/s, $0.059 per 1M output tokensfrom $0.060/hr
GLM 5.2753.3B MoE, ~456 GBNeeds more than one GPU. No single rentable card fits it.
DeepSeek R1 0528 Qwen3 8B8.2B, ~5.5 GBRTX 3060 12GB, ~42 tok/s, $0.40 per 1M output tokensfrom $0.060/hr
Qwen3.8 Flash Next180.0B MoE, ~108 GBH200 SXM, ~69 tok/s, $14.47 per 1M output tokensfrom $3.59/hr
GPT OSS 20B20.9B MoE, ~13 GBRTX A4000, ~98 tok/s, $0.25 per 1M output tokensfrom $0.088/hr

VRAM assumes Q4_K_M and a short context. Speed and cost are for one user at a time, per output token.

Frequently Asked Questions

What quantization should I use on a rented GPU?

Q4_K_M; every number on this page assumes it. Weights take about 4.8 bits each, so a 70B model needs roughly 46 GB including overhead instead of 140 GB at FP16, with little quality loss in chat and coding. With VRAM to spare, Q5_K_M or Q8_0 buys back a little quality for more memory and less speed. Go below Q4 only when it drops you onto a much cheaper GPU.

How long does it take to start a model on a rented GPU?

A few minutes for a small model, and up to 15 for a 70B. The machine itself is up in a minute or two; the rest is downloading weights, and a 70B at Q4 is about 40 GB, which takes a while on a host with a slow link. Billing starts when the machine boots, not when the model is ready, so stop the instance as soon as you're done.

Why is the cost per million tokens higher than an API?

The figures assume one user sending one request at a time, so most of the GPU sits idle. APIs batch many requests on the same card, so their cost per token is several times lower. Run vLLM with parallel requests and a rented GPU gets much closer. Renting still makes sense for private data, a fine-tuned model, or steady heavy use; for occasional chat an API is cheaper.