Run LLMs in the cloud
Live GPU prices from Vast.ai and RunPod, refreshed every hour. Rent a card by the hour, or start from a model and see the cheapest GPU it fits on.
By model size
Weights at Q4_K_M, on the cheapest NVIDIA GPU that fits with 20% VRAM headroom. Speed is for one user at a time; cost is per output token.
| Size | VRAM at Q4 | Cheapest GPU | Speed | Per 1M output | Hourly | Deploy |
|---|---|---|---|---|---|---|
| 7–8BLlama / Qwen-class | VRAM at Q4~5.3 GB | Cheapest GPURTX 3060 12GB | Speed~44 tok/s | Per 1M output tokens$0.38 | from $0.060/hr | |
| 12–14BQwen / Gemma-class | VRAM at Q4~9.2 GB | Cheapest GPURTX 3060 12GB | Speed~25 tok/s | Per 1M output tokens$0.66 | from $0.060/hr | |
| 27–32BQwen / Gemma-class | VRAM at Q4~21 GB | Cheapest GPURTX A600048 GB | Speed~24 tok/s | Per 1M output tokens$3.88 | from $0.33/hr | |
| 70BLlama / Qwen-class | VRAM at Q4~46 GB | Cheapest GPUA100 80GB | Speed~29 tok/s | Per 1M output tokens$10.22 | from $1.06/hr | |
| 100–120B MoE~12B active | VRAM at Q4~79 GB | Cheapest GPURTX PRO 600096 GB | Speed~106 tok/s | Per 1M output tokens$3.54 | from $1.35/hr | |
| 200B+235B MoE, ~22B active | VRAM at Q4~155 GB | Cheapest GPUB200192 GB | Speed~184 tok/s | Per 1M output tokens$9.02 | from $5.98/hr |
Each row uses the top of its size range. A model gets a price here when it fits on one rentable GPU (the largest right now has 288 GB); anything bigger needs several GPUs.
Popular models
The open models people run most, each on the cheapest NVIDIA GPU that fits it at Q4_K_M.
| Model | Params | VRAM | Cheapest GPU | Speed | Per 1M output | Hourly | Deploy |
|---|---|---|---|---|---|---|---|
| Qwen3 Coder 30B A3B Instruct30.5B MoE, ~19 GB | RTX 3090, ~166 tok/s, $0.28 per 1M output tokens | $0.28 | from $0.17/hr | ||||
| Qwen3.6 35B A3B36.0B MoE, ~22 GB | RTX A6000, ~125 tok/s, $0.73 per 1M output tokens | $0.73 | from $0.33/hr | ||||
| Qwen3 8B8.2B, ~5.5 GB | RTX 3060 12GB, ~42 tok/s, $0.40 per 1M output tokens | $0.40 | from $0.060/hr | ||||
| Qwen2.5 7B Instruct7.6B, ~5.0 GB | RTX 3060 12GB, ~47 tok/s, $0.36 per 1M output tokens | $0.36 | from $0.060/hr | ||||
| Qwen3.8 27B27.8B, ~17 GB | RTX 3090, ~35 tok/s, $1.33 per 1M output tokens | $1.33 | from $0.17/hr | ||||
| Gemma 4 E4B IT8.0B, ~5.3 GB | RTX 3060 12GB, ~44 tok/s, $0.38 per 1M output tokens | $0.38 | from $0.060/hr | ||||
| Qwen3 32B32.8B, ~20 GB | RTX A6000, ~25 tok/s, $3.73 per 1M output tokens | $3.73 | from $0.33/hr | ||||
| Gemma 4 31B IT31.3B, ~20 GB | RTX A6000, ~25 tok/s, $3.74 per 1M output tokens | $3.74 | from $0.33/hr | ||||
| Qwen3.5 9B9.7B, ~6.4 GB | RTX 3060 12GB, ~37 tok/s, $0.46 per 1M output tokens | $0.46 | from $0.060/hr | ||||
| Gemma 4 26B A4B IT25.8B MoE, ~16 GB | RTX 3090, ~174 tok/s, $0.27 per 1M output tokens | $0.27 | from $0.17/hr | ||||
| Qwen3 4B4.0B, ~2.9 GB | RTX 3060 12GB, ~81 tok/s, $0.21 per 1M output tokens | $0.21 | from $0.060/hr | ||||
| Qwen3 4B Instruct 25074.0B, ~2.9 GB | RTX 3060 12GB, ~81 tok/s, $0.21 per 1M output tokens | $0.21 | from $0.060/hr | ||||
| DeepSeek V4 Flash290.9B MoE, ~175 GB | B300, ~250 tok/s, $7.70 per 1M output tokens | $7.70 | from $6.94/hr | ||||
| Gemma 4 12B IT12.0B, ~8.2 GB | RTX 3060 12GB, ~28 tok/s, $0.59 per 1M output tokens | $0.59 | from $0.060/hr | ||||
| Qwen2.5 Coder 32B Instruct32.8B, ~21 GB | RTX A6000, ~24 tok/s, $3.76 per 1M output tokens | $3.76 | from $0.33/hr | ||||
| Qwen3 0.6B752M, ~0.9 GB | RTX 3060 12GB, ~269 tok/s, $0.062 per 1M output tokens | $0.062 | from $0.060/hr | ||||
| Qwen3 30B A3B Instruct 250730.5B MoE, ~19 GB | RTX 3090, ~166 tok/s, $0.28 per 1M output tokens | $0.28 | from $0.17/hr | ||||
| LFM2.5 2.6B2.7B, ~2.0 GB | RTX 3060 12GB, ~115 tok/s, $0.15 per 1M output tokens | $0.15 | from $0.060/hr | ||||
| NVIDIA Nemotron 3.5 Lightning 30B A3B BF1631.6B MoE, ~19 GB | RTX 3090, ~170 tok/s, $0.27 per 1M output tokens | $0.27 | from $0.17/hr | ||||
| Llama 3.1 8B Instruct8.0B, ~5.3 GB | RTX 3060 12GB, ~44 tok/s, $0.38 per 1M output tokens | $0.38 | from $0.060/hr | ||||
| Inkling Small266.0B MoE, ~160 GB | B300, ~399 tok/s, $4.84 per 1M output tokens | $4.84 | from $6.94/hr | ||||
| GLM 5.3 Flash321.3B MoE, ~195 GB | B300, ~217 tok/s, $8.89 per 1M output tokens | $8.89 | from $6.94/hr | ||||
| Muse Glimmer 30B29.8B, ~18 GB | RTX 3090, ~33 tok/s, $1.40 per 1M output tokens | $1.40 | from $0.17/hr | ||||
| Llama 3.2 3B Instruct3.2B, ~2.1 GB | RTX 3060 12GB, ~110 tok/s, $0.15 per 1M output tokens | $0.15 | from $0.060/hr | ||||
| Gemma 4 E2B IT5.1B, ~3.4 GB | RTX 3060 12GB, ~68 tok/s, $0.25 per 1M output tokens | $0.25 | from $0.060/hr | ||||
| Llama 3.2 1B Instruct1.2B, ~0.8 GB | RTX 3060 12GB, ~285 tok/s, $0.059 per 1M output tokens | $0.059 | from $0.060/hr | ||||
| GLM 5.2753.3B MoE, ~456 GB | Needs more than one GPU. No single rentable card fits it. | ||||||
| DeepSeek R1 0528 Qwen3 8B8.2B, ~5.5 GB | RTX 3060 12GB, ~42 tok/s, $0.40 per 1M output tokens | $0.40 | from $0.060/hr | ||||
| Qwen3.8 Flash Next180.0B MoE, ~108 GB | H200 SXM, ~69 tok/s, $14.47 per 1M output tokens | $14.47 | from $3.59/hr | ||||
| GPT OSS 20B20.9B MoE, ~13 GB | RTX A4000, ~98 tok/s, $0.25 per 1M output tokens | $0.25 | from $0.088/hr | ||||
VRAM assumes Q4_K_M and a short context. Speed and cost are for one user at a time, per output token.
Frequently Asked Questions
- What quantization should I use on a rented GPU?
Q4_K_M; every number on this page assumes it. Weights take about 4.8 bits each, so a 70B model needs roughly 46 GB including overhead instead of 140 GB at FP16, with little quality loss in chat and coding. With VRAM to spare, Q5_K_M or Q8_0 buys back a little quality for more memory and less speed. Go below Q4 only when it drops you onto a much cheaper GPU.
- How long does it take to start a model on a rented GPU?
A few minutes for a small model, and up to 15 for a 70B. The machine itself is up in a minute or two; the rest is downloading weights, and a 70B at Q4 is about 40 GB, which takes a while on a host with a slow link. Billing starts when the machine boots, not when the model is ready, so stop the instance as soon as you're done.
- Why is the cost per million tokens higher than an API?
The figures assume one user sending one request at a time, so most of the GPU sits idle. APIs batch many requests on the same card, so their cost per token is several times lower. Run vLLM with parallel requests and a rented GPU gets much closer. Renting still makes sense for private data, a fine-tuned model, or steady heavy use; for occasional chat an API is cheaper.