skt·AXK2ForCausalLM

A.X K2 — Hardware Requirements & GPU Compatibility

Chat

A.X K2 is a 691.7B-parameter open language model from skt. It supports a context window of up to 262,144 tokens. At Q4_K_M it needs about 418.90 GB of VRAM — see which GPUs and Macs can run it below.

134.1K downloads 106 likes 160.1K quant downloads262K context

Specifications

Publisher
skt
Parameters
691.7B
Architecture
AXK2ForCausalLM
Context Length
262,144 tokens
Vocabulary Size
163,840
Release Date
2026-07-28
License
Apache 2.0

Get Started

HuggingFace

skt/A.X-K2

How Much VRAM Does A.X K2 Need?

Select a quantization to see compatible GPUs below.

QuantizationBitsVRAM
Q2_Kest.3.40297.9 GB
Q3_K_Mest.3.90341.1 GB
IQ4_XS4.30375.7 GB
Q4_K_Mest.4.80418.9 GB
Q5_K_Mest.5.70496.7 GB
Q6_Kest.6.60574.5 GB
Q8_0est.8.00695.6 GB
BF16est.16.001387.3 GB

est.= calculated VRAM estimate; no published GGUF file found for that quantization yet. Other rows are verified against real community uploads.

Which GPUs Can Run A.X K2?

Q4_K_M · 418.9 GB

A.X K2 (Q4_K_M) requires 418.9 GB of VRAM to load the model weights. For comfortable inference with headroom for KV cache and system overhead, 545+ GB is recommended. Using the full 262K context window can add up to 454.9 GB, bringing total usage to 873.8 GB. No single GPU has enough memory — multi-GPU or cluster setups are needed.

Which Devices Can Run A.X K2?

Q4_K_M · 418.9 GB

2 devices with unified memory can run A.X K2, including NVIDIA DGX H100, NVIDIA DGX A100 640GB.

Where to Download A.X K2

Community quantizations of this model — GGUF for llama.cpp, Ollama, and LM Studio, plus AWQ/MLX variants where available.

Frequently Asked Questions

How much VRAM does A.X K2 need?

A.X K2 requires 418.9 GB of VRAM at Q4_K_M, or 1387.3 GB at BF16. Full 262K context adds up to 454.9 GB (873.8 GB total).

VRAM = Weights + KV Cache + Overhead

Weights = 691.7B × 4.8 bits ÷ 8 = 415 GB

KV Cache + Overhead ≈ 3.9 GB (at 2K context + ~0.3 GB framework)

KV Cache + Overhead ≈ 458.8 GB (at full 262K context)

VRAM usage by quantization

418.9 GB
873.8 GB

Learn more about VRAM estimation →

Can NVIDIA GeForce RTX 5090 run A.X K2?

No — A.X K2 requires at least 297.9 GB at Q2_K, which exceeds the NVIDIA GeForce RTX 5090's 32 GB of VRAM.

What's the best quantization for A.X K2?

For A.X K2, Q4_K_M (418.9 GB) offers the best balance of quality and VRAM usage. Q5_K_M (496.7 GB) provides better quality if you have the VRAM. The smallest option is Q2_K at 297.9 GB.

VRAM requirement by quantization

Q2_K
297.9 GB
IQ4_XS
375.7 GB
Q4_K_M ★
418.9 GB
Q5_K_M
496.7 GB
Q6_K
574.5 GB
BF16
1387.3 GB

★ Recommended — best balance of quality and VRAM usage.

Learn more about quantization →

Can I run A.X K2 on a Mac?

A.X K2 requires at least 297.9 GB at Q2_K, which exceeds the unified memory of most consumer Macs. You would need a Mac Studio or Mac Pro with a high-memory configuration.

Can I run A.X K2 locally?

Yes — A.X K2 can run locally on consumer hardware. At Q4_K_M quantization it needs 418.9 GB of VRAM. Popular tools include Ollama, LM Studio, and llama.cpp.

What's the download size of A.X K2?

At Q4_K_M, the download is about 415.01 GB. The full-precision BF16 version is 1383.38 GB. The smallest option (Q2_K) is 293.97 GB.

Which GPUs can run A.X K2?

No single consumer GPU has enough VRAM to run A.X K2 at Q4_K_M (418.9 GB). Multi-GPU or professional hardware is required.

Which devices can run A.X K2?

3 devices with unified memory can run A.X K2 at Q4_K_M (418.9 GB), including Mac Studio (M3 Ultra, 512GB), NVIDIA DGX A100 640GB, NVIDIA DGX H100. Apple Silicon Macs use unified memory shared between CPU and GPU, making them well-suited for local LLM inference.