Ling 3.0 Flash VL — Hardware Requirements & GPU Compatibility
VisionLing-3.0-flash-VL is inclusionAI's native multimodal model, built on Ling-3.0-flash, with 124 billion total and 5.5 billion active parameters per token in a sparse mixture-of-experts design. It accepts image and video input and targets multimodal reasoning, long-video understanding and interface-driven agent tasks. The backbone has 42 layers alternating KDA and Gated MLA layers at a 5:1 ratio. The card reports 42 on the Artificial Analysis Intelligence Index v4.1.1, 4 points above Ling-3.0-flash. Despite the small active count, all weights must fit in memory, so it needs multi-GPU or a very large unified-memory machine even when quantized. The card's reference recipe uses four 141 GB-class GPUs. The configured context window is 131,072 tokens, while the card states support for up to 256K. It is released under the MIT license, permitting commercial use. It was published in September 2026 as the vision extension of the Ling-3.0-flash language model.
Specifications
- Publisher
- Inclusion AI
- Family
- Ling
- Parameters
- 124.8B
- Architecture
- BailingMoeV3VLForConditionalGeneration
- Context Length
- 131,072 tokens
- Vocabulary Size
- 157,184
- Release Date
- 2026-09-04
- License
- MIT
Get Started
HuggingFace
Run in cloud
Fits on RTX PRO 6000 (96 GB) (19 GB headroom) · Q4_K_M
- Generation speed
- ~185 tok/s
- generation speed
- Cost per 1M output tokens
- $1.51
- per 1M output tokens
How Much VRAM Does Ling 3.0 Flash VL Need?
Select a quantization to see compatible GPUs below.
| Quantization | Bits | VRAM | + Context | File Size | Quality |
|---|---|---|---|---|---|
| Q2_K | 3.40 | 54.2 GB | 109.7 GB | 53.06 GB | 2-bit quantization with K-quant improvements |
| Q3_K_S | 3.50 | 55.8 GB | 111.3 GB | 54.62 GB | 3-bit small quantization |
| Q3_K_M | 3.90 | 62.0 GB | 117.5 GB | 60.86 GB | 3-bit medium quantization |
| Q4_0 | 4.00 | 63.6 GB | 119.1 GB | 62.42 GB | 4-bit legacy quantization |
| Q4_K_M | 4.80 | 76.1 GB | 131.6 GB | 74.91 GB | 4-bit medium quantization — most popular sweet spot |
| Q5_K_M | 5.70 | 90.1 GB | 145.6 GB | 88.95 GB | 5-bit medium quantization — good quality/size tradeoff |
| Q6_K | 6.60 | 104.2 GB | 159.7 GB | 103.00 GB | 6-bit quantization, very good quality |
| Q8_0 | 8.00 | 126.0 GB | 181.5 GB | 124.85 GB | 8-bit quantization, near-lossless |
Which GPUs Can Run Ling 3.0 Flash VL?
Q4_K_M · 76.1 GBLing 3.0 Flash VL (Q4_K_M) requires 76.1 GB of VRAM to load the model weights. For comfortable inference with headroom for KV cache and system overhead, 99+ GB is recommended. Using the full 131K context window can add up to 55.5 GB, bringing total usage to 131.6 GB. No consumer GPU has enough memory.
Rent an NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition (96 GB) from $1.00/hr.
Which Devices Can Run Ling 3.0 Flash VL?
Q4_K_M · 76.1 GB18 devices with unified memory can run Ling 3.0 Flash VL, including NVIDIA DGX H100, NVIDIA DGX A100 640GB, NVIDIA Jetson AGX Thor Developer Kit.
Runs great
— Plenty of headroomDecent
— Enough memory, may be tightWhere to Download Ling 3.0 Flash VL
Community quantizations of this model — GGUF for llama.cpp, Ollama, and LM Studio, plus AWQ/MLX variants where available.
Related Models
Frequently Asked Questions
- How much VRAM does Ling 3.0 Flash VL need?
Ling 3.0 Flash VL requires 76.1 GB of VRAM at Q4_K_M, or 250.9 GB at BF16. Full 131K context adds up to 55.5 GB (131.6 GB total).
VRAM = Weights + KV Cache + Overhead
Weights = 124.8B × 4.8 bits ÷ 8 = 74.9 GB
KV Cache + Overhead ≈ 1.2 GB (at 2K context + ~0.3 GB framework)
Fit ratings and hardware model lists check this model with room for a 16K-token context, which needs a little more memory.
KV Cache + Overhead ≈ 56.7 GB (at full 131K context)
VRAM usage by quantization
Q4_K_M76.1 GBQ4_K_M + full context131.6 GB- Can NVIDIA GeForce RTX 5090 run Ling 3.0 Flash VL?
No — Ling 3.0 Flash VL requires at least 35.5 GB at IQ2_XXS, which exceeds the NVIDIA GeForce RTX 5090's 32 GB of VRAM.
- What's the best quantization for Ling 3.0 Flash VL?
For Ling 3.0 Flash VL, Q4_K_M (76.1 GB) offers the best balance of quality and VRAM usage. Q4_K_L (77.7 GB) provides better quality if you have the VRAM. The smallest option is IQ2_XXS at 35.5 GB.
VRAM requirement by quantization
IQ2_XXS35.5 GBQ2_K54.2 GBQ3_K_L65.2 GBQ4_K_M ★76.1 GBQ4_K_L77.7 GBBF16250.9 GB★ Recommended — best balance of quality and VRAM usage.
- Can I run Ling 3.0 Flash VL on a Mac?
Ling 3.0 Flash VL requires at least 35.5 GB at IQ2_XXS, which exceeds the unified memory of most consumer Macs. You would need a Mac Studio or Mac Pro with a high-memory configuration.
- Can I run Ling 3.0 Flash VL locally?
Yes — Ling 3.0 Flash VL can run locally on consumer hardware. At Q4_K_M quantization it needs 76.1 GB of VRAM. Popular tools include Ollama, LM Studio, and llama.cpp.
- How fast is Ling 3.0 Flash VL?
At Q4_K_M, Ling 3.0 Flash VL can reach ~109 tok/s on AMD Instinct MI350X. Speed depends mainly on GPU memory bandwidth. Real-world results typically within ±20%.
tok/s = (bandwidth GB/s ÷ model GB) × efficiency
Example: NVIDIA B200 → 8000 ÷ 76.1 × 0.65 = ~333 tok/s
Estimated speed at Q4_K_M (76.1 GB)
~333 tok/s~333 tok/s~290 tok/sReal-world results typically within ±20%. Speed depends on batch size, quantization kernel, and software stack.
- What's the download size of Ling 3.0 Flash VL?
At Q4_K_M, the download is about 74.91 GB. The full-precision BF16 version is 249.70 GB. The smallest option (IQ2_XXS) is 34.33 GB.
- Which GPUs can run Ling 3.0 Flash VL?
No single consumer GPU has enough VRAM to run Ling 3.0 Flash VL at Q4_K_M (76.1 GB). Multi-GPU or professional hardware is required.
- Which devices can run Ling 3.0 Flash VL?
19 devices with unified memory can run Ling 3.0 Flash VL at Q4_K_M (76.1 GB), including ASUS Ascent GX10, Asus ROG Flow Z13 (2025, Ryzen AI Max+ 395, 128 GB), Beelink GTR9 Pro (Ryzen AI Max+ 395, 128 GB), Framework Desktop (Ryzen AI Max+ 395, 128 GB). Apple Silicon Macs use unified memory shared between CPU and GPU, making them well-suited for local LLM inference.