Inclusion AI·Ling·BailingMoeV3VLForConditionalGeneration

Ling 3.0 Flash VL — Hardware Requirements & GPU Compatibility

Vision

Ling-3.0-flash-VL is inclusionAI's native multimodal model, built on Ling-3.0-flash, with 124 billion total and 5.5 billion active parameters per token in a sparse mixture-of-experts design. It accepts image and video input and targets multimodal reasoning, long-video understanding and interface-driven agent tasks. The backbone has 42 layers alternating KDA and Gated MLA layers at a 5:1 ratio. The card reports 42 on the Artificial Analysis Intelligence Index v4.1.1, 4 points above Ling-3.0-flash. Despite the small active count, all weights must fit in memory, so it needs multi-GPU or a very large unified-memory machine even when quantized. The card's reference recipe uses four 141 GB-class GPUs. The configured context window is 131,072 tokens, while the card states support for up to 256K. It is released under the MIT license, permitting commercial use. It was published in September 2026 as the vision extension of the Ling-3.0-flash language model.

14.9K downloads 118 likes 16.5K quant downloads131K context

Specifications

Publisher
Inclusion AI
Family
Ling
Parameters
124.8B
Architecture
BailingMoeV3VLForConditionalGeneration
Context Length
131,072 tokens
Vocabulary Size
157,184
Release Date
2026-09-04
License
MIT

Get Started

Run in cloud

Fits on RTX PRO 6000 (96 GB) (19 GB headroom) · Q4_K_M

Generation speed
~185 tok/s
generation speed
Cost per 1M output tokens
$1.51
per 1M output tokens
Compare GPUs →
or

How Much VRAM Does Ling 3.0 Flash VL Need?

Select a quantization to see compatible GPUs below.

QuantizationBitsVRAM
Q2_K3.4054.2 GB
Q3_K_S3.5055.8 GB
Q3_K_M3.9062.0 GB
Q4_04.0063.6 GB
Q4_K_M4.8076.1 GB
Q5_K_M5.7090.1 GB
Q6_K6.60104.2 GB
Q8_08.00126.0 GB

Which GPUs Can Run Ling 3.0 Flash VL?

Q4_K_M · 76.1 GB

Ling 3.0 Flash VL (Q4_K_M) requires 76.1 GB of VRAM to load the model weights. For comfortable inference with headroom for KV cache and system overhead, 99+ GB is recommended. Using the full 131K context window can add up to 55.5 GB, bringing total usage to 131.6 GB. No consumer GPU has enough memory.

Rent an NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition (96 GB) from $1.00/hr.

Which Devices Can Run Ling 3.0 Flash VL?

Q4_K_M · 76.1 GB

18 devices with unified memory can run Ling 3.0 Flash VL, including NVIDIA DGX H100, NVIDIA DGX A100 640GB, NVIDIA Jetson AGX Thor Developer Kit.

Where to Download Ling 3.0 Flash VL

Community quantizations of this model — GGUF for llama.cpp, Ollama, and LM Studio, plus AWQ/MLX variants where available.

Related Models

Frequently Asked Questions

How much VRAM does Ling 3.0 Flash VL need?

Ling 3.0 Flash VL requires 76.1 GB of VRAM at Q4_K_M, or 250.9 GB at BF16. Full 131K context adds up to 55.5 GB (131.6 GB total).

VRAM = Weights + KV Cache + Overhead

Weights = 124.8B × 4.8 bits ÷ 8 = 74.9 GB

KV Cache + Overhead ≈ 1.2 GB (at 2K context + ~0.3 GB framework)

Fit ratings and hardware model lists check this model with room for a 16K-token context, which needs a little more memory.

KV Cache + Overhead ≈ 56.7 GB (at full 131K context)

VRAM usage by quantization

76.1 GB
131.6 GB

Learn more about VRAM estimation →

Can NVIDIA GeForce RTX 5090 run Ling 3.0 Flash VL?

No — Ling 3.0 Flash VL requires at least 35.5 GB at IQ2_XXS, which exceeds the NVIDIA GeForce RTX 5090's 32 GB of VRAM.

What's the best quantization for Ling 3.0 Flash VL?

For Ling 3.0 Flash VL, Q4_K_M (76.1 GB) offers the best balance of quality and VRAM usage. Q4_K_L (77.7 GB) provides better quality if you have the VRAM. The smallest option is IQ2_XXS at 35.5 GB.

VRAM requirement by quantization

IQ2_XXS
35.5 GB
Q2_K
54.2 GB
Q3_K_L
65.2 GB
Q4_K_M ★
76.1 GB
Q4_K_L
77.7 GB
BF16
250.9 GB

★ Recommended — best balance of quality and VRAM usage.

Learn more about quantization →

Can I run Ling 3.0 Flash VL on a Mac?

Ling 3.0 Flash VL requires at least 35.5 GB at IQ2_XXS, which exceeds the unified memory of most consumer Macs. You would need a Mac Studio or Mac Pro with a high-memory configuration.

Can I run Ling 3.0 Flash VL locally?

Yes — Ling 3.0 Flash VL can run locally on consumer hardware. At Q4_K_M quantization it needs 76.1 GB of VRAM. Popular tools include Ollama, LM Studio, and llama.cpp.

How fast is Ling 3.0 Flash VL?

At Q4_K_M, Ling 3.0 Flash VL can reach ~109 tok/s on AMD Instinct MI350X. Speed depends mainly on GPU memory bandwidth. Real-world results typically within ±20%.

tok/s = (bandwidth GB/s ÷ model GB) × efficiency

Example: NVIDIA B200 → 8000 ÷ 76.1 × 0.65 = ~333 tok/s

Estimated speed at Q4_K_M (76.1 GB)

~333 tok/s
~333 tok/s
~290 tok/s

Real-world results typically within ±20%. Speed depends on batch size, quantization kernel, and software stack.

Learn more about tok/s estimation →

What's the download size of Ling 3.0 Flash VL?

At Q4_K_M, the download is about 74.91 GB. The full-precision BF16 version is 249.70 GB. The smallest option (IQ2_XXS) is 34.33 GB.

Which GPUs can run Ling 3.0 Flash VL?

No single consumer GPU has enough VRAM to run Ling 3.0 Flash VL at Q4_K_M (76.1 GB). Multi-GPU or professional hardware is required.

Which devices can run Ling 3.0 Flash VL?

19 devices with unified memory can run Ling 3.0 Flash VL at Q4_K_M (76.1 GB), including ASUS Ascent GX10, Asus ROG Flow Z13 (2025, Ryzen AI Max+ 395, 128 GB), Beelink GTR9 Pro (Ryzen AI Max+ 395, 128 GB), Framework Desktop (Ryzen AI Max+ 395, 128 GB). Apple Silicon Macs use unified memory shared between CPU and GPU, making them well-suited for local LLM inference.