OpenGVLab·InternVLChatModel

InternVL3 1B — Hardware Requirements & GPU Compatibility

Vision

InternVL3-1B is OpenGVLab's smallest vision-language model in the InternVL3 series, combining an InternViT-300M-448px-V2.5 vision encoder with a Qwen2.5-0.5B language backbone through an MLP projector. Unlike earlier InternVL generations, it uses Native Multimodal Pre-Training, interleaving image-text and video-text data with text-only corpora in a single pretraining stage rather than adapting a finished text model, and adds Variable Visual Position Encoding for better long-context handling. The series extends its capabilities to tool use, GUI agents, industrial image analysis, and 3D vision perception. At under 1 billion parameters, it is small enough to run on a modest consumer GPU or even a CPU. It is released under the Apache 2.0 license, permitting unrestricted commercial and research use. It was published in April 2025, as the smallest member of a family that scales up to InternVL3-78B.

120.6K downloads 84 likes 2.3K quant downloads

Specifications

Publisher
OpenGVLab
Parameters
938M
Architecture
InternVLChatModel
Release Date
2025-04-10
License
Apache 2.0

Get Started

How Much VRAM Does InternVL3 1B Need?

Select a quantization to see compatible GPUs below.

QuantizationBitsVRAM
Q2_K3.400.4 GB
Q3_K_S3.500.5 GB
Q3_K_M3.900.5 GB
Q4_04.000.5 GB
Q4_K_M4.800.6 GB
Q5_K_M5.700.7 GB
Q6_K6.600.8 GB
Q8_08.001.0 GB

Which GPUs Can Run InternVL3 1B?

Q4_K_M · 0.6 GB

InternVL3 1B (Q4_K_M) requires 0.6 GB of VRAM to load the model weights. For comfortable inference with headroom for KV cache and system overhead, 1+ GB is recommended. 52 GPUs can run it, including NVIDIA GeForce RTX 5090, NVIDIA GeForce RTX 3090 Ti.

Runs great

— Plenty of headroom
NVIDIA GeForce RTX 5090~1879 tok/sNVIDIA GeForce RTX 3090 Ti~1057 tok/sNVIDIA GeForce RTX 4090~1057 tok/sNVIDIA GeForce RTX 5080~1007 tok/sNVIDIA GeForce RTX 3090~982 tok/sNVIDIA GeForce RTX 3080 Ti~957 tok/sNVIDIA GeForce RTX 5070 Ti~939 tok/sNVIDIA GeForce RTX 5090 Laptop GPU~939 tok/sAMD Radeon RX 7900 XTX~929 tok/sNVIDIA GeForce RTX 3080~797 tok/sAMD Radeon RX 7900 XT~774 tok/sNVIDIA GeForce RTX 4080 SUPER~772 tok/sNVIDIA GeForce RTX 4080~752 tok/sNVIDIA GeForce RTX 4070 Ti SUPER~705 tok/sNVIDIA GeForce RTX 5070~705 tok/sNVIDIA TITAN RTX~705 tok/sNVIDIA GeForce RTX 2080 Ti~646 tok/sNVIDIA GeForce RTX 3070 Ti~638 tok/sAMD Radeon RX 9070~619 tok/sAMD Radeon RX 9070 XT~619 tok/sAMD Radeon RX 7800 XT~604 tok/sNVIDIA GeForce RTX 4090 Laptop GPU~604 tok/sAMD Radeon RX 7900 GRE~557 tok/sNVIDIA GeForce RTX 4070~528 tok/sNVIDIA GeForce RTX 4070 SUPER~528 tok/sNVIDIA GeForce RTX 4070 Ti~528 tok/sNVIDIA GeForce GTX 1080 Ti~508 tok/sAMD Radeon RX 6800~496 tok/sAMD Radeon RX 6800 XT~496 tok/sAMD Radeon RX 6900 XT~496 tok/sNVIDIA GeForce RTX 3060 Ti~470 tok/sNVIDIA GeForce RTX 3070~470 tok/sNVIDIA GeForce RTX 5060~470 tok/sNVIDIA GeForce RTX 5060 Ti 16GB~470 tok/sNVIDIA GeForce RTX 5060 Ti 8GB~470 tok/sIntel Arc A770 16GB~452 tok/sAMD Radeon RX 7700 XT~418 tok/sAMD Radeon RX 9070 GRE~418 tok/sIntel Arc A750~413 tok/sNVIDIA GeForce RTX 3060 12GB~377 tok/sAMD Radeon RX 6700 XT~372 tok/sIntel Arc B580~368 tok/sAMD Radeon RX 9060 XT 16GB~310 tok/sIntel Arc B570~307 tok/sNVIDIA GeForce RTX 4060 Ti 16GB~302 tok/sNVIDIA GeForce RTX 4060 Ti 8GB~302 tok/sNVIDIA GeForce RTX 4060~285 tok/sAMD Radeon RX 7600~279 tok/sAMD Radeon RX 7600 XT~279 tok/sAMD Radeon RX 9050~279 tok/sNVIDIA GeForce RTX 3060 8GB~252 tok/sNVIDIA GeForce RTX 3050 8GB~235 tok/s

Which Devices Can Run InternVL3 1B?

Q4_K_M · 0.6 GB

59 devices with unified memory can run InternVL3 1B, including NVIDIA DGX H100, NVIDIA DGX A100 640GB.

Runs great

— Plenty of headroom
NVIDIA DGX H100~28097 tok/sNVIDIA DGX A100 640GB~17101 tok/sMac Studio (M3 Ultra, 256GB)~925 tok/sMac Studio (M3 Ultra, 512GB)~925 tok/sMac Studio (M3 Ultra, 96GB)~925 tok/sMac Pro M2 Ultra (192 GB)~903 tok/sMac Studio M2 Ultra (192 GB)~903 tok/sMacBook Pro 16" M5 Max (128 GB)~693 tok/sMac Studio M4 Max (128 GB)~617 tok/sMac Studio M4 Max (64 GB)~617 tok/sMacBook Pro 16" M4 Max (48 GB)~617 tok/sMacBook Pro 16" M4 Max (64 GB)~617 tok/sMac Studio M4 Max (36 GB)~463 tok/sMacBook Pro 14" M4 Max (36 GB)~463 tok/sMacBook Pro 16" M3 Max (48 GB)~463 tok/sMacBook Pro 14-inch (M5 Pro)~347 tok/sMac Mini M4 Pro (24 GB)~308 tok/sMac Mini M4 Pro (48 GB)~308 tok/sMacBook Pro 14" M4 Pro (24 GB)~308 tok/sMacBook Pro 16" M4 Pro (24 GB)~308 tok/sASUS Ascent GX10~286 tok/sNVIDIA DGX Spark~286 tok/sNVIDIA Jetson AGX Thor Developer Kit~286 tok/sAsus ROG Flow Z13 (2025, Ryzen AI Max+ 395, 128 GB)~268 tok/sBeelink GTR9 Pro (Ryzen AI Max+ 395, 128 GB)~268 tok/sFramework Desktop (Ryzen AI Max+ 395, 128 GB)~268 tok/sGMKtec EVO-X2 (Ryzen AI Max+ 395, 128 GB)~268 tok/sHP Z2 Mini G1a (Ryzen AI Max+ PRO 395, 128 GB)~268 tok/sHP ZBook Ultra G1a 14 (Ryzen AI Max+ PRO 395, 128 GB)~268 tok/sMinisforum MS-S1 MAX (Ryzen AI Max+ 395, 128 GB)~268 tok/sSnapdragon X2 Elite Extreme Copilot+ PC~239 tok/sNVIDIA Jetson AGX Orin 32GB~215 tok/sNVIDIA Jetson AGX Orin 64GB~215 tok/sMacBook Pro 14-inch (M5)~173 tok/siPad Pro M5 13" (16 GB)~173 tok/sSnapdragon X Elite Copilot+ PC~142 tok/sMac Mini M4 (16 GB)~136 tok/sMac Mini M4 (32 GB)~136 tok/sMacBook Air 13" M4 (16 GB)~136 tok/sMacBook Air 13" M4 (24 GB)~136 tok/sMacBook Air 15" M4 (16 GB)~136 tok/sMacBook Air 15" M4 (24 GB)~136 tok/sMacBook Pro 14" M4 (16 GB)~136 tok/siPad Pro M4 13" (16 GB)~136 tok/sAMD Ryzen AI 9 HX 370 (Strix Point) Laptop~116 tok/sMacBook Air 13" M3 (16 GB)~116 tok/sMacBook Air 13" M3 (24 GB)~116 tok/sMacBook Air 13" M3 (8 GB)~116 tok/sIntel Core Ultra 9 288V (Lunar Lake) Laptop~110 tok/sNVIDIA Jetson Orin NX 16GB~107 tok/sNVIDIA Jetson Orin Nano 8GB (Super)~107 tok/sApple iPhone 17 Pro~87 tok/siPhone 17 Pro Max~87 tok/siPhone 17~77 tok/siPhone Air~77 tok/siPhone 15 ProiPhone 15 Pro MaxiPhone 16 ProiPhone 16 Pro Max

Where to Download InternVL3 1B

Community quantizations of this model — GGUF for llama.cpp, Ollama, and LM Studio, plus AWQ/MLX variants where available.

Frequently Asked Questions

How much VRAM does InternVL3 1B need?

InternVL3 1B requires 0.6 GB of VRAM at Q4_K_M, or 2.1 GB at BF16.

VRAM = Weights + KV Cache + Overhead

Weights = 938M × 4.8 bits ÷ 8 = 0.6 GB

VRAM usage by quantization

0.6 GB

Learn more about VRAM estimation →

What's the best quantization for InternVL3 1B?

For InternVL3 1B, Q4_K_M (0.6 GB) offers the best balance of quality and VRAM usage. Q5_K_S (0.7 GB) provides better quality if you have the VRAM. The smallest option is IQ2_XXS at 0.3 GB.

VRAM requirement by quantization

IQ2_XXS
0.3 GB
IQ3_XS
0.4 GB
Q4_0
0.5 GB
IQ4_NL
0.6 GB
Q4_K_M ★
0.6 GB
BF16
2.1 GB

★ Recommended — best balance of quality and VRAM usage.

Learn more about quantization →

Can I run InternVL3 1B on a Mac?

InternVL3 1B requires at least 0.3 GB at IQ2_XXS, which exceeds the unified memory of most consumer Macs. You would need a Mac Studio or Mac Pro with a high-memory configuration.

Can I run InternVL3 1B locally?

Yes — InternVL3 1B can run locally on consumer hardware. At Q4_K_M quantization it needs 0.6 GB of VRAM. Popular tools include Ollama, LM Studio, and llama.cpp.

How fast is InternVL3 1B?

At Q4_K_M, InternVL3 1B can reach ~7742 tok/s on AMD Instinct MI350X. On NVIDIA GeForce RTX 4090: ~1057 tok/s. Speed depends mainly on GPU memory bandwidth. Real-world results typically within ±20%.

tok/s = (bandwidth GB/s ÷ model GB) × efficiency

Example: NVIDIA B200 → 8000 ÷ 0.6 × 0.65 = ~8387 tok/s

Estimated speed at Q4_K_M (0.6 GB)

~8387 tok/s
~1057 tok/s
~8387 tok/s
~7742 tok/s

Real-world results typically within ±20%. Speed depends on batch size, quantization kernel, and software stack.

Learn more about tok/s estimation →

What's the download size of InternVL3 1B?

At Q4_K_M, the download is about 0.56 GB. The full-precision BF16 version is 1.88 GB. The smallest option (IQ2_XXS) is 0.26 GB.

Which GPUs can run InternVL3 1B?

52 consumer GPUs can run InternVL3 1B at Q4_K_M (0.6 GB). Top options include AMD Radeon RX 6700 XT, AMD Radeon RX 6800, AMD Radeon RX 6800 XT. 52 GPUs have plenty of headroom for comfortable inference.

Which devices can run InternVL3 1B?

59 devices with unified memory can run InternVL3 1B at Q4_K_M (0.6 GB), including AMD Ryzen AI 9 HX 370 (Strix Point) Laptop, ASUS Ascent GX10, Apple iPhone 17 Pro, Asus ROG Flow Z13 (2025, Ryzen AI Max+ 395, 128 GB). Apple Silicon Macs use unified memory shared between CPU and GPU, making them well-suited for local LLM inference.