NVIDIA·Nemotron

Nemotron H 47B Reasoning 128K — Hardware Requirements & GPU Compatibility

ChatReasoning

Nemotron H 47B Reasoning 128K is a 46.8B-parameter open language model from NVIDIA in the Nemotron family. At BF16 it needs about 102.94 GB of VRAM — see which GPUs and Macs can run it below.

372 downloads 21 likes

Specifications

Publisher
NVIDIA
Family
Nemotron
Parameters
46.8B
Release Date
2025-05-22
License
Other

Get Started

How Much VRAM Does Nemotron H 47B Reasoning 128K Need?

Select a quantization to see compatible GPUs below.

QuantizationBitsVRAM
BF16est.16.00102.9 GB

est.= calculated VRAM estimate; no published GGUF file found for that quantization yet. Other rows are verified against real community uploads.

Which GPUs Can Run Nemotron H 47B Reasoning 128K?

BF16 · 102.9 GB

Nemotron H 47B Reasoning 128K (BF16) requires 102.9 GB of VRAM to load the model weights. For comfortable inference with headroom for KV cache and system overhead, 134+ GB is recommended. No single GPU has enough memory — multi-GPU or cluster setups are needed.

Which Devices Can Run Nemotron H 47B Reasoning 128K?

BF16 · 102.9 GB

11 devices with unified memory can run Nemotron H 47B Reasoning 128K, including NVIDIA DGX H100, NVIDIA DGX A100 640GB, MacBook Pro 16" M5 Max (128 GB).

Related Models

Frequently Asked Questions

How much VRAM does Nemotron H 47B Reasoning 128K need?

Nemotron H 47B Reasoning 128K requires 102.9 GB of VRAM at BF16.

VRAM = Weights + KV Cache + Overhead

Weights = 46.8B × 16 bits ÷ 8 = 93.6 GB

KV Cache + Overhead 9.3 GB (at 2K context + ~0.3 GB framework)

VRAM usage by quantization

102.9 GB

Learn more about VRAM estimation →

Can NVIDIA GeForce RTX 5090 run Nemotron H 47B Reasoning 128K?

No — Nemotron H 47B Reasoning 128K requires at least 102.9 GB at BF16, which exceeds the NVIDIA GeForce RTX 5090's 32 GB of VRAM.

Can I run Nemotron H 47B Reasoning 128K on a Mac?

Nemotron H 47B Reasoning 128K requires at least 102.9 GB at BF16, which exceeds the unified memory of most consumer Macs. You would need a Mac Studio or Mac Pro with a high-memory configuration.

Can I run Nemotron H 47B Reasoning 128K locally?

Yes — Nemotron H 47B Reasoning 128K can run locally on consumer hardware. At BF16 quantization it needs 102.9 GB of VRAM. Popular tools include Ollama, LM Studio, and llama.cpp.

How fast is Nemotron H 47B Reasoning 128K?

At BF16, Nemotron H 47B Reasoning 128K can reach ~43 tok/s on AMD Instinct MI350X. Speed depends mainly on GPU memory bandwidth. Real-world results typically within ±20%.

tok/s = (bandwidth GB/s ÷ model GB) × efficiency

Example: NVIDIA B2008000 ÷ 102.9 × 0.65 = ~51 tok/s

Estimated speed at BF16 (102.9 GB)

~51 tok/s
~51 tok/s
~43 tok/s

Real-world results typically within ±20%. Speed depends on batch size, quantization kernel, and software stack.

Learn more about tok/s estimation →

What's the download size of Nemotron H 47B Reasoning 128K?

At BF16, the download is about 93.58 GB.

Which GPUs can run Nemotron H 47B Reasoning 128K?

No single consumer GPU has enough VRAM to run Nemotron H 47B Reasoning 128K at BF16 (102.9 GB). Multi-GPU or professional hardware is required.

Which devices can run Nemotron H 47B Reasoning 128K?

18 devices with unified memory can run Nemotron H 47B Reasoning 128K at BF16 (102.9 GB), including ASUS Ascent GX10, Asus ROG Flow Z13 (2025, Ryzen AI Max+ 395, 128 GB), Beelink GTR9 Pro (Ryzen AI Max+ 395, 128 GB), Framework Desktop (Ryzen AI Max+ 395, 128 GB). Apple Silicon Macs use unified memory shared between CPU and GPU, making them well-suited for local LLM inference.