Skip to content

Nemotron 3 Ultra

NVIDIA · Open weights · 262K context · released Jun 4, 2026 · checked 2h ago

Ways to buy 3

Way to buyIn / out per 1M
DeepInfra
fp4 · cached $0.1
$0.5 / $2.2Cheapest
Baseten
fp4 · cached $0.12
$0.6 / $2.4+14% vs cheapest
Venice
fp8 · cached $0.188
$0.625 / $3.13+35% vs cheapest

USD per 1M tokens. Regions are separate routes.

Through a gateway

Your own numbers

100M input and 20M output tokens a month: $94.00 direct.

A dash: not on that gateway's own model list. No published token fee: Martian Gateway.

Related models

In / out per 1M
Nemotron 3.5 Lightning
NVIDIA · 1M
$0.065 / $0.18
Nemotron 3.5 Content Safety
NVIDIA · 131K
$0.2 / $0.2
Nemotron 3 Super
NVIDIA · 262K
$0.085 / $0.4
Nemotron 3 Nano 30B A3B
NVIDIA · 262K
$0.05 / $0.2
Gemini 3.8 Flash
Google · 1MCheapest: $0.75 on Google Gemini API (flex)
$0.75 / $3.75
Muse Spark 1.3
Meta · 1M
$1.25 / $4.25
Muse Spark 1.1
Meta · 1M
$1.25 / $4.25
GLM 5.3
Z.ai · 1.31M
$0.55 / $1.7

About Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE).

Takes text in and returns text. First listed Sep 26, 2026. Up to 183K output tokens. Blended price $0.925 per 1M tokens at 3 input to 1 output. Model id nemotron-3-ultra-550b-a55b.

Price history

No list-price history: this model is sold by hosts rather than its lab.

Across fru.devCompany profile

Questions

How much does Nemotron 3 Ultra cost?

$0.5 per million input tokens and $2.2 per million output tokens, with cached input at $0.1.

What is the context window of Nemotron 3 Ultra?

262K tokens, with up to 183K tokens of output.

Where is Nemotron 3 Ultra cheapest?

Of 3 routes tracked, DeepInfra is cheapest at $0.5 input and $2.2 output per million tokens.

Price changes by email

Thursdays, only in weeks when an LLM API list price changed.

Double opt-in. Unsubscribe any time.