Ways to buy 5
| Way to buy | Input | Output | In / out per 1M | Cached | Vs cheapest |
|---|---|---|---|---|---|
| Darkbloom int4 | $0.065 | $0.18 | $0.065 / $0.18Cheapest | n/a | Cheapest |
| Io Net Default region · cached $0.033 | $0.067 | $0.19 | $0.067 / $0.19+4% vs cheapest | $0.033 | +4% vs cheapest |
bf16 · cached $0.04 | $0.07 | $0.2 | $0.07 / $0.2+9% vs cheapest | $0.04 | +9% vs cheapest |
| Phala Default region · cached $0.04 | $0.07 | $0.2 | $0.07 / $0.2+9% vs cheapest | $0.04 | +9% vs cheapest |
bf16 · cached $0.04 | $0.08 | $0.2 | $0.08 / $0.2+17% vs cheapest | $0.04 | +17% vs cheapest |
USD per 1M tokens. Regions are separate routes.
Through a gateway
Your own numbers100M input and 20M output tokens a month: $10.10 direct.
LiteLLM$10.10+0%
Vercel AI Gateway$10.10+0%
Eden AI$10.66+5.5%
OpenRouter$10.66+5.5%
Cloudflare AI GatewayNot offeredNot offered
HeliconeNot offeredNot offered
Kong AI GatewayNot offeredNot offered
Neon AI GatewayNot offeredNot offered
PortkeyNot offeredNot offered
RequestyNot offeredNot offered
TrueFoundry AI GatewayNot offeredNot offered
A dash: not on that gateway's own model list. No published token fee: Martian Gateway.
Related models
| In / out per 1M | Cheapest route | ||||||
|---|---|---|---|---|---|---|---|
NVIDIA · 262K | $0.5 | $2.2 | $0.5 / $2.2 | $0.925 | 262K | on DeepInfra (fp4) | · |
NVIDIA · 131K | $0.2 | $0.2 | $0.2 / $0.2 | $0.2 | 131K | on DeepInfra (bf16) | · |
NVIDIA · 262K | $0.085 | $0.4 | $0.085 / $0.4 | $0.164 | 262K | on DeepInfra (bf16) | · |
NVIDIA · 262K | $0.05 | $0.2 | $0.05 / $0.2 | $0.088 | 262K | on Crusoe (fp8) | · |
OpenAI · 1.05MCheapest: $0.1 on OpenAI (flex) | $0.1 | $0.5 | $0.1 / $0.5 | $0.2 | 1.05M | $0.1 on OpenAI (flex) | · |
OpenAI · 1.05MCheapest: $0.1 on OpenAI (flex) | $0.1 | $0.5 | $0.1 / $0.5 | $0.2 | 1.05M | $0.1 on OpenAI (flex) | · |
Meta · 1M | $0.1 | $0.2 | $0.1 / $0.2 | $0.125 | 1M | Lab list price | · |
Z.ai · 1.31M | $0.045 | $0.14 | $0.045 / $0.14 | $0.069 | 1.31M | on InferenceNet (fp4) | · |
About Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total.
Takes text in and returns text. First listed Sep 26, 2026. Up to 131K output tokens. Blended price $0.094 per 1M tokens at 3 input to 1 output. Model id nemotron-3-5-lightning.
Price history
No list-price history: this model is sold by hosts rather than its lab.
Across fru.devCompany profile
Questions
How much does Nemotron 3.5 Lightning cost?
$0.065 per million input tokens and $0.18 per million output tokens.
What is the context window of Nemotron 3.5 Lightning?
1M tokens, with up to 131K tokens of output.
Where is Nemotron 3.5 Lightning cheapest?
Of 5 routes tracked, Darkbloom is cheapest at $0.065 input and $0.18 output per million tokens.