Skip to content

Llama 3.3 70B Instruct

Meta · Open weights · 131K context · released Dec 6, 2024 · checked 7h ago

Ways to buy 14

Way to buyIn / out per 1M
DeepInfra
turbo, fp8
$0.1 / $0.32Cheapest
Novita AI
bf16
$0.135 / $0.4+30% vs cheapest
AkashML
fp8 · cached $0.1
$0.2 / $0.52+81% vs cheapest
Parasail
fp8 · cached $0.11
$0.22 / $0.5+87% vs cheapest
SambaNova
Default region
$0.45 / $0.9+263% vs cheapest
Groq
Default region · cached $0.295
$0.59 / $0.79+313% vs cheapest
CoreWeave
fp16 · cached $0.71
$0.71 / $0.71+358% vs cheapest
Google Vertex AI
Default region
$0.72 / $0.72+365% vs cheapest
Google Vertex AI
us-central1
$0.72 / $0.72+365% vs cheapest
Databricks Mosaic AI Gateway + Foundation Model APIs
From the platform pricing page
$0.5 / $1.5+384% vs cheapest
IBM watsonx.ai
From the platform pricing page
$0.753 / $0.753+386% vs cheapest
Cloudflare Workers AI
fp8
$0.293 / $2.25+405% vs cheapest
Snowflake Cortex AI (AI_COMPLETE)
From the platform pricing page
$0.864 / $0.864+457% vs cheapest
Together AI
Default region
$1.04 / $1.04+571% vs cheapest

USD per 1M tokens. Regions are separate routes.

Through a gateway

Your own numbers

100M input and 20M output tokens a month: $16.40 direct.

No published token fee: Martian Gateway, Portkey, TrueFoundry AI Gateway.

Related models

In / out per 1M
Muse Spark 1.3
Meta · 1M
$1.25 / $4.25
Muse Spark 1.3 Contributor
Meta · 1M
$0.1 / $0.2
Muse Spark 1.2 Contributor
Meta · 1M
$0.1 / $0.2
Muse Glimmer 30B
Meta · 131K
$0.3 / $1.1
DeepSeek V4.1 Flash
DeepSeek · 1MCheapest: $0.131 on Morph
$0.15 / $0.6
GPT-6 Luna
OpenAI · 1.05MCheapest: $0.1 on OpenAI (flex)
$0.1 / $0.5
GPT-6 Luna Pro
OpenAI · 1.05MCheapest: $0.1 on OpenAI (flex)
$0.1 / $0.5
Qwen3.8 Omni Flash
Qwen (Alibaba) · 1M
$0.15 / $0.47

About Llama 3.3 70B Instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).

Takes text in and returns text. First listed Sep 23, 2026. Up to 16K output tokens. Blended price $0.155 per 1M tokens at 3 input to 1 output. Model id llama-3-3-70b-instruct.

Price history

No list-price history: this model is sold by hosts rather than its lab.

Questions

How much does Llama 3.3 70B Instruct cost?

$0.1 per million input tokens and $0.32 per million output tokens.

What is the context window of Llama 3.3 70B Instruct?

131K tokens, with up to 16K tokens of output.

Where is Llama 3.3 70B Instruct cheapest?

Of 14 routes tracked, DeepInfra is cheapest at $0.1 input and $0.32 output per million tokens.

Sources: OpenRouter's public model catalog and each provider's own pricing page, read every morning at 05:00 UTC; gateway fees re-read every Monday. List prices: your contract may differ. Logos via logo.dev; trademarks belong to their owners.

Price changes by email

Thursdays, only in weeks when an LLM API list price changed.

Double opt-in. Unsubscribe any time.