Skip to content

Qwen3.8 2.4T A95B

Qwen (Alibaba) · Open weights · 1M context · released Aug 12, 2026 · checked 10h ago

Ways to buy 7

Way to buyIn / out per 1M
Qwen (Alibaba) API
The lab's own API · cached $0.25
$2 / $6List price
DeepInfra
fp4 · cached $0.2
$2 / $6Same as list
Modal
Default region · cached $0.25
$2 / $6Same as list
Novita AI
Default region · cached $0.25
$2 / $6Same as list
SiliconFlow
fp8 · cached $0.25
$2 / $6Same as list
Together AI
Default region · cached $0.25
$2 / $6Same as list
Venice
Default region · cached $0.25
$2 / $6Same as list

USD per 1M tokens. Regions are separate routes.

Through a gateway

Your own numbers

100M input and 20M output tokens a month: $320.00 direct.

No published token fee: Martian Gateway, Portkey, TrueFoundry AI Gateway.

Related models

In / out per 1M
Qwen3.8 Max Prime
Qwen (Alibaba) · 1M
$4 / $12
Qwen3.8 Omni Flash
Qwen (Alibaba) · 1M
$0.15 / $0.47
Qwen3.8 Max (0902)
Qwen (Alibaba) · 1M
$2 / $6
Qwen3.8 Flash
Qwen (Alibaba) · 1M
$0.15 / $0.47
GPT-5.6 Sol
OpenAI · 1.05MCheapest: $2 on OpenAI (flex)
$2 / $10
Muse Spark 1.3
Meta · 1M
$1.25 / $4.25
Gemini 3.8 Flash
Google · 1MCheapest: $0.75 on Google Gemini API (flex)
$0.75 / $3.75
Gemini 3.1 Pro Preview
Google · 1MCheapest: $2.25 on Google Vertex AI (global-flex)
$2 / $12

About Qwen3.8 2.4T A95B

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total.

Takes text in and returns text. First listed Sep 23, 2026. Up to 131K output tokens. Blended price $3 per 1M tokens at 3 input to 1 output. Overall score 5 on leadersboard.fru.dev (rank 27). Model id qwen3-8-2-4t-a95b.

Price history

$3Sep 23, 2026Output $6Input $2

List price since Sep 23, 2026, from daily reads of the OpenRouter catalog.

Across fru.devCompany profile

Questions

How much does Qwen3.8 2.4T A95B cost?

$2 per million input tokens and $6 per million output tokens, with cached input at $0.25.

What is the context window of Qwen3.8 2.4T A95B?

1M tokens, with up to 131K tokens of output.

Where is Qwen3.8 2.4T A95B cheapest?

Of 7 routes tracked, DeepInfra is cheapest at $2 input and $6 output per million tokens.

Sources: OpenRouter's public model catalog and each provider's own pricing page, read every morning at 05:00 UTC; gateway fees re-read every Monday. List prices: your contract may differ. Logos via logo.dev; trademarks belong to their owners.

Price changes by email

Thursdays, only in weeks when an LLM API list price changed.

Double opt-in. Unsubscribe any time.