Skip to content

DeepSeek V4.1 Flash

DeepSeek · Open weights · 1M context · released Sep 10, 2026 · checked 6h ago

Ways to buy 26

Way to buyIn / out per 1M
Morph
Default region, cheapest · cached $0.0015
$0.075 / $0.3-50% vs list
DekaLLM
Default region · cached $0.01
$0.04 / $0.49-42% vs list
OpenInference
fp4 · cached $0.01
$0.1 / $0.5-24% vs list
Relace
Default region · cached $0.01
$0.1 / $0.5-24% vs list
DeepInfra
fp8 · cached $0.0042
$0.14 / $0.42-20% vs list
Wafer
Default region · cached $0.06
$0.099 / $0.6-15% vs list
DeepSeek API
The lab's own API · cached $0.003
$0.15 / $0.6List price
Sail Research
fp4 · cached $0.01
$0.13 / $0.75+9% vs list
StreamLake
fp8 · cached $0.0033
$0.165 / $0.66+10% vs list
CoreWeave
fp8 · cached $0.03
$0.2 / $0.65+19% vs list
Fireworks AI
Default region · cached $0.007
$0.22 / $0.66+26% vs list
GMICloud
fp8 · cached $0.0045
$0.225 / $0.9+50% vs list
Krea
fp8 · cached $0.006
$0.225 / $0.9+50% vs list
Phala
Default region · cached $0.0055
$0.276 / $1.1+84% vs list
Novita AI
fp8 · cached $0.0057
$0.285 / $1.14+90% vs list
AtlasCloud
fp8 · cached $0.03
$0.3 / $1.2+100% vs list
Baseten
fp8 · cached $0.03
$0.3 / $1.2+100% vs list
DigitalOcean
Default region · cached $0.006
$0.3 / $1.2+100% vs list
Makora
fp8 · cached $0.006
$0.3 / $1.2+100% vs list
Modal
Default region · cached $0.03
$0.3 / $1.2+100% vs list
NextBit
fp8 · cached $0.006
$0.3 / $1.2+100% vs list
Parasail
fp8 · cached $0.006
$0.3 / $1.2+100% vs list
Qwen (Alibaba)
Default region · cached $0.03
$0.3 / $1.2+100% vs list
SiliconFlow
fp8 · cached $0.006
$0.3 / $1.2+100% vs list
Together AI
Default region · cached $0.006
$0.3 / $1.2+100% vs list
Venice
fp8 · cached $0.0075
$0.375 / $1.5+150% vs list

USD per 1M tokens; batch $0.112 in, $0.336 out on the lab's API. Regions are separate routes.

Through a gateway

Your own numbers

100M input and 20M output tokens a month: $27.00 direct.

No published token fee: Martian Gateway, Portkey, TrueFoundry AI Gateway.

Related models

In / out per 1M
DeepSeek V4 Flash Vision Exp
DeepSeek · 1M
$0.216 / $0.647
DeepSeek V4 Pro 0813
DeepSeek · 1MCheapest: $0.525 on Baidu (fp8)
$0.66 / $1.98
DeepSeek V4 Flash 0731
DeepSeek · 1.31M
$0.053 / $0.158
DeepSeek V4 Flash 0423
DeepSeek · 1M
$0.049 / $0.098
Gemini 3.1 Flash Lite
Google · 1MCheapest: $0.281 on Google Vertex AI (global-flex)
$0.25 / $1.5
Gemini 3.1 Flash Lite Preview
Google · 1MCheapest: $0.281 on Google Gemini API (flex)
$0.25 / $1.5
GPT-6 Luna
OpenAI · 1.05MCheapest: $0.1 on OpenAI (flex)
$0.1 / $0.5
GPT-6 Luna Pro
OpenAI · 1.05MCheapest: $0.1 on OpenAI (flex)
$0.1 / $0.5

About DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture.

Takes text, image in and returns text. First listed Sep 23, 2026. Up to 393K output tokens. Blended price $0.262 per 1M tokens at 3 input to 1 output. Overall score 3 on leadersboard.fru.dev (rank 33). Model id deepseek-v4-1-flash.

Price history

$0.3Sep 23, 2026Output $0.6Input $0.15

List price since Sep 23, 2026, from daily reads of the OpenRouter catalog.

SeenModelChange
Sep 25, 2026DeepSeek V4.1 FlashInput $0.15 to $0.3, Output $0.6 to $1.2, Cached input $0.003 to $0.006+100%
Sep 25, 2026DeepSeek V4.1 FlashCached input $0.006 to $0.06 on wafer+900%
Sep 25, 2026DeepSeek V4.1 FlashInput $0.199 to $0.165, Output $0.794 to $0.66, Cached input $0.004 to $0.0033 on streamlake/fp8-17%
Sep 25, 2026DeepSeek V4.1 FlashInput $0.05 to $0.1, Output $0.25 to $0.5, Cached input $0.005 to $0.01 on relace+100%
Sep 25, 2026DeepSeek V4.1 FlashInput $0.072 to $0.075, Output $0.288 to $0.3, Cached input $0.0014 to $0.0015 on morph+4%
Sep 25, 2026DeepSeek V4.1 FlashInput $0.1 to $0.04, Output $1 to $0.49 on dekallm-60%
Sep 24, 2026DeepSeek V4.1 FlashInput $0.15 to $0.3, Output $0.6 to $1.2, Cached input $0.003 to $0.006+100%
Sep 24, 2026DeepSeek V4.1 FlashInput $0.15 to $0.3, Output $0.6 to $1.2, Cached input $0.015 to $0.03 on qwen+100%
Sep 24, 2026DeepSeek V4.1 FlashInput $0.285 to $0.225, Output $1.14 to $0.9, Cached input $0.0057 to $0.0045 on gmicloud/fp8-21%

Across fru.devCompany profile

Questions

How much does DeepSeek V4.1 Flash cost?

$0.15 per million input tokens and $0.6 per million output tokens, with cached input at $0.003, or $0.112 and $0.336 through the Batch API.

What is the context window of DeepSeek V4.1 Flash?

1M tokens, with up to 393K tokens of output.

Where is DeepSeek V4.1 Flash cheapest?

Of 26 routes tracked, Morph is cheapest at $0.075 input and $0.3 output per million tokens.

Has DeepSeek V4.1 Flash's price changed?

Yes. The output price went from $0.6 to $1.2, first seen Sep 25, 2026.

Sources: OpenRouter's public model catalog and each provider's own pricing page, read every morning at 05:00 UTC; gateway fees re-read every Monday. List prices: your contract may differ. Logos via logo.dev; trademarks belong to their owners.

Price changes by email

Thursdays, only in weeks when an LLM API list price changed.

Double opt-in. Unsubscribe any time.