Skip to content

GLM 5.3 Flash

Z.ai · Open weights · 1.31M context · released Aug 26, 2026 · checked 6h ago

Ways to buy 32

Way to buyIn / out per 1M
InferenceNet
fp4 · cached $0.01
$0.045 / $0.14Cheapest
DeepInfra
fp4 · cached $0.015
$0.075 / $0.25+73% vs cheapest
Relace
Default region · cached $0.02
$0.07 / $0.28+78% vs cheapest
GMICloud
fp8 · cached $0.018
$0.09 / $0.3+107% vs cheapest
OpenInference
fp4 · cached $0.02
$0.07 / $0.36+107% vs cheapest
Wafer
Default region · cached $0.06
$0.089 / $0.35+124% vs cheapest
Morph
Default region · cached $0.02
$0.098 / $0.343+132% vs cheapest
Sail Research
fp8 · cached $0.029
$0.045 / $0.6+167% vs cheapest
Sail Research
us, fp8 · cached $0.029
$0.045 / $0.6+167% vs cheapest
Decart
fp4 · cached $0.025
$0.128 / $0.425+194% vs cheapest
Phala
fp8 · cached $0.025
$0.128 / $0.425+194% vs cheapest
Novita AI
fp8 · cached $0.026
$0.132 / $0.44+204% vs cheapest
StreamLake
fp8 · cached $0.028
$0.141 / $0.47+225% vs cheapest
Io Net
fp8 · cached $0.029
$0.143 / $0.475+228% vs cheapest
AtlasCloud
fp8 · cached $0.03
$0.15 / $0.5+245% vs cheapest
Baseten
fp8 · cached $0.03
$0.15 / $0.5+245% vs cheapest
CoreWeave
nvfp4 · cached $0.05
$0.15 / $0.5+245% vs cheapest
Crusoe
fp4 · cached $0.03
$0.15 / $0.5+245% vs cheapest
DigitalOcean
Default region · cached $0.03
$0.15 / $0.5+245% vs cheapest
Fireworks AI
Default region · cached $0.03
$0.15 / $0.5+245% vs cheapest
Friendli
Default region · cached $0.03
$0.15 / $0.5+245% vs cheapest
Inceptron
fp8 · cached $0.07
$0.15 / $0.5+245% vs cheapest
Modal
nvfp4 · cached $0.03
$0.15 / $0.5+245% vs cheapest
Near AI
fp8 · cached $0.035
$0.15 / $0.5+245% vs cheapest
Parasail
fp8 · cached $0.03
$0.15 / $0.5+245% vs cheapest
Reka
Default region · cached $0.03
$0.15 / $0.5+245% vs cheapest
SiliconFlow
fp8 · cached $0.03
$0.15 / $0.5+245% vs cheapest
Together AI
Default region · cached $0.03
$0.15 / $0.5+245% vs cheapest
Venice
Default region · cached $0.03
$0.15 / $0.5+245% vs cheapest
Z.ai
fp8 · cached $0.03
$0.15 / $0.5+245% vs cheapest
NextBit
fp8 · cached $0.033
$0.165 / $0.55+280% vs cheapest
Cloudflare Workers AI
Default region · cached $0.03
$0.3 / $1+591% vs cheapest

USD per 1M tokens. Regions are separate routes.

Through a gateway

Your own numbers

100M input and 20M output tokens a month: $7.30 direct.

No published token fee: Martian Gateway, Portkey, TrueFoundry AI Gateway.

Related models

In / out per 1M
GLM 5.3 Prime
Z.ai · 1M
$2.8 / $8.8
GLM 5.3 FlashX
Z.ai · 1M
$0.37 / $1.25
GLM 5.3
Z.ai · 1.31M
$0.561 / $1.76
GLM 5.2
Z.ai · 1M
$0.561 / $1.76
Muse Spark 1.3 Contributor
Meta · 1M
$0.1 / $0.2
Muse Spark 1.2 Contributor
Meta · 1M
$0.1 / $0.2
DeepSeek V4 Flash 0731
DeepSeek · 1.31M
$0.053 / $0.158
Qwen3.7 Flash
Qwen (Alibaba) · 1M
$0.03 / $0.13

About GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai.

Takes text, image, video in and returns text. First listed Sep 23, 2026. Up to 131K output tokens. Blended price $0.069 per 1M tokens at 3 input to 1 output. Model id glm-5-3-flash.

Price history

No list-price history: this model is sold by hosts rather than its lab.

Across fru.devCompany profile

Questions

How much does GLM 5.3 Flash cost?

$0.045 per million input tokens and $0.14 per million output tokens, with cached input at $0.01.

What is the context window of GLM 5.3 Flash?

1.31M tokens, with up to 131K tokens of output.

Where is GLM 5.3 Flash cheapest?

Of 32 routes tracked, InferenceNet is cheapest at $0.045 input and $0.14 output per million tokens.

Has GLM 5.3 Flash's price changed?

Yes. The output price went from $0.28 to $0.5, first seen Sep 25, 2026.

Sources: OpenRouter's public model catalog and each provider's own pricing page, read every morning at 05:00 UTC; gateway fees re-read every Monday. List prices: your contract may differ. Logos via logo.dev; trademarks belong to their owners.

Price changes by email

Thursdays, only in weeks when an LLM API list price changed.

Double opt-in. Unsubscribe any time.