Z.ai: GLM 5.3 FlashX
by Kilo Gateway · family glm · listed Sep 2026RTV
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
Current rates USD per 1M tokens
Input$0.37
Output$1.25
Cache read$0.075
Cache write—
Against the market
On the record
- Model id
z-ai/glm-5.3-flashx- Context window
- 1.0M tokens
- Max output
- 131K tokens
- Input modalities
- text, image, video
- Output modalities
- text
- Reasoning
- yes
- Tool calls
- yes
- Open weights
- no
- Knowledge cutoff
- —
- Released
- 2026-09-18
- Last updated
- 2026-09-18
- Provider docs
- kilo.ai
Back of the envelope
≈
Also listed at
- Vercel AI Gateway$0.37 in · $1.25 out
- Zhipu AI$0.37 in · $1.25 out
- OpenRouter$0.37 in · $1.25 out
- Z.AI$0.37 in · $1.25 out
- ZenMux$0.375 in · $1.25 out