Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: z.ai/blog/glm-5.3-flash
Available now across all official platforms:
Weights: huggingface.co/zai-org/GLM-5.3-Flash
API: docs.z.ai/guides/llm/glm-5.3-flash
Coding Plan: z.ai/subscribe
ZCode: zcode.z.ai/en
Chat: chat.z.ai
AutoClaw: autoclaw.z.ai
Z.ai
@Zai_org
Standard API Pricing for GLM-5.3-Flash (per 1M tokens)
- Input: $0.15
- Output: $0.50
- Cached input: $0.03
- Input: $0.15
- Output: $0.50
- Cached input: $0.03
2:12 PM UTC · Aug 26, 2026 · 281.7K Views
60651.9K136
Z.ai
@Zai_org
On the Z.ai Code Bench, which measures real-world coding performance, GLM-5.3-Flash clearly outperforms GLM-5.2 at every effort level and performs on par with Claude Opus 4.8.

2:12 PM UTC · Aug 26, 2026 · 131.5K Views
122081154
Z.ai
@Zai_org
Architectural enhancements, combined with an optimized pre-training corpus, enable GLM-5.3-Flash to deliver greater intelligence with less compute.

2:12 PM UTC · Aug 26, 2026 · 96.6K Views
425795102
