GLM 5.3 Flash
Z.AI馃嚚馃嚦
z-ai/glm-5.3-flashGLM-5.3-Flash is Z.ai's first natively multimodal model in the GLM-5 family (320B total parameters, 18B active mixture-of-experts, MIT license, 1M-token context), released August 26, 2026 with a hybrid sparse-plus-linear attention architecture that keeps long-context behavior accurate while cutting compute overhead. It is built for efficient coding and long-horizon agent tasks, outperforming GLM-5.2 at roughly one-tenth the price while approaching flagship-class coding and agentic benchmark performance.
Cost rate
0.1x
Context
1M
Released
Aug 26, 2026
Input
TextImageVideo
Output
Text
Support
ReasoningTool callingStructured outputs
Best at
The categories where this model ranks highest.
Software & IT Services
#24 路 top 6%
Mathematical
#29 路 top 7%
Medicine & Healthcare
#33 路 top 9%
Image understanding
#35 路 top 21%
Performance
Median latency and throughput measured across recent requests.
Throughput60 tok/s
Latency