GLM 5.3 Flash
z-ai/glm-5.3-flashGLM-5.3-Flash is Z.ai's first natively multimodal model in the GLM-5 family (320B total parameters, 18B active mixture-of-experts, MIT license, 1M-token context), released August 26, 2026 with a hybrid sparse-plus-linear attention architecture that keeps long-context behavior accurate while cutting compute overhead. It is built for efficient coding and long-horizon agent tasks, outperforming GLM-5.2 at roughly one-tenth the price while approaching flagship-class coding and agentic benchmark performance.
Best at
The categories where this model ranks highest.
Mathematical
#12 路 top 3%
Software & IT Services
#19 路 top 5%
Image understanding
#30 路 top 20%
Life, Physical, & Social Science
#31 路 top 8%
Performance
Median latency and throughput measured across recent requests.