GLM 5.3 FlashX
Z.AI🇨🇳
z-ai/glm-5.3-flashxGLM-5.3-FlashX is Z.ai's speed-optimized variant of GLM-5.3-Flash, a natively multimodal MoE model (around 320B total / 18B active parameters, MIT-licensed, 1M-token context) built on the same hybrid sparse-and-linear attention architecture. It generates at up to 200 tokens/sec for latency-sensitive coding and long-horizon agent tasks that still need the full 1M-token context, at a premium over the standard GLM-5.3-Flash endpoint.
Kostenfaktor
0.3x
Kontext
1M
Veröffentlicht
18. Sept. 2026
Eingabe
TextImageVideo
Ausgabe
Text
Unterstützung
SchlussfolgernTool-AufrufeStrukturierte Ausgaben