GLM 5.3 FlashX
Z.AI🇨🇳
z-ai/glm-5.3-flashxGLM-5.3-FlashX is Z.ai's speed-optimized variant of GLM-5.3-Flash, a natively multimodal MoE model (around 320B total / 18B active parameters, MIT-licensed, 1M-token context) built on the same hybrid sparse-and-linear attention architecture. It generates at up to 200 tokens/sec for latency-sensitive coding and long-horizon agent tasks that still need the full 1M-token context, at a premium over the standard GLM-5.3-Flash endpoint.
Mức chi phí
0.3x
Ngữ cảnh
1M
Phát hành
18 thg 9, 2026
Đầu vào
TextImageVideo
Đầu ra
Text
Hỗ trợ
Suy luậnGọi công cụĐầu ra có cấu trúc