GLM 4.7 Flash
Z.AI馃嚚馃嚦
z-ai/glm-4.7-flashGLM-4.7-Flash is a 30B-class SOTA model from Z.ai that balances performance and efficiency, offering a faster, cheaper alternative to the full GLM-4.7 model. It targets general reasoning and coding tasks where speed and cost matter more than maximum capability.
Cost rate
0.1x
Context
200K
Released
Jan 19, 2026
Input
Text
Output
Text
Support
ReasoningTool callingStructured outputs
Best at
The categories where this model ranks highest.
Software & IT Services
#181 路 top 44%
Mathematical
#185 路 top 47%
Business, Management, & Finance
#194 路 top 48%
Legal & Government
#199 路 top 52%
Performance
Median latency and throughput measured across recent requests.
Throughput116 tok/s
Latency