GLM 4.6V
Z.AI馃嚚馃嚦
z-ai/glm-4.6vGLM-4.6V is Z.ai's vision-language model that treats images, video, and tools as first-class agent inputs, extending training context to 128K tokens with native multimodal function calling. It is designed for agentic multimodal workflows such as screenshot- and document-driven tool use.
Cost rate
0.2x
Context
131K
Released
Dec 8, 2025
Input
ImageTextVideo
Output
Text
Support
ReasoningTool callingStructured outputs
Best at
The categories where this model ranks highest.
Mathematical
#88 路 top 23%
Image understanding
#97 路 top 64%
Medicine & Healthcare
#101 路 top 27%
Life, Physical, & Social Science
#107 路 top 27%
Performance
Median latency and throughput measured across recent requests.
Throughput57 tok/s
Latency