Compare models
Put models side by side on price, context, capabilities, and benchmarks.
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s.
Claude Haiku 4.5 is Anthropic's fast, affordable small model in the Claude 4 generation, balancing strong instruction-following and coding ability with low latency. It is designed for high-throughput agentic and chat applications where speed and cost efficiency are priorities.
Benchmarks
Head-to-head scores across reasoning, coding, and agentic indices, plus per-domain rankings.