Сравнить модели
Сравнивайте модели рядом по цене, контексту, возможностям и бенчмаркам.
GLM-5.3-FlashX is Z.ai's speed-optimized variant of GLM-5.3-Flash, a natively multimodal MoE model (around 320B total / 18B active parameters, MIT-licensed, 1M-token context) built on the same hybrid sparse-and-linear attention architecture. It generates at up to 200 tokens/sec for latency-sensitive coding and long-horizon agent tasks that still need the full 1M-token context, at a premium over the standard GLM-5.3-Flash endpoint.
Claude Haiku 4.5 is Anthropic's fast, affordable small model in the Claude 4 generation, balancing strong instruction-following and coding ability with low latency. It is designed for high-throughput agentic and chat applications where speed and cost efficiency are priorities.
Бенчмарки
Сравнительные показатели по индексам рассуждений, программирования и агентности, а также рейтинги по отдельным доменам.