Compare models
Put models side by side on price, context, capabilities, and benchmarks.
Claude Haiku 4.5 is Anthropic's fast, affordable small model in the Claude 4 generation, balancing strong instruction-following and coding ability with low latency. It is designed for high-throughput agentic and chat applications where speed and cost efficiency are priorities.
Document analysis#31Software & IT Services#102Mathematical#102Writing, Literature, & Language#110
AuthorAnthropic
CountryπΊπΈ United States
ReleasedOct 15, 2025
Overview
Cost rate
1x
Context
200K
Input
Output
Reasoning
Tool calling
Structured outputs
Pricing/M tokens
Input$1.00
Cached input$0.10
Output$5.00
Context
Context length200,000
Max output tokens64,000
Knowledge cutoff-
Performance
Throughput86 tok/s
Latency0.75s
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception.
AuthorInception
CountryπΊπΈ United States
ReleasedAug 31, 2026
Overview
Cost rate
0.1x
Context
260K
Input
Output
Reasoning
Tool calling
Structured outputs
Pricing/M tokens
Input$0.04
Cached input$0.00
Output$0.15
Context
Context length260,000
Max output tokens65,536
Knowledge cutoff-
Performance
Throughput-
Latency-
Benchmarks
Head-to-head scores across reasoning, coding, and agentic indices, plus per-domain rankings.
Intelligence
62.5
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
62.1
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
59.5