Mercury 2
Inception馃嚭馃嚫
inception/mercury-2Mercury 2 is Inception Labs' diffusion-based LLM that generates responses in parallel rather than token-by-token, delivering very high throughput (1000+ tokens/second) with a 128K context window. It offers tunable reasoning and native tool use, matching Claude Haiku/Gemini Flash-class quality at much faster speeds and lower cost, ideal for latency-sensitive agentic and coding tasks.
Cost rate
0.2x
Context
128K
Released
Mar 4, 2026
Input
Text
Output
Text
Support
ReasoningTool callingStructured outputs
Best at
The categories where this model ranks highest.
Entertainment, Sports, & Media
#198 路 top 50%
Software & IT Services
#206 路 top 52%
Business, Management, & Finance
#211 路 top 54%
Medicine & Healthcare
#219 路 top 59%
Performance
Median latency and throughput measured across recent requests.
Throughput925 tok/s
Latency