Mercury 2
Inception🇺🇸
inception/mercury-2Mercury 2 is Inception Labs' diffusion-based LLM that generates responses in parallel rather than token-by-token, delivering very high throughput (1000+ tokens/second) with a 128K context window. It offers tunable reasoning and native tool use, matching Claude Haiku/Gemini Flash-class quality at much faster speeds and lower cost, ideal for latency-sensitive agentic and coding tasks.
Fascia di costo
0.2x
Contesto
128K
Rilasciato
4 mar 2026
Input
Text
Output
Text
Supporto
RagionamentoChiamata di strumentiOutput strutturati