Mercury 2
Inception🇺🇸
inception/mercury-2Mercury 2 is Inception Labs' diffusion-based LLM that generates responses in parallel rather than token-by-token, delivering very high throughput (1000+ tokens/second) with a 128K context window. It offers tunable reasoning and native tool use, matching Claude Haiku/Gemini Flash-class quality at much faster speeds and lower cost, ideal for latency-sensitive agentic and coding tasks.
Tarifa de coste
0.2x
Contexto
128K
Lanzamiento
4 mar 2026
Entrada
Text
Salida
Text
Compatibilidad
RazonamientoLlamada a herramientasSalidas estructuradas
En qué destaca
Las categorías en las que este modelo obtiene la mejor clasificación.
Software y servicios de TI
#219 · top 53%
Entretenimiento, deportes y medios
#219 · top 53%
Negocios, gestión y finanzas
#227 · top 56%
Medicina y salud
#242 · top 63%
Rendimiento
Latencia y rendimiento medianos medidos en solicitudes recientes.
Rendimiento697 tok/s
Latencia