Mercury 2.5
Inception🇺🇸
inception/mercury-2.5Mercury 2.5 is Inception Labs' diffusion-based LLM (dLLM), generating tokens in parallel rather than sequentially to reach about 1,107 tokens/second with a 260K-token context window, and is billed as the largest diffusion language model trained to date. It delivers a roughly 40% intelligence gain over Mercury 2 (comparable to cost-optimized frontier models like Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna Low) with tunable reasoning, parallel tool calls, and schema-aligned JSON, suiting latency-sensitive search agents, voice pipelines, and coding subagents.
Fascia di costo
0.1x
Contesto
260K
Rilasciato
8 set 2026
Input
Text
Output
Text
Supporto
RagionamentoChiamata di strumentiOutput strutturati