Mercury 2.5
Inception🇺🇸
inception/mercury-2.5Mercury 2.5 is Inception Labs' diffusion-based LLM (dLLM), generating tokens in parallel rather than sequentially to reach about 1,107 tokens/second with a 260K-token context window, and is billed as the largest diffusion language model trained to date. It delivers a roughly 40% intelligence gain over Mercury 2 (comparable to cost-optimized frontier models like Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna Low) with tunable reasoning, parallel tool calls, and schema-aligned JSON, suiting latency-sensitive search agents, voice pipelines, and coding subagents.
비용 비율
0.1x
컨텍스트
260K
출시일
2026년 9월 8일
입력
Text
출력
Text
지원
추론도구 호출구조화된 출력