Ling 3.0 Flash is InclusionAI's (Ant Group) successor to Ling 2.6 Flash, a hybrid-reasoning Mixture-of-Experts model (~124B total, ~5.1B active) combining Kimi Delta Attention with Multi-Head Latent Attention for efficient long-range memory. It supports both thinking and non-thinking modes with a native 262K context, suited for fast general-purpose coding and agentic tasks.
作者InclusionAI
国家/地区🇨🇳 China
发布日期2026年7月23日
概览
费用比率
0.1x
上下文
262K
输入
输出
推理
工具调用
结构化输出
定价/M tokens
输入$0.02
缓存输入$0.00
输出$0.06
上下文
上下文长度262,144
最大输出令牌32,768
知识截止日期-
性能
吞吐量389 tok/s
延迟2.44s
DeepSeek V4 Flash (0731) is the July 31, 2026 public-beta checkpoint of DeepSeek's V4 Flash model, DeepSeek's fast, cost-efficient V4-generation model tuned for agentic workflows. It offers the same lightweight, high-throughput design as other V4 Flash releases.
生命科学、自然科学与社会科学#79法律与政府#81写作、文学与语言#82医学与医疗保健#82
作者DeepSeek
国家/地区🇨🇳 China
发布日期2026年7月31日
概览
费用比率
0.1x
上下文
1.3M
输入
输出
推理
工具调用
结构化输出
定价/M tokens
输入$0.14
缓存输入$0.03
输出$0.28
上下文
上下文长度1,310,720
最大输出令牌131,072
知识截止日期-
性能
吞吐量122 tok/s
延迟1.1s
基准测试
推理、编程和智能体指数的对比得分,以及各领域排名。
智能
62.5
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
62.1
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
59.5