Ling 3.0 Flash is InclusionAI's (Ant Group) successor to Ling 2.6 Flash, a hybrid-reasoning Mixture-of-Experts model (~124B total, ~5.1B active) combining Kimi Delta Attention with Multi-Head Latent Attention for efficient long-range memory. It supports both thinking and non-thinking modes with a native 262K context, suited for fast general-purpose coding and agentic tasks.
開発元InclusionAI
国🇨🇳 China
リリース2026年7月23日
概要
コスト率
0.1x
コンテキスト
262K
入力
出力
推論
ツール呼び出し
構造化出力
価格/M tokens
入力$0.02
キャッシュ入力$0.00
出力$0.06
コンテキスト
コンテキスト長262,144
最大出力トークン32,768
知識のカットオフ-
パフォーマンス
スループット389 tok/s
レイテンシ2.44s
DeepSeek V4 Flash (0731) is the July 31, 2026 public-beta checkpoint of DeepSeek's V4 Flash model, DeepSeek's fast, cost-efficient V4-generation model tuned for agentic workflows. It offers the same lightweight, high-throughput design as other V4 Flash releases.
生命科学・物理科学・社会科学#79法律・行政#81文章・文学・言語#82医療・ヘルスケア#82
開発元DeepSeek
国🇨🇳 China
リリース2026年7月31日
概要
コスト率
0.1x
コンテキスト
1.3M
入力
出力
推論
ツール呼び出し
構造化出力
価格/M tokens
入力$0.14
キャッシュ入力$0.03
出力$0.28
コンテキスト
コンテキスト長1,310,720
最大出力トークン131,072
知識のカットオフ-
パフォーマンス
スループット122 tok/s
レイテンシ1.1s
ベンチマーク
推論、コーディング、エージェント性能の指数を直接比較したスコアと、分野別ランキング。
知能
62.5
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
62.1
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
59.5