Ling 3.0 Flash is InclusionAI's (Ant Group) successor to Ling 2.6 Flash, a hybrid-reasoning Mixture-of-Experts model (~124B total, ~5.1B active) combining Kimi Delta Attention with Multi-Head Latent Attention for efficient long-range memory. It supports both thinking and non-thinking modes with a native 262K context, suited for fast general-purpose coding and agentic tasks.
개발사InclusionAI
국가🇨🇳 China
출시일2026년 7월 23일
개요
비용 비율
0.1x
컨텍스트
262K
입력
출력
추론
도구 호출
구조화된 출력
가격/M tokens
입력$0.02
캐시된 입력$0.00
출력$0.06
컨텍스트
컨텍스트 길이262,144
최대 출력 토큰32,768
지식 기준일-
성능
처리량389 tok/s
지연 시간2.44s
DeepSeek V4 Flash (0731) is the July 31, 2026 public-beta checkpoint of DeepSeek's V4 Flash model, DeepSeek's fast, cost-efficient V4-generation model tuned for agentic workflows. It offers the same lightweight, high-throughput design as other V4 Flash releases.
생명과학, 물리과학 및 사회과학#79법률 및 행정#81작문, 문학 및 언어#82의료 및 헬스케어#82
개발사DeepSeek
국가🇨🇳 China
출시일2026년 7월 31일
개요
비용 비율
0.1x
컨텍스트
1.3M
입력
출력
추론
도구 호출
구조화된 출력
가격/M tokens
입력$0.14
캐시된 입력$0.03
출력$0.28
컨텍스트
컨텍스트 길이1,310,720
최대 출력 토큰131,072
지식 기준일-
성능
처리량122 tok/s
지연 시간1.1s
벤치마크
추론, 코딩, 에이전트 지수의 맞대결 점수와 도메인별 순위를 확인하세요.
지능
62.5
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
62.1
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
59.5